Audio processing system, audio processing method, and program

The audio processing system simplifies the generation and adjustment of reference data by displaying audio and pitch sequences on a common timeline, enhancing user interaction and accuracy in audio signal analysis.

CN114446266BActive Publication Date: 2025-07-15YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111270719.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-06
Filing Date
2021-10-29
Publication Date
2025-07-15
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

In the prior art, when generating reference data, it is necessary to estimate the playing position of the music, and it is difficult to quickly confirm and correct the analysis results of the reference signal.

Method used

The audio analysis unit analyzes the audio signal, determines the time series of pronunciation index and pitch, and displays it based on the common time axis on the display device of the display control unit, allowing the user to intuitively confirm and correct the analysis results.

Benefits of technology

It realizes convenient confirmation and correction of the process of generating reference data by users, reduces the influence of the sound components of the performance device, and improves the accuracy of the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114446266B_ABST
    Figure CN114446266B_ABST
Patent Text Reader

Abstract

Enables the user to easily confirm and correct the analysis results of the reference signal during the process of generating reference data. The audio processing system (10) includes an audio analysis unit (43) and a display control unit (36). The audio analysis unit (43) performs an analysis process of analyzing a reference signal (Zr) that includes the audio of the musical instrument (80). The analysis process includes the following processes: a process of determining the time series of the pronunciation index, which is an index of the accuracy of the audio component of the musical instrument (80) included in the reference signal (Zr); and a process of determining the time series of the pitch related to the audio component of the musical instrument (80). The display control unit (36) causes the time series of the pronunciation index and the time series of the pitch to be displayed on the display device (15) based on a common time axis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for analyzing an audio signal. Background Art

[0002] Conventionally, a technique has been proposed for causing music playback to follow the performance by a performer. For example, Patent Document 1 discloses a technique in which the performance position in a piece of music is estimated by analyzing an audio signal, and the automatic performance of the piece of music is controlled in accordance with the estimation result, where the audio signal represents musical sounds produced by the performance of the piece of music.

[0003] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2017-207615

[0004] For the estimation of the performance position, it is necessary to prepare in advance reference data to be used when comparing with the audio signal. When generating the reference data, it is necessary to analyze a reference signal representing musical sounds produced by a prior performance and to perform a process of correcting the analysis result in accordance with an instruction from a user. Against this background, there is a demand for a technique that enables a user to easily confirm and correct the analysis result of the reference signal in the process of generating the reference data. Summary of the Invention

[0005] In order to solve the above problems, an audio processing system according to one aspect of the present invention includes an audio analysis unit and a display control unit. The audio analysis unit performs an analysis process for analyzing an audio signal of an audio including a first sound source. The analysis process includes the following processes: a first process of determining a time series of a pronunciation index, which is an index of the accuracy of the audio component of the first sound source included in the audio signal; and a process of determining a time series of a pitch related to the audio component of the first sound source. The display control unit causes the time series of the pronunciation index and the time series of the pitch to be displayed on a display device based on a common time axis.

[0006] An audio processing method according to one aspect of the present invention performs an analysis process for analyzing an audio signal of an audio including a first sound source, and includes the following processes: a first process of determining a time series of a pronunciation index, which is an index of the accuracy of the audio component of the first sound source included in the audio signal; and a process of determining a time series of a pitch related to the audio component of the first sound source. The audio processing method causes the time series of the pronunciation index and the time series of the pitch to be displayed on a display device based on a common time axis.

[0007] One aspect of the present invention relates to a program that causes a computer system to function as an audio analysis unit and a display control unit. The audio analysis unit performs an analysis process of analyzing an audio signal of an audio including a first sound source. The analysis process includes the following processes: a first process of determining a time series of pronunciation indexes, which are indexes of the accuracy of the audio component of the first sound source included in the audio signal; and a process of determining a time series of pitches related to the audio component of the first sound source. The display control unit causes the time series of the pronunciation indexes and the time series of the pitches to be displayed on a display device based on a common time axis. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a block diagram illustrating the configuration of a playback system according to the first embodiment.

[0009] Figure 2 is a schematic diagram of music data.

[0010] Figure 3 is an explanatory diagram of the configuration and state of an operation device.

[0011] Figure 4 is a block diagram illustrating the functional configuration of an audio processing system.

[0012] Figure 5 is an explanatory diagram of the relationship between the playback of a playback part by a performance device and first and second instructions.

[0013] Figure 6 is a flowchart illustrating a specific process flow of playback control processing.

[0014] Figure 7 is an explanatory diagram related to the state of an operation device according to the second embodiment.

[0015] Figure 8 is an explanatory diagram related to the state of an operation device according to the third embodiment.

[0016] Figure 9 is a block diagram illustrating the functional configuration of an audio processing system according to the fourth embodiment.

[0017] Figure 10 is an explanatory diagram of the operation of an editing processing unit.

[0018] Figure 11 is a flowchart illustrating a specific process flow of editing processing.

[0019] Figure 12 is a block diagram illustrating the configuration of a playback system according to the fifth embodiment.

[0020] Figure 13is a block diagram illustrating the functional configuration of the audio processing system according to the fifth embodiment.

[0021] Figure 14 is a block diagram illustrating the specific configuration of the audio analysis unit.

[0022] Figure 15 is a schematic diagram of the confirmation screen.

[0023] Figure 16 is a schematic diagram showing the state of change of the confirmation screen.

[0024] Figure 17 is a schematic diagram showing the state of change of the confirmation screen.

[0025] Figure 18 is a schematic diagram showing the state of change of the confirmation screen.

[0026] Figure 19 is a flowchart illustrating the specific process of the adjustment process.

[0027] Figure 20 is a schematic diagram of the confirmation screen of the modified example. Detailed Embodiment

[0028] A: First Embodiment

[0029] Figure 1 is a block diagram illustrating the configuration of the playback system 100 according to the first embodiment. The playback system 100 is provided in the audio space where the user U is located. The user U is, for example, a performer who plays a specific part (hereinafter referred to as the "performance part") of a piece of music using an instrument 80 such as a stringed instrument.

[0030] The playback system 100 is a computer system that plays the piece of music in parallel with the performance of the performance part by the user U. Specifically, the playback system 100 plays the parts other than the performance part (hereinafter referred to as the "playback parts") among the multiple parts that make up the piece of music. The performance part is, for example, one or more parts that make up the main melody of the piece of music. The playback parts are, for example, one or more parts that make up the accompaniment of the piece of music. As understood from the above description, by performing the performance of the performance part by the user U and the playback of the playback parts by the playback system 100 in parallel, the performance of the piece of music is realized. In addition, the performance part and the playback parts may be common parts of the piece of music. Alternatively, the performance part may make up the accompaniment of the piece of music, and the playback parts may make up the main melody of the piece of music.

[0031] The playback system 100 includes a sound processing system 10 and a musical performance device 20. The sound processing system 10 and the musical performance device 20 are configured separately and communicate with each other by wire or wirelessly. Alternatively, the sound processing system 10 and the musical performance device 20 may be configured integrally.

[0032] The performance device 20 is a playback device that plays the playback part of the music under the control of the sound processing system 10. Specifically, the performance device 20 is an automatic performance instrument that performs automatic performance of the playback part. For example, an automatic performance instrument (for example, an automatic performance piano) of a different type from the instrument 80 played by the user U is used as the performance device 20. As understood from the above description, automatic performance is a form of "playback".

[0033] The performance device 20 of the first embodiment has a driving mechanism 21 and a sound-generating mechanism 22. The sound-generating mechanism 22 is a mechanism for generating musical sounds. Specifically, the sound-generating mechanism 22 has a string-striking mechanism for each key that makes a string (sound source) sound in conjunction with the displacement of each key of the keyboard, similar to a keyboard instrument of a natural musical instrument. The driving mechanism 21 drives the sound-generating mechanism 22 to perform automatic performance of the music. The sound-generating mechanism 22 is driven by the driving mechanism 21 in accordance with the instruction from the sound processing system 10, thereby realizing automatic performance of the playback part.

[0034] The sound processing system 10 is a computer system that controls the playback of the playback part performed by the performance device 20, and has a control device 11, a storage device 12, a sound pickup device 13, and an operation device 14. The sound processing system 10 is implemented by, for example, a mobile terminal device such as a smartphone or a tablet terminal, or a mobile or fixed terminal device such as a personal computer. In addition, the sound processing system 10 can be implemented by a single device or a plurality of devices that are separately configured from each other.

[0035] The control device 11 is a single or multiple processors that control the various elements of the sound processing system 10. Specifically, the control device 11 is composed of one or more processors such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0036] The storage device 12 is one or more memories that store the programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is constituted by a known recording medium such as a magnetic recording medium or a semiconductor recording medium, for example, or by a combination of multiple recording media. Additionally, a removable recording medium that can be detached from and attached to the audio processing system 10, or a recording medium capable of writing to or reading from via a communication network (for example, cloud storage) can be used as the storage device 12.

[0037] The storage device 12 stores music data D that specifies the time series of a plurality of notes constituting a piece of music for each piece of music. Figure 2 It is a schematic diagram of the music data D. The music data D includes reference data Da and performance data Db. The reference data Da specifies the time series of the notes of the performance part played by the user U. Specifically, the reference data Da specifies the pitch, the sounding period, and the sounding intensity (velocity) for each of the plurality of notes of the performance part. On the other hand, the performance data Db specifies the time series of the notes of the playback part played by the playback device 20. Specifically, the performance data Db specifies the pitch, the sounding period, and the sounding intensity for each of the plurality of notes of the playback part.

[0038] Each of the reference data Da and the performance data Db is, for example, time series data in the MIDI (Musical Instrument Digital Interface) format in which instruction data indicating the sounding or silencing of a musical sound and time data specifying the time point of the action indicated by the instruction data are arranged in time series. The instruction data indicates an action such as sounding or silencing by specifying the pitch and intensity, for example. The time data specifies the interval between successive instruction data, for example. The period from when the sounding of a specific pitch is indicated by the instruction data until the silencing of that pitch is indicated by the subsequent instruction data of that instruction data is the sounding period related to the note of that pitch.

[0039] Figure 1The sound pickup device 13 picks up the musical sound emitted from the musical instrument 80 by the performance of the user U, and generates an audio signal Z representing the waveform of the musical sound. For example, a microphone is used as the sound pickup device 13. For convenience, the A / D converter that converts the audio signal Z generated by the sound pickup device 13 from analog to digital is not shown in the figure. In addition, the user U can also sing the performance part of the music piece with respect to the sound pickup device 13. When the user U sings the performance part, the sound pickup device 13 generates an audio signal Z representing the waveform of the singing voice of the user U. As understood from the above description, "performance" includes not only the narrow sense of performance using the musical instrument 80, but also the singing of the music piece by the user U.

[0040] In addition, in the first embodiment, a structure in which the sound pickup device 13 is mounted on the audio processing system 10 is illustrated, but the sound pickup device 13 separate from the audio processing system 10 can also be connected to the audio processing system 10 in a wired or wireless manner. Alternatively, the audio processing system 10 can receive the output signal from an electric musical instrument such as an electric stringed instrument as the audio signal Z. As understood from the above description, the sound pickup device 13 can also be omitted from the audio processing system 10.

[0041] The operation device 14 is an input device that receives an instruction from the user U. As Figure 3 illustrated, the operation device 14 in the first embodiment has a movable part 141 that moves by the operation of the user U. The movable part 141 is an operation pedal that can be operated by the user U's foot. For example, a pedal-type MIDI controller is used as the operation device 14. The user U can operate the operation device 14 at a desired time point in parallel with the performance while playing the musical instrument 80 with both hands. In addition, a touch panel that detects the contact of the user U can be used as the operation device 14.

[0042] The operation device 14 switches between the released state and the operated state in response to the operation of the user U. The released state is a state in which the operation device 14 is not operated by the user U. Specifically, the released state is a state in which the user U does not step on the movable part 141. The released state is also represented as a state in which the movable part 141 is at the position H1. On the other hand, the operated state is a state in which the operation device 14 is operated by the user U. Specifically, the operated state is a state in which the user U steps on the movable part 141. The operated state is also represented as a state in which the movable part 141 exists at a position H2 different from the position H1. The released state is an example of the "first state", and the operated state is an example of the "second state".

[0043] Figure 4FIG. 0 is a block diagram illustrating the functional structure of the audio processing system 10. The control device 11 implements a plurality of functions (performance analysis unit 31, playback control unit 32, and instruction reception unit 33) for controlling the playback of the playback part performed by the performance device 20 by executing a program stored in the storage device 12.

[0044] The performance analysis unit 31 estimates the performance position X within the music piece by analyzing the audio signal Z supplied from the pickup device 13. The performance position X is the time point at which the user U actually performs within the music piece. The estimation of the performance position X is repeatedly executed in parallel with the performance of the performance part by the user U and the playback of the playback part by the performance device 20. That is, the performance position X is estimated at each of a plurality of time points on the time axis. The performance position X moves backward within the music piece as time passes.

[0045] Specifically, the performance analysis unit 31 calculates the performance position X by comparing the reference data Da of the music piece data D and the audio signal Z with each other. For the estimation of the performance position X performed by the performance analysis unit 31, any known analysis technique (sheet music positioning technique) can be arbitrarily adopted. For example, the analysis technique disclosed in Japanese Patent Application Laid-Open No. 2016-099512 is used for the estimation of the performance position X. In addition, the performance analysis unit 31 may estimate the performance position X using a statistical estimation model such as a deep neural network or a hidden Markov model.

[0046] The playback control unit 32 causes the performance device 20 to play each note specified by the performance data Db. That is, the playback control unit 32 causes the performance device 20 to perform the automatic performance of the playback part. Specifically, the playback control unit 32 moves the position to be played within the music piece (hereinafter, referred to as "playback position") Y backward in time, and sequentially supplies the instruction data corresponding to the playback position Y in the performance data Db to the performance device 20. That is, the playback control unit 32 acts as a sequencer that sequentially supplies each instruction data included in the performance data Db to the performance device 20. The process in which the playback control unit 32 causes the performance device 20 to play the playback part is executed in parallel with the performance of the performance part by the user U.

[0047] The playback control unit 32 causes the playback of the playback voice by the playback device 20 to follow the performance of the music piece by the user U according to the estimation result of the performance position X by the performance analysis unit 31. That is, the automatic performance of the playback voice by the playback device 20 proceeds at the same tempo as the performance of the performance voice by the user U. For example, when the advancement of the performance position X (i.e., the performance speed by the user U) is fast, the playback control unit 32 increases the advancement speed of the playback position Y (the playback speed by the playback device 20), and when the advancement of the performance position X is slow, the playback control unit 32 decreases the advancement speed of the playback position Y. That is, in order to synchronize with the advancement of the performance position X, the automatic performance of the playback voice is executed at the same performance speed as the performance by the user U. Therefore, the user U can perform the performance voice with the feeling that the playback device 20 performs the playback voice in matching with his / her own performance.

[0048] As described above, in the first embodiment, the playback of multiple notes of the playback voice follows the performance of the musical instrument 80 by the user U, and thus the intention (e.g., performance expression) or preference of the user U can be appropriately reflected in the playback of the playback voice.

[0049] The instruction receiving unit 33 receives the first instruction Q1 and the second instruction Q2 from the user U. The first instruction Q1 and the second instruction Q2 are generated by the user U's operation on the operation device 14. The first instruction Q1 is an instruction to temporarily stop the playback of the playback voice by the playback device 20. The second instruction Q2 is an instruction to resume the playback of the playback voice stopped by the first instruction Q1.

[0050] Specifically, the instruction receiving unit 33 receives the operation of the user U to switch the operation device 14 from the released state to the operated state as the first instruction Q1. That is, the user U gives the first instruction Q1 to the audio processing system 10 by stepping on the movable part 141 of the operation device 14. For example, the instruction receiving unit 33 determines the time point when the movable part 141 starts to move from the position H1 (released state) toward the position H2 (operated state) as the time point of the first instruction Q1. In addition, a structure can also be conceived in which the instruction receiving unit 33 determines the time point when the movable part 141 reaches a position midway between the position H1 and the position H2 as the time point of the first instruction Q1, or a structure in which the instruction receiving unit 33 determines the time point when the movable part 141 reaches the position H2 as the time point of the first instruction Q1.

[0051] In addition, an instruction receiving unit 33 receives an operation in which a user U causes the operation device 14 to transition from an operation state to a released state as a second instruction Q2. That is, the user U gives the second instruction Q2 to the audio processing system 10 by releasing the movable part 141 of the operation device 14 from the state of stepping on the movable part 141. For example, the instruction receiving unit 33 determines the time point when the movable part 141 starts to move from the position H2 (operation state) toward the position H1 (released state) as the time point of the second instruction Q2. In addition, a structure can also be conceived in which the instruction receiving unit 33 determines the time point when the movable part 141 reaches a position midway between the position H2 and the position H1 as the time point of the second instruction Q2, or a structure in which the instruction receiving unit 33 determines the time point when the movable part 141 reaches the position H1 as the time point of the second instruction Q2.

[0052] The user U can give the first instruction Q1 and the second instruction Q2 at any time point during the performance of the performance part. Therefore, the interval between the first instruction Q1 and the second instruction Q2 is a variable length corresponding to the intention of the user U. For example, the user U gives the first instruction Q1 before the start of a rest period in the music, and gives the second instruction Q2 at a time point after the rest period of the desired time length has elapsed.

[0053] Figure 5 It is an explanatory diagram of the relationship between the playback of the playback part by the performance device 20 and the first instruction Q1 and the second instruction Q2. The pronunciation periods of the respective notes specified by the performance data Db and the pronunciation periods of the respective notes actually played by the performance device 20 are recorded together in Figure 5 .

[0054] Figure 5 The note N1 of is one of the multiple notes specified by the performance data Db and corresponds to the first instruction Q1. Specifically, the note N1 is the note played by the performance device 20 at the time point of the first instruction Q1 among the multiple notes of the playback part. After generating the first instruction Q1, the playback control unit 32 causes the performance device 20 to continuously play the note N1 until the end point of the pronunciation period specified by the performance data Db for the note N1. For example, the playback control unit 32 supplies instruction data instructing the silencing of the note N1 to the performance device 20 at the end point of the pronunciation period of the note N1. As understood from the above description, the playback of the note N1 does not stop immediately at the time point of the first instruction Q1, but continues after the generation of the first instruction Q1 until the end point specified by the performance data Db. In addition, the note N1 is an example of the "first sound".

[0055] Figure 5The note N2 is the note immediately following note N1 among the multiple notes specified by the performance data Db. After stopping the playback of note N1, the playback control unit 32 uses the second instruction Q2 issued by the user U as an opportunity to start the playback of note N2 on the performance device 20. That is, regardless of the position of the start point of the sounding period specified by the performance data Db for note N2 and the time length of the interval between note N1 and note N2 specified by the performance data Db, the playback of note N2 is started on the condition that the second instruction Q2 is generated. Specifically, when the instruction receiving unit 33 receives the second instruction Q2, the playback control unit 32 supplies the instruction data of note N2 of the performance data Db to the performance device 20. Therefore, the playback of note N2 starts immediately after the second instruction Q2. In addition, note N2 is an example of the "second sound".

[0056] Figure 6 It is a flowchart exemplifying the specific process of the operation (hereinafter referred to as "playback control process") Sa in which the control device 11 controls the performance device 20. The playback control process Sa starts on the occasion of an instruction from the user U.

[0057] If the playback control process Sa is started, the control device 11 determines whether the waiting data W is in a valid state (Sa1). The waiting data W is data (e.g., a flag) indicating a state in which the playback of the playback voice part is temporarily stopped by the first instruction Q1 and is stored in the storage device 12. Specifically, the waiting data W is set to a valid state (e.g., W = 1) when the first instruction Q1 is generated and is set to an invalid state (e.g., W = 0) when the second instruction Q2 is generated. The waiting data W can also be referred to as data indicating a state of waiting for the resumption of the playback of the playback voice part.

[0058] When the waiting data W is not in a valid state (Sa1: NO), the control device 11 (performance analysis unit 31) estimates the performance position X by analyzing the audio signal Z supplied from the sound pickup device 13 (Sa2). The control device 11 (playback control unit 32) advances the playback of the playback voice part performed by the performance device 20 according to the estimation result of the performance position X (Sa3). That is, the control device 11 controls the playback of the playback voice part performed by the performance device 20 to follow the performance of the performance voice part by the user U.

[0059] The control device 11 (instruction receiving unit 33) determines whether the first instruction Q1 is received from the user U (Sa4). When the first instruction Q1 is received (Sa4: YES), the control device 11 (playback control unit 32) causes the performance device 20 to continuously play the note N1 that was being played at the time when the first instruction Q1 was received until the end of the sounding period specified by the performance data Db (Sa5). Specifically, the control device 11 advances the playback position Y at the same speed (rhythm) as the time point when the first instruction Q1 was generated. When the playback position Y reaches the end of the sounding period of the note N1, the instruction data indicating the muting of the note N1 is supplied to the performance device 20. If the above processing is executed, the control device 11 changes the wait data W from the invalid state to the valid state (W = 1) (Sa6). In addition, the update of the wait data W may be executed before step Sa5 (Sa6).

[0060] If the wait data W is set to the valid state, the result of the determination in step Sa1 is affirmative. When the wait data W is in the valid state (Sa1: YES), the estimation of the performance position X (Sa2), the playback control of the playback voice part (Sa3), and the processing related to the note N1 (Sa4 - Sa6) are not executed. That is, taking the first instruction Q1 from the user U as an opportunity, the playback control of the playback voice part linked to the performance position X is stopped. In addition, when the first instruction Q1 is not received (Sa4: NO), the processing related to the note N1 (Sa5, Sa6) is not executed.

[0061] The control device 11 (instruction receiving unit 33) determines whether the second instruction Q2 is received from the user U (Sa7). When the second instruction Q2 is received (Sa7: YES), the control device 11 (playback control unit 32) causes the performance device 20 to play the note N2 immediately following the note N1 (Sa8). Specifically, the control device 11 updates the playback position Y to the start point of the note N2. That is, through the second instruction Q2, the playback of the playback voice part that was stopped by the first instruction Q1 is restarted. The control device 11 changes the wait data W from the valid state to the invalid state (W = 0) (Sa9). As described above, if the wait data W is set to the invalid state, the result of the determination in step Sa1 is negative. Therefore, taking the second instruction Q2 as an opportunity, the estimation of the performance position X (Sa2) and the playback control of the playback voice part (Sa3) are restarted. In addition, the update of the wait data W may be executed before step Sa8 (Sa8).

[0062] The control device 11 determines whether to end the playback of the playback part performed by the playback device 20 (Sa10). For example, when the playback until the end point of the playback part is completed, or when the user U instructs the end, the control device 11 determines to end the playback of the playback part. When the playback of the playback part is not ended (Sa10: NO), the control device 11 advances the process to step Sa1 and repeats the processes exemplified above (Sa1 - Sa9). On the other hand, when the control device 11 determines to end the playback of the playback part (Sa10: YES), the playback control process Sa ends.

[0063] As described above, in the first embodiment, after playing the note N1 corresponding to the first instruction Q1 and stopping the playback of the note N1, taking the second instruction Q2 issued by the user U as an opportunity, the playback of the note N2 immediately following the note N1 is started. Therefore, the interval between the playback of the note N1 and the playback of the note N2 (for example, the time length during the rest in the music piece) can be changed corresponding to the respective time points of the first instruction Q1 and the second instruction Q2.

[0064] In addition, in the first embodiment, regarding the playback of the note N1 being played at the time point when the first instruction Q1 is generated, the playback continues until the end point of the note N1 specified by the performance data Db even after the first instruction Q1 is generated. Therefore, compared with the structure in which the playback of the note N1 is stopped at the time point when the first instruction Q1 is generated, the playback of the note N1 can be appropriately continued corresponding to the content of the performance data Db.

[0065] In the first embodiment, the user U operates the operation device 14, whereby the interval between the note N1 and the note N2 can be changed to an appropriate time length corresponding to the intention or preference of the user U. In the first embodiment, in particular, the first instruction Q1 is generated by switching the operation device 14 from the released state to the operated state, and after maintaining this operated state, the second instruction Q2 is generated by switching the operation device 14 from the operated state to the released state at a desired time point after the first instruction Q1 is generated. That is, the first instruction Q1 and the second instruction Q2 are generated by a series of operations of switching the operation state from the released state to the operated state and then switching back to the released state. Therefore, compared with the structure in which operations of switching the operation device 14 from the released state to the operated state are required respectively for the first instruction Q1 and the second instruction Q2, the operation of the user U on the operation device 14 is simplified.

[0066] B: Second Embodiment

[0067] A description will be given of the second embodiment. In addition, in each of the embodiments exemplified below, elements having the same functions as those in the first embodiment are denoted by the same reference numerals as those used in the description of the first embodiment, and their detailed descriptions are appropriately omitted.

[0068] In the first embodiment, when the first instruction Q1 is generated, the playback position Y is advanced at the same speed as the time point of the first instruction Q1, and when the playback position Y reaches the end point of the note N1, the playback of the note N1 is stopped. The playback control unit 32 of the second embodiment variably controls the advancing speed of the playback position Y after the generation of the first instruction Q1 (i.e., the playback speed of the playback voice part) in accordance with the operation speed V1 of the movable unit 141. The operation speed V1 is the speed at which the movable unit 141 moves from the position H1 corresponding to the released state toward the position H2 corresponding to the operated state. For example, the average value of a plurality of speeds calculated during the period in which the movable unit 141 moves from the position H1 to the position H2 is the operation speed V1.

[0069] Figure 7 It is an explanatory diagram related to the state of the operation device 14 of the second embodiment. As Figure 7 illustrated, at the time point when the movable unit 141 starts moving from the position H1 toward the position H2, the instruction receiving unit 33 receives the first instruction Q1. The playback control unit 32 controls the advancing speed of the playback position Y after the generation of the first instruction Q1 in accordance with the operation speed V1 of the movable unit 141.

[0070] Specifically, the faster the operation speed V1 is, the higher the advancing speed of the playback position Y is made by the playback control unit 32. For example, as Figure 7 illustrated, the advancing speed of the playback position Y when the operation speed V1 is the speed V1_H is greater than the advancing speed of the playback position Y when the operation speed V1 is the speed V1_L (V1_L < V1_H). Therefore, the faster the operation speed V1 is, the shorter the duration of the note N1 is. For example, the duration of the note N1 when the operation speed V1 is the speed V1_H is shorter than the duration of the note N1 when the operation speed V1 is the speed V1_L.

[0071] In the second embodiment, the same effects as those in the first embodiment are also achieved. In the second embodiment, since the duration of the musical note N1 is controlled in accordance with the operation speed V1, there is an advantage that the user U can adjust the duration of the musical note N1. Further, in the second embodiment, the operation device 14 for giving the first instruction Q1 and the second instruction Q2 is also used for adjusting the duration of the musical note N1. Therefore, compared with a configuration in which the giving of the first instruction Q1 and the second instruction Q2 and the adjustment of the duration of the musical note N1 are each operated by the user U on separate devices, there is also an advantage that the operation by the user U is simplified.

[0072] C: Third Embodiment

[0073] In the first embodiment, the playback of the musical note N2 starts immediately after the second instruction Q2. In the second embodiment, the time from the second instruction Q2 until the start of the playback of the musical note N2 (hereinafter referred to as "delay time") is controlled to be variable in accordance with the operation speed V2. The operation speed V2 is the speed at which the movable part 141 moves from the position H2 corresponding to the operation state toward the position H1 corresponding to the release state. For example, the average value of a plurality of speeds calculated during the period in which the movable part 141 moves from the position H2 to the position H1 is the operation speed V2.

[0074] Figure 8 It is an explanatory diagram related to the state of the operation device 14 of the third embodiment. As Figure 8 illustrated, at the time point when the movable part 141 starts to move from the position H2 toward the position H1, the instruction receiving unit 33 receives the second instruction Q2. The playback control unit 32 controls the delay time to be variable in accordance with the operation speed V2.

[0075] Specifically, the playback control unit 32 shortens the delay time as the operation speed V2 becomes faster. For example, as Figure 8 illustrated, the delay time when the operation speed V2 is the speed V2_L is longer than the delay time when the operation speed V2 is the speed V2_H (V2_H > V2_L). Therefore, the slower the operation speed V2, the later the time point at which the playback of the musical note N2 starts becomes on the time axis.

[0076] In the third embodiment, the same effects as those in the first embodiment are also achieved. In the third embodiment, since the time point at which the playback of note N2 starts is controlled in accordance with the operation speed V2, there is an advantage that the user U can adjust the start point of the first note N2 after the restart of the playback voice part. In addition, in the third embodiment, the operation device 14 for giving the first instruction Q1 and the second instruction Q2 also serves for adjusting the start point of note N2. Therefore, compared with the configuration in which the user U operates separate devices for giving the first instruction Q1 and the second instruction Q2 and for adjusting the start point of note N2, there is also an advantage that the operation by the user U is simplified. In addition, the configuration of the second embodiment can be applied to the third embodiment.

[0077] D: Fourth Embodiment

[0078] Figure 9 FIG. is a block diagram illustrating the functional configuration of the audio processing system 10 according to the fourth embodiment. The control device 11 according to the fourth embodiment functions also as an edit processing unit 34 on the basis of the same elements (performance analysis unit 31, playback control unit 32, and instruction reception unit 33) as those in the first embodiment. The edit processing unit 34 edits the performance data Db stored in the storage device 12 in accordance with an instruction from the user U. The operations of the elements other than the edit processing unit 34 are the same as those in the first embodiment. Therefore, in the fourth embodiment, the same effects as those in the first embodiment are also achieved. In addition, the configuration of the second embodiment or the third embodiment can be applied to the fourth embodiment.

[0079] Figure 10 FIG. is an explanatory diagram of the operation of the edit processing unit 34. In Figure 10 , the notes N1 and N2 specified by the performance data Db of the playback voice part are illustrated. Similar to the foregoing embodiments, the first instruction Q1 and the second instruction Q2 are generated by the user U at an arbitrary time point. Therefore, a time difference L is generated between the start point of the note N2 specified by the performance data Db and the time point of the second instruction Q2. The edit processing unit 34 edits the performance data Db so as to reduce the time difference L.

[0080] Figure 11 FIG. is a flowchart illustrating a specific process (hereinafter referred to as "edit process") Sb in which the edit processing unit 34 edits the performance data Db. For example, every time the playback (the foregoing playback control process Sa) of the playback voice part by the performance device 20 is repeated a specified number of times, the edit process Sb is executed. In addition, the edit process Sb may be started on the basis of an instruction from the user U.

[0081] If the editing process Sb starts, the editing processing unit 34 calculates the dispersion Δ of the time difference L of the past specified number of play control processes Sa (Sb1). The dispersion Δ is a statistic representing the degree of dispersion related to multiple time differences L. For example, the dispersion, standard deviation, or range of distribution of multiple time differences L is used as the dispersion Δ.

[0082] The editing processing unit 34 determines whether the dispersion Δ is greater than the threshold value Δth (Sb2). When the dispersion Δ is greater than the threshold value Δth, it is presumed that the user U practices playing the music while consciously varying the waiting time until the playback of the note N2 restarts. Therefore, it is inappropriate to edit the performance data Db corresponding to multiple time differences L. On the other hand, when the dispersion Δ is less than the threshold value Δth, it is presumed that the multiple time differences L are values that conform to the intention or preference of the user U (i.e., appropriate values inherent to the user U).

[0083] Considering the above tendency, when the dispersion Δ is less than the threshold value Δth (Sb2: NO), the editing processing unit 34 edits the performance data Db corresponding to multiple time differences L (Sb3 - Sb4). On the other hand, when the dispersion Δ is greater than the threshold value Δth (Sb2: YES), the editing processing unit 34 ends the editing process Sb without performing the editing of the performance data Db (Sb3, Sb4).

[0084] In the editing of the performance data Db, the editing processing unit 34 calculates the average time difference La by averaging multiple time differences L (Sb3). Then, the editing processing unit 34 changes the start point of the note N2 specified by the performance data Db by the amount of the average time difference La (Sb4). For example, when the average time difference La is negative, the editing processing unit 34 moves the start point of the note N2 specified by the performance data Db forward by an amount of time corresponding to the average time difference La. In addition, when the average time difference La is positive, the editing processing unit 34 moves the start point of the note N2 specified by the performance data Db backward by an amount of time corresponding to the average time difference La. That is, when there is a tendency for the user U to sufficiently ensure the waiting time immediately before the note N2, the start point of the note N2 specified by the performance data Db is changed backward, and when there is a tendency for the waiting time to be short, the start point of the note N2 specified by the performance data Db is changed forward.

[0085] As understood from the above description, in the fourth embodiment, the performance data Db is edited corresponding to the time difference L in the performance of the performance part by the user U. Therefore, the tendency inherent to each user U can be reflected in the performance data Db.

[0086] E: Fifth Embodiment

[0087] Figure 12 1 is a block diagram illustrating the configuration of a playback system 100 according to a fifth embodiment. The sound processing system 10 according to the fifth embodiment includes the same elements (control device 11, storage device 12, sound pickup device 13, and operation device 14) as the sound processing system 10 according to the first embodiment, and further includes a display device 15. The display device 15 displays an image instructed by the control device 11. The display device 15 is, for example, a liquid crystal display panel or an organic EL display panel.

[0088] The user U1 plays the musical instrument 80 in the same manner as in the first embodiment. The musical instrument 80 is a natural musical instrument such as a string instrument that produces sound by the performance performed by the user U1. On the other hand, the performance device 20 of the fifth embodiment is an automatic musical instrument that functions as an electronic musical instrument that can be manually performed by the user U2 in addition to functioning as a playback device that performs automatic performance of the playback part of the music. Specifically, the performance device 20 has a driving mechanism 21 and a sounding mechanism 22 in the same manner as the performance devices 20 of the aforementioned various modes.

[0089] The manual performance by the user U2 is a performance performed by, for example, pressing keys on a keyboard or by the body movements of the user U2. The sound generating mechanism 22 operates in conjunction with the performance by the user U2, thereby producing musical sounds from the performance device 20. In addition, the performance device 20 sequentially outputs instruction data d indicating instructions for the performance to the sound processing system 10 in parallel with the performance by the user U2. The instruction data d specifies, for example, pitch and intensity of sound generation to specify an action such as sound generation or sound silencing.

[0090] Figure 13 1 is a block diagram illustrating a functional structure of a sound processing system 10 according to a fifth embodiment. The control device 11 according to the fifth embodiment executes a program stored in a storage device 12, thereby functioning as a preparation processing unit 35 and a display control unit 36 in addition to the same elements as those of the first embodiment (performance analysis unit 31, playback control unit 32, and instruction receiving unit 33).

[0091] The preparation processing unit 35 generates music data D (reference data Da and performance data Db) used in the playback control process Sa. Specifically, the preparation processing unit 35 generates music data D corresponding to the performance of the musical instrument 80 performed by the user U1 and the performance of the performance device 20 performed by the user U2. The display control unit 36 causes the display device 15 to display various images.

[0092] The preparation processing unit 35 includes a first recording unit 41, a second recording unit 42, and an audio analysis unit 43. The audio analysis unit 43 generates reference data Da used in the playback control process Sa. Specifically, the audio analysis unit 43 performs an adjustment process Sc ( Figure 15 The reference data Da generated by the sound analysis unit 43 is stored in the storage device 12.

[0093] During the period before the adjustment process Sc is executed (hereinafter referred to as the "preparation period"), the user U1 and the user U2 play the music together. Specifically, during the preparation period, the user U1 plays the performance part of the music through the musical instrument 80, and the user U2 plays the playback part of the music through the performance device 20. The adjustment process Sc is a process of generating reference data Da using the results of the music played by the user U1 and the user U2 during the preparation period. In addition, the user U1 may sing the performance part of the music facing the sound pickup device 13.

[0094] The first recording unit 41 acquires the acoustic signal Z generated by the sound collecting device 13 during the preparation period. In the following, the acoustic signal Z acquired by the first recording unit 41 during the preparation period is referred to as a “reference signal Zr” for convenience. The first recording unit 41 stores the reference signal Zr in the storage device 12 .

[0095] During the preparation period, in addition to the musical sound produced by the musical instrument 80 through the performance by the user U1, the musical sound produced by the performance device 20 through the performance by the user U2 also reaches the sound pickup device 13. Therefore, the reference signal Zr is an acoustic signal including the acoustic components of the musical instrument 80 and the acoustic components of the performance device 20. In addition, the musical instrument 80 is an example of the "first sound source", and the performance device 20 is an example of the "second sound source". In the case where the user U1 sings the performance part, the user U1 corresponds to the "first sound source".

[0096] The second recording section 42 obtains the performance data Db representing the performance of the performance device 20 during the preparation period. Specifically, the second recording section 42 generates the performance data Db in a MIDI format in which the instruction data d sequentially supplied from the performance device 20 corresponding to the performance performed by the user U2 and the time data specifying the interval between the previous and next instruction data d are arranged in a time series. The second recording section 42 saves the performance data Db to the storage device 12. The performance data Db stored in the storage device 12 is used for the playback control process Sa as illustrated in the above-mentioned various modes. In addition, the performance data Db obtained by the second recording section 42 can also be edited by the editing processing section 34 illustrated in the fourth embodiment.

[0097] As understood from the above description, during the preparation period, the reference signal Zr and the performance data Db are saved to the storage device 12. The sound analysis unit 43 generates the reference data Da through the adjustment process Sc, which utilizes the reference signal Zr acquired by the first recording unit 41 and the performance data Db acquired by the second recording unit 42.

[0098] Figure 14 is a block diagram illustrating the specific structure of the sound analysis unit 43. As Figure 14 illustrated, the sound analysis unit 43 includes an index calculation unit 51, a period estimation unit 52, a pitch estimation unit 53, and an information generation unit 54. In addition, the operation device 14 of the fifth embodiment receives an indication related to the value of each of a plurality of variables (α, β, γ) applicable to the adjustment process Sc from the user U (U1 or U2). That is, the user U can set or change the value of each variable by operating the operation device 14.

[0099] [Index calculation unit 51]

[0100] The index calculation unit 51 calculates the pronunciation index C(t). The symbol t refers to one time point on the time axis. That is, the index calculation unit 51 determines the time series of the pronunciation index C(t) corresponding to different time points t on the time axis. The pronunciation index C(t) is an index of the accuracy (likelihood or probability) that the reference signal Zr contains the sound components of the musical instrument 80. That is, at the time point t on the time axis, the higher the accuracy that the sound components of the musical instrument 80 are included in the reference signal Zr, the larger the value of the pronunciation index C(t) is set. The index calculation unit 51 of the first embodiment includes a first analysis unit 511, a second analysis unit 512, a first operation unit 513, and a second operation unit 514.

[0101] The first analysis unit 511 calculates the first index B1(t,n) by analyzing the reference signal Zr. The symbol n refers to any one of the N pitches P1 to PN (n = 1 to N). Specifically, the first analysis unit 511 calculates the N first indexes B1(t,1) to B1(t,N) corresponding to different pitches Pn for each time point t on the time axis. That is, the first analysis unit 511 calculates the time series of the first index B1(t,n).

[0102] The first index B1(t,n) corresponding to the pitch Pn is an index of the accuracy with which the acoustic component of the pitch Pn of the performance device 20 or the musical instrument 80 is included in the reference signal Zr, and is set to a value within the range of 0 or more and 1 or less. For one or both of the performance device 20 and the musical instrument 80, the greater the intensity of the acoustic component of the pitch Pn, the greater the value to which the first index B1(t,n) is set. As understood from the above description, the first index B1(t,n) can also be alternatively referred to as an index of the reference signal Zr related to the intensity of the acoustic component of each pitch Pn. For the calculation of the first index B1(t,n) by the first analysis unit 511, a known acoustic analysis technique (particularly, pitch estimation technique) can be arbitrarily adopted.

[0103] The second analysis unit 512 calculates the second index B2(t,n) by analyzing the performance data Db. Specifically, the second analysis unit 512 calculates N second indices B2(t,1) to B2(t,N) corresponding to different pitches Pn for each time point t on the time axis. That is, the second analysis unit 512 calculates the time series of the second index B2(t,n).

[0104] The second index B2(t,n) corresponding to the pitch Pn is an index corresponding to the pronunciation intensity of the note of the pitch Pn specified by the performance data Db at the time point t, and is set to a value within the range of 0 or more and 1 or less. The greater the pronunciation intensity of the note of the pitch Pn specified by the performance data Db, the greater the value of the second index B2(t,n). When there is no note of the pitch Pn at the time point t, the second index B2(t,n) is set to 0.

[0105] For the calculation of the second index B2(t,n), a variable α set corresponding to an instruction from the user U is applied. For example, the second analysis unit 512 calculates the second index B2(t,n) by the operation of the following mathematical formula (1).

[0106]

Mathematical formula 1

[0107] B2(t, n) = 1 - exp{-c·v(t,n)·α} (1)

[0108] The symbol ν(t,n) in the mathematical formula (1) is a value corresponding to the pronunciation intensity of the note of the pitch Pn specified by the performance data Db at the time point t on the time axis. When the time point t on the time axis is within the pronunciation period of the note of the pitch Pn, the intensity ν(t,n) is set to the pronunciation intensity of the note specified by the performance data Db. On the other hand, when the time point t on the time axis is outside the pronunciation period of the pitch Pn, the intensity ν(t,n) is set to 0. The symbol c in the mathematical formula (1) is a coefficient and is set to a prescribed positive number.

[0109] As understood from Equation (1), when the variable α is small, even if the intensity ν(t, n) is a large value, the second index B2(t, n) is set to a small value. On the other hand, when the variable α is large, even if the intensity ν(t, n) is a small value, the second index B2(t, n) is set to a large value. That is, there is a tendency that the smaller the variable α, the smaller the second index B2(t, n) becomes.

[0110] Figure 14 The first arithmetic unit 513 calculates the pronunciation index E(t, n) by subtracting the second index B2(t, n) from the first index B1(t, n). The pronunciation index E(t, n) is an index of the accuracy (likelihood or probability) that the sound component of the pitch Pn of the musical instrument 80 is included in the reference signal Zr at the time point t. Specifically, the first arithmetic unit 513 calculates the pronunciation index E(t, n) by the operation of the following Equation (2).

[0111]

Equation 2

[0112] E(t, n) = max{0, B1(t, n) - B2(t, n)} (2)

[0113] max{a, b} in Equation (2) refers to the operation of selecting the larger one of the numerical values a and the numerical value b. As understood from Equation (2), the pronunciation index E(t, n) is a numerical value within the range of 0 or more and 1 or less. The greater the intensity of the sound component of the pitch Pn of the musical instrument 80, the larger the pronunciation index E(t, n) is set. That is, the pronunciation index E(t, n) can also be referred to as an index related to the intensity of the sound component (pitch Pn) of the musical instrument 80 of the reference signal Zr.

[0114] As described above, the acoustic components of both the performance device 20 and the musical instrument 80 affect the first index B1(t,n). On the other hand, only the acoustic components of the performance device 20 affect the second index B2(t,n). Therefore, in Equation (2), the operation of subtracting the second index B2(t,n) from the first index B1(t,n) is equivalent to the process of suppressing the influence of the acoustic components of the performance device 20 from the first index B1(t,n). That is, the pronunciation index E(t,n) is equivalent to an index related to the intensity of the acoustic components of the musical instrument 80 (pitch Pn) in the acoustic components of the reference signal Zr. As described above, there is a tendency that the larger the variable α, the larger the value set for the second index B2(t,n). Therefore, the variable α is a variable for controlling the degree of suppressing the influence of the acoustic components of the performance device 20 from the first index B1(t,n). That is, the larger the variable α (the larger the second index B2(t,n)), the more the influence of the acoustic components of the performance device 20 is suppressed in the pronunciation index E(t,n).

[0115] The second arithmetic unit 514 calculates the pronunciation index C(t) based on the pronunciation index E(t,n) calculated by the first arithmetic unit 513. Specifically, the second arithmetic unit 514 calculates the pronunciation index C(t) by the following Equation (3).

[0116]

Equation 3

[0117] C(t) = max{E(t,1), E(t,2), …, E(t,N)} (3)

[0118] As understood from Equation (3), the maximum value of the N pronunciation indexes E(t,1) to E(t,N) corresponding to different pitches Pn is selected as the pronunciation index C(t) at the time point t. As understood from the above description, the pronunciation index C(t) is an index of the accuracy with which the acoustic components of the musical instrument 80 corresponding to any one of the N pitches P1 to PN are included in the reference signal Zr. The larger the value of the variable α (the larger the second index B2(t,n)), the more the influence of the acoustic components of the performance device 20 on the pronunciation index C(t) is suppressed. That is, the pronunciation index C(t) during the period when the acoustic components of the performance device 20 predominantly exist becomes a small value. On the other hand, the pronunciation index C(t) during the period when there are no acoustic components of the performance device 20 hardly changes even when the value of the variable α is changed.

[0119] [Period estimation unit 52]

[0120] Figure 14The period estimation unit 52 estimates the period during which the sound components of the musical instrument 80 exist on the time axis (hereinafter referred to as the "performance period"). Specifically, the period estimation unit 52 calculates the sounding index G(t) using the pronunciation index C(t) calculated by the index calculation unit 51. The sounding index G(t) is an index indicating whether there is a sound component (sounding / non-sounding) of the musical instrument 80 at time t. The sounding index G(t) at the time point t when there is a sound component of the musical instrument 80 is set to the value g1 (for example, g1 = 1), and the sounding index G(t) at the time point t when there is no sound component of the musical instrument 80 is set to the value g0 (for example, g0 = 0). The period composed of one or more time points t where the sounding index G(t) is the value g1 corresponds to the performance period. The performance period is an example of the "pronunciation period".

[0121] When estimating the performance period, the first HMM (Hidden Markov Model) is used. The first HMM is a state transition model composed of a sounding state corresponding to sounding (value g1) and a non-sounding state corresponding to non-sounding (value g0). Specifically, the period estimation unit 52 calculates the sequence of the maximum likelihood states generated by the first HMM through Viterbi Detection as the sounding index G(t).

[0122] The probability of the sounding state occurring in the first HMM (hereinafter referred to as the "sounding probability") Λ is expressed by the following mathematical formula (4). The symbol σ in the mathematical formula (4) is a sigmoid function. The probability of the non-sounding state occurring is (1 - Λ). In addition, the probability of maintaining the sounding state or the non-sounding state between two adjacent time points t on the time axis is set to a prescribed constant (for example, 0.9).

[0123]

Mathematical formula 4

[0124] Λ = σ{C(t) - β} (4)

[0125] As understood from the mathematical formula (4), for the calculation of the sounding probability Λ, the variable β set corresponding to the instruction from the user U is applied. Specifically, the larger the variable β, the smaller the sounding probability Λ is set. Therefore, there is a tendency that the larger the variable β, the easier it is for the sounding index G(t) to be set to the value g0, and as a result, the performance period is more likely to be a short period. On the other hand, the smaller the variable β, the larger the sounding probability Λ is set. Therefore, there is a tendency that the sounding index G(t) is easily set to the value g1, and as a result, the performance period is likely to be a long period.

[0126] In addition, as described above, the larger the variable α is, the smaller the pronunciation index C(t) becomes during the period when the sound components of the performance device 20 are dominant. As understood from Equation (4), the smaller the pronunciation index C(t) is, the smaller the sounding probability Λ is set. Therefore, there is a tendency that the larger the variable α is, the easier it is for the sounding index G(t) to be set to the value g0, and as a result, the performance period tends to be short. As understood from the above description, the variable α affects not only the pronunciation index C(t) but also the performance period. That is, the variable α affects both the pronunciation index C(t) and the sounding index G(t). On the other hand, the variable β only affects the sounding index G(t).

[0127] [Pitch estimation unit 53]

[0128] The pitch estimation unit 53 determines the pitch K(t) of the sound components of the musical instrument 80. That is, the pitch estimation unit 53 determines the time series of the pitch K(t) corresponding to different time points t on the time axis. The pitch K(t) can be set to any one of the N pitches P1 to PN.

[0129] When estimating the pitch K(t), the second HMM is used. The second HMM is a state transition model composed of N states corresponding to different pitches Pn. The probability density function ρ(x|μ n , κ n ) of the observation probability x of the pitch Pn is the von Mises-Fisher distribution represented by the following Equation (5).

[0130]

Equation 5

[0131] ρ(x|μ n , κ n ) ∝ exp{κ n x T μ n / (||x||||μ n ||)} (5)

[0132] The notation T in Equation (5) refers to the transpose of a matrix, and the notation || || refers to the norm. The notation μ n in Equation (5) refers to the position parameter, and the notation κ n refers to the concentration parameter. The position parameter μ n and the concentration parameter κ n are set by machine learning using the pronunciation index E(t, n).

[0133] In the second HMM, the transition probability λ(n1, n2) from pitch Pn1 to pitch Pn2 is represented by the following equation (6) (n1 = 1 to N, n2 = 1 to N, n1 ≠ n2). For all combinations of selecting two pitches Pn (Pn1, Pn2) from N pitches P1 to PN, the transition probability λ(n1, n2) is obtained through equation (6).

[0134]

Equation 6

[0135] λ(n1, n2) = {I + γ·τ(n1, n2)} / (1 + γ) (6)

[0136] The notation I in equation (6) refers to the N-dimensional identity matrix. The notation τ(n1, n2) refers to the probability of the transition from pitch Pn1 to pitch Pn2, which is set through machine learning of existing musical scores. In addition, at the time point t when the pitch indicator G(t) has the value g0 (no pitch), the transition probability from pitch Pn1 to pitch Pn2 is set to the identity matrix I, and the observation probability x is set to a specified constant.

[0137] As can be understood from equation (6), for the estimation of pitch K(t), the variable γ set according to the instruction from the user U is applied. Specifically, the smaller the variable γ, the closer the transition probability λ(n1, n2) is to the identity matrix I, so it is difficult to generate a transition from pitch Pn1 to pitch Pn2. On the other hand, the larger the variable γ, the greater the influence of the transition probability τ(n1, n2) on the transition probability λ(n1, n2), so it is easy to generate a transition from pitch Pn1 to pitch Pn2.

[0138] [Information generation unit 54]

[0139] The information generation unit 54 determines M pronunciation points T1 to TM on the time axis and the pitch Fm (F1 to FM) of each pronunciation point Tm (m = 1 to M). The number M of pronunciation points Tm in the music piece is a variable value. Specifically, the information generation unit 54 determines the time point t at which the pitch K(t) changes as the pronunciation point Tm during the performance period when the pitch indicator G(t) has the value g1. In addition, the information generation unit 54 determines the pitch K(Tm) of each pronunciation point Tm as the pitch Fm.

[0140] As described above, the smaller the variable γ, the more difficult it is to generate a transition of pitch Pn, so the number M of pronunciation points Tm and pitch Fm becomes smaller. On the other hand, the larger the variable γ, the easier it is to generate a transition of pitch Pn, so the number M of pronunciation points Tm and pitch Fm becomes larger. That is, the variable γ can also be referred to as a parameter for controlling the number M of pronunciation points Tm and pitch Fm.

[0141] In addition, the information generation unit 54 stores reference data Da including the pronunciation index E(t,n) calculated by the index calculation unit 51 (second calculation unit 514), each pronunciation point Tm on the time axis, and the voiced index G(t) calculated by the period estimation unit 52 in the storage device 12.

[0142] Figure 13 The display control unit 36 of to Figure 13 displays the analysis result obtained by the sound analysis unit 43 described above on the display device 15. Specifically, the display control unit 36 Figure 15 displays the confirmation screen 60 illustrated in Figure 15 on the display device 15. The confirmation screen 60 is an image for the user U to confirm the analysis result obtained by the sound analysis unit 43.

[0143] The confirmation screen 60 includes a first area 61 and a second area 62. A common time axis At is set for the first area 61 and the second area 62. The time axis At is an axis extending horizontally. In addition, the time axis At may be displayed as an image that the user U can visually confirm, or may not be displayed on the confirmation screen 60. The display interval of the confirmation screen 60 in the music changes according to an instruction (for example, an enlargement / reduction instruction) from the user U to the operation device 14.

[0144] The performance period 621 on the time axis At in the second area 62 and the period other than the performance period 621 (hereinafter referred to as "non-performance period") 622 are displayed in different display manners. For example, the performance period 621 and the non-performance period 622 are displayed in different hues. The performance period 621 is a range where the voiced index G(t) is set to the value g1 on the time axis At. On the other hand, the non-performance period 622 is a range where the voiced index G(t) is set to the value g0 on the time axis At. As understood from the above description, the performance period 621 estimated by the period estimation unit 52 is displayed on the confirmation screen 60.

[0145] The converted image 64 is displayed in the first area 61. The converted image 64 is an image that displays the time series of the pronunciation index C(t) calculated by the index calculation unit 51 based on the time axis At. Specifically, the part corresponding to the time point t on the time axis At in the converted image 64 is displayed in a display manner corresponding to the pronunciation index C(t). The "display manner" refers to the property of the image that can be visually recognized by the observer. For example, in addition to the three attributes of color, namely hue (tone), saturation, and brightness (gray level), patterns or shapes are also included in the concept of "display manner". For example, the gray level concentration of the part corresponding to the time point t in the converted image 64 is controlled corresponding to the pronunciation index C(t). Specifically, the part corresponding to the time point t with a large pronunciation index C(t) in the converted image 64 is displayed in dark gray, and the part corresponding to the time point t with a small pronunciation index C(t) in the converted image 64 is displayed in light gray.

[0146] The staff 65, a plurality of indication images 67, and a plurality of note images 68 are displayed in the second area 62. The staff 65 is composed of five straight lines parallel to the time axis At. Each straight line constituting the staff 65 represents a different pitch. That is, a pitch axis Ap representing pitch is set in the second area 62. The pitch axis Ap is an axis extending in the vertical direction orthogonal to the time axis At. In addition, the pitch axis Ap can be displayed as an image that can be visually confirmed by the user U, or can be displayed on the confirmation screen 60.

[0147] Each indication image 67 is an image representing one pronunciation point Tm generated by the information generation unit 54. That is, the time series of the pronunciation points Tm is represented by a plurality of indication images 67. Specifically, the indication image 67 corresponding to the pronunciation point Tm is a vertical line arranged at the position corresponding to the pronunciation point Tm on the time axis At. Therefore, a plurality of indication images 67 corresponding to different pronunciation points Tm are arranged on the time axis At.

[0148] Each note image 68 is an image representing one pitch Fm generated by the information generation unit 54. For example, an image representing the note head of a note is exemplified as the note image 68. The time series of the pitch Fm is represented by a plurality of note images 68. Since the pitch Fm is set for each pronunciation point Tm, the note images 68 are arranged for each pronunciation point Tm (that is, for each indication image 67). Specifically, the note image 68 representing the pitch Fm at the pronunciation point Tm is arranged on the line of the indication image 67 representing the pronunciation point Tm in the direction of the time axis At. In addition, the note image 68 representing the pitch Fm is arranged at the position corresponding to the pitch Fm in the direction of the pitch axis Ap. That is, each note image 68 is arranged at a position overlapping or close to the staff 65.

[0149] As illustrated above, on the confirmation screen 60, the time series of the pronunciation index C(t) (converted image 64), the time series of the pronunciation points Tm (indicator image 67), and the time series of the pitch Fm (note image 68) are displayed based on the common time axis At. Therefore, the user U can visually and intuitively confirm the temporal relationship among the pronunciation index C(t), the pronunciation points Tm, and the pitch Fm.

[0150] In addition, the confirmation screen 60 includes a plurality of operation images 71 (71a, 71b, 71c) and an operation image 72. Each operation image 71 is an operating element that the user U can operate through the operating device 14. Specifically, the operation image 71a is an image that receives an instruction from the user U to change the variable α. The operation image 71b is an image that receives an instruction from the user U to change the variable β. The operation image 71c is an image that receives an instruction from the user U to change the variable γ.

[0151] The index calculation unit 51 (second analysis unit 512) changes the variable α in response to an instruction from the user U for the operation image 71a. The index calculation unit 51 calculates the pronunciation index C(t) by an operation using the changed variable α. The display control unit 36 updates the converted image 64 of the confirmation screen 60 each time the pronunciation index C(t) is calculated. As described above, there is a tendency that as the variable α increases, the pronunciation index C(t) during the period when the sound component of the performance device 20 predominantly exists becomes a smaller value, and as a result, the sound index G(t) is more likely to be set to the value g0. Therefore, as Figure 16 illustrated, as the variable α increases, the gray scale during the period when the sound component of the performance device 20 predominantly exists in the converted image 64 changes to a lighter gray scale, and the non-performance period 622 expands. During the preparation period, the user U operates the operation image 71a while confirming the confirmation screen 60 so that the performance period 621 approaches the period when the user U1 plays the musical instrument 80.

[0152] The period estimation unit 52 changes the variable β in response to an instruction from the user U for the operation image 71b. The period estimation unit 52 calculates the sounding probability Λ by an operation using the changed variable β. The display control unit 36 updates the performance period 621 of the confirmation screen 60 each time the sounding probability Λ is calculated. As described above, there is a tendency that as the variable β increases, the sound index G(t) is more likely to be set to the value g0. Therefore, as Figure 17As illustrated, as the variable β increases, the non-playing period 622 expands. During the preparation period, the user U operates the operation image 71b while confirming the confirmation screen 60 in such a way that the playing period 621 approaches the period when the user U1 plays the musical instrument 80.

[0153] The pitch estimation unit 53 changes the variable γ in response to an instruction from the user U for the operation image 71c. The pitch estimation unit 53 calculates the transition probability λ(n1,n2) through an operation using the changed variable γ. Each time the transition probability λ(n1,n2) is calculated, the display control unit 36 updates the instruction image 67 and the note image 68 on the confirmation screen 60. As described above, as the variable γ increases, the transition probability λ(n1,n2) increases. Therefore, as Figure 18 illustrated, as the variable γ increases, the number of instruction images 67 (articulation points Tm) and the number of note images 68 (pitch Fm) increase. The user U operates the operation image 71b while confirming the confirmation screen 60 in such a way as to approximate the playing content of the user U1 during the preparation period.

[0154] The operation image 72 is an image that receives an instruction for storing the reference data Da from the user U. The information generation unit 54 stores, as the reference data Da, the content of the analysis (articulation index E(t,n), each articulation point Tm, and sounding index G(t)) at the time point when the operation image 72 is operated, in the storage device 12.

[0155] Figure 19 is a flowchart exemplifying the specific process of the adjustment process Sc. After acquiring the playing data Db and the reference signal Zr during the preparation period, the adjustment process Sc is started on the occasion of an instruction from the user U for, for example, the operation device 14.

[0156] If the adjustment process Sc is started, the sound analysis unit 43 executes an analysis process Sc1 for analyzing the reference signal Zr. The analysis process Sc1 includes an index calculation process Sc11, a period estimation process Sc12, a pitch estimation process Sc13, and an information generation process Sc14. The index calculation process Sc11 is an example of the "first process", the period estimation process Sc12 is an example of the "second process", and the pitch estimation process Sc13 is an example of the "third process". In addition, the variable α is an example of the "first variable", the variable β is an example of the "second variable", and the variable γ is an example of the "third variable".

[0157] The index calculation unit 51 calculates the pronunciation index C(t) using the performance data Db and the reference signal Zr (index calculation process Sc11). As described above, the index calculation process Sc11 includes the calculation of the first index B1(t,n) by the first analysis unit 511, the calculation of the second index B2(t,n) by the second analysis unit 512, the calculation of the pronunciation index E(t,n) by the first operation unit 513, and the calculation of the pronunciation index C(t) by the second operation unit 514. In the index calculation process Sc11, a variable α set according to an instruction from the user U is applied.

[0158] The period estimation unit 52 estimates the performance period 621 by calculating the voiced index G(t) using the pronunciation index C(t) (period estimation process Sc12). For the period estimation process Sc12, a variable β set according to an instruction from the user U is applied. In addition, the pitch estimation unit 53 estimates the pitch K(t) of the acoustic component of the musical instrument 80 (pitch estimation process Sc13). In the pitch estimation process Sc13, a variable γ set according to an instruction from the user U is applied. Moreover, the information generation unit 54 determines a plurality of pronunciation points Tm on the time axis and the pitch Fm for each pronunciation point Tm (information generation process Sc4).

[0159] The display control unit 36 displays a confirmation screen 60 showing the result of the analysis process Sc1 exemplified above on the display device 15 (Sc2). Specifically, the performance period 621 on the time axis At, the conversion image 64 representing the pronunciation index C(t), the indication image 67 representing each pronunciation point Tm, and the note image 68 representing each pitch Fm are displayed on the confirmation screen 60.

[0160] The acoustic analysis unit 43 determines whether the operation image 71 (71a, 71b, 71c) is operated (Sc3). That is, it is determined whether a change in the variable (α, β, or γ) is instructed from the user U. When the operation image 71 is operated (Sc3: YES), the acoustic analysis unit 43 executes the analysis process Sc1 applied with the changed variables (α, β, or γ) and updates the confirmation screen 60 corresponding to the result of the analysis process Sc1 (Sc2). In addition, for the calculation of the first index B1(t,n) by the first analysis unit 511, it is only necessary to execute it once in the index calculation process Sc11 after the start of the analysis process Sc1.

[0161] When the operation image 71 is not operated (Sc3: NO), the sound analysis unit 43 determines whether the operation image 72 is operated (Sc4). That is, it is determined whether the user U has indicated the determination of the reference data Da. When the operation image 72 is not operated (Sc4: NO), the sound analysis unit 43 advances the process to step Sc3. On the other hand, when the operation image 72 is operated (Sc4: YES), the information generation unit 54 saves the result of the analysis process Sc1 at the current time point (pronunciation index E(t,n), each pronunciation point Tm, and sounding index G(t)) as the reference data Da in the storage device 12 (Sc5). The adjustment process Sc ends by saving the reference data Da.

[0162] As described above, in the first embodiment, the time series of the pronunciation index C(t) (converted image 64) and the time series of the pitch Fm (note image 68) are displayed based on the common time axis At. Therefore, the user U can easily confirm and correct the analysis result of the reference signal Zr in the process of generating the reference data Da. Specifically, the user U can visually and intuitively confirm the time relationship between the pronunciation index C(t) and the pitch Fm.

[0163] In addition, in the first embodiment, the pronunciation index C(t) is calculated by subtracting the second index B2(t,n) calculated by analyzing the performance data Db from the first index B1(t,n) calculated by analyzing the reference signal Zr. Therefore, the pronunciation index C(t) with the influence of the sound components of the performance device 20 reduced can be calculated. That is, the pronunciation index C(t) that emphasizes the sound components of the musical instrument 80 can be calculated. In addition, since the variable α set according to the instruction from the user U is applied to the calculation of the second index B2(t,n), the user U can adjust the pronunciation index C(t) to be suitable for the performance of the musical instrument 80 during the preparation period.

[0164] In the first embodiment, based on the time series of the pronunciation index C(t) (converted image 64) and the time series of the pitch Fm (note image 68), the performance period 621 is also displayed based on the common time axis At. Therefore, the user U can visually and intuitively confirm the time relationship among the pronunciation index C(t), the pitch Fm, and the performance period 621. In addition, since the variable β set according to the instruction from the user U is applied to the period estimation process Sc12, the user U can adjust the performance period 621 to be suitable for the performance of the musical instrument 80 during the preparation period.

[0165] In the first embodiment, based on the time series of the pronunciation index C(t) (converted image 64) and the time series of the pitch Fm (note image 68), the time series of the pronunciation point Tm (indicator image 67) is also displayed based on the common time axis At. Therefore, the user U can visually and intuitively confirm the temporal relationship among the pronunciation index C(t), the pitch Fm, and the pronunciation point Tm. In addition, the variable γ set corresponding to the instruction from the user U is applied to the pitch estimation process Sc13, so that the user U can adjust the pronunciation point Tm to be suitable for the performance of the musical instrument 80 during the preparation period.

[0166] F: Modification Example

[0167] Hereinafter, specific modification methods added to each of the above-exemplified methods are illustrated. Two or more methods arbitrarily selected from the following illustrations can be appropriately combined within a non-contradictory range.

[0168] (1) In each of the above-described methods, the instruction receiving unit 33 receives the operation that converts the operation device 14 from the released state to the operating state as the first instruction Q1, but the method of the first instruction Q1 is not limited to the above illustration. For example, a specific action performed by the user U is detected as the first instruction Q1. For the detection of the action of the user U, various detection devices such as a photographing device or an acceleration sensor are used. For example, the instruction receiving unit 33 determines various actions such as the action of the user U raising one hand, the action of raising the musical instrument 80, or the action of breathing (e.g., the action of inhaling) as the first instruction Q1. The breathing of the user U is, for example, the breath (ventilation) when playing a wind instrument as the musical instrument 80. The operation speed V1 of the second embodiment is comprehensively represented by the speed of the action of the user U determined as the first instruction Q1.

[0169] It is also possible to include specific data representing the first instruction Q1 (hereinafter referred to as "first data") in the performance data Db. The first data is, for example, data representing a fermata within a musical piece. The instruction receiving unit 33 determines that the first instruction Q1 has occurred when the playback position Y reaches the time point of the first data. As understood from the above description, the first instruction Q1 is not limited to the instruction from the user U. In addition, when the dispersion Δ is greater than the threshold Δth in the editing process Sb, the editing processing unit 34 may attach the first data to the note N1.

[0170] (2) In each of the above-described methods, the instruction receiving unit 33 receives an operation for changing the operation device 14 from the operating state to the released state as the second instruction Q2. However, the method of the second instruction Q2 is not limited to the above examples. For example, similar to the first instruction Q1 of the first embodiment, the instruction receiving unit 33 may receive an operation for changing the operation device 14 from the released state to the operating state as the second instruction Q2. That is, two operations including stepping on and releasing the movable portion 141 can be detected as the first instruction Q1 and the second instruction Q2.

[0171] In addition, a specific action of the user U may be detected as the second instruction Q2. For the detection of the action of the user U, various detection devices such as a photographing device or an acceleration sensor are used, for example. For example, the instruction receiving unit 33 determines various actions such as an action of the user U putting down one hand, an action of lowering the musical instrument 80, or an action of breathing (for example, an exhalation action) as the second instruction Q2. The breathing of the user U is, for example, the breath (ventilation) when playing a wind instrument as the musical instrument 80. The operation speed V2 of the second embodiment is comprehensively represented by the speed of the action of the user U determined as the second instruction Q2.

[0172] Specific data representing the second instruction Q2 (hereinafter, referred to as "second data") may be included in the performance data Db. The second data is, for example, data representing a fermata within a musical piece. The instruction receiving unit 33 determines that the second instruction Q2 has occurred when the playback position Y reaches the time point of the second data. As understood from the above description, the second instruction Q2 is not limited to an instruction from the user U.

[0173] As exemplified above, a structure is assumed in which one of a series of actions of the user U in pairs is received as the first instruction Q1 and the other is received as the second instruction Q2. For example, an action of the user U raising one hand is received as the first instruction Q1, and an action of putting down one hand after this action is received as the second instruction Q2. In addition, an action of the user U raising the musical instrument 80 is received as the first instruction Q1, and an action of lowering the musical instrument 80 after this action is received as the second instruction Q2. Similarly, an action of the user U inhaling is received as the first instruction Q1, and an action of exhaling after this action is received as the second instruction Q2.

[0174] However, the first instruction Q1 and the second instruction Q2 do not need to be the same type of actions performed by the user U. That is, separate actions that the user U can perform independently of each other may be detected as the first instruction Q1 and the second instruction Q2. For example, the instruction receiving unit 33 may detect an operation on the operation device 14 as the first instruction Q1, and detect other actions such as an action of raising the musical instrument 80 or an action of breathing as the second instruction Q2.

[0175] (3) In each of the above-described modes, the automatic musical instrument has been exemplified as the performance device 20, but the structure of the performance device 20 is not limited to the above examples. For example, a sound source system can be adopted as the performance device 20, which includes: a sound source device that generates an audio signal of musical sound corresponding to an instruction from the audio processing system 10; and a sound playback device that plays the musical sound represented by the audio signal. The sound source device is implemented as a hardware sound source or a software sound source. In addition, the same applies to the performance device 20 of the fifth embodiment.

[0176] (4) In the fifth embodiment, each time a change in a variable made by the user U is executed (Sc3: YES), all of the analysis process Sc1 is executed, but the conditions for executing each process (Sc11 - Sc14) included in the analysis process Sc1 are not limited to the above examples. In the following description, it is assumed that the user U changes a variable by operating the operation image 71 (71a, 71b, 71c). Specifically, the user U selects the operation image 71 by operating the operation device 14 and moves the operation image 71 while maintaining the selected state. The value of the variable is changed to a value corresponding to the position of the operation image 71 at the time when the selection is released. That is, the release of the selection of the operation image 71 indicates the determination of the value of the variable.

[0177] It is assumed that the user U changes the variable α by operating the operation image 71a. During the movement of the operation image 71a in the selected state, the index calculation unit 51 updates the pronunciation index C(t) at any time by repeatedly performing the index calculation process Sc11. The display control unit 36 updates the conversion image 64 in accordance with the updated pronunciation index C(t) each time the index calculation process Sc11 is executed. That is, in parallel with the movement of the operation image 71a (change in the variable α), the index calculation process Sc11 and the update of the conversion image 64 are executed. In the state where the operation image 71a is selected, the period estimation process Sc12, the pitch estimation process Sc13, and the information generation process Sc14 are not executed. When the selection of the operation image 71a is released, the period estimation process Sc12, the pitch estimation process Sc13, and the information generation process Sc14 are executed using the pronunciation index C(t) at that time, and the confirmation screen 60 is updated in accordance with the processing results. In the above structure, in the state where the operation image 71a is selected, the period estimation process Sc12, the pitch estimation process Sc13, and the information generation process Sc14 are not executed, so the processing load of the adjustment process Sc can be reduced.

[0178] Consider a case where the user U changes the variable β by operating the operation image 71b. During the movement of the operation image 71b in the selected state, the period estimation unit 52 updates the sounding index G(t) at any time by repeatedly performing the period estimation process Sc12. The display control unit 36 updates the performance period 621 of the confirmation screen 60 each time the period estimation process Sc12 is executed. In the state where the operation image 71a is selected, the index calculation process Sc11, the pitch estimation process Sc13, and the information generation process Sc14 are not executed. When the selection of the operation image 71b is released, the pitch estimation process Sc13 and the information generation process Sc14 are executed using the sounding index G(t) at that time point, and the confirmation screen 60 is updated according to the processing results.

[0179] Consider a case where the user U changes the variable γ by operating the operation image 71c. During the movement of the operation image 71c in the selected state, the analysis process Sc1 and the update (Sc2) of the confirmation screen 60 are not executed. When the selection of the operation image 71c is released, the pitch estimation process Sc13 applied to the changed variable γ and the information generation process Sc14 applied to the pitch K(t) calculated by the pitch estimation process Sc13 are executed. According to the above configuration, the number of times of the pitch estimation process Sc13 and the information generation process Sc14 can be reduced, so the processing load of the adjustment process Sc can be reduced.

[0180] (5) The form of the conversion image 64 representing the pronunciation index C(t) is not limited to the above examples. For example, as Figure 20 illustrated, the waveform on the time axis can be displayed as the conversion image 64 on the display device 15. The amplitude at the time point t on the time axis At in the waveform of the conversion image 64 is set corresponding to the pronunciation index C(t). For example, a waveform with a large amplitude is displayed as the conversion image 64 at the time point t where the pronunciation index C(t) is large. In addition, the conversion image 64 can also be displayed overlapping the musical score 65.

[0181] (6) The process by which the index calculation unit 51 calculates the pronunciation index C(t) is not limited to the process illustrated in the fifth embodiment. For example, the index calculation unit 51 can calculate the pronunciation index E(t,n) by subtracting the amplitude spectrum of the musical sound of the performance device 20 from the amplitude spectrum of the reference signal Zr. The amplitude spectrum of the musical sound of the performance device 20 is generated, for example, by known sound source processing for generating a musical sound signal representing the musical sound specified by the performance data Db and frequency analysis such as discrete Fourier transform for the musical sound signal. The amplitude spectrum after the subtraction operation corresponds to a series of N pronunciation indexes E(t,n) corresponding to different pitches Pn. The degree of subtraction operation on the amplitude spectrum diagram of the musical sound of the performance device 20 can be adjusted corresponding to the variable α.

[0182] (7) The process by which the performance period estimation unit 52 estimates the performance period is not limited to the process illustrated in the fifth embodiment. For example, the performance period estimation unit 52 estimates the period during which the signal strength in the reference signal Zr is greater than the threshold as the performance period. The threshold can be adjusted corresponding to the variable β. In addition, the pitch estimation unit 53's process of estimating the pitch K(t) is not limited to the foregoing illustration either.

[0183] (8) For example, the audio processing system 10 can also be implemented by a server device that communicates with a terminal device such as a smartphone or a tablet terminal. For example, the terminal device has: a sound pickup device 13 that generates an audio signal Z corresponding to the performance by the user U; and a performance device 20 that plays a music piece corresponding to an instruction from the audio processing system 10. The terminal device transmits the audio signal Z generated by the sound pickup device 13, and the first instruction Q1 and the second instruction Q2 corresponding to the actions of the user U to the audio processing system 10 via a communication network. The audio processing system 10 causes the performance device 20 of the terminal device to play the performance part of the music piece corresponding to the performance position X estimated from the audio signal Z, the first instruction Q1 and the second instruction Q2 received from the terminal device. In addition, the performance analysis unit 31 can also be mounted on the terminal device. The terminal device transmits the performance position X estimated by the performance analysis unit 31 to the audio processing system 10. In the above structure, the performance analysis unit 31 can be omitted from the audio processing system 10. The audio processing system 10 of the fifth embodiment is similarly implemented by a server device. For example, the audio processing system 10 generates reference data Da through an analysis process Sc1 applied to the reference signal Zr and the performance data Db received from the terminal device, and transmits the reference data Da to the terminal device.

[0184] (9) The functions of the audio processing system 10 illustrated above are implemented, as described above, by the cooperation of a single or multiple processors constituting the control device 11 and a program stored in the storage device 12. The program related to the present invention can be provided in a form stored in a computer-readable recording medium and installed in a computer. The recording medium is, for example, a non-transitory recording medium, preferably an optical recording medium (optical disc) such as a CD-ROM, and also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. In addition, as a non-transitory recording medium, it includes any recording medium other than a transitory, propagating signal, and volatile recording media may not be excluded. Also, in a structure where a transmission device transmits a program via a communication network, the storage device that stores the program in the transmission device is equivalent to the foregoing non-transitory recording medium.

[0185] G: Appendix

[0186] According to the method exemplified above, the following structure can be grasped, for example.

[0187] One embodiment (Embodiment 1) of the present invention relates to an audio processing system having an audio analysis unit and a display control unit. The audio analysis unit performs an analysis process of analyzing an audio signal of an audio including a first sound source. The analysis process includes the following processes: a first process of determining a time series of pronunciation indexes, which are indexes of the accuracy of the audio components of the first sound source included in the audio signal; and a process of determining a time series of pitches related to the audio components of the first sound source. The display control unit causes the time series of the pronunciation indexes and the time series of the pitches to be displayed on a display device based on a common time axis. In the above embodiment, since the time series of the pronunciation indexes and the time series of the pitches are displayed based on a common time axis, it is easy for the user to confirm and correct the analysis result of the audio signal in the process of generating reference data. Specifically, the user can visually and intuitively confirm the time relationship between the pronunciation indexes and the pitches.

[0188] In a specific example of Embodiment 1 (Embodiment 2), the audio signal includes audio components of a first sound source and audio components of a second sound source. The first process includes the following processes: a process of calculating a first index corresponding to the intensity of the audio signal; a process of calculating a second index corresponding to the intensity of the audio components of the second sound source according to performance data that designates the pronunciation intensity of each note for the second sound source; and a process of subtracting the second index from the first index. The pronunciation index is an index corresponding to the result of the subtraction operation. In the above embodiment, the pronunciation index is calculated according to the result obtained by subtracting the second index calculated corresponding to the performance data from the first index corresponding to the intensity of the audio signal including the audio components of the first sound source and the audio components of the second sound source. Therefore, it is possible to calculate a pronunciation index in which the influence of the audio components of the second sound source is reduced (that is, a pronunciation index that emphasizes the audio components of the first sound source).

[0189] In a specific example of Embodiment 2 (Embodiment 3), in the process of calculating the second index, a first variable set corresponding to an instruction from the user is applied, and the second index changes corresponding to the first variable. According to the above embodiment, the user can adjust the pronunciation index to be suitable for the existing pronunciation content (for example, performance content) of the first sound source.

[0190] In the specific example (Mode 4) described in any one of Modes 1 to 3, the analysis process further includes a second process for determining the pronunciation period of the sound component of the first sound source, and the display control unit causes the pronunciation period to be displayed based on the time axis. According to the above mode, the user can visually and intuitively determine the temporal relationship between the pronunciation index, pitch, and the pronunciation period of the first sound source.

[0191] In the specific example (Mode 5) of Mode 4, in the second process, a second variable set corresponding to an instruction from the user is applied, and the pronunciation period changes accordingly with the second variable. According to the above mode, the user can adjust the pronunciation period to be suitable for the existing pronunciation content (e.g., performance content) of the first sound source.

[0192] In the specific example (Mode 6) described in any one of Modes 1 to 5, the analysis process further includes a third process for determining the time series of the pitch of the sound component of the first sound source, and the display control unit causes the pronunciation point, i.e., the time point when the pitch changes, to be displayed based on the time axis. According to the above mode, the user can visually and intuitively determine the temporal relationship between the pronunciation index, pitch, and the pronunciation point of the first sound source.

[0193] In the specific example (Mode 7) of Mode 6, for the third process, a third variable set corresponding to an instruction from the user is applied, and the number of pronunciation points changes accordingly with the third variable. According to the above mode, the user can adjust the pronunciation points to be suitable for the existing pronunciation content (e.g., performance content) of the first sound source.

[0194] An audio processing method according to one mode (Mode 8) of the present invention is implemented by a computer system, which executes an analysis process for analyzing an audio signal containing the audio of a first sound source, and includes the following processes: a first process for determining the time series of a pronunciation index, which is an index of the accuracy of the audio signal containing the audio component of the first sound source; and a process for determining the time series of the pitch related to the audio component of the first sound source, and this audio processing method causes the time series of the pronunciation index and the time series of the pitch to be displayed on a display device based on a common time axis. In addition, Figure 19 The analysis process Sc1 is an example of the "analysis process" in Mode 8. Additionally, Figure 19 The step Sc2 is an example of the "process of displaying the time series of the pronunciation index and the time series of the pitch on a display device based on a common time axis" in Mode 8.

[0195] One aspect (Aspect 9) of the present invention relates to a program that causes a computer system to function as an audio analysis unit and a display control unit. The audio analysis unit performs an analysis process of analyzing an audio signal of audio including a first sound source. The analysis process includes the following processes: a first process of determining a time series of pronunciation indexes, which are indexes of the accuracy of the audio signal including the audio component of the first sound source; and a process of determining a time series of pitches related to the audio component of the first sound source. The display control unit displays the time series of the pronunciation indexes and the time series of the pitches on a display device based on a common time axis.

[0196] Description of reference numerals

[0197] 100... Playback system, 10... Audio processing system, 11... Control device, 12... Storage device, 13... Pickup device, 14... Operation device, 15... Display device, 20... Performance device, 21... Drive mechanism, 22... Sound generation mechanism, 31... Performance analysis unit, 32... Playback control unit, 33... Instruction reception unit, 34... Editing processing unit, 35... Preparation processing unit, 36... Display control unit, 41... First recording unit, 42... Second recording unit, 43... Audio analysis unit, 51... Index calculation unit, 52... Period estimation unit, 53... Pitch estimation unit, 54... Information generation unit, 60... Confirmation screen, 61... First area, 62... Second area, 64... Conversion image, 65... Staff notation, 67... Instruction image, 68... Note image, 71(71a, 71b, 71c), 72... Operation image, 80... Musical instrument, 141... Movable part, 511... First analysis unit, 512... Second analysis unit, 513... First arithmetic unit, 514... Second arithmetic unit, 621... Performance period, 622... Non-performance period.

Claims

1. An audio processing system having an audio analysis unit and a display control unit, The audio analysis unit performs an analysis process of analyzing audio signals of an audio for a first sound source and an audio for a second sound source, and the analysis process includes the following processes: a process of determining a time series of a first pitch related to an audio component of the first sound source and a time series of a second pitch related to an audio component of the second sound source; and a first process for determining a time series of pronunciation metrics, which is an index of the accuracy of the audio components corresponding to each pitch in the time series of the first pitch of the first sound source contained in the audio signal, wherein the display control unit causes the time series of the pronunciation metrics and the time series of the first pitch to be displayed on a display device based on a common time axis, the first process includes the following processes: a process of calculating a first metric corresponding to the intensity of the audio components corresponding to each pitch in the time series of the first pitch and the time series of the second pitch of the audio signal; a process of calculating a second metric corresponding to the intensity of the audio components corresponding to each pitch in the time series of the second pitch of the second sound source according to performance data specifying the pronunciation intensity of each note for the second sound source; and a process of subtracting the second metric from the first metric on a common time axis, wherein the pronunciation metric is a metric corresponding to the result of the subtraction process.

2. The audio processing system according to claim 1, wherein in the process of calculating the second metric, a first variable set corresponding to an instruction from a user is applied, and the second metric changes corresponding to the first variable.

3. The audio processing system according to claim 1 or 2, wherein the analysis process further includes a second process for determining the pronunciation period of the audio components of the first sound source, and the display control unit causes the pronunciation period to be displayed based on the time axis.

4. The audio processing system according to claim 3, wherein in the second process, a second variable set corresponding to an instruction from a user is applied, and the pronunciation period changes corresponding to the second variable.

5. The audio processing system according to claim 1 or 2, wherein the display control unit causes the pronunciation points, i.e., the time points at which each pitch in the time series of the first pitch changes, to be displayed based on the time axis.

6. The audio processing system according to claim 5, wherein when determining the time series of the first pitch related to the audio components of the first sound source, a third variable set corresponding to an instruction from a user is applied, and the number of pronunciation points changes corresponding to the third variable.

7. An audio processing method implemented by a computer system, Perform parsing processing that parses audio signals of an audio device storing a first sound source and an audio device storing a second sound source, including the following processing: a process of determining a time series of a first pitch related to the sound component of the first sound source and a time series of a second pitch related to the sound component of the second sound source; and a first process for determining a time series of pronunciation metrics, which is an index of the accuracy of the audio components corresponding to each pitch in the time series of the first pitch of the first sound source contained in the audio signal, executing a display control process that causes the time series of the pronunciation metrics and the time series of the first pitch to be displayed on a display device based on a common time axis, the first process includes the following processes: a process of calculating a first metric corresponding to the intensity of the audio components corresponding to each pitch in the time series of the first pitch and the time series of the second pitch of the audio signal; A process of calculating a second index corresponding to the intensity of an acoustic component corresponding to each pitch in the time series of the second pitch of the second sound source according to performance data specifying the pronunciation intensity of each note for the second sound source; And A process of subtracting the second index from the first index on a common time axis, The pronunciation index is an index corresponding to the result of the subtraction process.

8. The acoustic processing method according to claim 7, wherein, In the process of calculating the second index, a first variable set corresponding to an instruction from a user is applied, The second index changes corresponding to the first variable.

9. The acoustic processing method according to claim 7 or 8, wherein, The analysis process further includes a second process of determining the pronunciation period of the acoustic component where the first sound source exists, The display control process causes the pronunciation period to be displayed based on the time axis.

10. The acoustic processing method according to claim 9, wherein, In the second process, a second variable set corresponding to an instruction from a user is applied, The pronunciation period changes corresponding to the second variable.

11. The acoustic processing method according to claim 7 or 8, wherein, The display control process causes the pronunciation point, which is the time point when each pitch in the time series of the first pitch changes, to be displayed based on the time axis.

12. The acoustic processing method according to claim 11, wherein, When determining the time series of the first pitch related to the acoustic component of the first sound source, a third variable set corresponding to an instruction from a user is applied, The number of pronunciation points changes corresponding to the third variable.

13. A recording medium storing a program that causes a computer system to function as an acoustic analysis unit and a display control unit, The audio analysis unit performs an analysis process of analyzing audio signals of an audio for a first sound source and an audio for a second sound source, and the analysis process includes the following processes: a process of determining a time series of a first pitch related to an audio component of the first sound source and a time series of a second pitch related to an audio component of the second sound source; And a first process that determines a time series of a pronunciation index, which is an index of the accuracy of an acoustic component corresponding to each pitch in the time series of the first pitch of the first sound source included in the acoustic signal, The display control unit causes the time series of the pronunciation index and the time series of the first pitch to be displayed on a display device based on a common time axis, The first process includes the following processes: A process of calculating a first index corresponding to the intensity of an acoustic component corresponding to each pitch in the time series of the first pitch and the second pitch of the acoustic signal; A process of calculating a second index corresponding to the intensity of an acoustic component corresponding to each pitch in the time series of the second pitch of the second sound source according to performance data specifying the pronunciation intensity of each note for the second sound source; And A process of subtracting the second index from the first index on a common time axis, The pronunciation index is an index corresponding to the result of the subtraction process.

Citation Information

Patent Citations

  • Information providing device

    JP2016099512A

  • Automatic playing system and automatic playing method

    JP2017207615A

  • Control method and control device

    WO2018016638A1