Audio analysis systems, electronic musical instruments and audio analysis methods

CN116762124BActive Publication Date: 2026-09-18YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202280011529.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-05
Filing Date
2022-01-21
Publication Date
2026-09-18
Estimated Expiration
2042-01-21

Smart Images

  • Figure CN116762124B_ABST
    Figure CN116762124B_ABST
Patent Text Reader

Abstract

The acoustic analysis system includes: an indication receiving unit that receives an indication of a target timbre; an acquisition unit that acquires a first acoustic signal containing multiple acoustic components corresponding to different timbres; and an acoustic analysis unit that selects one or more reference signals from among multiple reference signals representing different performance tones. The reference rhythmic pattern representing the time variation of the signal intensity of one or more reference signals is similar to the analysis rhythmic pattern, which represents the time variation of the intensity of the acoustic component corresponding to the target timbre among the multiple acoustic components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique for analyzing audio signals. Background Technology

[0002] Techniques for analyzing the characteristics of acoustic signals representing the performance notes of a musical piece have been proposed in the past. For example, Patent Document 1 discloses a technique for automatically generating musical pieces using machine learning technology.

[0003] Patent Document 1: International Publication No. 2020 / 166094 Summary of the Invention

[0004] For example, in situations such as composing a piece of music or practicing playing an instrument, users sometimes desire a pattern similar to a pattern repeated with a specific timbre in a particular piece of music. However, searching for a suitable pattern is time-consuming and requires musical expertise, making it practically difficult. In view of the above, one objective of this invention is to reduce the effort required of users to explore patterns for playing with a specific timbre.

[0005] To address the above-mentioned issues, one aspect of the present invention relates to an audio analysis system comprising: an instruction receiving unit that receives an instruction for a target timbre; an acquisition unit that acquires a first audio signal including multiple audio components corresponding to different timbres; and an audio analysis unit that selects one or more reference signals from a plurality of reference signals representing different performance tones, wherein a reference rhythmic pattern representing the time variation of the signal intensity of the one or more reference signals is similar to an analysis rhythmic pattern representing the time variation of the intensity of the audio component corresponding to the target timbre among the plurality of audio components.

[0006] To address the above-mentioned problems, one aspect of the present invention relates to an electronic musical instrument comprising: an instruction receiving unit that receives an instruction for a target timbre; an acquisition unit that acquires a first acoustic signal including multiple acoustic components corresponding to different timbres; an acoustic analysis unit that selects one or more reference signals from a plurality of reference signals representing different performance notes; a performance device that receives performance from a user; and a playback control unit that causes a playback system to play the performance notes represented by the selected one or more reference signals and musical notes corresponding to the performance received by the performance device, wherein a reference rhythm pattern representing the time variation of the signal intensity of the one or more reference signals is similar to an analysis rhythm pattern representing the time variation of the intensity of the acoustic component corresponding to the target timbre among the plurality of acoustic components.

[0007] To address the above-mentioned issues, one aspect of the present invention relates to an acoustic analysis method that receives an indication of a target timbre, obtains a first acoustic signal including multiple acoustic components corresponding to different timbres, selects one or more reference signals from a plurality of reference signals representing different performance notes, and the reference rhythmic pattern representing the time variation of the signal intensity of the one or more reference signals is similar to the analytical rhythmic pattern representing the time variation of the intensity of the acoustic component corresponding to the target timbre among the plurality of acoustic components. Attached Figure Description

[0008] Figure 1 This is a block diagram illustrating the structure of an electronic musical instrument according to an embodiment.

[0009] Figure 2 This is a block diagram illustrating the functional structure of an electronic musical instrument.

[0010] Figure 3 This is a block diagram illustrating the specific structure of the audio resolution section.

[0011] Figure 4 This is an explanatory diagram of the separation section.

[0012] Figure 5 This is an explanatory diagram related to the analysis of rhythm patterns.

[0013] Figure 6 This is a flowchart illustrating the specific process of generating a parsed rhythm pattern.

[0014] Figure 7 This is an explanatory diagram of the operation of the selection section.

[0015] Figure 8 This is a schematic diagram illustrating the analysis of an image.

[0016] Figure 9 This is a schematic diagram illustrating the analysis of an image.

[0017] Figure 10 This is a flowchart illustrating the specific process of audio analysis and processing.

[0018] Figure 11 This is a block diagram illustrating the structure of an information processing system.

[0019] Figure 12 This is a block diagram illustrating the functional structure of an information processing system.

[0020] Figure 13 It is a flowchart of the process by which the control device of an information processing system creates a trained model through machine learning.

[0021] Figure 14 This is an illustration of the generation of the basis matrix by the information processing system.

[0022] Figure 15 This is an explanatory diagram of the generation of a reference rhythm pattern by an information processing system.

[0023] Figure 16 This is a block diagram illustrating the specific structure of the audio analysis unit according to the second embodiment.

[0024] Figure 17 This is a flowchart illustrating the specific process of the audio analysis processing in the second embodiment.

[0025] Figure 18 This is an explanatory diagram of the selection section in the third embodiment.

[0026] Figure 19 This is a block diagram illustrating the structure of the performance system according to the fourth embodiment.

[0027] Figure 20 This is an explanatory diagram of the selection section in the fifth embodiment.

[0028] Figure 21 It is a block diagram illustrating the specific structure of a trained model.

[0029] Figure 22 This is a flowchart illustrating the specific process of the audio analysis processing in the fifth embodiment.

[0030] Figure 23 This is a block diagram illustrating the functional structure of the information processing system according to the fifth embodiment.

[0031] Figure 24 This is a block diagram illustrating the structure of the performance system according to the sixth embodiment. Detailed Implementation

[0032] A: Implementation Method 1

[0033] Figure 1 This is a block diagram illustrating the structure of an electronic musical instrument 10 according to an embodiment of the present invention. The electronic musical instrument 10 is an audio analysis system that performs the functions of playing musical notes corresponding to the user's playing and analyzing the audio signal S1 representing the playing notes of a specific piece of music.

[0034] The electronic musical instrument 10 includes a control device 11, a storage device 12, a communication device 13, an operation device 14, a playing device 15, a sound source device 16, a sound playback device 17, and a display device 19. Furthermore, the electronic musical instrument 10 can be implemented not only as a single device, but also as multiple devices constructed separately from each other.

[0035] The control device 11 consists of one or more processors that control various elements of the electronic musical instrument 10. The control device 11 may be, for example, a CPU (Central Processing Unit), SPU (Sound Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).

[0036] It consists of one or more processors, such as Integrated Circuits.

[0037] Storage device 12 is a single or multiple memory devices that store the program executed by control device 11 and various data used by control device 11. Storage device 12 may be composed of known recording media such as magnetic recording media or semiconductor recording media, or a combination of multiple recording media. In addition, a removable recording medium that can be detached from electronic musical instrument 10, or a recording medium that can be written to or read from by control device 11 via communication network 90 such as the Internet (e.g., cloud storage) may also be used as storage device 12.

[0038] Storage device 12 stores the audio signal S1, which is the object of analysis for the electronic musical instrument 10. The audio signal S1 is a signal containing multiple audio components, including musical tones produced by different instruments. Furthermore, the audio signal S1 may include audio components of the voice produced by a singer. The audio signal S1 is stored in storage device 12, for example, as a music file transmitted from a music transmission device (not shown) to the electronic musical instrument 10. The audio signal S1 is an example of a "first audio signal." Alternatively, the audio signal S1 may be read from a recording medium such as an optical disc and supplied to the electronic musical instrument 10 by a playback device.

[0039] The communication device 13 communicates with other devices via the communication network 90. ​​For example, the communication device 13 communicates with the information processing system 40, which will be described later. Furthermore, the presence or absence of a wireless band in the communication line between the communication device 13 and the communication network 90 is arbitrary. Additionally, the communication device 13, which is separate from the electronic musical instrument 10, can be an information terminal such as a smartphone or tablet.

[0040] The operating device 14 is an input device that receives instructions from the user. The operating device 14 may be, for example, multiple operating components operated by the user, or a touch panel that detects the user's touch. The user can operate the operating device 14 to indicate a desired instrument (hereinafter referred to as the "target instrument") from a variety of instruments to the electronic instrument 10. The timbre of a musical tone differs for each type of instrument; therefore, the user's instruction to the instrument is an example of "timbre instruction." Furthermore, the target instrument is an example of "target timbre."

[0041] The playing device 15 is an input device for receiving performances from the user. Specifically, the playing device 15 is a keyboard with multiple keys 151 arranged to correspond to different pitches. The user plays music by sequentially operating the desired keys 151. That is, the electronic musical instrument 10 is an electronic keyboard instrument.

[0042] The sound source device 16 generates an acoustic signal corresponding to the performance performed on the playing device 15. Specifically, the sound source device 16 generates an acoustic signal representing the timbre of a key 151 among the multiple keys 151 of the playing device 15 pressed by the user. Alternatively, the function of the sound source device 16 can be implemented by the control device 11 by executing a program stored in the storage device 12. That is, the sound source device 16 can also be omitted.

[0043] The sound playback device 17 plays the musical tones represented by the audio signals generated by the sound source device 16. The sound playback device 17 is, for example, a speaker or headphones. In this embodiment, the sound source device 16 and the sound playback device 17 function as a playback system 18 that plays musical tones corresponding to the user's performance. The display device 19 displays images under the control of the control device 11. The display device 19 is, for example, a liquid crystal display panel.

[0044] Figure 2 This is a block diagram illustrating the functional structure of the electronic musical instrument 10. The control device 11 of the electronic musical instrument 10 performs multiple functions (acquisition unit 111, instruction receiving unit 112, sound analysis unit 113, prompting unit 114, and playback control unit 115) by executing a program stored in the storage device 12. Furthermore, the functions of the control device 11 can be implemented by multiple devices configured separately from each other, or some or all of the functions of the control device 11 can be implemented by dedicated circuitry.

[0045] The acquisition unit 111 acquires the audio signal S1. Specifically, the acquisition unit 111 sequentially reads each sample of the audio signal S1 from the storage device 12. In addition, the acquisition unit 111 can acquire the audio signal S1 from an external device that can communicate with the electronic musical instrument 10.

[0046] The instruction receiving unit 112 receives instructions from the user regarding the operating device 14. Specifically, the instruction receiving unit 112 receives instructions from the user regarding the target musical instrument and generates instruction data D representing the target musical instrument.

[0047] Figure 3 This is a block diagram illustrating the functional structure of the audio resolution unit 113. The audio resolution unit 113 includes a separation unit 1131, a resolution unit 1132, and a selection unit 1133.

[0048] Figure 4 This is an explanatory diagram of the separation unit 1131. The separation unit 1131 generates an audio signal S2 by separating the sound source of the audio signal S1. Specifically, the separation unit 1131 separates the audio signal S2, which represents the audio component corresponding to the target instrument indicated by the user, from the multiple audio components of the audio signal S1 that correspond to different musical instruments. That is, the audio signal S2 is a signal that relatively emphasizes the audio component of the target instrument among the multiple audio components of the audio signal S1 relative to the audio components other than the target instrument. The audio signal S2 is an example of a "second audio signal".

[0049] In the generation of the audio signal S2 by the separation unit 1131, a trained model M is used. Specifically, the separation unit 1131 inputs the combination of the audio signal S1 and the indicator data D, i.e., the input data X, into the trained model M, and outputs the audio signal S2 from the trained model M. The trained model M is a model that has learned the relationship between the combination of the audio signal S1 and the indicator data D and the audio signal S2 through machine learning.

[0050] The trained model M is, for example, composed of a deep neural network (DNN). The trained model M can utilize any form of neural network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Furthermore, the trained model M can be constructed from a combination of various deep neural networks. Additionally, supplementary elements such as long short-term memory (LSTM) can be incorporated into the trained model M.

[0051] The trained model M is implemented by having the control device 11 execute a program that generates an audio signal S2 based on the combination of the audio signal S1 and the instruction data D, i.e., the input data X, and by combining this program with multiple variables (e.g., weighting values ​​and biases) applied to the operation. The program for implementing the trained model M and the multiple variables are stored in the storage device 12. The values ​​of the multiple variables for the trained model M are preset through machine learning.

[0052] Figure 3 The analysis unit 1132 generates the analytical rhythm pattern Y by analyzing the audio signal S2. Figure 5 This is an explanatory diagram related to the analysis of rhythmic pattern Y. Figure 5 The notation f refers to frequency, and the notation t refers to time. The analysis unit 1132 generates an analytical rhythmic pattern Y for each of the multiple periods (hereinafter referred to as unit periods) T into which the audio signal S2 is divided on the time axis. The unit period T is, for example, a time length equivalent to a predetermined number of measures in the music (e.g., 1 measure, 4 measures, or 8 measures).

[0053] The rhythmic pattern Y is composed of M coefficient sequences y1 to yM corresponding to different timbres. The coefficient sequence ym corresponding to the m-th timbre (m = 1 to M) among the M timbres is a non-negative numerical sequence representing the time variation of the signal intensity (e.g., amplitude or power) related to the acoustic component of that timbre in the acoustic signal S2. Furthermore, the timbre varies, for example, for each type of instrument and each pitch of the musical note. Therefore, the coefficient sequence ym can also be described as the time variation of the intensity related to the acoustic component corresponding to the combination of instrument and pitch.

[0054] The analysis unit 1132 generates an analytical rhythmic pattern Y based on the acoustic signal S2 by utilizing the known non-negative matrix factorization (NMF) of the basis matrix B. The basis matrix B is a non-negative matrix containing M frequency characteristics b1 to bM corresponding to the timbre of musical tones produced by different instruments. The frequency characteristic bm corresponding to the acoustic component of the m-th instrument is a series of intensities (base vector) of that acoustic component on the frequency axis. Specifically, the frequency characteristic bm is, for example, an amplitude spectrum or a power spectrum. The basis matrix B, pre-generated through machine learning, is stored in the storage device 12.

[0055] As understood from the above explanation, the analytical rhythmic pattern Y is a non-negative coefficient matrix (activation matrix) corresponding to the basis matrix B. That is, each coefficient sequence ym of the analytical rhythmic pattern Y is a time variation of the weighted values ​​(activation degrees) of the frequency characteristics bm within the basis matrix B. Each coefficient sequence ym can also be referred to as the rhythmic pattern of the audio signal S2 related to the m-th timbre.

[0056] Figure 6 This is a flowchart illustrating the specific process of the analysis unit 1132 generating the analytical rhythm pattern Y. It is executed for each unit period T of the audio signal S2. Figure 6 The processing.

[0057] The analysis unit 1132 generates an observation matrix O(Sa1) for the unit period T of the acoustic signal S2. The observation matrix O is as follows: Figure 5 As shown, it is a non-negative matrix representing the time series of the frequency characteristics of the acoustic signal S2. Specifically, the time series (spectral graph) of the amplitude spectrum or power spectrum within a unit period T is generated as the observation matrix O.

[0058] The analysis unit 1132 calculates the analytical rhythm pattern Y based on the observation matrix O by utilizing the nonnegative matrix decomposition of the basis matrix B stored in the storage device 12 (Sa2). Specifically, the analysis unit 1132 calculates the analytical rhythm pattern Y in a manner that makes the product BY of the basis matrix B and the analytical rhythm pattern Y similar to (ideally consistent with) the observation matrix O.

[0059] Figure 7 yes Figure 3 The diagram illustrates the operation of the selection unit 1133. The storage device 12 stores N reference signals R1 to RN representing different performance notes and N reference rhythm patterns Z1 to ZN corresponding to the different reference signals Rn (n = 1 to N). Each reference rhythm pattern Zn consists of M coefficient sequences z1 to zM corresponding to different timbres of musical notes produced by a specific instrument. For example, the coefficient sequence zm of the reference rhythm pattern Zn is the m-th rhythm pattern of the n-th instrument.

[0060] N reference signals R1 to RN each represent a different part of a musical piece. Specifically, each reference signal Rn represents a part of the musical piece suitable for repeated performance (i.e., a loop). In this embodiment, a reference rhythm pattern Zn is generated based on each of the N reference signals R1 to RN.

[0061] The selection unit 1133 compares each of the N reference rhythm patterns Z1 to ZN with the analytical rhythm pattern Y. Specifically, the selection unit 1133 calculates the similarity Qn by comparing each reference rhythm pattern Zn with the analytical rhythm pattern Y. In the following explanation, the correlation coefficient, an indicator of the correlation between the reference rhythm pattern Zn and the analytical rhythm pattern Y, is used as an example of the similarity Qn. Therefore, the more correlated the reference rhythm pattern Zn and the analytical rhythm pattern Y are, the larger the similarity Qn becomes. That is, the similarity Qn is an indicator of the degree of similarity between the reference rhythm pattern Zn and the analytical rhythm pattern Y.

[0062] The selection unit 1133 selects one or more reference signals Rn from N reference signals R1 to RN based on the calculated similarity Qn, and outputs the selected reference signal Rn to the prompting unit 114 and the playback control unit 115. Specifically, the selection unit 1133 selects multiple reference signals Rn with a similarity Qn greater than a predetermined threshold, or a predetermined number of reference signals Rn that are ranked higher in descending order of similarity Qn.

[0063] As understood from the above explanation, the sound analysis unit 113 (selection unit 1133) selects multiple reference signals Rn from N reference signals R1 to RN whose reference rhythm pattern Zn is similar to the analysis rhythm pattern Y. Furthermore, the selection unit 1133 can select a predetermined number of reference signals Rn for each unit period T of the sound signal S1, or it can select a predetermined number of reference signals Rn in descending order of the average similarity over the entire range of unit periods T of the sound signal S1.

[0064] Figure 2 The prompting unit 114 displays the results of the analysis performed by the audio analysis unit 113 on the display device 19. Specifically, the prompting unit 114 prompts the user with multiple reference signals Rn selected by the selection unit 1133. The prompting unit 114 of the first embodiment will... Figure 8 or Figure 9 The resolved image is displayed on the display device 19. The resolved image is an image that displays the reference signal Rn in a sorted manner.

[0065] Figure 8 The analytical image is a representation of each reference signal Rn, which corresponds to a reference rhythmic pattern Zn similar to the analytical rhythmic pattern Y of the target instrument, "Drum". Similarly, Figure 9 The analytical image is an image representing each reference signal Rn, which corresponds to a reference rhythm pattern Zn similar to the analytical rhythm pattern Y of the target instrument, "Guitar".

[0066] Users can refer to Figure 8 or Figure 9 The analytical image visually identifies the reference signal Rn among multiple reference signals Rn that corresponds to the reference rhythmic pattern Zn, which is similar to the analytical rhythmic pattern Y of the target instrument. For example, the reference... Figure 8 From the analyzed image, the user can identify the reference signal Rn corresponding to the reference rhythm pattern Zn, which is most similar to the analyzed rhythm pattern Y of the target instrument "Drum". Furthermore, Figure 8 and Figure 9The strings such as "DrumPattern01" are the label names of the reference signal Rn, and the numbers such as "1" on the left side of this string indicate the order corresponding to the similarity Qn. Therefore, in Figure 8 and Figure 9 In the diagram, "DrumPattern01" and "GuitarRiff01" are the reference signals Rn with the highest similarity to Qn.

[0067] Figure 2 The playback control unit 115 controls the playback of musical sounds produced by the playback system 18. Specifically, the playback control unit 115 instructs the playback system 18 (specifically, the sound source device 16) to produce sound in accordance with the operation of the performance device 15. In addition, the playback control unit 115 causes the playback system 18 to play the performance sound represented by one reference signal Rn selected by the user from the analytical image among a plurality of reference signals Rn selected by the selection unit 1133.

[0068] Figure 10 This is a flowchart illustrating the specific process of the processing (audio analysis processing) performed by the control device 11. For example, audio analysis processing is performed based on an instruction from the user for the electronic musical instrument 10.

[0069] If audio processing begins, the acquisition unit 111 acquires audio signal S1 (Sb1). The instruction receiving unit 112 waits for the user to specify the target instrument (Sb2: NO). If the instruction receiving unit 112 receives the specification of the target instrument (Sb2: YES), the separation unit 1131 separates audio signal S2 (Sb3) from audio signal S1.

[0070] The analysis unit 1132 generates an observation matrix O for each of the multiple unit periods T that divide the audio signal S2 on the time axis (see reference). Figure 5 (Sb4). The analysis unit 1132 calculates the analytical rhythm pattern Y based on each observation matrix O by utilizing the non-negative matrix decomposition of the basis matrix B stored in the storage device 12 (Sb5).

[0071] The selection unit 1133 calculates the similarity Qn between the reference rhythm pattern Zn and the analytical rhythm pattern Y associated with each of the N reference signals R1 to RN (Sb6). The selection unit 1133 selects multiple reference signals Rn from the N reference signals R1 to RN whose reference rhythm pattern Zn is similar to the analytical rhythm pattern Y (Sb7).

[0072] The prompting unit 114 displays the tag names of each reference signal Rn selected by the selection unit 1133 in descending order of similarity Qn on the display device 19 (Sb8). The playback control unit 115 waits for the user to select a reference signal Rn (Sb9: NO). If the user selects any of the multiple reference signals Rn displayed on the display device 19 (Sb9: YES), the playback control unit 115 plays the performance tone represented by the reference signal Rn by supplying the reference signal Rn to the playback system 18 (Sb10).

[0073] Figure 1 The information processing system 40 generates a trained model M, which is used when the separation unit 1131 generates the audio signal S2. Figure 11 This is a block diagram illustrating the structure of an information processing system 40. The information processing system 40 includes a control device 41, a storage device 42, and a communication device 43. Furthermore, the information processing system 40 can be implemented not only as a single device but also as multiple devices configured separately from each other.

[0074] The control device 41 consists of one or more processors that control the various elements of the information processing system 40. The control device 41 consists of one or more processors such as CPU, SPU, DSP, FPGA, or ASIC. The communication device 43 communicates with the electronic musical instrument 10 via the communication network 90.

[0075] Storage device 42 is a single or multiple memory devices that store the program executed by control device 41 and various data used by control device 41. Storage device 42 may be composed of, for example, a known recording medium such as magnetic recording medium or semiconductor recording medium, or a combination of multiple recording media. Alternatively, a removable recording medium that can be attached to or detached from information processing system 40, or a recording medium that can be written to or read from by control device 41 via communication network 90 (e.g., cloud storage) may be used as storage device 42.

[0076] Figure 12 This is a block diagram illustrating the functional structure of the information processing system 40. The control device 41 functions as multiple elements (training data acquisition unit 51 and learning processing unit 52) ​​for creating a trained model M through machine learning by executing programs stored in the storage device 42.

[0077] The learning processing unit 52 creates a trained model M using teacher-led machine learning with multiple training data TDs. The training data acquisition unit 51 acquires multiple training data TDs. Specifically, the training data acquisition unit 51 acquires multiple training data TDs stored in the storage device 42 from the storage device 42.

[0078] Multiple training data TDs each as follows Figure 12 As shown, the training signal is composed of a combination of training input data Xt and training audio signal S2t. The training input data Xt is data formed by combining the training audio signal S1t and training indication data Dt. The training audio signal S1t is a known signal containing multiple acoustic components corresponding to different musical instruments. The training audio signal S1t is an example of a "first training audio signal".

[0079] The training instruction data Dt is data that specifies any of a variety of instruments. The training instruction data Dt is an example of "training instruction data". The training acoustic signal S2t is a known signal representing the acoustic component of the training acoustic signal S1t that corresponds to the instrument represented by the training instruction data Dt. The training acoustic signal S2t is an example of "second training acoustic signal".

[0080] Figure 13 This is a flowchart illustrating the specific process of the control device 41 creating a trained model M through machine learning (hereinafter referred to as the learning process) Sc. The learning process Sc also refers to the method for generating the trained model M.

[0081] If learning processing Sc begins, the training data acquisition unit 51 acquires any one of the multiple training data TDs stored in the storage device 42 (hereinafter referred to as "selected training data TD") (Sc1). The learning processing unit 52, as... Figure 12 As shown, the input data Xt of the selected training data TD is input into the initial or temporary model (hereinafter referred to as "temporary model") M0 (Sc2), and the audio signal S2 (Sc3) output by the temporary model M0 for the input is obtained.

[0082] The learning processing unit 52 calculates a loss function representing the error between the acoustic signal S2 generated by the temporary model M0 and the acoustic signal S2t of the selected training data TD (Sc4). The learning processing unit 52 updates multiple variables of the temporary model M0 to reduce (ideally minimize) the loss function (Sc5). For updating the multiple variables corresponding to the loss function, for example, error backpropagation is used.

[0083] The learning processing unit 52 determines whether the specified termination condition is met (Sc6). The termination condition is, for example, that the loss function is less than a specified threshold, or that the change in the loss function is less than a specified threshold. If the termination condition is not met (Sc6: NO), the training data acquisition unit 51 selects the previously unselected training data TD as the new selected training data TD (Sc1). That is, the learning processing unit 52 repeatedly updates multiple variables of the temporary model M0 (Sc1 to Sc5) until the termination condition is met. If the termination condition is met (Sc6: YES), the learning processing unit 52 terminates the update of the specified multiple variables of the temporary model M0 (Sc1 to Sc5). The temporary model M0 at the time point when the termination condition is met is determined as the trained model M. That is, the multiple variables of the trained model M are determined as the values ​​at the end time point of the learning processing Sc.

[0084] As understood from the above explanation, the trained model M, based on the latent relationship between the input data Xt and the audio signal S2t of multiple selected training data TD, outputs a statistically reasonable audio signal S2 for unknown input data X. That is, the trained model M is a model that, as described above, learns the relationship between the training input data Xt and the training audio signal S2t through machine learning.

[0085] The information processing system 40 sends the trained model M, created through the above process, from the communication device 43 to the electronic musical instrument 10 (Sc7). Specifically, the learning processing unit 52 sends multiple variables of the trained model M from the communication device 43 to the electronic musical instrument 10. The control device 11 of the electronic musical instrument 10 saves the trained model M received from the information processing system 40 to the storage device 12. Specifically, multiple variables of the trained model M are specified and stored in the storage device 12.

[0086] in addition, Figure 1 The information processing system 40 generates a base matrix B and a reference rhythm pattern Zn used by the analysis unit 1132 and the selection unit 1133. Figure 14 This is an explanatory diagram of the generation of the basis matrix B by the information processing system 40. Figure 15 This is an explanatory diagram of the generation of the reference rhythm pattern Zn by the information processing system 40. The basis matrix B and the reference rhythm pattern Zn are generated, for example, through the following process.

[0087] Control device 41 such as Figure 14As shown, N reference signals R1 to RN stored in storage device 42 are read out. Control device 41 generates observation matrix On based on each reference signal Rn. Observation matrix On, like the aforementioned observation matrix O, is a non-negative matrix representing the time series (spectral graph) of the frequency characteristics of the reference signal Rn.

[0088] Next, the control device 41 generates an observation matrix OT by connecting the N observation matrices O1 to ON on the time axis. The control device 41 then generates a basis matrix B based on the observation matrix OT through nonnegative matrix decomposition. As understood from the above explanation, the basis matrix B contains frequency characteristics bm corresponding to all types of timbre contained in the N reference signals R1 to RN.

[0089] Next, the control device 41, as Figure 15 As shown, the reference rhythm pattern Zn is calculated based on each observation matrix On by utilizing the nonnegative matrix decomposition of the generated basis matrix B. Specifically, the control device 41 calculates the reference rhythm pattern Zn in a manner that makes the product BZn of the basis matrix B and the reference rhythm pattern Zn approximately (ideally, identical) to the observation matrix On. The information processing system 40 sends the basis matrix B and N reference rhythm patterns Z1 to ZN generated through the above process to the electronic musical instrument 10 from the communication device 43. The control device 11 of the electronic musical instrument 10 stores the basis matrix B and the N reference rhythm patterns Z1 to ZN received from the information processing system 40 in the storage device 12.

[0090] As explained above, in the first embodiment, a reference rhythm pattern Zn is selected from among a plurality of reference signals Rn that is similar to the analytical rhythm pattern Y of the instrument (target instrument) indicated by the user. This reduces the time required for the user to explore the desired rhythm pattern for their specified instrument, thus improving the efficiency of music creation or performance practice.

[0091] Furthermore, in the first embodiment, multiple reference signals Rn are appropriately selected in accordance with the similarity Qn between the reference rhythm patterns Zn of each of the N reference signals R1 to RN and the analytical rhythm pattern Y of the instrument indicated by the user.

[0092] Furthermore, in the first embodiment, it is possible to determine the order in which the reference rhythmic pattern Zn is similar to the analytical rhythmic pattern Y of the target instrument for multiple reference signals Rn. Thus, the user can, for example, create or practice playing music in accordance with this order.

[0093] Furthermore, in the first embodiment, the user can refer to Figure 8 or Figure 9The analytical image is used to visually identify the reference signal Rn among multiple reference signals Rn that corresponds to the reference rhythm pattern Zn, which is similar to the analytical rhythm pattern Y of the target instrument.

[0094] B: Implementation Method 2

[0095] Next, the second embodiment will be described. Furthermore, in the embodiments illustrated below, elements that have the same function and structure as those in the first embodiment will be omitted with the same reference numerals as those in the first embodiment.

[0096] Figure 16 This is a block diagram illustrating the specific structure of the sound analysis unit 113 according to the second embodiment. The sound analysis unit 113 of the second embodiment is a structure that removes the separation unit 1131 from the same elements (separation unit 1131, analysis unit 1132, and selection unit 1133) as in the first embodiment. Specifically, in the first embodiment, an sound signal S2 emphasizing the sound components of the target instrument is generated by the separation unit 1131, which is separate from the analysis unit 1132. In contrast, in the second embodiment, the sound components of the target instrument are emphasized during the process of the analysis unit 1132 generating the analytical rhythm pattern Y.

[0097] Figure 17 This is a flowchart illustrating the specific process of the processing (audio analysis processing) performed by the control device 11 in the second embodiment.

[0098] If audio analysis processing begins, the acquisition unit 111 acquires the audio signal S1 (Sd1). The analysis unit 1132 generates an observation matrix O (Sd2) for each of the multiple unit periods T that divide the audio signal S1 on the time axis. In the first embodiment, the observation matrix O is a non-negative matrix corresponding to the audio signal S2 after the sound source is separated. In contrast, in the second embodiment, the observation matrix O is a non-negative matrix representing the time series of the frequency characteristics of the audio signal S1. Specifically, the observation matrix O is generated as the time series (spectral graph) of the amplitude spectrum or power spectrum of the unit period T.

[0099] Next, the analysis unit 1132 calculates the analytical rhythmic pattern Y based on the observation matrix O using the nonnegative matrix decomposition of the basis matrix B (Sd3). The basis matrix B is labeled with instrument names. Specifically, each of the M frequency characteristics b1 to bM constituting the basis matrix B is associated with an instrument name. That is, the series of intensities of the sound component of which instrument the m-th frequency characteristic among the M frequency characteristics b1 to bM represents is known.

[0100] The instruction receiving unit 112 awaits the user's designation of the target instrument (Sd4: NO). If the instruction receiving unit 112 receives the designation of the target instrument (Sd4: YES), the analysis unit 1132 sets each element of one or more coefficient sequences ym corresponding to instruments other than the target instrument among the M coefficient sequences y1 to yM constituting the analytical rhythm pattern Y to 0 (Sd5). Thus, the analytical rhythm pattern Y becomes a non-negative coefficient matrix in which each element of the coefficient sequence ym corresponding to instruments other than the target instrument is 0.

[0101] If the above processing is performed, the control device 11 performs the processing steps Sb6 to Sb10 in the same manner as in the first embodiment. Therefore, the second embodiment also achieves the same effect as the first embodiment.

[0102] C: Third Implementation Method

[0103] Figure 18 This is an explanatory diagram of the selection unit 1133 in the third embodiment. The selection unit 1133 generates a compressed analytical rhythm pattern Y' by compressing the analytical rhythm pattern Y on the time axis. Specifically, the selection unit 1133 calculates the average or sum of multiple elements of the coefficient sequence ym for each of the M coefficient sequences y1 to yM constituting the analytical rhythm pattern Y, thereby generating the compressed analytical rhythm pattern Y'. Therefore, the compressed analytical rhythm pattern Y' is composed of M coefficients y'1 to y'M corresponding to different timbres. That is, the coefficient y'm is the average or sum of multiple elements of the coefficient sequence ym. The coefficient y'm corresponding to the m-th timbre among the M timbres is a non-negative value representing the intensity related to the acoustic component of that timbre.

[0104] Similarly, the selection unit 1133 generates a compressed reference rhythm pattern Z'n based on each of the N reference rhythm patterns Z1 to ZN. The N compressed reference rhythm patterns Z'1 to Z'N are stored in the storage device 12. The compressed reference rhythm pattern Z'n is generated by compressing the reference rhythm pattern Zn along the time axis. Specifically, the selection unit 1133 calculates the average or sum of each element of the coefficient sequence zm for each of the M coefficient sequences z1 to zM constituting the reference rhythm pattern Zn, thereby generating the compressed reference rhythm pattern Z'n. Therefore, the compressed reference rhythm pattern Z'n consists of M coefficients z'1 to z'M corresponding to different timbres of a musical tone produced by a specific instrument. That is, the coefficient z'm is the average or sum of multiple elements of the coefficient sequence zm. The coefficient z'm corresponding to the m-th timbre among the M timbres is a non-negative value representing the intensity related to the acoustic component of that timbre.

[0105] The selection unit 1133 calculates the similarity Qn by comparing each of the N compressed reference rhythm patterns Z'1 to Z'N with the compressed analytical rhythm pattern Y'. As understood from the above description, the selection unit 1133 in the aforementioned method calculates the similarity Qn by comparing the reference rhythm pattern Zn with the analytical rhythm pattern Y. In contrast, the selection unit 1133 in the third embodiment calculates the similarity Qn by comparing the compressed reference rhythm pattern Z'n (compressed in the time axis direction) with the compressed analytical rhythm pattern Y' (compressed in the time axis direction) with the analytical rhythm pattern Y.

[0106] In the third embodiment described above, the same effects as in the first embodiment are achieved. Furthermore, the structures of the first and second embodiments are also applied to the third embodiment.

[0107] D: Fourth Implementation Method

[0108] Figure 19 This is a block diagram illustrating the structure of the performance system 100 according to the fourth embodiment. The performance system 100 includes an electronic musical instrument 10 and an information device 80. The information device 80 is, for example, a device such as a smartphone or tablet terminal. The information device 80 is connected to the electronic musical instrument 10, for example, via a wired or wireless means.

[0109] The information device 80 is implemented by a computer system having a control device 81, a storage device 82, a display device 83, and an operation device 84. The control device 81 consists of one or more processors that control the various elements of the information device 80. For example, the control device 81 consists of one or more processors such as a CPU, SPU, DSP, FPGA, or ASIC.

[0110] Storage device 82 is a single or multiple memory devices that store programs executed by control device 81 and various data used by control device 81. Storage device 82 may be composed of known recording media such as magnetic recording media or semiconductor recording media, or a combination of multiple recording media. In addition, a removable recording medium that can be detached from information device 80, or a recording medium that can be written to or read from by control device 81 via, for example, communication network 90 (e.g., cloud storage) may also be used as storage device 82.

[0111] The display device 83 displays images based on the control given by the control device 81. The operating device 84 is an input device that receives instructions from the user. Specifically, the operating device 84 receives instructions from the user regarding the target musical instrument.

[0112] The control device 81 executes the program stored in the storage device 82 to achieve the same functions as the control device 11 of the electronic musical instrument 10 in the first embodiment (acquisition unit 111, instruction receiving unit 112, sound analysis unit 113, prompting unit 114, and playback control unit 115). The reference signal Rn, the basis matrix B, and the trained model M used by the sound analysis unit 113 are stored in the storage device 82. In addition, the sound signal S1 is also stored in the storage device 82. On the other hand, in the electronic musical instrument 10 of the fourth embodiment, the functions exemplified in the first embodiment (acquisition unit 111, instruction receiving unit 112, sound analysis unit 113, prompting unit 114, and playback control unit 115) can be omitted. Furthermore, the division of functions between the electronic musical instrument 10 and the information device 80 can be appropriately changed relative to the above examples. For example, some of the functions of the acquisition unit 111, instruction receiving unit 112, sound analysis unit 113, prompting unit 114, and playback control unit 115 can be carried in the information device 80, and other functions can be carried in the electronic musical instrument 10. That is, as a whole, the performance system 100 can realize the multiple functions shown above.

[0113] The acquisition unit 111 acquires the audio signal S1 stored in the storage device 82. The instruction receiving unit 112 receives instructions from the user regarding the operation device 84. The audio analysis unit 113, similar to the first embodiment, determines a plurality of reference signals Rn based on the audio signal S1 and the instruction data D. The prompting unit 114 displays the plurality of reference signals Rn selected by the audio analysis unit 113 on the display device 83. The playback control unit 115 supplies one reference signal Rn selected by the user from the plurality of reference signals Rn to the electronic musical instrument 10, causing the playback system 18 to play the performance sound. Furthermore, the prompting unit 114 and the playback control unit 115 can be mounted on the electronic musical instrument 10. For example, the prompting unit 114 can display the analyzed image on the display device 19, similar to the first embodiment.

[0114] As understood from the above description, the same effects as in the first embodiment are achieved in the fourth embodiment. Furthermore, the structures of the second or third embodiments are also applied to the fourth embodiment.

[0115] In the fourth embodiment, for example, the trained model M constructed by the information processing system 40 is transmitted to the information device 80, and the trained model M is stored in the storage device 82. In the above structure, an authentication processing unit (not shown) that authenticates the legitimacy of the user of the information device 80 (that is, a pre-registered legitimate user) can be mounted on the information processing system 40. When the legitimacy of the user is authenticated by the authentication processing unit, the trained model M is automatically (i.e., without instruction from the user) transmitted to the information device 80.

[0116] E: Fifth Implementation

[0117] Figure 20 This is an explanatory diagram of the selection unit 1133. In the fifth embodiment, the selection unit 1133 is input with input data Xa, which is a combination of the parsed rhythm pattern Y and the reference rhythm pattern Zn. The selection unit 1133 outputs a similarity Qn corresponding to the input data Xa.

[0118] For the generation of similarity Qn by the selection unit 1133 in the fifth embodiment, a trained model Ma is used. Specifically, the selection unit 1133 inputs input data Xa into the trained model Ma and outputs similarity Qn from the trained model Ma. The trained model Ma is a model that has learned the relationship between the combination of the parsed rhythm pattern Y and the reference rhythm pattern Zn and the similarity Qn through machine learning.

[0119] The trained model Ma can be composed of any form of deep neural network, such as a recurrent neural network or a convolutional neural network. For example, the trained model Ma can be composed of a combination of a recurrent neural network and a convolutional neural network.

[0120] The trained model Ma is implemented by a combination of a program that causes the control device 11 to perform a calculation to generate a similarity Qn based on the input data Xa, and multiple variables (e.g., weighting values ​​and biases) applied to this calculation. The program for implementing the trained model Ma and the multiple variables are stored in the storage device 12. The values ​​of the multiple variables for the trained model Ma are pre-set through machine learning.

[0121] Figure 21 This is a block diagram illustrating the specific structure of the trained model Ma. The trained model Ma consists of a first model Ma1 and a second model Ma2. Input data Xa is input into the first model Ma1.

[0122] Model Ma1 generates feature data Xaf based on input data Xa. Model Ma1 is a trained model that learns (trains) the relationship between input data Xa and feature data Xaf. Feature data Xaf represents data corresponding to the differences between the parsed rhythmic pattern Y and the reference rhythmic pattern Zn. Model Ma1 can be constructed, for example, from a convolutional neural network.

[0123] The second model, Ma2, generates a similarity score Qn based on the feature data Xaf. Ma2 is a trained model that learns (trains) the relationship between the feature data Xaf and the similarity score Qn. For example, Ma2 is composed of a recurrent neural network. Furthermore, Ma2 can incorporate additional elements such as Long Short-Term Memory (LSTM) or Gated Recurrent Units (GRUs).

[0124] Figure 22 This is a flowchart illustrating the specific process of the processing (audio analysis processing) performed by the control device 11 in the fifth embodiment. In the fifth embodiment, Figure 10 In the illustrated first embodiment, step Sb6 is replaced by steps Se1 and Se2. The contents of the processes Sb1 to Sb5 and Sb7 to Sb10 are the same as in the first embodiment.

[0125] The selection unit 1133 combines the reference rhythm pattern Zn and the analytical rhythm pattern Y associated with each of the N reference signals R1 to RN to generate input data Xa1 to XaN. The selection unit 1133 inputs each input data Xan (n = 1 to N) into the trained model Ma (Se1) and outputs the similarity Qn (Se2) corresponding to each input data Xa1 to XaN. In the fifth embodiment, the same effect as in the first embodiment is achieved.

[0126] The trained model Ma shown above was generated by the information processing system 40. Figure 23 This is a block diagram illustrating the functional structure of the information processing system 40 related to the generation of the trained model Ma. The control device 41 operates as multiple elements (training data acquisition unit 51a and learning processing unit 52a) for creating the trained model Ma through machine learning by executing programs stored in the storage device 42.

[0127] The learning processing unit 52a creates a trained model Ma using teacher-led machine learning with multiple training data TDa. The training data acquisition unit 51a acquires multiple training data TDa. Specifically, the training data acquisition unit 51a acquires multiple training data TDa stored in the storage device 42.

[0128] Multiple training datasets TDa each as follows Figure 23As shown, the training input data Xat and the training similarity data Qnt are combined. The training input data Xat is obtained by combining the training analytical rhythm pattern Yt and the training reference rhythm pattern Znt. The training analytical rhythm pattern Yt is a known coefficient matrix consisting of multiple coefficient sequences corresponding to different timbres. The reference rhythm pattern Znt is an example of a "training reference rhythm pattern", and the analytical rhythm pattern Yt is an example of a "training analytical rhythm pattern".

[0129] The training reference rhythm pattern Znt is a known coefficient matrix consisting of multiple coefficient sequences corresponding to different timbres of musical notes produced by a specific instrument. The training similarity Qnt is a value pre-associated with the training input data Xat. Specifically, the similarity Qnt between the parsed rhythm pattern Yt of the input data Xat and the training reference rhythm pattern Znt is associated with the training input data Xat. The similarity Qnt is an example of "training similarity".

[0130] The learning processing unit 52a inputs the input data Xat of each of the multiple training data TTas into a temporary model, updating multiple variables of the temporary model to reduce (ideally minimize) the loss function between the similarity Q output by the model and the similarity Qnt of the training data TTas. That is, the trained model Ma learns the relationship between the input data Xat and the similarity Qnt. Therefore, based on the potential relationship between the input data Xat and the similarity Q of the multiple training input data Xats, the trained model Ma outputs a statistically reasonable similarity Qn for the unknown input data Xan.

[0131] F: Implementation Method 6

[0132] Figure 24 This is a block diagram illustrating the structure of the performance system 100 according to the sixth embodiment. The performance system 100, like the fourth embodiment, includes an electronic musical instrument 10 and an information device 80. The structures of the electronic musical instrument 10 and the information device 80 are the same as in the fourth embodiment.

[0133] The information processing system 40 stores multiple trained models Ma corresponding to different music genres. In the learning process used to create the trained models Ma corresponding to each music genre, training data TDa, including input data Xat specific to the music genre, is used. That is, multiple training data TDa are prepared separately for each music genre, and trained models Ma are created through separate learning processes for each genre. "Music genre" refers to a classification (category) of musical music from a musical perspective. For example, rock, pop, jazz, trance, or hip-hop are representative examples of music genres.

[0134] Information device 80 selectively acquires any one of the multiple trained models Ma stored in information processing system 40 via communication network 200. Specifically, information device 80 acquires one trained model Ma from the multiple trained models 60 that corresponds to a specific music genre from information processing system 40. For example, information device 80 acquires the trained model Ma corresponding to the music genre indicated by the genre tag contained in audio signal S1 (music file) from information processing system 40. Genre tag is a tag indicating a specific music genre in music files such as MP3 files or AAC (Advanced Audio Coding) files. Alternatively, information device 80 estimates the music genre of the music by parsing audio signal S1. Any known technique is used for estimating the music genre. Information device 80 acquires the trained model Ma corresponding to that music genre from information processing system 40. The trained model Ma obtained from the information processing system 40 is stored in the storage device 82 and used for processing the similarity Qn output by the selection unit 1133.

[0135] As understood from the above description, this variation achieves the same effects as the first to fifth embodiments. Furthermore, in the sixth embodiment, a trained model Ma is created for each music genre, thus offering the advantage of obtaining a high-precision similarity Qn compared to using a common trained model Ma that is independent of music genres.

[0136] Furthermore, the above description illustrates a structure where the information processing system 40 stores multiple trained models Ma corresponding to different music genres. However, the information device 80 can also retrieve and store multiple trained models Ma from the information processing system 40. That is, multiple trained models Ma are stored in the storage device 82 of the information device 80. The sound analysis unit 113 selectively uses any one of the multiple trained models Ma to calculate the similarity Qn.

[0137] G: Variation Example

[0138] The embodiments of the present invention have been described above, but the present invention is not limited to the embodiments described above, and various modifications can be made. Hereinafter, specific modifications that can be applied to the foregoing methods are illustrated. Two or more methods arbitrarily selected from the following examples can be appropriately combined within a non-contradictory scope.

[0139] (1) In the aforementioned methods, the audio signal S2 corresponding to the instrument indicated by the user is separated from the multiple audio components corresponding to different musical instruments in the audio signal S1, but the audio component of the singing voice among the multiple audio components can also be separated.

[0140] (2) In the aforementioned methods, the correlation between the reference rhythm pattern Zn and the analytical rhythm pattern Y is used as an example to represent the similarity Qn. However, the selection unit 1133 can also calculate the distance between the reference rhythm pattern Zn and the analytical rhythm pattern Y as the similarity Qn. In the above structure, the more similar the reference rhythm pattern Zn and the analytical rhythm pattern Y are to each other, the smaller the similarity Qn becomes. Furthermore, the distance between the reference rhythm pattern Zn and the analytical rhythm pattern Y can be arbitrarily used as a distance index such as cosine distance or KL divergence.

[0141] (3) In the aforementioned methods, the selection unit 1133 selects multiple reference signals Rn that are similar to the reference rhythm pattern Zn and the analytical rhythm pattern Y among the N reference signals R1 to RN, but the selection unit 1133 may also select 1 reference signal Rn.

[0142] (4) In the aforementioned methods, the reference signal Rn is a portion that typically contains the playing tone of a single instrument, but it can also be a portion that contains the playing tone of two or more different instruments.

[0143] (5) In the second embodiment, the elements of one or more coefficient sequences ym that correspond to instruments other than the target instrument among the M coefficient sequences y1 to yM constituting the analytical rhythm pattern Y are set to 0, but it is also possible not to set each element to 0.

[0144] (6) In the aforementioned methods, the information processing system 40 creates a trained model M, but the functions of the information processing system 40 (training data acquisition unit 51 and learning processing unit 52) ​​can also be implemented in the information device 80 of the fourth embodiment. Furthermore, in the aforementioned methods, the information processing system 40 generates a basis matrix B and a reference rhythm pattern Zn, but the functions of the information processing system 40 that generates the basis matrix B and the reference rhythm pattern Zn can also be implemented in the information device 80 of the fourth embodiment.

[0145] (7) In the foregoing methods, a deep neural network is exemplified as the trained model M, but the trained model M is not limited to a deep neural network. For example, a statistical inference model such as HMM (Hidden Markov Model) or SVM (Support Vector Machine) can be used as the trained model M. Furthermore, in the foregoing methods, a teacher-led machine learning method using multiple training data TDs is exemplified as the learning process Sc, but a trained model M can also be created using unteacher-free machine learning that does not require training data TDs, or reinforcement learning that maximizes rewards. For example, unteacher-free machine learning utilizes well-known clustering methods.

[0146] (8) The functions illustrated in the aforementioned embodiments (acquisition unit 111, instruction receiving unit 112, sound analysis unit 113, prompting unit 114, playback control unit 115) are implemented, as described above, through the coordinated operation of one or more processors constituting the control device (11, 81) and the program stored in the storage device (12, 82). The program described above can be provided and installed in a computer in the form of a computer-readable recording medium. The recording medium is, for example, a non-transitory recording medium, preferably an optical recording medium (optical disc) such as a CD-ROM, and also includes any known form of recording medium such as semiconductor recording media or magnetic recording media. Furthermore, as a non-transitory recording medium, it includes any recording medium other than a transient propagating signal, and may also include volatile recording media. Additionally, in a structure where a transmission device transmits a program via a communication network, the recording medium storing the program in that transmission device is equivalent to the aforementioned non-transitory recording medium.

[0147] (9) In the aforementioned methods, the similarity Qn is calculated by comparing the analytical rhythm pattern Y and the reference rhythm pattern Zn, but the method for calculating the similarity Qn is not limited to this example. For example, the selection unit 1133 can determine the similarity Qn by retrieving the similarity Qn corresponding to the combination of the feature quantity extracted from the acoustic signal S2 and the feature quantity extracted from the reference signal Rn (hereinafter referred to as "feature quantity data") from a table. In this table, the similarity Qn is registered for each of the multiple feature quantity data. Furthermore, the feature quantity of the acoustic signal S2 and the reference signal Rn is, for example, time series data representing the frequency characteristics of the played tone. For example, time series data representing the frequency characteristics of MFCC (Mel-Frequency Cepstrum Coefficient), MSLS (Mel-Scale Log Spectrum), or CQT (Constant-Q Transform) are shown as such feature quantities.

[0148] (10) In the aforementioned fifth embodiment, an example is shown where the trained model Ma for generating similarity Qn based on input data Xa is composed of a deep neural network. The type of trained model Ma is not limited to the example above. For example, statistical estimation models such as HMM (Hidden Markov Model) or SVM (Support Vector Machine) can be used as trained model Ma. Specific examples of trained model Ma are shown below.

[0149] (10-1)HMM

[0150] Hidden Markov Models (HMMs) are statistical inference models that connect multiple latent states corresponding to different similarity values ​​Qn. The feature data is a combination of features extracted from the audio signal S2 and features extracted from the reference signal Rn, taken as time-series inputs to the HMM. For example, the feature data might be data within an interval corresponding to one measure of a musical piece.

[0151] The selection unit 1133 inputs the time series of feature data into the trained model Ma, which is constructed by the HMM illustrated above. Based on the condition that multiple feature data have been observed, the selection unit 1133 uses the HMM to estimate the time series of the maximum likelihood similarity Qn. For example, dynamic programming algorithms such as the Viterbi algorithm are used to estimate the similarity Qn.

[0152] Hidden Markov Models (HMMs) are created using teacher-led machine learning that incorporates multiple training datasets, including similarity scores (Qn). In machine learning, the transition probabilities and output probabilities of each latent state are repeatedly updated by outputting the time series of the maximum likelihood similarity scores (Qn) for multiple feature datasets.

[0153] (10-2)SVM

[0154] An SVM is prepared for each combination of selecting two values ​​from multiple possible values ​​of similarity Qn. For the SVM corresponding to the combination of two values, a hyperplane in a multidimensional space is created through machine learning. The hyperplane is the boundary surface that separates the space of the feature data distribution corresponding to one of the two values ​​from the space of the feature data distribution corresponding to the other value. The trained model involved in this variation consists of multiple SVMs corresponding to different combinations of values ​​(multi-class SVM).

[0155] The selection unit 1133 inputs feature data to each of the multiple SVMs. For each combination, the selection unit chooses one of two values ​​related to the presence of feature data in either of the two spaces separated by the hyperplane. The same value selection is performed for each of the multiple SVMs corresponding to different combinations. The selection unit 1133 selects the value selected most frequently by the multiple SVMs and determines this value as the similarity Qn.

[0156] As understood from the above examples, the selection unit 1133 involved in this variation functions as an element that inputs feature data into a trained model and outputs a similarity Qn from the trained model. The similarity Qn is an index of the degree of similarity between the feature data extracted from the audio signal S2 and the feature data extracted from the reference signal Rn.

[0157] (11) In the aforementioned fifth embodiment, a teacher-led machine learning approach using multiple training data TDas is illustrated, but a trained model Ma can also be created through reinforcement learning that maximizes the reward. For example, if the similarity Q output by the temporary model Ma0 for the input data Xat of each training data TDa is consistent with the similarity Qnt of that training data TDa, the learning processing unit 52a sets the reward function to "+1"; if they are inconsistent, the reward function is set to "-1". The learning processing unit 52a repeatedly updates multiple variables of the temporary model Ma0 in a manner that maximizes the sum of the reward functions set for multiple training data TDas, thereby creating a trained model Ma.

[0158] (12) In the first embodiment, an audio signal S2 corresponding to the input data X is generated using a trained model M that has learned the relationship between input data X, including audio signal S1 and indicator data D, and audio signal S2. However, the structure and method for generating audio signal S2 based on input data X are not limited to the examples above. For example, a reference table in which different input data X are associated with audio signals S2 can be used for generating audio signals S2 by the separation unit 1131. The reference table is a data table that records the correspondence between input data X and audio signals S2, for example, stored in the storage device 12. The separation unit 1131 retrieves the input data X corresponding to the combination of audio signal S1 and indicator data D from the reference table, and obtains the audio signal S2 associated with the input data X from the multiple audio signals S2.

[0159] (13) In the fifth and sixth embodiments, a similarity Qn corresponding to the input data Xa is generated using a trained model Ma that learns the relationship between input data Xa, including the parsing rhythm pattern Y and the reference rhythm pattern Zn, and the similarity Qn. However, the structure and method for generating the similarity Qn based on the input data Xa are not limited to the examples above. For example, a reference table in which each of the multiple input data Xa is associated with a similarity Qn can also be used to generate the similarity Qn by the selection unit 1133. The reference table is a data table that records the correspondence between input data Xa and similarity Qn, for example, stored in the storage device 12. The selection unit 1133 retrieves the input data Xa corresponding to the combination of the parsing rhythm pattern Y and the reference rhythm pattern Zn from the reference table, and obtains the similarity Qn associated with the input data Xa among the multiple similarity Qn based on the reference table.

[0160] (14) Among the aforementioned methods, the method in which the instruction receiving unit 112 receives instructions from the user for the target instrument is exemplified, but the instruction receiving unit 112 may also receive instructions from outside the user for the target instrument. For example, it is also conceivable that the instruction receiving unit 112 receives instructions from an external device for the target instrument, or that the instruction receiving unit 112 receives instructions generated by the internal processing of the electronic instrument 10.

[0161] (15) Among the foregoing embodiments, an electronic keyboard instrument is exemplified as electronic musical instrument 10, but the electronic musical instrument is not limited to the above examples. For example, the present invention is also applicable to electronic musical instruments such as electronic string instruments (e.g., electronic guitar or electronic violin), electronic drums, and electronic tube instruments (e.g., electronic saxophone, electronic clarinet or electronic flute).

[0162] F: Appendix

[0163] Based on the examples above, for instance, understand the following structure.

[0164] One aspect (Aspect 1) of the present invention relates to an acoustic analysis system comprising: an instruction receiving unit that receives an instruction for a target timbre; an acquisition unit that acquires a first acoustic signal including multiple acoustic components corresponding to different timbres; and an acoustic analysis unit that selects one or more reference signals from a plurality of reference signals representing different performance tones, wherein a reference rhythmic pattern representing the time variation of the signal intensity of the one or more reference signals is similar to an analytical rhythmic pattern representing the time variation of the intensity of the acoustic component corresponding to the target timbre among the plurality of acoustic components. Based on the above structure, one or more reference signals from a plurality of reference signals whose reference rhythmic pattern is similar to the analytical rhythmic pattern of the target timbre are selected. This reduces the time required for the user to explore the desired rhythmic pattern for their specified timbre, thus improving the efficiency of, for example, music creation or performance practice.

[0165] In a specific example of Method 1 (Method 2), the sound analysis unit includes: a separation unit that separates a second sound signal from the first sound signal into a sound component that corresponds to the target timbre; an analysis unit that calculates the analysis rhythm pattern of the second sound signal; and a selection unit that selects one or more reference signals from the plurality of reference signals whose reference rhythm pattern is similar to the analysis rhythm pattern calculated by the analysis unit.

[0166] In a specific example of Method 2 (Method 3), the separation unit outputs the second audio signal by inputting the first audio signal and the indication data representing the target timbre to the trained model. The trained model has learned the relationship between the combination of the first training audio signal, which includes multiple audio components corresponding to different timbres, and the training indication data representing the timbre, and the second training audio signal. The second training audio signal represents the audio component among the multiple audio components of the first training audio signal that corresponds to the timbre represented by the training indication data.

[0167] According to a specific example of Method 2 or Method 3 (Method 4), the analysis unit calculates a coefficient matrix as the analysis rhythm pattern based on the second audio signal by utilizing non-negative matrix decomposition of the basis matrix representing multiple frequency characteristics corresponding to different timbres.

[0168] In a specific example of Method 2 (Method 5), the analysis unit calculates the coefficient matrix based on the second audio signal by using non-negative matrix decomposition of the basis matrix representing the frequency characteristics of the sound corresponding to different timbres, and sets each element of the coefficient sequence corresponding to the timbres other than the target timbres in the multiple coefficient sequences contained in the calculated coefficient matrix to 0, thereby generating the analysis rhythm pattern.

[0169] According to a specific example of any one of methods 2 to 5 (method 6), the selection unit calculates the similarity between the reference rhythm pattern and the analytical rhythm pattern for each of the plurality of reference signals, and selects one or more reference signals from the plurality of reference signals based on the similarity. In the above methods, one or more reference signals are appropriately selected corresponding to the similarity between the reference rhythm pattern of each of the plurality of reference signals and the analytical rhythm pattern of the target timbre.

[0170] In a specific example of Method 6 (Method 7), the selection unit outputs the similarity by inputting input data including the reference rhythm pattern and the analytical rhythm pattern into the trained model. The trained model has learned the relationship between the training input data including the training reference rhythm pattern and the training analytical rhythm pattern and the training similarity between the training reference rhythm pattern and the training analytical rhythm pattern.

[0171] In a specific example of Method 7 (Method 8), the selection unit outputs the similarity by inputting the input data into the trained model corresponding to a specific music genre among multiple trained models corresponding to different music genres.

[0172] In a specific example of Method 8 (Method 9), the trained model among the plurality of trained models corresponding to a music genre is created by machine learning using multiple training data corresponding to that music genre.

[0173] In a specific example of any of methods 7 to 9 (method 10), the trained model comprises: a first model consisting of a convolutional neural network that generates feature data based on the input data; and a second model consisting of a recurrent neural network that generates similarity based on the feature data.

[0174] In a specific example (method 11) of any of methods 2 to 5, the reference rhythm pattern includes multiple coefficient sequences corresponding to different timbres, the analytical rhythm pattern includes multiple coefficient sequences corresponding to different timbres, the selection unit generates a compressed reference rhythm pattern by averaging or summing multiple elements of each of the multiple coefficient sequences of the reference rhythm pattern, and generates a compressed analytical rhythm pattern by averaging or summing multiple elements of each of the multiple coefficient sequences of the analytical rhythm pattern, calculates the similarity between the compressed reference rhythm pattern and the compressed analytical rhythm pattern, and selects one or more reference signals from the multiple reference signals based on the similarity.

[0175] In a specific example (method 12) of any of methods 6 to 11, the more than one reference signal is two or more reference signals, and a prompting unit is provided that displays information related to the two or more reference signals in an order corresponding to the similarity on the display device. In the above methods, the user can grasp the order in which the reference rhythmic patterns among the multiple reference signals are similar to the analytical rhythmic patterns of the target timbre. Therefore, the user can, for example, perform music creation or performance practice in accordance with this order.

[0176] In a specific example (method 13) of any of methods 2 to 12, for each of the multiple unit periods into which the second audio signal is divided on the time axis, the analysis unit calculates the analysis rhythm pattern, and the selection unit selects one or more reference signals.

[0177] In a specific example (method 14) of any of methods 1 to 11, a prompting unit is also provided, which prompts the user to the one or more reference signals selected by the audio analysis unit. According to the above method, the user can visually grasp the one or more reference signals selected by the audio analysis unit.

[0178] One aspect (aspect 15) of the present invention relates to an electronic musical instrument comprising: an instruction receiving unit that receives an instruction for a target timbre; an acquisition unit that acquires a first acoustic signal including a plurality of acoustic components corresponding to different timbres; an acoustic analysis unit that selects one or more reference signals from a plurality of reference signals representing different performance notes; a performance device that receives performance by a user; and a playback control unit that causes a playback system to play the performance notes represented by the selected one or more reference signals and musical notes corresponding to the performance received by the performance device, wherein a reference rhythm pattern representing the time variation of the signal intensity of the one or more reference signals is similar to an analysis rhythm pattern representing the time variation of the intensity of the acoustic component corresponding to the target timbre among the plurality of acoustic components.

[0179] One aspect (aspect 16) of the present invention relates to an acoustic analysis method that receives an indication of a target timbre, obtains a first acoustic signal including multiple acoustic components corresponding to different timbres, selects one or more reference signals among multiple reference signals representing different performance notes, and the reference rhythmic pattern representing the time variation of the signal intensity of the one or more reference signals is similar to the analysis rhythmic pattern, which represents the time variation of the intensity of the acoustic component among the multiple acoustic components corresponding to the target timbre.

[0180] One aspect of the present invention (aspect 17) involves a program that causes a computer to function as a unit that receives an instruction for a target timbre; an acquisition unit that acquires a first acoustic signal that includes multiple acoustic components corresponding to different timbres; and an acoustic analysis unit that selects one or more reference signals from a plurality of reference signals representing different performance tones, wherein a reference rhythmic pattern representing the time variation of the signal intensity of the one or more reference signals is similar to an analytical rhythmic pattern representing the time variation of the intensity of the acoustic component corresponding to the target timbre among the plurality of acoustic components.

[0181] Explanation of the label

[0182] 10…Electronic musical instrument; 11, 81…Control device; 12, 82…Storage device; 13…Communication device; 14, 84…Operating device; 15…Performing device; 16…Sound source device; 17…Sound playback device; 18…Playback system; 19, 83…Display device; 40…Information processing system; 90…Communication network; 100…Performing system; 111…Acquisition unit; 112…Instruction receiving unit; 113…Sound analysis unit; 114…Prompt unit; 115… Playback control unit, 1131… separation unit, 1132… analysis unit, 1133… selection unit, D… instruction data, Dt… instruction data for training, M… trained model, O… observation matrix, similarity… Qn (Q1~QN), Rn (R1~RN)… reference signal, S1, S2… audio signal, S1t, S2t… audio signal for training, T… unit period, Y… analysis rhythm pattern, Zn (Z1~ZN)… reference rhythm pattern.

Claims

1. An audio analysis system, comprising: The receiving unit receives instructions for the target timbre. The acquisition unit acquires a first audio signal containing multiple audio components corresponding to different timbres; The separation unit, by inputting the first acoustic signal and the indication data representing the target timbre into the first trained model, outputs a second acoustic signal representing the acoustic component corresponding to the target timbre. The first trained model has learned the relationship between the combination of the first training acoustic signal, which includes multiple acoustic components corresponding to different timbres, and the training indication data representing the timbre, and the second training acoustic signal. The second training acoustic signal represents the acoustic component among the multiple acoustic components of the first training acoustic signal that corresponds to the timbre represented by the training indication data. The analysis unit calculates the analytical rhythm pattern, which represents the temporal variation of the intensity of the acoustic component corresponding to the target timbre in the second acoustic signal; The selection unit compares the reference rhythm pattern, which represents the temporal variation of the signal intensity of multiple reference signals representing different performance notes, with the analytical rhythm pattern, and selects one or more reference signals from among the multiple reference signals based on the comparison result.

2. The audio analysis system according to claim 1, wherein, The analytical unit calculates a coefficient matrix as the analytical rhythm pattern based on the second audio signal by using nonnegative matrix decomposition of the basis matrix representing multiple frequency characteristics corresponding to different timbres.

3. The audio analysis system according to claim 1, wherein, The selection unit calculates the similarity between the reference rhythm pattern and the parsed rhythm pattern for each of the plurality of reference signals. From the plurality of reference signals, one or more reference signals are selected based on the similarity.

4. The audio analysis system according to claim 3, wherein, The selection unit outputs the similarity by inputting input data, including the reference rhythm pattern and the analytical rhythm pattern, into the second trained model. The second trained model learns the relationship between the training input data, including the training reference rhythm pattern and the training analytical rhythm pattern, and the training similarity between the training reference rhythm pattern and the training analytical rhythm pattern.

5. The audio analysis system according to claim 4, wherein, The selection unit outputs the similarity by inputting the input data into the second trained model corresponding to a specific music genre from among multiple second trained models corresponding to different music genres.

6. The audio analysis system according to claim 5, wherein, The second trained model among the plurality of second trained models corresponding to a music genre is created by using machine learning with multiple training data corresponding to that music genre.

7. The audio analysis system according to any one of claims 4 to 6, wherein, The second trained model includes: The first model, composed of a convolutional neural network, generates feature data based on the input data; and The second model, which consists of a recurrent neural network, generates similarity based on the feature data.

8. The audio analysis system according to claim 1, wherein, The reference rhythm pattern contains multiple coefficient sequences corresponding to different timbres. The analytical rhythm pattern contains multiple coefficient sequences corresponding to different timbres. The selection unit generates a compressed reference rhythm pattern by averaging or summing multiple elements of each of the multiple coefficient sequences of the reference rhythm pattern. For each of the multiple coefficient sequences of the analytical rhythm pattern, a compressed analytical rhythm pattern is generated by averaging or summing multiple elements of the coefficient sequence. The similarity between the compressed reference rhythmic pattern and the compressed parsed rhythmic pattern is calculated. From the plurality of reference signals, one or more reference signals are selected based on the similarity.

9. The audio analysis system according to any one of claims 3 to 6, wherein, The more than one reference signal refers to more than two reference signals. It also has a prompting section that displays information related to the two or more reference signals in an order corresponding to the similarity on the display device.

10. The audio analysis system according to claim 1, wherein, For each of the multiple unit periods into which the second audio signal is divided on the time axis, The analytical unit calculates the analytical rhythm pattern. The selection unit selects one or more reference signals.

11. The audio analysis system according to claim 1, wherein, It also has a prompting unit that prompts the user to the one or more reference signals selected by the selection unit.

12. An electronic musical instrument, which has: The receiving unit receives instructions for the target timbre. The acquisition unit acquires a first audio signal containing multiple audio components corresponding to different timbres; The separation unit, by inputting the first acoustic signal and the indication data representing the target timbre into the first trained model, outputs a second acoustic signal representing the acoustic component corresponding to the target timbre. The first trained model has learned the relationship between the combination of the first training acoustic signal, which includes multiple acoustic components corresponding to different timbres, and the training indication data representing the timbre, and the second training acoustic signal. The second training acoustic signal represents the acoustic component among the multiple acoustic components of the first training acoustic signal that corresponds to the timbre represented by the training indication data. The analysis unit calculates the analytical rhythm pattern, which represents the temporal variation of the intensity of the acoustic component corresponding to the target timbre in the second acoustic signal; The selection unit compares the reference rhythm pattern of the time variation of the signal intensity of each of the multiple reference signals representing different performance notes with the analytical rhythm pattern, and selects one or more reference signals from the multiple reference signals based on the comparison result. A performance device that receives performances from the user; as well as The playback control unit enables the playback system to play the performance tones represented by the selected one or more reference signals and the musical tones corresponding to the performance received by the performance device.

13. A method for sound resolution, implemented through a computer system. Accept the target tone instruction. Obtain the first audio signal, which contains multiple audio components corresponding to different timbres. The first trained model outputs a second audio signal representing the audio component corresponding to the target timbre by inputting the first audio signal and indicator data representing the target timbre into the first trained model. The first trained model has learned the relationship between the combination of the first training audio signal (which includes multiple audio components corresponding to different timbres) and the training indicator data representing the timbre, and the second training audio signal. The second training audio signal represents the audio component among the multiple audio components of the first training audio signal that corresponds to the timbre represented by the training indicator data. Calculate the analytical rhythmic pattern, which represents the temporal variation in intensity of the acoustic component corresponding to the target timbre in the second acoustic signal. The reference rhythmic pattern, which represents the temporal variation of the signal intensity of multiple reference signals representing different performance notes, is compared with the analytical rhythmic pattern. Based on the result of this comparison, one or more reference signals are selected from the multiple reference signals.

Citation Information

Patent Citations

  • Information processing device, information processing method, and information processing program

    WO2020166094A1

  • Unit rhythm extraction method from musical acoustic signal, musical piece structure estimation method using this method, and replacing method of percussion instrument pattern in musical acoustic signal

    JP2010054802A

  • Acoustic analyzer

    JP2015079110A