Auditory feedback system, auditory feedback method, and program
Patent Information
- Application Number
- PCT/JP2025/006229
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2026-09-03
Smart Images

Figure JP2025006229_03092026_PF_FP_ABST
Abstract
Description
Auditory feedback system, auditory feedback method, and program
[0001] This disclosure relates to technologies for sensory feedback, particularly auditory feedback.
[0002] Humans enrich their lives and achieve smooth communication through the manipulation of sound, such as speech, singing, and playing musical instruments. In this sound generation process, sounds predicted by the brain are produced through the articulatory organs and movements of fingers, arms, and other parts of the body. These sounds can be monitored in real time through sensory feedback such as hearing, somatosensory, tactile, and visual. This sensory feedback enables the acquisition of speech and the learning of musical instruments, and allows for the correction of generation errors in real time.
[0003] Auditory feedback plays a particularly important role in sound generation. It is known that when the sound predicted by the brain is not accurately fed back to the auditory system, sound generation becomes difficult, or responses that cancel out the discrepancy between prediction and reality are observed. For example, when a person speaks while listening to a recording of their own speech, delayed by approximately 200 milliseconds, through earphones, they experience difficulty speaking (see Non-Patent Literature 1). This is called delayed auditory feedback (DAF). DAF has also been confirmed to cause changes in musical performance (Non-Patent Literature 2).
[0004] Evaluating an individual's auditory feedback ability in sound production is important in singing and instrumental performance training. Recent research has revealed that the influence of DAF (Dynamic Auditory Feedback) on speech is greater than that on instrumental performance.
[0005] Lee, BS, "Effects of delayed speech feedback.", J. Acoust. Soc. Am., 22, 824-826 (1950).Gates, A., Bradshaw, JL & Nettleton, NC, "Effect of different delayed auditory feedback intervals on a music performance task.", Perception & Psychophysics 15, 21-25 (1974).
[0006] However, while the effects of auditory feedback on singing and playing an instrument differ, it is thought that some of the mechanisms are shared within the brain.
[0007] This disclosure aims to provide a technology that enables the understanding of differences in delayed auditory feedback capabilities for multiple sound sources, such as singing and instrumental performance sounds.
[0008] The auditory feedback system of this disclosure comprises at least one of a first delayed sound generator and a second delayed sound generator, and an evaluation device. The first delayed sound generator generates a first delayed sound, which is an acoustic signal that has a predetermined delay in relation to the acoustic signals of multiple spoken utterances. The second delayed sound generator generates a second delayed sound, which is an acoustic signal that has a predetermined delay in relation to the acoustic signals of multiple played performances. The evaluation device causes a user wearing earphones to perceive at least one of sound, light, or vibration, which is output multiple times at predetermined intervals, and after multiple perceptions, the user is instructed to perform multiple utterances or multiple performances while listening to the first or second delayed sound through the earphones, with a predetermined interval as the target, and the evaluation device evaluates the multiple utterances or multiple performances performed.
[0009] According to the auditory feedback system of the embodiment of this disclosure, it is possible to grasp the differences in the delayed auditory feedback capabilities of multiple sound sources, such as singing and instrumental performance sounds.
[0010] Figure 1 is a diagram showing an example of the functional configuration of the auditory feedback system according to the first embodiment. Figure 2 is a diagram showing an example of the first processing flow of the auditory feedback system according to the first embodiment. Figure 3 is a diagram showing an example of the timing of the sound output from the earphones in the first processing flow example. Figure 4 is a diagram showing an example of the second processing flow of the auditory feedback system according to the first embodiment. Figure 5 is a diagram showing an example of the third processing flow of the auditory feedback system according to the first embodiment. Figure 6 is a diagram showing an example of the timing of the sound output from the earphones in the third processing flow example. Figure 7 is a diagram showing an example of the delay settings of the first delayed sound generator and the second delayed sound generator. Figure 8 is a diagram showing an example of the functional configuration of the auditory feedback system according to the second embodiment. Figure 9 is a diagram showing an example of the arrangement of singers and instrumentalists in a concert hall. Figure 10 is a diagram showing an example of the functional configuration of the auditory feedback system according to the third embodiment. Figure 11 is a diagram showing another example of the functional configuration of the auditory feedback system according to the third embodiment. Figure 12 is a diagram showing an example of the functional configuration of the auditory feedback system according to the fourth embodiment. Figure 13 is a diagram showing an example of the functional configuration of the auditory feedback system according to a modification of the first embodiment. Figure 14 is a diagram showing an example of the functional configuration of a computer.
[0011] The embodiments of this disclosure will be described in detail below. Components having the same function will be numbered the same, and redundant explanations will be omitted.
[0012] <<First Embodiment>>
[0013] Figure 1 shows an example of the functional configuration of an auditory feedback system according to the first embodiment. The auditory feedback system 100 of this embodiment includes a tone generation device 10, microphones 21 and 22, a first delayed sound generation device 31, a second delayed sound generation device 32, an amplifier 40, earphones 50, and an evaluation device 60. However, the auditory feedback system of this disclosure may be configured to have at least one of the first delayed sound generation device 31 and the second delayed sound generation device 32. That is, if only the first delayed sound generation device 31 is present, the microphones 22 and the second delayed sound generation device 32 are not required. If only the second delayed sound generation device 32 is present, the microphones 21 and the first delayed sound generation device 31 are not required. Furthermore, as will be described later, the tone generation device 10 is not an essential component. The function of the amplifier 40 may also be provided by the first delayed sound generator 31, the second delayed sound generator 32, or the earphone 50.
[0014] The auditory feedback system 100 uses a first delayed sound generator 31 and a second delayed sound generator 32 to add a delay to the sound of singing (speech) collected by the microphone 21 and the sound of a musical instrument, such as an electronic piano, and makes the delayed sound heard by the user through earphones 50 worn in the user's ears. In this disclosure, this is also referred to as "feedback." The auditory feedback system 100 uses an evaluation device 60 to evaluate the singing (speech) or performance performed based on the feedbacked sound.
[0015] (Tone Sound Generator 10) The tone sound generator 10 generates multiple tone sounds at predetermined intervals. A tone sound is an acoustic signal of a predetermined tone. The frequency, interval, and number of times the tone sound is presented may be configured to be arbitrarily set by the user. The generated tone sounds are output to the amplifier 40. The generated tone sounds may also be configured to be output to the evaluation device 60. Since the tone sounds indicate the timing to prompt the subject (singer or performer) to sing or play, this timing may be signaled by light. That is, as shown in Figure 1, the auditory feedback system 100 may be configured to have a light-emitting device 11 that emits light at predetermined intervals, either in place of the tone sound generator 10 or together with the tone sound generator 10. The light-emitting device 11 of this disclosure may be configured to emit light using a so-called light, or it may adopt a form that displays an image corresponding to light on a predetermined display screen. That is, it shall include a device that allows the timing to prompt the subject (singer or performer) to sing or play to be grasped visually. Similarly, this timing may be signaled by vibration. For example, the system may be configured to include a vibration generator 12 that emits light at predetermined intervals instead of the tone sound generator 10 or the light-emitting device 11, or together with the tone sound generator 10 or the light-emitting device 11. To facilitate understanding, this disclosure will describe the system using the tone sound generator 10. Although the description of the usage examples of the light-emitting device 11 and the vibration generator 12 will be omitted, the system may be configured to perform evaluation using either the light-emitting device 11 or the vibration generator 12. These points also apply to the second to fourth embodiments and the modifications of the first embodiment described later.
[0016] (Microphones 21, 22) Microphone 21 converts the input speech into an acoustic signal. The converted acoustic signal is output to the first delayed sound generator 31. Microphone 22 converts the input sound of a musical instrument into an acoustic signal and outputs it to the second delayed sound generator 32. For example, if the instrument is such that the acoustic signal generated by the performance can be output directly to the second delayed sound generator 32, such as an electronic piano, the system may be configured to output directly to the second delayed sound generator 32 without using microphone 22.
[0017] (First delayed sound generator 31) The first delayed sound generator 31 generates a first delayed sound, which is an acoustic signal obtained by applying a predetermined delay to the acoustic signals of a plurality of uttered speeches. For example, the content of the utterance may be "pa", but utterances other than this are also acceptable. The first delayed sound generator 31 can be implemented by, for example, using the Delay function of an effector for musical instruments, or a software application provided with a delay function for a personal computer. This point is the same for the second delayed sound generator 32.
[0018] As the predetermined delay, the settings of no delay (delay time zero) and applying delay (delay time > 0) may be configured to be arbitrarily settable by the user. Fig. 7 is a diagram showing an example of delay settings of the first delayed sound generator and the second delayed sound generator. The first delayed sound generator 31 may be configured such that its delay amount can be arbitrarily set, for example, as shown in the column of the first delayed sound generator 31 in Fig. 7, to be 0 to 200 milliseconds indicated by conditions A to F. The delay amounts in Fig. 7 are illustrative, and delay amounts larger or smaller than those in this example may be set. The first delayed sound generator 31 may be configured to apply a predetermined delay to the acoustic signals of utterances after a predetermined number of times among the plurality of uttered speeches. That is, the timing to apply a delay may be configured to be arbitrarily settable. For example, when eight tone sounds are presented and eight sounds are to be generated immediately after that, the setting may be made such that a delay is applied from the generation of the third sound. These points are the same for the second delayed sound generator 32 described later. The generated first delayed sound is output to the amplifier 40.
[0019] (Second Delayed Sound Generator 32) The second delayed sound generator 32 generates a second delayed sound, which is an acoustic signal that has been delayed by a predetermined amount of time relative to the acoustic signal of a performance that has been played multiple times. The sound of the performance may be, for example, "C," but other performances are also acceptable. The generated second delayed sound is output to the amplifier 40. The type of instrument that is the sound source of the acoustic signal input to the second delayed sound generator 32 does not matter. For example, it may be an instrument that can input sound directly to the second delayed sound generator 32 from an output terminal, such as an electronic piano. For example, it may be an instrument that does not have an output terminal, such as a flute. In the latter case, the microphone 22 is used to output the generated acoustic signal to the second delayed sound generator 32. As described above, predetermined delay settings such as no delay (delay time zero) or a delay (delay time > 0) may be configured to be arbitrarily set by the user. The second delayed sound generator 32 may be configured to allow its delay amount to be arbitrarily set, for example, as shown in the row of the second delayed sound generator 32 in Figure 7, such as 0 to 200 milliseconds as indicated by conditions A to F. The delay amounts in Figure 7 are illustrative examples, and the delay amount may be set to be larger or smaller than this example.
[0020] (Amplifier 40) The amplifier 40 amplifies the tone sound input from the tone sound generator 10, the first delayed sound input from the first delayed sound generator 31, and the second delayed sound input from the second delayed sound generator 32 into an acoustic signal of appropriate magnitude. The amplified acoustic signal is output to the earphone 50. The amplifier 40 may be configured to output the amplified acoustic signal to the evaluation device 60.
[0021] (Earphone 50) The earphone 50 outputs an input acoustic signal amplified by the amplifier 40. In the auditory feedback system 100 of the present disclosure, the earphone 50 outputs a plurality of tone sounds, and then continuously outputs the first delayed sound, the second delayed sound, or both. The earphone of the present disclosure is not limited to a type inserted into an ear canal, and includes so-called headphones, which are a type of earphone that covers the outer ear. However, as for the earphone 50 used in the auditory feedback system 100, headphones covering the outer ear are more preferable in order to prevent the user from directly hearing the sound of his / her own utterance or played sound.
[0022] The auditory feedback system 100 applies delay to the singing (utterance) collected by the microphone 21, the performance sound collected by the microphone 22, or the performance sound directly input to the second delayed sound generating device 32 by using the first delayed sound generating device 31 or the second delayed sound generating device 32 respectively, and feeds back the first delayed sound and the second delayed sound to the user's ears through the earphone 50. At this time, masking noise such as white noise may be added to the sound from the earphone 50 to reduce the influence of bone-conducted sound during utterance.
[0023] (Evaluation Device 60) The evaluation device 60 evaluates acoustic signals of a plurality of times of uttered speech input from the earphone 50, the amplifier 40, or the tone sound generating device 10, or acoustic signals of a plurality of times of played musical instrument sound input, or both. That is, a user wearing the earphone 50 is caused to sense the sound of the tone sound generating device 10 that outputs a plurality of times at predetermined intervals, the light from the light emitting device 11, or the vibration from the vibration generating device 12. After the user senses a plurality of times of sound, light, or vibration through auditory sense, visual sense, or tactile sense, the user is caused to perform a plurality of times of utterance or a plurality of times of performance while listening to the first delayed sound or the second delayed sound through the earphone 50 with the predetermined interval as a target. The evaluation device 60 evaluates the plurality of times of utterance or the plurality of times of performance performed by the user. Specific examples of the evaluation method will be described later.
[0024] The auditory feedback system 100 performs the auditory feedback method of this embodiment by carrying out the first to third processing flows illustrated in Figures 2, 4, and 5. Hereinafter, an example of the processing flow of the auditory feedback method in the auditory feedback system 100 will be described in order of procedure with reference to the figures.
[0025] <Example of First Processing Flow> Figure 2 is a diagram showing an example of the first processing flow of the auditory feedback system according to the first embodiment. As the auditory feedback method of the auditory feedback system 100, it is assumed that the user is wearing earphones 50 in their ears. The user is also asked to sing (speak) the same number of times at the same interval as the tone sound after listening to the tone sound. In this state, the processing is carried out in the following steps.
[0026] The tone generation device 10 generates and outputs tone sounds (step S1). The tone generation involves generating multiple tone sounds at predetermined intervals and outputting them to the amplifier 40. For example, a short tone sound such as 1 kHz is output twice at 400 millisecond intervals. This tone sound is output from the earphones 50. If a light-emitting device 11 is provided, it emits light towards the user at predetermined intervals in place of the tone generation device 10, or together with the tone generation device 10, and outputs a light signal corresponding to the light emitted by the light-emitting device 11 to the evaluation device 60. Similarly, if a vibration-generating device 12 is provided, it generates vibrations towards the user at predetermined intervals in place of the tone generation device 10 or the light-emitting device 11, or together with them, and outputs a vibration signal corresponding to the vibrations generated by the vibration-generating device 12 to the evaluation device 60.
[0027] Following step S1, the user makes their first utterance, for example, "pa," at the interval of the tone sounds presented by the tone sound generator 10. The utterance may be something other than "pa." As a result, the microphone 21 receives the user's utterance and converts the received utterance into an acoustic signal (step S2-1). The converted acoustic signal is output to the first delayed tone generator 31.
[0028] The first delayed sound generator 31 generates an acoustic signal (hereinafter also referred to as the "first delayed sound") with a predetermined delay, such as 150 milliseconds (step S3-1). The generated first delayed sound is output to the amplifier 40. The amplifier 40 amplifies the first delayed sound to a predetermined level, and the amplified first delayed sound is output from the earphone 50. The user listens to the first delayed sound output from the earphone 50. If the user has not reached a predetermined number of utterances (No. in step S4-1), the user returns to step S2-1 and makes the next utterance.
[0029] If the user has made a predetermined number of utterances (Yes in step S4-1), the evaluation device 60 performs an evaluation (step S5-1).
[0030] Figure 3 shows an example of the timing of sound output from the earphones in the first processing flow example. In this example, it is assumed that the tone generator 10 presents a 1 kHz tone twice, with an interval of 400 milliseconds. Figure 3 shows an example where, after the user has finished listening to the second tone, the user, based on their own perception, makes two utterances with an interval of 400 milliseconds, 400 milliseconds later. In this example, the first delayed sound generator 31 is set to a delay amount of 150 milliseconds, and the user hears the sound of their own utterance from the earphones 50 150 milliseconds after each utterance. In Figure 3, the black triangles indicate the timing when the user hears the tone from the earphones 50. The white circles indicate the timing when the user utters a sound. The black circles indicate the timing when the user hears the sound of their own utterance from the earphones 50.
[0031] The evaluation performed by the evaluation device 60 in step S5-1 is, for example, to output the acoustic signals of multiple utterances as a graph or numerical value. For example, only the timing of the user's utterances in Figure 3 may be output as a graph or numerical value, or the timing of the tone sound and the first delayed sound in Figure 3 may be output simultaneously. Alternatively, instead of timing, the waveform of the sound may be output.
[0032] The evaluation performed by the evaluation device 60 in step S5-1 may also be an output of the difference (hereinafter also referred to as the "first difference") between the evaluation result of the acoustic signals of multiple spoken utterances when a predetermined delay is introduced and the evaluation result of the acoustic signals of multiple spoken utterances when no predetermined delay is introduced.
[0033] <Example of Second Processing Flow> Figure 4 is a diagram showing an example of the second processing flow of the auditory feedback system according to the first embodiment. As the auditory feedback method of the auditory feedback system 100, it is assumed that the user is wearing earphones 50 in their ears. Furthermore, after listening to the tone, the user is asked to play the same number of times at the same interval as the tone. In this state, the processing is carried out in the following steps.
[0034] Following step S1, the user plays a note, for example, "C," for the first time at the interval of the tone sounds presented by the tone sound generator 10. The note played does not have to be "C." If a microphone 22 is used, the microphone 22 receives the user's performance and converts the received speech into an acoustic signal (step S2-2). The converted acoustic signal is output to the second delayed sound generator 32. If an instrument that can be directly connected to the second delayed sound generator 32 without using the microphone 22 is used, the acoustic signal generated by the performance is output from the instrument to the second delayed sound generator 32.
[0035] The second delayed sound generator 32 generates an acoustic signal (hereinafter also referred to as the "second delayed sound") with a predetermined delay, such as 150 milliseconds (step S3-2). The generated second delayed sound is output to the amplifier 40. The amplifier 40 amplifies the second delayed sound to a predetermined level, and the amplified second delayed sound is output from the earphones 50. The user listens to the second delayed sound output from the earphones 50. If the user has not reached the predetermined number of plays (No. in step S4-2), they return to step S2-2 and perform the next play.
[0036] If the user has played the instrument a predetermined number of times (Yes in step S4-2), the evaluation device 60 performs an evaluation (step S5-2).
[0037] The example of the timing of the sound output from the earphone 50 in the second processing flow example is the same as when the first delayed sound generator 31 is replaced with the second delayed sound generator 32 in Figure 3. That is, in the case of the second processing flow example, in Figure 3, the black triangle indicates the timing when the user hears the tone sound from the earphone 50. The white circle indicates the timing when the user plays. The black circle indicates the timing when the user hears the sound of their own performance from the earphone 50.
[0038] The evaluation performed by the evaluation device 60 in step S5-2 is, for example, to output the sound signals of multiple performances as a graph or numerical value. For example, as explained using Figure 3, only the timing of the user's performance may be output as a graph or numerical value, or the timing of the tone sound and the second delayed sound may be output simultaneously. Alternatively, instead of timing, the sound waveform may be output.
[0039] The evaluation performed by the evaluation device 60 in step S5-2 may also be an output of the difference (hereinafter also referred to as the "second difference") between the evaluation result of the sound signals of multiple performances with a predetermined delay and the evaluation result of the sound signals of multiple performances without a predetermined delay.
[0040] <Example of Third Processing Flow> Figure 5 is a diagram showing an example of the third processing flow of the auditory feedback system according to the first embodiment. As the auditory feedback method of the auditory feedback system 100, it is assumed that the user is wearing earphones 50 in their ears. Furthermore, after listening to a tone, the user is asked to sing (speak) and play an instrument at the same interval and number of times as the tone. In this state, processing is carried out in the following steps.
[0041] In the example flow shown in Figure 5, following step S1 described above, the processes in steps S2-1 to S4-1 and steps S2-2 to S4-2 are carried out in parallel.
[0042] Once the user has completed a predetermined number of utterances and performances (Yes in step S4-1 and Yes in step S4-2), the evaluation device 60 performs an evaluation (step S5-3).
[0043] Figure 6 shows an example of the timing of sound output from the earphones in the third processing flow example. In this example, it is assumed that the tone generator 10 presents a 1 kHz tone twice, with an interval of 400 milliseconds. Figure 6 shows an example where the user, after listening to the second tone, speaks twice at a 400-millisecond interval, after 400 milliseconds have elapsed according to the user's perception. In this example, the first delayed sound generator 31 is set to a delay of 150 milliseconds, so the user hears the sound of their own speech from the earphones 50 150 milliseconds after each utterance.
[0044] Figure 6 shows an example where, in parallel, after the second tone sound has finished being heard, two performances are given at 400-millisecond intervals, with the user's perception of 400 milliseconds elapsed. In this example, the second delayed sound generator 32 is set to a delay of 150 milliseconds, so the user hears the sound of their own performance through the earphones 50 150 milliseconds after each performance.
[0045] The evaluation performed by the evaluation device 60 in step S5-3 is, for example, to output the acoustic signals of multiple utterances or multiple performances as a graph or numerical value. For example, in Figure 6, only the timing of the user's utterances and performances may be output, or the tone, the timing of the first delayed sound, and the timing of the second delayed sound may be output simultaneously. Alternatively, instead of timing, the waveform of the sound may be output.
[0046] The evaluation performed by the evaluation device 60 in step S5-3 may be the output of either one or both of the first difference and the second difference described above. The evaluation performed by the evaluation device 60 in step S5-3 may also be the output of the difference between the intervals between multiple utterances and the intervals between multiple performances.
[0047] Suppose that when the same person uses the auditory feedback system 100, for example, the interval U1 of the speech sound in Figure 6 becomes 500 milliseconds, and the interval U2 of the musical performance becomes 450 milliseconds. This indicates that speech is more susceptible to auditory feedback than musical performance. Evaluation can also be performed by acquiring interval data when the delay is zero (hereinafter also referred to as the "baseline"). That is, by subtracting the baseline interval data from the interval data when a predetermined delay (>0) is actually generated, the effect of the delay on speech or the delay on musical performance can be evaluated based on the amount of delay actually introduced. Step S5-3 may be configured to output this information as a graph or numerical value.
[0048] The larger the intervals between the generated sounds of speech or musical instrument performance (intervals U1 and U2 in Figure 3), the more likely it is that auditory feedback is influencing the individual. In other words, it can be judged that the individual is listening carefully to the sound from the earphones. Examples of methods for evaluating auditory feedback ability include outputting the intervals between generated sounds of speech or musical instrument performance as graphs or numerical values, or outputting the difference between these generated sounds.
[0049] In the third processing flow example, for example, as shown in conditions C and E in Figure 7, the first delayed sound generator 31 can generate the first delayed sound with a smaller delay than the second delayed sound. For example, as shown in conditions D and F in Figure 7, the second delayed sound generator 32 can generate the second delayed sound with a smaller delay than the first delayed sound. For example, as shown in conditions A and B in Figure 7, the first delayed sound generator 31 and the second delayed sound generator 32 can generate the first and second delayed sounds in such a way that the same delay occurs.
[0050] As described above, by using the auditory feedback system 100 of this disclosure, it is possible to grasp the differences in delayed auditory feedback ability for multiple sound sources, such as singing and instrumental performance sounds. Specifically, by having a subject use the auditory feedback system 100 separately for singing (speaking) and playing an instrument, obtaining evaluation results, and comparing them, it is possible to grasp the differences in delayed auditory feedback ability. Alternatively, by having a subject use the auditory feedback system 100 for singing (speaking) and playing an instrument simultaneously and comparing them, it is possible to grasp the differences in delayed auditory feedback ability.
[0051] In speech, for example, consecutive utterances of the Japanese syllabary such as "aiueo" or "papipupepo" are used, while in musical instrument performance, for example, "doremifaso" is used as a basic tone sequence. However, it has not been clearly established whether the difficulty levels of these speech and musical instrument performance tasks are equal. For example, if the effect of auditory feedback is evaluated using "papipupepo" and "doremifaso," and the result is that speech is more affected, it is difficult to determine whether this effect is purely due to auditory feedback, or whether it is because the speech "papipupepo" is a more difficult task than the musical instrument performance "doremifaso."
[0052] Therefore, a method is needed to eliminate the influence of task difficulty. In this disclosure, it is possible to employ a method that combines multiple sounds, such as "papi" or "dore." However, in order to minimize the influence of task difficulty in speech and instrument playing, in other words, to more accurately evaluate the pure effect of auditory feedback, it is preferable to use single syllables or single sounds.
[0053] <<Second Embodiment>> The auditory feedback system 100 can be used for singing ability evaluation. FIG. 8 is a diagram showing an example of the functional configuration of an auditory feedback system according to a second embodiment. In FIG. 8, instead of outputting the auditory feedback ability itself, data on intervals of generated sound during DAF for a large number of people and data on singing proficiency can be collected and modeled by machine learning. Examples of the "singing proficiency" include evaluation scores from singing instructors, and the maximum and average scores of karaoke. By utilizing this model, it is also possible to estimate singing proficiency from the interval of the sound generated by DAF and output the result.
[0054] In FIG. 8, a large number of singers (α 1 , α 2 , α 3 , …, α n ) are asked to utter (h 1 , h 2 , h 3 , …, h n ) using the auditory feedback system 100, and evaluation results (H 1 , H 2 , H 3 , …, H n ) are obtained in advance. The singing proficiency (J 1 , α 2 , α 3 , …, α n ) of the singers (α 1 , J 2 , J 3 , …, J n ) acquired in advance is combined with the above result to create training data ((H 1 , J 1 ), (H 2 , J 2 ), (H 3 , J 3 ), …, (H n , J n )). The singing ability learning device 200 has a learning model that estimates singing ability based on training data. The created training data is input to, for example, the singing ability learning device 200 to cause the device to perform machine learning, and a singing ability evaluation model W, which is a trained model, is generated.
[0055] For example, if the delay amount is set to condition B in Figure 7, that is, if both the first delayed sound generator 31 and the second delayed sound generator 32 are set to a delay of 150 milliseconds, and if the singing ability evaluation model W generated by the singing ability learning device 200 indicates that "the greater the interval between the generated sounds of speech, the better the singing," then when singing (speech) is input to the singing ability evaluation model W, the result that the greater the interval between the generated sounds of speech, the better the singing will be output.
[0056] <<Third Embodiment>> The auditory feedback system 100 can be used in concert hall design. Specifically, it can utilize the difference in the influence of auditory feedback between singing and instrumental performance and reflect it in the design of the concert hall. Figure 9 shows an example of the arrangement of singers and instrumentalists in a concert hall. It is known that singers are more strongly affected by the delay of the generated sound (reflection and reverberation) than instrumentalists. Therefore, it is desirable for the singer's position to have less reflection and reverberation. For example, suppose there is a concert hall with two areas, stage X and audience seating Y, as shown in Figure 9. Singer α is in charge of singing, and instrumentalist β 1 , β 2 The person in charge of playing the instrument will be α. Singer α will be positioned, for example, in the center of the stage X, close to the audience Y. Instrumentalist β 1 , β 2 For example, this refers to a position on stage X that is further away from the audience seating area Y than from the singer α, and is located slightly to the left or right of the center.
[0057] Figure 10 shows an example of the functional configuration of the auditory feedback system according to the third embodiment. The concert hall design device 300 is a device that outputs positional information of sound reflectors and speakers to be placed, or the size and structure of the stage X and audience seating Y, in response to the input of evaluation results from the auditory feedback system 100. When designing a concert hall assuming the placement described in Figure 9, singer α, instrumentalist β 1 , instrumentalist β 2By having each of the three individuals use the auditory feedback system 100, the evaluation results, which represent auditory feedback ability, are obtained in advance. In Figure 10, singer α is made to utter "h" to obtain the evaluation result H. Instrumentalist β 1 ni m 1 The performance was evaluated and the result was M 1 Obtain the instrumentalist β. 2 ni m 2 The performance was evaluated and the result was M 2 The following data is obtained. In addition to this data, data on the size and shape of the desired stage X and audience seating Y, and other necessary data are input into the concert hall design device 300 to obtain concert hall design data Z1. The concert hall design data Z1 can include, for example, information on the positions of sound reflectors and speakers to be placed. Furthermore, for example, when designing a stage X and audience seating Y suitable for an orchestra and vocalists of a predetermined number, the evaluation results of the auditory feedback system 100 can be input into the concert hall design device 300 to obtain a proposed structural design for the stage X and audience seating Y themselves. In other words, the results of the auditory feedback system 100 of this disclosure can also be reflected in the design of the concert hall.
[0058] Figure 11 shows another example of the functional configuration of the auditory feedback system according to the third embodiment. In Figure 10, singer α was only responsible for singing, but in Figure 11, an example of use is assumed where singer α is responsible for both singing and playing an instrument, a so-called singer-songwriter. When singer α performs singing and playing an instrument simultaneously, as shown in Figure 11, the auditory feedback system 100 is used to have singer α perform speech h and performance m, and feedback results H and M are obtained. In addition to this data, data on the size and shape of the desired stage X and audience seating Y, and other necessary data are input to the concert hall design device 300 to obtain concert hall design data Z2. Concert hall design data Z2 can be obtained, for example, data on the size and shape of the stage X and audience seating Y modified based on the auditory feedback capability results, or position information of sound reflectors and speakers to be placed.
[0059] The auditory feedback system 100 can input evaluation results into the concert hall design device 300 and apply them to the design of the concert hall, not only when singer α is responsible for both singing and playing instruments, as shown in Figure 11, but also when one or more instrumentalists participate in addition to singer α, as shown in Figure 10.
[0060] <<Fourth Embodiment>> The auditory feedback system 100 can be used in communication network design. Specifically, the difference in the influence of auditory feedback between singing and playing an instrument can be utilized and reflected in the communication network design. Figure 12 is a diagram showing an example of the functional configuration of the auditory feedback system according to the fourth embodiment. The online session server device 400 is a system that, upon receiving the evaluation results of the auditory feedback system 100 for multiple subjects, takes these results into consideration, adds an appropriate delay to the acoustic signals input from each subject (singer or musician) as an online session, and transmits them to participants or viewers participating in the online session. In this example, singer α and three instrumentalists β 1 ~β 3 However, each has terminal T 1 ~T 4 It is connected to network N and can access the online session server device 400 connected to network N. (Singer α, three instrumentalists β) 1 ~β 3 Prior to this, participants were asked to use the auditory feedback system 100, and then spoken h and played m. 1 ~m 3 Evaluation results H, M obtained from 1 ~M 3 The data is obtained. By inputting these results into the online session server device 400, they can be reflected in the communication network design. In a live music session over a network, delays may occur due to the limited network bandwidth. In this case, since the singer is susceptible to auditory feedback, the online session server device 400 takes the delay of the singer α and inputs it to the instrumentalist β. 1 ~β 3By setting it to a lower level, the impact of auditory feedback on the singer can be reduced.
[0061] <<Modification of the First Embodiment>> The auditory feedback system of the present disclosure may be configured as shown in Figure 13, as shown in the auditory feedback system 101. Figure 13 is a diagram showing an example of the functional configuration of an auditory feedback system according to a modification of the first embodiment. The auditory feedback system 101 has a control device 70 in addition to the configuration of the auditory feedback system 100.
[0062] The control device 70 controls the tone sound generator 10, the first delayed sound generator 31, and the second delayed sound generator 32. Similar to the auditory feedback system 100, the tone sound generator 10 may be replaced by or combined with a light-emitting device 11 or a vibration-generating device 12. That is, the control device 70 controls the frequency of the tone sound produced by the tone sound generator 10 and the interval between multiple tone sounds. The control device 70 controls the number of flashes and intervals of the light-emitting device 11. The control device 70 controls the number of vibrations and intervals of the vibration-generating device 12. The control device 70 controls the presence or absence of delay, the delay time, and the timing of the delay for the first delayed sound generator 31 and the second delayed sound generator 32. The user can input desired settings for the tone generation device 10, light emission device 11, vibration generation device 12, first delayed sound generation device 31, and second delayed sound generation device 32 to the control device 70, and the auditory feedback system 101 controls the tone generation device 10, light emission device 11, vibration generation device 12, first delayed sound generation device 31, and second delayed sound generation device 32 according to the settings input by the user. In addition, the control device 70 may be configured to control the tone generation device 10, light emission device 11, vibration generation device 12, first delayed sound generation device 31, and second delayed sound generation device 32 using predetermined settings even if the user has not made any desired settings.
[0063] [Processors, Programs, Recording Media] The functions realized by the components described herein may be implemented in a circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the functions described herein. A processor includes transistors and other circuits and is considered a circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in memory.
[0064] In this specification, circuitry, unit, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0065] If the hardware is a processor that is considered to be a type of circuitry, then the circuitry, means, or unit is a combination of hardware and software used to constitute the hardware and / or processor.
[0066] The various processes described above can be carried out by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 14, and then causing the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc. to operate.
[0067] The program describing this process can be recorded on a computer-readable recording medium. Any computer-readable recording medium can be used, such as a magnetic recording device, optical disc, magneto-optical recording medium, or semiconductor memory.
[0068] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0069] A computer executing such a program may, for example, first store the program recorded on a portable storage medium or a program transferred from a server computer in its own storage device. Then, when processing is to be executed, the computer reads the program stored on its own storage medium and executes the processing according to the read program. Alternatively, the computer may directly read the program from the portable storage medium and execute the processing according to that program, or it may sequentially execute the processing according to the received program each time a program is transferred to it from a server computer. Furthermore, the processing may be executed using a so-called ASP (Application Service Provider) type service, where the processing function is realized only by issuing execution instructions and obtaining results, without transferring the program from the server computer to this computer.In addition, the processing may be executed using a so-called SaaS (Software as a Service) type service, where a part of the server computer is made available to the user along with the program. Furthermore, the term "program" in this form includes information used for processing by an electronic computer that is equivalent to a program (data, etc., that is not a direct instruction to the computer but has the property of defining the computer's processing).
[0070] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware.
Claims
1. An auditory feedback system comprising at least one of a first delayed sound generating device and a second delayed sound generating device, and an evaluation device, wherein the first delayed sound generating device generates a first delayed sound which is an acoustic signal that has a predetermined delay to the acoustic signal of multiple spoken utterances, the second delayed sound generating device generates a second delayed sound which is an acoustic signal that has a predetermined delay to the acoustic signal of multiple played performances, and the evaluation device causes a user wearing earphones to perceive at least one of sound, light, or vibration which is output multiple times at predetermined intervals, and after the perception, the user is made to listen to the first delayed sound or the second delayed sound through the earphones, with the predetermined interval as the target, while performing the multiple utterances or the multiple played performances, and the auditory feedback system evaluates the performed multiple utterances or the multiple played performances.
2. The auditory feedback system according to claim 1, wherein the first delayed sound generation device causes a predetermined delay in the acoustic signals of utterances from a predetermined number of times onward among the multiple utterances, and the second delayed sound generation device causes a predetermined delay in the acoustic signals of performances from a predetermined number of times onward among the multiple performances.
3. The auditory feedback system according to claim 1, comprising both the first delayed sound generating device and the second delayed sound generating device, wherein the first delayed sound generating device generates the first delayed sound with a delay smaller than that of the second delayed sound, or the second delayed sound generating device generates the second delayed sound with a delay smaller than that of the first delayed sound.
4. The auditory feedback system according to claim 1, wherein the evaluation by the evaluation device is an output of the acoustic signals of the multiple utterances or the multiple performances as a graph or numerical value.
5. The auditory feedback system according to claim 1, wherein the evaluation by the evaluation device outputs a first difference, which is the difference between the evaluation result of the acoustic signals of the multiple spoken utterances when the predetermined delay is introduced and the evaluation result of the acoustic signals of the multiple spoken utterances when the predetermined delay is not introduced; outputs a second difference, which is the difference between the evaluation result of the acoustic signals of the multiple played performances when the predetermined delay is introduced and the evaluation result of the acoustic signals of the multiple played performances when the predetermined delay is not introduced; or outputs of at least one of the first difference and the second difference.
6. The auditory feedback system according to claim 1, comprising both the first delayed sound generation device and the second delayed sound generation device, wherein the evaluation by the evaluation device is the output of the difference between the interval between the multiple utterances and the interval between the multiple performances.
7. An auditory feedback method performed by an auditory feedback system having at least one of a first delayed sound generating device and a second delayed sound generating device, and an evaluation device, wherein the first delayed sound generating device generates a first delayed sound which is an acoustic signal that has a predetermined delay to the acoustic signals of multiple spoken utterances, the second delayed sound generating device generates a second delayed sound which is an acoustic signal that has a predetermined delay to the acoustic signals of multiple played performances, and the evaluation device causes a user wearing earphones to perceive at least one of sound, light, or vibration which is output multiple times at predetermined intervals, and after the perception, the user is made to listen to the first delayed sound or the second delayed sound through the earphones, with the predetermined interval as the target, while performing the multiple utterances or the multiple played performances, and the performed multiple utterances or the multiple played performances are evaluated.
8. A program for causing a computer to function with the auditory feedback method described in claim 7.