Acoustic device

The acoustic device addresses the issue of unintended speech recognition in information terminals by utilizing a machine learning model to register and recognize user voice features, analyze voice commands, and cancel noise, thereby improving speech recognition accuracy and suppressing malfunctions.

JP2025081730APending Publication Date: 2025-05-27SEMICON ENERGY LAB CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025032681
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-09
Filing Date
2025-03-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Information terminals, such as smartphones, may mistakenly recognize speech from individuals other than the user, leading to unintended operations.

Method used

An acoustic device equipped with a sound detection unit, a sound separation unit, a sound determination unit, and a processing unit, which uses a machine learning model to register and recognize the voice features of the user, analyze voice commands, and generate signals to execute instructions while canceling noise.

Benefits of technology

The acoustic device effectively suppresses malfunctions in information terminals by ensuring that only registered voice commands are executed, thereby enhancing the precision of speech recognition and reducing noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081730000001_ABST
    Figure 2025081730000001_ABST
Patent Text Reader

Abstract

To provide an acoustic device which suppresses malfunction of an information terminal and can highly precisely recognize voice, an operation method of the device, an information processing system and an information processing method.SOLUTION: An acoustic device 10 includes: a sound detection unit 11 having a function for detecting sound 21; a sound separation unit 12 for separating the sound which the sound detection unit detects into voice and sound other than the voice; a sound determination unit 13 for registering a feature amount of the sound; a processing unit 15 for determining whether or not the feature amount of voice which the sound separation unit separates is the one registered by a machine learning model such as a neural network model, analyzing a command included in the voice if the feature amount of the voice is the one registered, generating a command signal indicating a content of the command and performing processing to cancel sound other than the voice to the sound other than the voice which the sound separation unit separates; a transmission / reception unit 16 for synthesizing the sound having been subjected to processing by the processing unit and sound which an information terminal 22 emits; and a sound output unit 17 for emitting the synthesized sound to the outside of the acoustic device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to an acoustic device and an operation method thereof. One aspect of the present invention relates to an information processing system and an information processing method.

Background Art

[0002] In recent years, the development of speech recognition technology has been advanced. By speech recognition, for example, when a user of an information terminal such as a smartphone speaks, the information terminal can execute an instruction included in the speech.

[0003] In order to improve the accuracy of speech recognition, it is preferable to cancel noise. Patent Document 1 discloses a headset that can cancel noise included in a voice signal.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] When an information terminal performs speech recognition, for example, the information terminal may recognize the speech of a person other than the user, and as a result, the information terminal may perform an operation unintended by the user.

[0006] One aspect of the present invention is to provide an acoustic device capable of suppressing malfunction of an information terminal. One aspect of the present invention is to provide an acoustic device capable of canceling noise. One aspect of the present invention is to provide an acoustic device that enables an information terminal to perform high-precision speech recognition. One aspect of the present invention is to provide a novel acoustic device.

[0007] One aspect of the present invention aims to provide an information processing system with suppressed malfunction. One aspect of the present invention aims to provide an information processing system capable of canceling noise. One aspect of the present invention aims to provide an information processing system capable of performing high-precision speech recognition. One aspect of the present invention aims to provide a novel information processing system.

[0008] One aspect of the present invention aims to provide an operation method for an acoustic device capable of suppressing malfunction of an information terminal. One aspect of the present invention aims to provide an operation method for an acoustic device capable of canceling noise. One aspect of the present invention aims to provide an operation method for an acoustic device that enables an information terminal to perform high-precision speech recognition. One aspect of the present invention aims to provide a novel operation method for an acoustic device.

[0009] One aspect of the present invention aims to provide an information processing method with suppressed malfunction. One aspect of the present invention aims to provide an information processing method capable of canceling noise. One aspect of the present invention aims to provide an information processing method capable of performing high-precision speech recognition. One aspect of the present invention aims to provide a novel information processing method.

[0010] Note that the description of these problems does not prevent the existence of other problems. Note that one aspect of the present invention does not need to solve all of these problems. Note that other problems can be extracted from the descriptions in the specification, drawings, claims, etc.

Means for Solving the Problems

[0011] One aspect of the present invention has a sound detection unit, a sound separation unit, a sound determination unit, and a processing unit. The sound detection unit has a function of detecting a first sound. The sound separation unit has a function of separating the first sound into a second sound and a third sound. The sound determination unit has a function of registering a feature amount of the sound. The sound determination unit has a function of determining whether or not the feature amount of the second sound is registered using a machine learning model. When the feature amount of the second sound is registered, the processing unit has a function of analyzing an instruction included in the second sound and generating a signal representing the content of the instruction. The processing unit has a function of generating a fourth sound by performing processing for canceling the third sound on the third sound. It is an acoustic device.

[0012] Alternatively, in the above aspect, the learning of the machine learning model may be performed using supervised learning in which voice is used as learning data and a label indicating whether or not to perform registration is used as teacher data.

[0013] Alternatively, in the above aspect, the machine learning model may be a neural network model.

[0014] Alternatively, in the above aspect, the fourth sound may be a sound having a reverse phase with respect to the third sound.

[0015] Alternatively, one aspect of the present invention is to detect a first sound, separate the first sound into a second sound and a third sound, determine whether or not the feature amount of the second sound is registered using a machine learning model, and when the feature amount of the second sound is registered, analyze an instruction included in the second sound and generate a signal representing the content of the instruction, and perform processing for canceling the third sound on the third sound to generate a fourth sound. It is an operation method of an acoustic device.

[0016] Alternatively, in the above aspect, the learning of the machine learning model may be performed using supervised learning in which voice is used as learning data and a label indicating whether or not to perform registration is used as teacher data.

[0017] Alternatively, in the above aspect, the machine learning model may be a neural network model.

[0018] Alternatively, in the above aspect, the fourth sound may be a sound having a phase opposite to that of the third sound.

Advantages of the Invention

[0019] According to one aspect of the present invention, an acoustic device capable of suppressing malfunction of an information terminal can be provided. According to one aspect of the present invention, an acoustic device capable of canceling noise can be provided. According to one aspect of the present invention, an acoustic device that enables an information terminal to perform high-precision speech recognition can be provided. According to one aspect of the present invention, a novel acoustic device can be provided.

[0020] According to one aspect of the present invention, an information processing system with suppressed malfunction can be provided. According to one aspect of the present invention, an information processing system capable of canceling noise can be provided. According to one aspect of the present invention, an information processing system capable of performing high-precision speech recognition can be provided. According to one aspect of the present invention, a novel information processing system can be provided.

[0021] According to one aspect of the present invention, an operation method of an acoustic device capable of suppressing malfunction of an information terminal can be provided. According to one aspect of the present invention, an operation method of an acoustic device capable of canceling noise can be provided. According to one aspect of the present invention, an operation method of an acoustic device that enables an information terminal to perform high-precision speech recognition can be provided. According to one aspect of the present invention, a novel operation method of an acoustic device can be provided.

[0022] According to one aspect of the present invention, an information processing method with suppressed malfunction can be provided. According to one aspect of the present invention, an information processing method capable of canceling noise can be provided. According to one aspect of the present invention, an information processing method capable of performing high-precision speech recognition can be provided. According to one aspect of the present invention, a novel information processing method can be provided.

[0023] Note that the description of these effects does not preclude the existence of other effects. Note that one aspect of the present invention does not necessarily have to have all of these effects. Note that other effects can be extracted from the descriptions in the specification, drawings, claims, etc.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0025] Hereinafter, embodiments will be described with reference to the drawings. However, the embodiments can be implemented in many different ways, and it is easily understood by those skilled in the art that the forms and details can be variously changed without departing from the spirit and its scope. Therefore, the present invention is not construed as being limited to the description of the following embodiments.

[0026] In the configuration of the invention described below, the same reference numerals are commonly used for the same parts or parts having the same functions among different drawings, and the repeated description thereof is omitted.

[0027] Also, the ordinal numbers "first", "second", "third", etc. used in this specification and the like are attached to avoid confusion of components and are not numerically limiting.

[0028] (Embodiment) In this embodiment, an acoustic device according to an aspect of the present invention and its operation method will be described. Also, an information processing system including the acoustic device according to an aspect of the present invention and an information processing method using the information processing system will be described.

[0029] <Configuration example of acoustic device> An acoustic device according to an aspect of the present invention can be, for example, an earphone or a headphone. An acoustic device according to an aspect of the present invention includes a sound detection unit, a sound separation unit, a sound determination unit, a processing unit, a transmission / reception unit, and a sound output unit. Here, the sound detection unit can be configured to include, for example, a microphone. Also, the sound output unit can be configured to include, for example, a speaker.

[0030] An acoustic device according to an aspect of the present invention is electrically connected to an information terminal such as a smartphone. Here, the acoustic device according to an aspect of the present invention and the information terminal may be connected by wire, or may be wirelessly connected by Bluetooth (registered trademark), Wi-Fi (registered trademark), or the like. It can be said that an information processing system according to an aspect of the present invention is configured by the acoustic device according to an aspect of the present invention and the information terminal.

[0031] Before using the acoustic device according to one aspect of the present invention, the feature amount (voiceprint) of the voice is registered in advance. For example, the feature amount of the voice of the user of the acoustic device according to one aspect of the present invention is registered. The feature amount of the voice can be, for example, the frequency characteristics of the voice. For example, it can be the frequency characteristics obtained by performing a Fourier transform on voice data which is data representing the voice. Also, as the feature amount of the voice, for example, Mel-Frequency Cepstrum Coefficients (MFCC) can be used.

[0032] When the sound detection unit detects a sound during the use of the acoustic device according to one aspect of the present invention, the sound separation unit separates the sound into voice and sounds other than voice. Here, the sounds other than voice can be, for example, environmental sounds, for example, noise.

[0033] Next, for the voice separated by the sound separation unit, the sound determination unit performs feature amount extraction and determines whether the extracted feature amount is registered. If it is registered, the processing unit analyzes the command included in the voice and generates a command signal which is a signal representing the content of the command. Note that the analysis of the command can be performed using language processing such as morphological analysis. The generated command signal is output to the transmission / reception unit.

[0034] On the other hand, if the feature amount extracted by the sound determination unit is not registered, the generation of the command signal is not performed.

[0035] After that, the processing unit performs processing for canceling the sound other than voice separated by the sound separation unit. For example, the processing unit generates a sound with a reverse phase to the sound.

[0036] Next, the transmission / reception unit synthesizes the sound processed by the processing unit and the sound emitted by the information terminal and outputs it to the sound output unit. Here, the sound emitted by the information terminal can be, for example, the music when the information terminal is playing music.

[0037] The sound output to the sound output unit is emitted outside the sound device according to one aspect of the present invention. The user of the sound device according to one aspect of the present invention can listen to the combined sound of the sound detected by the sound detection unit and the sound output by the sound output unit. As described above, the sound output by the sound output unit can include, in addition to the sound emitted by the information terminal, for example, the noise included in the sound detected by the sound detection unit with an opposite phase. As a result, the user of the sound device according to one aspect of the present invention can listen to, for example, a sound with noise canceled.

[0038] Also, when the processing unit generates a command signal and outputs it to the transmission / reception unit, that is, when the feature amount of the voice separated by the voice separation unit is registered, the transmission / reception unit outputs the command signal to the information terminal. The information terminal executes the command represented by the command signal. For example, when the information terminal is playing music and the command signal represents a command to "change the type of music", the music played by the information terminal can be changed to the specified one. The above is an example of the operation method of the sound device according to one aspect of the present invention.

[0039] Only when the feature amount of the voice separated by the voice separation unit is registered, by generating a command signal by the processing unit, it is possible to suppress malfunction of the information terminal compared to the case of generating a command signal regardless of the presence or absence of registration. For example, when registering the feature amount of the voice of the user of the information terminal in the sound device according to one aspect of the present invention, it is possible to suppress an operation unintended by the user of the information terminal from being performed in response to the voice of a person other than the user of the information terminal.

[0040] Here, the registration of the feature amount of the voice and the determination as to whether or not the feature amount of the voice input to the voice determination unit is registered can be performed using, for example, a machine learning model. As the machine learning model, for example, using a neural network model is preferable because inference can be performed with high accuracy. As the neural network model, for example, CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), etc. can be used. Also, as a learning method of the machine learning model, for example, supervised learning can be used.

[0041] When using supervised learning, for example, the feature amount of voice can be used as learning data, and a label indicating whether to perform registration or not can be used as teacher data.

[0042] When using supervised learning, the learning can be performed in two stages: the first learning and the second learning. That is, after the first learning, the second learning can be performed as additional learning.

[0043] In the first learning, for all the learning data, a label indicating "not to perform registration" is assigned as teacher data. In the first learning, it is preferable to use the feature amounts of voices of a plurality of people as learning data. In particular, for example, it is preferable to prepare the learning data of male voices and female voices without bias, and also to prepare the learning data of various voice qualities such as high voices and low voices among male and female voices without bias. Thereby, the inference using the learning result described later, that is, the determination as to whether the feature amount of the voice input to the voice determination unit is registered or not, can be performed with high accuracy.

[0044] In the second learning, for all the learning data, a label indicating "to perform registration" is assigned as teacher data. That is, by the second learning, the registration of the feature amount of the voice can be performed.

[0045] In the second learning, for example, the feature amount of the voice of the user of the acoustic device according to an aspect of the present invention is used as learning data. As the learning data, it is preferable to use the feature amounts of voices uttered by the same person in various uttering methods without bias. In addition, it is preferable to increase the number of learning data by changing parameters such as the pitch of the voice for the voice data acquired as learning data. As described above, the inference using the learning result, that is, the determination as to whether the feature amount of the voice input to the voice determination unit is registered or not, can be performed with high accuracy.

[0046] The first learning can be performed, for example, before the shipment of the acoustic device according to one aspect of the present invention. On the other hand, the second learning can be performed, for example, after the shipment of the acoustic device according to one aspect of the present invention. As a result, the second learning can be performed, for example, by the user of the acoustic device according to one aspect of the present invention himself / herself. As described above, in the acoustic device according to one aspect of the present invention, the user can register the feature amount of the voice by himself / herself.

[0047] By performing the learning described above, the sound determination unit can determine whether or not the feature amount of the voice separated by the sound separation unit is registered. Specifically, when a voice is input to the sound determination unit, the sound determination unit can infer whether or not the feature amount of the voice input to the sound determination unit is registered based on the learning result.

[0048] By performing the determination as to whether or not the feature amount of the voice is registered using a machine learning model, a more accurate determination can be made than when performing the determination without using a machine learning model. As a result, for example, it is possible to suppress an information terminal electrically connected to the acoustic device according to one aspect of the present invention from executing an instruction included in a voice for which the feature amount is not registered. In addition, for example, it is possible to suppress an information terminal electrically connected to the acoustic device according to one aspect of the present invention from not executing an instruction included in a voice for which the feature amount is registered. That is, the information terminal electrically connected to the acoustic device according to one aspect of the present invention can perform high-precision voice recognition.

[0049] FIG. 1A is a diagram showing a configuration example of an acoustic device 10, which is an acoustic device according to one aspect of the present invention. In FIG. 1A, for the purpose of explaining the functions of the acoustic device 10, in addition to the acoustic device 10, a sound 21, an information terminal 22, and an ear 23 are shown. Here, the information terminal 22 can be, for example, a smartphone. In addition, the information terminal 22 can be a portable electronic device such as a tablet terminal, a laptop PC, or a portable (portable) game machine. Note that the information terminal 22 may be an electronic device other than a portable electronic device.

[0050] The audio device 10 includes a sound detection unit 11, a sound separation unit 12, a sound determination unit 13, a memory unit 14, a processing unit 15, a transmission / reception unit 16, and a sound output unit 17.

[0051] Here, the transmission / reception unit 16 is electrically connected to the information terminal 22. The audio device 10 and the information terminal 22 may be connected by wire, or may be wirelessly connected by Bluetooth (registered trademark), Wi-Fi (registered trademark), or the like. It can be said that an information processing system according to an aspect of the present invention is configured by the audio device 10 and the information terminal 22.

[0052] In FIG. 1A, the arrows indicate the flow of data, signals, etc. However, the flow shown in FIG. 1A is an example and is not limited to what is shown in FIG. 1A. The same applies to other figures.

[0053] The sound detection unit 11 has a function of detecting sound. For example, it has a function of detecting sound 21 including human voice. The sound detection unit 11 can be configured to include, for example, a microphone.

[0054] The sound separation unit 12 has a function of separating the sound detected by the sound detection unit 11 according to characteristics. For example, when the sound detection unit 11 detects sound 21 including human voice, it has a function of separating sound 21 into voice and sound other than voice. Here, the sound other than voice can be, for example, environmental sound and can be, for example, noise.

[0055] The sound separation unit 12 has a function of separating, for example, the sound detected by the sound detection unit 11 based on the frequency of the sound. For example, human speech is mainly composed of frequency components of 0.2 kHz or higher and 4 kHz or lower. Therefore, for example, by separating the sound detected by the sound detection unit 11 into a sound with a frequency of 0.2 kHz or higher and 4 kHz or lower and a sound with a frequency other than that, it is possible to separate the speech from other sounds. It is said that the intermediate frequency of human speech is around 1 kHz. Therefore, for example, by separating the sound detected by the sound detection unit 11 into a sound with a frequency around 1 kHz and a sound with a frequency other than that, it may be possible to separate the speech from other sounds. For example, it may be separated into a sound with a frequency of 0.5 kHz or higher and 2 kHz or lower and a sound with a frequency other than that. Also, for example, the frequency for performing sound separation may be changed according to the type of sound detected by the sound detection unit 11. For example, when the sound detection unit 11 detects a sound including a female voice, a sound with a higher frequency than when it detects a sound including a male voice may be separated as speech. By changing the frequency for performing sound separation according to the type of sound detected by the sound detection unit 11, for example, the sound detected by the sound detection unit 11 can be separated into speech and other sounds with high accuracy.

[0056] The sound determination unit 13 has a function of performing feature quantity extraction on the sound separated by the sound separation unit 12. Specifically, for example, it has a function of performing feature quantity extraction on the speech separated by the sound separation unit 12. Note that the feature quantity of speech can be called a voiceprint.

[0057] The feature quantity can be, for example, a frequency characteristic. For example, it can be a frequency characteristic obtained by performing a Fourier transform on sound data, which is data representing sound. Also, as a feature quantity of sound, for example, MFCC can be used.

[0058] The extracted feature quantity can be registered. For example, a voiceprint can be registered. From the above, it can be said that the sound determination unit 13 has a function of registering the feature quantity of sound. The registration result can be stored in the storage unit 14.

[0059] In addition, the sound determination unit 13 has a function of determining whether or not the extracted feature amount is registered. The registration of the feature amount and the above determination can be performed using, for example, a machine learning model. When using, for example, a neural network model as the machine learning model, it is preferable because inference can be performed with high accuracy. As the neural network model, for example, CNN, RNN, etc. can be used. Further, as a learning method of the machine learning model, for example, supervised learning can be used.

[0060] The processing unit 15 has a function of performing processing on the sound output by the sound separation unit 12, for example. For example, it has a function of analyzing an instruction included in the sound output by the sound separation unit 12 and generating an instruction signal which is a signal representing the content of the instruction. Note that the analysis of the instruction can be performed using, for example, language processing such as morphological analysis.

[0061] In addition, the processing unit 15 has a function of performing processing for canceling noise or the like among the sounds output by the sound separation unit 12. For example, by generating a sound having a phase opposite to that of the noise or the like, the noise or the like output by the sound separation unit 12 can be canceled.

[0062] Here, the processing unit 15 has a function of performing processing based on the determination result of the sound determination unit 13. For example, when the sound separation unit 12 outputs a voice, an instruction signal can be generated only when the feature amount of the voice is registered.

[0063] The transmission / reception unit 16 has a function of synthesizing the sound processed by the processing unit 15 and the sound emitted by the information terminal 22. Here, the sound emitted by the information terminal 22 can be, for example, the music when the information terminal 22 is playing music.

[0064] Also, when the processing unit 15 generates a command signal, the command signal can be received by the transceiver unit 16. The transceiver unit 16 has a function of outputting the received command signal to the information terminal 22. The information terminal 22 has a function of executing the command represented by the command signal. For example, when music is being played on the information terminal 22 and the command signal represents a command to "change the type of music", the music played by the information terminal 22 can be changed to the specified one.

[0065] As described above, the command signal is generated only when, for example, the feature amount of the voice separated by the voice separation unit 12 is registered. Thereby, it is possible to suppress malfunction of the information terminal 22 as compared with the case where the command signal is generated regardless of registration or not. For example, when registering the feature amount of the voice of the user of the information terminal 22 in the audio device 10, it is possible to suppress an operation unintended by the user of the information terminal 22 from being performed in response to the voice of a person other than the user of the information terminal 22.

[0066] The sound output unit 17 has a function of emitting the sound synthesized by the transceiver unit 16 to the outside of the audio device 10. The user of the audio device 10 can hear the synthesized sound of the sound detected by the sound detection unit 11 and the sound output by the sound output unit 17 with the ear 23. As described above, the sound output by the sound output unit 17 can include, in addition to the sound emitted by the information terminal 22, for example, the noise or the like included in the sound detected by the sound detection unit 11 with an inverse phase. As a result, the user of the audio device 10 can hear, for example, a sound with noise or the like canceled. Note that the sound output unit 17 can be configured to include, for example, a speaker.

[0067] FIG. 1B1 and FIG. 1B2 are diagrams showing a specific example of the audio device 10. As shown in FIG. 1B1, the audio device 10 can be an earphone. Specifically, it can be an earphone worn by the user of the information terminal 22. Also, as shown in FIG. 1B2, the audio device 10 can be a headphone. Specifically, it can be a headphone worn by the user of the information terminal 22.

[0068] <Operation example of the audio device> Hereinafter, an example of the operation method of the acoustic device 10 will be described. FIGS. 2A and 2B are diagrams showing an example of a method for registering the feature amount of sound when the sound determination unit 13 has a function of determining whether or not the feature amount of sound is registered using a machine learning model. Specifically, it is a diagram showing an example of a method for registering the feature amount of sound using supervised learning.

[0069] First, as shown in FIG. 2A, the sound determination unit 13 extracts a feature amount from the sound data 31. For example, the frequency characteristics of the sound represented by the sound data 31 are used as the feature amount. For example, the frequency characteristics obtained by performing a Fourier transform on the sound data 31 can be used as the feature amount. Also, for example, MFCC can be used as the feature amount.

[0070] Thereafter, data representing the extracted feature amount with a label 32 indicating "not to be registered" is input to a generator 30 provided in the sound determination unit 13. The generator 30 is a program using a machine learning model. The generator 30 performs learning using the data representing the feature amount extracted from the sound data 31 as learning data and the label 32 as teacher data, and outputs a learning result 33. The learning result 33 can be stored in the storage unit 14. Note that when the generator 30 is a program using a neural network model, the learning result 33 can be a weight coefficient.

[0071] It is preferable to use the voices of a plurality of people as the sound data 31 which is the learning data. In particular, for example, it is preferable to prepare sound data of male voices and female voices without bias, and also to prepare sound data of various voice qualities such as high voices and low voices among male and female voices without bias and perform learning. Thereby, the inference using the learning result described later, that is, the determination as to whether or not the feature amount of the sound input to the sound determination unit 13 is registered can be performed with high accuracy.

[0072] Next, as shown in FIG. 2B, the sound determination unit 13 extracts feature amounts from the sound data 41. It is preferable that the feature amounts be of the same type as the feature amounts used as learning data in FIG. 2A. For example, when MFCC is extracted from the sound data 31 and used as learning data, it is preferable to also extract MFCC from the sound data 41.

[0073] Thereafter, data representing the extracted feature amounts with a label 42 which is a label indicating "perform registration" attached thereto is input to the generator 30 in which the learning result 33 is read. The generator 30 performs learning using the data representing the feature amounts extracted from the sound data 41 as learning data and the label 42 as teacher data, and outputs a learning result 43. The learning result 43 can be stored in the storage unit 14. When the generator 30 is a program using a neural network model, the learning result 43 can be a weight coefficient.

[0074] In FIGS. 2A and 2B, a label indicating "perform registration" is shown described as "registration ○", and a label indicating "do not perform registration" is shown described as "registration ×". The same description is made in other drawings as well.

[0075] The sound data 41 which is learning data is, for example, the voice of the user of the acoustic device 10. When using voice as the sound data 41, it is preferable to perform learning using the feature amounts of voices uttered by the same person by various utterance methods without bias. Further, for the voice data acquired as the sound data 41, it is preferable to increase the number of the sound data 41 by changing parameters such as the pitch of the voice and perform learning. As described above, inference using the learning result described later, that is, determination as to whether or not the feature amounts of the sound input to the sound determination unit 13 are registered can be performed with high accuracy.

[0076] As described above, after the sound determination unit 13 performs learning using the feature amounts of sounds that are not registered as shown in FIG. 2A as learning data, it can perform learning using the feature amounts of sounds that are registered as shown in FIG. 2B as learning data. That is, learning can be performed in two stages: the first learning and the second learning. Specifically, after performing the first learning shown in FIG. 2A, the second learning shown in FIG. 2B can be performed as additional learning.

[0077] The first learning can be performed, for example, before the shipment of the audio device 10. On the other hand, the second learning can be performed, for example, after the shipment of the audio device 10. As a result, the second learning can be performed, for example, by the user of the audio device 10 himself / herself. As described above, in the audio device 10, the user can register the feature amounts of sounds himself / herself.

[0078] By performing the learning shown above, the sound determination unit 13 can, for example, determine whether or not the feature amounts of the sounds separated by the sound separation unit 12 are registered. Specifically, when a sound is input to the sound determination unit 13, the sound determination unit 13 can infer whether or not the feature amounts of the input sound are registered based on the learning result 43.

[0079] By performing the determination as to whether or not the feature amounts of sounds are registered using a machine learning model, a more accurate determination can be made than when performing the determination without using a machine learning model. As a result, for example, it is possible to suppress an information terminal 22 electrically connected to the audio device 10 from executing an instruction included in a sound whose feature amounts are not registered. Also, for example, it is possible to suppress the information terminal 22 electrically connected to the audio device 10 from not executing an instruction included in a sound whose feature amounts are registered. That is, the information terminal 22 electrically connected to the audio device 10 can perform highly accurate speech recognition.

[0080] Next, an example of the operation method during the use of the audio device 10 will be described. FIG. 3 is a flowchart showing an example of the operation method during the use of the audio device 10. FIGS. 4A to 4C, and FIGS. 5A and 5B are schematic diagrams for explaining the details of each step shown in FIG. 3. Hereinafter, it is assumed that the registration of the sound feature amount has already been performed by the method shown in FIGS. 2A and 2B, and the following description will be given.

[0081] When the sound detection unit 11 detects a sound (step S01), the sound separation unit 12 separates the detected sound for each characteristic. For example, when the sound detection unit 11 detects a sound including a human voice, the sound separation unit 12 separates the detected sound into a voice and a sound other than the voice (step S02). As described above, the sound other than the voice can be, for example, environmental sound, for example, noise.

[0082] A specific example of step S02 is shown in FIG. 4A. As described above, the sound separation unit 12 has a function of separating, for example, the sound detected by the sound detection unit 11 based on the frequency of the sound. FIG. 4A shows an example in which the sound 21 detected by the sound detection unit 11 and input to the sound separation unit 12 is separated into a sound 21a and a sound 21b based on the frequency.

[0083] As described above, human voice is mainly composed of, for example, frequency components of 0.2 kHz or higher and 4 kHz or lower. Therefore, for example, by separating the sound detected by the sound detection unit 11 into sounds with frequencies of 0.2 kHz or higher and 4 kHz or lower and sounds with other frequencies, the voice can be separated from other sounds. It is said that the intermediate frequency of human voice is around 1 kHz. Therefore, for example, by separating the sound detected by the sound detection unit 11 into sounds with frequencies around 1 kHz and sounds with other frequencies, the voice can also be separated from other sounds. For example, it may be separated into sounds with frequencies of 0.5 kHz or higher and 2 kHz or lower and sounds with other frequencies. Further, for example, the frequency for sound separation may be changed according to the type of sound detected by the sound detection unit 11. For example, when the sound detection unit 11 detects a sound including a female voice, a sound with a higher frequency than when it detects a sound including a male voice may be separated as the voice. By changing the frequency for sound separation according to the type of sound detected by the sound detection unit 11, for example, the sound detected by the sound detection unit 11 can be separated into the voice and other sounds with high accuracy.

[0084] Hereinafter, it will be described assuming that the sound 21a is a voice and the sound 21b is a sound other than the voice.

[0085] After the sound separation unit 12 separates the sound 21 into the sound 21a that is a voice and the sound 21b that is a sound other than the voice, the sound determination unit 13 performs feature amount extraction on the sound 21a and determines whether or not the extracted feature amount is registered (step S03). Specifically, as shown in FIG. 4B, for example, the sound 21a is input to the generator 30 in which the learning result 43 is read, and the generator 30 outputs the data 24 indicating the presence or absence of registration, whereby it can be determined whether or not the feature amount extracted from the sound 21a is registered.

[0086] When the feature amount extracted from the sound 21a is registered, the processing unit 15 analyzes the instruction included in the sound 21a and generates an instruction signal, which is a signal representing the content of the instruction (steps S04 and S05). The analysis of the instruction can be performed using language processing such as morphological analysis. On the other hand, when the feature amount extracted from the sound 21a is not registered, the analysis of the instruction and the generation of the instruction signal are not performed (step S04).

[0087] In FIG. 4C, as a specific example of the process shown in step S05, a case where the instruction included in the sound 21a is "change the type of music" is shown. As shown in FIG. 4C, when the sound 21a including the instruction "change the type of music" is input to the processing unit 15, an instruction signal 25 representing the instruction "change the type of music" is output. The instruction signal 25 is output to the transmission / reception unit 16. In FIG. 4C, for example, the statement "change the type of music to xxxxx" is shown as "change the type of music To: xxxxx". The same applies to other figures.

[0088] Note that, for example, to include the instruction "change the type of music" in the sound 21a, a person with a registered voiceprint may utter words to the effect of "change the type of music". The sound detection unit 11 detects the sound including the words as the sound 21, and the sound separation unit 12 separates the voice included in the sound 21 as the sound 21a, so that the instruction "change the type of music" can be included in the sound 21a. Therefore, it can be said that the acoustic device 10 has a function of performing voice recognition.

[0089] Next, the processing unit 15 performs processing for canceling the sound 21b, which is a sound other than the voice separated by the sound separation unit 12 (step S06). For example, as shown in FIG. 5A, the sound 21b is input to the processing unit 15, and a sound 26 with the phase inverted from that of the sound 21b is output.

[0090] After that, the transmission / reception unit 16 synthesizes the sound 26, which is the sound processed by the processing unit 15, and the sound emitted by the information terminal 22, and outputs the synthesized sound to the sound output unit 17 (step S07). Here, the sound emitted by the information terminal 22 can be, for example, the music being played by the information terminal 22 when the information terminal 22 is playing music.

[0091] Also, when the processing unit 15 generates the command signal 25 and outputs it to the transmission / reception unit 16, that is, when the feature amount of the sound 21a, which is the voice separated by the sound separation unit 12, is registered, the transmission / reception unit 16 outputs the command signal 25 to the information terminal 22 (steps S08 and S09).

[0092] A specific example of steps S07 to S09 is shown in FIG. 5B. FIG. 5B shows an example in which the sound 26, which is the sound with the phase of the sound 21b inverted, the command signal 25 representing the command "change the type of music", and the sound 27 emitted from the information terminal 22 are input to the transmission / reception unit 16. The transmission / reception unit 16 synthesizes the sound 26 and the sound 27 and outputs the synthesized sound to the sound output unit 17. The sound input to the sound output unit 17 is emitted outside the sound device 10. The user of the sound device 10 can listen to the synthesized sound of the sound 21 detected by the sound detection unit 11, the sound 26 and the sound 27 output by the sound output unit 17, with the ear 23.

[0093] As described above, the sound 26 separates the sound 21b, which is a component such as noise included in the sound 21, and is, for example, a sound with the phase inverted. Therefore, the user of the sound device 10 can listen to, for example, a sound with the noise canceled.

[0094] Also, when the command signal 25 is input to the transmission / reception unit 16, the transmission / reception unit 16 outputs the command signal 25 to the information terminal 22. The information terminal 22 executes the command represented by the command signal 25. For example, when the information terminal 22 is playing music and the command signal 25 represents the command "change the type of music", the information terminal 22 can change the music it is playing to the specified one. The above is an example of the operation method of the sound device 10.

[0095] Only when the feature amount of sound such as voice separated by the sound separation unit 12 is registered, the processing unit 15 generates the command signal 25, so that the malfunction of the information terminal 22 can be suppressed more than when the command signal 25 is generated regardless of whether there is registration or not. For example, when registering the feature amount of the voice of the user of the information terminal 22 in the audio device 10, it is possible to suppress the occurrence of an operation unintended by the user of the information terminal 22 in response to the voice of a person other than the user of the information terminal 22.

[0096] In the operation method shown in FIG. 3 and the like, regardless of the content of the command represented by the command signal 25, the transmission / reception unit 16 outputs the command signal 25 to the information terminal 22, but one aspect of the present invention is not limited to this. Depending on the content of the command, the transmission / reception unit 16 may output the command signal 25 to an entity other than the information terminal 22.

[0097] FIG. 6 is a flowchart showing an example of an operation method when using the audio device 10, and is a modified example of the operation method shown in FIG. 3. The operation method shown in FIG. 6 is different from the operation method shown in FIG. 3 in that step S05 is replaced with step S05a and step S09 is replaced with step S09a.

[0098] In step S05a, the command included in the sound 21a, which is a voice, separated by the sound separation unit 12 is analyzed, and a command signal 25 representing the content of the command and the output destination of the command is generated. The output destination of the command can be determined according to, for example, the type of the command. Also, in step S09a, the transmission / reception unit 16 outputs the command signal 25 to a predetermined output destination.

[0099] Specific examples of step S07, step S08, and step S09a shown in FIG. 6 are shown in FIGS. 7A and 7B. FIG. 7A shows a case where the command signal 25 represents a command to "change the type of music". In this case, the transmission / reception unit 16 outputs the command signal 25 to the information terminal 22, and the information terminal 22 can change the music to be played to the specified one.

[0100] FIG. 7B shows an example in which the command signal 25 represents a command to "change the volume". In this case, the transmission / reception unit 16 outputs the command signal 25 to the sound output unit 17, and the sound output unit 17 can change the volume of the sound 27 emitted from the information terminal 22.

[0101] Note that the output destination of the command signal 25 may be made specifiable by, for example, the user of the audio device 10. For example, a person having a registered voiceprint may be able to specify the output destination of the command signal 25 by uttering words for specifying the output destination of the command signal 25.

[0102] In the operation method shown in FIG. 3 and the like, when the separated sound 21 separated by the sound separation unit 12 includes the voice sound 21a, the processing unit 15 does not perform the process of canceling the sound 21a even if the feature amount extracted from the sound 21a is not registered, but one aspect of the present invention is not limited to this. When the feature amount extracted from the sound 21a is not registered, the processing unit 15 may perform a process of canceling not only the non-voice sound 21b but also the voice sound 21a.

[0103] FIG. 8 is a flowchart showing an example of an operation method when the audio device 10 is in use, and is a modified example of the operation method shown in FIG. 3. The operation method shown in FIG. 8 is different from the operation method shown in FIG. 3 in that when the feature amount extracted from the sound 21a is not registered (step S04), step S06a is performed instead of step S06. FIG. 9 is a schematic diagram for explaining the details of step S06a.

[0104] In step S06a, the processing unit 15 performs a process of canceling all of the sound 21 detected by the sound detection unit 11. For example, as shown in FIG. 9, the sound 21 is input to the processing unit 15, and a sound having a phase inverted from that of the sound 21 is output as the sound 26.

[0105] Also, when the feature amount extracted from the sound 21a is not registered, the processing unit 15 may perform a process of reducing the volume of the sound 21a.

[0106] FIG. 10 is a flowchart showing an example of an operation method when the acoustic device 10 is in use, and is a modified example of the operation method shown in FIG. 8. The operation method shown in FIG. 10 is different from the operation method shown in FIG. 8 in that step S06a is replaced by step S06b.

[0107] FIG. 11 is a schematic diagram for explaining the details of step S06b. In step S06b, the processing unit 15 performs processing to reduce the magnitude of the sound 21a that is voice among the sounds 21 separated by the sound separation unit 12 and cancel the sound 21b that is a sound other than voice. For example, as shown in FIG. 11, the sounds 21a and 21b are input to the processing unit 15. Then, the processing unit 15 performs processing to invert the phase of the sound 21a and reduce the amplitude, and also performs processing to invert the phase of the sound 21b. The sound processed by the processing unit 15 is output as the sound 26.

[0108] As described above, by using the method shown in the present embodiment, malfunction of the information terminal 22 can be suppressed. In addition, since noise and the like can be canceled, the information terminal 22 can perform highly accurate speech recognition.

Explanation of Reference Numerals

[0109] 10: Acoustic device, 11: Sound detection unit, 12: Sound separation unit, 13: Sound determination unit, 14: Storage unit, 15: Processing unit, 16: Transmission / reception unit, 17: Sound output unit, 21: Sound, 21a: Sound, 21b: Sound, 22: Information terminal, 23: Ear, 24: Data, 25: Command signal, 26: Sound, 27: Sound, 30: Generator, 31: Sound data, 32: Label, 33: Learning result, 41: Sound data, 42: Label, 43: Learning result

Claims

[Claim 1] The device includes a sound detection unit, a sound separation unit, a sound determination unit, and a processing unit, The sound detection unit has a function of detecting a first sound, the sound separation unit has a function of separating the first sound into a second sound and a third sound, The sound determination unit has a function of registering a feature amount of a sound, the sound determination unit has a function of determining whether or not the feature amount of the second sound is registered by using a machine learning model; the processing unit has a function of analyzing a command included in the second sound and generating a signal representing the content of the command when the feature amount of the second sound is the registered feature amount of the second sound; The audio device, wherein the processing unit has a function of generating a fourth sound by performing processing on the third sound to cancel the third sound.

Citation Information

Patent Citations

  • Harmful customer detection system, its method and harmful customer detection program

    JP2010113167A

  • Content reproduction device

    JP2013213893A

  • Method and system for validating personalized account identifiers using biometric authentication and self-learning algorithms

    JP2014191823A

  • Voice processing device, voice processing method, and program

    JP2016075740A

  • Music playback device and music playback program

    JP2016130751A