Audio signal acquisition system and audio signal acquisition method

The audio signal acquisition system with a body-worn housing, sensor, and microphone effectively distinguishes and processes the user's voice, addressing misinterpretation issues in existing voice recognition systems.

JP2026068550APending Publication Date: 2026-04-22JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
JVC KENWOOD CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing voice recognition systems fail to distinguish between the voice of a user wearing earphones and that of others nearby, leading to misinterpretation of external voices as commands.

Method used

An audio signal acquisition system with a housing in contact with the user's body, incorporating a vibration detection sensor and microphone, controlled by a unit to enable speech recognition based on sensor detection, distinguishing user voice from external noise.

Benefits of technology

Accurately identifies and processes the user's voice while filtering out external noise, preventing misinterpretation of nearby voices as commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068550000001_ABST
    Figure 2026068550000001_ABST
Patent Text Reader

Abstract

To appropriately identify and process speech spoken by a user wearing earphones or similar devices. [Solution] The voice signal acquisition system 1 comprises a housing 11 used in close contact with the human body, a vibration detection sensor 13 positioned in the housing 11 at a location in close contact with the human body, a microphone 17 positioned in the housing 11 to collect sound, and a control unit 50 that controls the system to enable voice recognition processing of the sound collected by the microphone 17 based on the detection results of the sensor 13.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio signal acquisition system and an audio signal acquisition method.

Background Art

[0002] Devices that use AI (Artificial Intelligence) for the voice recognized by voice recognition processing are becoming widespread. Such devices may use a microphone arranged in a small device connected to, for example, a smartphone. In addition, a microphone is generally attached to TWS (True Wireless Stereo) earphones to pick up the speech of the user wearing them.

[0003] Techniques for reducing noise and picking up the voice of a user wearing headphones have been disclosed (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technique described in Patent Document 1 reduces the noise of the sound to be picked up. The technique described in Patent Document 1 does not determine whether the sound to be picked up is the voice of a person wearing earphones or the voice of a person around. For this reason, the voice of a person other than the user wearing earphones speaking near the earphones may be recognized as an audio command for various processes performed via the earphones.

[0006] This disclosure is made in view of the above, and aims to appropriately identify the voice spoken by a user wearing earphones or the like, and to perform speech recognition processing. [Means for solving the problem]

[0007] To solve the above-mentioned problems and achieve the objective, the audio signal acquisition system according to this disclosure comprises a housing used in close contact with the human body, a vibration detection sensor positioned in the housing at a location in close contact with the human body, a microphone positioned in the housing to collect sound, and a control unit that controls the system to enable speech recognition processing of the sound collected by the microphone based on the detection results of the sensor.

[0008] The audio signal acquisition method according to this disclosure is an audio signal acquisition method executed by a control unit that controls an audio signal acquisition system comprising a housing used in close contact with the human body, a sensor positioned in the housing at a location in close contact with the human body, and a microphone positioned in the housing for collecting sound, wherein the control unit controls the system to enable speech recognition processing of the sound collected by the microphone based on the detection result of the sensor. [Effects of the Invention]

[0009] According to this disclosure, the system can appropriately identify the voice spoken by a user wearing earphones or similar devices and perform speech recognition processing. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 is a schematic diagram showing an example of the configuration of an audio signal acquisition system according to an embodiment. [Figure 2] Figure 2 is a schematic front view of the earphone according to the embodiment. [Figure 3] Figure 3 is a block diagram showing an example of the configuration of an earphone according to the embodiment. [Figure 4] Figure 4 is a block diagram showing an example configuration of an information terminal device according to an embodiment. [Figure 5]Figure 5 is a flowchart showing the processing flow in the audio signal acquisition system according to the embodiment. [Modes for carrying out the invention]

[0011] Embodiments of the audio signal acquisition system 1 according to this disclosure will be described in detail below with reference to the attached drawings. However, the present invention is not limited to the following embodiments.

[0012] [Embodiment] <Audio signal acquisition system> Figure 1 is a schematic diagram showing an example configuration of an audio signal acquisition system according to an embodiment. The audio signal acquisition system 1 includes an earphone 10, which is an example of an audio signal acquisition device, and an information terminal device 40. The earphone 10 and the information terminal device 40 are connected by wire or wireless communication. In the following description, the earphone 10 and the information terminal device 40 will be described as being connected by wireless communication. The audio signal acquisition system 1 is not limited to a system consisting of an earphone 10 and an information terminal device 40, but may also be headphones equipped with functions equivalent to those necessary for carrying out the present invention in the information terminal device 40 and earphone 10, or glasses-type terminals such as AR (Augmented Reality) goggles. The audio signal acquisition system 1 controls whether or not to perform speech recognition processing on the sound picked up by the microphone 17 located on the earphone 10, which is controlled by the information terminal device 40.

[0013] <Earphones (Audio Signal Acquisition Device)> The earphone 10 has a housing 11 that is used in close contact with the human body. The human body refers to the user wearing the earphone 10, specifically the user's ear. In the following description, the human body will be referred to as the user. The earphone 10 may be, for example, a wireless earphone, a wired earphone, or headphones. The form of the earphone 10 may be any form, for example, an open-ear type or an in-ear type. Figures 1 and 2 show an example of a TWS (True Wireless Stereo) open-ear type earphone. The earphone 10 has a left earphone 10L for the left ear, which is worn on the left ear, and a right earphone 10R for the right ear, which is worn on the right ear. The left earphone 10L outputs the left channel data of the audio signal. The right earphone 10R outputs the right channel data of the audio signal. In the following description, when it is not necessary to distinguish between the left earphone 10L and the right earphone 10R, they will be described as earphone 10.

[0014] The earphone 10 will be described using Figures 1 to 3. Figure 1 shows a schematic side view of the earphone 10 according to the embodiment, and is a view of the earphone 10 as seen from the left or right direction of the user when the earphone 10 is worn by the user. In Figure 1, the left earphone 10L is a view of the left side of the user when it is worn on the user's left ear, and the right earphone 10R is a view of the right side of the user when it is worn on the user's right ear. Figure 2 is a schematic front view of the earphone according to the embodiment, and is a view of the earphone 10 as seen from the front of the user when it is worn by the user. In Figure 2, the right earphone 10R is shown, and the left earphone 10L has a similar configuration. Figure 3 is a block diagram showing an example of the configuration of the earphone according to the embodiment. The earphone 10 acquires various audio data, including music data, from, for example, an information terminal device 40. The earphone 10 includes a sensor 13, an operating unit 15, a camera 16, a microphone 17, an amplification unit 18, an audio output unit 19, a communication unit 27, a battery 28, a power supply unit 29, and a control unit 30.

[0015] The housing 11 of the earphone 10 is worn on the user's ear. The housing 11 of the earphone 10 is used in close contact with the user's ear or the area around the ear. The housing 11 defines the external shape of the earphone 10. The housing 11 comprises a main body 11a, a first ear hook 11b, and a second ear hook 11c. The main body 11a is the part that is positioned inside the user's auricle when the user wears the earphone 10. The first ear hook 11b is the part that is positioned between the helix of the auricle and the head when the user wears the earphone 10. The second ear hook 11c is the part that is positioned behind the auricle when the user wears the earphone 10. The main body 11a, the first ear hook 11b, and the second ear hook 11c are integrally formed. The housing 11 houses a sensor 13, an operation unit 15, a camera 16, a microphone 17, an amplification unit 18, an audio output unit 19, a communication unit 27, a battery 28, a power supply unit 29, and a control unit 30. The arrangement of the sensor 13, operation unit 15, camera 16, microphone 17, amplification unit 18, audio output unit 19, communication unit 27, battery 28, power supply unit 29, and control unit 30 in the housing 11 is an example.

[0016] Sensor 13 is positioned in the housing 11 of the earphone 10 in a location that is in close contact with the human body, such as on or around the user's ear. Sensor 13 is a sensor that detects vibrations caused by speech from a user wearing the earphone 10. Sensor 13 is, for example, an acceleration sensor. Sensor 13 detects vibrations caused by speech from a user wearing the earphone 10. Sensor 13 is positioned in close contact with the location where vibrations are generated by the user's speech. If the earphone 10 is an over-ear type earphone as shown in Figures 1 and 2, sensor 13 is positioned on the over-ear portion, etc. If the earphone 10 is a canal type earphone, sensor 13 is positioned in close contact with the side of the tragus, etc. In this embodiment, sensor 13 is positioned on the first over-ear portion 11b. Sensor 13 outputs sensor data, which is the detection result, to the detection unit 31. Sensor 13 is positioned on at least one of the left earphone 10L and the right earphone 10R. The sensor 13 does not perform detection when the earphones 10 are stored and charging in an earphone case (not shown), or when the power of the earphones 10 is turned off.

[0017] The operation unit 15 is a touch sensor disposed in the main body unit 11a. The operation unit 15 can receive various operations such as reproduction and stop of music data in the earphone 10. The operation unit 15 can receive various operations related to voice recognition, video recognition, etc. in the earphone 10. The operation unit 15 outputs an operation signal indicating the received various operations to the operation control unit 32.

[0018] The camera 16 photographs the periphery of the earphone 10. The photographing of the camera 16 is controlled by the photographing control unit 33. The camera 16 outputs the photographed video to the photographing control unit 33. The camera 16 is disposed in the main body unit 11a so as to face the front of the user when the user wears the earphone 10 on the ear, that is, so that the front of the user becomes the photographing direction. The camera 16 is disposed in at least one of the left earphone 10L and the right earphone 10R.

[0019] The microphone 17 picks up the sound around the earphone 10. The microphone 17 can pick up, for example, a voice that is a voice command for the earphone 10. The microphone 17 outputs the picked-up voice to the voice input control unit 34. The microphone 17 is disposed in the main body unit 11a.

[0020] The amplification unit 18 amplifies the voice output from the voice output unit 19. The amplification unit 18, for example, performs D / A conversion and amplification on the channel data of the voice including the voice data acquired from the information terminal device 40. The voice output unit 19 is disposed in the main body unit 11a.

[0021] The voice output unit 19 outputs the voice to be viewed in the user's ear. The voice output unit 19 outputs, for example, a voice based on a voice signal. The voice output control unit 35 performs control to output, for example, a voice signal indicating the recognition result of the object recognition process in the information terminal device 40 described later as a voice. The voice output unit 19 causes the voice signal amplified by the amplification unit 18 to be output from the voice output unit 19. The voice output unit 19 is disposed in the main body unit 11a.

[0022] The communication unit 27 is a communication unit. The communication unit 27 is capable of wireless communication including, for example, Bluetooth®, Wi-Fi®, or NFMI (Near Field Magnetic Induction). In this embodiment, the communication unit 27 can be paired with the information terminal device 40 via Bluetooth connection. The communication unit 27 is located, for example, on the second ear hook portion 11c.

[0023] The communication unit 27 is connected to the information terminal device 40 so as to be able to send and receive data. The communication unit 27 receives data, for example, music data, from the information terminal device 40. In this embodiment, the communication unit 27 transmits to the information terminal device 40 sensor data indicating the detection result of the sensor 13 detected by the detection unit 31, audio data based on the audio signal picked up by the microphone 17 acquired by the audio input control unit 34, and video captured by the camera 16 acquired by the shooting control unit 33. In this embodiment, the communication unit 27 receives text data indicating the recognition result of the object recognition process from the information terminal device 40.

[0024] The battery 28 supplies power for the earphone 10 to operate. The battery 28 is a rechargeable battery that can be repeatedly charged and discharged. The battery 28 is, for example, a nickel-metal hydride rechargeable battery, a lithium-ion rechargeable battery, or a lithium polymer rechargeable battery. The battery 28 is built into the earphone 10. The charging and discharging of the battery 28 is controlled by the power control unit 39. The battery 28 is located, for example, in the second ear hook portion 11c.

[0025] The power supply unit 29 can be connected to the power supply of an earphone case (not shown) and supplies power to the battery 28. The power supply unit 29 is connected to the power control unit 39. The power supply unit 29 may also be able to supply power from a DC power source such as a mobile battery or from a contactless charging device.

[0026] <Earphone control unit> The control unit 30 is a processing unit composed of, for example, a CPU (Central Processing Unit). The control unit 30 loads a program stored in a storage unit (not shown) into memory and executes the instructions contained in the program. The control unit 30 includes an internal memory (not shown) used for temporary data storage, etc. The control unit 30 has a detection unit 31, an operation control unit 32, a shooting control unit 33, an audio input control unit 34, an audio output control unit 35, a communication control unit 37, and a power supply control unit 39.

[0027] The detection unit 31 detects vibrations based on sensor data, which is the detection result of the sensor 13. The detection unit 31 detects vibrations based on speech uttered by the user wearing the earphones 10, based on sensor data, which is the detection result of the sensor 13. The detection unit 31 detects vibrations caused by speech uttered by the user wearing the earphones 10, based on sensor data, which is the detection result of the sensor 13. In this embodiment, the detection unit 31 transmits the sensor data to the information terminal device 40.

[0028] The vibrations detected by the detection unit 31, which are generated by the speech of the user wearing the earphones 10, include vibrations based on the movement of the user's mouth when they speak, or vibrations of the vocal cords caused by the user's speech. The sensor 13 detects these vibrations, which are transmitted through the user's head and other parts of the body.

[0029] The detection unit 31 may perform filtering to exclude vibrations caused by head movements of the user wearing the sensor 13, or vibrations caused by walking, from detection. In this case, the detection unit 31 may, for example, exclude low-frequency vibrations caused by the user's head movements or walking.

[0030] The detection unit 31 may work in conjunction with the user's information terminal device 40 and detect vibrations caused by speech using a vibration model that has been trained to recognize vibrations caused by speech.

[0031] The operation control unit 32 acquires operation signals from the operation unit 15 in response to operations performed on the operation unit 15. The operation control unit 32 outputs control signals corresponding to the acquired operation signals to each part of the earphone 10 and the information terminal device 40.

[0032] The shooting control unit 33 controls the shooting by the camera 16. The shooting control unit 33 acquires the video captured by the camera 16. In this embodiment, the shooting control unit 33 uses the communication unit 27 via the communication control unit 37 to transmit the captured video to the information terminal device 40.

[0033] The audio input control unit 34 acquires the audio signal picked up by the microphone 17. The audio input control unit 34 performs A / D conversion on the audio signal picked up by the microphone 17 and acquires it as audio data.

[0034] The audio output control unit 35 controls the output of audio from the audio output unit 19. For example, the audio output control unit 35 controls the output of audio data such as music data from the audio output unit 19. For example, the audio output control unit 35 controls the output of audio data indicating the recognition result of the object recognition process in the information terminal device 40, which will be described later, from the audio output unit 19. The audio output control unit 35 outputs the audio data acquired from the information terminal device 40 to the amplification unit 18.

[0035] The communication control unit 37 communicates wirelessly with the information terminal device 40 by controlling the communication unit 27. For example, the communication control unit 37 controls the information terminal device 40 to receive data including music data. In this embodiment, the communication control unit 37 transmits sensor data indicating the detection result of the sensor 13, audio data based on the audio signal picked up by the microphone 17, and video footage captured by the camera 16 to the information terminal device 40. In this embodiment, the communication control unit 37 controls the information terminal device 40 to receive text data indicating the recognition result of the object recognition process.

[0036] The power control unit 39 is a charging circuit that charges the battery 28. The power control unit 39 is connected to the battery 28 and the power supply unit 29. The power control unit 39 supplies power from the power supply unit 29 to the battery 28.

[0037] <Information terminal device> The information terminal device 40 will be described using Figure 4. Figure 4 is a block diagram showing an example configuration of the information terminal device according to the embodiment. The information terminal device 40 is a portable electronic device used by the user of the earphone 10, such as a smartphone or a tablet device. The information terminal device 40 detects vibrations caused by the user's speech using the sensor 13 of the earphone 10, and controls the microphone 17 located on the earphone 10 to enable speech recognition processing of the sound picked up by the microphone. The information terminal device 40 comprises a communication unit 41 and a control unit 50.

[0038] The communication unit 41 is a communication unit. The communication unit 41 is connected to the earphone 10 so as to be able to send and receive data. For example, the communication unit 41 transmits data including music data to the earphone 10. In this embodiment, the communication unit 41 receives sensor data indicating the detection result of the sensor 13, audio data based on the audio signal picked up by the microphone 17, and video captured by the camera 16 from the earphone 10. In this embodiment, the communication unit 41 transmits audio data indicating the recognition result of the object recognition process to the earphone 10. The communication unit 41 is capable of short-range wireless communication, including Bluetooth, for example. In this embodiment, the communication unit 41 can be paired with the earphone 10 via Bluetooth connection.

[0039] The communication unit 41 is capable of communicating with an external AI server. The communication unit 41 communicates with the external AI server using wireless communication such as Wi-Fi.

[0040] The video acquired by the information terminal device 40 from the earphone 10 can be used for various purposes. For example, the video may be recorded. For example, an object may be recognized from the video, and audio data indicating the recognition result may be fed back to the earphone 10. The object recognition process may use an external AI server provided by the AI ​​generation service provider. In this case, for example, the video acquired from the earphone 10 is sent to the external AI server along with instructions such as prompts corresponding to audio commands to perform the object recognition process. The AI ​​server performs the object recognition process and sends text data indicating the object recognition result to the information terminal device 40. The information terminal device 40 receives text data, audio data, etc. indicating the object recognition result from the AI ​​server.

[0041] The user's spoken voice, which is the audio acquired by the information terminal device 40 from the earphone 10, can be used for various purposes. For example, the voice may be recorded. For example, if the recognized voice is a predetermined phrase that is an audio command to perform object recognition processing, it may be used as a trigger to start processing that recognizes an object from the video acquired from the earphone 10 and displays the recognition result.

[0042] The prescribed phrases are, for example, "What is that?" or "Tell me its name." The instructions corresponding to these prescribed voice commands are, for example, instructions to recognize an object from the video, and to search for and output the name of the recognized object or information about the object.

[0043] The control unit 50 is an arithmetic processing unit, for example, composed of a CPU. The control unit 50 loads a program stored in a storage unit (not shown) into memory and executes the instructions contained in the program. The control unit 50 includes an internal memory (not shown) which is used for temporary data storage, etc. In this embodiment, the control unit 50 implements the function of a control unit for the voice signal acquisition system 1. Based on the detection results of the sensor 13, the control unit 50 controls the system to enable voice recognition processing for the voice picked up by the microphone 17. The control unit 50 has a communication control unit 51 and a permission / failure control unit 52.

[0044] The communication control unit 51 communicates wirelessly with the earphone 10 by controlling the communication unit 41. For example, the communication control unit 51 controls the earphone 10 to transmit data including music data. In this embodiment, the communication unit 41 controls the earphone 10 to receive sensor data indicating the detection result of the sensor 13, audio data based on the audio signal picked up by the microphone 17, and video footage captured by the camera 16. In this embodiment, the communication unit 41 controls the earphone 10 to transmit audio data indicating the recognition result of the object recognition process.

[0045] The communication control unit 51 communicates wirelessly with the AI ​​server by controlling the communication unit 41. The communication control unit 51 controls the AI ​​server to send instructions corresponding to video and audio commands. The communication unit 41 controls the AI ​​server to receive text data from the AI ​​server indicating the object recognition results based on the video.

[0046] The permission / recognition control unit 52 controls whether or not to perform speech recognition processing on the sound picked up by the microphone 17 placed on the earphone 10, based on vibrations caused by the speech of the user wearing the earphone 10. The permission / recognition control unit 52 controls to enable speech recognition processing on the sound picked up by the microphone 17, based on the detection result of the sensor 13 placed on the earphone 10. If vibrations caused by the speech of the user wearing the earphone 10 are detected, the permission / recognition control unit 52 controls to enable speech recognition processing on the sound picked up by the microphone 17 placed on the earphone 10. If the microphone 17 placed on the earphone 10 detects sound and the sensor 13 detects vibrations caused by the speech of the user wearing the earphone 10, the permission / recognition control unit 52 controls to enable speech recognition processing on the sound picked up by the microphone 17.

[0047] The permission control unit 52 may control the audio picked up by the microphone 17 placed in the earphone 10 to enable speech recognition processing for the audio during the period when the sensor 13 detects vibration. As a result, only the audio during the period when vibration caused by the user's speech is detected is subject to speech recognition processing.

[0048] The permission control unit 52 may control the system to enable speech recognition processing for speech during periods when the sound picked up by the microphone 17 located on the earphone 10 and the vibrations detected by the sensor 13 are synchronized. As a result, only speech from the period when the voice of the user wearing the earphone 10 and the vibrations caused by speech are detected in sync will be subject to speech recognition processing.

[0049] Enabling speech recognition processing means, for example, starting or executing speech recognition processing. When a predetermined voice command, which is a trigger for starting object recognition processing, is recognized by the speech recognition processing, the video acquired from the earphone 10 is sent to an external AI server along with instructions corresponding to the recognized voice command. The AI ​​server performs object recognition processing based on the video acquired from the earphone 10 and the instructions, and sends text data indicating the recognition result to the information terminal device 40. The information terminal device 40 acquires the text data indicating the recognition result from the AI ​​server, converts it into audio data, and sends the audio data indicating the recognition result to the earphone 10. The earphone 10 outputs the recognition result as audio from the audio output unit 19.

[0050] Sensor 13 may be placed in both the left earphone 10L and the right earphone 10R. In such a case, if the detection results of sensors 13L and 13R, which are located in the housings 11L of the left earphone 10L and 11R of the right earphone 10R, are synchronized, the control unit 52 may control the system to enable speech recognition processing for the sound picked up by the microphone 17 located in the earphone 10 based on the detection results of sensors 13L and 13R. This allows for more appropriate detection of vibrations caused by the speech of the user wearing the earphone 10.

[0051] The permission control unit 52 controls the system to disable speech recognition processing for the sound picked up by the microphone 17, based on the detection results of the sensor 13 located on the earphone 10. The permission control unit 52 controls the system to disable speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 if no vibrations caused by the speech of the user wearing the earphone 10 are detected. The permission control unit 52 controls the system to disable speech recognition processing for the sound picked up by the microphone 17 if the microphone 17 located on the earphone 10 detects sound, and the sensor 13 does not detect vibrations caused by the speech of the user wearing the earphone 10.

[0052] Disabling speech recognition processing means, for example, stopping the speech recognition process. When speech recognition processing is disabled, the object recognition process is stopped because the predetermined sound that triggers the initiation of object recognition processing is no longer recognized.

[0053] <Audio signal acquisition method> Next, information processing in the audio signal acquisition system 1 will be explained using Figure 5. Figure 5 is a flowchart showing the processing flow in the audio signal acquisition system according to the embodiment. For example, when the earphone 10 is activated and the voice recognition function or object recognition function is turned ON, the processing shown in the flowchart of Figure 5 is executed. During the execution of the processing shown in the flowchart of Figure 5, sensor data, which is the detection result of the sensor 13, is transmitted to the information terminal device 40. During the execution of the processing shown in the flowchart of Figure 5, sound is picked up by the microphone 17.

[0054] The control unit 50 determines, based on the feasibility control unit 52, whether or not sound has been detected (step S101). The control unit 50 determines, based on the feasibility control unit 52, whether or not sound picked up by the microphone 17 placed on the earphone 10 has been detected. The process in step S101 may also be to determine, based on the feasibility control unit 52, whether or not sound corresponding to a voice command has been detected. If the control unit 50 determines, based on the feasibility control unit 52, that sound has been detected (step S101; Yes), the control unit 50 proceeds to step S102. If the control unit 50 does not determine, based on the feasibility control unit 52, that sound has been detected (step S101; No), the control unit 50 proceeds to step S106.

[0055] If the control unit 50 determines that sound has been detected (step S101; Yes), the control unit 50 uses the feasibility control unit 52 to determine whether or not vibration has been detected (step S102). The control unit 50 uses the feasibility control unit 52 to determine whether or not vibration caused by speech from the user wearing the earphones 10 has been detected, based on the detection results of the sensor 13 placed on the earphones 10. Since the vibration detected in step S102 is caused by speech from the user wearing the earphones 10, the control unit 50 determines whether or not vibration was detected during a period that coincides with the period in which sound was detected. If the control unit 50 uses the feasibility control unit 52 to determine that vibration has been detected (step S102; Yes), the control unit 50 proceeds to step S103. If the control unit 50 does not use the feasibility control unit 52 to determine that vibration has been detected (step S102; No), the control unit 50 proceeds to step S104.

[0056] If it is determined that vibration has been detected (step S102; Yes), the control unit 50 enables speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 using the feasibility control unit 52 (step S103). The sound for which speech recognition processing is enabled in step S103 is the sound detected in step S101. The control unit 50 proceeds to step S106.

[0057] In step S103, once speech recognition processing is enabled, the control unit 50 performs speech recognition processing on the sound picked up by the microphone 17. If the control unit 50 recognizes a predetermined voice command, which is a trigger that starts object recognition processing, it performs object recognition processing using an external AI server.

[0058] If the control unit 50 does not determine that vibration has been detected (step S102; No), the control unit 50, in turn, determines whether or not speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 is enabled, as determined by the feasibility control unit 52 (step S104; Yes). If the control unit 50 determines, as determined by the feasibility control unit 52, that speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 is enabled (step S104; No), the control unit 50 proceeds to step S105. If the control unit 50 does not determine, as determined by the feasibility control unit 52, that speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 is enabled (step S104; No), the control unit 50 proceeds to step S106.

[0059] If the control unit 50 determines that speech recognition processing is enabled for the sound picked up by the microphone 17 located on the earphone 10 (step S104; Yes), the control unit 50 disables speech recognition processing for the sound picked up by the microphone 17 located on the earphone 10 (step S105). The control unit 50 proceeds to step S106.

[0060] The control unit 50 determines whether to terminate (step S106). The control unit 50 determines to terminate, for example, when an operation to terminate the voice recognition processing or object recognition processing in the earphone 10 is detected, or when an operation to turn off the power of the earphone 10 is detected. If the control unit 50 determines to terminate (step S106; Yes), it terminates the process in this flowchart. If the control unit 50 does not determine to terminate (step S106; No), it executes the process in step S101 again.

[0061] If the microphone 17 located on the earphone 10 detects sound (step S101; Yes), and the sensor 13 determines that it has detected vibrations caused by the user's speech (step S102; Yes), then speech recognition processing is enabled in step S103.

[0062] Step S102 may be set to Yes if the microphone 17 placed in the earphone 10 detects sound (step S101; Yes) and the sensor 13 continuously detects vibrations caused by the user's speech for a predetermined period of time.

[0063] If the result is Yes in step S101 and Yes in step S102, enabling speech recognition processing in step S103, and then No in step S106, and in subsequent loops, if the result is Yes in step S101 and No in step S102, and then Yes in step S104, then speech recognition processing is disabled in step S105. As a result, speech recognition processing is enabled only during the period when the sound picked up by the microphone 17 and the vibration detected by the sensor 13 are synchronized.

[0064] <Effects> As described above, this embodiment enables speech recognition processing of the sound picked up by the microphone 17 based on the detection results of the sensor 13 placed on the earphone 10. According to this embodiment, the sensor 13 can be used to identify the voice spoken by the user wearing the earphone 10, and speech processing can be performed. This embodiment can prevent the recognition of voices spoken by someone other than the user wearing the earphone 10 in the vicinity of the earphone 10 as voice commands for various processes performed through the earphone 10. According to this embodiment, the sensor 13 can be used to appropriately identify the voice spoken by the user wearing the earphone 10.

[0065] In this embodiment, vibrations caused by speech from a user wearing the earphones 10 can be detected. According to this embodiment, the sensor 13 can be used to more appropriately identify the voice spoken by a user wearing the earphones 10 or the like.

[0066] In this embodiment, speech recognition processing can be performed on the audio picked up by the microphone 17 during the period when the sensor 13 detects vibration. According to this embodiment, speech recognition processing can be disabled for audio picked up by the microphone 17 during the period when the sensor 13 does not detect vibration. According to this embodiment, it is possible to identify the voice spoken by a user wearing earphones 10 or the like and perform speech recognition processing.

[0067] In this embodiment, it is possible to perform speech recognition processing on speech during the period when the sound picked up by the microphone 17 and the vibration detected by the sensor 13 are synchronized. According to this embodiment, it is possible to identify the speech spoken by a user wearing earphones 10 or the like and perform speech recognition processing.

[0068] In this embodiment, if the detection results of the sensors 13 located in the housings 11 of the left earphone 10L and the right earphone 10R are synchronized, speech recognition processing can be performed on the sound picked up by the microphone 17. Synchronization of the detection results of the sensors 13 located in the housings 11 of the left earphone 10L and the right earphone 10R means that the detection results of the sensors 13 located in the housings 11 of the left earphone 10L and the right earphone 10R detect similar vibrations during the same period. According to this embodiment, it is possible to identify the voice spoken by a user wearing the earphones 10, etc., and perform speech recognition processing.

[0069] The components of the illustrated audio signal acquisition system are functionally conceptual and do not necessarily have to be physically configured as shown. In other words, the specific form of each device is not limited to that shown, and all or part of them may be functionally or physically distributed or integrated in any unit depending on the processing load and usage conditions of each device.

[0070] The configuration of the audio signal acquisition system is implemented, for example, as software, such as a program loaded into memory. In the above embodiment, these were described as functional blocks implemented through the cooperation of hardware or software. That is, these functional blocks can be implemented in various forms using hardware alone, software alone, or a combination thereof.

[0071] The components described above include those that are easily conceivable by those skilled in the art, and those that are substantially identical. Furthermore, the components described above can be combined as appropriate. In addition, various omissions, substitutions, or modifications of the components are possible without departing from the spirit of the present invention.

[0072] In the above embodiment, the operation unit 15 and the operation control unit 32 are not essential components.

[0073] In the above description, the earphone 10 was described as having a left earphone 10L and a right earphone 10R capable of communicating with the information terminal device 40, but it is not limited to this. The earphone 10 may also be a master unit in which the left earphone 10L or the right earphone 10R can communicate with the portable electronic device 100, and a slave unit in which the right earphone 10R or the left earphone 10L can communicate with the left earphone 10L or the right earphone 10R.

[0074] In the above, the control unit 50 of the information terminal device 40 implements a function to enable speech recognition processing for the sound picked up by the microphone 17 placed on the earphone 10 using the enable / disable control unit 52 of the control unit 50, but it is not limited to this. For example, the audio input control unit 34 of the control unit 30 of the earphone 10 may implement a function equivalent to the enable / disable control unit 52. In this case, the audio input control unit 34 of the control unit 30 of the earphone 10 recognizes a predetermined voice command, which is a trigger to start object recognition processing by speech recognition processing. When the voice input control unit 34 recognizes a voice command, the control unit 30 of the earphone 10 sends the video captured by the camera 16 along with the voice text corresponding to the recognized voice command to an external AI server. The AI ​​server performs object recognition processing based on the video acquired from the earphone 10 and the instructions corresponding to the voice command, and sends the recognition result to the earphone 10. The earphone 10 acquires the recognition result from the AI ​​server, converts it into an audio signal, and outputs the recognition result as audio from the audio output unit 19. [Explanation of Symbols]

[0075] 1. Audio signal acquisition system 10. Earphones (Audio signal acquisition device) 11 cabinets 13 sensors 15 Control section 16 cameras 17 Microphone 18 Amplification section 19 Audio output section 27 Communications Department 28 batteries 29 Power supply section 30 Control Unit 31 Detection unit 32 Operation Control Unit 33. Image capture control unit 34. Voice Input Control Unit 35 Audio Output Control Unit 37 Communication Control Unit 39 Power supply control unit 40 Information terminal devices 41 Communications Department 50 Control Unit 51 Communication Control Unit 52 Approval / Rejection Control Unit

Claims

1. A housing that is used in close contact with the human body, A vibration-detecting sensor is positioned in the aforementioned housing at a location in close contact with the human body, A microphone for picking up sound is placed in the aforementioned housing, A control unit that controls the microphone to enable speech recognition processing based on the detection results of the aforementioned sensor, An audio signal acquisition system equipped with the following features.

2. The sensor is positioned in close contact with the area of ​​the human body where vibrations occur due to speech. The audio signal acquisition system according to claim 1.

3. The control unit controls the sound picked up by the microphone to enable speech recognition processing for the sound during the period in which the sensor detects vibration. The audio signal acquisition system according to claim 1.

4. The control unit controls the system to enable speech recognition processing for the speech during the period when the speech picked up by the microphone and the vibration detected by the sensor are synchronized. The audio signal acquisition system according to claim 1.

5. The aforementioned housing consists of multiple housings used in close contact with the human body. The sensor is placed in each of the plurality of housings, The control unit, when the detection results of the sensors arranged in the plurality of housings are synchronized, controls the system to enable speech recognition processing of the sound picked up by the microphone based on the detection results of the sensors. The audio signal acquisition system according to claim 1.

6. The aforementioned housing is the housing for an earphone. The audio signal acquisition system according to claim 1.

7. A method for acquiring an audio signal, performed by a control unit that controls an audio signal acquisition system comprising a housing used in close contact with the human body, a sensor positioned in the housing at a location in close contact with the human body, and a microphone positioned in the housing for collecting sound, wherein the control unit controls the system, Based on the detection results of the aforementioned sensor, the system controls the microphone to enable speech recognition processing of the sound it has picked up. Method for acquiring audio signals.

Citation Information

Patent Citations

  • Voice sensing using multiple microphones

    JP2020102867A