Voice acquisition device and voice acquisition method
The voice acquisition system effectively addresses the challenge of identifying and extracting voice data from a moving sound source in noisy environments by using positional tracking and machine learning to isolate and extract voice data accurately from multiple voices.
Patent Information
- Application Number
- JP2025076077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-01
- Publication Date
- 2025-09-22
- Estimated Expiration
- 2045-05-01
AI Technical Summary
Existing systems struggle to accurately identify and acquire voice data from a moving sound source in a noisy environment, particularly when multiple voices are emitted almost simultaneously, as they require fixed RFID tags at seating positions.
A voice acquisition system that utilizes a tag positioning unit to measure the position of a movable wireless tag with a sound source, a sound source positioning unit to measure sound source positions based on waveform data, and a voice acquisition unit to extract voice data from the tag's position, combined with learning devices that perform machine learning using simulated noisy environments to enhance voice extraction.
Enables the robust acquisition of voice data from a predetermined sound source in a noisy environment, even when multiple voices are present, by leveraging machine learning and positional tracking to isolate and extract voice data accurately.
Smart Images

Figure 0007742969000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice acquisition device. and How to acquire audio By law Regarding. [Background technology]
[0002] A radio frequency identification (RFID) tag and microphone array for identifying a person speaking during a conference call in a conference room is disclosed in Patent Document 1. In Patent Document 1, a controller identifies a person who is speaking over the phone to a remote participant in a conference call attended by multiple people. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Patent No. 7,995,731 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in Patent Document 1, an RFID tag must be fixedly placed in advance at the seating position of each person (sound source) seated in the conference room. Therefore, if a person acting as a sound source moves their seating position, it is not possible to obtain the voice emitted by that person from the waveform data of multiple voices simultaneously acquired using a microphone array.
[0005] This problem is not limited to conference calls in a conference room. For example, a noisy environment (hereinafter referred to as a "noisy environment") containing multiple voices emitted almost simultaneously by multiple people may occur in a real space. In such a case, there is a problem in that it is not possible to obtain, from the noise waveform data, a voice emitted by a predetermined sound source (e.g., a predetermined person or a predetermined robot) that can move through the noisy environment.
[0006] In view of the above circumstances, the present invention aims to provide a speech acquisition device, a learning device, a speech acquisition method, and a learning method that are capable of acquiring speech emitted from a predetermined sound source that can move in a noisy environment from noise waveform data. [Means for solving the problem]
[0007] One aspect of the present invention is a voice acquisition device that includes a tag positioning unit that measures the position of a wireless tag that is movable together with a predetermined sound source among the sound sources of each sound in noise that includes multiple sounds; a sound source positioning unit that acquires waveform data of the noise from a microphone array and measures the position of the sound source of each sound based on the acquired waveform data of the noise; and a voice acquisition unit that acquires waveform data of the sound emitted from the sound source present at the position of the wireless tag from the waveform data of the noise based on a match between the position of the wireless tag and the position of the sound source of each sound.
[0008] One aspect of the present invention is a learning device that includes a learning unit that performs learning of a learning model using noise waveform data including multiple sounds emitted from each sound source in a simulated noisy environment and the positions of predetermined sound sources in the simulated noisy environment as learning data, and using feature quantities of the sounds emitted from the predetermined sound sources as correct labels.
[0009] One aspect of the present invention is a learning device that includes a learning unit that performs learning of a learning model using noise waveform data including multiple sounds emitted from each sound source in a simulated noisy environment and feature values of waveform data of sounds emitted from a predetermined sound source among the sound sources as learning data, and waveform data of sounds emitted from the predetermined sound source in a simulated quiet environment as a correct answer label.
[0010] One aspect of the present invention is a voice acquisition method executed by the above-mentioned voice acquisition device, comprising the steps of measuring the position of a wireless tag that is movable together with a predetermined sound source among the sound sources of each voice in noise containing multiple voices; acquiring waveform data of the noise from a microphone array and measuring the position of the sound source of each voice based on the acquired waveform data of the noise; and acquiring waveform data of the voice emitted from the sound source present at the position of the wireless tag from the waveform data of the noise based on a match between the position of the wireless tag and the position of the sound source of each voice.
[0011] One aspect of the present invention is a learning method executed by the above-mentioned learning device, which includes a step of executing learning of a learning model using noise waveform data including multiple sounds emitted from each sound source in a simulated noisy environment and the position of a predetermined sound source in the simulated noisy environment as learning data, and using the feature quantities of the sounds emitted from the predetermined sound source as correct labels.
[0012] One aspect of the present invention is a learning method executed by the above-mentioned learning device, which includes a step of executing learning of a learning model using noise waveform data including multiple sounds emitted from each sound source in a simulated noisy environment and features of waveform data of sounds emitted from a predetermined sound source among the sound sources as learning data, and using waveform data of sounds emitted from the predetermined sound source in a simulated quiet environment as a correct answer label. [Effects of the Invention]
[0013] According to the present invention, it is possible to obtain, from noise waveform data, a sound emitted from a predetermined sound source that can move in a noisy environment. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a voice acquisition system according to an embodiment. [Figure 2]1 is a diagram illustrating an example of the configuration of a wireless tag according to an embodiment. [Figure 3] FIG. 2 is a sequence diagram illustrating an example of the operation of the voice capturing system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing an example of the configuration of a voice capturing system 1 according to an embodiment. Hereinafter, a predetermined sound source (opt-in target) from which acquisition of voice waveform data is permitted will be referred to as a "permission target." In contrast, a sound source (opt-out target) from which acquisition of voice waveform data is not permitted will be referred to as a "non-permission target." The voice capturing system 1 is a system that captures voice emitted from a predetermined sound source (permission target) in a noisy environment from noise waveform data.
[0016] The voice capturing system 1 includes a learning device 2, a plurality of wireless measurement devices 3, a communication line 4, a microphone array 5, a voice capturing device 6, and a robot 7. Note that, although the voice capturing system 1 includes two wireless measurement devices 3 in FIG. 1 as an example, it may include more wireless measurement devices 3.
[0017] The learning device 2 is a device that executes machine learning of a learning model (deep model). The learning device 2 includes a storage device 21, a learning unit 22, and a communication unit 23. The communication line 4 is, for example, the Internet. The microphone array 5 includes an array of multiple microphones. For example, the multiple microphones are arranged in a line. The voice acquisition device 6 is a device that acquires voice emitted from a licensed subject (a predetermined sound source) from noise waveform data. The voice acquisition device 6 includes a communication unit 61, a storage device 62, a tag positioning unit 63, a sound source positioning unit 64, a voice acquisition unit 65, and a recording processing unit 66.
[0018] External calibration is performed in advance on the wireless measurement device 3. For example, the wireless measurement device 3 is installed in a predetermined position and orientation in the real space, which is predetermined as the position of the wireless measurement device 3. Note that world coordinates (xw, yw, zw) may be conveniently defined for the real space (noisy environment) in which the wireless measurement device 3 and the microphone array 5 are installed.
[0019] The microphone array 5 is installed in a predetermined orientation at a predetermined position in real space. The microphone array 5 may be installed, for example, near the wireless measuring device 3. For the microphone array 5, microphone array coordinates (xa, ya, za) may be conveniently determined.
[0020] Some or all of the functional units of the voice capture device 6 are realized as software by a processor such as a CPU (Central Processing Unit) executing a program stored in a storage unit having a non-volatile storage medium (non-transitory storage medium). The program may be recorded on a computer-readable storage medium. Examples of computer-readable storage media include portable media such as flexible disks, magneto-optical disks, ROMs (Read Only Memory), and CD-ROMs (Compact Disc Read Only Memory), and non-transitory storage media such as hard disks built into computer systems.
[0021] Some or all of the functional units of the voice acquisition device 6 may be realized using hardware including electronic circuits (electronic circuits or circuitry) using, for example, an LSI (Large Scale Integrated circuit), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0022] The robot 7 is a robot that can provide a predetermined service, for example, a robot that can converse with the sound source 201 (license target).
[0023] The sound source 201 is, for example, a person. The sound source 201 may be, for example, a robot different from the robot 7, and may be another robot capable of conversation. In the following, the sound source 201 is, for example, a person. Each sound source 201 may move in a noisy environment. Here, the predetermined sound source 201 may move in a noisy environment together with the wireless tag 101. In the following, the predetermined sound source 201 (licensed object) is, for example, sound source 201-2. In addition, the unlicensed objects are, for example, sound source 201-1, sound source 201-3, sound source 201-4, and sound source 201-5.
[0024] Next, the voice capturing system 1 will be described in detail. The storage device 21 stores one or more learning models in advance. The learning model (deep model) has, for example, a convolutional neural network. The storage device 21 also stores one or more trained models generated from the one or more learning models.
[0025] The learning unit 22 generates one or more trained models from one or more learning models using a machine learning technique. The machine learning technique is, for example, supervised learning. The learning unit 22 updates (optimizes) parameters of a convolutional neural network included in the learning model using, for example, backpropagation.
[0026] For example, the learning unit 22 generates a trained model (hereinafter referred to as a "feature extraction model") that extracts features of speech waveform data from noise waveform data from the training model. In other words, the feature extraction model is a filter (frequency filter) that extracts, from the noise waveform data, features (characteristic values) of a predetermined speech contained in the frequency components at each time in the noise waveform data.
[0027] In the learning stage of the feature extraction model, the learning data (explanatory variables) of the feature extraction model are waveform data of noise including multiple sounds emitted from each sound source 201 in a simulated noisy environment and the position of a predetermined sound source 201 (permission target) in the simulated noisy environment. For example, the learning data (explanatory variables) used for learning the feature extraction model are generated in advance by simulating a noisy environment in which the microphone array 5, the sound source 201-2, and a sound source 201 other than the sound source 201-2 are arranged. In addition, the correct answer label (objective variable) used for learning the feature extraction model is the feature of the waveform data of the sound emitted from the sound source 201-2 (permission target) in the simulated quiet environment.
[0028] The learning unit 22 inputs learning data including the waveform data of noise and the position of the sound source 201-2 to the feature extraction model. The learning unit 22 updates the parameters of the convolutional neural network provided in the feature extraction model so that the features obtained from the feature extraction model approach the correct labels used for learning the feature extraction model.
[0029] In the inference stage after the learning stage of the feature extraction model, the feature extraction model is used to extract features of the voice emitted from the licensed subject. The explanatory variables of the feature extraction model are noise waveform data including multiple voices emitted from each sound source 201 in a noisy environment and the position of the sound source 201-2 (licensed subject) in the noisy environment. That is, the explanatory variables of the feature extraction model are the noise waveform data and the position of the wireless tag 101 (position of the sound source 201-2). In addition, the objective variable of the feature extraction model is the feature (voice feature) of the waveform data of the voice emitted from the sound source 201-2 (licensed subject) located at the position of the wireless tag 101.
[0030] Furthermore, for example, the learning unit 22 generates a trained model (hereinafter referred to as a "waveform extraction model") that extracts speech waveform data from noise waveform data from the training model. In other words, the waveform extraction model is a filter (frequency filter) that extracts predetermined speech waveform data contained in the frequency components at each time in the noise waveform data from the noise waveform data.
[0031] In the learning stage of the waveform extraction model, the learning data (explanatory variables) of the waveform extraction model are waveform data of noise including multiple sounds emitted from each sound source 201 in a simulated noisy environment and feature quantities of waveform data of sound emitted from sound source 201-2 (approval target) among the sound sources 201. For example, a noisy environment in which the microphone array 5, sound source 201-2 (sound source 201-2 in the embodiment), and sound sources 201 other than sound source 201-2 are arranged is simulated, thereby generating in advance the learning data (explanatory variables) used for learning the waveform extraction model. In addition, the ground truth label used for learning the waveform extraction model is the waveform data of sound emitted from sound source 201-2 (approval target) in a simulated quiet environment.
[0032] The learning unit 22 inputs learning data including noise waveform data and feature quantities of waveform data of voice emitted from the sound source 201-2 (license target) to the waveform extraction model. The learning unit 22 updates the parameters of the convolutional neural network provided in the waveform extraction model so that the voice waveform data obtained from the waveform extraction model approaches the correct label used for training the waveform extraction model.
[0033] In the inference stage after the learning stage of the waveform extraction model, the waveform extraction model is used to extract waveform data of the voice emitted from the licensed subject. The explanatory variables of the waveform extraction model are waveform data of noise including multiple voices emitted from each sound source 201 in a noisy environment and feature quantities of the waveform data of the voice emitted from sound source 201-2 (licensed subject). In addition, the objective variable of the feature quantity extraction model is the waveform data of the voice emitted from sound source 201-2 (licensed subject).
[0034] The communication unit 23 communicates with other devices via the communication line 4. The communication unit 23 transmits the feature extraction model and the waveform extraction model to the voice acquisition device 6. The communication unit 23 may acquire waveform data of new voice acquired by the voice acquisition device 6 using the waveform extraction model from the communication unit 61. The acquired waveform data of new voice may be used to continuously update (optimize) the parameters of the convolutional neural network provided in each trained model.
[0035] The wireless measuring device 3 is an anchor device. The wireless measuring device 3 measures the distance from the wireless measuring device 3 to the wireless tag 101 based on a wireless signal. The distance measurement method is not limited to a specific measurement method. For example, the wireless measuring device 3 may measure the distance from the wireless measuring device 3 to the wireless tag 101 based on the time it takes for the wireless signal to travel back and forth between the wireless measuring device 3 and the wireless tag 101 (round trip time). For example, if the time of the wireless tag 101 and the time of the wireless measuring device 3 are synchronized, the wireless measuring device 3 may measure the distance from the wireless measuring device 3 to the wireless tag 101 based on the difference between the time the wireless tag 101 transmits a wireless signal and the time the wireless signal is received by the wireless measuring device 3. Each wireless measuring device 3 transmits the distance measurement result to the tag positioning unit 63 via the communication line 4 and the communication unit 61.
[0036] The wireless measuring device 3 may measure the angle of arrival (AOA) of the wireless signal transmitted from the wireless tag 101. That is, the wireless measuring device 3 may measure the angle of the wireless tag 101 relative to the position of the wireless measuring device 3 based on the wireless signal. Each wireless measuring device 3 may transmit the measurement result of each angle to the tag positioning unit 63 via the communication line 4 and the communication unit 61.
[0037] In a noisy environment, the microphone array 5 collects waveform data of noise containing multiple sounds emitted almost simultaneously from multiple sound sources 201. For example, the microphone array 5 may collect waveform data of noise containing multiple sounds by scanning (beamforming) the periphery of the microphone array 5 using a spatial filter (sound enhancement filter).
[0038] The voice capturing device 6 acquires waveform data of the voice emitted from the sound source 201-2 in a noisy environment from the waveform data of noise collected by the microphone array 5. That is, the voice capturing device 6 stores waveform data of the voice emitted from the authorization target in a noisy environment.
[0039] The robot 7 interacts with the licensed subject based on, for example, the voice of the licensed subject collected using the microphone array 5. More specifically, the robot 7 acquires, from the voice acquisition device 6, waveform data of the voice emitted from the sound source 201-2 (licensed subject), which is obtained from the collected waveform data of noise. The robot 7 may interact with the sound source 201-2 using a large language model (LLM) based on the waveform data of the voice emitted from the sound source 201-2.
[0040] 2 is a diagram showing an example of the configuration of a wireless tag 101 in an embodiment. The wireless tag 101 includes a transmitting unit 102, a memory 103, and an operation unit 104. The transmitting unit 102 transmits a wireless signal (radio wave) including identification information of the wireless tag 101 to the wireless measuring device 3. The transmitting unit 102 may transmit the wireless signal to the wireless measuring device 3, for example, when the operation unit 104 is turned on. The memory 103 stores the identification information of the wireless tag 101 in advance.
[0041] The operation unit 104 accepts operations by the sound source 201-2 holding the wireless tag 101. The sound source 201-2 (licensed subject) may use the wireless tag 101 to notify the voice acquisition device 6 of a period during which waveform data of the sound emitted from the sound source 201-2 can be acquired by turning the operation unit 104 on and off. Here, the sound source 201-2 turns on the operation unit 104 during the period during which waveform data of the sound from the sound source 201-2 can be acquired. The length of the period during which waveform data of the sound can be acquired is not limited to a specific time length, but is, for example, 5 seconds. The sound source 201-2 turns off the operation unit 104 during the period during which waveform data of the sound emitted from the sound source 201-2 cannot be acquired.
[0042] Sound source 201-2 may repeatedly turn on and off operation unit 104. This allows the feature amount of waveform data of the sound emitted from sound source 201-2 (license target) to be continuously updated by sound acquisition unit 65. Furthermore, sound source 201-2 may turn on operation unit 104 when sound source 201-2 and other sound sources 201 are not located in the same direction with microphone array 5 as the origin.
[0043] Next, the voice capturing device 6 will be described in detail. The communication unit 61 executes communication via the communication line 4. The communication unit 61 acquires a feature extraction model and a waveform extraction model from the learning device 2. The communication unit 61 acquires distance information between the wireless tag 101 and each wireless measurement device 3 from each wireless measurement device 3. The communication unit 61 may acquire information on the arrival direction (arrival angle) of the wireless signal arriving at the wireless measurement device 3 from the wireless tag 101 from each wireless measurement device 3 instead of acquiring distance information.
[0044] The communication unit 61 acquires the collected waveform data of noise from the microphone array 5. The communication unit 61 records the collected waveform data of noise in the storage device 62. The communication unit 61 may output to the robot 7 the waveform data of the voice acquired by the voice acquisition unit 65 using the waveform extraction model (waveform data of the voice to be authorized).
[0045] The storage device 62 stores the collected noise waveform data. The storage device 62 stores audio waveform data (audio waveform data of the licensed subject) acquired from the noise waveform data. Here, the acquired audio waveform data may be associated with identification information of the wireless tag 101. In FIG. 1, the audio waveform data emitted from the sound source 201-2 (licensed subject) may be associated with identification information of the wireless tag 101 and stored in the storage device 62.
[0046] The tag positioning unit 63 measures the position of the wireless tag 101 that is movable together with a predetermined sound source 201-2. For example, the tag positioning unit 63 may perform Ultra-Wideband (UWB) positioning of the wireless tag 101 using each wireless measurement device 3. The tag positioning unit 63 measures the position of the wireless tag 101 based on a measurement result of the distance between the wireless tag 101 and each wireless measurement device 3, for example, with the position of the wireless measurement device 3 or the microphone array 5 as the origin. The tag positioning unit 63 may measure the position of the wireless tag 101 based on a measurement result of the arrival direction (arrival angle) of a wireless signal arriving at each wireless measurement device 3 from the wireless tag 101, for example, with the position of the wireless measurement device 3 or the microphone array 5 as the origin. The tag positioning unit 63 tracks the position of the wireless tag 101 that moves together with the sound source 201-2 (licensed object).
[0047] The sound source location unit 64 acquires waveform data of noise from the microphone array 5. The sound source location unit 64 measures the position of the sound source 201 of each sound based on the acquired waveform data of noise. The sound source location unit 64 estimates the arrival direction of each sound using the microphone array 5 as the origin. The sound source location unit 64 measures the position of the sound source 201 of each sound based on the arrival direction of each sound. For example, the sound source location unit 64 detects waveform data of a sound that matches or is similar to the waveform data of each sound included in waveform data of noise that has arrived at a different wireless measurement device 3. The sound source location unit 64 may measure the position of the sound source 201 of each sound using triangulation based on the measurement results of the arrival direction (arrival angle) of the waveform data of the matching or similar sound.
[0048] The voice acquiring unit 65 acquires waveform data of the voice emitted from the sound source 201-2 from the waveform data of noise based on the coincidence between the position of the wireless tag 101 and the position of the sound source 201-2. That is, the voice acquiring unit 65 acquires waveform data of the voice emitted from the sound source 201-2 that is located at the same position as the wireless tag 101.
[0049] Furthermore, the voice acquiring unit 65 may use a feature extraction model to extract features (voice feature vectors) of the waveform data of the voice emitted from the voice source 201-2 from the waveform data of the voice emitted from the voice source 201-2. The voice acquiring unit 65 may use a waveform extraction model in which the feature of the waveform data of the voice emitted from the voice source 201-2 is one of the explanatory variables to extract waveform data of new voice emitted from the voice source 201-2 from the waveform data of noise.
[0050] The recording processing unit 66 records waveform data of sound emitted from the sound source 201-2 (license target) that is present at a position that coincides with the position of the wireless tag 101 in the storage device 62. The recording processing unit 66 may further record waveform data of new sound extracted based on the feature amount of the waveform data of the sound emitted from the sound source 201-2 in the storage device 62. In other words, the recording processing unit 66 may further record waveform data of new sound emitted from the sound source 201-2 (license target) in the storage device 62.
[0051] Next, an example of the operation of the voice capturing system 1 will be described. 3 is a sequence diagram showing an example of the operation of the voice capturing system 1 in the embodiment. The microphone array 5 transmits waveform data of noise including multiple sounds emitted almost simultaneously from multiple sound sources 201 to the voice capturing device 6 for each microphone constituting the microphone array 5 (step S101). The sound source positioning unit 64 estimates the arrival direction of each sound, with the microphone array 5 as the origin (origin of polar coordinates) (step S102). The sound source positioning unit 64 measures the position of the sound source 201 of each sound based on the arrival direction of each sound (step S103).
[0052] The wireless tag 101 transmits a wireless signal to the wireless measuring device 3 (step S104). The wireless measuring device 3 measures the distance between the wireless tag 101 held in the licensed object and each wireless measuring device 3 using the wireless signal (step S105). The tag positioning unit 63 measures the position of the wireless tag 101 based on the distance measurement result, for example, with the position of the microphone array 5 as the origin (step S106).
[0053] The audio acquisition unit 65 acquires waveform data of the audio emitted from the audio source 201-2 (license target) from the waveform data of noise based on the coincidence between the position of the wireless tag 101 and the position of the audio source 201 (step S107). The recording processing unit 66 records the waveform data of the audio emitted from the audio source 201-2 in the storage device 62 (step S108).
[0054] The speech acquisition unit 65 uses the feature extraction model to extract features (speech feature vectors) of the waveform data of the speech emitted from the sound source 201-2 from the waveform data of the speech emitted from the sound source 201-2 (step S109). The speech acquisition unit 65 uses the waveform extraction model to extract waveform data of new speech emitted from the sound source 201-2 from the waveform data of noise (step S110). The recording processing unit 66 further records the waveform data of the new speech emitted from the sound source 201-2 in the storage device 62 (step S111).
[0055] As described above, the tag positioning unit 63 measures the position of the wireless tag 101 that is movable together with the predetermined sound source 201-2 among the sound sources 201 of each sound in noise containing multiple sounds. The sound source positioning unit 64 acquires waveform data of the noise from the microphone array 5. The sound source positioning unit 64 measures the position of the sound source 201 of each sound based on the acquired waveform data of the noise. The sound acquiring unit 65 acquires waveform data of the sound emitted from the sound source 201 (sound source 201-2) present at the position of the wireless tag 101 from the waveform data of the noise based on a match between the position of the wireless tag 101 and the position of the sound source 201 of each sound. That is, the sound acquiring unit 65 acquires waveform data of the sound emitted from the sound source 201-2 present at a position that coincides with the position of the wireless tag 101.
[0056] In this way, the audio acquisition unit 65 acquires waveform data of the audio emitted from the sound source 201 (sound source 201-2) located at the position of the wireless tag 101 from the waveform data of noise based on the coincidence between the position of the wireless tag 101 and the position of the sound source 201 of each audio.
[0057] This makes it possible to acquire, from the waveform data of noise, the sound emitted from a predetermined sound source 201 (sound source 201-2) that can move in a noisy environment. It is possible to protect the privacy of an unauthorized subject. Even if the sound of the sound source 201-2 (licensed subject) is not pre-registered in the sound acquisition unit 65, it is possible to acquire the sound emitted from the sound source 201-2 from the waveform data of noise. Furthermore, even if the face of the sound source 201-2 (licensed subject) is not pre-registered in the storage device 62 for face authentication, it is possible to acquire the sound emitted from the sound source 201-2 from the waveform data of noise.
[0058] Here, since the position of sound source 201-2 holding wireless tag 101 is known, it is possible to acquire waveform data of the sound of sound source 201-2 more robustly than with beamforming that forms a spatial filter, even if multiple sound sources 201 (people) are located in the same direction with respect to microphone array 5. In other words, since the position of sound source 201-2 holding wireless tag 101 is known, it is possible to acquire waveform data of the sound of sound source 201-2 more robustly than with beamforming that forms a spatial filter, even if multiple sounds emitted almost simultaneously from sound source 201-2 and other sound sources 201 arrive at microphone array 5 from the same arrival direction.
[0059] The learning unit 22 may use noise waveform data including multiple sounds emitted from each sound source 201 in a simulated noisy environment and the position of a predetermined sound source 201 (sound source 201-2) in the simulated noisy environment as learning data, and may execute learning of a learning model (feature extraction model) using sound features of the sounds emitted from the predetermined sound source 201 (sound source 201-2) as correct labels.
[0060] The learning unit 22 may use, as learning data, noise waveform data including multiple sounds emitted from each sound source 201 in a simulated noisy environment and sound features of waveform data of sounds emitted from a predetermined sound source 201 (sound source 201-2) among the sound sources 201, and may perform learning of a learning model (waveform extraction model) using, as a correct answer label, the waveform data of sounds emitted from the sound source 201 (sound source 201-2) in a simulated quiet environment.
[0061] The speech acquisition unit 65 (feature extraction unit) may extract features (speech features) of the waveform data of speech emitted from the sound source 201-2 from the acquired speech waveform data using a feature extraction model. The speech acquisition unit 65 (waveform extraction unit) may further acquire new speech waveform data from noise waveform data using a waveform extraction model based on the extracted speech features. Here, the speech acquisition unit 65 (waveform extraction unit) acquires new speech waveform data having the extracted speech features from the noise waveform data. The speech features may be newly extracted from the new speech waveform data. In other words, the speech features may be updated.
[0062] This makes it possible to know the feature quantities (audio feature quantities) of the waveform data of the sound emitted by sound source 201-2 holding wireless tag 101, and therefore it is possible to acquire the waveform data of the sound of sound source 201-2 more robustly than with beamforming that forms a spatial filter, even if multiple sound sources 201 (people) are positioned in the same direction with respect to microphone array 5. In other words, since it is possible to know the feature quantities of the waveform data of the sound emitted by sound source 201-2 holding wireless tag 101, it is possible to acquire the waveform data of the sound of sound source 201-2 more robustly than with beamforming that forms a spatial filter, even if multiple sounds emitted almost simultaneously from sound source 201-2 and another sound source 201 arrive at microphone array 5 from the same arrival direction.
[0063] (Variation) The predetermined sound source may be an unlicensed object. In this modification, the sound source 201-2 is an unlicensed object. The voice capturing system 1 may be a system that deletes a voice emitted from a predetermined sound source (unlicensed object) in a noisy environment from collected noise waveform data.
[0064] When the sound source 201-2 located at the same position as the wireless tag 101 is an unauthorized subject, the recording processing unit 66 may delete the waveform data of the sound emitted from the sound source 201-2 from the waveform data of noise so as not to record the waveform data of the sound emitted from the sound source 201-2 in the storage device 62. The recording processing unit 66 may also prevent the waveform data of the sound extracted using the waveform extraction model from being recorded in the storage device 62. In other words, the recording processing unit 66 may prevent the waveform data of the sound emitted from the sound source 201-2 (unauthorized subject) from being recorded in the storage device 62.
[0065] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0066] 1...voice acquisition system, 2...learning device, 3...wireless measuring device, 4...communication line, 5...microphone array, 6...voice acquisition device, 7...robot, 21...storage device, 22...learning unit, 23...communication unit, 61...communication unit, 62...storage device, 63...tag positioning unit, 64...sound source positioning unit, 65...voice acquisition unit, 66...recording processing unit, 101...wireless tag, 102...transmitting unit, 103...memory, 104...operation unit, 201...sound source
Claims
1. a tag positioning unit that measures the position of a wireless tag that is movable together with a predetermined sound source among the sound sources of each sound in the noise that includes a plurality of sounds; a sound source positioning unit that acquires waveform data of the noise from a microphone array and measures a position of a sound source of each of the sounds based on the acquired waveform data of the noise; a sound acquisition unit that acquires waveform data of a sound emitted from a sound source present at the position of the wireless tag from the waveform data of the noise based on a match between the position of the wireless tag and the position of a sound source of each sound; An audio capturing device comprising:
2. The voice acquisition device according to claim 1 , wherein the voice acquisition unit extracts a feature amount of the voice waveform data from the acquired voice waveform data.
3. The speech acquisition device according to claim 2 , wherein the speech acquisition unit further acquires new speech waveform data from the noise waveform data based on the extracted feature amount.
4. The speech acquisition device according to claim 3 , wherein the speech acquisition unit acquires the waveform data of the new speech having the extracted feature amount from the waveform data of the noise.
5. A voice capturing method executed by a voice capturing device, measuring the position of a wireless tag that is movable together with a predetermined sound source among the sound sources of each sound in the noise including a plurality of sounds; acquiring waveform data of the noise from a microphone array, and determining a sound source position of each of the sounds based on the acquired waveform data of the noise; acquiring waveform data of a sound emitted from a sound source present at the position of the wireless tag from the waveform data of the noise based on a match between the position of the wireless tag and the position of a sound source of each sound; An audio acquisition method including:
Citation Information
Patent Citations
Information processing unit, information processing method, and program
JP2012038131A
Image forming apparatus, image forming system, and information processing method
JP2020127104A
Communication server and communication system
JP2023045371A
Speech recognition device, speech recognition system, and speech recognition method
WO2019130399A1
Tag interrogator and microphone array for identifying a person speaking in a room
US7995731B2