Recording device, recording system, and recording method thereof
The recording device and system effectively separate and utilize a user's voice from multiple voices by using a bone conduction microphone for the user's voice and an air conduction microphone for others, with a voice separation unit processing these signals for effective isolation and utilization.
Patent Information
- Application Number
- JP2023204187
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-12-01
AI Technical Summary
Existing recording devices equipped with bone conduction and air conduction microphones struggle to separate and utilize a user's voice from the voices of multiple people in a recording environment.
The recording device and system incorporate a bone conduction microphone to capture the user's voice and an air conduction microphone to capture the voices of multiple people, with a voice separation unit processing these signals to isolate the user's voice and store it separately from the other voices.
This solution enables effective separation and utilization of the user's voice from the voices of others, allowing for separate playback or text conversion of the user's voice and the voices of surrounding people.
Smart Images

Figure 0007696096000001 
Figure 0007696096000002 
Figure 0007696096000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a recording device, a recording system, and a recording method thereof that include a bone conduction microphone and an air conduction microphone and record a user's voice.
Background Art
[0002] Conventionally, bone conduction microphones that pick up a user's voice by utilizing vibrations of the user's jaw or head bones have become widespread. The sound pickup characteristics of bone conduction microphones (for example, the frequency characteristics of the picked-up sound) are different from those of general air conduction microphones that utilize air vibrations. Generally, bone conduction microphones are known to be more suitable for sound pickup, for example, in an environment with relatively a lot of noise. In addition, technologies that utilize the differences in the sound pickup characteristics of such bone conduction microphones and air conduction microphones have been developed.
[0003] For example, when a user sings a song, the singing sound simultaneously input from a bone conduction microphone (that is, a bone conduction microphone) and an air conduction microphone (that is, an air conduction microphone) is recorded in a recording unit as air conduction recording data and bone conduction recording data, and there is a recording and playback system that synchronously plays back the recording data of the combination of the recorded air conduction recording data and bone conduction recording data (see Patent Document 1). Thereby, the user can listen to his or her own singing sound played back without a sense of discomfort.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Incidentally, the recorded data generated by a recording device equipped with a bone conduction microphone and an air conduction microphone may include the voices of a plurality of people. For example, this may be the case when recording a conversation between the user of the recording device (e.g., the person who possesses the recording device) and the people present around them. On the other hand, for recorded data containing the voices of such a plurality of people, there may be a case where the user wants to separate and utilize the user's voice from the voices of the people around them. The utilization of such voices includes, for example, separately playing the user's voice and the voices of the people around them, or separately converting the user's voice and the voices of the people around them into text.
[0006] On the other hand, according to the prior art described in Patent Document 1, by utilizing the difference in the sound collection characteristics of the bone conduction microphone and the air conduction microphone, the user can listen to their own recorded singing voice without discomfort. However, in that prior art, in the case where recording is performed in an environment where a plurality of people are speaking, separating and utilizing the user's voice from the voices of the people around them is not assumed at all.
[0007] Therefore, a main object of the present disclosure is to provide a recording device, a recording system, and their recording methods that enable the separation and utilization of the voice of a user from the voices of a plurality of people including the user when the voices of the plurality of people are collected using a bone conduction microphone and an air conduction microphone.
Means for Solving the Problems
[0008] The recording device of the present disclosure includes a bone conduction microphone that collects the voice of a user and generates a bone conduction sound signal, an air conduction microphone that collects the voices of a plurality of people including the user and generates an air conduction sound signal, a voice separation unit that generates a separated sound signal by executing a process for separating the voice of the user based on the bone conduction sound signal from the voices of the plurality of people based on the air conduction sound signal, and a storage unit that stores bone conduction sound data including the voice of the user generated based on the bone conduction sound signal and separated sound data including the voices of the people other than the user generated based on the separated sound signal, respectively.
[0009] The recording system of the present disclosure includes a microphone set used by each of a plurality of users, and a server that respectively acquires voice signals from each of the microphone sets. Each of the microphone sets includes a bone conduction microphone that picks up the voice of each user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including each user to generate an air conduction sound signal. The server includes: an audio data acquisition unit that respectively acquires bone conduction audio data based on the bone conduction sound signal and air conduction audio data based on the air conduction sound signal from each of the microphone sets; an air conduction audio data synthesis unit that generates synthesized air conduction audio data by synthesizing the plurality of air conduction audio data acquired by the audio data acquisition unit; an audio data separation unit that generates separated audio data by executing a process for separating the voice of the user based on the bone conduction audio data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction audio data; and a storage unit that stores the bone conduction audio data and the separated audio data respectively.
[0010] The recording method of the present disclosure is a recording method of a recording device. The recording device includes a bone conduction microphone that picks up the voice of a user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including the user to generate an air conduction sound signal. A separated sound signal is generated by executing a process for separating the voice of the user based on the bone conduction sound signal from the voices of the plurality of people based on the air conduction sound signal. Bone conduction audio data including the voice of the user is generated based on the bone conduction sound signal, and separated audio data including the voices of the people other than the user is generated based on the separated sound signal. The bone conduction audio data and the separated audio data are respectively stored.
[0011] The recording method of the present disclosure is a recording method of a recording system. The recording system includes a microphone set used by each of a plurality of users, and a server that respectively acquires voice signals from the microphone sets. Each microphone set includes a bone conduction microphone that picks up the voice of each user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including each user to generate an air conduction sound signal. The server respectively acquires bone conduction sound data based on the bone conduction sound signal and air conduction sound data based on the air conduction sound signal from each microphone set, generates synthesized air conduction sound data by synthesizing a plurality of the air conduction sound data acquired by the voice data acquisition unit, and executes a process for separating the voice of the user based on the bone conduction sound data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction sound data to generate separated sound data, and is configured to respectively store the bone conduction sound data and the separated sound data.
Effects of the Invention
[0012] According to the present disclosure, when voices of a plurality of people including a user are picked up using a bone conduction microphone and an air conduction microphone, it is possible to separate and use the voice of the user and the voices of the people around the user.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0014] The first invention made to solve the above problems includes a bone conduction microphone that picks up the user's voice and generates a bone conduction sound signal, an air conduction microphone that picks up the voices of a plurality of people including the user and generates an air conduction sound signal, and a voice separation unit that generates a separated sound signal by executing a process for separating the user's voice based on the bone conduction sound signal from the voices of the plurality of people based on the air conduction sound signal, and a storage unit that stores bone conduction sound data including the user's voice generated based on the bone conduction sound signal and separated sound data including the voices of the people other than the user generated based on the separated sound signal, respectively.
[0015] According to this, when the voices of a plurality of people including the user are picked up using the bone conduction microphone and the air conduction microphone, the bone conduction sound data including the user's voice and the separated sound data including the voices of the people other than the user are stored respectively, so that it is possible to separate and use the user's voice and the voices of the people around him / her.
[0016] The second invention further includes a bone conduction sound transmission path through which the bone conduction sound signal is transmitted and an air conduction sound transmission path through which the air conduction sound signal is transmitted, and the voice separation unit is configured to generate the separated sound signal by subtracting the bone conduction sound signal transmitted through the bone conduction sound transmission path from the air conduction sound signal transmitted through the air conduction sound transmission path.
[0017] According to this, a separated sound signal including the voices of people other than the user can be generated with a simple configuration.
[0018] Further, in the third invention, the bone conduction sound transmission path is configured to further include a low-pass filter that extracts a voice signal based on the user's voice from the bone conduction sound signal used for generating the separated sound signal.
[0019] According to this, only the user's voice can be separated from the voices of a plurality of persons including the user with higher accuracy by the voice signal based on the user's voice extracted by the low-pass filter.
[0020] Further, in the fourth invention, the bone conduction sound transmission path is configured to further include a noise canceller that removes or reduces a noise component in the bone conduction sound signal used for generating the separated sound signal.
[0021] According to this, only the user's voice can be separated from the voices of a plurality of persons including the user with higher accuracy by the voice signal in which the noise component is removed or reduced by the noise canceller.
[0022] Further, in the fifth invention, the configuration further includes a voice reproduction unit that can selectively reproduce the voice based on the bone conduction sound data and the voice based on the separated sound data.
[0023] According to this, the user's voice and the voices of the surrounding persons can be separated and reproduced.
[0024] Further, in the sixth invention, the configuration further includes a voice recognition unit capable of executing text conversion processing by selectively executing voice recognition processing on the voice based on the bone conduction sound data and the voice based on the separated sound data.
[0025] According to this, the user's voice and the voices of the surrounding persons can be separated and converted into text.
[0026] Further, the seventh invention includes a microphone set used by each of a plurality of users, and a server that respectively acquires voice signals from the microphone sets. Each of the microphone sets includes a bone conduction microphone that picks up the voice of each user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including each user to generate an air conduction sound signal. The server includes a voice data acquisition unit that respectively acquires bone conduction sound data based on the bone conduction sound signal and air conduction sound data based on the air conduction sound signal from each of the microphone sets, an air conduction sound data synthesis unit that generates synthesized air conduction sound data by synthesizing the plurality of air conduction sound data acquired by the voice data acquisition unit, and a voice data separation unit that generates separated sound data by executing a process for separating the voice of the user based on the bone conduction sound data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction sound data, and a storage unit that stores the bone conduction sound data and the separated sound data respectively.
[0027] According to this, when the voices of a plurality of users are picked up using microphone sets each having a bone conduction microphone and an air conduction microphone, bone conduction sound data including the voice of each user and separated sound data including the voice of a specific user (i.e., a user selected from among the plurality of users) are respectively stored, so that it is possible to separate and use the voice of a specific user and the voices of other users.
[0028] Further, the eighth invention further includes an input device used by any one of the plurality of users. The server acquires information on at least one of the microphone sets selected by an input operation to the input device, and the voice data separation unit generates the separated sound data based on the bone conduction sound data corresponding to the microphone sets other than the at least one selected microphone set.
[0029] According to this, the user can separate and use the voice of a specific user (including himself / herself) selected by himself / herself and the voices of other users.
[0034] Further, the invention of 9 relates to a recording method of a recording device, wherein the recording device includes a bone conduction microphone that picks up the user's voice to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of persons including the user to generate an air conduction sound signal, and executes a process for separating the user's voice based on the bone conduction sound signal from the voices of the plurality of persons based on the air conduction sound signal to generate a separated sound signal, generates bone conduction sound data including the user's voice based on the bone conduction sound signal, and generates separated sound data including the voices of the persons other than the user based on the separated sound signal, and is configured to store the bone conduction sound data and the separated sound data respectively.
[0035] According to this, when the voices of a plurality of persons including the user are picked up using the bone conduction microphone and the air conduction microphone, the bone conduction sound data including the user's voice and the separated sound data including the voices of the persons other than the user are respectively stored, so that it is possible to separate and use the user's voice and the voices of the surrounding persons.
[0036] Further, the invention of 10 relates to a recording method of a recording system, wherein the recording system includes a microphone set used by each of a plurality of users respectively, and a server that respectively acquires voice signals from the respective microphone sets, each microphone set includes a bone conduction microphone that picks up the voice of each user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of persons including each user to generate an air conduction sound signal, the server respectively acquires bone conduction sound data based on the bone conduction sound signal and air conduction sound data based on the air conduction sound signal from the respective microphone sets, generates synthesized air conduction sound data by synthesizing the acquired plurality of air conduction sound data, executes a process for separating the voice of the user based on the bone conduction sound data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction sound data to generate separated sound data, and is configured to store the bone conduction sound data and the separated sound data respectively.
[0037] According to this, when voices of a plurality of users are picked up using a microphone set having a bone conduction microphone and an air conduction microphone respectively, bone conduction sound data including the voices of each user and separated sound data including the voice of a specific user (that is, a user selected from among the plurality of users) are each stored, so that it becomes possible to separate and use the voice of the specific user and the voices of other users.
[0040] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0041] (First Embodiment) FIG. 1 is a perspective view showing a usage state of a recording device 1 according to a first embodiment of the present disclosure. FIG. 2 is an exploded perspective view of the earphone 2 shown in FIG. 1.
[0042] As shown in FIG. 1, the recording device 1 includes an earphone 2 having a sound collection function, and a recording device main body 3 (hereinafter referred to as "device main body 3") that performs signal processing for recording the sound collected by the earphone 2. Here, an example is shown in which the earphone 2 is inserted into the right ear 5 of a user (that is, a person who holds the recording device 1). However, the earphone 2 may be inserted into the left ear of the user and used. Further, the recording device 1 may include a pair of left and right earphones 2.
[0043] In the recording device 1, the device main body 3 is connected to the earphone 2 via a signal line 4. However, the earphone 2 and the device main body 3 may be communicably connected to each other via a wireless signal based on Wifi (registered trademark), Bluetooth (registered trademark), or the like.
[0044] Further, the earphone 2 may further have the same function (or at least a part of the function) as the device main body 3. That is, in the recording device 1, the device main body 3 (or at least a part of its components) may be integrated inside the housing of the earphone 2. Note that the device main body 3 may include a part of the components of the earphone 2 (for example, an air conduction microphone 11 described later).
[0045] As shown in Fig. 2, the earphone 2 includes an outer housing 9, a bone conduction microphone 10 and an air conduction microphone 11 housed in the outer housing 9, a speaker 12, an inner housing 13, and a microphone rubber 15.
[0046] The outer housing 9 constitutes the outer shell of the earphone 2. The outer housing 9 includes a case 9A that constitutes the inner part (rear part) of the user's ear and a cover 9B that constitutes the outer part of the user's ear. The cover 9B is overlapped on the case 9A from the outside to constitute the outer housing 9. The outer housing 9 is provided with a cable cap 18 for leading the signal line 4 to the outside.
[0047] The bone conduction microphone 11 is a device provided on the user's ear so as to be able to pick up the vocal cord vibrations mainly transmitted to the vocal cord vibration transmission part, and mainly picks up the voice of the user wearing the earphone 2. The bone conduction microphone 10 includes an element (for example, a vibration detection element) for converting sound vibrations into electrical signals. That is, the bone conduction microphone 10 can pick up sounds such as the user's voice and generate electrical signals. The electrical signals generated by the bone conduction microphone 10 are input to the device main body 3 via the signal line 4. Since the bone conduction microphone 10 picks up the vocal cord vibrations propagating through the user's body as sound vibrations, it has the property of being less affected by ambient sounds other than the user's voice.
[0048] The air conduction microphone 11 is a microphone provided with an element (for example, a vibration detection element) for converting the vibration of air as sound vibrations into electrical signals. The air conduction microphone 11 picks up the user's voice and ambient sounds generated around the user (in this embodiment, the voices of people around the user). The electrical signals generated by the air conduction microphone 11 are input to the device main body 3 via the signal line 4. However, the bone conduction microphone 10 and the air conduction microphone 11 are respectively connected to the device main body 3 via different transmission paths constituting the signal line 4.
[0049] The speaker 12 converts the electrical signal input from the apparatus main body 3 into sound such as the user's voice and outputs it. A known configuration can be adopted for the speaker 12. Also, a bone conduction receiver (bone conduction receiver) may be used instead of the speaker 12. For example, when the apparatus main body 3 outputs electrical signals related to the user's pre-recorded voice or ambient sound, the speaker 12 outputs those sounds (i.e., playback sounds). Note that the sound output from the speaker 12 may include not only pre-recorded sounds but also any sound generated by the apparatus main body 3 (or acquired by the apparatus main body 3 from the outside).
[0050] The inner housing 13 is a member for guiding the playback sound of the speaker 12 to the external auditory canal, and is provided with a passage 13A for guiding the playback sound of the speaker 12 to the external auditory canal. The speaker 12 is disposed on one end side of the passage 13A of the inner housing 13.
[0051] On the other end side of the passage 13A, a cylindrical tube portion 13B extending to the outside of the outer housing 9 is provided via an opening 9C provided in the outer housing 9 (specifically, the case 9A). In the present embodiment, a tip rubber 16 that abuts against the user's external auditory canal over the entire circumference is provided at the extending end of the tube portion 13B. By the tip rubber 16 abutting against the external auditory canal over the entire circumference, the airtightness of the external auditory canal is ensured, and the playback sound from the speaker 12 is effectively transmitted to the user.
[0052] The microphone rubber 15 elastically holds the bone conduction microphone 10. The microphone rubber 15 connects a site where vocal cord vibration is transmitted through the bone in the vicinity of the ear 5 (hereinafter referred to as the vocal cord vibration transmission site) and the bone conduction microphone 10. Thereby, the vocal cord vibration transmitted to the vocal cord vibration transmission site is transmitted to the bone conduction microphone 10, and the bone conduction microphone 10 picks up sound by detecting the vocal cord vibration propagated to the bone and skin around the ear via the microphone rubber 15.
[0053] The air-conduction microphone 11 is disposed between the exterior housing 9 and the speaker 12. In order to prevent the sound vibration emitted by the speaker 12 from being transmitted to the air-conduction microphone 11, a resin member 17 is provided between the air-conduction microphone 11 and the speaker 12. The air-conduction microphone 11 mainly picks up the voice of the user as the air propagates the voice to the air-conduction microphone 11 near the ear.
[0054] The device main body 3 executes predetermined processing on the sound picked up by the earphone 2 (i.e., the generated electrical signal), and stores (i.e., records) the processed sound as audio data. Further, the device main body 3 can be used by the user by playing back the recorded voice (i.e., functioning as a voice recorder) or converting the voice into text.
[0055] The device main body 3 is constituted by a terminal (here, a smartphone) provided with a processor such as a CPU, a memory such as a RAM and a ROM, a storage such as an SSD and an HDD, a display such as a touch panel, a network interface, and audio input / output terminals. However, the configuration of the device main body 3 is not limited to a smartphone, and can also be constituted by, for example, a tablet provided with audio input / output terminals or various computers. Further, the device main body 3 can also be realized by a combination of a logic device and various analog devices.
[0056] FIG. 3 is a functional block diagram showing the configuration of the recording device 1 according to the first embodiment. FIG. 4 is an explanatory diagram showing a usage example of the recording device 1 shown in FIG. 1. FIG. 5 is an explanatory diagram showing an example of a setting screen ((A) voice recorder setting screen 56, (B) speech-to-text setting screen 59) in the recording device 1.
[0057] As shown in FIG. 3, in the recording device 1, an electrical signal (hereinafter referred to as a "bone conduction sound signal") generated by the bone conduction microphone 10 is input to the device main body 3 via the signal line 4. Similarly, an electrical signal (hereinafter referred to as an "air conduction sound signal") generated by the air conduction microphone 11 is input to the device main body 3 via the signal line 4. The bone conduction sound signal is mainly a signal based on the user's voice. Also, the air conduction sound signal is a signal based on ambient sound and the voices of multiple people (including the user's voice). Although not shown, the device main body 3 may include an AD converter that converts the analog signals input from the bone conduction microphone 10 and the air conduction microphone 11 into digital signals, respectively, and a DA converter that converts the digital signal output to the speaker 12 into an analog signal.
[0058] The device main body 3 has a bone conduction sound transmission path 21 that transmits the bone conduction sound signal input from the bone conduction microphone 10. Also, the device main body 3 has an air conduction sound transmission path 22 that transmits the air conduction sound signal input from the air conduction microphone 11. Further, the device main body 3 has a connection path 23 that connects between the bone conduction sound transmission path 21 and the air conduction sound transmission path 22.
[0059] Also, the device main body 3 has a recording processing unit 25 that executes processing for recording the sounds respectively picked up by the bone conduction microphone 10 and the air conduction microphone 11 as audio data. Further, the device main body 3 has an audio reproduction unit 26 that reproduces the recorded audio data (i.e., outputs it from the speaker 12), and an audio recognition unit 27 that performs audio recognition of the recorded audio data. Furthermore, the device main body 3 has a communication unit 29 for communicating with an external device (for example, an external server).
[0060] The bone conduction transmission path 21 is provided with an amplification unit 31 that amplifies the bone conduction sound signal generated by the bone conduction microphone 10. The bone conduction sound signal amplified by the amplification unit 31 is input to channel CH1 of the recording processing unit 25. The recording processing unit 25 is provided with a first separation processing unit 40 that performs processing for generating only the user's voice from the bone conduction sound signal. The first separation processing unit 40 includes an LPF 33 (low-pass filter) for extracting a voice signal related to the user's voice from the bone conduction sound signal amplified by the amplification unit 31 (that is, for excluding ambient sound and echo components). The voice signal extracted by the LPF 33 is sent to the second separation processing unit 41. Note that in the first separation processing unit 40, instead of the LPF 33, a noise canceller that removes or reduces the noise component in the bone conduction sound signal may be used.
[0061] The second separation processing unit 41 performs processing for generating separated sound data from which the user's voice has been removed from the voices of a plurality of people based on the air conduction sound signal. In the second separation processing unit 41, bone conduction sound data 45 is generated based on the voice signal extracted by the LPF 33. The generated bone conduction sound data 45 is stored in the storage unit 42 provided in the recording processing unit 25. Also, in the second separation processing unit 41, the voice signal extracted by the LPF 33 is input to an arithmetic unit 36 (an example of a voice separation unit) that generates a separated sound signal via the connection path 23.
[0062] On the other hand, the air conduction transmission path 22 is provided with an amplification unit 35 that amplifies the air conduction sound signal generated by the air conduction microphone 11. The amplified voice signal is input to channel CH2 of the recording processing unit 25. In the second separation processing unit 41, air conduction sound data 61 is generated based on the amplified voice signal. The generated air conduction sound data 61 is stored in the storage unit 42. Also, in the second separation processing unit 41, the voice signal amplified by the amplification unit 35 is input to the arithmetic unit 36 via the connection path 23.
[0063] The arithmetic unit 36 generates a separated sound signal by subtracting the sound signal (i.e., bone-conducted sound signal) extracted by the LPF 33 from the sound signal (i.e., air-conducted sound signal) amplified by the amplifier 35. Further, the arithmetic unit 36 generates separated sound data 46 based on the separated sound signal. The generated separated sound data 46 is stored in the storage unit 42.
[0064] In this way, in the apparatus main body 3, the bone-conducted sound signal from the bone-conducted sound transmission path 21 and the separated sound signal from the air-conducted sound transmission path 22 are input to different channels (i.e., channel CH1, channel CH2) of the recording processing unit 25, respectively. Each of the transmission paths 21-23, the amplifiers 31, 35, and the arithmetic unit 36 can be constituted by electronic elements and circuits, respectively.
[0065] The recording processing unit 25 executes recording processing on the bone-conducted sound signal and the separated sound signal input from different channels CH1 and CH2, respectively. By the recording processing, bone-conducted sound data 45, separated sound data 46, and air-conducted sound data 61 are generated from the bone-conducted sound signal, the separated sound signal, and the air-conducted sound signal, respectively. The generated bone-conducted sound data 45, separated sound data 46, and air-conducted sound data 61 are stored in the storage unit 42, respectively.
[0066] The recording processing by the recording processing unit 25 can be realized by at least one processor executing a predetermined control program (e.g., recording software). Regarding the processing of the electrical signal (i.e., sound signal) by the recording processing unit 25, known processing can be adopted, and for example, the sampling rate, bit depth, and recording format are set in advance.
[0067] The storage unit 42 includes a storage device such as a storage for storing data and information necessary for the processing of the recording device 1.
[0068] The voice playback unit 26 executes a playback process for playing back voice data (here, bone conduction voice data 45 and separated voice data 46) stored in the storage unit 42. Through this playback process, a corresponding voice signal is generated. The generated voice signal is sent to the earphone 2, and the corresponding voice is output from the speaker 12.
[0069] The playback process by the voice playback unit 26 can be realized by at least one processor executing a predetermined control program (for example, voice playback software). Note that for the processing of voice data by the voice playback unit 26, known processing can be adopted.
[0070] The voice recognition unit 27 executes a voice recognition process (an example of text conversion processing) for recognizing the voice included in the voice data stored in the storage unit 42. Through this voice recognition process, the voice-recognized voice data is converted into text, and the corresponding text data is generated. The generated text data is stored in the storage unit 42.
[0071] The voice recognition process by the voice recognition unit 27 can be realized by at least one processor executing a predetermined control program (for example, voice recognition software). For the voice recognition process by the voice recognition unit 27, a voice recognition engine equipped with a machine learning model generated in advance may be used. Note that for the processing of voice data by the voice recognition unit 27, known processing can be adopted.
[0072] The communication unit 29 performs wireless communication or wired communication with other devices via a communication network (not shown) according to a known communication protocol. The communication unit 29 may include a communication device equipped with an antenna, a communication circuit, and the like.
[0073] In this way, when the recording device 1 uses the earphone 2 (that is, the bone conduction microphone 10 and the air conduction microphone 11) to pick up the voices of a plurality of people including the user, the voice of the user (that is, the bone conduction sound data 45) and the voices of the people around the user (that is, the separated sound data 46 and the air conduction sound data 61) can be separated and used.
[0074] Next, based on FIG. 4 (see also FIG. 3), the recording device 1 will be described in terms of the method of picking up and recording the voices of a plurality of people including the user and the use of the recorded voices.
[0075] As shown in FIG. 4, the recording device 1 can be used to record the conversations of a plurality of people in, for example, a hospital ward and utilize the recorded data. In the example shown in FIG. 4, the voices of the first to third medical staff 51A - 51C (including doctors, nurses, etc.) and the voice of the patient 52 are recorded by the recording device 1. The first medical staff 51A is the user of the recording device 1 (shown by an icon in the figure). That is, only the first medical staff 51A wears the recording device 1 (at least the bone conduction microphone 10).
[0076] The voice of the first medical staff 51A is picked up by the bone conduction microphone 10 and the air conduction microphone 11 respectively. Also, the voices of the second and third medical staff 51B, 51C and the voice of the patient 52 are picked up by the air conduction microphone 11.
[0077] The picked-up voice of the first medical staff 51A is input as a bone conduction sound signal from the bone conduction microphone 10 to the device main body 3. Also, the voices of the first to third medical staff 51A - 51C and the voice of the patient 52 (hereinafter referred to as the voices of all people) are input as air conduction sound signals from the air conduction microphone 11 to the device main body 3.
[0078] In the apparatus main body 3, a bone conduction sound signal including the voice of the first medical staff member 51A is input from the bone conduction sound transmission path 21 to channel CH1 of the recording processing unit 25. The recording processing unit 25 can perform recording processing on the bone conduction sound signal. As a result, bone conduction sound data 45 mainly including the voice of the first medical staff member 51A is generated and stored in the storage unit 42.
[0079] Also, in the apparatus main body 3, an air conduction sound signal including the voices of the first to third medical staff members 51A - 51C is input from the bone conduction sound transmission path 21 to channel CH2 of the recording processing unit 25. The recording processing unit 25 can perform recording processing on the air conduction sound signal. As a result, air conduction sound data 61 mainly including the voices of the first to third medical staff members 51A - 51C is generated and stored in the storage unit 42.
[0080] Furthermore, in the recording processing unit 25, a separated sound signal including the voices of all members other than the first medical staff member 51A is generated by the arithmetic unit 36. The recording processing unit 25 can perform recording processing on the separated sound signal. As a result, separated sound data 46 mainly including the voices of the second and third medical staff members 51B and 51C (that is, ambient sounds other than the user's voice) is generated and stored in the storage unit 42.
[0081] After that, when the user selects a desired reproduction target and instructs reproduction in the recording apparatus 1, the voice reproduction unit 26 executes reproduction processing of the voice data corresponding to the selected reproduction target. Also, when the user selects a desired text conversion target and instructs text conversion in the recording apparatus 1, the voice recognition unit 27 executes voice recognition processing of the voice data corresponding to the selected text conversion target.
[0082] The recording apparatus 1 can display, on the display of the apparatus main body 3, a setting screen for the user to use the voice data (here, the bone conduction sound data 45, the separated sound data 46, and the air conduction sound data 61) stored in the storage unit 42.
[0083] For example, as shown in Fig. 5(A), the recording device 1 can display a voice recorder setting screen 56 on the touch panel display 55 (i.e., the display device and the input device) of the device main body 3. The user can perform an input operation regarding the reproduction process of the audio data on the voice recorder setting screen 56. On the voice recorder setting screen 56, the user can select any one of the bone conduction sound data 45 (corresponding to "user" in the figure), the separated sound data 46 (corresponding to "others (surroundings)" in the figure), and the air conduction sound data 61 (corresponding to "all (user + others)" in the figure).
[0084] Note that the recording device 1 may be provided with a display that does not have an input function instead of the touch panel display 55 as the display device. In that case, the recording device 1 can be provided with a known device (such as a keyboard, etc.) as the input device.
[0085] Fig. 5(A) shows an example where the user selects their own voice as the reproduction target. Therefore, when the user presses the execution button 57 (i.e., gives an instruction to reproduce), the audio reproduction unit 26 executes the reproduction process of the bone conduction sound data 45. Thereby, the user can confirm their own voice output from the speaker 12.
[0086] Similarly, the user can also select the voice of a person in the surroundings other than the user or the voice of all persons (i.e., the user and the persons in the surroundings other than the user) on the voice recorder setting screen 56. For example, after the user selects the voice of a person in the surroundings other than the user and then presses the execution button 57, the audio reproduction unit 26 executes the reproduction process of the separated sound data 46. Thereby, the user can confirm the ambient sound other than the user's voice output from the speaker 12 (i.e., the voices of the second and third medical staff 51B, 51C and the voice of the patient 52).
[0087] For example, as shown in FIG. 5(B), the recording device 1 can display a speech-to-text setting screen 59 on the touch panel display 55 of the device main body 3. The user can perform an input operation regarding the speech recognition process of the voice data on the speech-to-text setting screen 59.
[0088] In FIG. 5(B), an example is shown where the user selects the voices of ambient sounds other than the user's own voice (i.e., the voices of the second and third medical staff 51B, 51C and the voice of the patient 52) as the target for text conversion. Therefore, when the user presses the execute button 57 (i.e., gives an instruction for text conversion), the speech recognition unit 27 executes the speech recognition process of the separated voice data 46. As a result, the voices of the second and third medical staff 51B, 51C and the voice of the patient 52 are text-converted, and a text file is generated in a predetermined format. The generated text file is stored in the storage unit 42.
[0089] In the recording device 1, the processes related to the generation and storage of voice data may be started or stopped according to the user's instruction (i.e., input operation). Also, in the recording device 1, the user can perform an input operation to store only their own voice. In that case, only the bone-conducted voice data 45 is stored in the storage unit 42 of the recording device 1. On the other hand, the user can also perform an input operation to store only the voices of people other than themselves (i.e., the voices of surrounding people). In that case, only the separated voice data 46 is stored in the storage unit 42 of the recording device 1.
[0090] In this way, in the recording device 1, since the bone-conducted voice data 45 including the user's voice, the separated voice data 46 including the voices of people other than the user, and the air-conducted voice data 61 including the user's voice and the voices of the surrounding people are respectively stored in the storage unit 42, in addition to using the voices of all the people (including the user) around the user, it is possible to separate and use (here, voice playback and text conversion) the user's voice and the voices of the surrounding people.
[0091] (Second Embodiment) In the above-described first embodiment, an example was shown in which, when recording the voices (or conversations) of a plurality of persons, only the user uses the recording device 1 (that is, wears the bone conduction microphone 10). On the other hand, in the second embodiment described below, a case will be described in which a plurality of recording devices 1 are prepared and a plurality of persons each become a user of the corresponding recording device 1 (that is, a plurality of persons each wear the bone conduction microphone 10).
[0092] FIG. 6 is a functional block diagram showing the configuration of a recording system 100 including the recording device 1 according to the second embodiment. FIG. 7 is an explanatory diagram showing an example of use of the recording system 100 shown in FIG. 6. FIG. 8 is an explanatory diagram showing an example of a setting screen ((A) voice recorder setting screen 156, (B) speech-to-text setting screen 159) in the recording device 1 according to the second embodiment. In the recording device 1 shown in FIGS. 6 to 8, the same reference numerals are given to the same components as those in the recording device 1 according to the above-described first embodiment. Further, regarding the recording device 1 according to the second embodiment, matters not particularly mentioned below are the same as those in the recording device 1 according to the above-described first embodiment.
[0093] The recording system 100 includes recording devices 1-1 to 1-N respectively used by users U1 to UN (where N is the total number of users and is an integer of 2 or more), and a management server 101 (an example of a server) that manages the voice data respectively generated by the recording devices 1-1 to 1-N. Although not shown in FIG. 6, the recording devices 1-2 to 1-N have the same configuration as the recording device 1-1. The recording devices 1-1 to 1-N are each communicably connected to the management server 101 via a communication network 103. Hereinafter, when there is no need to distinguish the recording devices 1-1 to 1-N, they are collectively referred to as the recording device 1.
[0094] In the device main body 3 of the recording device 1, in the air conduction sound transmission path 22, the above-described arithmetic unit 36 is omitted. In this case, the second separation processing unit 41 performs processing for generating only the bone conduction sound data 45 and the air conduction sound data 61 without performing processing for generating separated sound data.
[0095] The recording processing unit 25 performs recording processing on the bone conduction sound signals and air conduction sound signals respectively input from different channels CH1 and CH2. Through this recording processing, bone conduction sound data 45 and air conduction sound data 61 are generated from the bone conduction sound signals and air conduction sound signals respectively. The generated bone conduction sound data 45 and air conduction sound data 61 are stored in the storage unit 42 respectively. The voice data stored in the storage unit 42 is transmitted to the management server 101 by the communication unit 29.
[0096] The recording processing by the recording processing unit 25 can be realized by at least one processor executing a predetermined control program (for example, recording software). Regarding the processing of electrical signals (that is, voice signals) by the recording processing unit 25, known processing can be adopted, and for example, the sampling rate, bit depth, and recording format are preset.
[0097] The management server 101 includes a communication unit 105, a control unit 106, and a storage unit 107.
[0098] The communication unit 105 communicates with each recording device 1-1 to 1-N via the communication network 103 according to a known communication protocol. The communication unit 105 may include a communication device equipped with an antenna, a communication circuit, etc.
[0099] The control unit 106 has a voice data acquisition unit 111, an air conduction sound data synthesis unit 112, and a voice data separation unit 113.
[0100] The voice data acquisition unit 111 acquires the bone conduction sound data 45 and the air conduction sound data 61 respectively through communication with each recording device 1-1 to 1-N. Those bone conduction sound data 45 and air conduction sound data 61 are respectively generated by the bone conduction microphone 10 and the air conduction microphone 11 (hereinafter, referred to as "microphone set" as necessary) in each recording device 1-1 to 1-N.
[0101] The air-conducted sound data synthesizing unit 112 generates synthesized air-conducted sound data by executing a process of synthesizing a plurality of air-conducted sound data 61 acquired by the voice data acquisition unit 111. In the present embodiment, the air-conducted sound data synthesizing unit 112 synthesizes all the air-conducted sound data acquired from each of the recording devices 1-1 to 1-N. However, the air-conducted sound data synthesizing unit 112 can also generate synthesized air-conducted sound data by synthesizing a part of those air-conducted sound data (for example, data selected by the user). The generated synthesized air-conducted sound data includes, for example, a synthesized voice in which the voices of all the users U1 to UN of the recording devices 1-1 to 1-N are synthesized.
[0102] The voice data separating unit 113 generates separated sound data by executing a process for separating the utterance of a user based on bone-conducted sound data corresponding to at least one bone-conducted microphone 10 from the synthesized voice based on the synthesized air-conducted sound data.
[0103] For example, when the total number of users N = 3 (that is, when there are three users), the synthesized voice includes the voices of users U1 to U3. For example, the voice data separating unit 113 can execute a process of separating the utterance of user U3 based on the bone-conducted sound data 45 from the synthesized voice. In this case, the utterance of user U3 to be separated corresponds to the voice corresponding to the bone-conducted sound signal generated by the bone-conducted microphone 10 of the recording device 1-3 used by user U3. As a result, the separated sound data includes the voices of users U1 and U2 excluding the voice of user U3. The voice data of users U1 and U2 is transmitted to at least one of the recording devices 1-1 to 1-3 of users U1 to U3 and is used (voice playback or text conversion) in that recording device.
[0104] Note that the total number of users N is not limited to 3 (persons) and various changes are possible. Also, the utterance of the user (that is, the voice to be separated) separated from the synthesized voice by the voice data separating unit 113 can be changed as appropriate.
[0105] At least a part of the functions of each of the units 111 to 113 in the control unit 106 can be realized by at least one processor executing a predetermined control program. Further, the control unit 106 can comprehensively control the operation of the management server 101.
[0106] The storage unit 107 includes a storage device such as a storage for storing data and information necessary for the processing of the management server 101. For example, the storage unit 107 stores voice data acquired from each of the recording devices 1-1 to 1-N, and voice data (such as synthesized bone-conducted voice data and separated voice data) generated by the processing of the management server 101.
[0107] Next, based on FIG. 7 (see also FIG. 6), the recording system 100 will be described in terms of the method of collecting the voices of a plurality of persons (here, users U1 to U8) and using the collected voices.
[0108] For example, as shown in FIG. 7, each of the recording devices 1-1 to 1-8 records the voices (or conversations) of a plurality of users U1 to U8 in a gathered state, and is used to utilize the recorded data. In FIG. 7, the voices of the users U1 to U8 are respectively picked up by the bone-conduction microphones 10 of the corresponding recording devices 1-1 to 1-8 (shown by icons in the figure). Also, the voices of the users U1 to U8 can be picked up by the air-conduction microphones 11 of all the recording devices 1-1 to 1-8.
[0109] The bone-conducted voice data 45 and the air-conducted voice data 61 generated by each of the recording devices 1-1 to 1-8 are respectively transmitted to the management server 101. The management server 101 generates synthesized air-conducted voice data from all the received air-conducted voice data 61. However, the management server 101 may generate synthesized air-conducted voice data from a part of the received air-conducted voice data 61 (for example, data selected by any one of the users). Also, when the air-conducted voice data 61 generated by any one of the recording devices 1-1 to 1-8 includes clear voices of all the users U1 to U8, the management server 101 can also use that one air-conducted voice data 61 instead of the synthesized air-conducted voice data.
[0110] Next, the management server 101 executes a process for separating the user's voice based on the bone conduction sound data corresponding to one or more recording devices (i.e., the bone conduction microphones 10) selected by the user (or the administrator of the recording system 100) from the synthesized voice based on the synthesized voice conduction sound data. As a result, the management server 101 generates separated sound data including only the voices of at least some of the users (i.e., the voices of other users are removed).
[0111] The generated separated sound data, the bone conduction sound data 45 and the air conduction sound data 61 obtained from each of the recording devices 1-1 to 1-8 are stored in the storage unit 107. Those data stored in the storage unit 107 are transmitted to the recording devices 1-1 to 1-8 in response to requests from the users U1 to U8.
[0112] When any one of the users U1 to U8 selects a desired playback target and instructs playback on the corresponding recording device 1-1 to 1-8, the audio playback unit 26 executes a playback process of the audio data corresponding to the selected playback target. Also, when any one of the users U1 to U8 selects a desired text conversion target and instructs text conversion, the speech recognition unit 27 executes a speech recognition process of the audio data corresponding to the selected text conversion target.
[0113] Similar to the case of the first embodiment, the recording device 1 can display, on the display of the device main body 3, a setting screen for the users U1 to U8 to use the audio data (here, the separated sound data) stored in the storage unit 107 of the management server 101.
[0114] For example, as shown in FIG. 8(A), the recording device 1-1 can display a voice recorder setting screen 156 on the touch panel display 55 of the device main body 3. The user U1 can perform an input operation regarding the playback process of the audio data on the voice recorder setting screen 156.
[0115] In Fig. 8(A), an example (an example of an input operation on the input device) is shown in which the user U1 selects his / her own voice as the reproduction target. Therefore, when the user U1 presses the execution button 157, information regarding the reproduction target selected by the user U1 (an example of information of the input operation) is transmitted to the management server 101. Here, the selection of the reproduction target by the user U1 is synonymous with the selection of at least one microphone set by the user U1. The management server 101 can acquire information regarding the reproduction target selected by the user U1 from the recording device 1-1. The voice data separation unit 113 executes a process for separating the user's voice (i.e., generation of separated voice data) based on bone conduction voice data included in a microphone set other than at least one microphone set (i.e., the reproduction target) selected by the user U1. Thereafter, the recording device 1-1 acquires the separated voice data generated by the management server 101. Therefore, the voice playback unit 26 executes a playback process of the separated voice data including only the voice of the user U1 (for example, the voices of other users U2 to U8 are removed). Thereby, the user U1 can confirm his / her own voice output from the speaker 12.
[0116] Similarly, the user U1 can select the voices of other users U2-U8 other than the user U1 on the voice recorder setting screen 156. Also, for example, when the user U1 wants to check a conversation among a plurality of users, the user U1 can also select the voices of those plurality of users on the voice recorder setting screen 156. When the user U1 presses the execution button 57 after the selection of the voice, the voice playback unit 26 executes a playback process of the separated voice data 46. Thereby, the user U1 can confirm the voice of a desired user output from the speaker 12.
[0117] Also, for example, as shown in Fig. 8(B), the recording device 1-1 can display a speech-to-text setting screen 159 on the touch panel display 55 of the device main body 3. The user U1 can perform an input operation regarding the speech recognition process of the voice data on the speech-to-text setting screen 159.
[0118] In FIG. 8(B), an example is shown in which the user U1 selects the voice of the user U1 (i.e., himself / herself) as the text conversion target. Therefore, when the user presses the execution button 57, the voice recognition unit 27 performs voice recognition processing on the separated voice data acquired from the management server 101 as described above. As a result, the voice of the user U1 is text-converted, and a text file is generated in a predetermined format. The generated text file is stored in the storage unit 42.
[0119] As described above, according to the recording system 100 including the recording device 1 according to the second embodiment, the bone conduction sound data including the voices of each user and the separated sound data including the voice of a specific user (i.e., the user selected from among a plurality of users) are stored in the storage unit 107, respectively. Therefore, it is possible to separate and use the voice of a specific user and the voices of other users.
[0120] (Third Embodiment) In the third embodiment, as in the case of the second embodiment, a case will be described in which a plurality of recording devices 1 are prepared and a plurality of persons each become a user of the corresponding recording device 1 (i.e., each person wears at least the bone conduction microphone 10).
[0121] FIG. 9 is a functional block diagram showing the configuration of a recording system 100 including the recording device 1 according to the third embodiment. In the recording device 1 and the recording system 100 shown in FIG. 9, the same components as those of the recording device 1 according to the first embodiment and the recording system 100 according to the second embodiment described above are denoted by the same reference numerals. Further, regarding the recording device 1 and the recording system 100 according to the third embodiment, matters not particularly mentioned below are the same as those of the recording device 1 according to the first embodiment and the recording system 100 according to the second embodiment.
[0122] The recording devices 1-1 to 1-N according to the third embodiment have the same configuration as the recording device 1 (see FIG. 3) according to the first embodiment described above.
[0123] In each of the recording devices 1-1 to 1-N, the bone conduction sound data 45 and the separated sound data 46 generated by the recording processing unit 25 are respectively transmitted to the management server 101 by the communication unit 29.
[0124] In the management server 101, the control unit 106 includes an audio data acquisition unit 111 and a bone conduction sound data synthesis unit 115.
[0125] The audio data acquisition unit 111 acquires the bone conduction sound data 45 and the separated sound data 46 respectively by communicating with each of the recording devices 1-1 to 1-N.
[0126] The bone conduction sound data synthesis unit 115 generates synthesized bone conduction sound data by synthesizing the bone conduction sound data corresponding to the bone conduction microphones 10 in two or more microphone sets. In the present embodiment, such two or more microphone sets can be selected by the user. The generated synthesized bone conduction sound data includes a synthesized voice in which the voices of two or more users are synthesized.
[0127] The generated synthesized bone conduction sound data, the bone conduction sound data 45 and the separated sound data 46 acquired from each of the recording devices 1-1 to 1-N are stored in the storage unit 107 and transmitted to the recording devices 1-1 to 1-N in response to requests from the users U1 to UN.
[0128] In the recording device 1-1 according to the third embodiment, for example, a setting screen similar to the voice recorder setting screen 156 shown in FIG. 8 described above can be displayed on the display of the device main body 3 (the same applies to the recording devices 1-2 to 1-N). The user U1 can select two or more users (i.e., playback targets) on the setting screen. Here, the selection of two or more playback targets (or synthesis targets) by the user U1 is synonymous with the selection of two or more microphone sets by the user U1. When the management server 101 acquires information regarding two or more playback targets from the recording device 1-1, the bone conduction sound data synthesis unit 115 synthesizes the bone conduction sound data corresponding to each bone conduction microphone 10 in the two or more microphone sets selected by the user U1 (i.e., the bone conduction sound data regarding the two or more selected playback targets) to generate synthesized bone conduction sound data.
[0129] After that, the recording device 1-1 acquires the synthesized bone conduction sound data generated by the management server 101. Subsequently, the voice playback unit 26 executes a playback process of the synthesized bone conduction sound data including only the voices of two or more users selected by the user U1 (for example, the voices of the other user U1 and the user U2 are synthesized). Thereby, the user U1 can confirm the voices of two or more users (i.e., the users selected by himself / herself) output from the speaker 12.
[0130] As described above, according to the recording system 100 including the recording device 1 according to the third embodiment, by storing the synthesized bone conduction sound data including the voices of a plurality of specific users and the air conduction sound data including the voices of a plurality of users in the storage unit 107, it is possible to separate and use the voice of a specific user and the voices of a plurality of users including that specific user.
[0131] As described above, the embodiments have been described as examples of the technology disclosed in the present application. However, the technology in the present disclosure is not limited to this, and can also be applied to embodiments with changes, replacements, additions, omissions, etc. Also, it is possible to form a new embodiment by combining the respective components described in the above embodiments.
Industrial Applicability
[0132] When the recording device according to the present disclosure uses a bone conduction microphone and an air conduction microphone to record the voices of a plurality of people including the user, it is possible to separate and use the user's voice from the voices of the people around the user. It is useful as a recording device, a recording system, and their recording methods that include a bone conduction microphone and an air conduction microphone and record the user's voice.
Description of Reference Numerals
[0133] 1: Recording device 2: Earphone 3: Recording device main body 4: Signal line 5: Ear 9: Exterior housing 9A: Case 9B: Cover 9C: Opening 10: Bone conduction microphone 11: Air conduction microphone 12: Speaker 13: Inner housing 13A: Passage 13B: Cylindrical part 15: Microphone rubber 16: Tip rubber 17: Resin member 18: Cable cap 21: Bone conduction sound transmission path 22: Air conduction sound transmission path 23: Connection path 25: Recording processing unit 26: Voice playback unit 27: Voice recognition unit 29: Communication unit 31: Amplification unit 33: LPF 35: Amplification unit 36: Arithmetic unit 40: First separation processing unit 41: Second separation processing unit 42: Memory unit 45: Bone conduction sound data 46: Separated sound data 51A - 51C: First to Third Medical Personnel 52: Patient 55: Touch Panel Display 56: Voice Recorder Setting Screen 57: Execution Button 59: Speech - to - Text Setting Screen 61: Air - Conduction Sound Data 100: Recording System 101: Management Server 103: Communication Network 105: Communication Unit 106: Control Unit 107: Memory Unit 111: Voice Data Acquisition Unit 112: Air - Conduction Sound Data Synthesis Unit 113: Voice Data Separation Unit 115: Bone - Conduction Sound Data Synthesis Unit 156: Voice Recorder Setting Screen 157: Execution Button 159: Speech - to - Text Setting Screen
Claims
1. A bone conduction microphone that picks up the user's voice and generates a bone conduction sound signal, A air conduction microphone that picks up the voices of a plurality of people including the user and generates an air conduction sound signal, An audio separation unit that generates a separated sound signal by performing processing for separating the user's voice based on the bone conduction sound signal from the voices of the plurality of people based on the air conduction sound signal, A storage unit that stores bone conduction sound data including the user's voice generated based on the bone conduction sound signal and separated sound data including the voices of the people other than the user generated based on the separated sound signal, respectively, A recording device comprising:
2. A bone conduction sound transmission path through which the bone conduction sound signal is transmitted, An air conduction sound transmission path through which the air conduction sound signal is transmitted, Further comprising: The recording device according to claim 1, wherein the audio separation unit generates the separated sound signal by subtracting the bone conduction sound signal transmitted through the bone conduction sound transmission path from the air conduction sound signal transmitted through the air conduction sound transmission path.
3. The recording device according to claim 2, further comprising a low-pass filter that extracts an audio signal based on the user's voice from the bone conduction sound signal used for generating the separated sound signal in the bone conduction sound transmission path.
4. The recording device according to claim 2, further comprising a noise canceller that removes or reduces a noise component in the bone conduction sound signal used for generating the separated sound signal in the bone conduction sound transmission path.
5. The recording device according to claim 1 or claim 2, further comprising an audio playback unit that can selectively play back the audio based on the bone conduction sound data and the audio based on the separated sound data.
6. A recording device according to claim 1 or claim 2, further comprising a speech recognition unit capable of executing a text conversion process by selectively executing a speech recognition process on speech based on the bone conduction sound data and speech based on the separated sound data.
7. A microphone set respectively used by each of a plurality of users, and a server that respectively acquires voice signals from each of the microphone sets, wherein each of the microphone sets includes a bone conduction microphone that picks up the voice of each user to generate a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including each user to generate an air conduction sound signal, wherein the server includes a voice data acquisition unit that respectively acquires bone conduction sound data based on the bone conduction sound signal and air conduction sound data based on the air conduction sound signal from each of the microphone sets, an air conduction sound data synthesis unit that generates synthesized air conduction sound data by synthesizing a plurality of the air conduction sound data acquired by the voice data acquisition unit, a voice data separation unit that generates separated sound data by executing a process for separating the voice of the user based on the bone conduction sound data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction sound data, and a storage unit that stores the bone conduction sound data and the separated sound data respectively, A recording system having the above components.
8. The recording system according to claim 7, further comprising an input device used by any one of the plurality of users, wherein the server acquires information on at least one of the microphone sets selected by an input operation to the input device, and the voice data separation unit generates the separated sound data based on the bone conduction sound data corresponding to the microphone sets other than the at least one selected microphone set.
9. A recording method of a recording device, The recording device includes a bone conduction microphone that picks up the user's voice and generates a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including the user and generates an air conduction sound signal, and is configured to generate a separated sound signal by performing a process for separating the user's voice based on the bone conduction sound signal from the voices of the plurality of people based on the air conduction sound signal, generate bone conduction sound data including the user's voice based on the bone conduction sound signal, and generate separated sound data including the voices of the people other than the user based on the separated sound signal, and record the bone conduction sound data and the separated sound data respectively. A recording method.
10. A recording method of a recording system, wherein the recording system includes a microphone set used by each of a plurality of users respectively, and a server that respectively acquires voice signals from the microphone sets, each of the microphone sets includes a bone conduction microphone that picks up the voice of each user and generates a bone conduction sound signal, and an air conduction microphone that picks up the voices of a plurality of people including each user and generates an air conduction sound signal, the server respectively acquires bone conduction sound data based on the bone conduction sound signal and air conduction sound data based on the air conduction sound signal from each of the microphone sets, generates synthesized air conduction sound data by synthesizing the acquired plurality of air conduction sound data, generates separated sound data by performing a process for separating the user's voice based on the bone conduction sound data corresponding to at least one of the microphone sets from the synthesized voice based on the synthesized air conduction sound data, and records the bone conduction sound data and the separated sound data respectively. A recording method.
Citation Information
Patent Citations
Singing sound recording and reproducing system
JP2010176041A
Speech apparatus and audio signal correction program
JP2017123554A
Signal processing device, signal processing method, and program
WO2020208926A1