Flesh conduction microphone system and voice acquisition method
The bone conduction microphone system addresses the issue of unstable sound collection by using a pressing unit and machine learning model to ensure consistent skin contact and convert bone conduction sounds into high-quality audible sounds.
Patent Information
- Application Number
- JP2021099704
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Existing bone conduction microphone systems face challenges in stably and reproducibly collecting bone conduction sounds due to dependence on the method of attachment and tightening, which affects sound quality.
A bone conduction microphone system with a propagation unit, microphone, processing unit, and pressing unit that ensures consistent contact with the skin surface using a predetermined pressing force, combined with a machine learning-based voice restoration model to convert bone conduction sounds into audible sounds.
Stable and highly reproducible collection of bone conduction sounds is achieved without requiring special technical knowledge, with improved sound quality and accuracy through the use of a pressing unit and machine learning model.
Smart Images

Figure 0007702721000001 
Figure 0007702721000002 
Figure 0007702721000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a bone conduction microphone system and a voice acquisition method.
Background Art
[0002] As a tool for realizing communication involving conversation without being restricted by the surrounding environment, for example, a bone conduction microphone system as disclosed in Japanese Patent Application Laid-Open No. 2008-42741 (hereinafter, Patent Document 1) can be mentioned.
[0003] For example, in an environment where vocalization is restricted such as a public facility, an environment where there are other people around, or a physical environment where normal voice cannot be vocalized due to an obstacle or the like, non-audible voice is collected using a bone conduction microphone system and converted into voice, thereby realizing communication involving conversation. Non-audible voice is voice (voiceless sound) that does not involve regular vibration of the vocal cords, and refers to vibration sound (breathing sound) that propagates through soft tissues in the body and is inaudible from the outside.
[0004] To collect such non-audible voice, the bone conduction microphone is, for example, mounted so as to be in close contact with the lower part of the auricle of the human body. Thereby, bone conduction sound that is vocalized in the vocal tract and transmitted through soft tissues in the body (muscles and fat other than bones) is collected.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
[0006] As a method of wearing a bone conduction microphone, there is a method of attaching it to the skin surface. In this method, the sound quality collected depends on the way of attachment. As another method of wearing a bone conduction microphone, there is also a method of wearing it by tightening it around the neck with a string or the like. However, even in this method, the sound quality collected depends on the way of tightening.
[0007] Therefore, one of the objectives of the present disclosure is to provide a bone conduction microphone system and a voice collection method that can stably collect highly reproducible bone conduction sounds.
[0008] According to an embodiment, a bone conduction microphone system is a bone conduction microphone system that collects bone conduction sounds, and includes a propagation unit that propagates bone conduction sounds by being in close contact with the skin surface, a microphone that is provided in contact with the propagation unit and converts the bone conduction sounds propagating through the propagation unit into electrical signals, a processing unit that executes a process of generating voice based on the electrical signals output from the microphone, and a pressing unit for pressing the propagation unit against the skin surface with a predetermined pressing force.
[0009] According to an embodiment, a method for collecting voice by bone conduction is a method for collecting voice using a microphone for collecting bone conduction sounds, and includes using a pressing unit to press the propagation unit against the skin surface with a predetermined pressing force to bring the propagation unit into close contact with the skin surface, propagating bone conduction sounds to the propagation unit in close contact with the skin surface, converting the bone conduction sounds propagating through the propagation unit into electrical signals by a microphone provided in contact with the propagation unit, and generating voice based on the electrical signals output from the microphone.
[0010] Further details will be described as embodiments below.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0012] <1. Overview of the Bone Conduction Microphone System and the Voice Collection Method>
[0013] (1) The bone conduction microphone system according to the embodiment is a bone conduction microphone system that collects bone conduction sound, and includes a propagation unit that propagates bone conduction sound by being in close contact with the skin surface, a microphone that is provided in contact with the propagation unit and converts the bone conduction sound propagating through the propagation unit into an electrical signal, a processing unit that executes a process of generating sound based on the electrical signal output from the microphone, and a pressing unit that presses the propagation unit against the skin surface with a predetermined pressing force.
[0014] Since the bone conduction microphone system includes the pressing unit, the propagation unit can be pressed against the skin surface with a predetermined pressing force without sticking the propagation unit to the skin surface or pressing it with a hand or the like. As a result, it is possible to stably collect bone conduction sound with high reproducibility without requiring special technology for using the bone conduction microphone.
[0015] (2) Preferably, the process of generating sound includes inputting the electrical signal as an input value to a voice restoration model to obtain an output value, and the voice restoration model is a machine learning model that is machine-learned to output an audible sound corresponding to the bone conduction sound with the electrical signal based on the bone conduction sound as the input value. This makes it possible to easily convert bone conduction sound into audible sound.
[0016] (3) Preferably, the voice restoration model is a machine learning model that is machine-learned using, as learning data, an electrical signal obtained by converting bone conduction sound of the same user and a voice sound corresponding to the bone conduction sound. The voice sound may be, for example, sound collected by a bone conduction microphone or air conduction sound recorded by a normal microphone. This makes it possible to reproduce bone conduction sound with the voice sound of the user who collected the sound.
[0017] (4) Preferably, the pressing part has a guide part for attaching a microphone to the object for collecting the bone conduction sound. Thereby, it becomes possible to attach the microphone at a position suitable for collecting the bone conduction sound. Therefore, it is possible to stably collect the bone conduction sound with high reproducibility without requiring special technology for attaching the microphone.
[0018] (5) Preferably, the processing part stores in advance, as a determination signal, an electrical signal based on the bone conduction sound for a specified sound, and outputs a determination result based on a comparison between the electrical signal obtained from the bone conduction sound propagated through the propagation part and the determination signal. When using, as the determination signal, an electrical signal based on the bone conduction sound collected when a microphone is attached at a position suitable for collecting the bone conduction sound, by comparing these, it is possible to determine whether the position of the microphone at the time of collecting the bone conduction sound is appropriate. By outputting the determination result, it becomes possible to adjust the position of the microphone, and as a result, it becomes possible to collect the bone conduction sound with high reproducibility.
[0019] (6) Preferably, the pressing part has an adjustment part capable of adjusting the pressing force. Thereby, it becomes easier to adjust the pressing. As a result, it becomes possible to press the microphone against the skin surface with an appropriate pressing force, and it becomes possible to collect the bone conduction sound with high reproducibility.
[0020] (7) The voice collection method using bone conduction according to the embodiment is a voice collection method using a bone conduction microphone for collecting bone conduction sound, which uses a pressing part to press the propagation part against the skin surface with a predetermined pressing force, makes the propagation part adhere to the skin surface, propagates the bone conduction sound to the propagation part that is adhered to the skin surface, converts the bone conduction sound propagating through the propagation part into an electrical signal by a microphone provided in contact with the propagation part, and generates voice based on the electrical signal output from the microphone. By using the pressing part, the propagation part can be pressed against the skin surface with a predetermined pressing force without sticking the propagation part to the skin surface or pressing it with a hand or the like. As a result, it is possible to stably collect bone conduction sound with high reproducibility without requiring special technology for using the bone conduction microphone.
[0021] <2. Examples of bone conduction microphone system and voice collection method>
[0022] FIG. 1 is a schematic diagram showing an example of the configuration of a bone conduction microphone system 100 according to the present embodiment. FIG. 2 is a schematic diagram showing the wearing state of the bone conduction microphone 1 on the user 2. In FIG. 2, a state in which the user 2 facing upward is viewed from directly above is shown.
[0023] The bone conduction microphone system 100 includes a bone conduction microphone 1. The bone conduction microphone 1 is worn on the user 2 and is used to collect the bone conduction sound of the user 2. The bone conduction sound refers to the sound obtained from the vibrations of the skin and muscles induced by vocalization. As an example, the bone conduction microphone 1 detects the vibrations of the skin and muscles around the user's throat and converts the vibrations into an electrical signal.
[0024] The bone conduction microphone system 100 includes a pressing part 3. The pressing part 3 presses the bone conduction microphone 1 against the skin surface of the user 2 with a predetermined pressure. Preferably, the pressing part 3 presses the bone conduction microphone 1 against the position 202 behind the auricle 201 of the user 2. Thereby, it becomes easier for the bone conduction microphone 1 to detect the vibrations of the skin and muscles around the user's throat, and it becomes easier to obtain voice.
[0025] The pressing part 3 has a guide part 31. The guide part 31 has a partially missing circular shape and is a belt-shaped member that is wound around the user's neck. In the examples of FIGS. 1 and 2, the guide part 31 is semi-circular.
[0026] The diameter of the guide part 31 is such that the distance L1 between the contact surfaces 1A of both bone conduction microphones 1 is smaller than the distance L2 between the left and right positions 202 of the mounted user 2. The distance L2 is assumed to be the distance between the left and right positions 202 of the smallest-sized user among various users assumed to be the objects for collecting bone conduction sound using the bone conduction microphone 1.
[0027] The pressing part 3 is formed of a material in which at least the guide part 31 has flexibility. As the flexible material, for example, a resin material or the like is used. The resin material includes, for example, at least one material selected from the group of polyvinyl chloride, polyethylene, polypropylene, ABS resin, AES resin, polycarbonate, modified polyphenylene ether, polyethylene terephthalate, polybutylene terephthalate, and nylon. Thereby, the formation of the arc-shaped guide part 31 becomes easy, and a predetermined pressing force as described later can be obtained due to the flexibility.
[0028] Adjusting parts 32 are provided at both ends of the guide part 31, and bone conduction microphones 1 are attached to both adjusting parts 32, respectively. Both bone conduction microphones 1 are attached to the adjusting parts 32 such that the contact surface 1A, which is the surface on the side that is brought into close contact with the skin surface, faces the center side of the arc of the guide part 31. Thereby, when attached to the user as described later, the bone conduction microphone 1 is pressed against the user's skin surface with a predetermined pressure.
[0029] Referring to FIG. 2, the bone conduction microphone 1 is mounted with the pressing portion 3 rotated around the neck 200B of the user 2 so that the inside of the apex 31A of the arc of the guide portion 31 contacts the position 203 directly behind the neck 200B of the user 2, and the bone conduction microphone 1 is directed left and right of the head 200A of the user 2. As a result, the bone conduction microphone 1 is positioned near the position 202 behind the left and right auricles 201 of the user 2 and is guided to an appropriate position.
[0030] Since the distance L1 is smaller than the distance L2, when the user 2 mounts the bone conduction microphone 1 around the neck 200B as shown in FIG. 2, the guide portion 31 slightly opens due to the flexibility of the pressing portion 3. For this reason, a force F directed inward is generated in the adjustment portion 32 extending from the guide portion 31. Thereby, the pressing portion 3 presses the bone conduction microphone 1 against the skin surface of the user 2 with a force F which is a predetermined pressing force at the position 202.
[0031] The adjustment portion 32 can adjust the pressing force for pressing the bone conduction microphone 1 against the skin surface of the user 2. As an example, the adjustment portion 32 can adjust the pressing force by attaching the bone conduction microphone 1 so as to be movable in the direction of arrow B. Arrow B is the direction in which the adjustment portion 32 faces from both ends of the guide portion 31, and as an example, it is the tangential direction at both ends of the circular guide portion 31.
[0032] FIG. 3 is an enlarged view of part A in FIG. 1 and is a schematic diagram showing an example of the adjustment portion 32. In FIG. 3, the adjustment portion 32 is shown as viewed in the direction of arrow C in FIG. 1. Referring to FIGS. 1 and 3, the adjustment portion 32 has a mounting member 21 attached to the bone conduction microphone 1. The bone conduction microphone 1 is fixedly attached to the mounting member 21 with a relative position. As an example, the bone conduction microphone 1 is attached to the mounting member 21 with an adhesive or the like.
[0033] At a position of the attachment member 21 different from the position where the bone conduction microphone 1 is attached, a fixing hole 21A is provided. The fixing hole 21A is a hole through which a rod-shaped fixing member 23 described later can be inserted.
[0034] The adjustment part 32 has a slide part 22 for attaching the attachment member 21. The slide part 22 is a part extending from the tip of the guide part 31. A slide hole 22A is provided in the slide part 22. The slide hole 22A is an elongated hole having a length W whose longitudinal direction coincides with the extending direction. The slide hole 22A is a hole through which the fixing member 23 described later can be inserted.
[0035] The adjustment part 32 has a rod-shaped fixing member 23. The fixing member 23 is a member that can fix the attachment member 21 to the slide part 22 by being inserted into the fixing hole 21A and the slide hole 22A. The fixing member 23 is a bolt as an example.
[0036] The position of the attachment member 21 is variable along the longitudinal direction of the slide hole 22A. That is, the position of the attachment member 21 is variable within a range of length W in the longitudinal direction. By fixing the attachment member 21 at an arbitrary position within the range of length W with the fixing member 23, the longitudinal position of the attachment member 21 with respect to the slide part 22 can be adjusted.
[0037] In the examples of FIGS. 1 and 2, since the guide part 31 is semicircular, the two adjustment parts 32 are parallel. In this case, when the bone conduction microphone 1 is worn on the user 2 as shown in FIG. 2, the direction of arrow B coincides with the front-rear direction D of the user 2. That is, in this case, the adjustment part 32 can adjust the position of the bone conduction microphone 1 in the front-rear direction D of the user 2.
[0038] By adjusting the position of the bone conduction microphone 1, it is possible to place the bone conduction microphone 1 at the position 202 of the user 2 suitable for collecting bone conduction sound, and it is also possible to adjust the angle of the bone conduction microphone 1 with respect to the user 2 at the position 202 or in the vicinity of the position 202. As a result, the force generated by the flexibility of the guide portion 31 acts on the angle at which the bone conduction microphone 1 is pressed against the skin surface of the user 2, that is, the pressing force for pressing the bone conduction microphone 1 against the skin surface of the user 2 can be adjusted.
[0039] FIG. 4 is a schematic cross-sectional view taken along line E-E of FIGS. 1 and 3 of the bone conduction microphone 1. Referring to FIG. 4, the bone conduction microphone 1 has a propagation portion 11. The propagation portion 11 propagates the bone conduction sound of the user 2 by being in close contact with the skin surface of the user 2.
[0040] The propagation portion 11 is inside the case 10 which is the outer shell, and is arranged so that the surface 11A on the contact surface 1A side contacts the surface 10A of the case 10 which becomes the contact surface 1A. Thereby, when the bone conduction microphone 1 is pressed against the skin surface of the user 2, the surface 11A of the propagation portion 11 comes into close contact with the skin surface of the user 2.
[0041] Preferably, as shown in FIG. 4, the propagation portion 11 is arranged in the case 10 such that the surface 11A is located outside the surface 10A and has a convex shape. Thereby, when the bone conduction microphone 1 is pressed against the skin surface of the user 2, the surface 11A of the propagation portion 11 comes into closer contact with the skin surface of the user 2.
[0042] The bone conduction microphone 1 has a microphone 12. The microphone 12 is provided in contact with the propagation portion 11 and converts the bone conduction sound propagating through the propagation portion 11 into an electrical signal. A communication line 13 is connected to the microphone 12, and the converted electrical signal is sent out via the communication line 13.
[0043] The bone conduction microphone system 100 includes a processing device 5. The processing device 5 may be a dedicated device composed of a general computer, or may be installed in a terminal device such as a smartphone.
[0044] The processing device 5 is connected to the bone conduction microphone 1 via a communication line 13 and can communicate therewith, and receives an electrical signal from the bone conduction microphone 1. The processing device 5 executes a process of generating sound based on the electrical signal from the bone conduction microphone 1. The communication between the processing device 5 and the bone conduction microphone 1 may be wired communication or wireless communication as shown in FIG. 1.
[0045] FIG. 5 is a block diagram showing an outline of the processing device 5. The processing device 5 is composed of a general computer having a processor 51 and a memory 52. The processor 51 is, for example, a CPU (Central Processing Unit).
[0046] The memory 52 may be a primary storage device or a secondary storage device. The memory 52 stores a program 521 executed by the processor 51. The processor 51 executes arithmetic processing by executing the program 521 stored in the memory 52.
[0047] The processing device 5 includes a communication device 53. The communication device 53 is, as an example, a communication module. The communication device 53 communicates with the microphone 12 via the communication line 13. The communication device 53 receives an electrical signal output from the microphone 12 by communicating with the microphone 12. The communication device 53 inputs the received electrical signal to the processor 51.
[0048] The processing device 5 is connected to a speaker 55 which is an example of an output device that outputs information based on audible sound. In this case, the audible sound itself is output from the output device. As another example, the output device may be a display. In this case, the output device outputs the audible sound after being converted into characters.
[0049] The arithmetic processing executed by the processor 51 includes a generation process 511. The generation process 511 includes generating voice based on the electrical signal output from the microphone 12. Generating voice includes, as an example, inputting the electrical signal as an input value to the voice restoration model 512 to obtain an output value. Thereby, the processor 51 can obtain voice as the output value from the voice restoration model 512.
[0050] The voice restoration model 512 is a machine learning model that is machine-learned to output audible sound corresponding to bone-conducted sound as an output value with the electrical signal based on bone-conducted sound as an input value. Thereby, the audible sound corresponding to the bone-conducted sound collected by the microphone 12 can be obtained as the output value from the voice restoration model 512. That is, the processor 51 can easily perform the restoration of voice from bone-conducted sound using the voice restoration model 512.
[0051] Preferably, the voice restoration model 512 is a machine learning model that is machine-learned to output the voice uttered by the user corresponding to the bone-conducted sound as an output value with the electrical signal converted from the bone-conducted sound of the same user as an input value. Thereby, the voice uttered by the user corresponding to the bone-conducted sound collected from the user by the microphone 12 can be obtained as the output value from the voice restoration model 512. The voice uttered corresponding to the bone-conducted sound is so-called clean voice, and may be, for example, audible sound collected (recorded) by air conduction using a normal microphone, or may be the bone-conducted sound collected by the microphone 12 with good sound quality.
[0052] By using the voice restoration model 512 whose output value is the voice uttered by the same user, the bone-conducted sound collected from user 2 can be reproduced by the clean voice of user 2, and more realistic reproducibility is realized.
[0053] The arithmetic processing executed by the processor 51 includes output processing 515. The output processing 515 includes causing an output device to output information based on the audible sound obtained by the generation processing 511. Specifically, the output processing 515 includes causing the audible sound obtained by the generation processing 511 to be output from the speaker 55. As a result, the audible sound corresponding to the bone-conducted sound collected by the microphone 12 is output from the speaker 55. Therefore, the bone-conducted sound of the user 2 is restored to an audible sound and can be heard using the speaker 55.
[0054] Preferably, the arithmetic processing executed by the processor 51 includes determination processing 513. The determination processing 513 includes outputting a determination result based on a comparison between the electrical signal output from the microphone 12 and the determination signal 514. The determination signal 514 is an electrical signal based on the bone-conducted sound of a prescribed sound and is stored in advance in the processor 51.
[0055] The determination signal 514 is an electrical signal obtained from the bone-conducted sound collected by the microphone 12 when the bone-conduction microphone 1 is pressed against the skin surface of the user 2 with an ideal pressing force. Specifically, a determination word (for example, "test") is prescribed in advance, and the electrical signal obtained from the bone-conducted sound collected by the microphone 12 when the user 2 utters the determination word with the bone-conduction microphone 1 pressed against the skin surface of the user 2 with an ideal pressing force is stored as the determination signal 514.
[0056] The determination result obtained by collecting the bone-conducted sound of the determination word uttered by the user 2 with the microphone 12 and comparing the obtained electrical signal with the determination signal 514 indicates whether the pressing force of the bone-conduction microphone 1 against the skin surface at the time of collecting the bone-conducted sound matches the ideal pressing force or is within an allowable range compared to the ideal pressing force. Therefore, by outputting the result of the determination processing 513, the pressing force can be adjusted by the adjustment unit 32 as necessary. As a result, the pressing force of the bone-conduction microphone 1 against the skin surface can be made the ideal pressing force, and the restoration accuracy to the audible sound can be improved.
[0057] When the processor 51 has a plurality of voice restoration models 512, preferably, the determination process 513 may include determining a voice restoration model 512 suitable for the user 2. The plurality of voice restoration models 512 may be, for example, the uttered voice as the output value being the uttered voice for each user or the uttered voice for each environment.
[0058] In this case, as an example, the determination signal 514 is prepared for each voice restoration model 512. In the determination process 513, the processor 51 collects the bone-conducted sound of the determination word uttered by the user 2 by the microphone 12, and determines the voice restoration model 512 to be used by comparing the obtained electrical signal with each of the plurality of determination signals 514. Thereby, for example, the bone-conducted sound collected by the microphone 12 is restored as the uttered voice corresponding to the user 2 or the uttered voice corresponding to the environment of the user 2.
[0059] Preferably, the arithmetic process executed by the processor 51 includes a learning process 516. The learning process 516 includes performing learning of the voice restoration model 512. In the learning process 516, the processor 51 uses, as learning data, the electrical signal obtained by converting the bone-conducted sound collected by the microphone 12 and the uttered voice corresponding to the bone-conducted sound for the same user.
[0060] The processor 51 may receive an input of the uttered voice corresponding to the bone-conducted sound from another device. Also, as shown in FIG. 5, the processing device 5 may be connected to a microphone 56 that collects voice by air conduction, and the input from the microphone 56 may be used as the uttered voice corresponding to the bone-conducted sound.
[0061] When the learning process 516 is performed by the processor 51, the voice restoration model 512 is learned using the combination of the myoelectric sound and the uttered sound by the user 2. When the voice restoration model 512 is a machine learning model that has been pre-trained using the combination of the myoelectric sound and the uttered sound by the user 2, the restoration accuracy by the voice restoration model 512 can be further improved by performing the learning process 516 by the processor 51.
[0062] When the voice restoration model 512 is not a machine learning model that has been pre-trained using the combination of the myoelectric sound and the uttered sound by the user 2, when the learning process 516 is performed by the processor 51, while the user 2 continues to use the myoelectric microphone system 100, the voice restoration model 512 is updated to a machine learning model that uses the myoelectric sound by the user 2 as an input value and the uttered sound by the user 2 as an output value. As a result, the myoelectric sound comes to be reproduced by the utterance of the user 2, and more realistic reproducibility is realized.
[0063] FIG. 6 is a flowchart showing an example of the flow of a voice collection method using a myoelectric sound collection microphone according to the present embodiment. Referring to FIG. 6, first, using the pressing portion 3, the propagation portion 11 of the myoelectric microphone 1 is pressed against the skin surface at the position 202 of the user 2 with a predetermined pressing force to bring the propagation portion 11 into close contact with the skin surface (step S100). By propagating the myoelectric sound to the propagation portion 11 in close contact with the skin surface, the myoelectric sound of the word for determination uttered by the user 2 is collected by the microphone 12 (step S101). The collected myoelectric sound is converted into an electrical signal by the microphone 12 (step S103).
[0064] The obtained electrical signal is compared with the determination signal 514 stored in advance, and a determination result is output (steps S105, S107). As an example, it is determined whether or not the value based on the obtained electrical signal is within a preset allowable range from the value based on the determination signal 514.
[0065] As a result of the comparison, if it is not within the allowable range (NO in step S105), the result is notified by error output or the like (step S107). Thereby, the pressing force of the bone conduction microphone 1 on the skin surface of the user 2 can be adjusted.
[0066] Preferably, in the voice acquisition method, steps S101 to S107 are repeated until the value based on the obtained electrical signal is within the allowable range. As a result, the pressing force of the bone conduction microphone 1 on the skin surface of the user 2 can be set to an ideal pressing force.
[0067] When the value based on the obtained electrical signal is within the allowable range (YES in step S105), the bone conduction sound of the user 2 is collected by the microphone 12 (step S109) and converted into an electrical signal (step S111). The obtained electrical signal is input as an input value to the voice restoration model 512, and the corresponding audible sound is obtained as the output value (step S113). Thereby, a voice corresponding to the bone conduction sound is generated. The generated voice is output from the speaker 55 as an audible sound (step S115).
[0068] The inventor conducted a verification experiment to evaluate the reproducibility of the bone conduction sound obtained by the voice acquisition method according to the embodiment. For the verification experiment, the bone conduction microphone system 100 shown in FIGS. 1, 3, 4, and 5 was used.
[0069] The evaluation of the reproducibility of the bone conduction sound was performed by comparing the audible sound obtained by restoring the bone conduction sound collected by the bone conduction microphone system 100 with the air conduction sound recorded by the air conduction microphone for the same utterance. It can be said that the closer the air conduction sound is, the higher the reproducibility of the audible sound, that is, the higher the reproducibility of the bone conduction sound collected.
[0070] In the verification experiment, bone conduction was collected under the following conditions 1 to 4. In each condition, the same voice restoration model 512 was used for restoring the bone conduction sound. The values to be compared were the time changes of the amplitude and frequency. The audible sound obtained from the bone conduction sound was output from the same speaker 55 and collected with the same microphone. Condition 1: Attach the bone conduction microphone 1 to position 202 and press it against the skin surface with a pressing force F1 that maintains the contact of the bone conduction microphone 1 with position 202. Condition 2: Attach the bone conduction microphone 1 to position 202 and press it against the skin surface with a pressing force F2 by the pressing portion 3. Condition 3: Attach the bone conduction microphone 1 so that it contacts only one side at a position 205 different from position 202, and press it against the skin surface with a pressing force F1 that maintains the contact of the bone conduction microphone 1 with position 205. Condition 4: Attach the bone conduction microphone 1 so that it contacts only one side at a position 205 different from position 202, and press it against the skin surface with a pressing force F2 by the pressing portion 3.
[0071] Figure 7 is a schematic diagram showing the wearing state of the bone conduction microphone 1 on the user 2 in Conditions 3 and 4. In the case of Conditions 3 and 4, the arc vertex 31A of the guide portion 31 is largely shifted backward from the position 203 directly behind the neck 200B of the user 2, and the guide portion 31 is worn around the neck 200B. As a result, the bone conduction microphone 1 is located at a position 205 behind position 202. Thereby, the spread in the guide portion 31 becomes smaller than when the vertex 31A is worn to hit the position 203. Therefore, the pressing force for pressing the bone conduction microphone 1 against the skin surface is smaller than when the vertex 31A is worn to hit the position 203.
[0072] Also, in Conditions 3 and 4, one of the bone conduction microphones 1 provided on both the left and right sides was brought into contact with position 205, and the other was in a state of hardly being in contact. Even in this state, it is possible to collect the bone conduction sound with the bone conduction microphone 1 on the side in contact with the skin surface.
[0073] Figures 8 to 11 respectively show the time variations of the amplitude and frequency of the audible sound obtained by restoring the bone-conducted sound obtained under Conditions 1 to 4. Figure 12 shows the time variations of the amplitude and frequency of the air-conducted sound. Comparing Figures 8, 9 with Figures 1, 11, it was found that both the amplitude and frequency appeared larger under Conditions 1 and 2 than under Conditions 3 and 4, and were closer to the amplitude and frequency in Figure 12. From this, it was found that contacting the bone-conduction microphone 1 to Position 202 rather than Position 205 has higher reproducibility.
[0074] Comparing each of Figures 8 and 9 with Figure 12, although it can be said that there is a certain degree of reproducibility in the case of Condition 1, it was found that Condition 2 has higher reproducibility. On the other hand, comparing each of Figures 10 and 11 with Figure 12, the reproducibility was extremely low under Condition 3, and under Condition 4, although the reproducibility was higher than that under Condition 3, it was found that the reproducibility was lower than that under Condition 1.
[0075] From the above results, it was verified that higher reproducibility bone-conducted sound is collected when the bone-conduction microphone 1 is brought into contact with the skin surface at Position 202, which is a position suitable for collecting bone conduction. Also, it was verified that higher reproducibility bone-conducted sound is collected when the bone-conduction microphone 1 is pressed against the skin surface with an appropriate pressing force F2. In addition, although the position where the bone-conduction microphone 1 is brought into contact is not appropriate, it was verified that reproducible bone-conducted sound can be collected using only one of the bone-conduction microphones 1.
[0076] In the bone-conduction microphone system 100 according to the embodiment, since the bone-conduction microphone 1 is configured to be worn on the user 2 using the pressing portion 3, when the user 2 appropriately wears it around the neck 200B, the bone-conduction microphone 1 comes into contact with Position 202 and is pressed against the skin surface with an appropriate pressing force by the pressing portion 3. Therefore, it was verified that in the bone-conduction microphone system 100, bone-conducted sound can be collected in a state like Condition 2 regardless of who uses it, and stable and highly reproducible bone-conducted sound is collected.
[0077] <3. Addendum> The present invention is not limited to the above-described embodiments, and various modifications are possible.
Explanation of Reference Numerals
[0078] 1: Bone conduction microphone 1A: Contact surface 2: User 3: Pressing part 5: Processing device 10: Case 10A: Surface 11: Propagation part 11A: Surface 12: Microphone 13: Communication line 21: Mounting member 21A: Fixing hole 22: Slide part 22A: Slide hole 23: Fixing member 31: Guide part 31A: Vertex 32: Adjusting part 51: Processor 52: Memory 53: Communication device 55: Speaker 56: Microphone 100: Bone conduction microphone system 200: User 200A: Head 200B: Neck 201: Pinna 511: Generation process 512: Voice restoration model 513: Judgment process 514: Judgment signal 515: Output process 516: Learning process 521: Program B: Arrow C: Arrow D: Direction F: Force F1: Pressing force F2: Pressing force L1: Distance L2: Distance
Claims
1. A bone conduction microphone system for collecting bone conduction sounds, comprising: a propagation unit that is adhered to the skin surface to propagate the bone conduction sound; a microphone provided in contact with the propagation unit for converting the bone conduction sound propagating through the propagation unit into an electrical signal; a processing unit that executes a process of generating sound based on the electrical signal output from the microphone; a pressing unit for pressing the propagation unit against the skin surface with a predetermined pressing force; wherein the processing unit stores in advance, as a determination signal, an electrical signal based on a bone conduction sound for a specified sound, and outputs a determination result based on a comparison between the electrical signal obtained from the bone conduction sound propagating through the propagation unit and the determination signal. A bone conduction microphone system.
2. The determination signal includes a determination signal that is an electrical signal obtained from a bone conduction sound collected by the microphone when a word for determination is uttered in a state where the microphone is pressed against the skin surface with an ideal pressing force, and the pressing unit has an adjustment unit capable of adjusting the pressing force based on the determination result. The bone conduction microphone system according to Claim 1.
3. The process of generating the sound includes inputting the electrical signal as an input value into a sound restoration model to obtain an output value, and the sound restoration model is a machine learning model that is machine-learned to output, as an output value, an audible sound corresponding to the bone conduction sound, with the electrical signal based on the bone conduction sound as an input value. The bone conduction microphone system according to Claim 1 or 2.
4. The sound restoration model is a machine learning model that is machine-learned using, as learning data, the electrical signal converted from the bone conduction sound for the same user and the uttered sound corresponding to the bone conduction sound. The bone conduction microphone system according to Claim 3.
5. The sound restoration model is a plurality of sound restoration models, the determination signal includes a plurality of determination signals that are electrical signals based on bone conduction sounds prepared for each of the plurality of sound restoration models, the processing unit outputs a determination result based on a comparison between the electrical signal obtained from the bone conduction sound propagating through the propagation unit and the plurality of determination signals, and determines, based on the determination result, a sound restoration model into which the input value is to be input from the plurality of sound restoration models. The bone conduction microphone system according to Claim 3.
6. The pressing part has a guide part for attaching the microphone to an object for collecting the bone-conducted sound. The bone-conducted microphone system according to any one of claims 1 to 5. **Claim 7** An audio acquisition method using a bone-conducted sound acquisition microphone for acquiring bone-conducted sound, using a pressing part to press a propagation part against the skin surface with a predetermined pressing force to bring the propagation part into close contact with the skin surface, propagating the bone-conducted sound to the propagation part in close contact with the skin surface, converting the bone-conducted sound propagating through the propagation part into an electrical signal by a microphone provided in contact with the propagation part, comprising: a processing unit generating audio based on the electrical signal output from the microphone, wherein the processing unit stores in advance, as a determination signal, an electrical signal based on a bone-conducted sound for a specified sound, and outputs a determination result based on a comparison between the electrical signal obtained from the bone-conducted sound propagating through the propagation part and the determination signal. Audio acquisition method.
Citation Information
Patent Citations
Contact type microphone
JP2006287810A
Flesh conducted sound pickup microphone
JP2008042741A
Information input device, specific frequency extraction method, and specific frequency extraction program
JP2014071823A
Communication device
JP2014143582A
Non-audible murmur input alarm device, method, and program
WO2008007616A1