Device and Method for Processing High Quality Voice signal Using Removing Ambient Noise based on Multi Sensor Signal Fusion
The multi-sensor signal fusion method effectively addresses noise interference in voice signal processing by using an accelerometer and voice microphone sensor to enhance voice clarity and robustness in noisy environments.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- INTUS CO LTD
- Filing Date
- 2022-09-19
- Publication Date
- 2026-07-21
AI Technical Summary
Conventional voice signal processing technologies face challenges in obtaining high-quality voice signals due to noise interference, particularly in environments where standard microphones fail to effectively filter both regular and irregular noise, and vocal cord microphones struggle with reduced clarity of high-frequency components.
A high-quality voice signal processing method using multi-sensor signal fusion, combining an accelerometer sensor and a voice microphone sensor to extract and remove noise based on speech interval information, synthesize low-frequency components, and restore voice signals by varying the synthesis ratio based on noise levels.
Enables robust voice signal processing in noisy environments by efficiently removing noise and enhancing voice clarity through multi-sensor fusion, ensuring precise voice signal restoration.
Smart Images

Figure 112022098282614-PAT00003_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to voice signal processing, and more specifically, to a high-quality voice signal processing apparatus and method based on multi-sensor signal fusion that enables voice signal processing robust to external noise environments through multi-sensor signal fusion using an accelerometer sensor (ACC) and a voice microphone sensor (MIC). Background Technology
[0002] Generally, a microphone is a means of converting a sender's voice into an electrical signal and transmitting it to a receiver.
[0003] These microphones include wired and wireless types, and most are designed to transmit sound coming from the user's mouth while mounted or positioned near the user's mouth.
[0004] However, due to the inconvenience of standard microphones—such as excessive noise, the inability to use them while wearing helmets or protective suits, and unclear voice transmission—and even among the general public as well as specialized personnel like security guards and special agents, there is a growing trend of using vocal cord microphones that transmit voice through the vibration of the vocal cords.
[0005] Unlike standard microphones, vocal cord microphones transmit voice signals through the vibration of the vocal cords, eliminating the need to speak loudly, making them useful for security personnel. Additionally, because they do not pick up noise, they can transmit clearer voice signals, making them useful for the general public as well.
[0006] Meanwhile, since vocal cord microphones collect vibration signals resulting from vocal cord resonance and convert them into electrical signals, they require a very high level of technical expertise, such as being perfectly protected from the external environment and capable of eliminating noise when collecting signals from vocal cord resonance.
[0007] Figure 1 is a frequency characteristic graph of a voice microphone and a vocal cord microphone, and Figure 2 is a configuration diagram showing the principle of noise removal through active noise canceling.
[0008] Generally, vocal cord microphones use inductive vibration sensors as a means of vibration conversion.
[0009] The inductive vibration sensor is composed of a structure including a vibrating membrane, a coil, and a permanent magnet. It utilizes the principle that when a light coil is connected to the vibrating membrane and the membrane and coil vibrate together, the magnetic field around the coil changes due to the permanent magnet in the center of the coil, thereby generating a voltage in the coil, to convert the vibration of the vocal cords into an electrical signal.
[0010] However, these inductive vibration sensors have a characteristic in which the frequency response decreases in proportion to the frequency. For this reason, inductive vibration sensors fail to properly transmit high-frequency components of speech compared to low frequencies, leading to a problem of reduced speech clarity.
[0011] Although technology using accelerometer sensors in vocal cord microphones is being introduced, this also has limitations in acquiring high-quality voice signals.
[0012] Meanwhile, although active noise canceling technology is used to acquire high-quality voice signals in microphone environments, it effectively responds to and processes regular low-frequency noise; however, it is insufficient for irregular high-frequency noise and has problems such as actually generating noise in certain environments.
[0013] Therefore, there is a need to develop new technology that can obtain high-quality voice signals by processing input voice signals containing noise. Prior art literature
[0014] Republic of Korea Published Patent No. 10-2021-0101644 Republic of Korea Registered Patent No. 10-0873094 Republic of Korea Published Patent No. 10-2018-0093363 The problem to be solved
[0015] The present invention aims to solve the problems of conventional voice signal processing technology by providing a high-quality voice signal processing device and method through ambient noise removal based on multi-sensor signal fusion, which enables voice signal processing robust to external noise environments through multi-sensor signal fusion using an accelerometer sensor (ACC) and a voice microphone sensor (MIC).
[0016] The present invention aims to provide a high-quality voice signal processing device and method through ambient noise removal based on multi-sensor signal fusion, which enables efficient voice signal processing by extracting and removing noise from the output signal of a voice microphone sensor (MIC) using speech interval information of an accelerometer sensor (ACC).
[0017] The present invention aims to provide a high-quality voice signal processing device and method through ambient noise removal based on multi-sensor signal fusion, which determines the level of noise extracted from the output signal of a voice microphone sensor (MIC), and synthesizes the low-frequency component of an accelerometer sensor (ACC) and the low-frequency component of a voice microphone sensor (MIC) by varying the synthesis ratio based on the determined noise level to improve the quality of the voice signal.
[0018] The present invention aims to provide a high-quality voice signal processing apparatus and method based on multi-sensor signal fusion through ambient noise removal, which can obtain a high-quality voice signal by first removing noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2) using speech interval information using the output signal of an accelerometer sensor (ACC), second removing noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2) using a beamforming algorithm, and third removing noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2) again using speech interval information.
[0019] The present invention aims to provide a high-quality voice signal processing device and method based on multi-sensor signal fusion and ambient noise removal, which enables precise voice signal processing by including more of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) during the synthesis process, if the noise level extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, including more of the low-frequency component of the accelerometer sensor (ACC), and if it is lower, including more of the low-frequency component of the voice microphone sensor (MIC).
[0020] Other objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0021] A high-quality voice signal processing device based on multi-sensor signal fusion and ambient noise removal according to the present invention for achieving the above-mentioned purpose comprises: a voice microphone sensor (MIC) that senses and outputs a voice signal of a speaker; an accelerometer sensor (ACC) that detects the vibration of the speaker's vocal cords and outputs a signal; a noise reduction processing MCU that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC), synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the synthesis ratio based on the noise level extracted from the output signal of the voice microphone sensor (MIC) using the vocalization segment information, and restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC); and a wireless communication module that outputs the restored voice signal externally.
[0022] Here, the noise reduction processing MCU is characterized by including more of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) in the synthesis process of the low-frequency component of the accelerometer sensor (ACC) and the voice microphone sensor (MIC), if the level of noise extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, and including more of the low-frequency component of the voice microphone sensor (MIC) (31) if it is lower.
[0023] The noise reduction processing MCU is characterized by using speech interval information to determine, extract, and remove the speech microphone sensor (MIC) output signal outside the speech interval as noise, and separating the speech microphone sensor (MIC) output signal within the speech interval into low-frequency and high-frequency components.
[0024] The noise reduction processing MCU is characterized by including: a vocalization section extraction unit that extracts a vocalization section according to vocal cord vibration using the output signal of an accelerometer sensor (ACC); an ACC low-frequency component processing unit that processes the low-frequency component signal of the accelerometer sensor (ACC); a MIC noise extraction and removal unit that determines and extracts / removes the voice microphone sensor (MIC) output signal other than the vocalization section as noise using vocalization section information; a noise level determination unit that determines the level of noise extracted from the output signal of the voice microphone sensor (MIC); a MIC low-frequency component processing unit and a MIC high-frequency component processing unit that separate and process the voice microphone sensor (MIC) output signal of the vocalization section into low-frequency components and high-frequency components; a MIC and ACC low-frequency component synthesis unit that synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the synthesis ratio based on the noise level determined by the noise level determination unit; and a voice signal restoration output unit that restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC).
[0025] A high-quality voice signal processing device based on multi-sensor signal fusion for achieving other purposes according to the present invention, comprising: first and second voice microphone sensors (MIC1)(MIC2) that sense and output a speaker's voice signal while spaced apart from each other; an accelerometer sensor (ACC) that detects vibration of the speaker's vocal cords and outputs a signal; a noise reduction processing MCU that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC), synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1)(MIC2) by varying the synthesis ratio based on the noise level extracted from the output signals of the first and second voice microphone sensors (MIC1)(MIC2) using the vocalization segment information, and restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the first and second voice microphone sensors (MIC1)(MIC2); and a wireless communication module that outputs the restored voice signal externally.
[0026] Here, the noise reduction processing MCU is characterized by including more of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) during the synthesis process of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) if the noise level extracted from the output signal of the first and second voice microphone sensors (MIC1) (MIC2) is higher than a reference value, and including more of the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) if it is lower.
[0027] The noise reduction processing MCU utilizes speech interval information to determine, extract, and remove the first and second voice microphone sensor (MIC1)(MIC2) output signals other than the speech interval as noise, and separates the first and second voice microphone sensor (MIC1)(MIC2) output signals of the speech interval into low-frequency and high-frequency components.
[0028] The method is characterized by first removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using voice segment information, second removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the first noise has been removed using a beamforming algorithm, and third removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the second noise has been removed again using voice segment information.
[0029] And the noise reduction processing MCU comprises: a vocalization segment extraction unit that extracts vocalization segments based on vocal cord vibration using the output signal of an accelerometer sensor (ACC); an ACC low-frequency component processing unit that processes the low-frequency component signal of the accelerometer sensor (ACC); a noise extraction and removal unit that performs first-order noise extraction and removal using vocalization segment information from the output signals of the first and second voice microphone sensors (MIC1)(MIC2), performs second-order noise extraction and removal using a beamforming algorithm, and performs third-order noise extraction and removal using vocalization segment information again; a noise level determination unit that determines the level of the first-order noise extracted by the noise extraction and removal unit; a MIC low-frequency component processing unit and a MIC high-frequency component processing unit that separate and process the output signals of the first and second voice microphone sensors (MIC1)(MIC2), on which third-order noise removal has been performed, into low-frequency and high-frequency components; and a unit that varies the synthesis ratio of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1)(MIC2) based on the noise level determined by the noise level determination unit. It is characterized by including a MIC and ACC low-frequency component synthesis unit that synthesizes, and a voice signal restoration output unit that restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the first and second voice microphone sensors (MIC1) (MIC2).
[0030] The noise extraction and removal unit is characterized by including a first noise extraction and removal unit that extracts and removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) by using voice interval information, a second noise extraction and removal unit that removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the first noise has been removed by using a beamforming algorithm, and a third noise extraction and removal unit that extracts and removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the second noise has been removed by using voice interval information again.
[0031] A method for processing a high-quality voice signal through ambient noise removal based on multi-sensor signal fusion according to the present invention for achieving another objective comprises: a step of extracting a vocalization segment based on vocal cord vibration using an output signal of an accelerometer sensor (ACC); a step of determining, extracting, and removing a voice microphone sensor (MIC) output signal other than the vocalization segment as noise using vocalization segment information, and separating the voice microphone sensor (MIC) output signal of the vocalization segment into a low-frequency component and a high-frequency component; a step of determining the level of noise extracted from the output signal of the voice microphone sensor (MIC); a step of synthesizing the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the sum ratio based on the determined noise level; and a step of restoring and outputting a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC).
[0032] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC), if the level of noise extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, the low-frequency component of the accelerometer sensor (ACC) is included more, and if it is lower, the low-frequency component of the voice microphone sensor (MIC) (31) is included more.
[0033] A high-quality voice signal processing method based on multi-sensor signal fusion and ambient noise removal according to the present invention for achieving another objective comprises: a step of extracting a vocalization segment based on vocal cord vibration using the output signal of an accelerometer sensor (ACC); a step of determining and extracting the output signals of the first and second voice microphone sensors (MIC1)(MIC2) other than the vocalization segment as noise and performing first-order noise removal using the vocalization segment information; a step of performing second-order noise removal using a beamforming algorithm on the output signals of the first and second voice microphone sensors (MIC1)(MIC2) after first-order noise removal; a step of determining the signals outside the vocalization segment as noise in the second-order noise-removed signal and performing third-order noise removal, and separating the output signals of the first and second voice microphone sensors (MIC1)(MIC2) in the vocalization segment into low-frequency and high-frequency components; a step of determining the level of noise extracted from the output signals of the first and second voice microphone sensors (MIC1)(MIC2) using the vocalization segment information; and based on the determined noise level The method is characterized by including: a step of synthesizing the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) by varying the synthesis ratio; and a step of restoring and outputting a voice signal by adding the synthesized low-frequency component and the high-frequency component of the first and second voice microphone sensors (MIC1) (MIC2).
[0034] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2), if the level of noise extracted from the output signal of the first and second voice microphone sensors (MIC1) (MIC2) is higher than a reference value, the low-frequency component of the accelerometer sensor (ACC) is further included, and if it is lower, the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) is further included. Effects of the invention
[0035] The high-quality voice signal processing device and method based on multi-sensor signal fusion and ambient noise removal according to the present invention, as described above, has the following effects.
[0036] First, multi-sensor signal fusion using an accelerometer (ACC) and a voice microphone (MIC) enables voice signal processing that is robust against external noise environments.
[0037] Second, by using the speech interval information of the accelerometer sensor (ACC), noise is extracted and removed from the output signal of the voice microphone sensor (MIC), thereby enabling efficient voice signal processing.
[0038] Third, the noise level extracted from the output signal of the voice microphone sensor (MIC) is determined, and based on the determined noise level, the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) are combined at different synthesis ratios to improve the quality of the voice signal.
[0039] Fourth, noise is first removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using speech interval information obtained from the output signal of the accelerometer sensor (ACC), noise is secondarily removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using a beamforming algorithm, and noise is thirdly removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) again using speech interval information to obtain a high-quality voice signal.
[0040] Fifth, in the synthesis process of the low-frequency components of the accelerometer sensor (ACC) and the voice microphone sensor (MIC), if the noise level extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, the low-frequency components of the accelerometer sensor (ACC) are included more, and if it is lower, the low-frequency components of the voice microphone sensor (MIC) are included more, thereby enabling precise voice signal processing. Brief explanation of the drawing
[0041] Figure 1 is a frequency characteristic graph of a voice microphone and a vocal cord microphone. Figure 2 is a configuration diagram illustrating the principle of noise removal through active noise canceling. FIG. 3 is a configuration diagram of a voice signal processing device according to a first embodiment of the present invention. FIG. 4 is a detailed configuration diagram of a noise reduction processing MCU according to a first embodiment of the present invention. FIG. 5 is a flowchart illustrating a high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion according to a first embodiment of the present invention. FIG. 6 is a configuration diagram of a voice signal processing device according to a second embodiment of the present invention. FIG. 7 is a detailed configuration diagram of a noise reduction processing MCU according to a second embodiment of the present invention. FIG. 8 is a flowchart illustrating a high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion according to a second embodiment of the present invention. Specific details for implementing the invention
[0042] Hereinafter, a preferred embodiment of the high-quality voice signal processing apparatus and method based on multi-sensor signal fusion and ambient noise removal according to the present invention will be described in detail as follows.
[0043] The features and advantages of the high-quality voice signal processing apparatus and method based on multi-sensor signal fusion and ambient noise removal according to the present invention will become apparent through the detailed description of each embodiment below.
[0044] FIG. 3 is a configuration diagram of a voice signal processing device according to a first embodiment of the present invention.
[0045] The terms used in this disclosure have been selected to be as widely used and general as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.
[0046] When a part of a specification is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "...part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.
[0047] The high-quality voice signal processing device and method based on multi-sensor signal fusion and ambient noise removal according to the present invention enables voice signal processing robust to external noise environments through multi-sensor signal fusion using an accelerometer sensor (ACC) and a voice microphone sensor (MIC).
[0048] To this end, the present invention may include a configuration for extracting and removing noise from a voice microphone sensor (MIC) output signal using speech interval information of an accelerometer sensor (ACC) in order to enable efficient voice signal processing.
[0049] The present invention may include a configuration for determining the level of noise extracted from the output signal of a voice microphone sensor (MIC) to improve the quality of a voice signal, and synthesizing the low-frequency component of an accelerometer sensor (ACC) and the low-frequency component of a voice microphone sensor (MIC) by varying the synthesis ratio based on the determined noise level.
[0050] The present invention may include a configuration that enables obtaining a high-quality voice signal by first removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using speech interval information obtained from the output signal of an accelerometer sensor (ACC), second removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using a beamforming algorithm, and third removing noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) again using speech interval information.
[0051] The present invention may include a configuration that enables precise voice signal processing by including more of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) when the noise level extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value during the synthesis process of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC), and including more of the low-frequency component of the voice microphone sensor (MIC) when it is lower.
[0052] The present invention enables fast voice signal processing by allowing signal processing processes such as noise extraction, noise level determination, low-frequency component synthesis, and voice signal restoration to be performed without a separate digital conversion process, as the accelerometer sensor (ACC) outputs a digital signal.
[0053] The configuration of a high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion according to the first embodiment of the present invention is as follows.
[0054] As shown in FIG. 3, the system includes a voice microphone sensor (MIC) (31) that senses and outputs a voice signal of a speaker, an accelerometer sensor (ACC) (32) that detects vibration of the vocal cords while in contact with the speaker's vocal cords and outputs a signal, a noise reduction processing MCU (33) that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC) (32), synthesizes the low-frequency component of the accelerometer sensor (ACC) (32) and the low-frequency component of the voice microphone sensor (MIC) (31) by varying the synthesis ratio based on the noise level extracted from the output signal of the voice microphone sensor (MIC) (31) using the vocalization segment information, and restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC) (31), and a wireless communication module (34) that outputs the restored voice signal externally.
[0055] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) (32) and the low-frequency component of the voice microphone sensor (MIC) (31), the noise reduction processing MCU (33) preferably includes more of the low-frequency component of the accelerometer sensor (ACC) (32) if the level of noise extracted from the output signal of the voice microphone sensor (MIC) (31) is higher than a reference value, and includes more of the low-frequency component of the voice microphone sensor (MIC) (31) if it is lower.
[0056] Then, the noise reduction processing MCU (33) uses the voice microphone sensor (MIC) (31) output signal outside the voice microphone sensor (MIC) (31) output signal to determine the voice microphone sensor (MIC) (31) output signal outside the voice microphone sensor (MIC) (31) output signal as noise, extracts and removes it, and separates the voice microphone sensor (MIC) (31) output signal in the voice microphone sensor (MIC) (31) output signal into low-frequency components and high-frequency components.
[0057] The detailed configuration of the noise reduction processing MCU (33) is as follows.
[0058] FIG. 4 is a detailed configuration diagram of a noise reduction processing MCU according to the first embodiment of the present invention.
[0059] As shown in FIG. 4, the noise reduction processing MCU (33) comprises: a vocalization section extraction unit (42) that extracts a vocalization section according to vocal cord vibration using the output signal of an accelerometer sensor (ACC) (32); an ACC low-frequency component processing unit (43) that processes the low-frequency component signal of the accelerometer sensor (ACC) (32); a MIC noise extraction and removal unit (41) that determines and extracts / removes the output signal of the voice microphone sensor (MIC) (31) other than the vocalization section as noise using vocalization section information; a noise level determination unit (44) that determines the level of noise extracted from the output signal of the voice microphone sensor (MIC) (31); a MIC low-frequency component processing unit (45) and a MIC high-frequency component processing unit (46) that separate and process the output signal of the voice microphone sensor (MIC) (31) of the vocalization section into low-frequency and high-frequency components; and an accelerometer based on the noise level determined by the noise level determination unit (44). It includes a MIC and ACC low-frequency component synthesis unit (47) that synthesizes the low-frequency component of the sensor (ACC) (32) and the low-frequency component of the voice microphone sensor (MIC) (31) by varying the synthesis ratio, and a voice signal restoration output unit (48) that restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC) (31).
[0060] A method for processing high-quality voice signals through ambient noise removal based on multi-sensor signal fusion according to the first embodiment of the present invention is described in detail as follows.
[0061] FIG. 5 is a flowchart illustrating a high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion according to a first embodiment of the present invention.
[0062] A method for processing a high-quality voice signal through ambient noise removal based on multi-sensor signal fusion according to the first embodiment of the present invention first extracts a vocalization segment based on vocal cord vibration using the output signal of an accelerometer sensor (ACC), as shown in FIG. 5. (S501)
[0063] Next, using the vocalization section information, the vocal microphone sensor (MIC) output signal outside the vocalization section is identified as noise, extracted, and removed, and the vocal microphone sensor (MIC) output signal within the vocalization section is separated into low-frequency and high-frequency components. (S502)
[0064] Then, the level of noise extracted from the output signal of the voice microphone sensor (MIC) is determined. (S503)
[0065] Next, based on the determined noise level, the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) are combined by varying the sum ratio. (S504)
[0066] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC), if the level of noise extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, it is desirable to include more of the low-frequency component of the accelerometer sensor (ACC), and if it is lower, to include more of the low-frequency component of the voice microphone sensor (MIC) (31).
[0067] Then, the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC) are added to restore and output a voice signal. (S505)
[0068] The configuration of a high-quality voice signal processing device based on multi-sensor signal fusion and ambient noise removal according to the second embodiment of the present invention is as follows.
[0069] FIG. 6 is a configuration diagram of a voice signal processing device according to a second embodiment of the present invention.
[0070] As shown in FIG. 6, first and second voice microphone sensors (MIC1)(MIC2)(61)(62) that sense and output a speaker's voice signal while spaced apart from each other, an accelerometer sensor (ACC)(63) that detects the vibration of the vocal cords while in contact with the speaker's vocal cords and outputs a signal, a noise reduction processing MCU (64) that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC)(63), synthesizes the low-frequency component of the accelerometer sensor (ACC)(63) and the low-frequency component of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) by varying the synthesis ratio based on the noise level extracted from the output signals of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) using the vocalization segment information, and adds the high-frequency component of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) to restore and output a voice signal, and It includes a wireless communication module (65) that outputs the restored voice signal externally.
[0071] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) (63) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62), the noise reduction processing MCU (64) preferably includes more of the low-frequency component of the accelerometer sensor (ACC) (63) if the level of noise extracted from the output signal of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62) is higher than a reference value, and includes more of the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62) if it is lower.
[0072] Then, the noise reduction processing MCU (64) uses the voice interval information to determine and extract the output signal of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) outside the voice interval as noise, and removes it, and separates the output signal of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) in the voice interval into a low-frequency component and a high-frequency component.
[0073] It is also preferable to first remove noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) using voice interval information obtained from the output signal of the accelerometer sensor (ACC)(63), secondly remove noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) from which the first noise has been removed using a beamforming algorithm, and thirdly remove noise from the output signals of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) from which the second noise has been removed again using voice interval information.
[0074] The detailed configuration of the noise reduction processing MCU (64) is as follows.
[0075] FIG. 7 is a detailed configuration diagram of a noise reduction processing MCU according to a second embodiment of the present invention.
[0076] As shown in FIG. 7, the noise reduction processing MCU (64) comprises a vocalization segment extraction unit (72) that extracts a vocalization segment according to vocal cord vibration using the output signal of an accelerometer sensor (ACC) (63), an ACC low-frequency component processing unit (73) that processes the low-frequency component signal of the accelerometer sensor (ACC) (63), a first noise extraction removal unit (71) that extracts and removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62) using vocalization segment information, a second noise extraction removal unit (74) that removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62) from which the first noise has been removed using a beamforming algorithm, and a third noise extraction unit that removes noise from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) (61) (62) from which the second noise has been removed again using vocalization segment information. It includes a removal unit (75), a noise level determination unit (78) that determines the level of noise extracted from the first noise extraction removal unit (71), a MIC low-frequency component processing unit (76) and a MIC high-frequency component processing unit (77) that separate and process the output signals of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) in which third-order noise removal has been performed into low-frequency and high-frequency components, a MIC and ACC low-frequency component synthesis unit (79) that synthesizes the low-frequency components of the accelerometer sensor (ACC)(63) and the low-frequency components of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62) by varying the synthesis ratio based on the noise level determined by the noise level determination unit (78), and a voice signal restoration output unit (80) that restores and outputs a voice signal by adding the synthesized low-frequency components and the high-frequency components of the first and second voice microphone sensors (MIC1)(MIC2)(61)(62).
[0077] A method for processing high-quality voice signals through ambient noise removal based on multi-sensor signal fusion according to the second embodiment of the present invention is described in detail as follows.
[0078] FIG. 8 is a flowchart illustrating a high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion according to a second embodiment of the present invention.
[0079] First, the output signal of the accelerometer sensor (ACC) is used to extract the vocalization segment based on vocal cord vibration. (S801)
[0080] Next, using the vocalization section information, the output signals of the first and second voice microphone sensors (MIC1) (MIC2) other than the vocalization section are determined to be noise, extracted, and first noise removal is performed. (S802)
[0081] Then, secondary noise removal is performed using a beamforming algorithm on the output signals of the first and second voice microphone sensors (MIC1)(MIC2) after primary noise removal has been performed. (S803)
[0082] Next, signals outside the vocalization section in the second noise-removed signal are identified as noise, and a third noise removal is performed, and the output signals of the first and second voice microphone sensors (MIC1)(MIC2) in the vocalization section are separated into low-frequency and high-frequency components. (S804)
[0083] Then, using the vocalization interval information, the level of noise extracted from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) is determined. (S805)
[0084] Next, based on the determined noise level, the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) are combined at different combination ratios. (S806)
[0085] Here, in the process of combining the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2), if the level of noise extracted from the output signal of the first and second voice microphone sensors (MIC1) (MIC2) is higher than a reference value, it is desirable to include more of the low-frequency component of the accelerometer sensor (ACC), and if it is lower, it is desirable to include more of the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2).
[0086] Then, the synthesized low-frequency component and the high-frequency components of the first and second voice microphone sensors (MIC1) (MIC2) are added to restore and output a voice signal. (S807)
[0087] The high-quality voice signal processing apparatus and method based on multi-sensor signal fusion according to the present invention described above enables voice signal processing robust to external noise environments through multi-sensor signal fusion using an accelerometer sensor (ACC) and a voice microphone sensor (MIC). It extracts and removes noise from the output signal of the voice microphone sensor (MIC) using speech interval information from the accelerometer sensor (ACC), determines the level of noise extracted from the output signal of the voice microphone sensor (MIC), and synthesizes the low-frequency components of the accelerometer sensor (ACC) and the low-frequency components of the voice microphone sensor (MIC) by varying the synthesis ratio based on the determined noise level to improve the quality of the voice signal.
[0088] As explained above, it will be understood that the present invention is implemented in a modified form without departing from the essential characteristics of the invention.
[0089] Therefore, the described embodiments should be considered in an illustrative rather than a limiting sense, and the scope of the invention is defined by the claims rather than the foregoing description, and all variations within the equivalent scope should be interpreted as being included in the invention. Explanation of the symbols
[0090] 31. Voice Microphone Sensor (MIC) 32. Accelerometer Sensor (ACC) 33. Noise Reduction Processing MCU 34. Wireless communication module
Claims
Claim 1 A voice microphone sensor (MIC) that senses and outputs a speaker's voice signal; an accelerometer sensor (ACC) that detects the speaker's vocal cord vibration and outputs a signal; a noise reduction processing MCU that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC), uses the vocalization segment information to determine and extract / remove the voice microphone sensor (MIC) output signal outside the vocalization segment as noise, synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the synthesis ratio based on the level of noise extracted from the output signal of the voice microphone sensor (MIC), and restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC); and a wireless communication module that outputs the restored voice signal externally; wherein, during the synthesis process of the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC), if the level of noise extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, the accelerometer sensor (ACC) A high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion, characterized by including more low-frequency components, and if low, including more low-frequency components of a voice microphone sensor (MIC). Claim 2 delete Claim 3 A high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion, wherein the noise reduction processing MCU uses speech interval information to determine, extract, and remove voice microphone sensor (MIC) output signals other than the speech interval as noise, and separates the voice microphone sensor (MIC) output signals within the speech interval into low-frequency and high-frequency components. Claim 4 In claim 1, the noise reduction processing MCU comprises: a vocalization segment extraction unit that extracts a vocalization segment according to vocal cord vibration using an output signal of an accelerometer sensor (ACC); an ACC low-frequency component processing unit that processes a low-frequency component signal of the accelerometer sensor (ACC); a MIC noise extraction and removal unit that determines, extracts, and removes a voice microphone sensor (MIC) output signal other than the vocalization segment as noise using vocalization segment information; a noise level determination unit that determines the level of noise extracted from the output signal of the voice microphone sensor (MIC); a MIC low-frequency component processing unit and a MIC high-frequency component processing unit that separate and process the voice microphone sensor (MIC) output signal of the vocalization segment into low-frequency components and high-frequency components; a MIC and ACC low-frequency component synthesis unit that synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the synthesis ratio based on the noise level determined by the noise level determination unit; and a voice signal restoration output unit that restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC). A high-quality voice signal processing device characterized by ambient noise removal based on multi-sensor signal fusion. Claim 5 A first and second voice microphone sensors (MIC1)(MIC2) spaced apart from each other to sense and output a speaker's voice signal; an accelerometer sensor (ACC) that detects vibration of the speaker's vocal cords and outputs a signal; a noise reduction processing MCU that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC), uses the vocalization segment information to determine the output signals of the first and second voice microphone sensors (MIC1)(MIC2) other than the vocalization segment as noise, extracts and removes them, synthesizes the low-frequency components of the accelerometer sensor (ACC) and the low-frequency components of the first and second voice microphone sensors (MIC1)(MIC2) by varying the synthesis ratio based on the noise level extracted from the output signals of the first and second voice microphone sensors (MIC1)(MIC2), and restores and outputs a voice signal by adding the synthesized low-frequency components and the high-frequency components of the first and second voice microphone sensors (MIC1)(MIC2); and a wireless communication module that outputs the restored voice signal externally; comprising noise reduction processing A high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion, characterized in that, in the process of synthesizing the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2), if the noise level extracted from the output signal of the first and second voice microphone sensors (MIC1) (MIC2) is higher than a reference value, the low-frequency component of the accelerometer sensor (ACC) is included further, and if it is lower, the low-frequency component of the first and second voice microphone sensors (MIC1) (MIC2) is included further. Claim 6 delete Claim 7 In claim 5, the noise reduction processing MCU utilizes speech interval information to determine, extract, and remove the first and second voice microphone sensor (MIC1)(MIC2) output signals other than the speech interval as noise, and separates the first and second voice microphone sensor (MIC1)(MIC2) output signals of the speech interval into low-frequency and high-frequency components, thereby forming a high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion. Claim 8 A high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion, characterized in that, in claim 7, noise is first removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using speech segment information, noise is secondarily removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the first noise has been removed using a beamforming algorithm, and noise is thirdly removed from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the second noise has been removed again using speech segment information. Claim 9 In claim 5, the noise reduction processing MCU comprises: a vocalization segment extraction unit that extracts a vocalization segment based on vocal cord vibration using the output signal of an accelerometer sensor (ACC); an ACC low-frequency component processing unit that processes the low-frequency component signal of the accelerometer sensor (ACC); a noise extraction and removal unit that performs first-order noise extraction and removal using vocalization segment information from the output signals of the first and second voice microphone sensors (MIC1)(MIC2), performs second-order noise extraction and removal using a beamforming algorithm, and performs third-order noise extraction and removal using vocalization segment information again; a noise level determination unit that determines the level of the first-order noise extracted by the noise extraction and removal unit; a MIC low-frequency component processing unit and a MIC high-frequency component processing unit that separate and process the output signals of the first and second voice microphone sensors (MIC1)(MIC2), on which third-order noise removal has been performed, into low-frequency and high-frequency components; and based on the noise level determined by the noise level determination unit, the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the first and second voice microphone sensors (MIC1)(MIC2) A high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion, characterized by including a MIC and ACC low-frequency component synthesis unit that synthesizes components with different synthesis ratios, and a voice signal restoration output unit that restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the first and second voice microphone sensors (MIC1) (MIC2). Claim 10 In claim 9, the noise extraction and removal unit comprises: a first noise extraction and removal unit that extracts and removes, as noise, from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) using voice interval information; a second noise extraction and removal unit that removes, as noise, from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the first noise has been removed using a beamforming algorithm; and a third noise extraction and removal unit that extracts and removes, as noise, from the output signals of the first and second voice microphone sensors (MIC1) (MIC2) from which the second noise has been removed using voice interval information again, thereby providing a high-quality voice signal processing device through ambient noise removal based on multi-sensor signal fusion. Claim 11 The method comprises: a step of extracting a vocalization segment based on vocal cord vibration using the output signal of an accelerometer sensor (ACC) in a vocalization segment extraction unit of a noise reduction processing MCU; a step of determining, extracting, and removing the voice microphone sensor (MIC) output signal outside the vocalization segment as noise using vocalization segment information in a MIC noise extraction and removal unit, and separating the voice microphone sensor (MIC) output signal within the vocalization segment into low-frequency and high-frequency components in a MIC low-frequency component processing unit and a MIC high-frequency component processing unit; a step of determining the level of noise extracted from the output signal of the voice microphone sensor (MIC) in a noise level determination unit; a step of synthesizing the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the summing ratio based on the determined noise level in a MIC and ACC low-frequency component synthesis unit; and a step of restoring and outputting a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC) in a voice signal restoration output unit; wherein the low-frequency component of the accelerometer sensor (ACC) and the voice microphone sensor (MIC) A high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion, characterized in that, during the synthesis process of low-frequency components, if the level of noise extracted from the output signal of the voice microphone sensor (MIC) is higher than a reference value, the low-frequency component of the accelerometer sensor (ACC) is included more, and if it is lower, the low-frequency component of the voice microphone sensor (MIC) is included more. Claim 12 delete Claim 13 A step of extracting a vocalization segment based on vocal cord vibration using the output signal of an accelerometer sensor (ACC) in the vocalization segment extraction unit of a noise reduction processing MCU; a step of determining the output signals of the first and second voice microphone sensors (MIC1)(MIC2) other than the vocalization segment as noise using the vocalization segment information in the noise extraction and removal unit, and extracting and performing first-order noise removal; a step of performing second-order noise removal using a beamforming algorithm on the output signals of the first and second voice microphone sensors (MIC1)(MIC2) from which first-order noise removal has been performed in the noise extraction and removal unit; a step of determining the signals outside the vocalization segment as noise from the signal from which second-order noise removal has been performed in the noise extraction and removal unit, and performing third-order noise removal, and separating the output signals of the first and second voice microphone sensors (MIC1)(MIC2) of the vocalization segment into low-frequency and high-frequency components in the MIC low-frequency component processing unit and the MIC high-frequency component processing unit; and a step of using the vocalization segment information in the noise level determination unit to A step of determining the level of noise extracted from the output signal of the 1st and 2nd voice microphone sensors (MIC1) (MIC2); a step of synthesizing the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the 1st and 2nd voice microphone sensors (MIC1) (MIC2) by varying the synthesis ratio based on the noise level determined by the MIC and ACC low-frequency component synthesis unit; a step of restoring and outputting a voice signal by adding the synthesized low-frequency component and the high-frequency component of the 1st and 2nd voice microphone sensors (MIC1) (MIC2) in the voice signal restoration output unit;A high-quality voice signal processing method through ambient noise removal based on multi-sensor signal fusion, comprising: a low-frequency component of an accelerometer sensor (ACC) and a low-frequency component of a first and second voice microphone sensor (MIC1)(MIC2); wherein, in the synthesis process of the low-frequency component of an accelerometer sensor (ACC) and the low-frequency component of a first and second voice microphone sensor (MIC1)(MIC2), if the noise level extracted from the output signal of the first and second voice microphone sensor (MIC1)(MIC2) is higher than a reference value, the low-frequency component of the accelerometer sensor (ACC) is further included, and if it is lower, the low-frequency component of the first and second voice microphone sensor (MIC1)(MIC2) is further included. Claim 14 delete