Noise removal device, noise removal method, and program

The noise removal device employs HPSS processing and a trained model to separate and identify noise in audio signals, effectively removing noise without degrading audio quality.

JP7758585B2Active Publication Date: 2025-10-22PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022010177
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-10-22
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Conventional noise removal methods fail to effectively eliminate low-frequency noise, leading to incomplete noise reduction and degradation of audio quality when attempting to remove low-frequency components.

Method used

A noise removal device and method utilizing HPSS processing to separate audio signals into harmonic and percussion sound signals, followed by a determination process using a trained model to distinguish between noise and percussion sounds, allowing selective output of either the harmonic signal or the processed audio signal based on the determination.

Benefits of technology

Accurately removes noise from audio signals while preserving audio quality by distinguishing between noise and percussion sounds, thereby enhancing noise removal accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758585000001
    Figure 0007758585000001
  • Figure 0007758585000002
    Figure 0007758585000002
  • Figure 0007758585000003
    Figure 0007758585000003
Patent Text Reader

Abstract

To highly accurately remove noise in a voice signal while deterioration of voice quality is suppressed.SOLUTION: A noise removal device includes: an acquisition part for acquiring a voice signal; a noise removal part for performing noise removal processing on the voice signal acquired by the acquisition part; an HPSS processing part for separating the voice signal acquired by the acquisition part into a harmonic signal and a percussion sound signal by HPSS processing; a determination part for determining whether the percussion sound signal separated by the HPSS processing part is noise or not; and an output control part for outputting the harmonic signal separated in the HPSS processing part when the percussion sound signal is determined to be noise by the determination part, while outputting the voice signal processed by the noise removal part when the determination part determines that the percussion sound signal is not noise.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a noise removal device, a noise removal method, and a program. [Background technology]

[0002] Conventionally, a noise canceling function for an audio signal has been used to improve audio quality. For example, in a conventional noise canceling function applicable to a radio receiving device, when removing pulse-type noise, the width and amplitude of high-frequency noise components are detected by filtering, and then the components are cut. For example, Patent Document 1 discloses a configuration including a noise detection unit that detects noise components and a noise cut unit that cuts high-frequency components of noise, with the aim of removing noise without creating a sense of discomfort. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 2002-64389 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional methods can only remove high-frequency noise, leaving low-frequency noise behind, resulting in insufficient noise removal. However, if you try to remove low-frequency components by lowering the noise detection range, the original sound will also be removed, resulting in a deterioration in audio quality.

[0005] The present disclosure has been devised in consideration of the above-described conventional circumstances, and aims to provide a noise removal device, a noise removal method, and a program that are capable of removing noise from an audio signal with high accuracy while suppressing degradation of audio quality. [Means for solving the problem]

[0006] The present disclosure provides a noise removal device having an acquisition unit that acquires an audio signal, a noise removal unit that performs noise removal processing on the audio signal acquired by the acquisition unit, an HPSS processing unit that separates the audio signal acquired by the acquisition unit into a harmonic signal and a percussion sound signal by HPSS processing, a determination unit that determines whether the percussion sound signal separated by the HPSS processing unit is noise, and an output control unit that, if the determination unit determines that the percussion sound signal is noise, outputs the harmonic signal separated by the HPSS processing unit, and, if the determination unit determines that the percussion sound signal is not noise, outputs the audio signal processed by the noise removal unit.

[0007] The present disclosure also provides a noise removal method having an acquisition step of acquiring an audio signal, a noise removal step of performing noise removal processing on the audio signal acquired in the acquisition step, an HPSS processing step of separating the audio signal acquired in the acquisition step into a harmonic signal and a percussion sound signal by HPSS processing, a determination step of determining whether the percussion sound signal separated in the HPSS processing step is noise or not, and an output control step of outputting the harmonic signal separated in the HPSS processing step if it is determined in the determination step that the percussion sound signal is noise, and outputting the audio signal processed in the noise removal step if it is determined in the determination step that the percussion sound signal is not noise.

[0008] The present disclosure also provides a program for causing a computer to function as an acquisition unit that acquires an audio signal, a noise removal unit that performs noise removal processing on the audio signal acquired by the acquisition unit, an HPSS processing unit that separates the audio signal acquired by the acquisition unit into a harmonic signal and a percussion sound signal by HPSS processing, a determination unit that determines whether the percussion sound signal separated by the HPSS processing unit is noise, and an output control unit that outputs the harmonic signal separated by the HPSS processing unit if the determination unit determines that the percussion sound signal is noise, and outputs the audio signal processed by the noise removal unit if the determination unit determines that the percussion sound signal is not noise.

[0009] Any combination of the above components, and conversion of the present disclosure into a method, device, system, storage medium, computer program, etc., are also valid aspects of the present disclosure. [Effects of the Invention]

[0010] According to the present disclosure, it is possible to remove noise from an audio signal with high accuracy while suppressing degradation of audio quality. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a system having a noise canceling function according to a first embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an example of a functional configuration of a radio wave state detection unit according to a first embodiment. [Figure 3] FIG. 1 is a block diagram showing an example of the functional configuration of a noise processing unit and a noise detection unit according to a first embodiment. [Figure 4] Diagram for explaining an example of noise [Figure 5] Diagram to explain speech signal separation using HPSS [Figure 6] Flowchart of audio processing according to the first embodiment [Figure 7] FIG. 1 is a schematic diagram illustrating a learning process for generating a trained model for a noise determination process according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, with appropriate reference to the accompanying drawings, embodiments specifically disclosing a noise removal device, a noise removal method, and a program according to the present disclosure will be described in detail. However, more detailed description than necessary may be omitted. For example, detailed description of well-known matters or redundant description of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure and are not intended to limit the subject matter recited in the claims.

[0013] <First Embodiment> The noise reduction device according to the present invention can be applied, for example, as a configuration included in a radio receiving device. In this embodiment, the noise reduction device will be described using a configuration applied to a radio receiving device. However, the scope of application of the present invention is not limited to radio receiving devices, and the present invention can also be applied to devices that output audio based on other audio signals.

[0014] [System Configuration] 1 is a block diagram showing an example of the configuration of a radio receiving device 100 including components related to a noise canceling function according to this embodiment. In Fig. 1, arrows indicate the flow of signals to each functional block.

[0015] Antenna 101 detects radio signals transmitted from an external source and passes them to receiving unit 102. There are no particular limitations on the number, orientation, shape, detectable frequency band, wavelength, or other configuration of antenna 101. Antenna 101 may be configured to detect radio signals from a specific direction, or may be configured to detect radio signals from all directions. The radio signals may be FM (Frequency Modulation), AM (Amplitude Modulation), or wide FM, but this embodiment will be described taking FM as an example.

[0016] The receiving unit 102 receives the radio signal detected by the antenna 101 and outputs a composite signal composed of signals of a predetermined frequency. The composite signal is output to the noise processing unit 103, the noise detection unit 104, and the radio wave state detection unit 105.

[0017] The noise processing unit 103 performs noise processing on the input composite signal. An example of noise processing is removal of high frequency components by known filtering processing. The signal processed by the noise processing unit 103 is output to the demodulation unit 106. A more detailed configuration example of the noise processing unit 103 will be described later with reference to FIG. 3.

[0018] The noise detection unit 104 detects whether or not noise is included in the composite signal output from the receiving unit 102. The noise detection unit 104 outputs a signal indicating the noise detection result to the noise processing unit 103 and the microcomputer unit 111. A more detailed configuration example of the noise detection unit 104 will be described later with reference to FIG. 3.

[0019] The radio wave condition detection unit 105 detects the radio wave condition based on the signal output from the receiving unit 102. The radio wave condition includes, for example, the field strength and multipath level of the radio signal. The radio wave condition detection unit 105 outputs a signal indicating the detection result of the radio wave condition to the microcomputer unit 111. A more detailed configuration example of the radio wave condition detection unit 105 will be described later with reference to FIG. 2.

[0020] The demodulation unit 106 demodulates the signal received from the noise processing unit 103 and outputs the demodulated signal as an audio signal to the delay correction unit 107 and the HPSS processing unit 108. The delay correction unit 107 performs correction by delaying the audio signal received from the demodulation unit 106 by a certain amount of time. The amount of correction here may be specified in advance or may be adjusted according to various processes described below. The amount of correction may also be referred to as the delay amount. The delay correction unit 107 outputs the delayed audio signal to the switch unit 109.

[0021] The HPSS processing unit 108 separates the audio signal received from the demodulation unit 106 into a harmonic signal and a percussive sound signal using HPSS (Harmonic / Percussive Sound Separation). Details of the HPSS processing, the harmonic signal, and the percussive sound signal will be described later. The harmonic signal is output to the switch unit 109, and the percussive sound signal is output to the noise determination AI unit 110.

[0022] Switch unit 109 is configured to receive as input the audio signal from delay correction unit 107 and the harmonic signal from HPSS processing unit 108, and to switch between either one and output it to audio amplifier 112 based on a signal from microcomputer unit 111. For convenience, the input of switch unit 109 to which the audio signal is input from delay correction unit 107 is also referred to as the first input terminal, and the input to which the harmonic signal is input from HPSS processing unit 108 is also referred to as the second input terminal.

[0023] The noise determination AI unit 110 uses a trained model generated in advance by a learning process to determine whether or not noise is contained in the radio signal received by the receiving unit 102. Details of the trained model used in this embodiment will be described later. The noise determination AI unit 110 receives the percussion sound signal output from the HPSS processing unit 108, and outputs a signal indicating the noise determination result to the microcomputer unit 111.

[0024] The microcomputer unit 111 controls the switching of the switch unit 109 based on various signals acquired from the noise detection unit 104, the radio wave condition detection unit 105, and the noise determination AI unit 110. The switching control of the switch unit 109 will be described later. The microcomputer unit 111 may be configured using at least one of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), or an FPGA (Field Programmable Gate Array), for example.

[0025] The audio amplifier 112 receives the audio signal from the switch unit 109 and amplifies the audio signal to a predetermined state. Then, the audio amplifier 112 outputs the amplified audio signal to the speaker 113. The conditions for amplifying the audio signal by the audio amplifier 112 may be set in advance or may be operated by the user as needed. The speaker 113 outputs the audio signal received from the audio amplifier 112 as sound.

[0026] Although not shown in FIG. 1, the radio receiving device 100 may further include a power supply unit, a switching unit for switching receivable frequencies, an operation unit for accepting operations by the user, and the like.

[0027] (Radio wave condition detector) Fig. 2 is a diagram showing an example of the functional configuration of the radio wave state detection unit 105 according to this embodiment. The radio wave state detection unit 105 detects the radio wave state based on the received signal from the receiving unit 102. In this embodiment, FM is assumed, and a configuration example for detecting the occurrence of multipath will be described. Note that the configuration shown in Fig. 2 is just an example, and other configurations for detecting radio wave states may also be used.

[0028] Field strength meter 201 measures the field strength of the signal received from receiving unit 102, i.e., the radio wave strength. Field strength meter 201 is, for example, an S meter. Conventionally, if multipath (in other words, multiple wave propagation) occurs in the received signal for some reason, an AC (alternating current) component occurs in the output of field strength meter 201. Therefore, radio wave state detection unit 105 used in this embodiment detects the AC component from the measurement result of field strength meter 201, and performs multipath detection based on the detection result.

[0029] A BPF (Band Pass Filter) 202 is applied to the detection result of the field strength meter 201. The BPF 202 passes a predetermined frequency band and removes other high-frequency and low-frequency components. For example, a frequency band around 20 kHz is set as the predetermined frequency band that the BPF 202 passes. The pass band of the BPF 202 is not particularly limited, but may be determined in advance. This extracts the AC component of the received signal. The output from the BPF 202 is then input to a peak detection unit 203.

[0030] The peak detection unit 203 detects the peak level of the signal output from the BPF 202. The peak detection unit 203 outputs information related to the detected peak level to the multipath detection unit 204. The multipath detection unit 204 detects multipaths based on the peak level detected by the peak detection unit 203. In this embodiment, the multipath detection unit 204 detects that multipaths have occurred when the peak level is equal to or greater than a predetermined threshold. Multipath detection may also be performed using multiple thresholds. For example, when detecting the peak level in dBμV units, two thresholds are set at 2 dBμV and 10 dBμV. Then, when the peak level is equal to or less than 2 dBμV, it may be determined that multipaths are small or not occurring. Furthermore, when the peak level is equal to or greater than 10 dBμV, it may be determined that multipaths that have a significant impact on audio quality have occurred. The predetermined threshold is not particularly limited, but may be defined in advance. Furthermore, when detecting multipaths, the detection is not limited to an actual value (a detection value expressed in the unit dBμV in the above), but a ratio (%) may also be used. The detection result by the multipath detection unit 204 is output to the microcomputer unit 111 as a detection result signal of the radio wave state detection unit 105.

[0031] (Noise processing section and noise detection section) 3 is a diagram showing an example of the functional configuration of noise processing unit 103 and noise detection unit 104 according to this embodiment. The composite signal from receiving unit 102 is input to noise processing unit 103 and noise detection unit 104, respectively.

[0032] In the noise processing unit 103, the input composite signal is input to a delay correction unit 301 and a noise removal unit 302. The delay correction unit 301 delays the input composite signal by a predetermined delay amount and outputs the signal to the switch unit 305. The delay amount by the delay correction unit 301 may correspond to the time required for processing in the noise removal unit 302. In other words, the timing of the output from the delay correction unit 301 and the output from the noise removal unit 302 to the switch unit 305 may be configured to match or approximately match.

[0033] In the noise removal unit 302, first, an LPF (Low Pass Filter) 303 is applied to the input composite signal. The LPF 303 passes a frequency band lower than a predetermined threshold and removes high-frequency components above that threshold. In other words, the LPF 303 removes high-frequency components corresponding to noise. The pass band of the LPF 303 is not particularly limited, but may be determined in advance. Furthermore, in the noise removal unit 302, the output from the LPF 303 is input to a signal maintaining unit 304. The signal maintaining unit 304 maintains the signal to which the LPF 303 has been applied for a certain period of time. The signal maintaining unit 304 then outputs the maintained signal to the switch unit 305.

[0034] Switch unit 305 is configured to receive the signal from delay correction unit 301 and the signal from noise removal unit 302 as input, and to switch between either one of them based on the signal from noise detection unit 104 and output it as an output signal. Therefore, depending on the noise detection result by noise detection unit 104, noise processing unit 103 outputs either a composite signal delayed by a certain time or a composite signal from which noise has been removed.

[0035] In the noise detection unit 104, first, an HPF (High Pass Filter) 311 is applied to the input composite signal. The HPF 311 passes frequencies higher than a predetermined threshold and removes low-frequency components below that threshold. In other words, it extracts high-frequency components corresponding to noise. The passband of the HPF 311 is not particularly limited and may be predetermined. The predetermined threshold defining the passband of the HPF 311 may be the same as that of the LPF 303 of the noise removal unit 302. Furthermore, in the noise detection unit 104, the output from the HPF 303 is input to an AGC (Auto Gain Control) unit 312. The AGC unit 312 adjusts the input signal so that it falls within a predetermined gain range. The AGC unit 312 outputs the adjusted signal to a pulse noise detection unit 313. The pulse noise detection unit 313 detects pulse noise based on the signal from the AGC unit 312. If pulse noise is detected in the signal, pulse noise detection unit 313 instructs switch unit 305 of noise processing unit 103 to switch so as to output the signal from noise removal unit 302. On the other hand, if pulse noise is not detected in the signal, pulse noise detection unit 313 instructs switch unit 305 of noise processing unit 103 to switch so as to output the signal from delay correction unit 301.

[0036] In this embodiment, the noise processing performed by the noise processing unit 103 and the noise detection unit 104 is also referred to as first noise processing. Note that the first noise processing method is not limited to the noise canceling function configured as shown in Fig. 3, and other known configurations may be used.

[0037] [noise] The radio signal handled in this embodiment may contain percussion sounds in addition to noise. Percussion sounds are composed of certain high-frequency components and have a structure similar to multipath noise, so they may be treated as noise when a noise canceling function is applied. For example, when a radio signal is displayed as an image signal, the percussion sound signal and pulse noise will appear similar in the image. Therefore, when a conventional noise canceling function is applied to percussion sounds, they may be removed in the same way as multipath noise, resulting in a decrease in the quality of the radio signal. One of the objectives of this embodiment is to properly recognize such percussion sounds and prevent erroneous processing by the noise canceling function.

[0038] First, the noise handled in this embodiment will be described. Fig. 4 shows an example of a signal containing noise, with the vertical axis representing frequency and the horizontal axis representing time. As shown in Fig. 4(b), the signal shown in Fig. 4(a) contains two noise components 404 and 405. Here, the explanation will be given assuming that a frequency threshold 403 is 5 kHz. Furthermore, frequencies greater than the threshold 403 are defined as a high-frequency component region 401, and frequencies equal to or less than the threshold 403 are defined as a low-frequency component region 402.

[0039] 3, that is, the first noise processing, removes noise by removing high frequency components in area 401. In other words, signals below threshold 403 are nearly identical to the original signal in order to maintain the audio quality of the original sound.

[0040] Furthermore, this embodiment uses HPSS technology. HPSS is a technology that enables separation of an audio signal into a harmonic signal and a percussive sound signal. For information on HPSS, see, for example, "Harmonic / percussive separation using median filtering," Fitzgerald, Derry., 13th International Conference on Digital Audio Effects (DAFX10), Graz, Austria, 2010.

[0041] HPSS processing will be explained using Figure 5. In Figure 5, the vertical axis represents frequency and the horizontal axis represents time. As shown in Figure 5(a), an audio signal contains two percussion sound components 501 and 502. By applying HPSS processing to the audio signal shown in Figure 5(a), it is separated into a harmonic signal and a percussion sound signal.

[0042] Figure 5(b) shows the harmonic signals separated by HPSS processing. Of the percussion sound component 501, high-frequency components 511 above a predetermined threshold are removed. Furthermore, signals equivalent to the percussion sound component 501 are also removed from the portion of the original sound below the predetermined threshold. Similarly, of the percussion sound component 502, high-frequency components 513 above a predetermined threshold are removed. Furthermore, signals equivalent to the percussion sound component 503 are also removed from the portion of the original sound below the predetermined threshold.

[0043] FIG. 5(c) shows a percussion sound signal separated by HPSS processing. Signal 521 corresponding to percussion sound component 501 is emphasized. Similarly, signal 522 corresponding to percussion sound component 502 is emphasized. As described above, percussion sounds and pulse noise have similar characteristics, and the percussion sound components 501 and 502 shown in FIG. 5(a) can also be described as pulse noise. In other words, the harmonic signal obtained by applying HPSS processing can be treated as a signal from which noise has been removed. In this embodiment, for convenience, noise removal by HPSS processing (in other words, separation of harmonic signals) is also referred to as second noise processing.

[0044] As mentioned above, the first noise processing method has difficulty in removing noise below a predetermined threshold, even when the signal contains noise. On the other hand, the harmonic components obtained by the second noise processing method, such as those shown in Fig. 5(b), also remove the original sound, so if the signal corresponds to a percussion sound rather than noise, the audio quality may be degraded.

[0045] Therefore, in this embodiment, a trained model obtained by machine learning is used to determine whether a signal such as that shown in Fig. 5(c) is a percussion sound or pulse noise, and the signal output is controlled based on the determination result.

[0046] [Learning process] The learning process for generating a trained model used in the noise determination AI unit 110 according to this embodiment will be described. Fig. 7 is a conceptual diagram showing the flow for generating a trained model according to this embodiment. Here, the process is roughly divided into a preprocessing phase for preparing training data and a learning process phase for generating a trained model using the training data.

[0047] In the description of the present embodiment, "learning" or "machine learning" refers to generating a "trained model" by performing learning using training data and an arbitrary learning algorithm. A trained model is updated as needed as learning progresses using multiple pieces of training data, and its output changes even when the input is the same. Therefore, the state of a trained model is not limited to a specific point in time. Here, a model used in learning is referred to as a "learning model," and a learning model that has undergone a certain level of learning is referred to as a "trained model." Specific examples of "training data" will be described later, but the configuration may vary depending on the learning algorithm used. Training data may include training data used for the training itself, verification data used to verify the trained model, and test data used to test the trained model. In the following description, the term "training data" is used to collectively refer to data related to training, and the term "training data" is used to refer to data used when performing the training itself. It is not intended to clearly classify training data, verification data, and test data contained in training data. For example, depending on the training, verification, and testing methods, all training data may also be training data.

[0048] The processing of each phase is realized by a processing unit of an information processing device (not shown) reading and executing various programs stored in a storage unit. An example of the information processing device is a PC (Personal Computer). The processing unit may be configured with a CPU, a GPU (Graphical Processing Unit), etc. The storage unit may be configured with a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), etc.

[0049] In the pre-processing phase, training data to be used in training is prepared. Here, an audio signal obtained as a signal received by the radio receiving device 100 is used as the original data. First, in step S701, a predetermined number of peaks are detected from the original data. Then, the original data is cut out for each predetermined time unit according to the peak positions. The number of peaks detected and the predetermined time unit for cutting out are not particularly limited. Next, in step S702, a percussion sound signal is extracted from each data by applying HPSS processing to each data cut out in step S701.

[0050] Next, in step S703, features are extracted from the percussion sound signal extracted in step S702. Feature extraction may be implemented, for example, by applying a known Mel filter bank to perform level extraction. Feature data of a predetermined number of dimensions corresponding to each data is then generated through feature extraction. Then, in step S704, the feature data generated in step S703 is labeled. For example, labeling is performed by assigning label information (classification information) indicating whether the audio signal, which is the original data, is noise or percussion sound. Here, two classifications, "noise" and "percussion sound," are used, but more detailed classifications may be made, for example, according to the level of noise. Training data consisting of pairs of feature data and label information is prepared through preprocessing.

[0051] Next, in the learning phase, the learning process and the verification operation are repeatedly performed using the above-described learning data, thereby generating a trained model with a certain level of accuracy. Before the radio receiving device 100 is configured, a trained model to be used in the noise determination AI unit 110 according to this embodiment is generated. Note that the trained model used in the noise determination AI unit 110 does not limit the trained model to be used at any point in time. Therefore, the learning process may be performed as appropriate, and the trained model held by the noise determination AI unit 110 may be updated with the trained model updated thereby.

[0052] In this embodiment, a configuration will be described in which inputs are classified using a third-order SVM (Support Vector Machine) method, which is one of machine learning learning methods. Note that the machine learning algorithm is not particularly limited, and classification may be performed using a known method such as a convolutional neural network (CNN) that uses a deep learning method using a neural network.

[0053] Through the above learning, a trained model is generated that receives feature data of an audio signal as input and outputs the classification of the feature data. Note that the noise determination AI unit 110 according to this embodiment performs preprocessing (e.g., conversion into feature data) on the percussion sound signal input from the HPSS processing unit 108 in accordance with the input format of the trained model. Therefore, if the trained model is configured to be able to process the percussion sound signal as is as an input, preprocessing of the percussion sound signal is not required.

[0054] [Processing flow] FIG. 6 is a flowchart showing a series of steps in the operation of the radio receiving device 100 according to this embodiment. The processing of each step is realized by cooperation between the components shown in FIG. 1. Furthermore, it is assumed that a learning process is executed before this processing flow is executed, and the trained model generated thereby is available for use by the noise determination AI unit 110. Note that the operation of the radio receiving device 100 is not necessarily limited to being executed serially as shown in the steps of the flowchart in FIG. 6, and some steps may be executed simultaneously in parallel in accordance with the functional configuration of FIG. 1.

[0055] The microcomputer unit 111 instructs the switch unit 109 to switch so as to output an audio signal from the side of the delay correction unit 107. As a result, the switch unit 109 switches to the first input terminal side (step S601).

[0056] The radio receiving device 100 starts receiving radio signals from the surrounding area through the antenna 101 (step S602).

[0057] The receiving unit 102 converts the radio signal received by the antenna 101 into a composite signal (step S603). Then, the receiving unit 102 outputs the converted composite signal to the noise processing unit 103, the noise detection unit 104, and the radio wave state detection unit 105.

[0058] The radio wave condition detection unit 105 detects the radio wave condition based on the composite signal input from the receiving unit 102 (step S604). In this embodiment, as described with reference to FIG. 2, multipath is detected in the composite signal, and if multipath is detected, it is determined that the radio wave condition has deteriorated. For example, it may be determined that the radio wave condition has deteriorated if the multipath level exceeds a predetermined threshold (e.g., 40%) for a certain period of time (e.g., 100 ms). If it is determined that the radio wave condition has deteriorated (step S604; YES), the processing of the radio receiving device 100 proceeds to step S609. On the other hand, if it is determined that the radio wave condition has not deteriorated (step S604; NO), the processing of the radio receiving device 100 proceeds to step S605.

[0059] The noise detection unit 104 detects the occurrence of noise based on the composite signal input from the receiving unit 102 (step S605). In this embodiment, as described with reference to FIG. 3, if pulse noise is included in the composite signal, it is determined that noise has been detected. For example, if the noise level exceeds a predetermined threshold (e.g., 0.03), it may be determined that noise has occurred. If it is determined that noise has been detected (step S605; YES), the processing of the radio receiving device 100 proceeds to step S609. On the other hand, if it is determined that noise has not been detected (step S605; NO), the processing of the radio receiving device 100 proceeds to step S606.

[0060] If no noise is detected in step S605, the noise detection unit 104 instructs the switch unit 305 of the noise processing unit 103 to switch so that output is performed from the delay correction unit 301 side. As a result, the noise processing unit 103 outputs a composite signal that has been delayed by a certain amount without undergoing noise removal processing to the demodulation unit 106. The demodulation unit 106 then demodulates the signal input from the noise processing unit 103 and outputs it as an audio signal (step S606). The demodulation unit 106 outputs the demodulated audio signal to the delay correction unit 107.

[0061] The delay correction unit 107 delays the audio signal input from the demodulation unit 106 by a certain amount of time (step S607). The amount of delay may be constant or may vary depending on predetermined conditions. The delay correction unit 107 then outputs the delayed audio signal to the switch unit 109.

[0062] The switch unit 109 outputs the audio signal input from the delay correction unit 107 to the audio amplifier 112. In this case, since the input of the switch unit 109 is the first input terminal side, the audio signal from the delay correction unit 107 side is output. More specifically, the audio signal is a signal that has not been subjected to noise removal (first noise processing) in the noise processing unit 103. The audio signal is then output as sound via the audio amplifier 112 and the speaker 113 (step S608). Then, the process proceeds to step S619.

[0063] The HPSS processor 108 starts the HPSS processing (step S609). Here, the HPSS processing will be described as taking a certain amount of time to execute.

[0064] If noise is detected in step S605, the noise detection unit 104 instructs the switch unit 305 of the noise processing unit 103 to switch to output from the noise removal unit 302. Then, the noise removal unit 302 controls the noise removal unit 302 to perform noise removal processing (first noise processing) on ​​the composite signal input from the receiving unit 102 for a predetermined time (step S610). Here, the predetermined time may be set based on the time required for HPSS processing or determination using a trained model. Then, the noise processing unit 103 outputs the composite signal that has been subjected to noise removal processing to the demodulation unit 106.

[0065] The demodulation unit 106 demodulates the signal input from the noise processing unit 103 and outputs it as an audio signal (step S611). At this time, the demodulation unit 106 outputs the demodulated audio signal to the delay correction unit 107 and also to the HPSS processing unit 108.

[0066] The delay correction unit 107 delays the audio signal input from the demodulation unit 106 by a certain amount of time (step S612). The amount of delay may be constant or may vary depending on predetermined conditions. For example, the amount of delay may be the amount of time required for the signal to be output from the HPSS processing unit 108. The delay correction unit 107 then outputs the delayed audio signal to the switch unit 109.

[0067] The HPSS processing unit 108 applies HPSS processing to the audio signal input from the demodulation unit 106, separating the audio signal into a harmonic signal and a percussion sound signal (step S613). Then, the HPSS processing unit 108 outputs the harmonic signal to the switch unit 109 and outputs the percussion sound signal to the noise determination AI unit 110.

[0068] The noise determination AI unit 110 receives the percussion sound signal from the HPSS processing unit 108 as input, applies the trained model, and performs a determination by classifying the input percussion sound signal (step S614). As described above, if preprocessing is required to input the percussion sound signal to the trained model, the noise determination AI unit 110 performs the preprocessing before inputting the percussion sound signal to the trained model. The noise determination AI unit 110 then outputs the determination result to the microcomputer unit 111.

[0069] The microcomputer unit 111 determines whether the inputs from the noise detection unit 104, the radio wave condition detection unit 105, and the noise determination AI unit 110 satisfy predetermined conditions (step S615). The predetermined conditions may be, for example, when the noise detection unit 104 detects that the noise level exceeds a predetermined threshold, or when the radio wave condition detection unit 105 detects that the multipath level exceeds a predetermined threshold for a certain period of time, and when the noise determination AI unit 110 classifies the percussion sound signal as not being "noise." In this embodiment, when the classification result is "noise," it means that the separated signal is "noise," as shown in FIG. 5(c). On the other hand, when the classification result is not "noise," it means that the separated signal is "percussion sound," as shown in FIG. 5(c). If the predetermined conditions are satisfied (step S615; YES), the processing of the radio receiving device 100 proceeds to step S616. On the other hand, if the predetermined condition is not satisfied (step S615; NO), the processing of the radio receiving device 100 proceeds to step S618.

[0070] The microcomputer unit 111 instructs the switch unit 109 to switch so as to output an audio signal from the HPSS processing unit 108. As a result, the switch unit 109 switches to the second input terminal side (step S616).

[0071] The switch unit 109 outputs the audio signal input from the HPSS processing unit 108 to the audio amplifier 112. In this case, since the input of the switch unit 109 is the second input terminal side, the audio signal from the HPSS processing unit 108 side is output. More specifically, the harmonic signal separated by HPSS in the HPSS processing unit 108 is output as a signal after noise removal processing (second noise processing). The audio signal is then output as sound via the audio amplifier 112 and speaker 113 (step S617). Then, the process proceeds to step S619.

[0072] The switch unit 109 outputs the audio signal input from the delay correction unit 107 to the audio amplifier 112. In this case, since the input of the switch unit 109 is the first input terminal side, the audio signal from the delay correction unit 107 side is output. More specifically, the audio signal that has been subjected to noise removal (first noise processing) in the noise processing unit 103 is output. The audio signal is then output as sound via the audio amplifier 112 and the speaker 113 (step S618). Then, the process proceeds to step S619.

[0073] The microcomputer unit 111 determines whether the audio output has ended (step S619). This determination may be made based on whether an instruction to end the audio output has been received through a user operation. If the audio output has ended (step S619; YES), this processing flow ends. On the other hand, if the audio output has not ended (step S619; NO), the processing of the radio receiving device 100 returns to step S601 and the processing is repeated. Note that the HPSS processing by the HPSS processing unit 108 and the determination processing using a trained model by the noise determination AI unit 110 are expected to impose a high processing load, so it is preferable not to operate them when no noise is occurring. However, if the processing load is low or the device is capable of executing higher-speed processing, they may be configured to operate constantly.

[0074] As described above, according to this embodiment, radio receiving device 100 includes antenna 101 and receiving unit 102 that receive radio signals, noise processing unit 103 that performs noise reduction processing on the radio signals, HPSS processing unit 108 that separates the radio signals into harmonic signals and percussion sound signals by HPSS processing, noise determination AI unit 110 that determines whether the percussion sound signals separated by HPSS processing unit 108 are noise, and microcomputer unit 111 that outputs the harmonic signals separated by HPSS processing unit 108 if the noise determination AI unit 110 determines that the percussion sound signals are noise, and outputs the audio signal processed by noise processing unit 103 if the noise determination AI unit 110 determines that the percussion sound signals are not noise. This makes it possible to highly accurately remove noise from audio signals while suppressing degradation of audio quality.

[0075] The noise determination AI unit 110 receives the percussion sound signals separated by the HPSS processing as input and performs a learning process using the classification of the percussion sound signals as output, and uses a trained model generated by the learning process to make the determination. This makes it possible to use the trained model that has undergone a certain level of learning to determine with high accuracy whether a component included in an audio signal is noise or a percussion sound.

[0076] The radio receiving device 100 further includes a radio wave condition detection unit 105 that detects the radio wave condition of the audio signal acquired by the antenna 101 and the receiving unit 102, and the microcomputer unit 111 controls output switching based on the detection result by the radio wave condition detection unit 105. This makes it possible to control output switching based on the radio wave condition in addition to the determination result by the noise determination AI unit 110.

[0077] Furthermore, the radio wave condition detection unit 105 detects the radio wave condition based on multipath noise contained in the audio signal acquired by the antenna 101 and the receiving unit 102. This makes it possible to control output switching based on the determination result of the noise determination AI unit 110 as well as the detection result of multipath noise.

[0078] The radio receiving device 100 further includes a noise detection unit 104 that detects noise contained in an audio signal based on high-frequency components of the audio signal acquired by the antenna 101 and the receiving unit 102, and the microcomputer unit 111 controls output switching based on the detection result by the noise detection unit 104. This makes it possible to control output switching based on the determination result of the noise determination AI unit 110 as well as the result of a conventional noise detection method.

[0079] Furthermore, when noise is detected by the noise detection unit 104, the noise processing unit 103 performs noise removal processing on the audio signal acquired by the antenna 101 and the receiving unit 102 and outputs the result. This makes it possible to control the output of a signal equivalent to the original sound without unnecessary noise processing when the noise detection unit 104 does not detect noise.

[0080] <Other embodiments> In the above embodiment, a radio receiving device having a noise canceling function using a noise reduction device according to the present invention has been described as an example. However, the present invention is not limited to this, and the characteristic configuration of the present invention can be applied to any device that acquires and outputs an audio signal. For example, an audio playback device capable of playing back storage media such as CDs and BDs (registered trademarks) may be equipped with the characteristic configuration of the present invention to remove noise from audio data during playback. It is expected that dust, dirt, etc. will physically adhere to the surface of such storage media, causing noise.

[0081] Furthermore, the present invention is not limited to application when outputting audio, but may be used for noise removal when converting an audio signal into another signal and then converting it back into an audio signal, for example.

[0082] Furthermore, in the above embodiment, an example of a configuration including noise processing unit 103 and noise detection unit 104 that perform the first noise processing, and radio wave condition detection unit 105 that detects the radio wave condition has been shown, but the present invention is not limited to this. For example, these components may be omitted, and a configuration including HPSS processing unit 108 and noise determination AI unit 110 may be used, in which if noise is determined to be noise by noise determination AI unit 110, the separated harmonic signal is output by HPSS processing unit 108, and if not, the original audio signal is output.

[0083] It is also possible to omit the noise determination AI unit 110. For example, it is also possible to configure the system so that when multipath noise is detected by the radio wave condition detection unit 105, the harmonic signals separated by the HPSS processing unit 108 are output, and when multipath noise is not detected, the original audio signal is output.

[0084] Alternatively, a configuration may be adopted in which a trained model for noise determination is applied as a function of the noise detection unit 104. In this case, a configuration may be adopted in which switching is performed to perform any one of the first noise processing, HSPP processing, and bandwidth control depending on the signal classification result by the noise detection unit 104.

[0085] In addition, the programs and applications for realizing the functions of one or more of the above-described embodiments can be supplied to a system or device using a network or storage medium, and one or more processors in the computer of the system or device can read and execute the programs.

[0086] Alternatively, it may be realized by a circuit that realizes one or more functions (for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array)).

[0087] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to these examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents may be made within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention. [Industrial Applicability]

[0088] The present disclosure is useful as a noise removal device, a noise removal method, and a program capable of removing noise contained in an audio signal. It is useful as. [Explanation of symbols]

[0089] 100...Radio receiving device 101...Antenna 102...Receiver 103...Noise processing section 104...Noise detection unit 105...Radio wave condition detection unit 106...Demodulation section 107...Delay correction unit 108...HPSS processing section 109...Switch section 110...Noise detection AI section 111...Microcomputer section 112...Audio amplifier 113...Speaker

Claims

1. an acquisition unit that acquires an audio signal; a noise removal unit that performs noise removal processing on the audio signal acquired by the acquisition unit; an HPSS processing unit that separates the audio signal acquired by the acquisition unit into a harmonic signal and a percussion sound signal by HPSS processing; a determination unit that determines whether the percussion sound signal separated by the HPSS processing unit is noise or a non-noise percussion sound; an output control unit that outputs a harmonic signal separated by the HPSS processing unit when the determination unit determines that the percussion sound signal is noise, and that outputs an audio signal processed by the noise removal unit when the determination unit determines that the percussion sound signal is a percussion sound that is not noise; A noise removal device having:

2. the determination unit performs determination using a trained model generated by performing a learning process using a percussion sound signal separated by HPSS processing as an input and a classification of the percussion sound signal as an output. The noise removal device according to claim 1 .

3. a radio wave condition detection unit that detects a radio wave condition of the audio signal acquired by the acquisition unit, the output control unit controls the output further based on the detection result by the radio wave state detection unit. The noise removal device according to claim 1 or 2.

4. The noise removal device according to claim 3 , wherein the radio wave condition detection unit detects the radio wave condition based on multipath noise contained in the audio signal acquired by the acquisition unit.

5. a noise detection unit that detects noise included in the audio signal based on high-frequency components of the audio signal acquired by the acquisition unit, 5. The noise removal device according to claim 1, wherein the output control unit controls the output further based on a detection result by the noise detection unit.

6. The noise removal device according to claim 5 , wherein, when the noise detection unit detects noise, the noise removal unit performs the noise removal process on the audio signal acquired by the acquisition unit and outputs the processed audio signal.

7. the noise removal device is provided in a radio receiving device, the audio signal is a radio signal; The noise removal device according to any one of claims 1 to 6.

8. an acquisition step of acquiring an audio signal; a noise removal process for performing noise removal processing on the audio signal acquired in the acquisition process; an HPSS processing step of separating the audio signal acquired in the acquisition step into a harmonic signal and a percussion sound signal by HPSS processing; a determination step of determining whether the percussion sound signal separated in the HPSS processing step is noise or a non-noise percussion sound; an output control step of outputting a harmonic signal separated in the HPSS processing step when the determination step determines that the percussion sound signal is noise, and outputting a sound signal processed in the noise removal step when the determination step determines that the percussion sound signal is a percussion sound that is not noise; A noise removal method comprising:

9. Computer, an acquisition unit that acquires an audio signal; a noise removal unit that performs noise removal processing on the audio signal acquired by the acquisition unit; an HPSS processing unit that separates the audio signal acquired by the acquisition unit into a harmonic signal and a percussion sound signal by HPSS processing; a determination unit that determines whether the percussion sound signal separated by the HPSS processing unit is noise or a non-noise percussion sound; an output control unit that outputs a harmonic signal separated by the HPSS processing unit when the determination unit determines that the percussion sound signal is noise, and that outputs a sound signal processed by the noise removal unit when the determination unit determines that the percussion sound signal is a percussion sound that is not noise; A program to function as a

Citation Information

Patent Citations

  • Noise removing device

    JP2002064389A

  • Acoustic analysis device

    JP2017067901A

  • Receiving device and receiving method

    JP2020096268A

  • Acoustic diagnosis method, acoustic diagnosis system and acoustic diagnosis program

    JP2021124887A