Signal processing method and signal processing device
The signal processing method addresses processing delays in cough detection systems by predicting and preparing for pronunciation actions like coughing, ensuring timely and accurate responses.
Patent Information
- Application Number
- JP2023217376
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing cough detection systems, such as those described in Patent Document 1, suffer from processing delays due to the execution of operations based on the presence or absence of a cough, leading to delayed responses.
A signal processing method that detects pre-articulation operations from sensor information, predicts the sound signal associated with these operations, and determines the processing content before receiving the actual sound signal, thereby reducing processing delays.
The method effectively suppresses processing delays by determining the content of signal processing for pronunciation actions like coughing before the actual sound is received, ensuring timely and accurate responses.
Smart Images

Figure 2025100184000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a signal processing method and a signal processing apparatus that determine and execute the processing content of a sound signal based on information acquired by a sensor.
Background Art
[0002] Patent Document 1 discloses a method for determining the presence or absence of a cough based on the acoustic feature amount of acoustic data received by a microphone array and image data obtained by photographing the scene where the sound is generated. The invention of Patent Document 1 discloses a method for controlling the operation of other devices according to the presence or absence of a cough.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the cough detection apparatus disclosed in Patent Document 1, for example, a predetermined operation is executed on an air cleaner depending on the presence or absence of a cough. Therefore, in the cough detection apparatus of Patent Document 1, the processing for a pronunciation action such as a cough is delayed.
[0005] An object of the present invention is to provide a signal processing method and a signal processing apparatus that determine the processing content for a user's pronunciation action such as a cough from information acquired by a sensor and suppress processing delay.
Means for Solving the Problems
[0006] The signal processing method according to an embodiment of the present invention receives sensor information from a sensor, detects a pre-articulation operation corresponding to an articulation operation from the sensor information, predicts a sound signal emitted by the articulation operation based on the detected pre-articulation operation, determines the content of processing for the predicted sound, receives a sound signal, and executes signal processing on the received sound signal based on the determined content of processing.
Advantages of the Invention
[0007] According to the signal processing method of the present invention, by detecting the pre-articulation operation of the user from the information acquired by the sensor and determining the content of processing of the sound signal based on the detected pre-articulation operation, it is possible to suppress the delay of processing, and thus there is an effect.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0009] Hereinafter, the signal processing device according to the embodiment of the present invention will be described with reference to the drawings. The same reference numerals are given to the same parts in each figure. In the second and subsequent embodiments, the description of matters common to the first embodiment will be omitted, and only the differences will be described. In particular, the same operations and effects due to the same configurations will not be sequentially mentioned for each embodiment.
[0010] 《First Embodiment》 FIG. 1 is a block diagram showing a basic configuration of a signal processing device 1 according to the first embodiment. The signal processing device 1 includes a sensor 11, a CPU 12, a DSP 13, a memory 14, a RAM 15, a user interface (I / F) 16, a microphone 17, a speaker 18, and a communication unit 19.
[0011] In the present embodiment, the sensor 11 is a camera that acquires an image of the user.
[0012] Note that the sensor 11 is not limited to a camera. The sensor 11 may be, for example, a distance measuring device such as a TOF (Time-of-Flight) sensor, a motion sensor, or a LiDAR (Light Detection And Ranging). The distance measuring device measures, for example, the distance to a specific part of the user's body and monitors the user's movement. Further, the sensor 11 may be a sensor provided in a biological information monitoring device such as a smartwatch or a body health meter. The biological information monitoring device monitors, for example, the user's breathing cycle.
[0013] The CPU 12 functions as a control unit that comprehensively controls the operation of the signal processing device 1 by reading the program for operation from the memory 14 into the RAM 15. Note that the program does not necessarily need to be stored in the memory 14 of the device itself. The CPU 12 may, for example, download it from a server or the like each time and read it into the RAM 15.
[0014] The DSP 13 is a signal processing unit that processes the image signal and the audio signal according to the control of the CPU 12.
[0015] The user I / F 16 accepts operations from the user. The user I / F 16 accepts operations such as, for example, adjusting the sensitivity of the microphone 17 or adjusting the volume of the speaker 18.
[0016] The microphone 17 acquires the user's voice. The speaker 18 outputs the voice received from a device on the remote side connected to the user via the Internet or the like.
[0017] The communication unit 19 transmits the image signal and the audio signal of the own device (near-end side) after being processed by the DSP 13 to a device of another device (remote side) connected via the Internet or the like. Also, the communication unit 19 receives an audio signal from the device on the remote side. The communication unit 19 outputs the received audio signal to the speaker 18. Thereby, the user of the signal processing device 1 can make a remote call with a remote location.
[0018] FIG. 2 is a functional block diagram of the signal processing device 1 according to the first embodiment. The functional configuration shown in FIG. 2 is realized by the CPU 12 and the DSP 13.
[0019] FIG. 3 is a flowchart showing the operation of the signal processing method according to the first embodiment.
[0020] Functionally, the signal processing device 1 includes an audio acquisition unit 100, an operation detection unit 200, a signal processing determination unit 300, and a signal processing unit 400.
[0021] FIG. 2 shows a configuration in the case of performing signal processing based on the image signal and the audio signal on the proximal side, and FIG. 3 shows its operation.
[0022] The motion detection unit 200 acquires the image signal received from the sensor 11 as a proximal user image (S001). The proximal user image is an example of the sensor information of the present invention.
[0023] Next, the motion detection unit 200 detects a pre-articulation motion corresponding to the articulation motion based on the acquired proximal user image (S002). The articulation motion refers to a motion highly likely to be accompanied by the generation of a specific sound. The pre-articulation motion refers to a motion highly likely to be taken before the articulation motion. For example, the articulation motion is a sneeze. When the articulation motion is a sneeze, the pre-articulation motions detected by the motion detection unit 200 include motions such as covering the mouth with the hand, forward and backward movement of the head, and rotational movement of the neck. That is, the pre-articulation motion in the present embodiment means a motion performed when emitting a specific unintended sound such as a sneeze, different from a gesture which is a motion performed with intention.
[0024] Next, the signal processing determination unit 300 predicts an audio signal (hereinafter referred to as an articulation motion sound) generated along with the articulation motion based on the detected pre-articulation motion (S003), and determines the content of the processing for the predicted articulation motion sound (S004). For example, when the motion detection unit 200 detects a motion of covering the mouth with the hand as a pre-articulation motion, the signal processing determination unit 300 predicts a sneeze or a cough as the articulation motion, and predicts the articulation motion sound emitted by the sneeze or the cough. Then, the content of the processing for the predicted articulation motion sound is determined. The signal processing determination unit 300 determines, for example, a gain reduction process as the content of the processing for the articulation motion sound emitted by a sneeze or a cough.
[0025] Next, the audio acquisition unit 100 receives an audio signal from the microphone 17 (S005).
[0026] Next, based on the content of the process determined in step S004, the signal processing unit 400 executes a process on the sound signal received in step S005 (S006) and ends the flow (END).
[0027] In the present embodiment, when the operation detection unit 200 detects a specific operation set in advance, the signal processing unit 400 cancels the process executed in step S006. The specific operation set in advance is an operation that serves as an indicator indicating that the user has finished the pronunciation operation. The specific operation set in advance is, for example, the operation of lowering the hand that was covering the mouth.
[0028] As described above, the user of the signal processing apparatus 1 can obtain a new customer experience in which they can concentrate on the meeting without having to perform a mute operation, suppress sneezing, coughing, etc. even when performing a pronunciation operation. Also, the user on the far - end side can obtain a new customer experience in which they can concentrate on the meeting without hearing the pronunciation operation sound, which is an unnecessary sound for the meeting. Further, the signal processing apparatus 1 can suppress the delay of signal processing by determining the content of the signal processing for the sound signal based on the detected pre - pronunciation operation. In particular, since the signal processing apparatus 1 of the present embodiment can determine the content of the signal processing before actually receiving the pronunciation operation sound, it can completely eliminate the delay in the process for the pronunciation operation sound.
[0029] Note that the pre - pronunciation operation, pronunciation operation, content of the process for the sound signal, and specific operation set in advance in the present invention are not limited to the examples described above.
[0030] As the relationship between pre-articulation movements and articulation movements, specifically, the following combinations can be considered. When a movement of covering the mouth with a hand or a movement of moving the head is detected as a pre-articulation movement, coughing, sneezing, or yawning is predicted as the articulation movement. Also, when a movement of moving the head is detected as a pre-articulation movement, casual conversation is predicted as the articulation movement. Further, when a movement of swaying the body is detected as a pre-articulation movement, a swinging movement of the hands and feet is predicted as the articulation movement. Additionally, when a movement of aligning both hands in front of the body is detected as a pre-articulation movement, a typing operation on the keyboard is predicted as the articulation movement.
[0031] Also, as the content of the processing for the sound signal, in addition to the gain reduction processing of the sound signal described above, processing for changing the frequency characteristics (equalizer) of the sound signal, compression processing of the sound signal, band-pass filter processing, or noise cancellation processing, etc. can be considered. Here, the compression processing is a process of compressing the amplitude value of the part when the amplitude value of the sound signal exceeds a specified threshold value. Also, the band-pass filter processing is a process of restricting the frequency band outside the human voice region by a band-pass filter. Further, the noise cancellation processing is a process of canceling the articulation movement sound by adding a sound signal with an opposite phase to the target articulation movement sound.
[0032] Also, as a preset specific movement, in addition to the movement of lowering the hand that was covering the mouth described above, a movement of directing the line of sight towards the camera, a movement of moving the mouth such as speaking, or a movement of returning to the original posture, etc. can be considered. Furthermore, when the voice acquisition unit 100 acquires a voice that seems to be speech, this may also be treated as a preset specific movement.
[0033] Note that combinations of pronunciation operations and pre - pronunciation operations, as well as combinations of specific preset operations, may be pre - listed and stored in the memory 14. Alternatively, the signal processing determination unit 300 may determine the content of the process using a trained model. The trained model is a model trained by a DNN (Deep Neural Network) for the relationship between pre - pronunciation operations, pronunciation operations, pronunciation operation sounds, and the content of the process. In this case, the signal processing determination unit 300 may input the pre - pronunciation operation into the trained model and cause the trained model to output the corresponding content of the process. The signal processing determination unit 300 may input the pre - pronunciation operation into the trained model and output information indicating the pronunciation operation sound. The signal processing determination unit 300 may input sensor information (image signal) into the trained model and output information indicating the pre - pronunciation operation, or input the image signal into the trained model and output information indicating the pronunciation operation sound, or input the image signal into the trained model and output the content of the process.
[0034] 《Second Embodiment》 Next, the operation of the signal processing device 1A in the second embodiment will be described. FIG. 4 is a functional block diagram of the signal processing device 1A according to the second embodiment. The functional configuration shown in FIG. 4 is realized by the CPU 12 and the DSP 13. The configurations common to FIG. 2 are denoted by the same reference numerals and the description thereof is omitted.
[0035] The functional configuration of the signal processing device 1A is different from that of the signal processing device 1 in that it includes a determination unit 500. The determination unit 500 determines whether the pronunciation operation sound predicted by the signal processing determination unit 300 matches the sound signal of the processing candidate detected from the sound signal received by the microphone 17. The sound signal of the processing candidate refers to a specific sound signal received by the microphone 17 after the motion detection unit 200 detects the pre - pronunciation operation of the user.
[0036] FIG. 5 is a flowchart showing the operation of the signal processing method according to the second embodiment. The processes of steps S001 - S006 in FIG. 5 are the same as the processes of steps S001 - S006 in FIG. 5, so the description thereof is omitted.
[0037] After the signal processing unit 400 executes signal processing in step S006, the determination unit 500 detects a sound signal to be processed from the sound signal received by the sound acquisition unit 100 (S007). For example, when the level every predetermined time (for example, every 1 second) exceeds a threshold value, the determination unit 500 detects the sound signal of the predetermined time as the sound signal to be processed.
[0038] Next, the determination unit 500 determines whether or not the pronunciation operation sound predicted by the signal processing determination unit 300 matches the sound signal to be processed (S008). The determination unit 500 determines whether or not the detected sound signal to be processed matches the feature amount of the predicted pronunciation operation sound based on the feature amount of the sound signal to be processed. The feature amount is, for example, a level, a spectral envelope, a mel spectrum, or a mel-frequency cepstral coefficient. When the similarity between the feature amount of the sound signal to be processed and the feature amount of the pronunciation operation sound exceeds a threshold value, the determination unit 500 determines that they match.
[0039] When they match (S008: YES), the signal processing unit 400 continues the signal processing. When they do not match (S008: NO), the signal processing unit 400 cancels the signal processing (S009) and ends the flow (END).
[0040] In the present embodiment, when the sound acquisition unit 100 or the motion detection unit 200 detects a specific preset motion, the signal processing unit 400 cancels the processing that has been continued after step S008.
[0041] Note that, in addition to the specific preset motion described above, the signal processing unit 400 may cancel the processing when the sound acquisition unit 100 stops detecting the sound signal to be processed.
[0042] Thereby, even when the motion detection unit 200 erroneously detects a motion that is not a pre-phonetic motion as a pre-phonetic motion and the signal processing unit 400 executes processing on a sound signal that does not require processing, if the sound signal to be processed is different from the predicted pronunciation operation sound, the signal processing unit 400 can immediately cancel the processing.
[0043] As described above, when the signal processing determination unit 300 makes an incorrect prediction of the pronunciation operation sound, the signal processing device 1A immediately cancels the signal processing. As a result, the signal processing device 1A can suppress the influence when the signal processing is erroneously executed while reducing the delay of the signal processing.
[0044] Note that in the present embodiment, the intensity of the processing for the sound signal in step S006 can be changed.
[0045] For example, in steps S003 and S004, assume that the signal processing determination unit 300 predicts a sneezing sound and determines a gain reduction process. In this case, in step S006, the signal processing unit 400 executes the gain reduction process with an intensity of 50%. Then, when it is determined that the sound signal of the processing candidate is a sneezing sound (S008: YES), the signal processing unit 400 raises the intensity of the gain reduction process to 100%.
[0046] As described above, the signal processing device 1A can not only suppress the delay of the signal processing, but also suppress the influence of a processing error even when the signal processing unit 400 executes processing on a sound signal that does not require processing. Further, the signal processing unit 400 can change the intensity of the processing before and after step S008 that determines whether the sound signal of the processing candidate matches the predicted pronunciation operation sound. Therefore, by setting the intensity of the processing low at the stage before step S008, the influence of the processing error can be suppressed to a minimum.
[0047] <<Third Embodiment>> In the first and second embodiments, the signal processing unit 400 executes processing on the sound signal received by the microphone 17 immediately after the signal processing determination unit 300 determines the content of the processing. However, when it is desired to avoid the signal processing unit 400 from erroneously executing the processing, the following processing may be executed.
[0048] FIG. 6 is a flowchart showing the operation of the signal processing method according to the third embodiment. Since the processes of steps S001 - S005, S007, and S008 in FIG. 6 are the same as those of steps S001 - S005, S007, and S008 in FIG. 5, the description thereof is omitted.
[0049] The determination unit 500 detects a sound signal as a processing candidate (S007) after the voice acquisition unit 100 receives a sound signal in step S005. Then, the determination unit 500 determines whether the detected sound signal as a processing candidate matches the pronunciation operation sound predicted by the signal processing determination unit 300 (S008). If they match (S008: YES), the signal processing unit 400 executes signal processing on the sound signal as a processing candidate (S010). On the other hand, if they do not match (S008: NO), the signal processing unit 400 does not execute the process and ends the flow (END).
[0050] As described above, in the signal processing method of the present embodiment, the content of the process on the sound signal is determined at the stage where the pre - pronunciation operation is detected, but the process is not executed until it is determined that the sound signal as a processing candidate matches the predicted pronunciation operation sound. Thereby, the signal processing apparatus according to the present embodiment does not erroneously execute a process on a sound signal (voice) necessary for a meeting, and can reduce the processing delay by the time until the content of the process is determined.
[0051] <<Fourth Embodiment>> In the first to third embodiments, the signal processing method when receiving a sound signal from one microphone 17 has been described. In the fourth embodiment, the signal processing method when receiving sound signals from a plurality of microphones 17 will be described.
[0052] FIG. 7 is a functional block diagram of the signal processing apparatus 1B according to the fourth embodiment. The functional configuration shown in FIG. 7 is realized by the CPU 12 and the DSP 13. In the present embodiment, the DSP 13 also functions as a signal processing unit that performs beamforming. As described above, the signal processing apparatus 1B includes a plurality of microphones 17.
[0053] FIG. 8 is a flowchart showing the operation of the signal processing method according to the fourth embodiment. Since the processes of steps S001, S003 - S007, and S009 in FIG. 8 are the same as those of steps S001, S003 - S007, and S009 in FIG. 5, the description thereof will be omitted.
[0054] After acquiring the user image on the proximal side in step S001, the motion detection unit 200 detects a pre-articulation motion corresponding to the articulation motion based on the user image on the proximal side (S021). At this time, the motion detection unit 200 also detects the direction of the pre-articulation motion together with the pre-articulation motion. The direction of the pre-articulation motion refers to the direction in which the user who performed the pre-articulation motion exists.
[0055] After detecting the sound signal of the processing candidate in step S007, the sound acquisition unit 100 detects the direction (arrival direction) in which the sound signal of the processing candidate arrives based on the time difference of the sound signals of the processing candidate received by the plurality of microphones 17 (S011). The arrival direction of the sound signal of the processing candidate refers to the direction in which the user who emitted the sound signal of the processing candidate exists.
[0056] Next, the determination unit 500 determines whether or not the arrival direction of the sound signal of the processing candidate matches the detection direction of the pre-articulation motion detected by the motion detection unit 200 (S012). If they match (S012: YES), the signal processing unit 400 continues the signal processing. On the other hand, if they do not match (S012: NO), the signal processing unit 400 cancels the signal processing (S009) and ends the flow (END).
[0057] Also, in the present embodiment, the plurality of microphones 17 may function as microphones whose directivity can be changed by beamforming. In this case, the signal processing unit 400 of the signal processing apparatus 1B may have a function as a beamforming processing unit. The beamforming processing unit performs beamforming on the sound signals received by the plurality of microphones 17. Beamforming is a process of forming a sound collection beam having directivity in a predetermined direction by adding a delay to and synthesizing the sound signals acquired by the plurality of microphones 17. A plurality of sound collection beams can be formed simultaneously.
[0058] Instead of determining the content of the processing for the sound signal, the beamforming processing unit may form a sound collection beam for a user other than the user who performed the pre-articulation movement based on the pre-articulation movement detected by the motion detection unit 200.
[0059] Further, the beamforming processing unit may form a non-sound collection beam (so-called null) with lower sensitivity than other directions for the user who performed the pre-articulation movement based on the pre-articulation movement detected by the motion detection unit 200.
[0060] As described above, the signal processing device 1B determines whether or not the arrival direction of the sound signal as a processing candidate matches the detection direction of the pre-articulation movement. When the signal processing determination unit 300 makes an incorrect prediction of the articulation operation sound, the signal processing is immediately cancelled. Thereby, the signal processing device 1B can suppress the influence when the signal processing is erroneously executed while reducing the delay of the signal processing. Further, since the signal processing device 1B has a function as a beamforming processing unit, it picks up only the sound signal (voice) necessary for the meeting and does not pick up the articulation operation sound which is an unnecessary sound for the meeting. Therefore, the user of the signal processing device 1B can obtain a customer experience of being able to concentrate on the meeting without hearing the articulation operation sound.
[0061] In the present embodiment, the intensity of the processing for the sound signal in step S006 may be set by the user of the signal processing device 1B.
[0062] For example, assume that the user of the signal processing device 1B sets the intensity of the processing in step S006 to 50%. And in steps S003 and S004, assume that the signal processing determination unit 300 predicts background noise and determines noise cancellation processing. In this case, in step S006, the signal processing unit 400 adds a sound signal with an inverted phase of 50% amplitude to the received sound signal. And when the arrival direction of the sound signal as a processing candidate matches the detection direction of the pre-articulation movement (S012: YES), the signal processing unit 400 adds a sound signal with an inverted phase of 100% amplitude to the received sound signal.
[0063] Embodiment 5 In the first to fourth embodiments, an example of processing a sound signal acquired by a microphone of the own device (near-end side) based on sensor information of a near-end user has been described. In the fifth embodiment, an example of processing a sound signal acquired by a microphone of another device (far-end side) based on sensor information of a far-end user will be described.
[0064] FIG. 9 is a functional block diagram of a signal processing apparatus 1C according to the fifth embodiment. The functional configuration shown in FIG. 9 is realized by a CPU 12 and a DSP 13.
[0065] The signal processing apparatus 1C also includes a sensor 11C and a microphone 17C on the far-end side. In this embodiment, the sensor 11C is a camera.
[0066] FIG. 10 is a flowchart showing the operation of a signal processing method according to the fifth embodiment.
[0067] The motion detection unit 200 acquires, as a far-end user image, an image signal received from the sensor 11C via the communication unit 19 (S101). The far-end user image is an example of the sensor information of the present invention.
[0068] Next, the motion detection unit 200 detects a pre-articulation motion corresponding to an articulation motion based on the acquired near-end user image (S102).
[0069] Next, the signal processing determination unit 300 predicts an articulation operation sound generated along with the articulation operation based on the detected pre-articulation operation (S103), and determines the content of processing for the predicted articulation operation sound (S104).
[0070] Next, the audio acquisition unit 100 acquires, as a far-end audio signal, an audio signal received from the microphone 17C via the communication unit 19 (S105).
[0071] Next, based on the determined processing content, the signal processing unit 400 performs signal processing on the sound signal received in step 105 (S106) and ends the flow (END).
[0072] As described above, the user of the signal processing apparatus 1C can obtain a customer experience that allows them to concentrate on the meeting without hearing the pronunciation operation sound, which is an unnecessary sound for the meeting, among the sound signals received on the remote side.
[0073] <<Modification Example 1>> FIG. 11 is a functional block diagram of a signal processing apparatus 1D according to Modification Example 1. The functional configuration shown in FIG. 11 is realized by the CPU 12 and the DSP 13. In Modification Example 1, the DSP 13 also functions as an image processing unit that performs framing processing to cut out the speaker's image from the image signal acquired by the sensor 11C.
[0074] The signal processing apparatus 1D according to Modification Example 1 is different from the signal processing apparatus 1C in that it includes an image processing unit 600 and a display 20.
[0075] The motion detection unit 200 receives the remote user image acquired by the sensor 11C via the communication unit 19 and detects the pre - pronunciation motion. The image processing unit 600 performs a process of zooming out the user who performed the pre - pronunciation motion and outputs an image in which the user is not reflected to the display 20.
[0076] Also, the signal processing apparatus 1D may be used not only for remote calls with a remote location but also for broadcast contents such as live TV broadcasts and live distributions.
[0077] As described above, the user of the signal processing apparatus 1D can obtain a new customer experience that allows them to concentrate on the meeting without being distracted by the remote - side user performing the pronunciation operation.
[0078] Finally, the description of this embodiment should be considered as illustrative in all respects and not restrictive. The scope of the present invention is indicated by the scope of the claims rather than the above-described embodiments. Further, the scope of the present invention is intended to include all modifications within the meaning and scope equivalent to the scope of the claims.
Description of Reference Numerals
[0079] 1, 1A, 1B, 1C, 1D: signal processing devices, 11, 11C: sensors, 12: CPU, 13: DSP, 14: memory, 15: RAM, 16: user interface, 17, 17C: microphones, 18: speaker, 19: communication unit, 20: display, 100: voice acquisition unit, 200: motion detection unit, 300: signal processing determination unit, 400: signal processing unit, 500: determination unit, 600: image processing unit
Claims
1. Receiving sensor information from a sensor, detecting a pre - pronunciation operation corresponding to a pronunciation operation from the sensor information, predicting a sound signal emitted by the pronunciation operation based on the detected pre - pronunciation operation, determining the content of processing for the predicted sound, receiving a sound signal, executing signal processing on the received sound signal based on the determined content of processing, A signal processing method.
2. The sensor information includes information received from the far - end side in a remote call, The signal processing method according to claim 1.
3. Executing the signal processing, detecting a candidate sound signal for processing from the received sound signal, determining whether the detected candidate sound signal for processing matches the predicted sound, if they match, continuing the signal processing, if they do not match, canceling the signal processing, The signal processing method according to claim 1 or claim 2.
4. Executing the signal processing, detecting a candidate sound signal for processing from the received sound signal, determining whether the detected candidate sound signal for processing matches the predicted sound, if they match, increasing the intensity of the signal processing, if they do not match, canceling the signal processing, The signal processing method according to claim 1 or 2.
5. detecting a candidate sound signal for processing from the received sound signal, determining whether the detected candidate sound signal for processing matches the predicted sound, if they match, executing the signal processing on the candidate sound signal for processing, if they do not match, not executing the signal processing, The signal processing method according to claim 1 or claim 2.
6. The signal processing includes gain reduction processing of the sound signal, modification processing of the frequency characteristics of the sound signal, compression processing of the sound signal, beamforming processing, or noise cancellation processing, including any one of them, The signal processing method according to claim 1 or claim 2.
7. When detecting a preset specific operation from the sensor information, canceling the signal processing, The signal processing method according to claim 1 or claim 2.
8. detecting a candidate sound signal for processing from the received sound signal, determining whether the detected candidate sound signal for processing matches the predicted sound based on at least one of the feature amount of the candidate sound signal for processing, the level of the candidate sound signal for processing, or the arrival direction of the candidate sound signal for processing, if they match, executing the signal processing on the candidate sound signal for processing, The signal processing method according to claim 1 or 2.
9. The sensor information includes an image acquired by a camera. The signal processing method according to claim 1 or claim 2.
10. A sensor, An operation detection unit that receives sensor information from the sensor and detects a pre-articulation operation corresponding to an articulation operation from the sensor information; A signal processing determination unit that predicts a sound emitted by the articulation operation based on the detected pre-articulation operation and determines the content of processing for the predicted sound; An audio acquisition unit that receives an audio signal; A signal processing device comprising: a signal processing unit that executes signal processing on the received audio signal based on the determined content of the processing. Signal processing device.
11. The sensor information includes information received from the far-end side in a remote call. The signal processing device according to claim 10.
12. The signal processing unit executes the signal processing. The determination unit detects a candidate sound signal for processing from the received audio signal, determines whether the detected candidate sound signal for processing matches the predicted sound, if they match, continues the signal processing, if they do not match, cancels the signal processing. The signal processing device according to claim 10 or claim 11.
13. The signal processing unit executes the signal processing. The determination unit detects a candidate sound signal for processing from the received audio signal, determines whether the detected candidate sound signal for processing matches the predicted sound, if they match, increases the intensity of the signal processing, if they do not match, cancels the signal processing. The signal processing device according to claim 10 or 11.
14. The determination unit detects a candidate sound signal for processing from the received audio signal, determines whether the detected candidate sound signal for processing matches the predicted sound, if they match, executes the signal processing on the candidate sound signal for processing, if they do not match, does not execute the signal processing. The signal processing device according to claim 10 or claim 11.
15. The signal processing is any one of gain reduction processing of the audio signal, change processing of frequency characteristics of the audio signal, compression processing of the audio signal, beamforming processing, or noise cancellation processing. including any of them. The signal processing device according to claim 10 or claim 11.
16. When the signal processing unit detects a preset specific operation from the sensor information, the signal processing is cancelled. The signal processing device according to claim 10 or claim 11.
17. The determination unit detects a candidate sound signal for processing from the received audio signal. Based on at least one of the feature amount of the sound signal of the processing candidate, the level of the sound signal of the processing candidate, or the arrival direction of the sound signal of the processing candidate, determine whether the detected sound signal of the processing candidate matches the predicted sound. If they match, execute the signal processing on the sound signal of the processing candidate. The signal processing apparatus according to claim 10 or 11.
18. The sensor information includes an image acquired by a camera. The signal processing apparatus according to claim 10 or claim 11.
Citation Information
Patent Citations
Cough detector, and method and program for detecting cough
JP2021003181A