Sound signal processing method and sound signal processing device
By acquiring the speaker's image and estimating their location information, a correction filter is generated for speech compensation, solving the problem of speech attenuation for distant speakers and achieving stable and high-precision speech acquisition.
Patent Information
- Application Number
- CN202111135988.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-09
- Filing Date
- 2021-09-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-09-27
AI Technical Summary
Existing technologies cannot effectively solve the problem of speech attenuation from distant speakers, resulting in the inability to acquire the speech of distant speakers at an appropriate level.
By acquiring an image of the speaker, estimating their location information, and generating a correction filter for compensation, the speech is processed using the correction filter to compensate for speech attenuation.
It enables the acquisition of speech at an appropriate level regardless of whether the speaker is near or far, thus improving the stability and accuracy of speech acquisition.
Smart Images

Figure CN114333873B_ABST
Abstract
Description
Technical Field
[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus that processes sound signals obtained by a microphone based on the location of the sound source. Background Technology
[0002] Patent document 1 discloses a sound processing system that detects the speaker's location information based on an image captured by a camera, and then enhances the speaker's speech based on the detected location information.
[0003] Patent Document 1: Japanese Patent Application Publication No. 2012-29209
[0004] The sound processing system in Patent Document 1 does not take into account the attenuation of the speech of a distant speaker. Therefore, the sound processing system in Patent Document 1 cannot acquire the speech of a distant speaker at an appropriate level. Summary of the Invention
[0005] Therefore, one embodiment of the present invention aims to provide a sound signal processing method and a sound signal processing apparatus that can acquire sound signals at an appropriate level, whether the speech of a speaker is far away or near.
[0006] The sound signal processing method takes a sound signal related to a speaker's speech as input, obtains a speaker image, estimates the speaker's position information based on the speaker image, generates a correction filter corresponding to the estimated position information to compensate for the attenuation of the speech, applies the filtering processing involved in the correction filter to the sound signal, and outputs the sound signal after the filtering processing.
[0007] The effects of the invention
[0008] According to one embodiment of the present invention, the voice of either a speaker at a distance or a speaker at a nearby location can be acquired at an appropriate level. Attached Figure Description
[0009] Figure 1 It is a block diagram representing the structure of a sound signal processing device.
[0010] Figure 2 It is a flowchart representing the actions of a sound signal processing method.
[0011] Figure 3 It is a block diagram representing the functional structure of a sound signal processing device.
[0012] Figure 4 This is a diagram showing an example of an image captured by camera 11.
[0013] Figure 5 This is a diagram representing an example of speaker location information.
[0014] Figure 6 This is a block diagram showing the functional structure of the sound signal processing unit 51.
[0015] Figure 7 This is a block diagram illustrating the functional structure of the sound signal processing unit 51, which represents the condition of obtaining reverberation characteristics.
[0016] Figure 8 This is a diagram illustrating an example of generating a correction filter in response to the recognition result of table T.
[0017] Figure 9 This is a flowchart illustrating the actions of a sound signal processing method that generates a correction filter based on attitude information.
[0018] Figure 10 It is a block diagram representing the functional structure of a sound signal processing device.
[0019] Figure 11 This is a diagram representing an example of attitude information.
[0020] Figure 12 This is a block diagram showing the functional structure of the sound signal processing unit 51.
[0021] Figure 13 This is a block diagram showing the functional structure of the sound signal processing unit 51 when the reverberation characteristics are obtained. Detailed Implementation
[0022] (First Embodiment)
[0023] Figure 1 This is a block diagram showing the structure of the sound signal processing device 1. Figure 2 It is a flowchart representing the actions of a sound signal processing method.
[0024] The sound signal processing device 1 includes a camera 11, a CPU 12, a DSP 13, a flash memory 14, a RAM 15, a user interface (I / F) 16, a speaker 17, six microphones 18A to 18F, and a communication unit 19. Furthermore, in this embodiment, the signal represents a digital signal.
[0025] Camera 11, speaker 17, and microphones 18A-18F are disposed, for example, above or below a display (not shown). Camera 11 acquires an image of the user in front of the display (not shown). Microphones 18A-18F acquire the voice of the user in front of the display (not shown). Speaker 17 outputs voice to the user in front of the display (not shown). Furthermore, the number of microphones is not limited to six. A single microphone may be used. In this embodiment, six microphones are used, forming an array microphone. DSP 13 performs beamforming processing on the sound signals acquired by microphones 18A-18F.
[0026] The CPU 12 functions as a control unit that centrally controls the operation of the audio signal processing device 1 by reading the operating program from the flash memory 14 into the RAM 15. Furthermore, the program does not need to be pre-stored in the flash memory 14 of this device. It can be read into the RAM 15 by the CPU 12, for example, by downloading it from a server each time.
[0027] DSP 13 is a signal processing unit that processes video signals and audio signals separately under the control of CPU 12. DSP 13 functions as an image processing unit, for example, performing frame-segmentation processing to cut out the speaker's image from the video signal. Additionally, DSP 13 also functions as a filtering processing unit, for example, performing correction filtering processing to compensate for attenuation of the speaker's voice.
[0028] The communication unit 19 transmits the image and audio signals processed by the DSP 13 to other devices. Additionally, the communication unit 19 receives image and audio signals from other devices. The communication unit 19 outputs the received image signal to a display (not shown). The communication unit 19 outputs the received audio signal to a speaker 17. The display shows the image captured by the camera of the other device. The speaker 17 outputs the voice of the speaker captured by the microphone of the other device. The other device is, for example, an audio signal processing device located at a remote location. Thus, the audio signal processing device 1 functions as a communication system for conducting voice conversations with a remote location.
[0029] Figure 3 This is a functional block diagram representing the sound signal processing device 1. The above functional structure is implemented by CPU 12 and DSP 13. For example... Figure 3 As shown, the sound signal processing device 1 functionally includes a sound signal input unit 50, a sound signal processing unit 51, an output unit 52, an image acquisition unit 100, a position estimation unit 101, and a filter generation unit 102.
[0030] The sound signal input unit 50 inputs sound signals from microphones 18A to 18F (S11). Meanwhile, the image acquisition unit 100 acquires an image containing the speaker's image from the camera 11 (S12). The position estimation unit 101 estimates the speaker's position information based on the acquired speaker image (S13).
[0031] Location information estimation includes facial recognition processing. Facial recognition processing involves identifying the position of the faces of multiple individuals based on images captured by camera 11 using a prescribed algorithm, such as a neural network. Hereinafter, in this embodiment, "speaker" refers to someone attending a meeting and currently speaking, "user" refers to someone attending the meeting, including the speaker, "non-user" refers to someone not attending the meeting, and "person" refers to everyone visible in camera 11.
[0032] Figure 4 This is a diagram showing an example of an image captured by camera 11. Figure 4 For example, camera 11 takes pictures of the faces of multiple people along the length (depth) of table T.
[0033] The table T is rectangular when viewed from above. Camera 11 photographs four users located to the left and right of the table T in the width direction, as well as non-users located further away from the table T.
[0034] The position estimation unit 101 identifies a person's face based on images captured by the camera 11. Figure 4 In the example, user A1, located in the lower left of the image, is speaking. The position estimation unit 101, based on multiple frames of images, identifies the face of user A1 as the speaker's face. Furthermore, other individuals A2 to A5 are also identified, but they are not the speakers. Therefore, the position estimation unit 101 identifies user A1's face as the speaker's face.
[0035] The position estimation unit 101 sets a bounding box (represented by a square) for the position of the identified speaker's face, as shown in the figure. The position estimation unit 101 calculates the distance to the speaker based on the size of the bounding box. A table or function representing the relationship between the size of the bounding box and the distance is pre-stored in the flash memory 14. The position estimation unit 101 compares the set size of the bounding box with the table stored in the flash memory 14 to calculate the distance to the speaker.
[0036] The position estimation unit 101 calculates the 2D coordinates (X, Y coordinates) of the set bounding box and its distance from the speaker, as the speaker's position information. Figure 5This is an example diagram representing the speaker's location information. The speaker's location information includes a label indicating the speaker, 2D coordinates, and distance. The 2D coordinates are X and Y coordinates (Cartesian coordinates) with a predetermined position (e.g., lower left) of the image captured by camera 11 as the origin. The distance is a value expressed in meters, for example. The location estimation unit 101 outputs the speaker's location information to the filter generation unit 102. Furthermore, when multiple speakers' faces are identified, the location estimation unit 101 outputs the location information of multiple speakers.
[0037] Furthermore, the position estimation unit 101 can estimate the position information of a person not only based on the image captured by the camera 11, but also based on the sound signals obtained by the microphones 18A to 18F. In this case, the position estimation unit 101 inputs the sound signals obtained by the microphones 18A to 18F from the sound signal input unit 50. For example, the position estimation unit 101 can determine the timing of the person's voice arriving at the microphone by calculating the correlation between the sound signals obtained by multiple microphones. The position estimation unit 101 can determine the direction of arrival of the person's voice based on the positional relationship of each microphone and the arrival timing of the voice. In this case, the position estimation unit 101 can also perform facial recognition only based on the image captured by the camera 11. For example, in Figure 4 For example, the position estimation unit 101 identifies the facial images of four users located on the left and right sides of the table T in the width direction, as well as non-users located further away from the table T. Furthermore, based on the aforementioned facial images, the position estimation unit 101 estimates the speaker's position information as the facial image that aligns with the direction of arrival of the speaker's voice.
[0038] Furthermore, the position estimation unit 101 can also estimate the person's body and position information based on the images captured by the camera 11. The position estimation unit 101 calculates the person's skeleton (bones) based on the images captured by the camera 11 using a prescribed algorithm such as a neural network. The skeleton includes eyes, nose, head, shoulders, hands, and feet. A table or function representing the relationship between the size and distance of the bones is pre-stored in the flash memory 14. The position estimation unit 101 compares the size of the identified bones with the table stored in the flash memory 14 to calculate the distance to the person.
[0039] Next, the filter generation unit 102 generates a correction filter (S14) corresponding to the speaker's position information. The correction filter includes a filter for compensating for speech attenuation. The correction filter includes, for example, gain correction, an equalizer, and beamforming. The speaker's speech attenuates more at greater distances. Furthermore, compared to the low-frequency components of the speaker's speech, the high-frequency components attenuate even more at greater distances. Therefore, the filter generation unit 102 generates a gain correction filter that increases the level of the sound signal as the distance value in the speaker's position information increases. Additionally, the filter generation unit 102 can also generate an equalizer filter that increases the level of high-frequency frequencies as the distance value in the speaker's position information increases. Furthermore, the filter generation unit 102 can also generate a correction filter that performs beamforming processing to direct the beam towards the speaker's coordinates.
[0040] The audio signal processing unit 51 performs filtering processing on the audio signal, involving the correction filter generated by the filter generation unit 102 (S15). The output unit 52 outputs the filtered audio signal to the communication unit 19 (S16). The audio signal processing unit 51 is, for example, composed of a digital filter. The audio signal processing unit 51 converts the audio signal into a signal on the frequency axis and changes the level of the signal at each frequency, thereby performing various filtering processes.
[0041] Figure 6 This is a block diagram illustrating the functional structure of the audio signal processing unit 51. The audio signal processing unit 51 comprises a beamforming processing unit 501, a gain correction unit 502, and an equalizer 503. The beamforming processing unit 501 filters and synthesizes the audio signals obtained from microphones 18A to 18F, thereby performing beamforming. The signal processing involved in beamforming can be various methods such as delay sum, Griffiths Jim type, Sidelobe Canceller type, or Frost type adaptive beamformer.
[0042] The gain correction unit 502 corrects the gain of the beamformed audio signal. The equalizer 503 adjusts the frequency response of the gain-corrected audio signal. The beamformer filter, the gain correction unit 502 filter, and the equalizer 503 filter correspond to all the correction filters. The filter generation unit 102 generates correction filters in accordance with the speaker's position information.
[0043] The filter generation unit 102 generates filter coefficients that form a directivity towards the speaker's position and sets them in the beamforming processing unit 501. As a result, the audio signal processing device 1 can acquire the speaker's voice with high accuracy.
[0044] Furthermore, the filter generation unit 102 sets the gain of the gain correction unit 502 based on the speaker's location information. As described above, the speaker's voice attenuates more as the distance increases. Therefore, the filter generation unit 102 generates a gain correction filter that increases the level of the sound signal as the distance value in the speaker's location information increases, and sets it in the gain correction unit 502. As a result, the sound signal processing device 1 can acquire the speaker's voice at a stable level regardless of the speaker's distance.
[0045] Furthermore, the filter generation unit 102 sets the frequency characteristics of the equalizer 503 based on the speaker's location information. As described above, the filter generation unit 102 generates an equalizer filter that increases the high-frequency level as the distance value in the speaker's location information increases. As a result, the audio signal processing device 1 can acquire the speaker's voice with stable sound quality regardless of the speaker's distance.
[0046] Furthermore, the filter generation unit 102 can also obtain information about the direction of arrival of the speech from the beamforming processing unit 501. As described above, the direction of arrival of the speech can be determined based on the sound signals from multiple microphones. The filter generation unit 102 can also set the gain of the gain correction unit 502 by comparing the speaker's position information and the information about the direction of arrival of the speech. For example, the filter generation unit 102 sets the gain value to be smaller if the difference (angle of departure) between the speaker's position shown in the speaker's position information and the direction of arrival of the speech is larger. That is, the filter generation unit 102 sets the gain to be inversely proportional to the angle of departure. Alternatively, the filter generation unit 102 can be set to make the gain decrease exponentially in relation to the angle of departure. Alternatively, the filter generation unit 102 can be set to make the gain become 0 when the angle of departure is above a predetermined threshold. As a result, the sound signal processing device 1 can obtain the speaker's speech with higher accuracy.
[0047] In addition, the filter generation unit 102 can also obtain the reverberation characteristics of the room and generate a correction filter in accordance with the obtained reverberation characteristics. Figure 7 This is a block diagram showing the functional structure of the sound signal processing unit 51 when the reverberation characteristics are obtained. Figure 7 The audio signal processing unit 51 shown also has an adaptive echo canceller (AEC) 701.
[0048] The AEC 701 estimates the echo component (reverberation component) returned to microphones 18A-18F from the sound output from speaker 17 and cancels the estimated echo component. The echo component is generated by adaptively filtering the signal output to speaker 17. The adaptive filter is an FIR filter that simulates the reverberation characteristics of the room using a prescribed adaptive algorithm. The adaptive filter generates the echo component by filtering the signal output to speaker 17 using this FIR filter.
[0049] The filter generation unit 102 acquires the reverberation characteristics (reverberation information) simulated by the adaptive filter of AEC 701. The filter generation unit 102 generates a correction filter corresponding to the acquired reverberation information. For example, the filter generation unit 102 calculates the power of the reverberation characteristics. The filter generation unit 102 sets the gain of the gain correction unit 502 corresponding to the power of the reverberation characteristics. As described above, the filter generation unit 102 can also be set to make the gain decrease exponentially in relation to the angle of departure. Alternatively, the filter generation unit 102 can be set such that the attenuation index decreases more slowly as the power of the reverberation characteristics increases. In the above case, the filter generation unit 102 sets a larger threshold as the power of the reverberation characteristics increases. If this threshold is increased, the directivity of the beam generated by the beamforming processing unit 501 becomes blunted. That is, the filter generation unit 102 blunts the directivity when the reverberation component is large. When the reverberation component is large, speech from directions other than the actual speaker's direction may arrive, thus reducing the accuracy of the estimated direction of arrival. That is, there may be people in directions other than the presumed direction of arrival, and sometimes the value of the aforementioned angle becomes larger. Therefore, the filter generation unit 102 blunts the directionality when the reverberation component is large, to prevent the speaker's voice from being unable to be obtained.
[0050] Furthermore, the filter generation unit 102 can also reflect the results of frame processing to the correction filter in addition to the position information of the person. User A1 uses the user I / F 16 to cut out a specific region from the image captured by camera 11. DSP 13 performs frame processing to cut out the specified region. The filter generation unit 102 sets the gain of the gain correction unit 502 in accordance with the boundary angle of the cut-out region and the direction of arrival of the speech. The filter generation unit 102 sets the gain to 0 when the direction of arrival of the speech exceeds the boundary angle of the cut-out region and detaches from the cut-out region. Alternatively, when the direction of arrival of the speech exceeds the boundary angle of the cut-out region and detaches from the cut-out region, the filter generation unit 102 can be configured to assign a gain closer to 0 the larger the deviation from the boundary angle. Furthermore, the boundary angle can be set to the left and right sides, or to any of the four directions (left, right, up, down). Thus, the sound signal processing device 1 can obtain the speaker's voice in the region specified by the user with high accuracy.
[0051] Furthermore, the filter generation unit 102 can also generate a correction filter corresponding to the recognition result of a specific object. For example, the position estimation unit 101 can identify the table T as a specific object. Figure 8 This diagram illustrates an example of generating a correction filter corresponding to the recognition result of table T. The position estimation unit 101 identifies table T as a specific object using a predetermined algorithm such as a neural network. The position estimation unit 101 outputs the position information of table T to the filter generation unit 102.
[0052] The filter generation unit 102 generates a correction filter corresponding to the position information of the table T. For example, such as Figure 8 As shown, filter coefficients are generated such that the direction is directional, forming regions S1 and S2 located to the left and right of table T in the width direction, which are higher than the position of table T. These coefficients are then set in the beamforming processing unit 501. Alternatively, the filter generation unit 102 may set the gain of the gain correction unit 502 in accordance with the difference (angle of departure) between the positions of regions S1 and S2 and the direction of arrival of the speech. The filter generation unit 102 sets a smaller gain value for larger angles of departure. Alternatively, the filter generation unit 102 may set the gain to decrease exponentially in accordance with the angle of departure. Alternatively, the filter generation unit 102 may set the gain to become 0 when the angle of departure is above a predetermined threshold. Alternatively, the filter generation unit 102 may determine whether the person's position is inside or outside regions S1 and S2, and set the gain of the gain correction unit 502 to become 0 when the person's position is outside.
[0053] Therefore, the sound signal processing device 1 can acquire, with high precision, the speech in regions S1 and S2 located on the left and right sides of the table, which are positioned higher than the table and separated by the table T in the width direction. For example, if it is Figure 8 For example, the voice signal processing device 1 can acquire the voices of users A1, A2, A4, and A5 without acquiring the voice of user A3.
[0054] Additionally, the filter generation unit 102 can also generate a correction filter that cuts off the corresponding person's speech when the distance between the person and the table is greater than a predetermined value. For example, in Figure 8 In the example, when user A3 speaks, the position estimation unit 101 estimates user A3's position as the speaker's position information. However, the filter generation unit 102 generates a correction filter that considers the distance to the speaker to be above a predetermined value and cuts off user A3's speech.
[0055] Furthermore, the specified value can be derived based on the identification results of a specific object. For example, in Figure 8 In the example, the filter generation unit 102 generates a correction filter that cuts off speech located further away from the table T.
[0056] (Second Implementation)
[0057] then, Figure 9 This is a flowchart illustrating the actions of a sound signal processing method that generates a correction filter based on attitude information. Figure 10 This is a block diagram illustrating the functional structure of the audio signal processing apparatus 1 in the case of generating a correction filter based on attitude information. In this example, the audio signal processing apparatus 1 has an attitude estimation unit 201 instead of a position estimation unit 101. Hardware structure and Figure 1 The structures shown are the same.
[0058] exist Figure 9 In this example, instead of the position estimation process of the position estimation unit 101 (S13), the posture estimation unit 201 estimates the speaker's posture information based on the acquired speaker image (S23). Other processing is similar to... Figure 2 The flowchart shown is the same.
[0059] The estimation of posture information includes speaker facial recognition processing. Similar to the estimation of position information, the speaker facial recognition processing involves identifying the position of the speaker's face based on images captured by camera 11 using a predetermined algorithm, such as a neural network. The posture estimation unit 201 identifies the speaker's face based on images captured by camera 11. Furthermore, the posture estimation unit 201 estimates the speaker's facing direction based on the positions of the eyes, mouth, and nose on the identified face. For example, the flash memory 14 stores tables or functions that correlate the deviations (offsets) of the eye, mouth, and nose positions relative to the face with posture information. The posture estimation unit 201 compares the offsets of the eye, mouth, and nose positions relative to the face with the tables stored in flash memory 14 to determine the speaker's posture. Additionally, if the posture estimation unit 201 identifies the position of the face but cannot identify the eyes, mouth, and nose, it estimates a backward-facing posture.
[0060] Figure 11 This is a diagram illustrating an example of posture information. A speaker's posture is information indicating the direction (angle) of their face to the left or right. For example, posture estimation unit 201 identifies user A1's posture as 15 degrees. In this example, posture estimation unit 201 identifies it as 0 degrees when facing forward, a positive angle when facing to the right, a negative angle when facing to the left, and 180 degrees (or -180 degrees) when facing directly behind.
[0061] Furthermore, the posture estimation unit 201 can also estimate the speaker's body and posture information based on the images captured by the camera 11. The posture estimation unit 201 uses a predetermined algorithm, such as a neural network, to identify the nasal bones and the bones of the body (head, shoulders, hands, feet, etc.) based on the images captured by the camera 11. Tables or functions that correlate the deviation (offset) of the nasal bones and body bones with posture information are pre-stored in the flash memory 14. The posture estimation unit 201 can compare the offset of the nasal bones relative to the body with the tables stored in the flash memory 14 to determine the speaker's posture.
[0062] The filter generation unit 102 generates a correction filter in accordance with the attitude information. The correction filter includes a filter for compensating for the attenuation of the level corresponding to the orientation of the face. The correction filter includes, for example, gain correction, equalization, and beamforming.
[0063] Figure 12 This is a block diagram showing the functional structure of the sound signal processing unit 51. Figure 12 The block diagram shown illustrates the relationship between the filter generation unit 102 and the input of attitude information. Figure 6 The block diagram shown has the same structure.
[0064] The speaker's voice exhibits the highest level when facing directly forward, and attenuates more as it moves to the left or right. Furthermore, the greater the left or right direction, the more the high frequencies are attenuated compared to the low frequencies. Therefore, the filter generation unit 102 generates a gain correction filter that increases the level of the audio signal as the left or right direction (angle) increases, and sets it in the gain correction unit 502. Additionally, the filter generation unit 102 can also generate an equalizer filter that increases the level of high frequencies or decreases the level of low frequencies as the left or right direction (angle) increases, and sets it in the equalizer 503.
[0065] Therefore, the sound signal processing device 1 can acquire the speaker's voice with a stable level and stable sound quality regardless of the speaker's posture.
[0066] Furthermore, the filter generation unit 102 can also control the directivity of the beamforming processing unit 501 based on attitude information. The reverberation component exhibits the lowest level when the speaker is facing directly forward, and becomes larger as the direction to the left or right increases. Therefore, the filter generation unit 102 can also determine that the reverberation component is large when the direction (angle) to the left or right is large, and thus dull the directivity. As a result, the sound signal processing device 1 can acquire the speaker's voice with high accuracy.
[0067] In addition, such as Figure 13 As shown, the filter generation unit 102 can also obtain reverberation information. Figure 13 Structure and Figure 7 The example is the same. The filter generation unit 102 obtains the reverberation information from the AEC 701. The filter generation unit 102 generates a correction filter corresponding to the obtained reverberation information. For example, the filter generation unit 102 calculates the power of the reverberation characteristic. The filter generation unit 102 may also set the gain of the gain correction unit 502 corresponding to the power of the reverberation characteristic.
[0068] The first embodiment of the sound signal processing apparatus 1 shows an example of generating a correction filter based on position information, while the second embodiment of the sound signal processing apparatus 1 generates a correction filter based on attitude information. Of course, the sound signal processing apparatus 1 can also generate a correction filter based on both position information and attitude information. However, the estimation speed of position information and the estimation speed of attitude information are sometimes different. The estimation speed of position information in the first embodiment of the sound signal processing apparatus 1 is faster than the estimation speed of attitude information in the second embodiment. In this case, the filter generation unit 102 can generate the correction filter at the respective timings when the position estimation unit 101 estimates the position information and when the attitude estimation unit 201 estimates the attitude information.
[0069] The descriptions of the first and second embodiments are illustrative in all respects and are not restrictive. The scope of the invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the invention includes all equivalents of the claims and all modifications within that scope.
[0070] Explanation of the label
[0071] 1…Sound signal processing device
[0072] 11…camera
[0073] 12…CPU
[0074] 13…DSP
[0075] 14… Flash Memory
[0076] 15…RAM
[0077] 16…User I / F
[0078] 17…speakers
[0079] 18A~18F… Microphones
[0080] 19…Ministry of Communications
[0081] 50…Audio signal input section
[0082] 51…Sound Signal Processing Department
[0083] 52… Output Department
[0084] 100…Image Acquisition Department
[0085] 101…Position Estimation Section
[0086] 102…Filter Generation Section
[0087] 201…Attitude Estimation Department
[0088] 501…Beamforming Processing Department
[0089] 502…Gain Correction Section
[0090] 503…Equalizer
[0091] 701…AEC
Claims
1. A sound signal processing method, Input the sound signal related to the speaker's voice. Obtain speaker image, The speaker's location information is inferred based on the speaker's image. Determine the direction of arrival of the speaker's voice. A correction filter is generated that compensates for the attenuation of the speech at least in accordance with the angle between the speaker's position as indicated by the estimated speaker's position information and the direction of arrival of the speaker's speech. The correction filter includes gain correction, and is configured such that the larger the angle of deviation, the smaller the gain value of the speaker's voice. The sound signal is subjected to the filtering process described in the correction filter. The sound signal after the filtering process is applied will be output.
2. The sound signal processing method according to claim 1, wherein, The location information includes the distance to the speaker. The correction filter includes a process for compensating for the level attenuation corresponding to the distance.
3. The sound signal processing method according to claim 1 or 2, wherein, Obtaining reverberation characteristics, The correction filter is generated in correspondence with the angle between the speaker's position and the direction of arrival of the speaker's voice as indicated by the estimated speaker's position information, and the obtained reverberation characteristics.
4. The sound signal processing method according to claim 1 or 2, wherein, The correction filter includes beamforming. The sound signal processing method obtains the reverberation characteristics. The directivity of the beamforming is changed in accordance with the obtained reverberation characteristics.
5. The sound signal processing method according to claim 1 or 2, wherein, The speaker's image is segmented into frames to cut out the specified region. The correction filter sets the gain of the sound signal in accordance with the boundary angle of the specified region and the direction of arrival of the speaker's voice.
6. The sound signal processing method according to claim 1 or 2, wherein, Identify specific objects. The correction filter is generated and corresponds to the angle between the speaker's position and the direction of arrival of the speaker's voice as shown by the estimated speaker's position information, and the recognition result of the specific object.
7. The sound signal processing method according to claim 1 or 2, wherein, The location information includes the distance to the speaker. The correction filter includes a process that cuts off the corresponding speaker's speech when the distance is above a predetermined value.
8. The sound signal processing method according to claim 1 or 2, wherein, The speaker's posture information is inferred based on the speaker's image. The correction filter is generated based on the estimated speaker position information, the angle between the speaker's position and the direction of arrival of the speaker's voice, and the posture information. The estimation speed of the position information is faster than the estimation speed of the attitude information. The correction filter is generated at its respective timing when the position information is estimated and when the attitude information is estimated.
9. A sound signal processing device, comprising: The sound signal input unit inputs a sound signal related to the speaker's speech. The image acquisition unit acquires the speaker's image; The location estimation unit estimates the location information of the speaker based on the speaker image and determines the direction of arrival of the speaker's voice. A filter generation unit generates a correction filter that compensates for the attenuation of the speech at least in accordance with the angle between the speaker's position indicated by the estimated speaker's position information and the direction of arrival of the speaker's speech. The correction filter includes gain correction, and is set such that the larger the angle, the smaller the gain value of the speaker's speech is set. The audio signal processing unit performs filtering processing on the audio signal in accordance with the correction filter. as well as The output section outputs the audio signal after the filtering process has been performed.
10. The sound signal processing apparatus according to claim 9, wherein, The location information includes the distance to the speaker. The correction filter includes a process for compensating for the level attenuation corresponding to the distance.
11. The sound signal processing apparatus according to claim 9 or 10, wherein, The filter generation unit obtains the reverberation characteristics and generates the correction filter in correspondence with the angle between the speaker's position and the direction of arrival of the speaker's voice as indicated by the estimated speaker's position information and the obtained reverberation characteristics.
12. The sound signal processing apparatus according to claim 9 or 10, wherein, The correction filter includes beamforming. The filter generation unit obtains the reverberation characteristics and modifies the directivity of the beamforming accordingly.
13. The sound signal processing apparatus according to claim 9 or 10, wherein, The sound signal processing device includes an image processing unit that performs frame-segmentation processing on the speaker's image to cut out a specified region. The filter generation unit sets the gain of the sound signal in accordance with the boundary angle of the specified region and the direction of arrival of the speaker's voice.
14. The sound signal processing apparatus according to claim 9 or 10, wherein, The location estimation unit identifies a specific object. The filter generation unit generates a correction filter that corresponds to the angle between the speaker's position and the direction of arrival of the speaker's voice as indicated by the estimated speaker's position information, and the recognition result of the specific object.
15. The sound signal processing apparatus according to claim 9 or 10, wherein, The location information includes the distance to the speaker. The correction filter includes a process that cuts off the corresponding speaker's speech when the distance is above a predetermined value.
16. The sound signal processing apparatus according to claim 9 or 10, wherein, The device includes a posture estimation unit that estimates the speaker's posture information based on the speaker image. The filter generation unit generates the correction filter based on the estimated speaker position information, the angle between the speaker's position and the direction of arrival of the speaker's voice, and the posture information. The estimation speed of the position information is faster than the estimation speed of the attitude information. The correction filter is generated at its respective timing when the position information is estimated and when the attitude information is estimated.
Citation Information
Patent Citations
Audio processing system
JP2012029209A
Sound direction positioning processing method, device and system, computer equipment and storage medium
CN111048113A
Speech Signal Enhancement Using Visual Information
US20140337016A1