Sound signal processing method and sound signal processing device
By acquiring the speaker's image information to infer the speaker's posture and generate a correction filter, the speech processing problem affected by posture is solved, and stable and high-quality speech output is achieved.
Patent Information
- Application Number
- CN202111133047.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-09
- Filing Date
- 2021-09-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-09-27
AI Technical Summary
Existing technologies fail to effectively consider the impact of the speaker's posture on speech processing, resulting in poor speech processing results.
By acquiring the speaker's image information, inferring his or her posture information, and generating a correction filter corresponding to the posture, the sound signal is filtered to compensate for the attenuation of the voice and adjust the frequency characteristics.
It achieves speech processing that adapts to the speaker's posture, improves the quality and accuracy of the speech signal, and ensures stable output of high-quality speech in different postures.
Smart Images

Figure CN114420144B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One embodiment of the present application relates to a sound signal processing method and a sound signal processing apparatus that process a sound signal acquired by a microphone based on a position of a sound source. BACKGROUND
[0002] A sound processing system disclosed in Patent Literature 1 detects position information of a speaker from an image captured by a camera, and performs processing of enhancing a voice of the speaker based on the detected position information.
[0003] Patent Literature 1: Japanese Patent Application Laid-Open No. 2012-29209
[0004] A voice of a speaker changes in correspondence with a posture of the speaker. However, the sound processing system of Patent Literature 1 does not take into account the posture of the speaker. SUMMARY
[0005] Therefore, an object of one embodiment of the present application is to provide a sound signal processing method and a sound signal processing apparatus that can appropriately acquire a voice of a speaker in correspondence with a posture of the speaker.
[0006] The sound signal processing method inputs a sound signal related to a voice of a speaker, acquires an image of the speaker, estimates posture information of the speaker from the image of the speaker, generates a correction filter corresponding to the estimated posture information, performs filter processing related to the correction filter on the sound signal, and outputs the sound signal on which the filter processing is performed.
[0007] EFFECT OF THE INVENTION
[0008] According to one embodiment of the present application, a voice of a speaker can be appropriately acquired in correspondence with a posture of the speaker. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 is a block diagram showing a configuration of a sound signal processing apparatus.
[0010] Figure 2 is a flowchart showing an action of a sound signal processing method.
[0011] Figure 3 is a block diagram showing a functional configuration of a sound signal processing apparatus.
[0012] Figure 4 is a diagram showing one example of an image captured by a camera 11.
[0013] Figure 5 is a diagram showing one example of position information of a speaker.
[0014] Figure 6 is a block diagram showing the functional structure of the sound signal processing section 51.
[0015] Figure 7 is a block diagram showing the functional structure of the sound signal processing section 51 in the case where the residual characteristics are acquired.
[0016] Figure 8 is a diagram showing an example of the case where the correction filter is generated in correspondence with the recognition result of the table T.
[0017] Figure 9 is a flowchart showing the action of the sound signal processing method in the case where the correction filter is generated based on the posture information.
[0018] Figure 10 is a block diagram showing the functional structure of the sound signal processing apparatus.
[0019] Figure 11 is a diagram showing one example of the posture information.
[0020] Figure 12 is a block diagram showing the functional structure of the sound signal processing section 51.
[0021] Figure 13 is a block diagram showing the functional structure of the sound signal processing section 51 in the case where the residual characteristics are acquired. DETAILED DESCRIPTION
[0022] (First Embodiment)
[0023] Figure 1 is a block diagram showing the structure of the sound signal processing apparatus 1. Figure 2 is a flowchart showing the action of the sound signal processing method.
[0024] The sound signal processing apparatus 1 has a camera 11, a CPU 12, a DSP 13, a flash memory 14, a RAM 15, a user interface (I / F) 16, a speaker 17, six microphones 18A to 18F, and a communication section 19. Further, in the present embodiment, the signal represents a digital signal.
[0025] The camera 11, the speaker 17, and the microphones 18A to 18F are disposed, for example, above or below a display (not shown). The camera 11 acquires an image of a user present in front of the display (not shown). The microphones 18A to 18F acquire a voice of the user present in front of the display (not shown). The speaker 17 outputs a voice to the user present in front of the display (not shown). Furthermore, the number of microphones is not limited to six. The microphones can be one microphone. The number of microphones in the present embodiment is six, and an array microphone is constituted. The DSP 13 performs a beamforming process on a sound signal acquired by the microphones 18A to 18F.
[0026] The CPU 12 functions as a control section that controls the operation of the sound signal processing apparatus 1 by reading a program for operation from the flash memory 14 to the RAM 15. Furthermore, the program does not need to be stored in the flash memory 14 of the present apparatus in advance. It can be read to the RAM 15 by the CPU 12, for example, each time from a server or the like.
[0027] The DSP 13 is a signal processing section that processes an image signal and a sound signal respectively under the control of the CPU 12. The DSP 13 functions as an image processing section that performs, for example, a frame division process of cutting out an image of a speaker from the image signal. In addition, the DSP 13 also functions as a filter processing section that performs, for example, a correction filter process for compensating for attenuation of a voice of the speaker.
[0028] The communication section 19 transmits an image signal and a sound signal processed by the DSP 13 to another apparatus. In addition, the communication section 19 receives an image signal and a sound signal from another apparatus. The communication section 19 outputs the received image signal to a display (not shown). The communication section 19 outputs the received sound signal to the speaker 17. The display displays an image acquired by the camera of another apparatus. The speaker 17 outputs a voice of a speaker acquired by the microphone of another apparatus. Another apparatus is, for example, a sound signal processing apparatus provided at a remote location. Thus, the sound signal processing apparatus 1 functions as a communication system for performing a voice conversation with a remote location.
[0029] Figure 3 is a functional block diagram showing the sound signal processing apparatus 1. The above-described functional structure is realized by the CPU 12 and the DSP 13. As shown in Figure 3 , the sound signal processing apparatus 1 functionally has a sound signal input section 50, a sound signal processing section 51, an output section 52, an image acquisition section 100, a position estimation section 101, and a filter generation section 102.
[0030] The sound signal input section 50 inputs a sound signal from the microphones 18A to 18F (Sll). In addition, the image acquisition section 100 acquires an image including a speaker image from the camera 11 (S12). The position estimation section 101 estimates position information of the speaker from the acquired speaker image (S13).
[0031] The estimation of the position information includes a person face recognition process. The person face recognition process is a process of recognizing positions of faces of a plurality of persons from an image captured by the camera 11 by a predetermined algorithm such as a neural network. Hereinafter, in the present embodiment, a speaker refers to a person who is participating in a conference and is currently speaking, a user refers to a person who is participating in the conference, including the speaker. A non-user refers to a person who is not participating in the conference, and a person refers to all persons who are captured by the camera 11.
[0032] Figure 4 is an example of an image captured by the camera 11. In Figure 4 In the example of
[0033] The table T is a rectangle when viewed from above. The camera 11 captures images of four users who are located on the left and right sides of the table T in the width direction and a non-user who is located farther than the table T.
[0034] The position estimation section 101 recognizes faces of persons from the image captured by the camera 11. In Figure 4 In the example of
[0035] The position estimation section 101 sets a bounding box as shown by a box in the drawing to the position of the recognized face of the speaker. The position estimation section 101 calculates a distance to the speaker based on the size of the bounding box. A table or a function or the like showing a relationship between the size of the bounding box and the distance is stored in advance in the flash memory 14. The position estimation section 101 compares the size of the set bounding box with the table stored in the flash memory 14, and calculates the distance to the speaker.
[0036] The position estimation section 101 calculates the 2-dimensional coordinates (X, Y coordinates) of the set bounding box and the distance to the speaker as the position information of the speaker. Figure 5Fig. 1 is a diagram showing an example of the speaker's position information. The speaker's position information includes a label name indicating the speaker, 2-dimensional coordinates, and a distance. The 2-dimensional coordinates are X, Y coordinates (rectangular coordinates) with a prescribed position (e.g., the lower left) of the image captured by the camera 11 as the origin. The distance is a value expressed in meters, for example. The position estimation section 101 outputs the speaker's position information to the filter generation section 102. In addition, the position estimation section 101 outputs the position information of a plurality of speakers in the case where the faces of the plurality of speakers are recognized.
[0037] In addition, the position estimation section 101 can estimate the position information of the person based not only on the image captured by the camera 11 but also on the sound signals acquired by the microphones 18A to 18F. In this case, the position estimation section 101 inputs the sound signals acquired by the microphones 18A to 18F from the sound signal input section 50. For example, the position estimation section 101 can find the timing at which the voice of the person reaches the microphones by finding the mutual correlation of the sound signals acquired by the plurality of microphones. The position estimation section 101 can find the direction from which the voice of the person comes based on the positional relationship of the microphones and the arrival timing of the voice. In this case, the position estimation section 101 can also perform face recognition only from the image captured by the camera 11. For example, in the case where the face of the person is recognized from the image captured by the camera 11, the position estimation section 101 can find the direction from which the voice of the person comes based on the position of the recognized face and the arrival timing of the voice. Figure 4 In the example of Fig. 1, the position estimation section 101 recognizes the face images of four users who are on the left and right sides across the table T and a non-user who is at a position farther than the table T. Then, the position estimation section 101 estimates the face image coinciding with the direction from which the voice of the speaker comes as the position information of the speaker from the face images described above.
[0038] In addition, the position estimation section 101 can estimate the position information of the person by estimating the body of the person from the image captured by the camera 11. The position estimation section 101 finds the skeleton of the person from the image captured by the camera 11 according to a prescribed algorithm such as a neural network. The skeleton includes the eyes, the nose, the head, the shoulders, and the hands and feet. A table or a function or the like indicating the relationship between the size of the skeleton and the distance is stored in advance in the flash memory 14. The position estimation section 101 compares the size of the recognized skeleton with the table stored in the flash memory 14 to find the distance to the person.
[0039] Next, the filter generating section 102 generates a correction filter corresponding to the position information of the speaker (S14). The correction filter includes a filter for compensating for attenuation of the voice. The correction filter includes, for example, a gain correction, an equalizer, and beamforming. The farther the distance of the speaker's voice, the more attenuated it is. In addition, the high frequency component of the speaker's voice is more attenuated the farther the distance compared to the low frequency component of the speaker's voice. Therefore, the filter generating section 102 generates a gain correction filter that increases the level of the sound signal the greater the distance value in the position information of the speaker. In addition, the filter generating section 102 can also generate a filter of an equalizer that increases the level of the high frequency the greater the distance value in the position information of the speaker. In addition, the filter generating section 102 can also generate a correction filter that performs beamforming processing that directs the directivity toward the coordinates of the speaker.
[0040] The sound signal processing section 51 applies the filter processing involved in the correction filter generated by the filter generating section 102 to the sound signal (S15). The output section 52 outputs the sound signal after the filter processing to the communication section 19 (S16). The sound signal processing section 51 is constituted by, for example, a digital filter. The sound signal processing section 51 converts the sound signal to a signal on the frequency axis and changes the level of the signal of each frequency, thereby performing various filter processing.
[0041] Figure 6 is a block diagram showing the functional structure of the sound signal processing section 51. The sound signal processing section 51 constitutes a beamforming processing section 501, a gain correction section 502, and an equalizer 503. The beamforming processing section 501 applies filter processing to the sound signals acquired by the microphones 18A to 18F respectively and combines them, thereby performing beamforming. The signal processing involved in the beamforming can be various methods such as a Delay Sum method, a Griffiths Jim type, a Sidelobe Canceller type, or a Frost type adaptive beamformer.
[0042] The gain correction section 502 corrects the gain of the sound signal after the beamforming processing. The equalizer 503 adjusts the frequency characteristic of the sound signal after the gain correction. The filters of the beamforming processing, the filter of the gain correction section 502, and the filter of the equalizer 503 correspond to all of the correction filters. The filter generating section 102 generates the correction filters corresponding to the position information of the speaker.
[0043] The filter generating section 102 generates filter coefficients that form directivity toward the position of the speaker and sets them in the beamforming processing section 501. Thus, the sound signal processing apparatus 1 can acquire the speaker's voice with high precision.
[0044] Further, the filter generating section 102 sets the gain of the gain correction section 502 based on the position information of the speaker. As described above, the farther the distance of the speaker's voice, the more attenuated it is. Therefore, the filter generating section 102 generates a gain correction filter that makes the greater the distance value in the position information of the speaker, the higher the level of the sound signal, and sets it to the gain correction section 502. Thus, the sound signal processing apparatus 1 can obtain the speaker's voice at a stable level regardless of the distance from the speaker.
[0045] Further, the filter generating section 102 sets the frequency characteristic of the equalizer 503 based on the position information of the speaker. As described above, the filter generating section 102 generates a filter of the equalizer that makes the greater the distance value in the position information of the speaker, the higher the level of the high frequency. Thus, the sound signal processing apparatus 1 can obtain the speaker's voice at a stable sound quality regardless of the distance from the speaker.
[0046] Further, the filter generating section 102 can obtain the information of the direction of arrival of the voice from the beamforming processing section 501. As described above, the direction of arrival of the voice can be found based on the sound signals of the plurality of microphones. The filter generating section 102 can set the gain of the gain correction section 502 by comparing the position information of the person and the information of the direction of arrival of the voice. For example, the filter generating section 102 sets the value of the gain smaller the greater the difference (off-angle) between the position of the speaker indicated by the position information of the speaker and the direction of arrival of the voice. That is, the filter generating section 102 sets the gain in inverse proportion to the off-angle. Alternatively, the filter generating section 102 can set so that the gain exponentially decreases in correspondence with the off-angle. Alternatively, the filter generating section 102 can set so that the gain becomes 0 when the off-angle is equal to or greater than a prescribed threshold. Thus, the sound signal processing apparatus 1 can obtain the speaker's voice with higher accuracy.
[0047] Further, the filter generating section 102 can obtain the reverberation characteristic of the room and generate a correction filter in correspondence with the obtained reverberation characteristic. Figure 7 is a block diagram showing the functional structure of the sound signal processing section 51 in the case where the reverberation characteristic is obtained. Figure 7 The sound signal processing section 51 shown further has an adaptive echo canceller (AEC) 701.
[0048] AEC 701 estimates the component of the sound output from speaker 17 that returns to microphones 18A to 18F (echo component) and cancels the estimated echo component. The echo component is generated by applying adaptive filtering to the signal output to speaker 17. The adaptive filter is an FIR filter that simulates the reverberation characteristics of a room using a predetermined adaptive algorithm. The adaptive filter generates the echo component by filtering the signal output to speaker 17 using this FIR filter.
[0049] The filter generation unit 102 obtains the reverberation characteristics (reverberation information) simulated by the adaptive filter of the AEC 701. The filter generation unit 102 generates a correction filter based on the obtained reverberation information. For example, the filter generation unit 102 determines the power of the reverberation characteristics. The filter generation unit 102 sets the gain of the gain correction unit 502 based on the power of the reverberation characteristics. As described above, the filter generation unit 102 may also be configured so that the gain decreases exponentially with the departure angle. Alternatively, the filter generation unit 102 may be configured so that the exponential decay slows down as the power of the reverberation characteristics increases. In this case, the filter generation unit 102 sets a larger threshold as the power of the reverberation characteristics increases. Increasing this threshold blunts the directivity of the beam generated by the beamforming processing unit 501. In other words, the filter generation unit 102 blunts the directivity when the reverberation component is large. When the reverberation component is large, speech arrives from directions other than the actual speaker, reducing the accuracy of the estimated arrival direction. That is, a person may be present in a direction other than the estimated arrival direction, and the value of the separation angle may become larger. Therefore, when the reverberation component is large, the filter generation unit 102 blunts the directivity to prevent the speaker's voice from being unable to be acquired.
[0050] Further, the filter generation section 102 can reflect the result of the frame division processing to the correction filter in addition to the position information of the person. The user Al performs an operation of cutting out a specific region from an image captured by the camera 11 using the user I / F 16. The DSP 13 performs the frame division processing of cutting out the specified region. The filter generation section 102 sets the gain of the gain correction section 502 in correspondence with the boundary angle of the cut-out region and the direction of arrival of the voice. The filter generation section 102 sets the gain to 0 in a case where the direction of arrival of the voice exceeds the boundary angle of the cut-out region and departs from the cut-out region. Alternatively, the filter generation section 102 can be configured such that the greater the excess of the boundary angle of the cut-out region, the closer the gain to 0 is given in a case where the direction of arrival of the voice exceeds the boundary angle of the cut-out region and departs from the cut-out region. Further, the boundary angle can be set to the left and right sides, or to the left and right and upper and lower sides. Thus, the sound signal processing apparatus 1 can obtain the voice of the speaker in the region specified by the user with high accuracy.
[0051] Further, the filter generation section 102 can generate the correction filter in correspondence with the recognition result of the specific object. For example, the position estimation section 101 can recognize a table T as the specific object. Figure 8 is an example of a case where the correction filter is generated in correspondence with the recognition result of the table T. The position estimation section 101 recognizes the table T as the specific object by a predetermined algorithm such as a neural network. The position estimation section 101 outputs the position information of the table T to the filter generation section 102.
[0052] The filter generation section 102 generates the correction filter in correspondence with the position information of the table T. For example, as shown in Figure 8 , filter coefficients are generated such that directivity is formed toward regions S1 and S2 on the left and right sides at a position above the table T and across the table T in the width direction, and are set to the beamforming processing section 501. Alternatively, the filter generation section 102 can set the gain of the gain correction section 502 in correspondence with the difference (off angle) between the positions of the regions S1 and S2 and the direction of arrival of the voice. The filter generation section 102 sets the value of the gain to be smaller as the off angle is larger. Alternatively, the filter generation section 102 can be configured such that the gain exponentially decreases in correspondence with the off angle. Alternatively, the filter generation section 102 can be configured such that the gain becomes 0 in a case where the off angle is equal to or larger than a predetermined threshold. Alternatively, the filter generation section 102 can determine whether the position of the person is inside or outside the regions S1 and S2, and set the gain of the gain correction section 502 in such a manner that the gain becomes 0 in a case where the position of the person is outside.
[0053] Thus, the sound signal processing apparatus 1 can acquire the voices of the areas S1 and S2 that are located above the table and on the left and right sides of the table T in the width direction with high precision. For example, if it is the case of Figure 8 , the sound signal processing apparatus 1 can acquire the voices of the users Al, A2, A4, and A5 without acquiring the voice of the user A3.
[0054] Further, the filter generating section 102 can generate a correction filter that cuts off the voice of the corresponding person when the distance between the person and the table is equal to or greater than a predetermined value. For example, in the case of Figure 8 , when the user A3 speaks, the position estimating section 101 estimates the position of the user A3 as the position information of the speaker. However, the filter generating section 102 generates a correction filter that cuts off the voice of the user A3 considering that the distance from the person is equal to or greater than a predetermined value.
[0055] Further, the predetermined value can be found based on the recognition result of the specific object. For example, in the case of Figure 8 , the filter generating section 102 generates a correction filter that cuts off the voice of the person located farther than the table T.
[0056] (Second Embodiment)
[0057] Next, Figure 9 is a flowchart showing the operation of the sound signal processing method when the correction filter is generated based on the posture information. Figure 10 is a block diagram showing the functional structure of the sound signal processing apparatus 1 when the correction filter is generated based on the posture information. The sound signal processing apparatus 1 of this example has a posture estimating section 201 instead of the position estimating section 101. The hardware structure is the same as that shown in Figure 1 .
[0058] In the case of Figure 9 , instead of the position estimation processing (S13) of the position estimating section 101, the posture estimating section 201 estimates the posture information of the speaker from the acquired speaker image (S23). The other processing is the same as the flowchart shown in Figure 2 .
[0059] The estimation of the posture information includes a face recognition process of the speaker. The face recognition process of the speaker is the same as the estimation of the position information, and is a process of recognizing the position of the face of the speaker from the image captured by the camera 11 by a prescribed algorithm such as a neural network. The posture estimation section 201 recognizes the face of the speaker from the image captured by the camera 11. In addition, the posture estimation section 201 estimates the direction of the orientation of the speaker from the position of the eyes, the position of the mouth, and the position of the nose in the recognized face, and the like. For example, the flash memory 14 stores a table or a function that associates the deviation (offset) of the positions of the eyes, the mouth, and the nose from the face with the posture information. The posture estimation section 201 compares the offset of the positions of the eyes, the mouth, and the nose from the face with the table stored in the flash memory 14, and calculates the posture of the speaker. Furthermore, in a case where the position of the face is recognized but the eyes, the mouth, and the nose cannot be recognized, the posture estimation section 201 estimates the posture as a rearward posture.
[0060] Figure 11 is a diagram that shows one example of the posture information. The posture of the speaker is information that shows the orientation (angle) of the face to the left and right. For example, the posture estimation section 201 recognizes the posture of the user Al as 15 degrees. In this example, the posture estimation section 201 recognizes 0 degrees in a case where the orientation is forward, a positive angle in a case where the orientation is to the right, a negative angle in a case where the orientation is to the left, and 180 degrees (or -180 degrees) in a case where the orientation is directly rearward.
[0061] In addition, the posture estimation section 201 can also estimate the posture information from the body of the speaker from the image captured by the camera 11. The posture estimation section 201 recognizes the nose skeleton and the skeleton of the body (the head, the shoulders, the hands, and the feet, and the like) from the image captured by the camera 11 by a prescribed algorithm such as a neural network. The flash memory 14 stores a table or a function that associates the deviation (offset) of the nose skeleton and the skeleton of the body with the posture information in advance. The posture estimation section 201 can compare the offset of the nose skeleton of the skeleton from the body with the table stored in the flash memory 14, and calculate the posture of the speaker.
[0062] The filter generation section 102 generates a correction filter corresponding to the posture information. The correction filter includes a filter for compensating for the level attenuated corresponding to the orientation of the face. The correction filter includes, for example, a gain correction, an equalizer, and beamforming.
[0063] Figure 12 is a block diagram that shows the functional structure of the sound signal processing section 51. Figure 12 The block diagram shown is the same structure as the block diagram shown in Figure 6 except for the point that the filter generation section 102 inputs the posture information.
[0064] The speaker's voice shows the highest level in the case of facing straight ahead, and the greater the orientation to the left and right (the more to the left and right), the more attenuated. Also, the greater the orientation to the left and right, the more high frequencies are attenuated compared to low frequencies. Therefore, the filter generation section 102 generates a gain correction filter that increases the level of the sound signal the greater the orientation (angle) to the left and right, and sets it in the gain correction section 502. Also, the filter generation section 102 can generate a filter of an equalizer that increases the level of high frequencies or decreases the level of low frequencies the greater the orientation (angle) to the left and right, and set it in the equalizer 503.
[0065] Thus, the sound signal processing apparatus 1 can obtain the speaker's voice at a stable level and a stable sound quality regardless of the speaker's posture.
[0066] Also, the filter generation section 102 can control the directivity of the beamforming processing section 501 based on the posture information. The reverberation component shows the lowest level in the case of the speaker facing straight ahead, and becomes greater the greater the orientation to the left and right. Therefore, the filter generation section 102 can dull the directivity in the case of a large orientation (angle) to the left and right, judging that the reverberation component is large. Thus, the sound signal processing apparatus 1 can obtain the speaker's voice with high precision.
[0067] Also, as shown in Figure 13 , the filter generation section 102 can obtain the reverberation information. Figure 13 The structure is the same as in the example of Figure 7 . The filter generation section 102 obtains the reverberation information from the AEC 701. The filter generation section 102 generates a correction filter corresponding to the obtained reverberation information. For example, the filter generation section 102 calculates the power of the reverberation characteristic. The filter generation section 102 can also set the gain of the gain correction section 502 corresponding to the power of the reverberation characteristic.
[0068] The sound signal processing apparatus 1 of the first embodiment shows an example of generating a correction filter based on position information, and the sound signal processing apparatus 1 of the second embodiment generates a correction filter based on posture information. Of course, the sound signal processing apparatus 1 can generate a correction filter based on both the position information and the posture information. However, the estimation speed of the position information and the estimation speed of the posture information are sometimes different. The estimation speed of the position information of the sound signal processing apparatus 1 of the first embodiment is faster than the estimation speed of the posture information of the second embodiment. In this case, the filter generation section 102 can generate a correction filter at each of the timings when the position estimation section 101 estimates the position information and when the posture estimation section 201 estimates the posture information.
[0069] The explanations of the first and second embodiments are illustrative in all respects and are not restrictive. The scope of the present application is not represented by the above-described embodiments but by the claims. Also, the scope of the present application includes all modifications within the equivalent meanings and the scope of the claims.
[0070] Explanation of reference numerals
[0071] 1 sound signal processing apparatus
[0072] 11 camera
[0073] 12 CPU
[0074] 13 DSP
[0075] 14 flash memory
[0076] 15 RAM
[0077] 16 user I / F
[0078] 17 speaker
[0079] 18A to 18F microphone
[0080] 19 communication section
[0081] 50 sound signal input section
[0082] 51 sound signal processing section
[0083] 52 output section
[0084] 100 image acquisition section
[0085] 101 position estimation section
[0086] 102 filter generation section
[0087] 201 attitude estimation section
[0088] 501 beamforming processing section
[0089] 502 gain correction section
[0090] 503 equalizer
[0091] 701 AEC
Claims
1. A sound signal processing method, Input the sound signal related to the speaker's voice, Get the speaker image, Inferring the speaker's posture information based on the speaker image, generating a correction filter corresponding to the estimated posture information, performing filtering processing involving the correction filter on the sound signal, Output the sound signal after the filtering process is performed. in, The posture information includes information indicating the leftward or rightward orientation of the face, and the correction filter is generated corresponding to the leftward or rightward orientation of the face. The correction filter includes processing to increase the high-frequency level or decrease the low-frequency level as the leftward or rightward orientation of the face increases.
2. The sound signal processing method according to claim 1, wherein: The posture information includes the orientation of the speaker's face, The correction filter includes a process of compensating for a level that is attenuated according to the orientation of the face.
3. The sound signal processing method according to claim 1 or 2, wherein: The correction filter includes an equalizer.
4. The sound signal processing method according to claim 1 or 2, wherein: The posture information includes information on a backward posture.
5. The sound signal processing method according to claim 1 or 2, wherein: The position information of the speaker is estimated based on the speaker image. generating the correction filter based on the position information, The estimation speed of the position information is faster than the estimation speed of the posture information. The correction filter is generated at respective timings when the position information is estimated and when the posture information is estimated.
6. A sound signal processing device comprising: a sound signal input unit for inputting a sound signal related to a speaker's voice; an image acquiring unit that acquires an image of a speaker; a posture estimating unit for estimating posture information of the speaker based on the speaker image; a filter generating unit configured to generate a correction filter corresponding to the estimated posture information; a sound signal processing unit that performs filtering processing using the correction filter on the sound signal; and an output unit that outputs the sound signal after the filtering process is performed, The posture information includes information indicating the leftward and rightward orientation of the face, and the correction filter is generated corresponding to the leftward and rightward orientation of the face. The correction filter includes processing to increase the high-frequency level or reduce the low-frequency level as the leftward and rightward orientation of the face increases.
7. The sound signal processing device according to claim 6, wherein: The posture information includes the orientation of the speaker's face, The correction filter includes a process of compensating for a level that is attenuated according to the orientation of the face.
8. The sound signal processing device according to claim 6 or 7, wherein: The correction filter includes an equalizer.
9. The sound signal processing device according to claim 6 or 7, wherein: The posture information includes information on a backward posture.
10. The sound signal processing device according to claim 6 or 7, wherein: A position estimating unit is provided for estimating position information of the speaker based on the speaker image. The filter generation unit generates the correction filter based on the position information. The estimation speed of the position information is faster than the estimation speed of the posture information. The correction filter is generated at respective timings when the position information is estimated and when the posture information is estimated.
Citation Information
Patent Citations
Audio processing system
JP2012029209A
Speech signal enhancement using visual information
EP2766901A1
Directivity control device, sound collection system, directivity control method, and directivity control program
JP2019103009A
Sound source-separating device and sound source -separating method
US20160064000A1