Video and audio conferencing equipment, terminal equipment, sound source localization method and medium

By placing the microphone array on different planes of the base and display screen in the audio-visual conferencing equipment, and combining the sound source positioning unit and Kalman filtering method, the problem that the microphone array cannot perceive the vertical direction changes of the sound source is solved, and three-dimensional precise positioning of the sound source is achieved, improving the camera tracking capability and user experience.

CN114095687BActive Publication Date: 2025-07-04C SKY MICROSYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010752044.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-30
Publication Date
2025-07-04
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

The microphone array of existing audio-visual conferencing devices can only sense changes in the horizontal direction of the sound source, but cannot sense changes in the vertical direction, which makes the camera unable to track the face of the person standing up and cannot achieve accurate positioning of the sound source.

Method used

Part of the microphone array is located at the first plane of the base and the other part is located at the second plane of the display screen. Combined with the sound source positioning unit and the Kalman filtering method, three-dimensional precise positioning of the sound source is achieved through time delay term optimization and regular term optimization.

Benefits of technology

It realizes precise positioning of the horizontal and vertical directions of the sound source, improves the accuracy and user experience of camera tracking, and enhances the functions of audio-visual conferencing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114095687B_ABST
    Figure CN114095687B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio-video conferencing device, a terminal device, a sound source localization method, and a medium. The audio-video conferencing device includes a base extending along a first plane, a display screen extending from the base along a second plane, a microphone array, and a sound source localization unit. At least some microphones in the microphone array are encapsulated in the base, and at least some microphones are encapsulated in the display screen. The sound source localization unit determines the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array. The embodiments of the present disclosure can identify changes in the vertical direction of the sound source and achieve accurate sound source localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of electronics, and more particularly, to an audio-video conferencing device, a terminal device, a sound source localization method, and a medium. Background Art

[0002] Currently, microphone arrays are widely used in audio-video conferencing devices to capture sound and calculate the sound source position based on the captured sound, so that the camera of the audio-video conferencing device rotates following the sound source, achieving the purpose of timely capturing the video of the speaker in the audio-video conference. Generally, the audio-video conferencing device includes a base with a number input keyboard thereon. Through this number input keyboard, users can input the conference number in a timely manner to connect to the conference. In addition, the audio-video conferencing device further includes a display screen separated from the base. The display screen has a camera. The camera captures the video of the participants, and the captured video is displayed on the display screen. The microphone array is encapsulated inside the base. According to the voices of the speakers collected by different microphones in the microphone array, the sound source position can be determined, so that the display screen rotates towards the sound source to better collect the conference sound.

[0003] Generally, the base is placed horizontally. Multiple microphones in the microphone array are located in a horizontal plane and cannot sense the change in the vertical direction of the sound source. If the speaker stands up while speaking, these microphones cannot sense the vertical component of the sound and cannot recognize that the speaker has stood up, and cannot make the camera track the face of the standing person. Summary of the Invention

[0004] In view of this, the present disclosure aims to be able to recognize the change in the vertical direction of the sound source and achieve accurate sound source localization.

[0005] To achieve this purpose, according to one aspect of the present disclosure, the present disclosure provides an audio-video conferencing device, including a base extending along a first plane, a display screen extending from the base along a second plane, a microphone array, and a sound source localization unit, wherein at least part of the microphones in the microphone array are encapsulated in the base, at least part of the microphones are encapsulated in the display screen, and the sound source localization unit determines the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array.

[0006] Optionally, there is an axis between the base and the display screen, and the axis rotates according to the determined three-dimensional spatial position of the sound source, so that the display screen faces the sound source.

[0007] Optionally, the microphone array is in the shape of a tetrahedron, and the microphone array has 4 microphones, which are respectively located at the four vertices of the tetrahedron.

[0008] Optionally, the microphone array includes multiple tetrahedrons.

[0009] According to one aspect of the present disclosure, a terminal device is provided, including a first part extending along a first plane, a second part extending from the first part along a second plane, a microphone array, and a sound source localization unit, wherein at least some microphones in the microphone array are encapsulated in the first part, at least some microphones are encapsulated in the second part, and the sound source localization unit determines the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array.

[0010] Optionally, there is an axis between the first part and the second part, and the axis rotates according to the determined three-dimensional spatial position of the sound source, so that the second part faces the sound source.

[0011] Optionally, the terminal device is a smart TV, the first part is a base, and the second part is a display screen.

[0012] Optionally, the terminal device is a smart speaker, the first part is an operating table, and the second part is a speaker box.

[0013] Optionally, the terminal device is a conversation robot, the first part is a robot foot, and the second part is a robot face.

[0014] According to one aspect of the present disclosure, a sound source localization method is provided, including:

[0015] Obtaining an initial sound source position;

[0016] Predicting the sound source position in the next period according to the initial sound source position as the target sound source position;

[0017] Substituting the target sound source position into the delay error polynomial corresponding to multiple candidate delays of the microphone pairs in the microphone array, and taking the candidate delay with the smallest value of the delay error polynomial among the multiple candidate delays as the target delay;

[0018] Constructing a delay equation of the microphone array according to the target delay of the microphone pairs in the microphone array, and determining the sound source position that minimizes the delay equation. If the distance from the determined sound source position to the target sound source position is not within a predetermined distance threshold, substituting the determined sound source position into the delay error polynomial again until the distance from the determined sound source position to the target sound source position is within the predetermined distance threshold.

[0019] Optionally, the obtaining of the initial sound source position includes:

[0020] Constructing a generalized cross-correlation function of the delay for the microphone pairs of the microphone array, and determining the delay that maximizes the generalized cross-correlation function;

[0021] For the microphone pair, a time delay expression based on the initial sound source position and the positions of the two microphones of the microphone pair is constructed such that the time delay expression is equal to the determined time delay that maximizes the generalized cross-correlation function, and the initial sound source position is obtained by solving.

[0022] Optionally, the generalized cross-correlation function is the integral in the frequency domain of the product of the frequency weighting function and the cross-spectrum of the sound signals respectively received by the microphone pair after phase-shifting according to the time delay.

[0023] Optionally, the time delay expression is the difference between the distances between the initial sound source position and the positions of the two microphones of the microphone pair divided by the speed of sound.

[0024] Optionally, the next period is the period during which the camera of the audio-visual conferencing device where the microphone array is located captures the next frame.

[0025] Optionally, the prediction of the sound source position in the next period according to the initial sound source position is performed according to the Kalman filtering method.

[0026] Optionally, the multiple candidate time delays are the time delays corresponding to the k-th largest peak of the generalized cross-correlation function of the time delays constructed for the microphone pairs of the microphone array, where k = 1,..., N, and N is a positive integer greater than or equal to 2.

[0027] Optionally, the time delay error term formula is the square of the difference between the time delay corresponding to the k-th largest peak of the generalized cross-correlation function and the time delay expression based on the initial sound source position and the positions of the two microphones of the microphone pair, multiplied by the gain adjustment coefficient of the microphone pair and divided by the k-th largest peak.

[0028] Optionally, the time delay equation is equal to the sum of the target time delays of the microphone pairs in the microphone array plus the product of the regularization term and the regularization term coefficient, and the regularization term is equal to the square of the difference between the sound source position to be determined and the target sound source position.

[0029] According to one aspect of the present disclosure, there is provided a computer storage medium including computer-executable code that, when executed by a processor, implements the method as described above.

[0030] Different from the prior art where all the microphones of the microphone array are located on a single plane, in the microphone array of the embodiments of the present disclosure, a part of the microphones are located on the first plane of the base, and another part are located on the second plane of the display screen. In this way, not only can the change in the horizontal direction of the sound source be sensed, but also the change in the vertical direction of the sound source can be sensed, realizing three-dimensional precise positioning of the sound source.

[0031] The general method for sound source localization in the prior art generally requires obtaining the time delay of the sound source reaching the microphone array first, and then using this time delay in combination with the array structure to determine the sound source position. The two processes are isolated. In the embodiments of the present disclosure, the two isolated processes are unified. First, the sound source position is determined according to the time delay, then the error in the time delay estimation is eliminated according to the deviation of the sound source position, and then the sound source position is re-determined according to the time delay after eliminating the error, and this is repeated continuously to achieve the purpose of accurate sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] By referring to the following drawings for the description of the embodiments of the present disclosure, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0033] Figure 1 is a perspective view of the structure of an existing audio-video conferencing device;

[0034] Figure 2 is a perspective view of the structure of an audio-video conferencing device according to an embodiment of the present disclosure;

[0035] Figure 3 is a block diagram of a terminal device according to an embodiment of the present disclosure;

[0036] Figure 4 is an electrical structure diagram of an audio-video conferencing device according to an embodiment of the present disclosure;

[0037] Figure 5 is a flowchart of a sound source localization method according to an embodiment of the present disclosure;

[0038] Figure 6 is a graph showing the variation of the generalized cross-correlation function of the time delay with the time delay;

[0039] Figure 7A is a localization trajectory diagram of the sound source localization method in the prior art;

[0040] Figure 7B is a localization trajectory diagram of the sound source localization method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following is a description of the present disclosure based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. To avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. Additionally, the drawings are not necessarily drawn to scale.

[0042] The following terms are used in this article.

[0043] Microphone array: An array formed by multiple microphones. Each microphone separately collects the sound signals emitted by the sound source. Since the angles of each microphone relative to the sound source are different, the directions of the detected sound signals are also different. Therefore, the position of the sound source can be uniquely determined based on the directions of the sound signals detected by multiple microphones.

[0044] Tetrahedron: Generally refers to a triangular pyramid. A triangular pyramid is a type of pyramid, a geometric body composed of four triangles. The base of the tetrahedron is a triangle, which has three vertices. The three sides of the triangle respectively form a triangle with the fourth vertex of the tetrahedron.

[0045] Microphone pair: A pair formed between any two microphones in the microphone array.

[0046] Time delay: The difference between the time when one microphone in the microphone pair receives the sound signal emitted by the same sound source and the time when the other microphone receives the sound signal emitted by the same sound source is called the time delay.

[0047] Candidate time delay: The time delay as a candidate.

[0048] Time delay error term formula: It refers to the arithmetic expression obtained by multiplying the difference between the time delays calculated in two different ways by the corresponding adjustment coefficient. In the two different ways of calculating the time delay in the embodiments of the present disclosure, one is based on the generalized cross-correlation function of the time delay. When the generalized cross-correlation function takes a certain value, the corresponding time delay can be deduced. The other is determined according to the distance difference divided by the speed of sound. Among them, the difference between the distances from the sound source position to the two microphones of the microphone pair, divided by the speed of sound, gives the difference in the time when the sound signal emitted by the sound source propagates to one microphone and then to the other microphone. In the embodiments of the present disclosure, for multiple candidate time delays, the above time delay error term formula is calculated respectively, and the values of the obtained time delay error term formulas are different. Since the time delay obtained by dividing the difference between the distances from the sound source position to the two microphones of the microphone pair by the speed of sound is the true time delay, the smaller the value of the above time delay error term formula, the closer the candidate time delay is to the true time delay. Therefore, by determining the candidate time delay with the smallest value of the time delay error term formula among multiple candidate time delays as the target time delay, the accuracy of determining the time delay can be improved.

[0049] Generalized cross-correlation (GCC) function: A method for calculating the time delay difference between signals using the correlation between signals. The generalized cross-correlation algorithm considers the free-field model and the microphone pair to calculate the time delay. It represents the cross-correlation of the sound signals of the same sound source received by the two microphones of the microphone pair as a function of the time delay. The generalized cross-correlation function can be the integral in the frequency domain of the product of the frequency weighting function and the cross-spectrum of the sound signals respectively received by the microphone pair after phase-shifting according to the time delay of the sound signals respectively received by the microphone pair, which will be described in detail later.

[0050] Frequency-domain weighting function: Before calculating the cross-correlation of two signals, the generalized cross-correlation function filters the signals, which is equivalent to multiplying by a weighting function in the frequency domain. This weighting function is called the frequency-domain weighting function.

[0051] Target time delay: The time delay determined from multiple candidate time delays according to the time delay error term formula and used as the basis for determining the sound source position.

[0052] Time delay equation: There are different microphone pairs in the microphone array, and each microphone pair has a target time delay. The overall equation of the microphone array constructed based on the target time delays of different microphone pairs is the time delay equation. The sound source position can be determined according to this time delay equation. For example, the sound source position when the time delay equation is minimized is used as the determined sound source position.

[0053] Time delay expression based on the sound source position and the positions of the two microphones in the microphone pair: That is, the expression obtained by dividing the difference between the distances from the sound source position to the positions of the two microphones by the speed of sound. It will be described in detail later.

[0054] Kalman filter: An algorithm that uses the linear system state equation to estimate the system state through the system input and output observation data. In the embodiments of the present disclosure, the Kalman filtering method is used to predict the sound source position in the next cycle from the sound source position in the current cycle.

[0055] Gain adjustment coefficient: A coefficient used to adjust the time delay corresponding to the k-th largest peak in the time delay error term formula, which is the coefficient of the difference between the time delays calculated by two methods. Different k values correspond to different gain adjustment coefficients.

[0056] Regular term: A constraint term. That is, when usually using an empirical function to find the minimum value to solve unknown variables, in order to reflect the actual constraints on the empirical function, a constraint term is usually added to the empirical function. After adding the constraint term and solving it by finding the minimum value, the accuracy of the solution result can be improved.

[0057] Gauss-Newton iterative algorithm: An iterative method for least squares to find regression parameters in a nonlinear regression model.

[0058] Figure 1Schematic diagram of an existing audio-visual conferencing device. The audio-visual conferencing device includes a base 20 and a display screen 10. The base 20 and the display screen 10 are separated. The base 20 has a number input keyboard 21 for inputting a conference number to connect to a telephone conference. By using this number input keyboard, a user can input a conference number or the respective numbers of participants to connect to an audio-visual conference. The display screen 10 includes a screen body 11 for displaying the images of participants beside the audio-visual conferencing device and the images of remote participants of other parties connected to the audio-visual conference. The screen body 11 can be a touch display screen body. Staff members beside the audio-visual conferencing device can click, switch, set the main speaker or the main speaking venue, etc. among the images displayed on the screen body 11 by touching. The display screen 10 may also include a camera 12 for photographing the person speaking beside the audio-visual conferencing device.

[0059] Inside the base 20, a microphone array 22 is encapsulated, as Figure 1 shown by the dashed line. The microphone array 22 includes a plurality of microphones 23 located in the same plane, which are respectively used for collecting the sound signals spoken by the person speaking (sound source). As Figure 1 shown in [reference], 3 microphones 23 are shown, and the angle between adjacent microphones 23 is 120°. Since the angles of the 3 microphones 23 from the sound source are different, the directions of the collected sound signals are also different. In this way, the position of the sound source can be determined according to the directions of the sound signals collected by the 3 microphones 23 respectively, so that the camera 12 can rotate following the sound source, making the camera 12 always face the sound source, achieving the purpose of clearly photographing the speaker in an audio-visual conference.

[0060] Generally, the base 20 is placed horizontally and the display screen 10 is placed vertically. The plurality of microphones 23 in the microphone array 22 are located in a horizontal plane and cannot sense the change in the vertical direction of the sound source. If the person speaking stands up while speaking, these microphones 23 cannot sense the vertical component of the sound and cannot recognize that the person has stood up, and cannot make the camera 12 track the face of the person who has stood up.

[0061] To overcome the problem that all the microphones in the microphone array of the existing technology are located in one plane and thus cannot sense the vertical movement of the sound source, as Figure 2As shown, in the embodiments of the present disclosure, the display screen 10 and the base 20 are no longer separate, but are joined together. However, the display screen 10 and the base 20 extend in two different directions. The base 20 extends on a first plane, and the first level may be horizontal. The display screen 10 extends from the base 20 along a second plane, and the second plane may be vertical, that is, perpendicular to the first plane, but may also intersect the first plane at any acute or obtuse angle. Since the display screen 10 and the base 20 are joined and located on different planes, a microphone array 22 can be designed. At least some of the microphones 23 in the microphone array 22 are encapsulated in the base 20 and are located on the first plane of the base 20, and at least some of the microphones 23 are encapsulated in the display screen 10 and are located on the second plane of the display screen 10. In this way, not only can the change in the horizontal direction of the sound source be sensed, but also the change in the vertical direction of the sound source can be sensed, realizing three-dimensional precise positioning of the sound source.

[0062] Figure 2 A microphone array in the shape of a tetrahedron is shown, as indicated by the dashed line. The microphone array 22 has 4 microphones 23, which are respectively located at the four vertices of the tetrahedron. Among them, 3 microphones are located at the three vertices of the bottom surface of the tetrahedron, and they are located in the first plane of the base 20; 1 microphone is located in the second plane of the display screen 10 and is located at the top vertex of the tetrahedron. Although Figure 2 Taking the example that 3 microphones are located in the first plane of the base 20 and 1 microphone is located in the second plane of the display screen 10, those skilled in the art understand that other numbers of microphones can be set on the first plane, such as 2 or 4; other numbers of microphones can be set on the second plane, such as 2, 3, or 4. In addition, the microphone array 22 does not need to be in the shape of a tetrahedron, and it can be in other shapes, as long as a certain number of microphones are respectively provided on the first plane and the second plane. However, the tetrahedron shape is relatively microphone-saving, and relatively accurate sound source positioning can be achieved while saving the number of microphones.

[0063] As Figure 2 shown, the base 20 and the display screen 10 are joined by a shaft 25. The microphone array 22 can be set as 2 tetrahedrons, and the 2 tetrahedrons are respectively arranged on both ends of the shaft 25. The advantage of doing this is to improve the accuracy of sound source positioning in different directions. It should be understood that the microphone array 22 can also be set as other numbers of tetrahedrons to improve the accuracy of sound source positioning in different directions.

[0064] Figure 4 is the internal structure diagram of the audio-visual conferencing device. In addition to Figure 2The display screen 11, camera 12, and microphone array 22 shown in the figure. Inside the audio and video conferencing device, there are also a sound source localization unit 24, a motor driver 27, a motor 26, a processor 28, a transceiver 29, and a memory 30. The sound source localization unit 24 determines the three-dimensional spatial position of the sound source based on the sound signals collected by the microphone array 22. Since some of the microphones 23 in the microphone array 22 are located on the first plane of the base 20 and some are located on the second plane of the display screen 10, in this way, it can not only sense the changes in the horizontal direction of the sound source but also sense the changes in the vertical direction of the sound source. Compared with the prior art that can only achieve two-dimensional localization of the sound source, it can achieve three-dimensional precise localization of the sound source. When the sound source localization unit 24 determines the position of the sound source, it sends a drive signal according to the position of the sound source, and drives the motor 26 to drive the shaft 25 to rotate through the motor driver 27. The shaft can rotate in a 360° direction. Let the shaft 25 rotate according to the determined three-dimensional spatial position of the sound source, so that the camera 12 faces the sound source.

[0065] The image of the person speaking captured by the camera 12 is transmitted to the processor 28. The sound signals collected by the microphone array 22 are transmitted to the sound source localization unit 24 for localization on the one hand, and also transmitted to the processor 28 for sound processing on the other hand. After the processor 28 processes the image signal and the sound signal, it is sent to the display screen 11 for display on the one hand, and transmitted to the audio and video conferencing devices in other participating conference venues through the transceiver 29 for display and playback on the audio and video conference venue devices in other participating conference venues. In addition, the processor 28 also receives the videos of the participants in other participating conference venues transmitted by the audio and video conferencing devices in other participating conference venues from the transceiver 29, and displays and plays them on the display screen 11 simultaneously with the audio and video of the person speaking in this venue. When displaying on the display screen 11, the display screen 11 can be divided into different sub-pictures, and the videos of people in different venues occupy different sub-pictures. Which venue's sound is played can be specified by the person next to the audio and video conferencing device on the touch display screen 11.

[0066] The implementation idea described above, which is to locate the microphones 23 of the microphone array 22 of the embodiments of the present disclosure on two different planes respectively to not only sense the change in the horizontal direction of the sound source but also sense the change in the vertical direction of the sound source, thereby achieving three-dimensional precise positioning of the sound source, is applied to an embodiment of an audio-video conferencing device. In fact, this implementation idea can be applied not only to audio-video conferencing devices but also to any terminal device that needs to sense the position of the sound source during human-computer interaction. The terminal device can be a smart home device, such as a smart TV or a smart speaker, or a smart human-computer interaction device, such as a chatbot or a voice search system. A smart TV needs to sense the changes in the horizontal and vertical directions of the person speaking to precisely locate the person, so that the TV screen always faces the person, improving the viewing experience. A smart speaker needs to sense the changes in the horizontal and vertical directions of the person speaking to precisely locate the person, making the front of the speaker face the person, improving the listening experience. A chatbot needs to sense the changes in the horizontal and vertical directions of the person speaking to precisely locate the person, making the face of the robot always face the human face, and even the robot can move along with the person when the person moves, improving the conversation experience of the person.

[0067] Figure 3 The block diagram of the general terminal device applied in the embodiments of the present disclosure is shown. It includes a first part 20' extending along a first plane, a second part 10' extending from the first part 20' along a second plane, a microphone array 22, and a sound source positioning unit (not shown). At least some of the microphones in the microphone array 22 are encapsulated in the first part 20', and at least some of the microphones are encapsulated in the second part 10'. In Figure 3 it is shown that microphones S2 - S4 are located in the first part 20', and microphone S1 is located in the second part 10'. However, in fact, there may be other numbers of microphones in the first part 20', and there may be other numbers of microphones in the second part 10'. As long as there are a certain number of microphones on the two planes respectively, the purpose of sensing the changes in the horizontal and vertical directions of the person speaking to precisely locate the person can be achieved. The sound source positioning unit is not shown. It can be inside the first part 20', or inside the second part 10', or there can be a part in each of them. Its function is to determine the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array.

[0068] When the terminal device is a smart TV, the first part 20' is the base of the TV, and the second part 10' is the display screen. There is an axis between the base and the display screen, and the axis rotates according to the determined three-dimensional spatial position of the sound source, so that the display screen faces the sound source. In this way, no matter in which direction the person watching the TV moves up, down, left, or right, the display screen moves with the user, making the face face the display screen directly, improving the user's TV viewing experience.

[0069] When the terminal device is a smart speaker, the first part 20' is an operation console, on which a smart speaker switch and various function keys can be arranged, and the second part 10' is a speaker box. There is an axis between the operation console and the speaker box, and the axis rotates according to the three-dimensional spatial position of the determined sound source, so that the speaker box faces the sound source. In this way, no matter in which direction the person listening to the speaker moves up, down, left or right, the speaker of the speaker box will move with the user, making the speaker face the user directly, and improving the clarity of the sound heard by the user.

[0070] When the terminal device is a conversation robot, the first part 20' is the robot's feet, and the second part 10' is the robot's face. The neck of the robot can rotate 360°. In this way, the neck of the robot can rotate according to the three-dimensional spatial position of the determined sound source, so that the robot's face faces the sound source, that is, faces the face of the person speaking, realizing face-to-face conversation and improving the user's conversation experience. In this embodiment, as the sound source moves horizontally in the direction determined by the sound source positioning unit, the robot's feet can be controlled to walk, so that the robot follows the person speaking and always speaks facing the face, improving the conversation experience.

[0071] Those skilled in the art should understand that the terminal device can be other devices that require human-machine voice interaction. As long as the device requires human-machine voice interaction, it needs to sense not only the change in the horizontal direction of the sound source, but also the change in the vertical direction of the sound source to achieve three-dimensional precise positioning of the sound source. The first part 20' and the second part 10' can be deployed in the terminal device according to the actual structure of the terminal device.

[0072] Figure 5 The flowchart of the sound source positioning method according to an embodiment of the present disclosure is shown. This method is executed by the sound source positioning unit 24, and it can be applied not only to Figure 2 the video and audio conferencing device shown, but also to Figure 3 the more general terminal device shown. As Figure 5 shown, this method includes:

[0073] Step 110, obtain the initial sound source position;

[0074] Step 120, predict the sound source position in the next cycle according to the initial sound source position as the target sound source position;

[0075] Step 130, substitute the target sound source position into the delay error term formula corresponding to multiple candidate delays of the microphone pairs in the microphone array, and use the candidate delay with the smallest value of the delay error term formula among the multiple candidate delays as the target delay;

[0076] Step 140: Construct a time-delay equation for the microphone array based on the target time delay of the microphone pairs in the microphone array, and determine the sound source position that minimizes the time-delay equation. If the distance from the determined sound source position to the target sound source position is not within a predetermined distance threshold, substitute the determined sound source position for the target sound source position and re-enter it into the time-delay error term formula until the distance between the determined sound source position and the target sound source position is within the predetermined distance threshold.

[0077] The following describes these steps in detail.

[0078] In one embodiment, step 110 may include: constructing a generalized cross-correlation function of time delay for the microphone pairs of the microphone array, and determining the time delay that maximizes the generalized cross-correlation function; for this microphone pair, constructing a time-delay expression based on the initial sound source position and the positions of the two microphones of this microphone pair, such that the time-delay expression is equal to the determined time delay that maximizes the generalized cross-correlation function, and solving to obtain the initial sound source position.

[0079] Any two microphones 23 in the microphone array 22 form a microphone pair. The sound signals received by the two microphones of this microphone pair from the same sound source are respectively:

[0080]

[0081] Among them, y1(t) is the sound signal received by the first microphone, s(t) is the sound signal emitted by the sound source, τ is the time taken for the sound signal emitted by the sound source to reach the first microphone, α1 is the attenuation coefficient of the sound signal emitted by the sound source when reaching the first microphone, and v1(t) is the noise signal interfering with the first microphone; y2(t) is the sound signal received by the second microphone, s(t) is the sound signal emitted by the sound source, τ is the time taken for the sound signal emitted by the sound source to reach the first microphone, φ 12 represents the difference in the time taken for the sound signal emitted by the sound source to reach the first microphone and the second microphone, that is, the time delay, α2 is the attenuation coefficient of the sound signal emitted by the sound source when reaching the second microphone, and v2(t) is the noise signal interfering with the second microphone.

[0082] The generalized cross-correlation (GCC) function is a method for calculating the time-delay difference of signals using the correlation between signals. The generalized cross-correlation algorithm considers the free-field model and the microphone pair for time-delay calculation, and it represents the cross-correlation of the sound signals of the same sound source received by the two microphones of the microphone pair as a function of time delay. The generalized cross-correlation function may be the integral in the frequency domain of the product of the frequency-weighting function and the cross-spectrum of the sound signals received by this microphone pair, phase-shifted according to the time delay of the sound signals received by this microphone pair, as shown in Equation 3.

[0083] (Formula 3)

[0084] r y1y2 R(p) is the generalized cross - correlation function. W(f) represents the frequency - domain weighting function. The generalized cross - correlation function filters the signals before calculating the cross - correlation of two signals. In terms of the frequency - domain representation, it is equivalent to multiplying by a weighting function in the frequency domain, and this weighting function is called the frequency - domain weighting function. p is the time delay for the sound signal of the sound source to reach the first microphone and the second microphone of the microphone pair. The phase shift is carried out according to this time delay. It is expressed as the cross - spectral density, that is

[0085]

[0086] Y1(f) represents the signal after the Fourier transform of y1(t), Y2*(f) represents the conjugate of the signal after the Fourier transform of y2(t), and E represents the expectation.

[0087] If is set to 1, Formula 3 degenerates into the cross - correlation algorithm. In the most commonly used PHAT (Phase Transform) algorithm, is set to

[0088]

[0089] Then Formula 3 can be transformed into

[0090]

[0091] The above - mentioned Formula 6 can be expressed as Figure 6 as shown in the curve graph. Figure 6 In it, the abscissa represents the time delay p, and the ordinate represents the generalized cross - correlation function r y1y2 (p) of the time delay. This generalized cross - correlation function has several peaks, and one of them is the largest peak. The time delay corresponding to this largest peak can be found in Figure 6 That is to say, when the time delay is this value, the generalized cross - correlation function is the largest, and the sound signals reaching the two microphones of the microphone pair are the most correlated. After obtaining this time delay, the initial sound - source position is solved by using the time - delay expression based on the initial sound - source position and the positions of the two microphones of this microphone pair. The time - delay expression based on the sound - source position and the positions of the two microphones of this microphone pair can be an expression obtained by dividing the difference between the distances from the sound - source position to the positions of the two microphones by the speed of sound to get the time delay.

[0092] Still taking Figure 2Taking the tetrahedral microphone array shown as an example, where microphone S1 is located on the second plane of the display screen 10, and microphones S2, S3, and S4 are located on the first plane of the base 20. Microphone S1 and microphone S2 form a microphone pair, microphone S1 and microphone S3 form a microphone pair, and microphone S1 and microphone S4 form a microphone pair. The microphone pairs formed inside microphones S2, S3, and S4 are not considered because the microphone pairs formed between them still cannot take into account the position change of the sound source in the vertical direction. Assuming the sound source position is Then the time delay can be calculated by the following formula

[0093] (Formula 7)

[0094] The above Formula 7 is the time delay expression constructed based on the initial sound source position and the positions of the two microphones in the microphone pair, where j = {2, 3, 4}, c represents the speed of sound, represents the two-norm, and represents the distance between two position coordinates. represents the distance between the sound source position coordinate and the position coordinate of microphone S1, represents the distance between the sound source position coordinate and the position coordinate of microphone S j The distance difference divided by the speed of sound gives the time difference for the sound of the sound source to travel to microphone S1 and to travel to microphone S j That is, the time delay. In the traditional algorithm, the corresponding time delay has been obtained according to the largest peak in Figure 6 which is equal to Substituting it into Formula 7, we get That is, the initial sound source position.

[0095] In step 120, according to the initial sound source position, predict the sound source position in the next cycle as the target sound source position.

[0096] The next cycle is the period when the camera 12 of the audio-visual conferencing device where the microphone array 22 is located captures the next frame. The camera 12 has to rotate with the face (sound source) of the person speaking. Therefore, the period for determining the sound source position can be the same as the period of rotation of the camera 12. In this way, in each cycle, determine the real-time position of the sound source, and rotate the camera 12 to capture the next frame according to this real-time position.

[0097] According to the initial sound source position, predicting the sound source position in the next cycle can be performed according to the Kalman filtering method. Since the Kalman filtering method is a known method, it will not be elaborated here.

[0098] In step 130, substitute the target sound source position into the time delay error term formula corresponding to multiple candidate time delays for the microphone pairs in the microphone array, and use the candidate time delay with the smallest value of the time delay error term formula among the multiple candidate time delays as the target time delay.

[0099] The candidate time delay is the time delay as a candidate. In an embodiment of the present disclosure, the candidate time delay is the time delay corresponding to the k-th largest peak of the generalized cross-correlation function of the time delays constructed for the microphone pairs of the microphone array, where k = 1,..., N, and N is a positive integer greater than or equal to 2. As Figure 6 shown, in the curve of the generalized cross-correlation function, in addition to the maximum peak, one peak on each of the left and right sides of the maximum peak is also relatively large. The time delays corresponding to these 3 peaks can all be used as candidate time delays. Different from the prior art where only the time delay of the maximum peak is used as the target time delay to calculate the sound source position accordingly, in the embodiment of the present disclosure, it is considered that other peaks that are not the maximum peak may be more suitable as the target time delay, and the sound source position calculated accordingly may be more accurate. Therefore, by first determining multiple candidate time delays, selecting one of them as the target time delay, using the target time delay to determine the sound source position, then using the determined sound source position to re-determine the target time delay among the multiple candidate time delays, and then determining the sound source position accordingly, through such a process of repeated multiple times, the accuracy of sound source localization is improved.

[0100] The time delay error term formula refers to the arithmetic expression obtained by multiplying the square of the difference between the time delay determined according to the above-mentioned generalized cross-correlation function method and the time delay determined according to the time delay expression based on the initial sound source position and the positions of the two microphones of the microphone pair by the corresponding adjustment coefficient. The former method can inversely deduce the corresponding time delay when the generalized cross-correlation function takes a certain value according to the generalized cross-correlation function of the time delay. The latter is determined according to the distance difference divided by the speed of sound. Among them, dividing the difference between the distances from the sound source position to the two microphones of the microphone pair by the speed of sound gives the difference in the time for the sound signal emitted by the sound source to propagate to one microphone and then to the other microphone. The above-mentioned time delay error term formula represents the influence of different candidate time delays on the time delay error. The smaller the value of the time delay error term formula, the smaller the determined time delay error. In the embodiment of the present disclosure, for multiple candidate time delays, the above-mentioned time delay error term formula is calculated respectively, and the obtained values of the time delay error term formula are different. Since the time delay obtained by dividing the difference between the distances from the sound source position to the two microphones of the microphone pair by the speed of sound is the true time delay, the smaller the value of the above-mentioned time delay error term formula, the closer the candidate time delay is to the true time delay. Therefore, by determining the candidate time delay with the smallest value of the time delay error term formula among the multiple candidate time delays as the target time delay, the accuracy of determining the time delay can be improved. That is, as formula 8 below:

[0101]

[0102] Formula 8

[0103] where represents the time delay determined according to the time delay expression of the positions of the two microphones of the microphone pair. Since the target sound source position is known, each microphone S1 and Sj The position is also known. Here is known. That is, the time delay determined by the above-mentioned generalized cross-correlation function method, namely, multiple candidate time delays corresponding to the first several large peaks. m represents the number of microphone pairs. The more microphones are used, the more accurate the positioning is. The minimum number of microphone pairs is three pairs because only in this way can a three-dimensional plane be just formed. However, in actual use, the number of microphone pairs is usually greater than three pairs to ensure a certain redundancy. is the adjustment coefficient, which is used to improve the priority of the time delays corresponding to the first several largest peaks, and it is equal to the gain adjustment coefficient corresponding to the microphone pair j , divided by the k-th largest peak of the microphone pair j . In this way, among multiple candidate time delays, the candidate time delay with the smallest value of the time delay error term can be used as the target time delay.

[0104] In step 140, according to the target time delays of the microphone pairs in the microphone array, construct the time delay equation of the microphone array, and determine the sound source position that minimizes the time delay equation. If the distance from the determined sound source position to the target sound source position is not within the predetermined distance threshold, substitute the determined sound source position for the target sound source position and re-substitute it into the time delay error term until the distance from the determined sound source position to the target sound source position is within the predetermined distance threshold.

[0105] As Figure 5 shown, step 140 may include: step 141, according to the target time delays of the microphone pairs in the microphone array, construct the time delay equation of the microphone array, and determine the sound source position that minimizes the time delay equation; step 142, determine whether the distance from the sound source position that minimizes the time delay equation to the target sound source position is within the predetermined distance threshold; if so, return to step 120 to re-predict the sound source position in the next cycle as the target sound source position; if not, execute step 143, that is, substitute the sound source position that minimizes the time delay equation for the target sound source position, return to step 130, re-substitute the target sound source position into the time delay error term, and iterate repeatedly until the distance from the determined sound source position to the target sound source position is within the predetermined distance threshold.

[0106] There are different microphone pairs in the microphone array, and each microphone pair has a target time delay. The overall formula of the microphone array constructed according to the target time delays of different microphone pairs is the time delay equation. According to this time delay equation, the sound source position can be determined. For example, the sound source position when the time delay equation is minimized is taken as the determined sound source position. The time delay equation can be as shown in formula 9:

[0107]

[0108] In Equation 9, the target time delays calculated for m pairs of microphones are summed to obtain , which can be used as the time delay equation. The sound source position at a certain time point can be directly optimized from Equation 9. However, doing so isolates the calculation of a certain position. Therefore, a regularization term is added after Equation 9 to associate each calculation with the sound source position calculated in the previous round and form a certain constraint, that is

[0109] (Equation 10)

[0110] Here, v represents a regularization term coefficient, is the regularization term and can be expressed as

[0111] (Equation 11)

[0112] Thus, in one embodiment, the time delay equation is equal to the sum of the target time delays of the microphone pairs in the microphone array plus the product of the regularization term multiplied by the regularization term coefficient , where the regularization term is equal to the square of the difference between the sound source position to be determined and the target sound source position .

[0113] When the target time delay is determined, is fixed, the target sound source position estimated by Kalman filtering is determined, and only is uncertain. Therefore, solving Equation 10 becomes a problem of determining when the value of is the smallest. It can be solved by Gauss-Newton iteration, which will not be elaborated here.

[0114] After determining the sound source position that minimizes the time delay equation, in step 142, it is determined whether the distance between the sound source position that minimizes the time delay equation and the target sound source position is within a predetermined distance threshold. If not, step 143 is executed, that is, the sound source position that minimizes the time delay equation is used to replace the target sound source position, and step 130 is returned. The target sound source position is re-substituted into the time delay error term formula and iterated repeatedly until the distance between the determined sound source position and the target sound source position is within the predetermined distance threshold. After consistency, it is considered that the accurate sound source position in the next cycle has been determined. The axis can drive the display screen 11 to rotate according to this sound source position during the next frame shooting, so that the camera 12 faces the sound source. Then, return to step 120, predict the sound source position in the next cycle as the target sound source position, and repeat steps 130-140 to obtain the accurate sound source position in the next cycle, and so on.

[0115] The present disclosure also provides a computer storage medium including computer executable code. When the computer executable code is executed by the processor 28, the following is achieved: Figure 5 The method shown.

[0116] The sound source localization method of the prior art uses the time delay corresponding to the maximum peak of the generalized cross-correlation function to be equal to the time delay obtained based on the sound source and each microphone position to solve the sound source position, and the simultaneous equations can be solved directly, but this approach is computationally intensive, difficult to simplify, and does not take into account error information, which easily leads to unsolvable situations. The disclosed embodiment unifies the two processes of determining the time delay and obtaining the sound source position using the determined time delay, effectively eliminating the error in the time delay estimation while accurately positioning in three dimensions, and on this basis, regards the entire sound source localization process as a whole, and adds a regularization term to improve the estimation accuracy of positioning.

[0117] Figure 7A It is a positioning trajectory diagram of the sound source localization method in the prior art. Figure 7B 23 shows the positions of two microphones in a microphone pair. In the actual test, the speaker moves around the field for one and a half circles. Figure 7A and 7B The lines respectively reflect the sound source positions determined at each time point in the prior art and the embodiment of the present disclosure connected in the order of time points. Figure 7A There are a lot of noise points, and the trajectory is seriously distorted. Figure 7B The middle trajectory is relatively smooth and better reflects the trajectory of the sound source movement.

[0118] The commercial value of the present disclosure

[0119] In the disclosed embodiment, a new type of deployment scheme for audio-visual conferencing equipment and terminals is proposed, and based on this scheme, a set of effective sound source localization algorithms is proposed, which associates the localization problem with the delay problem, uses the position of the previous cycle to assist the application of the generalized cross-correlation algorithm of the current cycle, and uses the regularization term to associate the information of the previous and next positions, so as to regard the sound source position estimation problem as a coherent whole rather than an isolated process, which has high application value. According to experiments, the disclosed embodiment improves the accuracy of sound source localization by 50%, and has extremely high commercial prospects.

[0120] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0121] It should be understood that the above description relates to specific embodiments of the present specification. Other embodiments are within the scope of the claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0122] It should be understood that an element described herein in the singular or shown only in one of the figures is not intended to limit the quantity of that element to one. Additionally, modules or elements described or shown herein as separate may be combined into a single module or element, and a module or element described or shown herein as a single one may be split into multiple modules or elements.

[0123] It should also be understood that the terms and expressions used herein are for descriptive purposes only, and one or more embodiments of the present specification should not be limited to these terms and expressions. Using these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.

Claims

1. An audio-visual conferencing device, comprising a base extending along a first plane, a display screen protruding from the base along a second plane, a microphone array, and a sound source localization unit, wherein, At least some of the microphones in the microphone array are encapsulated in the base, and at least some of the microphones are encapsulated in the display screen. The sound source positioning unit determines the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array; Among them, the sound source positioning unit is used for: Obtain the initial sound source position; According to the initial sound source position, predict the sound source position in the next cycle as the target sound source position; Substitute the target sound source position into the delay error term formula corresponding to multiple candidate delays of the microphone pairs in the microphone array, and take the candidate delay with the smallest value of the delay error term formula among the multiple candidate delays as the target delay. Among them, the multiple candidate delays are the delays corresponding to the k-th largest peak of the generalized cross-correlation function of the delays constructed for the microphone pairs of the microphone array, where k = 1,..., N, and N is a positive integer greater than or equal to 2; According to the target delays of the microphone pairs in the microphone array, construct the delay equation of the microphone array, and determine the sound source position that minimizes the delay equation. If the distance between the sound source position that minimizes the delay equation and the target sound source position is not within the predetermined distance threshold, substitute the determined sound source position to replace the target sound source position and re-substitute it into the delay error term formula until the distance between the determined sound source position and the target sound source position is within the predetermined distance threshold; Among them, the delay equation is equal to the sum of the target delays of the microphone pairs in the microphone array plus the product of the regularization term and the regularization coefficient.

2. The audio-visual conferencing device according to claim 1, wherein, There is an axis between the base and the display screen, and the axis rotates according to the determined three-dimensional spatial position of the sound source so that the display screen faces the sound source.

3. The audio-visual conferencing device according to claim 1, wherein, The microphone array is in the shape of a tetrahedron, and the microphone array has 4 microphones, which are respectively located at the four vertices of the tetrahedron.

4. The audio-visual conferencing device according to claim 3, wherein, The microphone array includes multiple tetrahedrons.

5. A terminal device, comprising a first part extending along a first plane, a second part extending from the first part along a second plane, a microphone array, and a sound source localization unit, wherein, At least some of the microphones in the microphone array are encapsulated in the first part, and at least some of the microphones are encapsulated in the second part. The sound source positioning unit determines the three-dimensional spatial position of the sound source according to the sound signals collected by the microphone array; Among them, the sound source positioning unit is used for: Obtain the initial sound source position; According to the initial sound source position, predict the sound source position in the next cycle as the target sound source position; Substitute the target sound source position into the delay error term formula corresponding to multiple candidate delays of the microphone pairs in the microphone array, and take the candidate delay with the smallest value of the delay error term formula among the multiple candidate delays as the target delay. Among them, the multiple candidate delays are the delays corresponding to the k-th largest peak of the generalized cross-correlation function of the delays constructed for the microphone pairs of the microphone array, where k = 1,..., N, and N is a positive integer greater than or equal to 2; Construct a time-delay equation for the microphone array according to the target time delay of the microphone pairs in the microphone array, and determine the sound source position that minimizes the time-delay equation. If the distance between the determined sound source position that minimizes the time-delay equation and the target sound source position is not within a predetermined distance threshold, substitute the determined sound source position for the target sound source position and re-enter it into the time-delay error term formula until the distance between the determined sound source position and the target sound source position is within the predetermined distance threshold; Wherein, the time-delay equation is equal to the sum of the target time delays of the microphone pairs in the microphone array plus the product of a regularization term and a regularization coefficient.

6. The terminal device according to claim 5, wherein, There is an axis between the first part and the second part, and the axis rotates according to the three-dimensional spatial position of the determined sound source, so that the second part faces the sound source.

7. The terminal device according to claim 5, wherein, The terminal device is a smart TV, the first part is a base, and the second part is a display screen.

8. The terminal device according to claim 5, wherein, The terminal device is a smart speaker, the first part is an operating table, and the second part is a speaker box.

9. The terminal device according to claim 5, wherein, The terminal device is a conversation robot, the first part is a robot foot, and the second part is a robot face.

10. A sound source localization method is applied to the sound source localization unit of an audio-visual conference device. The audio-visual conference device further includes a base extending along a first plane, a display screen protruding from the base along a second plane, and a microphone array, wherein, At least some of the microphones in the microphone array are encapsulated in the base, and at least some of the microphones are encapsulated in the display screen. The method includes: Obtain an initial sound source position; According to the initial sound source position, predict the sound source position in the next cycle as the target sound source position; Substitute the target sound source position into the time-delay error term formula corresponding to multiple candidate time delays for the microphone pairs in the microphone array, and use the candidate time delay with the smallest value of the time-delay error term formula among the multiple candidate time delays as the target time delay. Wherein, the multiple candidate time delays are the time delays corresponding to the k-th largest peak of the generalized cross-correlation function of the time delays constructed for the microphone pairs in the microphone array, where k = 1,..., N, and N is a positive integer greater than or equal to 2; Construct a time-delay equation for the microphone array according to the target time delay of the microphone pairs in the microphone array, and determine the sound source position that minimizes the time-delay equation. If the distance between the determined sound source position that minimizes the time-delay equation and the target sound source position is not within a predetermined distance threshold, substitute the determined sound source position for the target sound source position and re-enter it into the time-delay error term formula until the distance between the determined sound source position and the target sound source position is within the predetermined distance threshold; Wherein, the time-delay equation is equal to the sum of the target time delays of the microphone pairs in the microphone array plus the product of a regularization term and a regularization coefficient.

11. The method according to claim 10, wherein, The obtaining of the initial sound source position includes: Construct the generalized cross-correlation function of the time delays for the microphone pairs in the microphone array, and determine the time delay that maximizes the generalized cross-correlation function; For this microphone pair, construct a time-delay expression based on the initial sound source position and the positions of the two microphones in this microphone pair, such that the time-delay expression is equal to the determined time delay that maximizes the generalized cross-correlation function, and solve to obtain the initial sound source position.

12. The method according to claim 11, wherein The generalized cross-correlation function is the integral in the frequency domain of the product of a frequency weighting function and the cross-spectrum of the sound signals respectively received by the microphone pair, after phase-shifting according to the time delay of the sound signals respectively received by the microphone pair.

13. The method according to claim 11, wherein, The time delay expression is the difference between the distances between the initial sound source position and the positions of the two microphones of the microphone pair, divided by the speed of sound.

14. The method according to claim 10, wherein The next period is the period during which the camera of the audio-visual conferencing device where the microphone array is located captures the next frame.

15. The method according to claim 10, wherein Predicting the sound source position in the next period according to the initial sound source position is performed according to the Kalman filtering method.

16. The method according to claim 10, wherein, The time delay error term formula is the square of the difference between the time delay corresponding to the k-th largest peak of the generalized cross-correlation function and the time delay expression based on the initial sound source position and the positions of the two microphones of the microphone pair, multiplied by the gain adjustment coefficient corresponding to the microphone pair, and divided by the k-th largest peak.

17. The method according to claim 10, wherein The regularization term is equal to the square of the difference between the sound source position to be determined and the target sound source position.

18. A computer storage medium comprising computer-executable code which, when executed by a processor, implements the method according to any one of claims 10-17.

Citation Information

Patent Citations

  • Video conference tracking system and method based on microphone array

    CN107809596A

  • Sound source positioning method of audio equipment and audio equipment

    CN110691196A