Real-time speaker tracking and positioning method and system based on audio and video combination
Through the combined audio and video method, the audio frames are divided and the audio feature vectors are constructed, which are then combined with video data to locate the speaker. This solves the problem of high computational complexity of traditional methods and achieves higher positioning accuracy and real-time performance.
Patent Information
- Application Number
- CN202510907238.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-02
AI Technical Summary
The traditional speaker localization method based on sound source direction finding algorithm has high computational complexity, resulting in low accuracy and real-time performance of speaker real-time tracking and positioning.
A combined audio and video method is adopted to obtain the audio signal before the speaker speaks, divide it into reverberation sound frames and direct sound frames, and construct an audio feature vector by combining the frequency domain amplitude distribution and voiceprint coefficient. The characteristics of the current speaker and historical speakers are matched, and positioning is performed in combination with video data.
The error of reverberation sound in positioning is reduced, the accuracy and real-time performance of speaker audio feature matching are improved, the amount of calculation is reduced, and the accuracy and real-time performance of positioning are improved.
Smart Images

Figure CN120412649B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speaker positioning, and in particular to a method and system for real-time tracking and positioning of a speaker based on audio and video integration. Background Art
[0002] A videoconferencing system is a remote conferencing tool that uses transmission lines and multimedia equipment to transfer audio, video, and document data, enabling instant and interactive communication between participants in different locations. Real-time speaker tracking is a key feature of a videoconferencing system. Accurate speaker location allows participants in different locations to identify the current speaker. This, by displaying the current speaker's video, improves communication and enhances participation and interactivity. Furthermore, timely confirmation of the current speaker by participants avoids confusion caused by multiple participants speaking simultaneously, ensuring meeting order and efficiency.
[0003] The primary technical approach currently used in videoconferencing systems is to calculate the speaker's azimuth angle at the conference table using a sound source direction-finding algorithm, and then acquire video data within the azimuth angle's shooting angle range to further pinpoint the speaker. Traditional positioning methods employ a sound source direction-finding algorithm based on Time Difference of Arrival (TDOA) to determine the sound source azimuth angle for real-time speaker location. However, this method requires processing audio signals collected by multiple microphones. Limited by the complex iterative process of the sound source direction-finding algorithm, the calculation of the speaker's azimuth angle is complex and computationally intensive, making it difficult to accurately locate the sound source direction in a short period of time. This, in turn, reduces the accuracy and real-time performance of real-time speaker tracking and location. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of this application is to provide a method and system for real-time speaker tracking and positioning based on audio and video integration. The technical solutions adopted are as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for real-time tracking and positioning of a speaker based on audio and video integration, the method comprising the following steps:
[0006] Acquire an audio signal before the speaker starts speaking, and divide the audio signal into multiple audio frames;
[0007] All audio data in each audio frame are integrated to determine the short-time energy of each audio frame, so as to divide all audio frames into reverberation sound frames and direct sound frames; based on the amplitude distribution of each audio frame at different frequencies in the frequency domain, the distribution characteristic value of each audio frame is determined, and the comprehensive characteristic value of each audio frame is determined by combining the proportion of all reverberation sound frames in all direct sound frames;
[0008] Based on the information distribution of each audio frame, all voiceprint coefficients of each audio frame are determined. The short-term energy of each audio frame, the comprehensive eigenvalue and all voiceprint coefficients are combined to form the audio feature vector of each audio frame of the speaker. By analyzing the similarity of the audio feature vectors of all audio frames between the current speaker and the historical speakers, the matching degree between the current speaker and the historical speakers is determined to determine the direction angle of the current speaker.
[0009] The video data from the camera in the preset shooting angle range where the direction angle of the current speaker is located is obtained, the facial key points of all participants in the video data are marked, and the current speaker is tracked and located.
[0010] Preferably, the short-time energy of each audio frame is the sum of all audio data in each audio frame.
[0011] Preferably, dividing all audio frames into reverberation sound frames and direct sound frames includes:
[0012] The short-time energy of all audio frames is used as the input of the threshold segmentation algorithm, and the segmentation threshold is output. The audio frames with short-time energy greater than the segmentation threshold are recorded as direct sound frames, and all other audio frames are recorded as reverberation sound frames.
[0013] Preferably, the method for determining the distribution characteristic value of each audio frame is:
[0014] All amplitudes of the frequency domain signal of each audio frame are used as input of the threshold segmentation algorithm, the segmentation threshold is output, and the frequency corresponding to the segmentation threshold is used as the segmentation frequency;
[0015] The cumulative sum of all square amplitudes before the split frequency in the frequency domain signal of each audio frame is taken as the low-frequency energy, the cumulative sum of all square amplitudes after the split frequency in the frequency domain signal is taken as the high-frequency energy, and the ratio of the high-frequency energy to the low-frequency energy is taken as the distribution characteristic value of each audio frame.
[0016] Preferably, the expression of the comprehensive feature value of each audio frame is: Where, represents the comprehensive feature value of audio frame i; 、 Respectively represent the number of all reverberation sound frames and the number of all direct sound frames in the audio signal; represents the distribution characteristic value of audio frame i; exp( ) represents an exponential function with a natural constant as the base.
[0017] Preferably, the method for determining all voiceprint coefficients of each audio frame is:
[0018] Calculate the Mel-frequency cepstral coefficients of each audio frame, and use the Mel-frequency cepstral coefficients of the first preset number of orders as the voiceprint coefficients of each audio frame.
[0019] Preferably, the matching degree between the current speaker and the historical speakers is a result of taking the average of the similarities of all audio feature vectors between the current speaker and the historical speakers.
[0020] Preferably, the method for determining the direction angle of the current speaker is:
[0021] If the matching degree between the current speaker and the historical speaker is greater than a preset threshold, the direction angle of the historical speaker is used as the direction angle of the current speaker. Otherwise, the direction angle of the current speaker is obtained using the sound source direction finding algorithm.
[0022] Preferably, the step of marking facial key points of all participants in the video data and tracking and locating the current speaker includes:
[0023] A facial detector is used to mark the facial key points of each participant in the video data, all the facial key points of all the participants are used as input of the pre-trained neural network, the probability of all the participants speaking is output, and the position of the participant with the highest probability is used as the position of the current speaker.
[0024] In the second aspect, an embodiment of the present application also provides a real-time tracking and positioning system for a speaker based on audio and video combination, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned real-time tracking and positioning methods for a speaker based on audio and video combination.
[0025] This application has at least the following beneficial effects:
[0026] The present application integrates all audio data in each audio frame to divide all audio frames into reverberation sound frames and direct sound frames, and constructs a comprehensive feature value based on the amplitude distribution of each audio frame at different frequencies in the frequency domain and the proportion of all reverberation sound frames in all direct sound frames. It analyzes the audio characteristics of different speakers according to the differences in the proportions of direct sound and reverberation sound of speakers at different positions in the conference room, reduces the error introduced by reverberation sound to the audio recognition of speakers at different positions, improves the accuracy of subsequent matching of speaker audio features, and thus improves the accuracy and real-time performance of sound source direction finding; further, the present application obtains the direction angle for tracking and positioning the current speaker by matching the audio feature vectors of the current speaker and historical speakers, uses historical speaker data to avoid additional complex calculations, reduces the computational complexity of real-time tracking and positioning of speakers, improves positioning accuracy, reduces the time for sound source positioning of speakers, and improves the real-time performance of speaker tracking and positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] Figure 1 A flowchart of the steps of a method for real-time speaker tracking and positioning based on audio and video integration provided in one embodiment of the present application;
[0029] Figure 2 A schematic diagram of a real-time speaker tracking and positioning system provided in one embodiment of the present application. Figure 2 It includes: conference table 1, participant seats 2, camera 3, microphone array 4, display screen 5;
[0030] Figure 3 A schematic diagram of a microphone array provided in one embodiment of the present application is shown. Figure 3 It includes: the plane 6 where the conference table is located and the microphone array 7, the coordinate axes x, y, z, and the coordinate origin O. DETAILED DESCRIPTION
[0031] To further illustrate the technical means and effectiveness of this application to achieve the intended invention objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of the method and system for real-time speaker tracking and positioning based on audio and video integration proposed in this application. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0032] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0033] The specific scheme of the speaker real-time tracking and positioning method and system based on audio and video combination provided by this application is described in detail below with reference to the accompanying drawings.
[0034] See also Figure 1 , which shows a flowchart of a method for real-time speaker tracking and positioning based on audio and video integration provided by an embodiment of the present application, the method comprising the following steps:
[0035] Step S1: Acquire an audio signal before the speaker starts speaking, and divide the audio signal into multiple audio frames.
[0036] In this embodiment, the speaker real-time tracking and positioning system includes: a conference table, participant seats, a camera, a microphone array, and a display screen. The schematic diagram of the speaker real-time tracking and positioning system is shown in FIG. Figure 2 As shown. Since the microphone cross array has a simple structure and can cover all directions of the conference table, and can reduce mutual interference between microphones, the microphone array of this embodiment adopts a cross array, which includes a total of 4 microphones as array elements. The microphone array schematic diagram is shown as follows: Figure 3 As shown. A spatial rectangular coordinate system is established with the origin d centimeters above the center of the conference table, where the x-axis and y-axis are parallel to the wide and long sides of the conference table, respectively, and the z-axis is perpendicular to the plane of the conference table. The coordinates of the four microphones are (d, 0, d), (-d, 0, d), (0, d, d), and (0, -d, d). In this embodiment, d is 30. The number of microphones and the value of d can also be set by the implementer according to the conference scenario. This embodiment does not impose any special restrictions.
[0037] In this embodiment, seven cameras are installed outside the microphone array to capture participants at different angles around the conference table and obtain the shooting angle range of each camera. This embodiment uses a high-definition camera model HDC-5000, which is installed at the same height as the participants' heads to ensure clear facial images of the participants.
[0038] In this embodiment, the xOy plane in the spatial rectangular coordinate system is used as a reference, and the y-axis is rotated. Rotation toward the negative semi-axis direction of the x-axis is negative, and rotation toward the positive semi-axis direction of the x-axis is positive. The preset shooting angle ranges of each camera are obtained as [-157.5, -112.5), [-112.5, -67.5), [-67.5, -22.5), [-22.5, 22.5), [22.5, 67.5), [67.5, 112.5), and [112.5, 157.5]. The setting of the preset shooting angle range is artificial, and the implementer can also set it according to the specific situation. This embodiment does not impose any special restrictions.
[0039] In this embodiment, the audio data collected by each element in the microphone array and the video data collected by the camera are uploaded to the host computer via Ethernet. The audio data is encoded using PCM with a sampling frequency of 16kHz and a bit depth of 16 bits; the video data is encoded using H.264 with a resolution of 1920×1080 and a frame rate of 30fps. Data transmission uses the TCP / IP protocol to ensure reliable data transmission.
[0040] A preset number of audio data before the speaker starts speaking is used to form an audio signal, and the audio signal is divided into N audio frames. In this embodiment, each audio frame contains M audio data for use in finding the speaker's sound source. The preset number, the number of audio frames N, and the number of audio frames are all manually set. In this embodiment, the preset number is 6400, and the value of N is 20. Therefore, the value of M is 320. The implementer can also set it according to the specific situation. This embodiment does not impose any special restrictions.
[0041] Furthermore, in order to reduce spectral leakage, each audio frame is windowed. In this embodiment, a Hamming window is used to window the audio frame. In actual application, the implementer may also use other windowing methods such as a Haining window according to specific circumstances. This embodiment does not impose any special restrictions on the selection of the windowing method.
[0042] The Hamming window is a well-known technology, and its specific principle will not be described in detail.
[0043] Step S2: Integrate all audio data in each audio frame and determine the short-time energy of each audio frame to divide all audio frames into reverberation sound frames and direct sound frames; based on the amplitude distribution of each audio frame at different frequencies in the frequency domain, determine the distribution characteristic value of each audio frame, and combine the proportion of all reverberation sound frames in all direct sound frames to determine the comprehensive characteristic value of each audio frame.
[0044] In a conference room, the speaker's voice reflects multiple times throughout the room. When other participants speak, the sound waves travel directly to the microphone. They also reflect off hard surfaces like walls, ceilings, and floors, creating multi-reflected sound waves known as reverberation. Reverberation alters the time and frequency domain characteristics of the speaker's audio data, causing the audio features extracted from the mixed sound to differ from those of the direct sound. The superposition of reverberation and direct sound distorts the audio data, reducing the accuracy of subsequent speaker matching.
[0045] Therefore, this embodiment integrates all audio data in each audio frame to determine the short-time energy of each audio frame, thereby dividing all audio frames into reverberation sound frames and direct sound frames. Based on the amplitude distribution of each audio frame at different frequencies in the frequency domain, the distribution characteristic value of each audio frame is determined. In combination with the proportion of all reverberation sound frames in all direct sound frames, the comprehensive characteristic value of each audio frame is determined to reduce the error introduced by reverberation sound when matching audio data of speakers in different locations. The specific process is as follows:
[0046] (1) All audio data in each audio frame are integrated to determine the short-time energy of each audio frame to divide all audio frames into reverberation sound frames and direct sound frames, specifically:
[0047] The sum of all audio data in each audio frame is used as the short-time energy of each audio frame, which reflects the energy of the audio frame and is used to distinguish direct sound from reverberation sound. Since direct sound has higher energy and lower attenuation, the larger the short-time energy, the more likely the audio frame is to be a direct sound. Conversely, since reverberation sound has lower energy and faster attenuation, the smaller the short-time energy, the more likely the audio frame is to be the audio data corresponding to reverberation sound.
[0048] (2) Based on the short-time energy obtained in (1), the audio frames are classified into the following categories:
[0049] The short-time energy of all audio frames is used as the input of the threshold segmentation algorithm, and the segmentation threshold is output. The audio frames with short-time energy greater than the segmentation threshold are recorded as direct sound frames, and all other audio frames are recorded as reverberation sound frames.
[0050] It should be noted that there are many commonly used threshold segmentation algorithms. In this embodiment, the maximum inter-class variance algorithm is used to classify audio frames. In actual application, implementers can also use other threshold segmentation algorithms based on specific circumstances. Regarding the selection of threshold segmentation algorithms, this embodiment does not impose any special restrictions.
[0051] Among them, the maximum inter-class variance algorithm is a well-known technology, and its specific principle is not repeated here.
[0052] It is supplemented that, in this embodiment, all threshold segmentation algorithms involve the maximum inter-class variance algorithm.
[0053] (3) Further, based on the amplitude distribution of each audio frame at different frequencies in the frequency domain, the distribution characteristic value of each audio frame is determined, specifically:
[0054] The energy distribution of audio data from different speakers in the frequency domain is different. At the same time, the reverberation sound contains a larger proportion of low-frequency components, and the audio data at low frequencies is more interfered by the reverberation sound. Therefore, the reverberation sound frame provides less information in subsequent speaker matching, while the direct sound frame has less attenuation, and its energy is concentrated in fewer audio data frames. It contains more speaker audio feature information and plays a greater role in speaker matching.
[0055] The frequency domain characteristics of the speaker's audio data result in different energy distributions for the low-frequency and high-frequency components. Considering the statistical characteristics of energy distribution, in the frequency domain sequence of the speaker's audio data frame, its low-frequency and high-frequency components are usually concentrated in two fixed intervals, forming two peaks at the low and high frequencies of the frequency domain signal.
[0056] Therefore, based on the above analysis, all amplitudes of the frequency domain signal of each audio frame are used as the input of the threshold segmentation algorithm, the segmentation threshold is output, and the frequency corresponding to the segmentation threshold is used as the segmentation frequency;
[0057] The cumulative sum of all amplitude squares before the split frequency in the frequency domain signal of each audio frame is taken as the low-frequency energy, the cumulative sum of all amplitude squares after the split frequency in the frequency domain signal is taken as the high-frequency energy, and the ratio of high-frequency energy to low-frequency energy is taken as the distribution characteristic value of each audio frame, which is used to reflect the proportion of audio feature information belonging to the speaker when the speaker is speaking. If the ratio of high-frequency energy to low-frequency energy of the current audio frame is larger, the distribution characteristic value is larger, which means that the possibility that the current audio frame belongs to the audio feature information of the speaker is greater. Conversely, if the ratio of high-frequency energy to low-frequency energy of the current audio frame is smaller, that is, the distribution characteristic value is smaller, which means that the possibility that the audio feature information in the current audio frame belongs to the speaker is smaller.
[0058] (4) Further, based on the distribution characteristic value of each audio frame and the proportion of all reverberation sound frames in all direct sound frames, the comprehensive characteristic value of each audio frame is determined, specifically:
[0059] Comprehensive feature value of audio frame i The expression is: Where, 、 Respectively represent the number of all reverberation sound frames and the number of all direct sound frames in the audio signal; represents the distribution characteristic value of audio frame i; exp( ) represents an exponential function with a natural constant as the base.
[0060] According to the total eigenvalue of each audio frame, it can be understood that the comprehensive eigenvalue reflects the proportion of the speaker's audio information content provided by the audio frame. If the ratio between the number of reverberation sound frames and the number of direct sound frames is larger, that is, the proportion of the reverberation sound frames is larger, and the distribution eigenvalue of the current audio frame is smaller, it means that the current audio frame contains a smaller proportion of the speaker's audio feature information, and the corresponding comprehensive distribution value is smaller; conversely, if the ratio between the number of reverberation sound frames and the number of direct sound frames is smaller, that is, the proportion of the reverberation sound frames is smaller, and the distribution eigenvalue of the current audio frame is larger, it means that the current audio frame contains a larger proportion of the speaker's audio feature information, and the corresponding comprehensive distribution value is larger.
[0061] Step S3: Based on the information distribution of each audio frame, all voiceprint coefficients of each audio frame are determined, and the short-time energy, comprehensive eigenvalue and all voiceprint coefficients of each audio frame are combined to form the audio feature vector of each audio frame of the speaker; by analyzing the similarity of the audio feature vectors of all audio frames between the current speaker and the historical speakers, the matching degree between the current speaker and the historical speakers is determined to determine the direction angle of the current speaker.
[0062] Taking into account the inherent voiceprint features of the speaker, this embodiment calculates the Mel-frequency cepstral coefficients of each audio frame, and uses the Mel-frequency cepstral coefficients of the first preset number of orders as the voiceprint coefficients of each audio frame to characterize the inherent voiceprint features of the speaker.
[0063] It should be noted that the value of the preset number is set manually. In this embodiment, the value of the preset number is 12. In actual application, the implementer can also set it by himself according to the specific situation. This embodiment does not impose any special restrictions.
[0064] The calculation method of the Mel-frequency cepstral coefficient is a well-known technology, and its specific calculation process is not repeated here.
[0065] Furthermore, in order to integrate all the audio features of the speaker, the short-time energy, comprehensive eigenvalue and all voiceprint coefficients of each audio frame are combined to form the audio feature vector of each audio frame of the speaker for subsequent speaker matching.
[0066] Furthermore, by obtaining the speaker's direction angle, the speaker's position is located. The sound source direction finding algorithm is usually used to obtain the speaker's direction angle. However, considering that speakers often speak repeatedly during a meeting, the sound source direction finding algorithm has a high computational complexity. If the sound source direction finding algorithm is used to obtain the speaker's direction angle every time a speaker is changed, a huge amount of calculation will be generated, affecting the real-time performance of speaker tracking and positioning.
[0067] Therefore, this embodiment uses the audio data and direction angle of the speaker in the meeting as prior knowledge to perform speaker matching. There is no need to enter the location information of the participants in advance, and additional complex calculations are avoided. The time for sound source positioning for the speaker is reduced, and the real-time performance of speaker tracking and positioning based on audio and video joint is improved. That is, by analyzing the similarity of the audio feature vectors of all audio frames between the current speaker and the historical speakers, the matching degree between the current speaker and the historical speakers is determined to determine the direction angle of the current speaker. Specifically,
[0068] For the first speaker at the meeting, the speaker's audio feature vector is calculated, and the speaker's direction angle is obtained using the sound source direction finding algorithm. The speaker's audio feature vector and direction angle are used as historical speaker data and recorded in the speaker audio library.
[0069] After the speech is finished, it is determined whether the meeting is over. If the meeting is over, the conference video system is turned off. If the meeting is not over, the audio feature vector of the speaker is calculated and matched with the audio feature vectors of historical speakers in the speaker audio library.
[0070] Furthermore, for the remaining speakers after the first speaker in the meeting, the result of taking the average of the similarities of all audio feature vectors between the current speaker and the historical speakers is used as the matching degree between the current speaker and the historical speakers, reflecting the proportion of reverberation sound caused by the position and the similarity of the speakers' inherent voiceprint features.
[0071] It should be noted that there are many methods for measuring the similarity between vectors. In this embodiment, the cosine similarity of all audio feature vectors between the current speaker and the historical speakers is used as the similarity of all audio feature vectors between the current speaker and the historical speakers. In actual application, as other implementation methods, the implementer may also use other methods for measuring the similarity between vectors, such as the inverse of the Euclidean distance. This embodiment does not impose any special restrictions on the selection of methods for measuring the similarity between vectors.
[0072] The calculation method of cosine similarity is a well-known technology, and its specific calculation process will not be described in detail.
[0073] Furthermore, if the matching degree between the current speaker and the historical speaker is greater than a preset threshold, the direction angle of the historical speaker is used as the direction angle of the current speaker; otherwise, the direction angle of the current speaker is obtained using a sound source direction finding algorithm.
[0074] It should be noted that the value of the preset threshold is set manually. In this embodiment, the value of the preset threshold is 0.9. The implementer can also set it by himself according to the specific situation. This embodiment does not impose any special restrictions.
[0075] It should be noted that the sound source direction finding algorithm used in this embodiment is the Generalized Cross Correlation-Phase Transform (GCC-PHAT) algorithm. In actual application, as other implementation methods, the implementer may also use other sound source direction finding algorithms to obtain the direction angle.
[0076] The generalized cross-correlation-phase transformation algorithm is a well-known technology, and the specific principle of using it to localize the sound source and obtain the direction angle will not be described in detail.
[0077] Step S4: obtaining video data from a camera in a preset shooting angle range where the direction angle of the current speaker is located, marking facial key points of all participants in the video data, and tracking and locating the current speaker.
[0078] Based on the direction angle of the current speaker obtained in step S3, the direction angle of the current speaker is compared with the shooting angle ranges of the seven cameras to identify which shooting angle range the direction angle of the current speaker is located in. The video data obtained by the cameras in the shooting angle range in which the current speaker is located is used for speaker positioning analysis. The RetinaFace facial detector is used to annotate all facial key points of all participants in the video data.
[0079] Among them, the method of using the RetinaFace face detector to mark facial key points is a well-known technology, and its specific marking process is not repeated here.
[0080] Furthermore, a facial detector is used to mark the facial key points of each participant in the video data, and all the facial key points of all the participants are used as input to the pre-trained neural network. The probability of all the participants speaking is output, and the position of the participant with the highest probability is used as the position of the current speaker.
[0081] The training process of the pre-trained neural network is as follows: video data of different people is obtained from the face recognition dataset (YouTube Faces DB), and the speaking status of each person is set as a label 0 or 1, where 0 indicates speaking and 1 indicates not speaking. The RetinaFace facial detector is used to annotate the facial key points of different speakers in all video data to form a facial key point sequence. The facial key point sequence of each person is used as the input of the neural network. The cross-entropy loss function and the Adam function are used to train it to obtain a pre-trained neural network for identifying the probability of people speaking or not speaking.
[0082] It should be noted that, in this embodiment, the CNN convolutional neural network is used as the neural network used for training. In actual application, the implementer may also adopt other neural network models based on specific circumstances. This embodiment does not impose any special restrictions on the selection of the neural network model.
[0083] Among them, the face recognition dataset (YouTube Faces DB), CNN convolutional neural network, cross-entropy loss function and Adam function are all well-known technologies, and their specific principles are not repeated here.
[0084] Based on the same inventive concept as the above method, an embodiment of the present application also provides a real-time tracking and positioning system for a speaker based on audio and video combination, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned methods for real-time tracking and positioning of a speaker based on audio and video combination.
[0085] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0086] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0087] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A real-time speaker tracking and positioning method based on audio and video integration, characterized in that: The method comprises the following steps: Acquire an audio signal before the speaker starts speaking, and divide the audio signal into multiple audio frames; All audio data in each audio frame are integrated to determine the short-time energy of each audio frame, so as to divide all audio frames into reverberation sound frames and direct sound frames; based on the amplitude distribution of each audio frame at different frequencies in the frequency domain, the distribution characteristic value of each audio frame is determined, and the comprehensive characteristic value of each audio frame is determined by combining the proportion of all reverberation sound frames in all direct sound frames; Based on the information distribution of each audio frame, all voiceprint coefficients of each audio frame are determined. The short-term energy of each audio frame, the comprehensive eigenvalue and all voiceprint coefficients are combined to form the audio feature vector of each audio frame of the speaker. By analyzing the similarity of the audio feature vectors of all audio frames between the current speaker and the historical speakers, the matching degree between the current speaker and the historical speakers is determined to determine the direction angle of the current speaker. Obtain video data from a camera within a preset shooting angle range where the current speaker's direction angle is located, mark the facial key points of all participants in the video data, and track and locate the current speaker; The method for determining the distribution characteristic value of each audio frame is: All amplitudes of the frequency domain signal of each audio frame are used as input of the threshold segmentation algorithm, the segmentation threshold is output, and the frequency corresponding to the segmentation threshold is used as the segmentation frequency; The cumulative sum of all square amplitudes before the split frequency in the frequency domain signal of each audio frame is used as the low-frequency energy, the cumulative sum of all square amplitudes after the split frequency in the frequency domain signal is used as the high-frequency energy, and the ratio of the high-frequency energy to the low-frequency energy is used as the distribution characteristic value of each audio frame; The expression of the comprehensive feature value of each audio frame is: Where, represents the comprehensive feature value of audio frame i; 、 Respectively represent the number of all reverberation sound frames and the number of all direct sound frames in the audio signal; represents the distribution characteristic value of audio frame i; exp( ) represents an exponential function with a natural constant as the base.
2. The speaker real-time tracking and positioning method based on audio and video combination according to claim 1 is characterized in that: The short-time energy of each audio frame is the sum of all audio data in each audio frame.
3. The speaker real-time tracking and positioning method based on audio and video combination as claimed in claim 1, characterized in that: The method of dividing all audio frames into reverberation sound frames and direct sound frames includes: The short-time energy of all audio frames is used as the input of the threshold segmentation algorithm, and the segmentation threshold is output. The audio frames with short-time energy greater than the segmentation threshold are recorded as direct sound frames, and all other audio frames are recorded as reverberation sound frames.
4. The speaker real-time tracking and positioning method based on audio and video combination according to claim 1, characterized in that: The method for determining all voiceprint coefficients of each audio frame is as follows: Calculate the Mel-frequency cepstral coefficients of each audio frame, and use the Mel-frequency cepstral coefficients of the first preset number of orders as the voiceprint coefficients of each audio frame.
5. The speaker real-time tracking and positioning method based on audio and video combination according to claim 1 is characterized in that: The matching degree between the current speaker and the historical speakers is the result of taking the average of the similarities of all audio feature vectors between the current speaker and the historical speakers.
6. The speaker real-time tracking and positioning method based on audio and video combination according to claim 1, characterized in that: The method for determining the direction angle of the current speaker is: If the matching degree between the current speaker and the historical speaker is greater than a preset threshold, the direction angle of the historical speaker is used as the direction angle of the current speaker. Otherwise, the direction angle of the current speaker is obtained using the sound source direction finding algorithm.
7. The speaker real-time tracking and positioning method based on audio and video combination according to claim 1, characterized in that: The process of labeling the facial key points of all participants in the video data and tracking and locating the current speaker includes: A facial detector is used to mark the facial key points of each participant in the video data, all the facial key points of all the participants are used as input of the pre-trained neural network, the probability of all the participants speaking is output, and the position of the participant with the highest probability is used as the position of the current speaker.
8. A speaker real-time tracking and positioning system based on audio and video integration, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for real-time tracking and positioning of a speaker based on audio and video combination as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio noise reduction method and device, equipment and medium
CN112233688A
Audio noise reduction method and device thereof, equipment and medium
CN112233689A