A Speaker Tracking Method and System Based on Multimodal Information

Through the method of calculating scores and dynamic update of prior databases by multimodal information, the problems of low spokesperson tracking accuracy and lack of dynamic updates in the prior art are solved, and a spokesperson tracking system with higher accuracy and applicability are realized.

CN115131405BActive Publication Date: 2025-07-01SHENYANG AEROSPACE UNIVERSITY +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210792440.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-07-01
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

The existing spokesperson tracking technology is low in accuracy, especially when people are dense or hierarchical, and it is difficult to accurately determine the spokesperson, and it lacks dynamically updated systems and high confidence data entry capabilities.

Method used

A spokesperson tracking method based on multimodal information is adopted to calculate the lip movement score, phonological matching score and lip synchronization score of each face in the image, and combine the prior database to support early entry and dynamic update of vocal and face pairs.

Benefits of technology

It improves the accuracy and accuracy of spokesperson tracking, especially in people-intensive and hierarchical distribution scenarios, and supports dynamic updates of high confidence data pairs, enhancing the applicability and practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131405B_ABST
    Figure CN115131405B_ABST
Patent Text Reader

Abstract

The present invention discloses a speaker tracking method and system based on multimodal information, which relates to the field of speaker tracking. It can be applied to the online speaker tracking task of offline meetings or online meetings, and can quickly and accurately locate the speaker and give a close-up of the speaker; it can also be used for the non-online task of annotating the speaker in each part of the provided video. In the case where multiple faces appear in the same picture and each person takes turns speaking, the input image and the corresponding audio information are used to calculate the speech lip movement score, voice appearance matching score, and lip shape synchronization score of each face in the image, and the specific speaker is located according to the score of each face in the image. At the same time, it supports pre-recording and registering paired human voices and faces, and supports entering the paired human voices and faces with high pairing confidence into the prior database during use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speaker tracking, and particularly to a speaker tracking method and system based on multi-modal information. Background Art

[0002] Regarding the problem of "identifying the speaker in a multi-person image", existing methods either rely on some physical devices such as an array microphone for speaker localization, or rely on pre-registering the faces and voices of the participants, or simply use single-modal information such as face image information or voice information for speaker tracking. The accuracy of these speaker tracking methods is relatively low, and the situations where an array microphone must be used or pre-registration must be carried out limit their application scenarios.

[0003] The solution of Patent CN111263106A aims to solve the problem of quickly detecting the current speaker among multiple participants in a meeting scene. It proposes to obtain the position distribution of people by processing image information, then perform sound localization processing based on the microphone array, and finally combine the information of both to determine the position of the speaker and the corresponding face image. However, this method has strict requirements on the distribution of people. When people are dense or distributed hierarchically, it will be difficult to determine the real speaker mainly relying on the sound localization information of the microphone array.

[0004] Patent CN112633219A proposes to monitor the lip area of each person in real time and judge that the person with a lip area greater than the preset area threshold is speaking. The disadvantage of this method is that the accuracy is not high enough, and behaviors such as yawning, eating, and grinning will also cause the lip area to be higher than the threshold and thus be misjudged as the speaker.

[0005] The solution proposed by Patent CN112040119A requires pre-entering the face information and voice information of the person, and then it can detect the specific speaker in the picture, which has certain limitations.

[0006] Patent CN112487978A proposes two solutions: one is to compare the pre-entered information with the current face and voice data to judge whether they match; the other is to use the SyncNet model to extract the feature vectors of the face and voice, calculate the cosine similarity, and judge whether they match. This solution has better effects than the previous ones, but it has poor effects in the case of low resolution and blurred lip movements.

[0007] The above solutions are not sufficient for mining the sound and picture information in the video, and the technical means used are relatively simple and traditional. None of the solutions consider the correlation between human speech and facial features, resulting in low speaker tracking accuracy and poor effects in scenarios with blurred lip movements. At the same time, existing technical solutions use pre-recorded pairs of human faces and voices, but do not design a dynamically updated system, and do not record pairs of human faces and voices with a high enough pairing confidence in the pairing database during the usage process. Summary of the Invention

[0008] To solve the deficiencies of the prior art, for the task of locating the speaker in the video, the present invention proposes a speaker tracking method and system based on multimodal information, which calculates the speaking lip movement score, voice-face matching score, and lip shape synchronization score for each human face in the image using the input image and the corresponding audio information, and locates the specific speaker according to the scores of each human face in the image. At the same time, it supports pre-registering and pairing human voices and faces in advance, and also supports entering pairs of human voices and faces with a high pairing confidence into the prior database during the usage process.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] In a first aspect, the present invention provides a speaker tracking system based on multimodal information, the system includes: a voice identity information feature extraction module, a voice content information feature extraction module, an image facial information feature extraction module, an image content feature extraction module, a human face image quality calculation module, a human face detection and grouping module, a lip shape synchronization module, a speaking lip movement recognition module, a voice-face matching module, and a prior database.

[0011] Using the voice identity information feature extraction module, the voice identity information feature vector is extracted from the input audio.

[0012] Using the voice content information feature extraction module, the voice content information feature vector is extracted from the input audio.

[0013] Using the image facial information feature extraction module, r input images face 1 …face r are successively used to extract the frame-by-frame human face facial feature vectors, and each image is input into the human face image quality calculation module to calculate the quality score of each input image. The quality scores of the r images are concatenated on the channel dimension with the r frame-by-frame human face facial feature vectors to extract the human face facial feature vector.

[0014] Using the content feature extraction module of the image, splice r input images in the time dimension to obtain the spliced image vector; input each input image into the face image quality calculation module separately to obtain the quality score of each input image, copy and expand the quality score of each input image and splice and extract features with the image vector to obtain the face lip content feature vector.

[0015] The face image quality calculation module inputs a single-color face image into a convolutional neural network to obtain the image quality score.

[0016] The face detection and grouping module detects faces in the video segment frame by frame, gives the matrix information of each face, groups the face matrices belonging to the same person, and completes the face information for the frames lacking face information to obtain a complete sequence of face matrices.

[0017] The lip synchronization module inputs the face lip content feature vector and the speech content information feature vector, and uses the cosine similarity to calculate the similarity of the two feature vectors to obtain the lip synchronization score.

[0018] The speaking lip movement recognition module inputs the face lip content feature vector into one or more fully connected layers with activation functions, and then inputs it into a fully connected layer with a Sigmoid activation function to obtain the speaking lip movement score.

[0019] The voice and appearance matching module inputs the face appearance information feature vector and the voice identity information feature vector, and uses the L1 distance to calculate the distance between the two feature vectors to obtain the voice and appearance matching score.

[0020] The prior database supports pre-recording the prior database and recording the prior database during use, and preferentially uses the prior database for matching during the speaker tracking process.

[0021] The identity information feature extraction module of the said voice is specifically: for the input audio, extract the network filter bank (Filter Bank) feature v0 through a Mel filter; input the network filter bank feature v0 into the first convolutional neural network (ECAPA-TDNN), extract the intermediate vector v1 of w1 dimensions, perform L2 regularization on the intermediate vector v1, and extract the voice identity information feature vector emb through c1 fully connected layers. vid 。

[0022] The content information feature extraction module of the voice is specifically as follows: perform L2 regularization on the intermediate vector v1, and obtain an intermediate vector v2 with dimension w2 through c2 fully connected layers; pass the intermediate vector v2 through c3 fully connected layers to obtain an intermediate vector v3 with dimension w3; use residual connection to add the intermediate vectors v2 and v3 to obtain v4 = v2 + v3, and then pass through c4 fully connected layers to obtain the voice content information feature vector emb vct 。

[0023] The face information feature extraction module of the image is specifically as follows: sequentially input r input images face 1 …face r into the second convolutional neural network (Inception-V1), and extract an intermediate vector with dimension w4 and perform L2 regularization, and extract a feature vector with dimension w5 through c5 fully connected layers After processing the r input images, a feature vector z with shape (r, w5) will be obtained fid ; Input each input image face i individually into the face image quality calculation module to calculate the quality score q of each input image i ;

[0024] The r input images will obtain a quality score vector q with shape (r, 1); Concatenate the quality score vector q and the feature vector z fid to obtain a vector with shape (r, w5 + 1), and input it into the recurrent neural network (LSTM) to calculate an intermediate vector z1 with dimension w5 + 1; Pass the intermediate vector z1 through c6 fully connected layers to obtain the face feature vector emb that synthesizes the r input images fid 。

[0025] The content feature extraction module of the image is specifically as follows; Concatenate the r input images in the time dimension and keep other dimensions unchanged to obtain a vector with size (c, w * r, h), where c represents the number of channels of the input image. If the input is a color image, then c = 3; If the input is a grayscale image, then c = 1; where r represents the number of input images; w represents the number of pixels in the width of the input image; h represents the number of pixels in the height of the input image. The concatenated input image concatenation vector is x0;

[0026] Input each input image individually into the face image quality calculation module to obtain a quality score vector x1 with shape (r, 1);

[0027] Copy and expand the quality score vector x1 with shape (r, 1) to a quality score vector x2 with shape (1, w*r, h), where x2[1, i, j] = x1[i % w, 1], i ∈ [0, w*r), j ∈ [0, h); Concatenate the input image splicing vector x0 and the quality score vector x2 in the first dimension to obtain a feature vector x3 with shape (c + 1, w*r, h).

[0028] Input the feature vector x3 into the third convolutional neural network to extract a feature vector with dimension w6, denoted as x4; Perform L2 normalization on the intermediate vector x4 to obtain the content feature vector emb fct 。

[0029] The face image quality calculation module inputs a single-color face image into the fourth convolutional neural network (ResNet50) to obtain an intermediate vector v with dimension w7, and inputs this intermediate vector into a fully connected layer with a Sigmoid activation function to obtain the image quality score score quality ∈(0, 1);

[0030] The face detection and grouping module: Use a deep learning algorithm to detect all faces in each frame of the video segment to obtain the matrix information of each face represents the matrix information of the i-th face detected in the j-th frame; Group the face matrices of the same person in all frames according to the intersection over union of the face matrix information of adjacent frames. If and The intersection over union is greater than the set threshold, then it is determined that these two face matrices belong to the same person and will be grouped into the same group; Use linear interpolation to complete the missing face information frames according to the face matrix information of adjacent frames; According to the completed face matrix sequence Crop to obtain a face image sequence

[0031] The lip synchronization module inputs the face lip content feature vector emb fct and the speech content information feature vector emb vct , and uses the cosine similarity to calculate the similarity of the two feature vectors, which is the lip synchronization score score ct , where score ct ∈[-1, 1]; The higher the score, the better the match.

[0032] The voice and appearance matching module inputs the face appearance information feature vector emb fid and the voice identity information feature vector emb vid , and uses the L1 distance to calculate the distance between the two feature vectors, which is the voice and appearance matching score score id ; Among them, scoreid ≥0; The smaller the score, the better the match.

[0033] The said lip movement recognition module for speech inputs the face lip content feature vector emb fct into a fully connected layer with an activation function to obtain an intermediate vector a1 of dimension w8; The intermediate vector a1 is input into a fully connected layer with a Sigmoid activation function to obtain the speech lip movement score score talk ∈(0, 1). The higher the speech lip movement score, the higher the possibility that the face corresponding to the calculated face lip content feature vector is speaking;

[0034] The said prior database pre - inputs several face photos corresponding to the personnel, the human voice audio. The face photo sequence is input into the face appearance information feature extraction module of the image to obtain the face appearance information feature vector emb fid of each personnel. The human voice audio is denoised and input into the speech identity information feature extraction module to extract the speech identity information feature vector emb vid of each personnel. The vectors emb vid and emb fid are saved into the prior database. During the speaker tracking process, audio - visual matching based on the prior database is preferentially performed.

[0035] The said prior database supports inputting or updating during use. During use, the human voice - face pairs with high pairing confidence are input into the database. Specifically: When the matching speech identity information feature vector and the image face appearance information feature vector are found according to modules such as lip - shape synchronization, audio - visual matching, and speech lip movement detection, the vector pairs with matching scores higher than the input threshold are saved into the prior database;

[0036] The said speech identity information feature extraction module Model vid and the image face appearance information feature extraction module Model fid are jointly trained. The training process is as follows: The face pictures and human voice audios of the same personnel are respectively input into Model fid and Model vid to obtain emb fid and emb vid ;

[0037] The mean square error loss function Loss1 is used as shown in Equation (1):

[0038] Loss1 = MSE(emb fid , emb vid ) (1)

[0039] The said speech content information feature extraction module Model vctCo - train with the content information feature extraction module Model of the image fct Specifically, the co - training is as follows:

[0040] For the content information feature extraction module Model of speech vct All network parameters of the first convolutional neural network in it come from the parameters of the first convolutional neural network of the speech identity information feature extraction module Model vid During the training process, the numerical values of these parameters are fixed and do not participate in the parameter update in the backpropagation process;

[0041] Input the face image sequence and the human voice audio segment corresponding to the speech segment of the same person into Model fct and Model vct respectively, and obtain the lip content feature vector emb of the image fct and the speech content information feature vector emb based on the audio vct ; Input the human voice audio that has no corresponding relationship with the picture sequence into Model vct to obtain the unmatched speech content information feature vector emb′ vct; By maximizing the cosine similarity between emb fct and emb′ vct and minimizing the cosine similarity between emb fct and emb vct the two models learn the content information in the video; The loss function Loss2 is shown in Equation (2):

[0042] Loss2 = CosineSim(emb fct , emb vct ) - CosineSim(emb fct , - emb′ vct ) (2)

[0043] The speech lip movement recognition module is denoted as Model talk , and it is trained on the emb fct extracted by the content information feature extraction module of the image;

[0044] Specifically: Input the face image sequence of a person speaking into Model fct to obtain Input the face image sequence of a person not speaking into Model fct to obtain Input and into Model talk to obtain the corresponding speech lip movement score and Train the model using binary cross - entropy loss to minimize and maximize The loss function Loss3 is shown in Equation (3):

[0045]

[0046] On the other hand, the present invention provides a speaker tracking method based on multimodal information, which is implemented by using the above - mentioned speaker tracking system based on multimodal information, and includes the following steps:

[0047] S1: Obtain audio and video, and respectively use an audio acquisition device and a video acquisition device to obtain an audio segment and a video segment from time t to time t + s;

[0048] S2: Human voice judgment and speech feature extraction. Determine whether the audio segment contains human voice; if it does not contain human voice, it is determined that no one is speaking from time t to time t + s, and go to S9; if it contains human voice, input the audio segment into the speech identity information feature extraction module to obtain the speech identity information feature vector emb vid ; and input the audio segment into the speech content information feature extraction module to obtain the speech content information feature vector emb vct ;

[0049] S3: Extract the sequence of face images. Input the video segment frame by frame into the face detection and grouping module to obtain the sequence of face images

[0050] S4: Image feature extraction. Input the sequence of face images into the face image quality calculation module to obtain the image quality score corresponding to each frame of face image Input and into the face appearance information feature extraction module of the image to obtain the sequence of face appearance feature vectors Input and into the face lip content feature extraction module of the image to obtain the face lip content feature vector

[0051] S5: Retrieve all the recorded speech identity information feature vectors in the prior database and determine whether there is a recorded human voice similar to the speech identity feature vector emb vid ;

[0052] If there exists a recorded human voice vector emb′ vid similar to emb vid , then go to S6;

[0053] If there is no recorded human voice similar to the speech identity feature vector embvid If the input human voices are similar, go to S7;

[0054] S6: Retrieve the target facial feature vector corresponding to emb′ vid from the candidate sequence of facial information feature vectors in the given image and find out if there is a feature vector with a similarity higher than the matching threshold threshold . If there is, mark and output the corresponding face matrix sequence information; if not, it is determined that there is no face in the current picture that matches the corresponding human voice, and go to S9; match

[0055] S7: Input the pairing of the i-th person's in the image with emb vct into the lip synchronization module to obtain the lip synchronization score Input and emb vid into the audio-visual matching module to calculate the audio-visual matching score Input into the speech lip movement recognition module to calculate the speech lip movement score

[0056] Based on the lip synchronization score, the audio-visual matching score, and the speech lip movement score, calculate the weighted final score Compare the final score with the recognition threshold threshold score . If the score of each person's face image sequence is lower than the recognition threshold, it is determined that there is no face that matches the human voice, and go to S9; if the scores of one or more people's face image sequences are higher than the recognition threshold, the person with the highest score is recorded as the current speaker;

[0057] S8: If the final score of the current speaker is higher than the input threshold threshold record , register the emb vid corresponding to the current speaker and emb fid in the prior database;

[0058] S9: t = t + s, and return to step S1.

[0059] The beneficial effects of adopting the above technical solutions are as follows:

[0060] ​1. The present invention provides a speaker tracking method and system based on multimodal information, which comprehensively calculates the speaking lip movement score, lip shape synchronization score, and voice appearance matching score of the human voice and face, makes a judgment on the current speaker in the image, thereby supporting the entry of data pairs with high pairing confidence into the database during the operation process, and the database supports the pre-entry of registered and matched face and voice data pairs.

[0061] 2. By calculating the matching score between the input human voice and each face in the image, the present invention solves the problems of dense personnel and multiple people at the same angle that cannot be solved by the sound localization information relying on the microphone array in the traditional method.

[0062] 3. The present invention uses a multi-layer neural network to extract the deep information of the face image, which is more accurate than using the shallow lip area data to judge whether the face is speaking.

[0063] 4. The present invention not only supports the pre-entry of face and voice data pairs, but also supports the judgment of whether the newly appeared face and voice are paired during the use process, and can enter the data pairs with high confidence into the database for subsequent use.

[0064] 5. The present invention not only extracts the lip movement information of the face, calculates the lip shape synchronization score of the face and the voice, but also extracts the face identity information, and calculates the voice appearance matching score of the face and the voice according to the deep connection between the face appearance information and the voiceprint information of the voice, thereby improving the matching accuracy of the voice and the face when the image resolution is low and the lip movement is difficult to identify.

[0065] 6. The present invention comprehensively uses multi-dimensional information, not only uses the connection between the lip movement sequence and the audio content information, but also jointly uses the relationship between the face appearance information and the voiceprint information of the voice. Further improves the matching accuracy, alleviates the matching pressure in the scenario where the lip movement is not clear enough, and can identify to a certain extent whether the speaker in the picture is just lip-syncing. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a module diagram of a speaker tracking system based on multimodal information provided by an embodiment of the present invention;

[0067] Figure 2 It is a working flow chart of the voice identity information feature extraction module provided by an embodiment of the present invention;

[0068] Figure 3 It is a working flow chart of the voice content information feature extraction module provided by an embodiment of the present invention;

[0069] Figure 4 It is a working flow chart of the face appearance information feature extraction module of the image provided by an embodiment of the present invention;

[0070] Figure 5 This is the flowchart of the content feature extraction module for the images provided by the embodiments of the present invention;

[0071] Figure 6 This is the flowchart of the face image quality calculation module provided by the embodiments of the present invention;

[0072] Figure 7 This is the flowchart of the face detection and grouping completion module provided by the embodiments of the present invention;

[0073] Figure 8 This is the flowchart of the lip synchronization module provided by the embodiments of the present invention;

[0074] Figure 9 This is the flowchart of the speaking lip movement recognition module provided by the embodiments of the present invention;

[0075] Figure 10 This is the flowchart of the voice and appearance matching module provided by the embodiments of the present invention;

[0076] Figure 11 This is the flowchart of the prior database provided by the embodiments of the present invention;

[0077] Figure 12 This is the flowchart of a speaker tracking method based on multi-modal information provided by the embodiments of the present invention. Specific embodiments

[0078] The following combines the accompanying drawings and embodiments to further describe in detail the specific embodiments of the present invention.

[0079] This embodiment proposes a speaker tracking method based on multi-modal information, a system that calculates the speaking lip movement score, voice and appearance matching score, and lip synchronization score for each face in the image using the input image and audio information, can score each face in the image, and locate the specific speaker. At the same time, it supports pre-recording and registering paired human voices and faces in advance, and supports entering the paired human voices and faces with high confidence into the prior database during use.

[0080] To achieve the above object, the technical solution adopted in this embodiment is as follows:

[0081] In the first aspect, this embodiment provides a speaker tracking system based on multi-modal information, as Figure 1 shown. The system includes: an identity information feature extraction module for voice, a content information feature extraction module for voice, a facial appearance information feature extraction module for images, a content feature extraction module for images, a face image quality calculation module, a face detection and grouping module, a lip synchronization module, a speaking lip movement recognition module, a voice and appearance matching module, and a prior database.

[0082] The voice identity information feature extraction module, as Figure 2 shown, for the input audio, the network filter bank feature v0 is extracted through the Mel filter bank; the network filter bank feature v0 is input into the ECAPA-TDNN convolutional neural network model, and a 512-dimensional intermediate vector v1 is extracted. The intermediate vector v1 is subjected to L2 regularization, and through 4 fully connected layers, the voice identity information feature vector emb is extracted vid .

[0083] The voice content information feature extraction module, as Figure 3 shown, the intermediate vector v1 is subjected to L2 regularization, and through 5 fully connected layers, a 256-dimensional intermediate vector v2 is obtained; the intermediate vector v2 passes through 2 fully connected layers to obtain a 256-dimensional intermediate vector v3; using residual connection, the intermediate vectors v2 and v3 are added to obtain v4 = v2 + v3, and then through 1 fully connected layer, the voice content information feature vector emb is obtained vct .

[0084] The face information feature extraction module of the image, as Figure 4 shown, sequentially input r input images face 1 …face r into the Inception-V1 convolutional neural network, and a 512-dimensional intermediate vector is extracted and subjected to L2 regularization, and through 4 fully connected layers, a 128-dimensional face feature vector of each image is extracted After processing the r input images, a feature vector z with the shape of (r, 128) will be obtained fid ; each input image face i is separately input into the face image quality calculation module, and the quality score q of each input image is calculated i ;

[0085] The score is between 0 and 1, and the higher the score, the higher the image quality; the image quality is judged by the clarity of the face in the image and whether the pose is frontal, and is used to indicate whether the face in the picture is clear enough and has enough information to extract features;

[0086] The r input images obtain a quality score vector q with the shape of (r, 1); the quality score vector q and the feature vector z fid are concatenated to obtain a vector with the shape of (r, 129) dimensions, and input into the LSTM to obtain a 129-dimensional intermediate vector z1; the intermediate vector z1 passes through 1 fully connected layer to obtain the face feature vector emb that synthesizes the r input images fid .

[0087] The content feature extraction module of the image, such as Figure 5 shown in the figure, splices r input images in the time dimension and keeps other dimensions unchanged, obtaining a vector of size (c, w*r, h). Here, c represents the number of channels of the input image. If the input is a color image, then c = 3; if the input is a grayscale image, then c = 1; r represents the number of input images; w represents the number of pixels in the width of the input image; h represents the number of pixels in the height of the input image. The spliced input image splicing vector is x0;

[0088] Each input image is separately input into the face image quality calculation module to obtain a quality score vector x1 of shape (r, 1). The quality score of each input image is between 0 and 1, and the higher the score, the higher the image quality. The image quality includes image clarity and the face pose in the image.

[0089] The quality score vector x1 of shape (r, 1) is copied and expanded into a quality score vector x2 of shape (1, w*r, h), where x2[1, i, j] = x1[i%w, 1], i ∈ [0, w*r), j ∈ [0, h); the input image splicing vector x0 and the quality score vector x2 are spliced in the first dimension to obtain a feature vector x3 of shape (c + 1, w*r, h);

[0090] The feature vector x3 is input into a 17-layer two-dimensional convolutional network to extract a 128-dimensional feature vector x4; the feature vector x4 is subjected to L2 normalization to obtain the content feature vector emb fct .

[0091] The face image quality calculation module, such as Figure 6 shown in the figure, inputs a single-color face image into the ResNet50 convolutional neural network to obtain an intermediate vector v of 2048 dimensions. The intermediate vector v is input into a fully connected layer and passed through a Sigmoid layer to obtain the image quality score score quality ∈(0, 1);

[0092] In this embodiment, the face detection and grouping module uses the yolo-v5 or s3fd deep learning model for face detection. As Figure 7 shown in the figure, it detects all the faces in each frame of the video segment from time t to time t + s to obtain the matrix information of each face where i represents the i-th face detected in the current frame; j represents the j-th frame. denote the abscissa information of the upper left corner, the ordinate information of the upper left corner, the abscissa information of the lower right corner, and the ordinate information of the lower right corner of the matrix corresponding to the i-th face; group the face matrices of the same person in all frames according to the intersection-over-union ratio of the face matrix information of adjacent frames. If and have an intersection-over-union ratio greater than the set threshold, it is determined that these two face matrices belong to the same person and will be grouped into the same group to obtain the grouped face matrix sequence.

[0093] The lip synchronization module, as Figure 8 shown, inputs the face lip content feature vector emb fct and the speech content information feature vector emb vct , and uses the cosine similarity to calculate the similarity of the two feature vectors, which is the lip synchronization score score ct , where score ct ∈[-1, 1]; the higher the score, the better the match.

[0094] The speaking lip movement recognition module, as Figure 9 shown, inputs the face lip content feature vector emb fct into the fully connected layer with an activation function to obtain a 128-dimensional intermediate vector a1; inputs the intermediate vector a1 into the fully connected layer with a Sigmoid activation function to obtain the speaking lip movement score score talk ∈(0, 1). The higher the speaking lip movement score, the higher the possibility that the face corresponding to the calculated face lip content feature vector is speaking; not speaking can be either silence or chewing, smiling actions;

[0095] The voice and appearance matching module, as Figure 10 shown, inputs the face appearance information feature vector emb fid and the voice identity information feature vector emb vid , and uses the L1 distance to calculate the distance between the two feature vectors, which is the voice and appearance matching score score id ; where, score id ≥0; the smaller the score, the better the match;

[0096] The prior database, as Figure 11 shown, given several face photos and a piece of human voice audio corresponding to the person entered into the database; inputs the given sequence of face photos into the face appearance information feature extraction module of the image to obtain the face appearance information feature vector emb fid corresponding to each person; performs noise reduction processing on the given human voice audio, inputs it into the audio-based identity information feature extraction module, and extracts the voice identity information feature vector emb vid; Save the triple <ID, emb> composed of the personnel number and its vector into the prior database. During the speaker tracking process, prioritize the voice appearance matching based on the prior database; vid , emb fid > into the prior database. During the speaker tracking process, prioritize the voice appearance matching based on the prior database;

[0097] For the automatic update of the prior database during use, when the corresponding vector for the input data is not found in the database, and subsequently, through modules such as lip synchronization, voice appearance matching, and speaking lip movement detection, the matching "voice identity information feature vector emb vid " and "image appearance information feature vector emb fid " are found, then save the triple <ID, emb vid , emb fid > composed of the speaker number and vector with a matching score higher than the input threshold into the prior database;

[0098] The voice identity information feature extraction module and the image appearance information feature extraction module are jointly trained. The training process is as follows: The modules to be trained are the voice identity information feature extraction module Model vid , and the image appearance information feature extraction module Model fid . Input the face image and voice audio of the same person into each module to obtain emb vid and emb fid ; Among them, the 4 fully connected layers of Model vid and the 4 fully connected layers of Model fid share network parameters.

[0099] Use the mean squared error loss function Loss1 as shown in Equation (1):

[0100] Loss1 = MSE(emb fid , emb vid ) (1)

[0101] The voice content information feature extraction module Model vct and the image content information feature extraction module Model fct are jointly trained;

[0102] Specifically: The network parameter values of the ECAPC-TDNN layer in the voice content information feature extraction module Model vct are taken from the network parameters of the voice identity information feature extraction module Model vid . During the training process, these parameter values are fixed and do not change, and do not participate in the parameter update of backpropagation;

[0103] Input the face image sequence and the human voice audio segment corresponding to the speech segment of the same person into Model fct and Model vct respectively, and obtain the lip content feature vector emb of the image fct and the speech content information feature vector emb based on the audio vct ; Input the human voice audio that has no corresponding relationship with the picture sequence into Model vct to obtain the mismatched speech content information feature vector emb′ vct ; In order to make the mutually matching features emb fct and emb vct extracted from the same video close enough, and the mismatched features emb fct and emb′ vct far enough away, calculate the cosine similarity between the two; By maximizing the cosine similarity between emb fct and emb′ vct , and minimizing the similarity between emb fct and emb vct to let the two models learn the content information in the video; The loss function Loss2 is shown in Equation (2):

[0104] Loss2 = CosineSim(emb fct , emb vct ) - CosineSim(emb fct , -emb′ vct ) (2)

[0105] The speech lip movement recognition module is denoted as Model talk , and is trained on the emb fct extracted by the content information feature extraction module of the image;

[0106] Specifically: Input the face image sequence of the person speaking into Model fct to obtain Input the face image sequence of the person not speaking into Model fct to obtain Input and into Model talk to obtain the corresponding speech lip movement scores and Use binary cross-entropy loss to train the model, minimize and maximize The loss function Loss3 is shown in Equation (3):

[0107]

[0108] On the other hand, the present invention provides a speaker tracking method based on multimodal information, which is implemented by using the above-mentioned speaker tracking system based on multimodal information. As Figure 12 shown, it includes the following steps:

[0109] S1: Obtain a video segment from time t to time t + s through a pan-tilt camera, denoted as Obtain an audio segment from time t to time t + s through a microphone or an array microphone, denoted as

[0110] S2: For the audio segment Extract the energy level and zero-crossing rate to determine whether this segment contains human voices; if it does not contain human voices, then no one is speaking from time t to time t + s, t = t + s, and return to S1; if it contains human voices, then input the human voice audio into the voice identity information feature extraction module to obtain the voice identity information feature vector emb vid ; input the human voice audio into the voice content information feature extraction module to obtain the voice content information feature vector emb vct ;

[0111] S3: Input the video segment into the face detection and grouping module to obtain the face matrix information sequence of each person in each frame where i represents the face of the i-th person, and j represents the j-th frame, j ∈ [t, t + s];

[0112] S4: Since the faces in the video may be in a moving state, it cannot be guaranteed that the pictures of each frame are clear enough, and thus there may be a situation where the face detection module fails to recognize some faces in some frames. To address this problem, the linear interpolation method is used to complement the frames lacking face information based on the face matrix information of adjacent frames, obtaining the updated face matrix information sequence

[0113] Specifically: If the face matrices of the i-th person at time j1 and time j2 are detected and and the face of this person is not detected between time j1 and time j2, use the linear interpolation method to obtain the face matrix information corresponding to the i-th person at time k where,

[0114] If the first frame in which the face of the i-th person is detected is at time t first , and t first > t, then tfirst The face matrix information at a moment is assigned to the frames from moment t to moment t first ; if the last frame t final < t + s of the detected face is detected, then the face matrix information at moment t final is used to assign the face matrix information to the frames after moment t final .

[0115] S5: According to the face matrix sequence crop to obtain a face image sequence and input it into the face image quality calculation module to obtain the image quality score corresponding to each frame of the face image Input and into the face appearance information feature extraction module of the image to obtain a sequence of face appearance feature vectors Input and into the lip information feature extraction module of the image to obtain a sequence of face lip content feature vectors

[0116] S6: Retrieve all the recorded voice identity information feature vectors in the database and determine whether there is a vector emb′ vid satisfying L1(emb′ vid, emb vid ) < threshold vid , where L1(*) represents the L1 distance between two vectors and threshold vid is the distance threshold;

[0117] If there exists a recorded voice vector emb′ vid whose L1 distance from the vector emb vid is less than threshold vid , the pairing is successful. If multiple audio pairings are successful, the recorded voice vector with the closest L1 distance is denoted as emb′ vid ; enter S7;

[0118] If there is no recorded voice vector in the prior database whose L1 distance from the vector emb vid is less than the threshold, then enter S8;

[0119] S7: Take out the target face appearance information feature vector corresponding to emb′ vid , denoted as Traverse all the vectors in the face appearance information feature vector sequence in the given image, calculate the L1 distance from the target face appearance information feature vector , and check whether there is a vector satisfying vector, if any, then select the face information corresponding to the face feature vector with the smallest L1 distance between as the marking result; if not, then determine that the speaker is not in the picture;

[0120] Proceed to step S10;

[0121] S8: Sequentially input the of the i-th person in the image into the lip synchronization module to obtain the lip synchronization score vct Input and emb vid into the audio-visual matching module to calculate the audio-visual matching score Input into the speech lip movement recognition module to calculate the speech lip movement score

[0122] Integrate the lip synchronization score, audio-visual matching score, and speech lip movement score, and calculate the final score with weights Compare the final score with the recognition threshold threshold score , if the score of each person's face image sequence is lower than the recognition threshold, then it is determined that there is no face that matches the human voice; if the score of the face image sequence of one or more people is higher than the recognition threshold, then the one with the highest score is recorded as the current speaker;

[0123] S9: If the final score score of the current speaker is higher than the entry threshold threshold record , then register the current speaker number and its corresponding emb vid and emb fid in the prior database;

[0124] S10: t = t + s, return to step S1.

Claims

1. A speaker tracking system based on multimodal information, characterized in that: The system includes: a voice identity information feature extraction module, a voice content information feature extraction module, a face appearance information feature extraction module for images, a content feature extraction module for images, a face image quality calculation module, a face detection and grouping module, a lip synchronization module, a speaking lip movement recognition module, a voice and appearance matching module, and a prior database; Using the voice identity information feature extraction module, extract the voice identity information feature vector from the input audio; Using the voice content information feature extraction module, extract the voice content information feature vector from the input audio; Adopt the facial appearance information feature extraction module of the image, and sequentially input r input images … Extract the frame-by-frame facial appearance feature vectors, input each image into the facial image quality calculation module, calculate the quality score of each input image, splice the quality scores of the r images and the r frame-by-frame facial appearance feature vectors in the channel dimension, and extract the facial appearance feature vectors; Using the content feature extraction module for images, splice r input images in the time dimension to obtain a spliced image vector; input each input image separately into the face image quality calculation module to obtain the quality score of each input image, copy and expand the quality score of each input image and splice it with the image vector for splicing and feature extraction to obtain the face lip content feature vector; The face image quality calculation module inputs a single-color face image into a convolutional neural network to obtain an image quality score; The face detection and grouping module detects faces in the video segment frame by frame, gives the matrix information of each face, groups the face matrices belonging to the same person into one group, and completes the face information filling for the frames lacking face information to obtain a complete sequence of face matrices; The lip synchronization module inputs the face lip content feature vector and the voice content information feature vector, and uses the cosine similarity to calculate the similarity of the two feature vectors to obtain a lip synchronization score; The speaking lip movement recognition module inputs the face lip content feature vector into one or more fully connected layers with activation functions, and then inputs it into a fully connected layer with a Sigmoid activation function to obtain a speaking lip movement score; The voice and appearance matching module inputs the face appearance information feature vector and the voice identity information feature vector, and uses the L1 distance to calculate the distance between the two feature vectors to obtain a voice and appearance matching score; The prior database supports pre-entering the prior database and entering the prior database during use, and preferentially uses the prior database for matching during the speaker tracking process. Specifically: When the matching voice identity information feature vector and face appearance information feature vector are found according to the lip synchronization, voice and appearance matching, and speaking lip movement detection modules, save the vector pairs with matching scores higher than the entry threshold into the prior database; The speaker tracking system based on multi-modal information implements speaker tracking using the following method, including the following steps: S1: Obtain audio and video, and respectively obtain the audio segment and video segment from time t to time t + s using an audio acquisition device and a video acquisition device; S2: Human voice judgment and speech feature extraction, determine whether the audio segment contains a human voice; if it does not contain a human voice, it is determined that no one is speaking from time t to time t + s, and proceed to S9; if it contains a human voice, input the audio segment into the speech identity information feature extraction module to obtain a speech identity information feature vector ; and input the audio segment into the speech content information feature extraction module to obtain a speech content information feature vector ; S3: Facial image sequence extraction. Input the video clip frame by frame into the face detection and grouping module to obtain a facial image sequence ; S4: Image feature extraction. Input the face image sequence into the face image quality calculation module to obtain the image quality score corresponding to each frame of the face image . Input and into the facial feature information extraction module of the image to obtain a sequence of facial feature vectors . Input and into the content feature extraction module of the image to obtain the facial lip content feature vector . S5: Retrieve all the voice identity information feature vectors entered in the prior database and determine whether there is an entered voice similar to the voice identity information feature vector of the entered voice; If there exists an input voice vector similar to , then proceed to S6; If there is no recorded human voice similar to the voice identity information feature vector then proceed to S7; S6: Retrieve the target face feature vector corresponding to , and search for whether there is a feature vector in the candidate sequence of face information feature vectors in the given image whose similarity is higher than the matching threshold . If there is, mark and output the corresponding face matrix sequence information. If not, determine that there is no face in the current frame that matches the corresponding voice, and proceed to S9;​ S7: Sequentially input the i individual's paired with into the lip synchronization module to obtain the lip synchronization score ; Input paired with into the voice and appearance matching module to calculate the voice and appearance matching score ; Input into the speaking lip movement recognition module to calculate the speaking lip movement score ; Based on the comprehensive lip - shape synchronization score, voice - appearance matching score, and speaking lip - movement score, calculate the final score by weighted calculation ; Compare the final score with the recognition threshold , if the scores of the face image sequences of each person are all lower than the recognition threshold, it is determined that there is no face that matches the human voice, and go to S9; if the scores of the face image sequences of one or more people are higher than the recognition threshold, then the one with the highest score is recorded as the current speaker; S8: If the final score of the current speaker is higher than the input threshold , then the corresponding to the current speaker and will be registered in the prior database; S9: t = t + s, return to step S1.

2. The speaker tracking system based on multi-modal information according to claim 1, wherein: The identity information feature extraction module for the voice is specifically as follows: for the input audio, the network filter bank features are extracted through a Mel filter ; the network filter bank features are input into the first convolutional neural network, and the -dimensional intermediate vector is extracted. The L2 regularization is performed on the intermediate vector , and through fully connected layers, the voice identity information feature vector is extracted; The content information feature extraction module for the speech is specifically: regularize the intermediate vector by L2 regularization, and obtain, through fully connected layers, an intermediate vector with dimensions; pass the intermediate vector through fully connected layers to obtain an intermediate vector with dimensions; use a residual connection to add the intermediate vectors and to obtain , and then pass it through fully connected layers to obtain the speech content information feature vector .

3. The speaker tracking system based on multi-modal information according to claim 1, wherein: The appearance information feature extraction module of the image is specifically as follows: sequentially input r input images … into the second convolutional neural network, and extract the -dimensional intermediate vector and perform L2 regularization. Pass through fully connected layers to extract the -dimensional feature vector ; after processing the r input images, a feature vector with a shape of (r, ) will be obtained ; input each input image individually into the face image quality calculation module to calculate the quality score of each input image ; The quality score vector with a shape of (r, 1) is obtained from r input images ; The quality score vector and the feature vector are concatenated to obtain a vector with a shape of (r, +1), which is input into a recurrent neural network to calculate and obtain +1-dimensional intermediate vector ; The intermediate vector passes through fully connected layers to obtain a face feature vector that synthesizes r input images ; The content feature extraction module of the said image is specifically: splicing r input images in the time dimension while keeping other dimensions unchanged, to obtain a vector of size (c, w*r, h), where c represents the number of channels of the input image. If the input is a color image, then c = 3; if the input is a grayscale image, then c = 1; r represents the number of input images; w represents the number of pixels of the width of the input image; h represents the number of pixels of the height of the input image. The spliced input image splicing vector is ; Each input image is separately input into the face image quality calculation module to obtain a quality score vector with a shape of (r, 1). ; Copy and expand the mass score vector with shape (r, 1) to a mass score vector with shape (1, w * r, h); Concatenate the input image concatenation vector and the mass score vector in the first dimension to obtain a feature vector with shape (c + 1, w * r, h). ; Input the feature vector into the third convolutional neural network, and extract a feature vector with dimensions, denoted as ; perform L2 normalization on the intermediate vector to obtain the face lip content feature vector .

4. The speaker tracking system based on multi-modal information according to claim 1, wherein: The face image quality calculation module inputs a single-color face image into the fourth convolutional neural network to obtain an intermediate vector v of dimension, and inputs this intermediate vector into a fully connected layer with a Sigmoid activation function to obtain an image quality score ; The face detection and grouping module uses a deep learning algorithm to detect all faces in each frame of the video clip, obtaining the matrix information of each face , representing the matrix information of the -th face detected in the -th frame; grouping the face matrices of the same person in all frames according to the intersection over union of the face matrix information of adjacent frames. If the and intersection over union is greater than the set threshold, it is determined that these two face matrices belong to the same person and will be grouped into the same group; using linear interpolation to complete the frames with missing face information according to the face matrix information of adjacent frames; cropping the completed face matrix sequence to obtain the face image sequence .

5. The speaker tracking system based on multi-modal information according to claim 1, wherein: The lip synchronization module inputs the feature vector of the lip content of the human face and the feature vector of the speech content information , and calculates the similarity of the two feature vectors using the cosine similarity, which is the lip synchronization score , where ; the higher the score, the better the match; The speech lip movement recognition module inputs the content feature vector of the human face lip part into a fully connected layer with an activation function to obtain an intermediate vector of dimensions; the intermediate vector is input into a fully connected layer with a Sigmoid activation function to obtain a speech lip movement score . The higher the speech lip movement score, the higher the possibility that the face corresponding to the calculated content feature vector of the human face lip part is speaking.

6. The speaker tracking system based on multimodal information according to claim 1, wherein: The audio-visual matching module inputs the feature vector of the facial appearance information and the feature vector of the voice identity information , and uses the L1 distance to calculate the distance between the two feature vectors, which is the audio-visual matching score ; where ; the smaller the score, the better the match 7. The speaker tracking system based on multimodal information according to claim 1, wherein: The prior database pre - inputs several face photos, human voice audio corresponding to personnel. The face photo sequence is input into the face information feature extraction module of the image to obtain the face information feature vector corresponding to each personnel. , the human voice audio is denoised and input into the voice identity information feature extraction module to extract the voice identity information feature vector corresponding to each personnel. , the vectors and are saved into the prior database; during the speaker tracking process, the voice - face matching based on the prior database is preferentially performed.

8. The speaker tracking system based on multimodal information according to claim 1, wherein: The identity information feature extraction module of the voice and the face information feature extraction module of the image are jointly trained. The training process is as follows: The face pictures and voice audio of the same person are respectively input into and to obtain and , and a mean square error loss function is established; The content information feature extraction module of the speech and the content information feature extraction module of the image are co-trained; Specifically: the content information feature extraction module of the speech All network parameters of the first convolutional neural network in come from the parameters of the first convolutional neural network of the identity information feature extraction module of the speech, and the numerical values of these parameters are fixed during the training process and do not participate in the parameter update during the backpropagation process; Input the face image sequence and the human voice audio segment corresponding to the speech segment of the same person into the content information feature extraction module of the image and the content information feature extraction module of the speech respectively to obtain the face lip content feature vector of the image and the speech content information feature vector based on the audio ; Input the human voice audio that has no corresponding relationship with the picture sequence into the content information feature extraction module of the speech to obtain the mismatched speech content information feature vector ; By minimizing and the cosine similarity between, and maximizing and the cosine similarity between, let the two models learn the content information in the video to obtain the loss function ; The speech lip movement recognition module is denoted as , and is trained on the extracted by the content information feature extraction module of the image; Specifically: input the sequence of face images of the person speaking into the content information feature extraction module of the image to obtain ; input the sequence of face images of the person not speaking into the content information feature extraction module of the image to obtain ; input and into to obtain the corresponding speaking lip movement scores and ; train the model using binary cross-entropy loss to minimize and maximize to obtain the loss function .

Citation Information

Patent Citations

  • Picture tracking method and device for video conference

    CN111263106A

  • Conference speaker tracking method and device, computer equipment and storage medium

    CN112040119A

  • Conference speaker tracking method and device, computer equipment and storage medium

    CN112633219A

  • Voice and face image matching method and device, storage medium and electronic equipment

    CN111507218A

  • Multi-modal conference data structuring method and device and computer equipment

    CN114298170A