Audio and Video Person Recognition Method and System Based on Multimodal Biometric Consistency

Through the audio and video character recognition method with multimodal biometric consistency, combined with face, gait and voiceprint recognition, the problem of poor recognition effect in complex scenarios is solved, and higher identity recognition accuracy and applicability are achieved.

CN116612542BActive Publication Date: 2025-07-22XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310571748.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-07-22
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

Traditional character identity recognition methods mainly rely on single modal information, especially in complex scenarios, with limited recognition effects and difficulty in identifying objects wearing hats and other obstructions.

Method used

The audio and video character recognition method with multimodal biometric consistency is adopted, and the face area and human body area are extracted through a face detector, combined with face recognition, gait recognition and voiceprint recognition, and multimodal screening and consistency scoring methods are used to comprehensively use face features, gait features and voiceprint features for identity recognition.

Benefits of technology

It improves the accuracy of character identity recognition in complex scenarios, can effectively identify objects such as hats, and is suitable for scenarios such as community security, public safety management and smart homes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612542B_ABST
    Figure CN116612542B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio-visual person recognition method and system based on multi-modal biometric consistency, which relates to the field of person identity recognition. The present invention uses face detector and human body detector technologies to extract the face region and the human body region, and uses foreground and background separation technology to obtain the human body silhouette from the human body region; at the same time, deep learning technology is used to extract face features from the face region using face recognition, extract gait features from the human body region using gait recognition, and extract voiceprint features from audio frames using voiceprint recognition; furthermore, a novel multi-modal screening method and a multi-modal consistency scoring method are used to efficiently utilize multi-modal information including face features, gait features, and voiceprint features to more accurately identify the person's identity. And the method of the present invention is particularly suitable for use in complex scenarios such as community security, public safety management, and smart home scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of person identity recognition, and particularly to an audio - video person recognition method and system based on multi - modal biometric consistency. Background Art

[0002] Traditional person identity recognition methods mainly focus on visual information, mainly face recognition, which is related to clothing and body posture, limited to single - modal recognition, and generally have the following problems: (1) Single - modal information is limited, the information utilization efficiency is low, and the requirements for the recognition scene are high: Current pedestrian recognition algorithms mainly perform recognition based on single - modal information (such as features of image color, texture, depth, etc.). However, single - modal information has limitations and cannot comprehensively reflect the appearance and features of pedestrians. The recognition effect is limited in complex scenes, and different recognition scenes have different requirements, posing a great challenge to the generalization ability of the algorithm; (2) It is difficult to recognize objects with occlusions such as wearing hats: Due to the influence of external environment and personal privacy and other factors, pedestrians often wear hats, masks and other occlusions, which makes it difficult for the recognition algorithm to obtain complete pedestrian image information, thus reducing the recognition effect. Summary of the Invention

[0003] In view of the problems raised in the above - mentioned background art, the present invention provides an audio - video person recognition method and system based on multi - modal biometric consistency to improve the accuracy of person identity recognition in complex scenes.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] On the one hand, the present invention provides an audio - video person recognition method based on multi - modal biometric consistency, including:

[0006] Obtain the audio - video stream to be recognized for identity and perform pre - processing to separate the video stream data and the audio stream data;

[0007] For each frame of data in the video stream data, use a face detector to extract the face region and the corresponding face key points, and use a human body detector to extract the human body region corresponding to the face region within a time window before and after the frame;

[0008] Use a face recognition network to extract the face features of the face region and extract the gait features of the human body region;

[0009] For each frame of data in the audio stream data, extract the voiceprint features within a time window before and after the frame;

[0010] Perform multi - modal screening on the extracted face features, gait features and voiceprint features to obtain a set of candidate persons;

[0011] Perform multimodal consistency scoring for each person in the set of candidate persons, and return the identity of the person with the highest score as the identified person identity;

[0012] Perform identity annotation on the person in each frame of the audio-video stream according to the identified person identity, and output the audio-video stream after identity recognition.

[0013] Optionally, the extraction of the gait features of the human body region specifically includes:

[0014] Input the human body region corresponding to the face region into the foreground-background separation network, and output a sequence of human body silhouettes;

[0015] Input the sequence of human body silhouettes into the gait recognition network, and output the extracted gait features.

[0016] Optionally, for each frame of data in the audio stream data, the extraction of the voiceprint features within a time window before and after the frame specifically includes:

[0017] For each frame of data in the audio stream data, convert the sound signal sequence within a time window before and after the frame into a Mel spectrogram and perform MFCC feature extraction to extract the corresponding speech features;

[0018] Input the speech features into the speech recognition network to extract the corresponding voiceprint features.

[0019] Optionally, the multimodal screening of the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons specifically includes:

[0020] Calculate the cosine similarity between the extracted face features and each face feature in the face database, sort the multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_face1, C_face2,..., C_face K And the corresponding person identities;

[0021] Calculate the cosine similarity between the extracted gait features and each gait feature in the gait database, sort the multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_gait1, C_gait2,..., C_gait K And the corresponding person identities;

[0022] Calculate the cosine similarity between the extracted voiceprint features and each voiceprint feature in the voiceprint database, sort the multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_voice1, C_voice2,..., C_voice K And the corresponding person identities;

[0023] Take the union of the top K results returned by each of the three modalities of face features, gait features, and voiceprint features to obtain a set M of candidate persons.

[0024] Optionally, perform multi-modal consistency scoring on each person in the set of candidate persons, and return the identity of the person with the highest score as the identified person's identity, specifically including:

[0025] For the k-th person M in the set of candidate persons M k , compare the cosine similarity between its face features and gait features, and take the modality with the higher cosine similarity as the k basic modality of M, and use the cosine similarity value corresponding to the basic modality as the basic modality score Score_base k ;

[0026] Calculate the consistency score w between the face and gait based on the face region and the corresponding human body region f,g ;

[0027] Calculate the consistency score w between the face and voiceprint based on the face key points and the Mel spectrogram f,v ;

[0028] Record the consistency score between the gait and the voiceprint as w g,v ;

[0029] According to the consistency scores w f,g , w f,v and w g,v calculate the modality consistency score Score_coin under different basic modalities k ;

[0030] According to the basic modality score Score_base k and the modality consistency score Score_coin k calculate the total score Score k of the k-th person M k = Score_base k + Score_coin k ;

[0031] Return the identity of the person with the highest total score Score k as the identified person's identity.

[0032] On the other hand, the present invention provides an audio-visual person recognition system based on multi-modal biometric consistency, including:

[0033] A preprocessing module for obtaining an audio-visual stream to be identified and performing preprocessing to separate the video stream data and the audio stream data;

[0034] A face and human body region extraction module, which is used for each frame of data in the video stream data, to extract the face region and the corresponding face key points by using a face detector, and to extract the human body region corresponding to the face region within a time window before and after the frame by using a human body detector;

[0035] A face and gait feature extraction module, which is used to extract the face features of the face region by using a face recognition network and to extract the gait features of the human body region;

[0036] A voiceprint feature extraction module, which is used for each frame of data in the audio stream data to extract the voiceprint features within a time window before and after the frame;

[0037] A multi-modal screening module, which is used to perform multi-modal screening on the extracted face features, gait features and voiceprint features to obtain a set of candidate persons;

[0038] A multi-modal consistency scoring module, which is used to perform multi-modal consistency scoring on each person in the set of candidate persons and return the person identity with the highest score as the recognized person identity;

[0039] A person identity annotation module, which is used to annotate the identity of the person on each frame in the audio-video stream according to the recognized person identity and output the audio-video stream after identity recognition.

[0040] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio-video person recognition method based on multi-modal biometric consistency as described above is implemented.

[0041] Optionally, the memory is a non-transitory computer-readable storage medium.

[0042] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0043] The present invention provides an audio-video person recognition method and system based on multi-modal biometric consistency, which uses face detector and human body detector technologies to extract face regions and human body regions, and uses foreground-background separation technology to obtain human silhouettes from human body regions; at the same time, deep learning technology is used to extract face features from face regions by using face recognition, to extract gait features from human body regions by using gait recognition, and to extract voiceprint features from audio frames by using voiceprint recognition; further, a novel multi-modal screening method and a multi-modal consistency scoring method are used to efficiently utilize multi-modal information including face features, gait features and voiceprint features to more accurately identify person identities. And the method of the present invention is particularly suitable for use in complex scenarios such as community security, public safety management and smart home. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of a method for audio-visual person recognition based on multi-modal biometric consistency of the present invention;

[0046] Figure 2 It is a schematic diagram of the principle of a method for audio-visual person recognition based on multi-modal biometric consistency of the present invention;

[0047] Figure 3 It is a schematic diagram of the multi-modal screening process of a method for audio-visual person recognition based on multi-modal biometric consistency of the present invention. Specific embodiments

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0049] The purpose of the present invention is to provide a method and system for audio-visual person recognition based on multi-modal biometric consistency to improve the accuracy of person identity recognition in complex scenarios.

[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0051] Figure 1 and Figure 2 are respectively a flowchart and a schematic diagram of the principle of a method for audio-visual person recognition based on multi-modal biometric consistency of the present invention. Refer to Figure 1 and Figure 2 , a method for audio-visual person recognition based on multi-modal biometric consistency, includes:

[0052] Step 1: Obtain the audio-visual stream to be recognized for identity and perform preprocessing to separate the video stream data and the audio stream data.

[0053] Preprocess the input audio-visual stream to be identified for identity, including separating video stream data and audio stream data. Assume that there are n individuals with different behaviors in the current audio-visual stream scene, denoted as P1, P2, P3, ..., P n .

[0054] Step 2: For each frame of data in the video stream data, use a face detector to extract the face region and corresponding face key points, and use a human body detector to extract the human body region corresponding to the face region within a time window before and after the frame.

[0055] Taking the i-th frame in the video stream data as an example, use a face detector to detect m face regions in the i-th frame, namely F1, F2, F3, …, F m ; Use a human body detector to detect o human body regions that appear in the i-th frame, denoted as B1, B2, B3, …, B o .

[0056] Both the face detector and the human body detector can be trained using the yolov3 network. The difference lies in the different training sample sets used. The input of the face detector is video frame data, and the output is the face region in the video frame; the input of the human body detector is video frame data, and the output is the human body region in the video frame.

[0057] Step 3: Use a face recognition network to extract the face features of the face region and extract the gait features of the human body region.

[0058] The present invention uses a face recognition network to extract the face features of the face region, and inputs the human body region corresponding to the face region into a foreground-background separation network to output a human body silhouette sequence, and then inputs the human body silhouette sequence into a gait recognition network to output the extracted gait features. Among them, the network types of the foreground-background separation network and the gait recognition network can both be convolutional neural networks, and are trained using different training sample sets.

[0059] Traverse each face region in the i-th frame. Taking the x-th face as an example, crop the face region F x , and send it into a face recognition network and a face key point detection network respectively. Use a feature extraction algorithm to create a facial embedding face-embeding, representing the face feature vector f_face of a face x , and obtain the face key points landmark through the face key point detection network x ; For a time window W before and after this frame (the maximum length of this sliding window is 31 frames, the length before and after this frame is 15, and if not, it is filled with 0, and the sliding step size is 1), for the human body region B corresponding to the face region F x x ​Perform cropping and input it into the foreground and background separation network to obtain a sequence of human silhouettes \(W_{sil}\) of the same person within a time window. x =(S i-15 ,S i-14 ,...,S i ,…,,S i+14 ,S i+15 ). Input the sequence of human silhouettes \(W_{sil}\) x into the gait recognition network to obtain the gait feature \(f_{gait}\) x .

[0060] Step 4: For each frame of data in the audio stream data, extract the voiceprint features within a time window before and after the frame.

[0061] For each frame of data in the audio stream data, convert the sound signal sequence within a time window before and after the frame into a Mel spectrogram and perform MFCC feature extraction to extract the corresponding speech features; input the speech features into the speech recognition network to extract the corresponding voiceprint features. The speech recognition network can be trained using a convolutional neural network.

[0062] Specifically, convert the sound signal sequence \(W_{audio}\) x =(A i-15 ,A i-14 ,...,A i ,…,,A i+14 ,A i+15 ) within a time window \(W\) before and after the \(i\)-th frame into a Mel spectrogram MFCC i , perform MFCC feature extraction on it, and denote the extracted corresponding speech features as \(f_{audio}\) x ; input the speech features \(f_{audio}\) x into the speech recognition network to obtain the voiceprint feature \(f_{voice}\) x .

[0063] Step 5: Perform multi-modal screening on the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons.

[0064] The person database pre-established in the present invention includes: a face database with \(N_{Face}\) person face features face1, face2, …, face N_Face , a gait database with \(N_{Gait}\) person gait features gait1, gait2, …, gati N_Gait , and a voiceprint database with \(N_{Voice}\) person voiceprint features voice1, voice2, …, voice N_Voice .

[0065] Such as Figure 3As shown, the obtained face feature f_face x , gait feature f_gait x and voiceprint feature f_voice x are respectively matched with the modal features stored in the background person database of the corresponding modality, the cosine value of the included angle between the two feature vectors is calculated, and the cosine similarities C_face1, C_face2, …, C_face N_Face of the face feature modality, the cosine similarities C_gait1, C_gait2, …, C_gait N_Gait of the gait feature modality, and the cosine similarities C_voice1, C_voice2, …, C_voice N_Voice of the voiceprint feature modality are obtained respectively.

[0066] Calculate the cosine similarity between the extracted face feature f_face x and each face feature face1, face2, …, face N_Face in the face database. Sort the multiple cosine similarities C_face1, C_face2, …, C_face N_Face from high to low by value, and return the top K cosine similarity values C_face1, C_face2, …, C_face K and the corresponding person identities.

[0067] Calculate the cosine similarity between the extracted gait feature f_gait x and each gait feature gait1, gait2, …, gati N_Gait in the gait database. Sort the multiple cosine similarities C_gait1, C_gait2, …, C_gait N_Gait from high to low by value, and return the top K cosine similarity values C_gait1, C_gait2, …, C_gait K and the corresponding person identities.

[0068] Calculate the cosine similarity between the extracted voiceprint feature f_voice x and each voiceprint feature voice1, voice2, …, voice N_Voice in the voiceprint database. Sort the multiple cosine similarities C_voice1, C_voice2, …, C_voice N_Voice from high to low by value, and return the top K cosine similarity values C_voice1, C_voice2, …, C_voice K and the corresponding person identities.

[0069] The method for calculating the cosine similarity is as follows: normalize the feature vectors of each modality respectively; calculate the cosine value of the included angle between the two feature vectors as their cosine similarity.

[0070] Take the union of the top K results returned by each of the three modalities of face features, gait features, and voiceprint features to obtain the candidate person set M. That is, sort the cosine similarity values of each modality from high to low, and take the union of the top K persons obtained from each modality to form a candidate person set M with N_K persons = (M1, M2,..., M N_K )

[0071] Step 6: Perform multi-modal consistency scoring on each person in the candidate person set, and return the person identity with the highest score as the identified person identity.

[0072] The scoring rule of the multi-modal consistency scoring of the present invention is divided into a modality basic score and a modality consistency score. Since the confidence levels of face features and gait features are high, when setting the basic modalities, only these two modalities of face features and gait features are considered.

[0073] Only considering the face feature modality and the gait feature modality, select the modality with a higher cosine similarity as the basic modality, and use the cosine similarity corresponding to the basic modality as the modality basic score. When the cosine similarity value corresponding to a certain modality is greater than 0, it means that there is data for this modality and it can be added to the calculation of the modality consistency score. When the k-th candidate M in the candidate person set M k is selected for more than two of the face feature modality, gait feature modality, and voiceprint feature modality at the same time, that is, M k corresponding face cosine similarity C_face k , gait cosine similarity C_gait k , cosine similarity C_voice k are more than two greater than 0, the modality consistency score is increased. The specific calculation method of the modality consistency score is as follows:

[0074] 1) When the selected basic modality is the face feature modality:

[0075] ① If the face feature modality, gait feature modality, and voiceprint feature modality are all selected, the modality consistency score is:

[0076] Score_coin k = w f,g × C_gait k + w f,v × C_voice k ;

[0077] ② Only when the face feature modality and the gait feature modality are selected, the modality consistency score is:

[0078] Score_coin k = w f,g ×C_gait k ;

[0079] 2) When the selected basic modality is the gait feature modality, the modality consistency score includes:

[0080] ① Only when the gait and the face feature modality are selected, the modality consistency score is:

[0081] Score_coin k = w f,g ×C_face k ;

[0082] ② Only when the gait and the voiceprint feature modality are selected, the modality consistency score is:

[0083] Score_coin k = w g,v ×C_voice k .

[0084] Among them, w f,v is the face and voiceprint consistency score, defined as the relationship between the MFCC energy (voice amplitude) of each frame of sound and the opening of the mouth. Using the face key points landmark, when the mouth is detected to be closed but the MFCC amplitude is high, it means that this person is not speaking, and the score is 0, otherwise the score is 1; w g,v is the gait and voiceprint consistency score. Since there is no obvious connection between the walking posture and the person's voice, this item is set to 0; w f,g is the face and gait consistency score, defined as the proximity level between the face area and the human body area corresponding to the gait, with a value in (0, 1]. The closer the two areas are, the closer the score is to 1.

[0085] Therefore, step 6 specifically includes:

[0086] Step 6.1: For the kth person M in the candidate person set M k , compare the cosine similarities C_face k and C_gait k of its face features and gait features, and take the modality with the higher cosine similarity as the basic modality of M k , and use the cosine similarity value corresponding to the basic modality as the basic modality score Score_base k ;

[0087] Step 6.2: Calculate the face and gait consistency score w f,g ;

[0088] Specifically, obtain the face region F of the i-th frame x The center point is (X face , Y face ); the Mel spectrogram of the corresponding audio frame is MFCC i ; obtain the human body region B x The center point of which is (X body , Y body ), and the distances from this center point to the left and right, upper and lower boundaries of the human body region are denoted as D X , D Y ; then the consistency score between the face and gait is denoted as:

[0089]

[0090] Step 6.3: Calculate the consistency score w between the face and voiceprint according to the face key points and the Mel spectrogram f,v ;

[0091] Denote the opening and closing state of the lip key points in the face key points landmark x as State lips , where the value of 1 indicates that the lip key points are open, and the value of 0 indicates that the lip key points are closed. Then the consistency score between the face and voiceprint is denoted as:

[0092]

[0093] Step 6.4: Denote the consistency score between the gait and voiceprint as w g,v ;

[0094] Since there is no obvious relationship between the walking posture and the human voice, the consistency score between the gait and voiceprint is denoted as:

[0095] w g,v = 0;

[0096] Step 6.5: Calculate the modal consistency score Score_coin under different basic modalities according to the consistency scores w f,g , w f,v and w g,v ; k ;

[0097] ① When the basic modality is the face feature modality:

[0098]

[0099] ② When the basic modality is the gait feature modality:

[0100]

[0101] Step 6.6: According to the basic modality score Score_base kAnd the Score_coin consistent with the modality k Calculate the total score Score of the k-th person M k k = Score_base k k + Score_coin k ;

[0102] Step 6.7: Return the total score Score k The identity of the person with the highest score is used as the identified person's identity

[0103] After traversing the set M to be selected, the total score set S=(Score1, Score2,..., Score N_K ) of each person is obtained. Sort in descending order, and take the identity of the person with the highest score as the identity of the x-th person in the i-th frame

[0104] Step 7: Perform identity annotation on the people in each frame of the audio-visual stream according to the identified person's identity, and output the audio-visual stream after identity recognition

[0105] The input of the method of the present invention is a multi-person audio-visual stream to be identified, and the output is a video with identity annotation for each person in each frame of the audio-visual stream, which can be used for pedestrian video person recognition, but the applicable scenarios of the present invention are not limited to this

[0106] Based on the method provided by the present invention, the present invention also provides an audio-visual person recognition system based on multi-modal biometric consistency, including:

[0107] A preprocessing module, used to obtain the audio-visual stream to be identified and perform preprocessing, and separate the video stream data and the audio stream data

[0108] A face and human body region extraction module, used for each frame of data in the video stream data, using a face detector to extract the face region and the corresponding face key points, and using a human body detector to extract the human body region corresponding to the face region within a time window before and after the frame

[0109] A face and gait feature extraction module, used to extract the face features of the face region using a face recognition network, and extract the gait features of the human body region

[0110] A voiceprint feature extraction module, used for each frame of data in the audio stream data, to extract the voiceprint features within a time window before and after the frame

[0111] A multi-modal screening module, used to perform multi-modal screening on the extracted face features, gait features and voiceprint features to obtain a set of candidate people

[0112] The multi-modal consistency scoring module is used to perform multi-modal consistency scoring on each person in the set of candidate persons, and return the identity of the person with the highest score as the identified person's identity;

[0113] The person identity annotation module is used to annotate the identity of the person on each frame in the audio-visual stream according to the identified person's identity, and output the audio-visual stream after identity recognition.

[0114] Furthermore, the present invention also provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the computer program in the memory to execute the audio-visual person recognition method based on multi-modal biometric consistency.

[0115] In addition, when the computer program in the above-mentioned memory is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-transitory computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.

[0116] The present invention integrates face feature information belonging to visual information, walking gait feature information unique to the human body, and voiceprint feature information belonging to auditory information; at the same time, a novel modality screening method and a multi-modal fusion consistency scoring method are used to efficiently utilize visual and auditory information, realize multi-modal information complementarity, and improve the accuracy and robustness of identity recognition. The present invention can quickly and accurately identify the identities of different persons in a multi-person audio-visual stream, has broad application value, especially in the fields of community security, public safety management, and smart home, and has extremely high practical value and economic and social benefits.

[0117] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0118] In this article, specific examples are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. An audio-visual person recognition method based on multimodal biometric consistency, characterized in that, Including: Obtain the audio-visual stream of the identity to be recognized and perform preprocessing to separate the video stream data and the audio stream data; For each frame of data in the video stream data, use a face detector to extract the face region and the corresponding face key points, and use a human body detector to extract the human body region corresponding to the face region within a time window before and after the frame; Use a face recognition network to extract the face features of the face region and extract the gait features of the human body region; The extraction of the gait features of the human body region specifically includes: Input the human body region corresponding to the face region into a foreground-background separation network to output a human body silhouette sequence; Input the human body silhouette sequence into a gait recognition network to output the extracted gait features; For each frame of data in the audio stream data, extract the voiceprint features within a time window before and after the frame; Perform multi-modal screening on the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons; The multi-modal screening of the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons specifically includes: Calculate the cosine similarity between the extracted face features and each face feature in the face database, sort the multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_face1, C_face2,..., C_face K and the corresponding person identities; Calculate the cosine similarity between the extracted gait features and each gait feature in the gait library, sort multiple cosine similarities in descending order of values, and return the top K cosine similarity values C_gait1, C_gait2,..., C_gait K and the corresponding person identities; Calculate the cosine similarity between the extracted voiceprint features and each voiceprint feature in the voiceprint database, sort multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_voice1, C_voice2,..., C_voice K and the corresponding person identities; Take the union of the top K results returned by each of the three modalities of face features, gait features, and voiceprint features to obtain a set of candidate persons M; Perform multi-modal consistency scoring on each person in the set of candidate persons, and return the person identity with the highest score as the recognized person identity; The multi-modal consistency scoring of each person in the set of candidate persons and returning the person identity with the highest score as the recognized person identity specifically includes: For the k-th person M in the set M of candidates to be selected k , compare the cosine similarity of their facial features and gait features, and take the modality with the higher cosine similarity as the basis modality of M k , and take the cosine similarity value corresponding to the basis modality as the basis modality score Score_base k ; Calculate the consistency score w of face and gait based on the face region and the corresponding human body region f,g ; Calculate the consistency score w of face and voiceprint based on facial key points and Mel spectrogram f,v ; Denote the consistency score of gait and voiceprint as w g,v ; According to the consistency scores w f,g , w f,v and w g,v calculate the modal consistency score Score_coin for different basic modes k ; According to the basic modal score Score_base k and the modal consistency score Score_coin k Calculate the total score Score of the k-th character M k ; k = Score_base k + Score_coin k ; Return the total score Score k The identity of the person with the highest score is used as the identified person's identity; Perform identity annotation on the person in each frame of the audio-visual stream according to the recognized person identity, and output the audio-visual stream after identity recognition.

2. The method for audio-visual person recognition based on multimodal biometric consistency according to claim 1, wherein The extraction of the voiceprint features within a time window before and after the frame for each frame of data in the audio stream data specifically includes: For each frame of data in the audio stream data, convert the sound signal sequence within a time window before and after the frame into a Mel spectrogram and perform MFCC feature extraction to extract the corresponding speech features; Input the speech features into a speech recognition network to extract the corresponding voiceprint features.

3. An audio-visual person recognition system based on multimodal biometric consistency, characterized in that, Including: A preprocessing module for obtaining the audio-visual stream of the identity to be recognized and performing preprocessing to separate the video stream data and the audio stream data; A face and human body region extraction module for, for each frame of data in the video stream data, using a face detector to extract the face region and the corresponding face key points, and using a human body detector to extract the human body region corresponding to the face region within a time window before and after the frame; A face and gait feature extraction module for using a face recognition network to extract the face features of the face region and extracting the gait features of the human body region; The extraction of the gait features of the human body region specifically includes: Input the human body region corresponding to the face region into a foreground-background separation network to output a human body silhouette sequence; Input the human body silhouette sequence into a gait recognition network to output the extracted gait features; A voiceprint feature extraction module for, for each frame of data in the audio stream data, extracting the voiceprint features within a time window before and after the frame; A multi-modal screening module for performing multi-modal screening on the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons; Performing multi-modal screening on the extracted face features, gait features, and voiceprint features to obtain a set of candidate persons, specifically including: Calculate the cosine similarity between the extracted face features and each face feature in the face database, sort multiple cosine similarities in descending order of values, and return the top K cosine similarity values C_face1, C_face2,..., C_face K and the corresponding person identities; Calculate the cosine similarity between the extracted gait features and each gait feature in the gait library, sort multiple cosine similarities in descending order of value, and return the top K cosine similarity values C_gait1, C_gait2,..., C_gait K and the corresponding person identities; Calculate the cosine similarity between the extracted voiceprint features and each voiceprint feature in the voiceprint library, sort multiple cosine similarities in descending order of value, and return the first K cosine similarity values C_voice1, C_voice2,..., C_voice K and the corresponding person identities; Taking the union of the top K results returned by each of the three modalities of face features, gait features, and voiceprint features to obtain a set of candidate persons M; A multi-modal consistency scoring module for performing multi-modal consistency scoring on each person in the set of candidate persons and returning the identity of the person with the highest score as the identified person's identity; Performing multi-modal consistency scoring on each person in the set of candidate persons and returning the identity of the person with the highest score as the identified person's identity, specifically including: For the k-th person M in the set M of candidates to be selected k , compare the cosine similarity of their face features and gait features, and take the modality with the higher cosine similarity as M k 's basic modality, and use the cosine similarity value corresponding to the basic modality as the basic modality score Score_base k ; Calculate the consistency score w of face and gait based on the face region and the corresponding human body region f,g ; Calculate the consistency score w between the face and the voiceprint based on the facial key points and the Mel spectrogram f,v ; Denote the consistency score of gait and voiceprint as w g,v ; According to the consistency score w f,g , w f,v and w g,v calculate the modal consistency score Score_coin under different basic modes k ; According to the basic modal score Score_base k and the modal consistency score Score_coin k Calculate the total score Score of the k-th character M k as follows k = Score_base k + Score_coin k ; Return the total score Score k The identity of the person with the highest score is used as the identified person's identity; A person identity annotation module for annotating the identity of the person in each frame of the audio-video stream according to the identified person's identity and outputting the audio-video stream after identity recognition.

4. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio-video person recognition method based on multi-modal biometric consistency as described in any one of claims 1 to 2.

5. The electronic device according to claim 4, characterized in that, The memory is a non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Identity authentication method and device based on biological characteristics

    CN118734280A

  • Multi-modal risk data synthesis method for anti-fraud and risk identification

    CN119479026A