Human mouth shape and voice matching recognition method based on deep learning

By combining deep learning technology with image and audio recognition, the system identifies each person's lip movements and sound source location, separates and processes noise, and solves the problem of speech recognition difficulties in noisy environments using traditional hearing aids, achieving high-precision speech information extraction and recording.

CN120877754APending Publication Date: 2025-10-31SHENZHEN LESENBELL HEARING TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510990056.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional hearing aids struggle to accurately identify and distinguish target speech signals in noisy environments, especially when multiple people are communicating, as they cannot effectively suppress noise interference, resulting in low speech recognition accuracy and affecting the communication ability of people with hearing impairments.

Method used

Using a deep learning-based method, the image identifies the lip movements of each person in the surrounding environment, and combines this with a microphone array to identify the location of the sound source. The sound source location is then corrected to separate audio features. Deep learning is then performed to obtain speech information, and background noise is processed to extract each person's spoken speech and its text content.

Benefits of technology

It effectively solves the "cocktail party problem" of hearing aids in noisy environments, improves the accuracy and clarity of speech recognition, and can directionally receive and record speech information under far-field conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877754A_ABST
    Figure CN120877754A_ABST
Patent Text Reader

Abstract

The invention provides a deep learning-based human mouth shape and voice matching recognition method, which comprises the following steps of: recognizing respective speaking mouth shape characteristics of all people in a surrounding environment through an image, and recognizing a sound source position and an audio characteristic of the surrounding environment through a pickup array; according to the speaking mouth shape features, correcting the sound source position so as to separate the audio features to obtain the subordinate audio features of each person; performing deep learning on the audio features and the speaking mouth shape features of each person to obtain voice information sent by each person; carrying out background noise processing on the voice information to obtain the speaking voice and the text content of each person; personnel around the hearing aid in a noisy environment are identified in a visual and voice recognition mode, voice of the personnel is recorded and converted, the cocktail problem of the hearing aid in far-field recognition is effectively solved, voice interference is effectively restrained, and voice recognition accuracy and definition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition, and more particularly to a deep learning-based method for matching and recognizing human lip movements with speech. Background Technology

[0002] Hearing impairment, as an auditory sensory disorder, severely impacts an individual's language development, learning ability, and social participation. While hearing aids can assist in hearing, traditional hearing aids suffer from inconvenience and low recognition accuracy, particularly in noisy environments, such as when multiple people are speaking simultaneously. In these situations, the accuracy of speech recognition drops significantly, a phenomenon known as the "cocktail party problem." Furthermore, traditional speech recognition technologies struggle to accurately distinguish and identify target speech signals in far-field conditions, especially in the presence of multi-source interference, making it difficult to pinpoint the sound source. Consequently, in multi-person communication environments, hearing aids amplify various sounds, making it difficult for hearing-impaired individuals to hear clearly. This leads to a loss of auditory attention, making it difficult to focus on the events they are interested in, and in noisy environments, they are distracted by additional sounds, hindering effective communication with others and resulting in a high rate of hearing aid abandonment. Therefore, accurately identifying and processing sound in complex sound environments is crucial for hearing aids to achieve noise suppression and improve sound recognition clarity in the "cocktail party problem." Summary of the Invention

[0003] The purpose of this invention is to provide a deep learning-based method for matching and recognizing human lip movements and speech. This method involves identifying the lip movements of all individuals in the surrounding environment through image recognition, identifying the location and audio features of sound sources in the environment using a microphone array, correcting the sound source location based on the lip movement characteristics, and separating the audio features to obtain the audio features of each individual. Deep learning is then applied to the audio features and lip movement characteristics of each individual to obtain their speech information. Background noise is processed on the speech information to obtain the spoken speech and its text content for each individual. By using visual and auditory recognition methods, individuals in noisy environments are identified, and their speech is recorded and converted. This method can both receive speech information in a directional manner and convert the speech of surrounding individuals into text or record speech for later review, effectively solving the "cocktail party problem" in far-field recognition using hearing aids, effectively suppressing sound interference, and improving the accuracy and clarity of speech recognition.

[0004] This invention is achieved through the following technical solution:

[0005] Deep learning-based methods for human lip-syncing and speech recognition include:

[0006] Image recognition identifies the lip movements of all individuals in the surrounding environment; a microphone array identifies the location of sound sources and audio characteristics in the surrounding environment.

[0007] The sound source location is corrected based on the lip-sync characteristics.

[0008] Based on the corrected sound source location, the audio features are separated to obtain the audio features of each person's subordinates;

[0009] Deep learning is performed on the audio features and lip-sync features of each person to obtain the speech information of each person.

[0010] Background noise is processed on the speech information to obtain the spoken voice and its text content for each person.

[0011] Optionally, the image recognizes the individual lip-sync features of all people in the surrounding environment, including:

[0012] Capture dynamic images of all people in the surrounding environment, and perform lip-sync recognition on the dynamic images to obtain the lip-sync features of each person; wherein, the lip-sync features include the temporal variation characteristics of the lip-sync movements.

[0013] Optionally, identifying the location of sound sources in the surrounding environment via a microphone array includes:

[0014] Obtain the time difference when each microphone in the microphone array receives the same sound;

[0015] Based on the time difference and the spatial geometric position of all microphones in the microphone array, the sound source location of the surrounding environment is determined; wherein, the sound source location refers to the location of the person emitting the sound at different times in the surrounding environment.

[0016] Optionally, the audio characteristics of the surrounding environment are identified by a microphone array, including:

[0017] Acquire raw sound data collected by the microphone array from the surrounding environment;

[0018] The original sound data is subjected to voiceprint preprocessing, spectrum analysis, and noise reduction to extract audio features.

[0019] Optionally, correcting the sound source location based on the lip-sync features includes:

[0020] Extract the timing of each person's lip movements from the lip-reading features;

[0021] The timing of each person's lip-reading actions is compared with the location of people who made sounds at different times within the sound source location to determine whether the person actually made a sound at the corresponding time, thereby correcting the location of people who made sounds at different times in the surrounding environment.

[0022] Optionally, the timing of each person's lip-syncing actions is extracted from the lip-syncing features, including:

[0023] Based on the top-left x-coordinate, top-left y-coordinate, width, and height of the bounding box containing the face region of each person, the face region of each person is detected from the video stream, and the bounding box containing the face region is obtained. The bounding box containing the face region of each person is represented by the top-left x-coordinate, top-left y-coordinate, width, and height of the bounding box containing the face region of each person, and is specifically expressed as the following formula (1):

[0024]

[0025] In the above formula (1), This represents the position expression of the bounding box containing the face region of the i-th person at time t; Let x and y represent the top-left x-coordinate and top-left y-coordinate of the bounding box containing the face region of the i-th person at time t, respectively. Let represent the width and height of the bounding box containing the face region of the i-th person at time t, respectively;

[0026] Furthermore, the bounding box containing the face region of each person in the video stream is identified by an ID, so that the ID of the bounding box containing the face region of the same person is unique and the same in each frame of the video stream.

[0027] Determine the vertical and horizontal cropping ratios within the bounding box of a person's face region. Based on these ratios, extract the mouth region from the bounding box of each person's face region. The extracted mouth region is represented by the following formula (2):

[0028]

[0029] In the above formula (2), Let V represent the bounding box containing the face region of the i-th person at time t, from which the mouth region is extracted; α represents the first cropping ratio in the vertical direction; β represents the second cropping ratio in the vertical direction; γ represents the first cropping ratio in the horizontal direction; δ represents the second cropping ratio in the horizontal direction; V t Vt[] represents the video frame at time t in the video stream; Vt[] represents a rectangular region extracted from the video frame Vt.

[0030] The feature vector of the i-th person at time t is determined based on the mouth region sequence from frame tk to frame t. The speaking probability of the i-th person in the video frame at time t is determined based on the feature vector of the i-th person at time t, as specifically expressed by the following formula (3):

[0031]

[0032] In the above formula (3), Pspeak(i,t) represents the probability of the i-th person speaking in the video frame at time t; σ() represents the Sigmoid function; b represents the preset bias term; W T It is a 256-dimensional weight matrix; The feature vector used to determine the i-th person at time t is a lip fingerprint represented by 256 indicators; k represents the size of the time window. Represents the mouth region sequence from frame tk to frame t; ResNet3D() represents the 3D-ResNet18 network model; R 256 It represents a 256-dimensional real vector, corresponding to 256 indicators describing mouth movements.

[0033] In one embodiment, the timing of each person's lip-reading actions is compared with the locations of people who made sounds at different times within the sound source location to determine whether the person actually made a speaking sound at the corresponding time, thereby correcting the locations of people who made sounds at different times in the surrounding environment, including:

[0034] Based on the estimated direction of the sound source at time t estimated by the microphone array, the horizontal deflection angle of the person's face relative to the front of the camera at time t, the horizontal pixel coordinates of the center of the bounding box of the person's face region at time t, the width of the video frame, and the horizontal field of view of the camera, a spatial judgment function G(i,x,t) is constructed, which is specifically expressed as the following formula (4):

[0035]

[0036] In the above formula (4), This represents the estimated direction of the x-th sound source at time t, as estimated by the microphone array; ρ represents the bandwidth parameter, which controls the looseness of the matching, and its value is 15°. Used to determine the horizontal deflection angle of the face of the i-th person relative to the front of the camera at time t; represents the horizontal pixel coordinates of the center of the bounding box of the face region of the i-th person at time t; W represents the width of the video frame; Indicates the horizontal field of view of the camera;

[0037] The spatial judgment function is used to determine whether the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person. The output of the spatial judgment function is in the range of [0, 1]. The closer the output is to 1, the higher the probability that the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person.

[0038] Based on the activity level of the sound source at time t, the spatial judgment function, and the speaking probability of the person in the video frame at time t, the matching degree between the person and the sound source is determined, specifically expressed as the following formula (5):

[0039]

[0040] In the above formula (5), Ci,x represents the matching degree between the i-th person and the x-th sound source; A x (t) represents the activity level of the x-th sound source at time t, and its value ranges from [0, 1]. The larger the value of Ci,x, the higher the probability that the i-th person and the x-th sound source are the same speaker.

[0041] When Ci,x is greater than or equal to the preset threshold, the location of the i-th person is determined to be the same as that of the x-th sound source; when Ci,x is less than the preset threshold, proceed to step S6 below.

[0042] The estimated direction above can be corrected using the following formula (6).

[0043]

[0044] In the above formula (6), Indicates the estimated direction The corrected result; λ is the preset fusion weight.

[0045] Optionally, based on the corrected sound source location, the audio features are separated to obtain the audio features of each person's subordinates, including:

[0046] Obtain the speech time distribution characteristics and voiceprint characteristics of the person corresponding to the corrected sound source location;

[0047] Based on the speaking time distribution characteristics and voiceprint characteristics, the audio characteristics of the personnel's subordinates are separated from the audio characteristics using a convolutional neural network model;

[0048] or,

[0049] Deep learning is performed on the audio features and lip-sync features of each person to obtain the speech information emitted by each person, including:

[0050] Deep learning phoneme encoding is performed on the audio features of each person to obtain the phoneme weights corresponding to the audio features;

[0051] The lip-sync feature is used to train a lip-sync generation model. The trained lip-sync generation model is then input based on the phoneme weights. The original sound data collected by the microphone array is then labeled with the speaker.

[0052] Based on the speaker tags of the original sound data, the voice information spoken by each person is separated.

[0053] Optionally, background noise processing is performed on the speech information to obtain the spoken voice and its text content for each person, including:

[0054] The speech information is subjected to self-adaptive noise suppression by using a weighted overlapping additive filter and sub-band division based on a psychoacoustic model.

[0055] The speech information is isolated from background noise by spatial masking release processing, thereby extracting the speech of each person;

[0056] Semantic recognition is performed on the spoken speech to obtain the corresponding text content.

[0057] Optionally, it also includes: transmitting the voice of the corresponding person to the sound playback device worn by the user according to the user's directional voice reception request;

[0058] And / or, after storing the spoken voice and its text content, the spoken voice and / or the text content are transmitted to the user's device for playback and / or display according to the user's query request.

[0059] Compared with the prior art, the present invention has the following beneficial effects:

[0060] This application provides a deep learning-based human lip-reading and speech matching recognition method. It identifies the lip-reading features of all individuals in the surrounding environment using an image recognition system. A microphone array identifies the location and audio features of sound sources in the surrounding environment. Based on the lip-reading features, the sound source location is corrected, thereby separating the audio features to obtain the audio features of each individual. Deep learning is then applied to the audio features and lip-reading features of each individual to obtain their speech information. Background noise is processed on the speech information to obtain the spoken speech and its text content for each individual. Through visual and auditory recognition, individuals in noisy environments are identified, and their speech is recorded and converted. This method can both receive speech information in a directional manner and convert the speech of surrounding individuals into text or record speech for later review. It effectively solves the "cocktail party problem" in far-field recognition using hearing aids, effectively suppressing sound interference and improving the accuracy and clarity of speech recognition. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0062] Figure 1 The flowchart shows the deep learning-based human lip-reading and speech matching and recognition method provided by this invention.

[0063] Figure 2 A flowchart for identifying the location of sound sources in the surrounding environment for a microphone array.

[0064] Figure 3 A flowchart for identifying audio features of the surrounding environment for a microphone array.

[0065] Figure 4 This is a flowchart of deep learning for audio features and lip-sync features.

[0066] Figure 5 This is a schematic diagram of a scenario corresponding to the deep learning-based human lip-reading and speech matching and recognition method provided by the present invention. Detailed Implementation

[0067] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, not the entire structure. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0068] The terms “comprising” and “having”, and any variations thereof, used in this application are intended to cover non-exclusive inclusion. For example, a process, method, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0069] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0070] Please see Figure 1 As shown, an embodiment of this application provides a deep learning-based method for human lip-syncing and speech matching recognition. This deep learning-based method for human lip-syncing and speech matching recognition includes:

[0071] S100: Image recognition of the lip movements of all individuals in the surrounding environment; identification of sound source locations and audio characteristics in the surrounding environment through a microphone array;

[0072] S200 corrects the sound source location based on lip movements.

[0073] S300 separates the audio features based on the corrected sound source location to obtain the audio features of each person's subordinates;

[0074] S400 performs deep learning on the audio features and lip-sync features of each person's subordinates to obtain the speech information emitted by each person.

[0075] The S500 processes background noise from the speech information to obtain the spoken voice and its text content for each person.

[0076] The beneficial effects of the above embodiments are as follows: This deep learning-based human lip-reading and speech matching recognition method identifies the lip-reading features of all individuals in the surrounding environment through image recognition; it identifies the sound source location and audio features of the surrounding environment through a microphone array; it corrects the sound source location based on the lip-reading features, thereby separating the audio features to obtain the audio features of each individual; it performs deep learning on the audio features and lip-reading features of each individual to obtain the speech information emitted by each individual; it processes the speech information to remove background noise, obtaining the speech and text content of each individual; and it identifies and records the speech of individuals in noisy environments through visual and auditory recognition methods. This method can both receive speech information in a targeted manner and convert the speech of individuals in the surrounding environment into text or record speech for later review, effectively solving the "cocktail party problem" in far-field recognition of hearing aids, effectively suppressing sound interference and improving the accuracy and clarity of speech recognition.

[0077] In another embodiment, in S100, image recognition includes identifying the lip-sync features of all individuals in the surrounding environment, including:

[0078] S110 captures dynamic images of all people in the surrounding environment, performs lip-sync recognition on the dynamic images, and obtains the lip-sync features of each person; among which, the lip-sync features include the temporal variation characteristics of the lip-sync movements.

[0079] In practical applications, cameras can capture dynamic images of all individuals in the surrounding environment during multi-person conversations. This includes capturing individual facial images and / or a global distribution image of all individuals within the environment. This allows for real-time, continuous tracking and recognition of subtle facial movements, particularly facial expressions and lip movements. Facial landmark detection and simultaneous lip-sync recognition technologies are also used to analyze these dynamic images, revealing the temporal changes in each individual's lip movements, such as the angle and shape of their mouths. Since each individual's lip movements differ during speech, visually recognizing these temporal changes allows for individual differentiation and recording. This facilitates more refined speech recognition when multiple individuals speak simultaneously, leading to confusion. Furthermore, image recognition of all individuals in the surrounding environment enables spatial identification and localization of each person, providing an effective solution to spatial confusion during subsequent sound acquisition and recognition, resulting in clearer and more accurate sound recognition.

[0080] In another embodiment, please refer to Figure 2 As shown, in S100, the location of sound sources in the surrounding environment is identified through a microphone array, including:

[0081] S120, obtain the time difference when each microphone in the microphone array receives the same sound;

[0082] S130, determine the location of the sound source in the surrounding environment based on the time difference and the spatial geometric position of all microphones in the microphone array; where the sound source location refers to the location of the person emitting the sound at different times in the surrounding environment.

[0083] In practical applications, an array containing multiple microphones can be used to collect sound from the surrounding environment. Each microphone can be, but is not limited to, a microphone and a sound analysis and processing terminal. The microphone is used to collect audio data from the surrounding environment, while the sound analysis and processing terminal is used to process the audio data by performing spatial propagation time analysis, sound content recognition, and conversion. Since the spatial distance between different microphones within the array and the same person is not uniform, there is a time difference in the arrival time of the sound emitted by the same person (i.e., the same sound source) at different microphones (i.e., it is collected by different microphones). Generally speaking, the microphone closer to the sound source collects the sound earlier. In fact, the spatial distribution of all microphones in the microphone array is known and determined. Thus, the relative distance and relative orientation between any two microphones are also known and determined. By obtaining the time difference between any two microphones in the microphone array receiving the same sound from the same sound source, and based on the time difference, the speed of sound propagation, and the relative distance and relative orientation between the corresponding two microphones, combined with the triangulation method, the position of each sound source (i.e., each person) in the surrounding environment can be determined. This allows for the accurate marking of the individual positions of all people in the surrounding environment during their speech, providing a spatial positioning direction for subsequent marking and extraction of each person's speech from the collective sound collected from the surrounding environment, thereby improving the accuracy of far-field speech recognition.

[0084] In another embodiment, please refer to Figure 3 As shown, in S100, the audio characteristics of the surrounding environment are identified through a microphone array, including:

[0085] S140, acquire raw sound data collected by the microphone array from the surrounding environment;

[0086] S150 performs voiceprint preprocessing, spectrum analysis, and noise reduction on the raw sound data to extract audio features.

[0087] In practical applications, the microphone array can also collect raw sound data from the surrounding environment. This raw sound data includes the sounds emitted by all people speaking in the surrounding environment, as well as background noise. To extract each person's own voice from the raw sound data and considering that different people have different voiceprint characteristics, the raw sound data undergoes voiceprint preprocessing, spectrum analysis, and noise reduction to extract the audio features corresponding to each person. This provides an accurate basis for subsequent independent speech recognition for each person. Preferably, a hidden Markov model method and VQ clustering algorithm (such as K-means clustering algorithm) can be used to preprocess the raw sound data. Spectrum analysis is performed on the raw sound data containing background noise, and the spectrum of the pure background noise signal is subtracted to obtain the denoised spectrum. Then, audio features are extracted from the denoised spectrum. These audio features may include, but are not limited to, pitch period, formants, short-time average energy, or amplitude.

[0088] In another embodiment, in S200, correcting the sound source location based on lip-sync characteristics includes:

[0089] S210, extract the timing of each person's lip movements from lip-reading features;

[0090] S220 compares the timing of each person's mouth movements with the location of the person making the sound at different times, which is included in the sound source location, to determine whether the person actually made the sound at the corresponding time, thereby correcting the location of the person making the sound at different times in the surrounding environment.

[0091] In practical work, considering that the speech produced by a person during the speaking process is closely related to the person's lip movements, this correlation is mainly reflected in the fact that the lip movements made by a person to a certain extent determine the content of the speech. By visually recognizing the lip movements of a person, it is possible to determine whether a person has made a sound. In this way, by recognizing the occurrence of the lip movements of a person throughout the process of multiple people speaking simultaneously or interactively, it is possible to accurately determine whether each person has actually made a sound at a certain moment. Specifically, the timing of each person's lip movements is extracted from the lip-sync features. This timing includes all time points when each person makes a real lip-sync action during the speaking process. Then, the timing of each person's lip-sync action is compared with the location of the person who made a sound at different times (i.e., the corresponding location of the person who made a sound at different times during the speaking process) contained in the sound source location. If the person at the corresponding location is considered to have made a sound at a certain time point and also made a lip-sync action at the same time point, then the person is considered to have indeed made a sound at the aforementioned time point. Otherwise, the person is considered not to have made a sound at the aforementioned time point. This corrects the location of the person who made a sound at different times in the surrounding environment, which was previously identified by the microphone's sound splitting, ensuring that the corrected sound source location accurately reflects the time point when each person in the surrounding environment made a sound. This provides an accurate time division benchmark for subsequently separating and extracting each person's own audio features from the audio features.

[0092] In another embodiment, in S210, the timing of each person's lip-syncing actions is extracted from the lip-syncing features, including:

[0093] S211, based on the upper left corner x-coordinate, upper left corner y-coordinate, width, and height of the bounding box of the face region of each person, detect the face region of each person from the video stream to obtain the bounding box of the face region. The bounding box of the face region of each person is represented by the upper left corner x-coordinate, upper left corner y-coordinate, width, and height of the bounding box of the face region of each person, and is specifically expressed as the following formula (1):

[0094]

[0095] In the above formula (1), This represents the position expression of the bounding box containing the face region of the i-th person at time t; Let x and y represent the top-left x-coordinate and top-left y-coordinate of the bounding box containing the face region of the i-th person at time t, respectively. Let represent the width and height of the bounding box containing the face region of the i-th person at time t, respectively;

[0096] Furthermore, the bounding box containing the face region of each person in the video stream is identified by an ID, so that the ID of the bounding box containing the face region of the same person is unique and the same in each frame of the video stream.

[0097] S212, determine the vertical and horizontal cropping ratios within the bounding box of the person's face region, and extract the mouth region from the bounding box of each person's face region according to the vertical and horizontal cropping ratios. The extracted mouth region is represented by the following formula (2):

[0098]

[0099] In the above formula (2), Let V represent the bounding box containing the face region of the i-th person at time t, from which the mouth region is extracted; α represents the first cropping ratio in the vertical direction; β represents the second cropping ratio in the vertical direction; γ represents the first cropping ratio in the horizontal direction; δ represents the second cropping ratio in the horizontal direction; V t V represents the video frame at time t in the video stream; t [] indicates the video frame V t Extract a rectangular area from the middle;

[0100] Formula (2) above describes the operation of "cutting out" the mouth region from a face image, as follows:

[0101] At time t, the bounding box of the face region of the i-th person is a rectangle:

[0102] The method used to represent vertical cutting is, for example, cutting down 30% (α = 0.3) from the side of the rectangle closest to the top of the head, leaving 70% of the height (β = 0.7), to obtain the upper and lower boundaries of the lips;

[0103] The method used to represent horizontal cutting is, for example, cutting 20% ​​(γ = 0.2) from the left side of the face to the right of the rectangle mentioned above, leaving 80% of the width (δ = 0.8), to obtain the left and right boundaries of the lips;

[0104] S213. Based on the mouth region sequence from frame tk to frame t, determine the feature vector of the i-th person at time t. Based on the feature vector of the i-th person at time t, determine the speaking probability of the i-th person in the video frame at time t, specifically expressed as the following formula (3):

[0105]

[0106] In the above formula (3), Pspeak(i,t) represents the probability of the i-th person speaking in the video frame at time t; σ() represents the Sigmoid function; b represents the preset bias term; W T It is a 256-dimensional weight matrix; The feature vector used to determine the i-th person at time t is a lip fingerprint represented by 256 indicators; k represents the size of the time window. Represents the mouth region sequence from frame tk to frame t; ResNet3D() represents the 3D-ResNet18 network model; R 256 It represents a 256-dimensional real vector, corresponding to 256 indicators describing mouth movements.

[0107] In another embodiment, in S220, the timing of each person's lip-syncing is compared with the locations of people who made sounds at different times, as included in the sound source location, to determine whether the person actually made a speaking sound at the corresponding time, thereby correcting the locations of people who made sounds at different times in the surrounding environment, including:

[0108] S221, based on the estimated direction of the sound source at time t estimated by the microphone array, the horizontal deflection angle of the person's face relative to the front of the camera at time t, the horizontal pixel coordinates of the center of the bounding box of the person's face region at time t, the width of the video frame, and the horizontal field of view of the camera, a spatial judgment function G(i,x,t) is constructed, which is specifically expressed as the following formula (4):

[0109]

[0110] In the above formula (4), This represents the estimated direction of the x-th sound source at time t, as estimated by the microphone array; ρ represents the bandwidth parameter, which controls the looseness of the matching, and its value is 15°. Used to determine the horizontal deflection angle of the face of the i-th person relative to the front of the camera at time t; represents the horizontal pixel coordinates of the center of the bounding box of the face region of the i-th person at time t; W represents the width of the video frame; This indicates the horizontal field of view of the camera, and its value can be 60°;

[0111] The spatial judgment function is used to determine whether the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person. The output of the spatial judgment function is in the range of [0, 1]. The closer the output is to 1, the higher the probability that the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person.

[0112] S222, based on the activity level of the sound source at time t, the spatial judgment function, and the speaking probability of the person in the video frame at time t, the matching degree between the person and the sound source is determined, specifically expressed as the following formula (5):

[0113]

[0114] In the above formula (5), Ci,x represents the matching degree between the i-th person and the x-th sound source; A x (t) represents the activity level of the x-th sound source at time t, and its value ranges from [0, 1]. The larger the value of Ci,x, the higher the probability that the i-th person and the x-th sound source are the same speaker.

[0115] Among them, activity level A x (t) is used to indicate whether the x-th sound source is emitting sound at time t and the significance of its emission. It is a value between 0 and 1. For example, the value can be preset according to rules. For example, the closer the value is to 1, the more it indicates that the sound source is emitting sound (such as a person speaking clearly or a musical instrument playing continuously); the closer the value is to 0, the more it indicates that the sound source is silent or is just background noise (such as keyboard sounds or air conditioner noise).

[0116] When Ci,x is greater than or equal to the preset threshold, the location of the i-th person is determined to be the same as that of the x-th sound source; when Ci,x is less than the preset threshold, proceed to step S6 below.

[0117] S223, using the following formula (6), correct the estimated direction above.

[0118]

[0119] In the above formula (6), Indicates the estimated direction The corrected result; λ is the preset fusion weight.

[0120] The beneficial effects of the above process are as follows: First, joint visual and auditory analysis, utilizing both lip movements captured by the camera and sound source directions detected by the microphone array, significantly improves speaker localization accuracy through spatiotemporal consistency matching; Second, resistance to multi-person interaction interference: solves the problem of "who is speaking" in scenarios where multiple people are speaking simultaneously; Third, lightweight network design: employs efficient models such as 3D-ResNet18 to process lip features, meeting real-time requirements (such as video conferencing and live streaming scenarios), and reducing latency compared to traditional pure audio localization systems; Fourth, high robustness: the bandwidth parameter in the spatial judgment function can be dynamically adjusted according to the environment (e.g., increasing fault tolerance in noisy environments), and the fusion weights are automatically optimized based on sensor accuracy (relying on audio when the camera is poor, and on vision when the microphone is poor); Fifth, cross-scenario applicability: suitable for various scenarios such as conference rooms, smart homes, security monitoring, and in-vehicle voice systems.

[0121] In another embodiment, in S300, the audio features are separated based on the corrected sound source location to obtain the audio features of each person's subordinates, including:

[0122] S301, Obtain the speech time distribution characteristics and voiceprint characteristics of the person corresponding to the corrected sound source location;

[0123] S302, based on the speech time distribution characteristics and voiceprint characteristics, uses a convolutional neural network model to separate the audio features of the personnel's subordinates from the audio features.

[0124] The above analysis shows that the corrected sound source location accurately reflects the speaking sounds of all individuals in the surrounding environment at all corresponding time points. The speaking time distribution features and voiceprint features of each individual are extracted from the corrected sound source location. The speaking time distribution features refer to the distribution of the actual speaking time points / time intervals of each individual, while the voiceprint features refer to the unique voiceprint features of each individual. These features are then input into a convolutional neural network model to extract and separate the audio features of each individual, thus accurately separating the audio features of each individual during their speaking interactions in the surrounding environment. In practical applications, readily available deep learning-based supervised speech separation techniques can be used to minimize the cost function and separate the audio features of each individual, thereby solving the problem of multi-source signal interference detection. This is a conventional technique in this field and will not be described in detail here.

[0125] In another embodiment, please refer to Figure 4 As shown, in S400, deep learning is performed on the audio features and lip-sync features of each person's subordinates to obtain the speech information emitted by each person, including:

[0126] S401. Perform deep learning phoneme encoding processing on the audio features of each person's subordinates to obtain the phoneme weights corresponding to the audio features.

[0127] S402. Train a lip movement generation model using the lip movement features. Input the phoneme weights into the trained lip movement generation model to mark the speakers for the original sound data collected by the microphone array.

[0128] S403. Separate the speech information emitted by each person according to the speaker marking of the original sound data.

[0129] A phoneme is the smallest speech unit divided according to the natural attributes of speech. From an acoustic perspective, a phoneme is the smallest speech unit divided from the aspect of sound quality. From a physiological perspective, one pronunciation action forms one phoneme. For example, "ma" (mom) contains two pronunciation actions of "m" and "a", corresponding to two phonemes. The sounds emitted by the same pronunciation action are one phoneme, and the sounds emitted by different pronunciation actions are different phonemes. For example, in "ma-mi" (mommy), the two "m" pronunciation actions are the same, so they are the same phoneme, while the pronunciation actions of "a" and "i" are different, so they are different phonemes. Perform deep learning phoneme encoding processing on the audio features of each person using a convolutional neural network model to obtain the phoneme weights included in each person's own audio features. The above phoneme weights refer to the proportion of the quantity weights of all phonemes included in the person's audio features. For example, the proportion of the quantity weights of phonemes such as "m", "a", "i", etc. in the audio features. Then, use the lip movement features to train the lip movement generation model, and input the above phoneme weights into the trained lip movement generation model to distinguish and mark the speaker identities of the original sound data collected by the microphone array, that is, to calibrate which specific identity of the speaker each sound data in the original sound data comes from, and achieve a one-to-one identification and distinction of the original sound data generated in the "cocktail party problem" scenario of the surrounding environment for all the people present in the surrounding environment, so as to separate and extract the actual speech information emitted by each person, providing basic data for subsequent speech content recognition of each person.

[0130] In another embodiment, in S500, perform background noise processing on the speech information to obtain the speech and its text content of each person, including:

[0131] S501. Perform self-adaptive noise suppression on the speech information through a weighted overlap-add filter and subband division based on a psychoacoustic model.

[0132] S502. Isolate the background noise from the speech information through spatial masking release processing to extract the speech of each person.

[0133] S503 performs semantic recognition on the spoken speech to obtain the corresponding text content.

[0134] In practical work, the surrounding environment contains both human voices and background noise. While the method described above is used to separate and extract individual voice information for each person, background noise is unavoidable and can affect the accuracy of voice recognition. Therefore, it is necessary to first suppress and isolate noise in the voice information to prevent background noise from drowning out the actual speech and to improve the signal-to-noise ratio. Then, semantic recognition is performed on the spoken voice to obtain the corresponding text content, ensuring the accuracy of the speech-to-text conversion.

[0135] In practical applications, a subband adaptive noise suppression method can be used. This method employs a weighted overlapping additive filter and subband division based on a psychoacoustic model to perform adaptive noise suppression on speech information. Additionally, spatial masking release processing (SRM) combined with artificial intelligence and the physics of sound propagation is used to isolate background noise from speech information, thereby extracting the speech of each individual and enhancing the accuracy of speech recognition.

[0136] In another embodiment, it further includes: transmitting the voice of the corresponding person to the sound playback device worn by the user according to the user's directional voice reception request;

[0137] And / or, after storing the spoken voice and its text content, transmit the spoken voice and / or text content to the user's device for playback and / or display according to the user's query request.

[0138] In practice, users with hearing impairments can send a directional voice reception request to listen to the voice of a specific person in their surroundings. Considering that the above process has already separated the voice of each person in the surrounding environment and identified their identity and location, the directional voice reception request is parsed to determine the specific person the user wishes to hear. The specific person's voice is then extracted from the separated voices of all individuals and transmitted to the user's hearing aid or other audio playback device. Alternatively, after saving and storing the voice and text content of each person in the surrounding environment, a matching voice and / or text content is extracted from the saved voice and text data and transmitted to the hearing aid and / or smartphone for playback and / or display. This allows hearing-impaired individuals to promptly understand the surrounding environment's audio situation, effectively solving the "cocktail party problem."

[0139] Please see Figure 5 This invention relates to a multi-person conference scenario where the deep learning-based lip-reading and speech matching recognition method applies. The participants include both hearing-impaired and hearing-normal individuals (i.e., those with normal hearing). Hearing aids worn by the hearing-impaired individuals may include a microphone for collecting sound information in the multi-person conference scenario. A portable device carried by the hearing-impaired individual includes a camera, processor, memory, and display. This portable device may be, but is not limited to, a smartphone, and is communicatively connected to the microphone. The camera is used to capture images of the surrounding environment in the multi-person conference scenario; preferably, the camera has a 360-degree image capture function. The processor processes the sound information collected by the microphone and the surrounding environment images collected by the camera, thereby implementing the aforementioned deep learning-based lip-reading and speech matching recognition method. The memory stores the spoken speech and its text content obtained by the processor implementing the aforementioned lip-reading and speech matching recognition method. When a hearing-impaired person initiates a query request to the portable device, the portable device transmits the corresponding speech to the hearing aid worn by the hearing-impaired person and plays back the speech; and / or, when a hearing-impaired person initiates a query request to the portable device, the portable device displays the text content corresponding to the speech to the hearing-impaired person through its own display.

[0140] In summary, this deep learning-based human lip-reading and speech matching recognition method identifies the lip-reading features of all individuals in the surrounding environment through image recognition, identifies the location and audio features of sound sources in the surrounding environment through a microphone array, corrects the sound source location based on the lip-reading features, and separates the audio features to obtain the audio features of each individual. Deep learning is then applied to the audio features and lip-reading features of each individual to obtain their speech information. Background noise is processed on the speech information to obtain the spoken speech and its text content for each individual. Visual and auditory recognition methods are used to identify individuals in noisy environments and record and convert their speech. This method can both receive speech information in a targeted manner and convert the speech of surrounding individuals into text or record speech for later review, effectively solving the "cocktail party problem" in far-field recognition for hearing aids, effectively suppressing sound interference and improving the accuracy and clarity of speech recognition.

[0141] The above is only one specific embodiment of the present invention, and any improvements made based on the concept of the present invention shall be considered within the scope of protection of the present invention.

Claims

1. A deep learning-based method for human lip-syncing and speech recognition, characterized in that, include: Image recognition of the individual lip movements of all people in the surrounding environment; The location of sound sources and audio characteristics of the surrounding environment are identified by a microphone array; The sound source location is corrected based on the lip-sync characteristics. Based on the corrected sound source location, the audio features are separated to obtain the audio features of each person's subordinates; Deep learning is performed on the audio features and lip-sync features of each person to obtain the speech information of each person. Background noise is processed on the speech information to obtain the spoken voice and its text content for each person.

2. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 1, characterized in that: Image recognition of the individual lip movements of all people in the surrounding environment, including: Capture dynamic images of all people in the surrounding environment, and perform lip-sync recognition on the dynamic images to obtain the lip-sync features of each person; wherein, the lip-sync features include the temporal variation characteristics of the lip-sync movements.

3. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 1, characterized in that: Identifying the location of sound sources in the surrounding environment using a microphone array includes: Obtain the time difference when each microphone in the microphone array receives the same sound; Based on the time difference and the spatial geometric position of all microphones in the microphone array, the sound source location of the surrounding environment is determined; wherein, the sound source location refers to the location of the person emitting the sound at different times in the surrounding environment.

4. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 2, characterized in that: Identifying the audio characteristics of the surrounding environment through a microphone array includes: Acquire raw sound data collected by the microphone array from the surrounding environment; The original sound data is subjected to voiceprint preprocessing, spectrum analysis, and noise reduction to extract audio features.

5. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 1, characterized in that: Based on the lip-sync characteristics, the location of the sound source is corrected, including: Extract the timing of each person's lip movements from the lip-reading features; The timing of each person's lip-reading actions is compared with the location of people who made sounds at different times within the sound source location to determine whether the person actually made a sound at the corresponding time, thereby correcting the location of people who made sounds at different times in the surrounding environment.

6. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 5, characterized in that: Extracting the timing of each person's lip movements from the lip-reading features includes: Based on the top-left x-coordinate, top-left y-coordinate, width, and height of the bounding box containing the face region of each person, the face region of each person is detected from the video stream, and the bounding box containing the face region is obtained. The bounding box containing the face region of each person is represented by the top-left x-coordinate, top-left y-coordinate, width, and height of the bounding box containing the face region of each person, and is specifically expressed as the following formula (1): In the above formula (1), This represents the position expression of the bounding box containing the face region of the i-th person at time t; Let x and y represent the top-left x-coordinate and top-left y-coordinate of the bounding box containing the face region of the i-th person at time t, respectively. Let represent the width and height of the bounding box containing the face region of the i-th person at time t, respectively; Furthermore, the bounding box containing the face region of each person in the video stream is identified by an ID, so that the ID of the bounding box containing the face region of the same person is unique and the same in each frame of the video stream. Determine the vertical and horizontal cropping ratios within the bounding box of a person's face region. Based on these ratios, extract the mouth region from the bounding box of each person's face region. The extracted mouth region is represented by the following formula (2): In the above formula (2), Let V represent the bounding box containing the face region of the i-th person at time t, from which the mouth region is extracted; α represents the first cropping ratio in the vertical direction; β represents the second cropping ratio in the vertical direction; γ represents the first cropping ratio in the horizontal direction; δ represents the second cropping ratio in the horizontal direction; V t V represents the video frame at time t in the video stream; t [] indicates the video frame V t Extract a rectangular area from the middle; The feature vector of the i-th person at time t is determined based on the mouth region sequence from frame tk to frame t. The speaking probability of the i-th person in the video frame at time t is determined based on the feature vector of the i-th person at time t, as specifically expressed by the following formula (3): In the above formula (3), Pspeak(i,t) represents the probability of the i-th person speaking in the video frame at time t; σ() represents the Sigmoid function; b represents the preset bias term; W T It is a 256-dimensional weight matrix; The feature vector used to determine the i-th person at time t is a lip fingerprint represented by 256 indicators; k represents the size of the time window. Represents the mouth region sequence from frame tk to frame t; ResNet3D() represents the 3D-ResNet18 network model; R 256 It represents a 256-dimensional real vector, corresponding to 256 indicators describing mouth movements.

7. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 5, characterized in that: The timing of each person's lip-reading movements is compared with the locations of people who made sounds at different times within the sound source location to determine whether the person actually made a sound at the corresponding time, thereby correcting the locations of people who made sounds at different times in the surrounding environment, including: Based on the estimated direction of the sound source at time t estimated by the microphone array, the horizontal deflection angle of the person's face relative to the front of the camera at time t, the horizontal pixel coordinates of the center of the bounding box of the person's face region at time t, the width of the video frame, and the horizontal field of view of the camera, a spatial judgment function G(i,x,t) is constructed, which is specifically expressed as the following formula (4): In the above formula (4), This represents the estimated direction of the x-th sound source at time t, as estimated by the microphone array; ρ represents the bandwidth parameter, which controls the looseness of the matching, and its value is 15°. Used to determine the horizontal deflection angle of the face of the i-th person relative to the front of the camera at time t; represents the horizontal pixel coordinates of the center of the bounding box of the face region of the i-th person at time t; W represents the width of the video frame; Indicates the horizontal field of view of the camera; The spatial judgment function is used to determine whether the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person. The output of the spatial judgment function is in the range of [0, 1]. The closer the output is to 1, the higher the probability that the mouth movements of the i-th person captured by the camera and the position of the x-th sound source received by the microphone array belong to the same person. Based on the activity level of the sound source at time t, the spatial judgment function, and the speaking probability of the person in the video frame at time t, the matching degree between the person and the sound source is determined, specifically expressed as the following formula (5): In the above formula (5), Ci,x represents the matching degree between the i-th person and the x-th sound source; A x (t) represents the activity level of the x-th sound source at time t, and its value ranges from [0, 1]. The larger the value of Ci,x, the higher the probability that the i-th person and the x-th sound source are the same speaker. When Ci,x is greater than or equal to the preset threshold, it is determined that the i-th person and the x-th sound source are in the same location. When Ci,x is less than the preset threshold, proceed to step S6 below; The estimated direction above can be corrected using the following formula (6). In the above formula (6), Indicates the estimated direction The corrected result; λ is the preset fusion weight.

8. The deep learning-based human lip-reading and speech matching and recognition method as described in claim 1, characterized in that: Based on the corrected sound source location, the audio features are separated to obtain the audio features of each person's subordinates, including: Obtain the speech time distribution characteristics and voiceprint characteristics of the person corresponding to the corrected sound source location; Based on the speaking time distribution characteristics and voiceprint characteristics, the audio characteristics of the personnel's subordinates are separated from the audio characteristics using a convolutional neural network model; or, Deep learning is performed on the audio features and lip-sync features of each person to obtain the speech information emitted by each person, including: Deep learning phoneme encoding is performed on the audio features of each person to obtain the phoneme weights corresponding to the audio features; The lip-sync feature is used to train a lip-sync generation model. The trained lip-sync generation model is then input based on the phoneme weights. The original sound data collected by the microphone array is then labeled with the speaker. Based on the speaker tags of the original sound data, the voice information spoken by each person is separated.

9. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 1, characterized in that: The speech information is processed to remove background noise, resulting in the spoken speech and its text content for each person, including: The speech information is subjected to self-adaptive noise suppression by using a weighted overlapping additive filter and sub-band division based on a psychoacoustic model. The speech information is isolated from background noise by spatial masking release processing, thereby extracting the speech of each person; Semantic recognition is performed on the spoken speech to obtain the corresponding text content.

10. The deep learning-based human lip-sync and speech matching and recognition method as described in claim 1, characterized in that: It also includes: transmitting the voice of the corresponding person to the sound playback device worn by the user according to the user's directional voice reception request; And / or, after storing the spoken voice and its text content, the spoken voice and / or the text content are transmitted to the user's device for playback and / or display according to the user's query request.

Citation Information

Patent Citations

  • Animation fusion method and device

    CN114898019A

  • Voice separation method and device, equipment and storage medium

    CN115662463A

  • Speaker detection method, device and equipment and computer readable storage medium

    CN115937726A

  • Voice processing method and device, equipment and storage medium

    CN118447868A

  • Multi-mode speech recognition method, device and equipment and computer readable medium

    CN118748008A