Speaker detection method, device, storage medium and program product

By using lip information to determine the speaker and storing the voiceprint information in the video, the problem of noise-affected sound source positioning is solved, and accurate speaker tracking is achieved when lip information is lost, improving the accuracy and user experience of speaker detection.

CN114513622BActive Publication Date: 2025-08-15ALIBABA (CHINA) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210135075.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-08-15
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

In the prior art, when determining the speaker through sound source positioning, it is susceptible to background noise and room reverb, resulting in a decrease in accuracy of the speaker's position determination and the real speaker cannot be accurately determined under multiple users.

Method used

The speaker is determined through the user's lip information in the video and the voiceprint information is stored. When the lip information is lost, the stored voiceprint information is used to continue tracking the speaker, and combined with voiceprint recognition to improve accuracy.

Benefits of technology

When the lip information is lost, the speaker is continued to be tracked through the voiceprint information, reducing the mislocalization and improving the accuracy and user experience of speaker detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114513622B_ABST
    Figure CN114513622B_ABST
Patent Text Reader

Abstract

The present application provides a speaker detection method, device, storage medium, and program product. The method includes: determining a speaker from at least one user in a video based on lip information of at least one user, and storing the speaker's voiceprint information. If the speaker's lip information is detected to be lost, the speaker in the video is determined using the stored voiceprint information. After determining the speaker based on the lip information, the present application can store the speaker's voiceprint information and, if the lip information is lost, enable voiceprint recognition to continue tracking the speaker. This reduces the possibility of speakers being unable to be correctly located during video playback due to actions such as lowering their head or leaning sideways, thereby improving the accuracy of speaker detection and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a speaker detection method, device, storage medium, and program product. Background Art

[0002] A video conference is a meeting where multiple users in two or more locations engage in face-to-face conversations using communication devices and a network. During a video conference, different users may speak, and when a user speaks, it's necessary to identify the speaker from among the multiple users.

[0003] In the prior art, a speaker is generally determined from multiple users by sound source localization, that is, the speaker's position is determined by the source of the speaker's voice, and then the speaker is determined based on the speaker's position.

[0004] However, when determining the speaker's position, the sound source is easily affected by background noise and room reverberation, which reduces the accuracy of determining the speaker's position. After determining the speaker's position, the position may correspond to multiple users, and it is impossible to accurately determine which user is the real speaker, thereby reducing the accuracy of speaker determination. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to provide a speaker detection method, device, storage medium and program product to improve the accuracy of speaker detection.

[0006] In a first aspect, an embodiment of the present application provides a speaker detection method, comprising:

[0007] Determining a speaker from the at least one user based on lip information of the at least one user in the video, and storing voiceprint information of the speaker;

[0008] If it is detected that the speaker's lip information is missing, the speaker in the video is determined by using the stored voiceprint information.

[0009] Optionally, the method further includes:

[0010] During playback of the video, capturing a facial image of each user in the video;

[0011] At a preset position on the video playback interface, facial images of various users are displayed, and the facial image of the currently determined speaker is identified.

[0012] Optionally, the sound in the video is collected by a microphone array; and determining a speaker from at least one user based on lip information of at least one user in the video includes:

[0013] determining the direction of the sound source through the sound collected by the microphone array;

[0014] Determining the speaker's position range in the video image based on the sound source direction;

[0015] A speaker is determined from the at least one user according to lip information of the at least one user within the position range.

[0016] Optionally, determining a speaker from at least one user in the video based on lip information of the at least one user includes:

[0017] Get video clips from the video through a preset time window;

[0018] For any user, the lip information of the user is extracted from the multiple frames of the video clip, and whether the user is a speaker is determined based on the changes in the lip information in the multiple frames.

[0019] Optionally, the lip information is used to indicate the locations of key points of the lips; and extracting the lip information of the user from multiple frames of the video clip includes:

[0020] Extracting a facial image of the user from each frame of image, inputting the facial image into a facial key point recognition model to obtain a plurality of key points of the face;

[0021] Determine a lip image according to lip key points among multiple key points of the face;

[0022] The lip image is input into a lip key point recognition model to obtain multiple key points of the lip.

[0023] Optionally, determining whether the user is a speaker based on changes in lip information in multiple frames of images includes:

[0024] Determine the lip opening degree corresponding to each frame of the lip image according to multiple key points of the lip;

[0025] determining a lip movement state corresponding to each lip image based on a change in the lip opening degree of the lip image relative to the previous lip image;

[0026] Determine whether the user is a speaker based on lip movement states corresponding to the multiple frames of lip images.

[0027] Optionally, determining whether the user is a speaker based on lip movement states corresponding to multiple frames of lip images includes:

[0028] Inputting the lip movement states corresponding to the multiple frames of lip images into a machine learning model to determine whether the user is a speaker;

[0029] Alternatively, the lip movement states corresponding to the multiple frames of lip images are added together to obtain a cumulative change amount, and the cumulative change amount is compared with a preset threshold to determine whether the user is a speaker;

[0030] The lip movement state of the lip image is specifically the absolute value of the difference between the lip opening degree of the lip image and the lip opening degree of the previous lip image.

[0031] Optionally, the speaker determined from the at least one user is the first user; if it is detected that the lip information of the speaker is lost, determining the speaker in the video by using the stored voiceprint information includes:

[0032] If it is detected that the lip information of the first user is lost, tracking the body of the first user and determining whether the speaker is the first user based on the stored voiceprint information;

[0033] If so, the speaker in the video is determined based on the human body tracking results.

[0034] Optionally, if it is detected that the speaker's lip information is missing, determining the speaker in the video using stored voiceprint information includes:

[0035] When there are multiple current speakers, if it is detected that the lip information of the second user among the current speakers is lost, the voiceprint information corresponding to the speech collected after the lip information is lost is compared with the voiceprint information stored before the loss;

[0036] If the comparison result is inconsistent, it is determined whether the second user is in a state of continuing to speak during the process of losing lip information based on the lip information of other users in the video except the second user.

[0037] In a second aspect, another embodiment of the present application provides a speaker detection method, comprising:

[0038] Obtaining a real-time video stream of the conference, determining a speaker from the at least one user based on lip information of the at least one user in the real-time video stream, and storing voiceprint information of the speaker;

[0039] After storing the voiceprint information, if it is detected that the speaker's lip information is lost, the speaker in the video is determined by using the stored voiceprint information;

[0040] After the speaker is determined through lip information or voiceprint information, the current speaker is marked in the real-time video stream.

[0041] In a third aspect, an embodiment of the present application further provides a target user detection method, comprising:

[0042] Determining at least one user in the video, determining a target user from the at least one user based on image feature information of the at least one user in the video, and storing voice feature information of the target user in the video;

[0043] If it is detected that the image feature information of the current target user is lost, the target user in the video is determined by the stored sound feature information;

[0044] The image feature information includes at least one of the following: lip information, face information, action information, and information about items held by the user.

[0045] In a fourth aspect, an embodiment of the present application provides an electronic device, including:

[0046] at least one processor; and

[0047] a memory communicatively coupled to the at least one processor;

[0048] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method described in any one of the above aspects.

[0049] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method described in any one of the above aspects is implemented.

[0050] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any of the above aspects when executed by a processor.

[0051] The present application provides a speaker detection method, device, storage medium and program product, which can determine the speaker from at least one user based on the lip information of at least one user in the video, and store the voiceprint information of the speaker. When it is detected that the speaker's lip information is lost, the speaker in the video is determined by the stored voiceprint information, thereby realizing the connection detection of lip information and voiceprint information, and can enable voiceprint recognition to continue tracking the speaker after the lip information is lost, reducing the situation in which the speaker cannot be correctly located due to operations such as the speaker lowering his head or turning sideways during the video, thereby improving the accuracy of speaker detection and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0053] Figure 1A schematic diagram of an application scenario provided in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of system interaction provided in an embodiment of the present application;

[0055] Figure 3 A flowchart of a speaker detection method provided in an embodiment of the present application;

[0056] Figure 4 A flowchart of a speaker detection method provided in another embodiment of the present application;

[0057] Figure 5 A schematic diagram of an application for determining the degree of lip opening provided in an embodiment of the present application;

[0058] Figure 6 This is a schematic diagram of the application of lip movement status provided in an embodiment of the present application;

[0059] Figure 7 A schematic diagram illustrating the principle of a speaker determination method provided in an embodiment of the present application;

[0060] Figure 8 A flowchart of a speaker detection method provided in another embodiment of the present application;

[0061] Figure 9 A schematic diagram of an application of the speaker display method provided in an embodiment of the present application;

[0062] Figure 10 This is a schematic diagram of an application of a speaker display method provided in another embodiment of the present application;

[0063] Figure 11 A flowchart of a speaker detection method provided in another embodiment of the present application;

[0064] Figure 12 A schematic diagram illustrating the principle of the speaker detection method provided in an embodiment of the present application;

[0065] Figure 13 A flowchart of a target user detection method provided in another embodiment of the present application;

[0066] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0067] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0068] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.

[0069] First, let’s explain the terms involved in this application:

[0070] Sound source localization: Use a microphone array to calculate the angle and distance of the target speaker, thereby achieving target speaker tracking and subsequent voice directional pickup.

[0071] Neural network model: It is a complex network system formed by a large number of simple processing units (also called neurons) that are widely interconnected. It reflects many basic characteristics of human brain functions and is a highly complex nonlinear dynamic learning system.

[0072] The application scenarios and inventive concepts of this application are explained below.

[0073] The embodiment of the present application can be applied to any scenario where it is necessary to determine the speaker in a video. The video can be a real-time video stream during a live broadcast, or a pre-stored or downloaded video.

[0074] In a video conferencing scenario, there are often multiple participants. For example, a meeting may be set up with multiple venues, and each venue will film at least one participant. Alternatively, although the meeting is set up with only one or two venues, there are multiple participants in the venue. In this case, it is necessary to determine the current speaker in real time during the meeting, mark the current speaker in the video, or switch the screen to the current speaker, so that other participants can promptly understand who is currently speaking.

[0075] In entertainment scenarios, such as large-scale evening parties, the director is usually required to switch the screen to the current speaker, especially in language programs or occasions with multiple hosts. The screen can track the person who is speaking in real time. The speaker detection method provided in the embodiment of the present application can locate the current speaker in time and assist in achieving screen switching.

[0076] In the live broadcast scenario of products, especially the live broadcast scenario of multiple anchors, the speaker detection method provided in the embodiment of the present application can also be used to detect and switch the screen to the anchor currently speaking to facilitate the audience's viewing. It is also possible to pin the corresponding product link to the top or display the product currently being introduced in the video based on the product held by the anchor currently speaking, thereby improving the user experience.

[0077] For example, a video conference involves at least two access points, each of which has at least one user, and each user can speak during the conference. To ensure that each user in the conference clearly understands who is speaking, it is necessary to identify the speaker from among multiple users.

[0078] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present application. Figure 1 As shown, in a video conference, there may be multiple users at the conference site, including user A, user B, etc., and it is necessary to detect in real time which user is speaking.

[0079] Some technologies can use sound source localization (e.g., microphone array-based sound source localization) to identify the speaker in real time from multiple users. This involves determining the speaker's location based on the source of their voice, and then determining the speaker based on their location. However, when determining the speaker's location, the sound source is easily affected by background noise and room reverberation, reducing the accuracy of speaker location determination. Furthermore, when the speaker and other participants are close to or overlapping in the vertical direction of the sound source, sound source localization is prone to errors.

[0080] For example, when user A is speaking, the microphone array in the terminal device can locate the speaker through sound source localization. Since the output result of the sound source localization method locates the speaker's position along the longitudinal direction of the speaker's position, user B is located on one side of user A and intersects with user A in the longitudinal direction of their position. Therefore, when determining the speaker through sound source localization, it is impossible to determine whether the speaker is user A or user B. That is, after the speaker's position is determined, the position may correspond to multiple users, and it is impossible to accurately determine which user is the real speaker, thereby reducing the accuracy of speaker determination.

[0081] In other technologies, the speaker can also be determined through the lip information of each participant during a video conference. However, after the speaker is determined through lip information, the speaker may perform actions such as turning sideways or lowering the head while speaking, resulting in the loss of the speaker's lip information, making it impossible to determine the speaker, or mislocating the speaker to another user, reducing the accuracy of speaker determination.

[0082] Based on the above technical problems, the present application stores the speaker's voiceprint information after determining the speaker based on the lip information, and enables voiceprint recognition to continue tracking the speaker after the lip information is lost. This avoids the situation in which the speaker is mislocated to other users due to operations such as lowering the head or turning sideways during the video process, thereby improving the accuracy of speaker determination and thus improving the user's application experience. In addition, there is no need to pre-store the voiceprint information of each user before the video conference, which facilitates promotion and application.

[0083] Figure 2 This is a schematic diagram of a system interaction provided in an embodiment of the present application. Figure 2 As shown, the conferencing system may include multiple terminal devices. When conducting a video conference through the terminal devices, there may be at least two access points, each corresponding to a terminal device, and each access point having at least one user. During the video conference, users at each access point can speak. The terminal device can capture a video stream in real time and identify the speaker from the at least one user based on lip information of the at least one user in the real-time captured video stream.

[0084] In addition, after the speaker is determined, the speaker's voiceprint information can be stored. Later, when it is detected that the speaker's lip information is lost, the stored voiceprint information can be used to determine the speaker in the video, further improving the accuracy of speaker determination.

[0085] The terminal device can be a device capable of video conferencing, such as a smartphone, tablet, personal computer, or smart wearable device. Terminal devices at different access points can connect to the cloud, which can obtain real-time video streams captured by each terminal device through the cloud. The cloud can then determine the speaker from the at least one user based on lip information of the at least one user in the real-time video stream.

[0086] Furthermore, the speaker detection method provided in this application can also be applied in other scenarios. For example, based on the speaker recognition results, speech separation can be performed after a video conference to obtain speech data corresponding to different speakers. Alternatively, it can be applied to character recognition scenarios such as film and television dramas and short videos. By identifying the speaker, each character's speech segment can be segmented and used as video material.

[0087] The following detailed description of some embodiments of the present application is provided in conjunction with the accompanying drawings. The following embodiments and features thereof may be combined with one another unless they conflict with each other. Furthermore, the sequence of steps in the following method embodiments is provided for illustrative purposes only and is not intended to be a strict limitation.

[0088] Figure 3This is a flow chart of a speaker detection method provided in an embodiment of the present application. The execution subject of the method in this embodiment can be any device with data processing function, such as a terminal device or a server. Figure 3 As shown, the method may include:

[0089] Step 301: Determine a speaker from at least one user based on lip information of at least one user in a video, and store the speaker's voiceprint information.

[0090] In this embodiment, after acquiring a video, lip information of the users involved in the video can be extracted from the video, where the users involved can be one or more users appearing in the video. Then, the speaker can be determined from the users involved based on the lip information of the users involved.

[0091] In addition, the lip information can be used to indicate the locations of key points of the lips, which can be a sequence of multiple coordinate values.

[0092] The video may be acquired from different scenarios. In one example, the video may be a real-time video stream. For example, the video may be acquired in real time from a video conference scenario. That is, during a video conference, video stream data is acquired in real time, and then lip information of the users involved is extracted from the video stream data. The speaker is then determined from at least one user based on the extracted lip information of the user.

[0093] In another example, the video may be a non-real-time video stream, such as a TV series, short video, or other data stored locally on the device or retrieved from the cloud. Lip information of the users involved in the TV series, short video, or other data may then be extracted, and the speaker (i.e., different characters) may be determined from at least one user based on the extracted lip information. Subsequently, the speech segments of each character may be segmented based on the determined speaker and used as video material.

[0094] Step 302: If it is detected that the speaker's lip information is missing, the speaker in the video is determined through the stored voiceprint information.

[0095] In this embodiment, after the speaker is identified from at least one user, the speaker's voiceprint information can be obtained and stored. Subsequently, due to factors such as the speaker lowering their head or leaning to the side, the speaker's lip information may be lost. If the speaker's lip information is detected to be missing, the stored voiceprint information can be used to determine the speaker in the video. This avoids the situation where the speaker's lip information is missing and cannot be accurately determined, thereby improving the accuracy of speaker identification.

[0096] Furthermore, loss of a speaker's lip information refers to the inability to detect the speaker's lip information, which can also be understood as the inability to identify the speaker based on the lip information. Alternatively, when a speaker's lip information is determined using a model or algorithm, loss of lip information can mean that the model or algorithm cannot determine the lip information, or that the confidence level of the lip information determined by the model or algorithm is less than a certain threshold.

[0097] Furthermore, obtaining and storing the speaker's voiceprint information can specifically include: after determining the speaker through lip information, a speaker identifier (e.g., a speaker ID) can be created, and the voiceprint information and speaker identifier are associated to establish a speaker identifier library, wherein the speaker identifier library includes the voiceprint information corresponding to each speaker identifier. In addition, the speaker identifier can include at least one of uppercase letters, lowercase letters, numbers, and special symbols, which are not further limited here.

[0098] In addition, when it is detected that the speaker's lip information is lost and the speaker in the video cannot be determined through the stored voiceprint information, the lip information of at least one user can be obtained again from the video, and based on the lip information of at least one user in the video, a new speaker can be determined from at least one user, and the voiceprint information of the new speaker can be stored.

[0099] In summary, the speaker detection method provided in the embodiment of the present application can determine the speaker from at least one user based on the lip information of at least one user in the video, and store the voiceprint information of the speaker at the same time. When it is detected that the speaker's lip information is lost, the speaker in the video is determined by the stored voiceprint information, thereby realizing the connection detection of lip information and voiceprint information, and enabling voiceprint recognition to continue tracking the speaker after the lip information is lost, reducing the situation where the speaker cannot be correctly located due to operations such as the speaker lowering his head or turning sideways during the video, thereby improving the accuracy of speaker detection and improving user experience.

[0100] Figure 4 This is a flow chart of a speaker detection method provided by another embodiment of the present application. Figure 4 As shown, the sound in the video is collected by a microphone array. The step 301 of determining the speaker from at least one user based on lip information of at least one user in the video may include:

[0101] Step 401: Determine the direction of the sound source by collecting the sound through the microphone array.

[0102] Step 402: Determine the speaker's position range in the video image based on the sound source direction.

[0103] Step 403: Determine a speaker from the at least one user based on the lip information of the at least one user within the location range.

[0104] In this embodiment, when determining the speaker based on lip information in a video, since the video may involve multiple users, lip information must be extracted for each user, and then the speaker is determined based on the lip information of each user. This increases the speaker determination time and reduces the speaker determination efficiency. To address this issue, the speaker's position range in the video image (i.e., the speaker's approximate location) can be determined through sound source localization. Then, at least one candidate corresponding to the position range is obtained, and the speaker is determined from the candidates based on the lip information of the at least one candidate.

[0105] Furthermore, the specific process of determining the position range of the speaker in the video image by sound source localization can be: first determine the sound source direction through the sound collected by the microphone array, then determine the position range of the speaker in the video image corresponding to the sound source direction based on the pre-stored mapping relationship between the sound source direction and the image position in the video, and determine at least one user corresponding to the position range of the speaker in the video image, and then determine the speaker based on the lip information of the at least one user corresponding to the position range.

[0106] In summary, although the location range may correspond to multiple users, the number of users has been greatly reduced compared to all users involved in the video, thereby shortening the speaker identification time and improving the efficiency and accuracy of speaker recognition.

[0107] Optionally, the mapping relationship between the sound source direction and the image position in the video can be determined by the relative position relationship between the microphone array and the video shooting device, which is not limited in detail here.

[0108] In addition, in another embodiment, based on the lip information of at least one user in the video, determining the speaker from the at least one user can specifically include: obtaining a video clip in the video through a preset time window; for any user, extracting the lip information of the user from multiple frames of images in the video clip, and determining whether the user is the speaker based on the changes in the lip information in the multiple frames.

[0109] In this embodiment, when determining the speaker from at least one user in a video based on their lip information, to reduce the device's processing workload, video clips can be captured from the video using a preset time window at preset intervals. Lip information for each user can then be extracted from multiple frames of the video clips. Based on the changes in the lip information across the multiple frames, the speaker's identity can be determined. The preset duration and time window can be customized based on the actual application scenario, minimizing device processing overhead while ensuring accurate speaker identification. For example, the preset duration can be any value between 1 and 3 seconds, and the time window can be any value between 1 and 3 seconds, with 5 to 10 frames per second.

[0110] In summary, by setting the time window, the video can be divided into multiple groups of video clips, and then the changes in lip information can be determined based on each group of video clips to determine whether the user is the speaker. This reduces the workload of the device processing and improves the efficiency of speaker determination.

[0111] In one or more embodiments of the present application, the lip information is optionally used to indicate the locations of lip key points. Extracting the user's lip information from multiple frames of the video clip may include: extracting the user's facial image from each frame, inputting the facial image into a facial key point recognition model to obtain multiple facial key points; determining a lip image based on the lip key points among the multiple facial key points; and inputting the lip image into a lip key point recognition model to obtain multiple lip key points.

[0112] Specifically, after obtaining multiple frames of video clips, the user's face image can be extracted from each frame, and then the face image is input into a pre-trained facial key point recognition model for recognition to obtain multiple key points of the face.

[0113] Optionally, after extracting the user's facial image from each frame image, some facial images may be small, affecting the subsequent recognition accuracy. Therefore, after extracting the facial image, the facial image can be enlarged to a uniform size, and then the enlarged facial image can be input into the facial key point recognition model to obtain multiple key points of the face, so as to make the positioned key points more accurate, thereby improving the positioning accuracy of the facial image.

[0114] In addition, when extracting the user's facial image from each frame image, it can be the user's facial image that has been screened according to the sound source positioning method, that is, the user's facial image is determined in the position range of the speaker in the video image based on the sound source direction.

[0115] Furthermore, after determining multiple facial key points, lip key points can be determined from these key points to obtain the user's lip information. However, the lip key points determined from these key points are only rough lip locations, and using these lip key points as lip information to determine the speaker may reduce the accuracy of speaker identification. Therefore, a two-stage extraction approach can be adopted. After determining the lip key points from the multiple facial key points, further extraction can be performed from the lip key points to obtain more accurate lip key points as lip information. The number of key points can be customized based on the actual application scenario and is not specifically limited here.

[0116] Optionally, a lip image may be determined based on the lip key points among the multiple key points of the face, and then the lip image may be input into a lip key point recognition model to obtain multiple key points of the lip.

[0117] For example, the key points of the inner lips can be located during the secondary key point extraction. Accordingly, the facial image can be input into a facial key point recognition model to obtain multiple key points of the face. Then, 20 pairs of lip key points can be extracted from the multiple key points of the face to determine the lip image. The lip image can then be input into a lip key point recognition model to obtain three pairs of key points of the inner lips.

[0118] Optionally, the facial key point recognition model and the lip key point recognition model can be obtained by training a neural network model or other network models, and no specific limitation is made here.

[0119] In summary, the secondary extraction method is used to determine the key points of the lip, which improves the accuracy of lip information recognition and thus improves the accuracy of speaker determination.

[0120] In one or more embodiments of the present application, optionally, determining whether the user is a speaker based on changes in lip information in multiple frames of images may include: determining the degree of lip opening corresponding to each frame of lip image based on multiple key points of the lips; determining the lip movement state corresponding to each lip image based on changes in the degree of lip opening of the lip image relative to the previous lip image; and determining whether the user is a speaker based on the lip movement states corresponding to multiple frames of lip images.

[0121] Specifically, when the user speaks, the degree of lip opening changes greatly, and when the user does not speak, the degree of lip opening changes less. Therefore, for each frame of lip image, after determining multiple key points of the lips, the degree of lip opening corresponding to each frame image can be determined based on the multiple key points of the lips, which can be used as a basis for judging whether the user is speaking.

[0122] Further, Figure 5This is a schematic diagram of an application for determining the degree of lip opening provided in an embodiment of the present application, such as Figure 5 As shown in the figure, in this embodiment, P2, P3, P4, P6, P7, and P8 are three pairs of lip key points extracted in the second stage, and P1 and P5 are the positions of the mouth corners. A coordinate axis can be established based on P1, P5, P3, and P7. The coordinate values of P2, P3, P4, P6, P7, and P8 are then determined based on the coordinate axis. The lip opening degree corresponding to each frame of the image is then determined according to the expression ratio = |||P2-P8||+||P3-P7||+||P4-P6||, where |||| represents the 2-normal form.

[0123] After determining the lip opening degree corresponding to the lip image in each frame, the lip movement state corresponding to each lip image can be determined based on the change in lip opening degree of each lip image relative to the previous lip image. The lip movement state represents the degree of change in lip opening. When speaking, people open and close their lips, and the lip state varies significantly. Therefore, the lip movement state can be used to improve the accuracy of speaker identification. Optionally, the lip movement state of a lip image is specifically the absolute value of the difference between the lip opening degree of the lip image and the lip opening degree of the previous lip image.

[0124] Optionally, when determining the lip movement state corresponding to a lip image, the lip movement state corresponding to the lip image can also be determined based on the change in the degree of lip opening of each lip image relative to the lip images of the previous N frames (for example, the previous 2 frames, the previous 3 frames, etc.), thereby reducing the amount of computational processing required by the device.

[0125] For example, Figure 6 The application diagram of the lip movement state provided in the embodiment of the present application is as follows Figure 6 As shown in this embodiment, the X-axis represents the number of frames, and the Y-axis represents the lip movement status value of each lip image frame. The left half of the figure shows the lip movement status when not speaking, while the right half shows the lip movement status when speaking. It can be seen that there is a significant difference between the two. Therefore, based on the lip movement status value, it is possible to clearly determine whether the user is speaking, and based on this state, determine whether the user is speaking.

[0126] Optionally, the lip movement states corresponding to multiple frames of lip images can be input into a state machine to learn the distribution of speaking states and non-speaking states. The state machine can then be used to determine the speaker in the video.

[0127] In summary, lip movement can more accurately reflect whether the user is speaking than the static lip opening degree, further improving detection accuracy.

[0128] In one implementation, determining whether the user is a speaker based on the lip movement states corresponding to multiple frames of lip images may include: inputting the lip movement states corresponding to multiple frames of lip images into a machine learning model to determine whether the user is a speaker.

[0129] Specifically, the lip movement states corresponding to multiple frames of lip images can be directly input into a pre-trained machine learning model to determine whether the user is a speaker.

[0130] Optionally, multiple frames of lip images, or the lip opening degrees corresponding to multiple frames of lip images, may be directly input into a pre-trained machine learning model to determine whether the user is a speaker.

[0131] Among them, the machine learning model can be pre-trained through a neural network model or other network models, and is not specifically limited here.

[0132] By adopting a machine learning model to determine whether the user is the speaker, the efficiency and accuracy of speaker determination are improved.

[0133] In another implementation, determining whether the user is a speaker based on the lip movement states corresponding to multiple frames of lip images can include: adding the lip movement states corresponding to multiple frames of lip images to obtain a cumulative change, and comparing the cumulative change with a preset threshold to determine whether the user is a speaker.

[0134] Specifically, Figure 7 This is a schematic diagram of the principle of the speaker determination method provided in the embodiment of the present application, such as Figure 7 As shown, in this embodiment, a preset threshold number of lip image frames (for example, 10 lip image frames) can be grouped together, and a cumulative change is obtained for each group. The cumulative change is then compared with a preset threshold to determine whether the user is a speaker. Specifically, a preset number of lip image frames can be acquired first, and then the lip movement states corresponding to the preset number of lip images can be summed to obtain a cumulative change. The cumulative change is then compared with a preset threshold. If the cumulative change is greater than or equal to the preset threshold, the user is determined to be a speaker, indicating successful speaker location. If the cumulative change is less than the preset threshold, the user is determined not to be a speaker, indicating a failed speaker location, and a new preset number of lip image frames can be acquired for processing. The preset threshold can be customized based on the actual application scenario and is not further defined here.

[0135] Optionally, an expression can be passed: To determine the cumulative amount of change.

[0136] Among them, diS i Indicates the lip ratio of the current frame, pdiSi represents the lip ratio of the previous frame, deIta_diS represents the cumulative change of each group of lip images, and k represents the number of frames of lip images in each group of lip images.

[0137] In summary, whether the user is a speaker is determined based on the relationship between the cumulative amount of changes in lip movement states corresponding to multiple frames of lip images and a preset threshold. This comprehensively considers the actual changes in the user's lips when speaking, thereby improving the accuracy of speaker determination. The calculation is simple, effectively improving detection efficiency.

[0138] Figure 8 A flow chart of a speaker detection method provided in another embodiment of the present application is shown as follows: Figure 8 As shown, the method may further include:

[0139] Step 801: During the video playback process, capture the facial image of each user in the video.

[0140] In this embodiment, during video playback, facial images of each user in the video can be captured using existing image capture methods. Information such as the size and clarity of the captured facial images can be automatically determined based on the image capture method. This can also be customized by the user based on actual circumstances.

[0141] Step 802: Display the facial images of each user at a preset position on the video playback interface, and identify the facial image of the currently determined speaker.

[0142] In this embodiment, after capturing the facial images of each user in the video, the facial images of each user can be displayed at a preset location on the video playback interface. Furthermore, to enable each user to intuitively identify the speaker, the facial image of the currently identified speaker can be labeled. For example, the lips and / or small microphone can be displayed below the speaker icon, or directly on the speaker icon.

[0143] Optionally, instead of displaying the captured facial images of each user in the playback interface, the speaker can be displayed directly in the playback video interface. For example, a small microphone can be displayed next to the speaker and move with the speaker.

[0144] Optionally, you can also display only the speaker's video in the playback interface, omitting the others. When the speaker changes, the interface switches to the other speaker. Alternatively, if the conference has multiple venues, only the venue of the current speaker is displayed.

[0145] For example, Figure 9 This is a schematic diagram of the application of the speaker display method provided in the embodiment of the present application, such as Figure 9As shown, in this embodiment, five users may be included, namely user A, user B, user C, user D, and user E. After the facial images of the five users are captured in the video, the facial images of the five users can be displayed at the bottom of the video playback interface. In addition, if user A is the speaker, the lips and a small microphone can be displayed below the user A icon to remind other users that the speaker is user A.

[0146] In summary, by marking the currently determined speaker in the playback interface, each user can intuitively determine the speaker, thereby improving the user's application experience.

[0147] Furthermore, in another embodiment, the speaker determined from the at least one user is the first user. If it is detected that the lip information of the speaker is missing, determining the speaker in the video using the stored voiceprint information may include:

[0148] If it is detected that the lip information of the first user is lost, the body of the first user is tracked, and the stored voiceprint information is used to determine whether the speaker is the first user; if so, the speaker in the video is determined based on the body tracking result.

[0149] In this embodiment, when the first user is speaking, they may lower their head or lean sideways, causing lip information to be lost or key points to be inaccurately located. Voiceprint comparison can be performed using the pre-stored voiceprint information of the first user to determine whether the speaker is the first user. This reduces the possibility of mislocating other users when the speaker's lip information is lost. Furthermore, voiceprint comparison can only be used to determine whether the first user is still speaking. However, the first user is constantly active, and it is difficult to accurately locate the speaker in the video and capture the speaker's image for display based solely on voiceprints. Therefore, body tracking of the first user in the video can also be performed to determine the speaker's location in the video. This allows the user to display a box containing the speaker or similar instructions in real time, allowing the user to accurately identify the specific speaker and improve the user experience.

[0150] Optionally, there are multiple ways to track the body of a user in a video. For example, when storing the speaker's voiceprint information, the speaker's facial information or body information can be stored at the same time. Later, the first user can be tracked based on the stored speaker's facial information or body information to determine the speaker's position in the video. The viewer can also be shown a box where the speaker is located or given similar instructions to enable the user to accurately determine the specific speaker. Among them, the process of tracking the first user based on the stored speaker's facial information or body information can be achieved through a pre-trained body tracking model. In addition, this application only lists one specific method that can be implemented. Other methods of implementing body tracking are also within the scope of protection of this application and will not be discussed in detail here.

[0151] In summary, when identifying the speaker through voiceprint information, combined with body tracking, comprehensive identification can be performed. When the user moves, the speaker can also be identified to the user in real time, allowing the user to accurately determine the specific speaker, thereby improving the user experience.

[0152] Optionally, the method may further include: in the process of determining the speaker through voiceprint information, if the lip information of the first user is re-detected, determining whether the speaker is the first user through the voiceprint information and lip information; if the speaker is determined to be the first user through voiceprint information and lip information for a continuous preset time, switching to using lip information to identify the speaker.

[0153] In this embodiment, it is difficult to locate the speaker's position in the image using voiceprint information alone, so lip information is generally used to determine the speaker under normal circumstances. Therefore, in the process of determining the speaker using voiceprint information, the lip information of the first user can be re-detected at preset intervals. If the lip information of the first user can be re-detected, that is, after the lip information is retrieved, the voiceprint information can be discarded and the lip information can continue to be used for speaker recognition. In addition, when the lip information is just retrieved, it can be fused for a period of time (that is, the voiceprint information and lip information are used to jointly determine whether the speaker is the first user). If the speaker determined by the voiceprint information and lip information is consistent for a consecutive preset period of time, the lip information can be switched to be used for speaker recognition, thereby reducing the error rate of speaker recognition.

[0154] In summary, when retrieving lip information, the speaker is first determined by fusing the voiceprint information and the lip information. After the lip information is stable, the speaker is determined solely by the lip information. This avoids frequent switching between lip information and voiceprint information for identification, thereby improving the user experience.

[0155] Optionally, determining whether the speaker is the first user through voiceprint information and lip information may include: comparing the speaker's voiceprint information with the stored voiceprint information of the first user to determine the confidence that the speaker is the first user; determining the confidence that the speaker is the first user based on the lip information of the first user; performing a weighted summation of the confidence determined through the voiceprint information and the confidence determined through the lip information, and determining whether the speaker is the first user based on the obtained result.

[0156] Specifically, when comprehensively determining whether the speaker is the first user through voiceprint information and lip information, a confidence level can be determined respectively through the voiceprint information and the lip information, and then different weight values are set for the confidence level determined by the voiceprint information and the confidence level determined by the lip information, respectively. According to the set weight values, the confidence level determined by the voiceprint information and the confidence level determined by the lip information are weighted and summed, and whether the speaker is the first user is determined based on the obtained result.

[0157] Optionally, after weighted summing of the confidence determined by the voiceprint information and the confidence determined by the lip information is performed and the speaker is determined to be the first user based on the obtained result, the speaker (ie, the first user) can be tracked and displayed in real time. Figure 10 This is a schematic diagram of an application of a speaker display method provided in another embodiment of the present application, such as Figure 10 As shown, in this embodiment, continue with Figure 9 Taking the embodiment in as an example, the first user is user A. The speaker recognition result can be determined based on the sound source localization method and the lip information detection method. Then, when the lip information is lost, the speaker recognition result can be determined by voiceprint recognition. When determining the speaker through voiceprint recognition, it can be determined whether the lip information is detected at preset intervals. If it is determined that the lip information can be detected, the confidence level determined by the voiceprint information and the confidence level determined by the lip information can be weighted and summed to comprehensively determine the speaker recognition result (i.e., the speaker is user A). The lips and small microphone can then be displayed under the user A icon to remind other users that the speaker is user A.

[0158] Optionally, if the stored voiceprint information determines that the speaker is not the first user, the acquired voiceprint information can be compared with the stored voiceprint information of other users to determine whether it is another user who has spoken. If the comparison with the stored voiceprint information of other users determines that the current speaker is a second user, body tracking of the second user can be performed using facial information or body language. Specifically, the second user can be located in the video image and a box or similar indication of the speaker can be displayed to the viewer in real time. Furthermore, during the process of determining the speaker using voiceprint information, if the second user's lip information is re-detected, the speaker can be determined to be the second user using a combination of voiceprint information, lip information, and body tracking. After the speaker has been confirmed to be the second user using voiceprint information, lip information, and body tracking for a predetermined period of time, speaker identification can be switched to using lip information. If the current speaker cannot be determined using the stored voiceprint information of other users, lip information of each user in the video can be re-acquired, and the speaker can be determined from at least one user based on the lip information of at least one user in the video, and the speaker's voiceprint information can be stored.

[0159] In summary, by setting different weights for the confidence determined by the voiceprint information and the confidence determined by the lip information, the determination results of the two methods are comprehensively considered, thereby further improving the accuracy of the first user determination.

[0160] In one or more embodiments of the present application, there may be multiple users speaking at the same time. Optionally, if it is detected that the lip information of the speaker is lost, determining the speaker in the video through the stored voiceprint information may include: when there are multiple current speakers, if it is detected that the lip information of the second user among the current speakers is lost, comparing the voiceprint information corresponding to the speech collected after the lip information is lost with the voiceprint information stored before the loss; if the comparison result is inconsistent, determining whether the second user was in a state of continuing to speak during the process of lip information loss based on the lip information of the remaining users in the video except the second user.

[0161] Optionally, when multiple people are speaking simultaneously, the stored voiceprint information can be a mixture of multiple voiceprints. If the lip information of one or more speakers is lost, the voiceprint information corresponding to the current speech can be compared with the voiceprint information stored before the loss. If they match, it is determined that the speaker has not changed. If the speaker whose lip information is lost is recorded as the second user, the second user will continue to speak during the period when the lip information is lost.

[0162] If they are inconsistent, it is considered that the current speaker has changed, perhaps by adding, removing, or replacing a speaker. In this case, the status of the other users can be used to determine whether the second user whose lip information was lost is still speaking. If the status of the other users does not change before and after the second user's lip information is lost, it can be assumed that the second user has stopped speaking.

[0163] For example, users A, B, C, D, and E participate in a meeting. During the meeting, it is detected that users A, B, and C are speaking at the same time. Therefore, the voiceprint information corresponding to users A, B, and C can be stored. After user C's lip information is lost, if the voiceprint information corresponding to the collected speech is consistent with the previously stored voiceprint information, it means that users A, B, and C are still speaking. If they are inconsistent, the lip information of other users is recognized. If it is determined that users A and B are still speaking, and users D and E are still not speaking, it means that user C has stopped speaking. At this time, the current speakers are updated to include only users A and B.

[0164] In addition, when multiple people are speaking at the same time, technologies such as sound source direction can still be used for auxiliary identification. For example, the sound source direction, face recognition, lip information and voiceprint recognition can be used in sequence to make judgments and finally locate the speaker. The specific implementation principle is similar to the aforementioned embodiment, and the steps are not repeated here.

[0165] In summary, when the lip information of some speakers among multiple speakers is lost, voiceprint comparison and lip information recognition of other users can effectively determine whether the speaker whose lip information is lost has stopped speaking, thereby improving the recognition accuracy when multiple people are speaking at the same time and meeting the needs of meetings of different specifications.

[0166] In one or more embodiments of the present application, optionally, when multiple people are speaking at the same time, each speaker may be identified, for example, by displaying lips and / or a small microphone at each speaker icon.

[0167] When there are a large number of speakers, the identification strategy can be adjusted. Optionally, if multiple speakers are detected, the N most recent speakers are identified. Furthermore, when the number of speakers exceeds M, all speakers are unidentified. N and M can be set based on actual needs, for example, N = 2 and M = 4. This means that a maximum of two speakers can always be identified. If a third speaker speaks, the earliest speaker is replaced, retaining the status of two speakers. No identification is performed after more than four speakers have spoken.

[0168] For example, user A speaks first, then user B joins, and users A and B are in a state of speaking at the same time. At this time, small microphones are displayed under the icons of users A and B. Then user C joins, and users A, B, and C speak at the same time. At this time, user A can be unidentified and only users B and C can be identified. If users D and E also join at this time, and there are more than 4 users speaking at the same time, no user can be identified.

[0169] In summary, by limiting the number of speakers identified at the same time, users watching the video can focus on the person who started speaking most recently. Moreover, by canceling the identification of all speakers after the number of speakers exceeds a certain value, it is possible to avoid user perception confusion caused by too many identifications and improve the user experience.

[0170] Figure 11 This is a flow chart of a speaker detection method provided by another embodiment of the present application. The execution subject of the method in this embodiment can be any device with data processing function, such as a terminal device or a server. Figure 11 As shown, the method may include:

[0171] Step 1101: Acquire a real-time video stream of a conference, determine a speaker from at least one user based on lip information of at least one user in the real-time video stream, and store the speaker's voiceprint information.

[0172] Step 1102: After storing the voiceprint information, if it is detected that the speaker's lip information is lost, the speaker in the video is determined using the stored voiceprint information.

[0173] Step 1103: After the speaker is determined through the lip information or voiceprint information, the current speaker is marked in the real-time video stream being played.

[0174] This embodiment can be applied in a video conferencing scenario, where there are at least two access points, each corresponding to a terminal device, and each access point has at least one user. At the start of the conference, the real-time video stream of the conference can be directly acquired. Then, based on the lip information of at least one user in the real-time video stream, the speaker can be identified from the at least one user, and the speaker's voiceprint information can be stored. After the voiceprint information is stored, if the speaker's lip information is lost during the conversation, the stored voiceprint information can be used to determine the speaker in the video. Furthermore, after the speaker is identified using lip information or voiceprint information, the current speaker can be annotated in the live video stream. For example, the lips and / or small microphone can be displayed below the speaker icon, or directly above the speaker icon. Alternatively, instead of displaying the captured facial images of each user in the playback interface, the speaker can be displayed directly in the playback video interface. For example, a small microphone can be displayed next to the speaker, and follow the speaker's movements.

[0175] Optionally, after acquiring the real-time video stream of the conference, at least one candidate speaker may be determined by sound source localization, and then the speaker may be determined from the candidate speakers based on the lip information of the at least one candidate speaker.

[0176] Optionally, after acquiring the real-time video stream of the conference, video stream packets of a preset duration can be extracted for processing. For example, the preset duration can be 1 second, with 10 frames per second. Time T0 can represent the start of the video conference. Starting from time T0, candidate speakers are first identified through sound source localization. Then, facial images of each candidate speaker in frames 1 through 10 are extracted. Based on the facial images, lip information of each candidate speaker is determined. Further identification is performed based on the lip information of each candidate speaker to determine the speaker. Images from frames 11 through 20 can then be extracted and the aforementioned process repeated. Furthermore, after the speaker is identified, voiceprint information of the speaker can be extracted every 2 seconds and stored in a voiceprint database. If the speaker's lip information is subsequently detected to be missing at time T1, the acquired voiceprint information can be compared with the latest voiceprint information stored in the voiceprint database to determine whether the current speaker is the previously identified speaker.

[0177] Figure 12 The schematic diagram of the principle of the speaker detection method provided in the embodiment of the present application is as follows: Figure 12 As shown, in this embodiment, the sound can be first determined from the acquired real-time video stream, and then the sound can be identified by sound source localization to obtain at least one candidate speaker, and then the lip information of the at least one candidate speaker can be extracted, and the speaker can be judged based on the lip information of the at least one candidate speaker. If the target candidate speaker is determined to be the speaker, the speaker is identified on the playback interface; if the target candidate speaker is determined not to be the speaker, the speaker is not identified on the playback interface, and the speaker is judged based on the lip information of the next candidate speaker. By combining the sound source localization method and the lip information to determine the speaker in the video conference, the accuracy of speaker determination in the video conference is improved. By identifying the speaker in the video interface, each user in the video conference can accurately understand the speaker, thereby improving the user's conference experience.

[0178] Figure 13 This is a flow chart of a target user detection method provided by another embodiment of the present application. The execution subject of the method in this embodiment can be any device with data processing function, such as a terminal device or a server. Figure 13 As shown, the method may include:

[0179] Step 1301: Determine at least one user in the video, determine a target user from the at least one user based on image feature information of the at least one user in the video, and store sound feature information of the target user in the video.

[0180] Step 1302: If it is detected that the image feature information of the current target user is lost, the target user in the video is determined through the stored sound feature information.

[0181] The image feature information includes at least one of the following: lip information, face information, action information, and information about items held by the user.

[0182] In this embodiment, the video can also be a video in an entertainment scene (for example, multi-person singing, mixed instrument performance, singing and dancing performance, etc.) or a video in a product live broadcast scene. In different scenes, the image feature information corresponding to the user may be different.

[0183] For example, in a multi-person singing scene, the image feature information corresponding to the user may be lip information, and the target user is the person currently singing. By detecting the connection between the lips and the voice, the current singer can be accurately located in real time.

[0184] In a mixed instrument performance scenario, the image feature information corresponding to the user may be hand movement information, and the target user is the person currently playing the instrument. By detecting the connection between hand movements and sounds, the currently playing instrument can be accurately located in real time and directed to the corresponding screen.

[0185] In a singing and dancing performance scenario, the image feature information corresponding to the user may be body movement information, and the target user is the person currently performing, so that the camera can focus on the current singer and dancer while blurring the rest of the people waiting still, thereby improving the stage effect.

[0186] In the live broadcast scenario of products, the image feature information corresponding to the user may be the information of the item held by the user, and the target user is the anchor who is introducing the product, so that the audience can focus on the product currently being introduced. In some cases, when the product is blocked, sound detection can be enabled to prevent the screen from switching to other anchors, thereby improving the audience experience.

[0187] In summary, the method provided in this embodiment can first determine the target user from at least one user based on the image feature information of at least one user in the video, and store the sound feature information of the target user in the video. Then, when it is detected that the image feature information of the current target user is lost, the target user in the video is determined by the stored sound feature information, which reduces the situation where the target user cannot be correctly located due to the loss of image feature information, improves the accuracy of target user detection, and improves the user experience.

[0188] The various methods provided in the embodiments of this application can be applied to both the server and the terminal device, or some steps can be deployed on the server and some steps can be deployed on the terminal device. The implementation schemes provided in the various embodiments of this application can be referenced to each other and will not be described in detail.

[0189] Corresponding to the above method, an embodiment of the present application further provides a speaker detection device, the speaker detection device comprising:

[0190] The first determination module is configured to determine a speaker from at least one user in the video based on lip information of the at least one user, and store the voiceprint information of the speaker.

[0191] The first determination module is further configured to determine the speaker in the video using stored voiceprint information if it is detected that the speaker's lip information is missing.

[0192] Optionally, the device further includes a first display module, and the first display module is configured to:

[0193] During the playback of the video, a facial image of each user in the video is captured.

[0194] At a preset position on the video playback interface, facial images of various users are displayed, and the facial image of the currently determined speaker is identified.

[0195] Optionally, the sound in the video is collected by a microphone array; and the first determining module is further configured to:

[0196] The direction of the sound source is determined by the sound collected by the microphone array.

[0197] The position range of the speaker in the video image is determined according to the sound source direction.

[0198] A speaker is determined from the at least one user according to lip information of the at least one user within the position range.

[0199] Optionally, the first determining module is further configured to:

[0200] Get video clips from the video through a preset time window;

[0201] For any user, the lip information of the user is extracted from the multiple frames of the video clip, and whether the user is a speaker is determined based on the changes in the lip information in the multiple frames.

[0202] Optionally, the first determining module is further configured to:

[0203] The user's face image is extracted from each frame image, and the face image is input into a facial key point recognition model to obtain a plurality of key points of the face.

[0204] A lip image is determined according to lip key points among multiple key points of the human face.

[0205] The lip image is input into a lip key point recognition model to obtain multiple key points of the lip.

[0206] Optionally, the first determining module is further configured to:

[0207] The lip opening degree corresponding to each frame of the lip image is determined based on multiple key points of the lip.

[0208] The lip movement state corresponding to each lip image is determined based on the change in the lip opening degree of each lip image relative to the previous lip image.

[0209] Determine whether the user is a speaker based on lip movement states corresponding to the multiple frames of lip images.

[0210] Optionally, the first determining module is further configured to:

[0211] The lip movement states corresponding to the multiple frames of lip images are input into a machine learning model to determine whether the user is a speaker.

[0212] Alternatively, the lip movement states corresponding to multiple frames of lip images are added together to obtain a cumulative change amount, and the cumulative change amount is compared with a preset threshold to determine whether the user is a speaker.

[0213] The lip movement state of the lip image is specifically the absolute value of the difference between the lip opening degree of the lip image and the lip opening degree of the previous lip image.

[0214] Optionally, the first determining module is specifically configured to:

[0215] If it is detected that the lip information of the first user is lost, the body of the first user is tracked, and whether the speaker is the first user is determined by using the stored voiceprint information.

[0216] If so, the speaker in the video is determined based on the human body tracking results.

[0217] Optionally, the first determining module is further configured to:

[0218] In the process of determining the speaker through voiceprint information, if the lip information of the first user is re-detected, whether the speaker is the first user is determined through the voiceprint information and the lip information.

[0219] If the speaker is determined to be the first user through voiceprint information and lip information for a continuous preset time, the speaker is switched to being identified using lip information.

[0220] Optionally, the first determining module is specifically configured to:

[0221] When there are multiple current speakers, if it is detected that the lip information of the second user among the current speakers is lost, the voiceprint information corresponding to the speech collected after the lip information is lost is compared with the voiceprint information stored before the loss;

[0222] If the comparison result is inconsistent, it is determined whether the second user is in a state of continuing to speak during the process of losing lip information based on the lip information of other users in the video except the second user.

[0223] Another embodiment of the present application provides a speaker detection device, comprising:

[0224] The second determination module is used to obtain the real-time video stream of the conference, determine the speaker from the at least one user based on the lip information of the at least one user in the real-time video stream, and store the voiceprint information of the speaker.

[0225] The second determining module is further configured to determine the speaker in the video using the stored voiceprint information if it is detected that the speaker's lip information is lost after the voiceprint information is stored.

[0226] The second display module is used to mark the current speaker in the real-time video stream after the speaker is determined through lip information or voiceprint information.

[0227] The present application also provides a target user detection device, including:

[0228] The third determination module is used to determine at least one user in the video, determine a target user from the at least one user based on image feature information of the at least one user in the video, and store sound feature information of the target user in the video.

[0229] The third determination module is further configured to determine the target user in the video using stored sound feature information if it is detected that the image feature information of the current target user is lost.

[0230] The image feature information includes at least one of the following: lip information, face information, action information, and information about items held by the user.

[0231] The specific implementation principles and technical effects of each device provided in the embodiments of the present application can be found in the aforementioned embodiments and will not be repeated here.

[0232] Figure 14This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 14 As shown, the electronic device of this embodiment may include:

[0233] At least one processor 1401; and a memory 1402 communicatively connected to the at least one processor 1401;

[0234] The memory 1402 stores instructions that can be executed by the at least one processor 1401, and the instructions are executed by the at least one processor 1401 to enable the electronic device to execute the method as described in any of the above embodiments.

[0235] Optionally, the memory 1402 may be independent or integrated with the processor 1401 .

[0236] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.

[0237] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.

[0238] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented.

[0240] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.

[0241] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor. The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, etc.

[0242] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0243] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0244] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0245] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0246] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0247] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A speaker detection method, characterized in that: include: Determining a speaker from the at least one user based on lip information of the at least one user in the video, and storing the voiceprint information of the speaker; wherein the lip information is used to indicate the location of key points of the lip; If it is detected that the speaker's lip information is lost, the speaker in the video is determined by using the stored voiceprint information; wherein the lip information loss is used to indicate that the speaker's lip information cannot be detected; The method further comprises: During playback of the video, capturing a facial image of each user in the video; Displaying facial images of various users at a preset position on the video playback interface and identifying the facial image of the currently determined speaker; The speaker determined from the at least one user is a first user; if it is detected that the lip information of the speaker is lost, the speaker in the video is determined by using the stored voiceprint information, including: If it is detected that the lip information of the first user is missing, the first user is tracked based on the stored body information of the speaker, and the stored voiceprint information is used to determine whether the speaker is the first user; if so, the speaker in the video is determined based on the body tracking result; or, When there are multiple current speakers, if it is detected that the lip information of the second user among the current speakers is lost, the voiceprint information corresponding to the speech collected after the lip information is lost is compared with the voiceprint information stored before the loss; If the comparison result is inconsistent, the lip information of the remaining users except the second user in the video is used to determine whether the second user is in a state of continuing to speak during the process of losing the lip information. If the second user is in a state of continuing to speak, the second user stops speaking.

2. The method according to claim 1, characterized in that The sound in the video is collected by a microphone array; and the speaker is determined from the at least one user based on lip information of the at least one user in the video, including: determining the direction of the sound source through the sound collected by the microphone array; Determining the speaker's position range in the video image based on the sound source direction; A speaker is determined from the at least one user according to lip information of the at least one user within the position range.

3. The method according to claim 1, characterized in that Determining a speaker from at least one user according to lip information of the at least one user in the video includes: Get video clips from the video through a preset time window; For any user, the lip information of the user is extracted from the multiple frames of the video clip, and whether the user is a speaker is determined based on the changes in the lip information in the multiple frames.

4. The method according to claim 3, characterized in that Extracting lip information of the user from multiple frames of images of the video clip includes: Extracting a facial image of the user from each frame of image, inputting the facial image into a facial key point recognition model to obtain a plurality of key points of the face; Determine a lip image according to lip key points among multiple key points of the face; The lip image is input into a lip key point recognition model to obtain multiple key points of the lip.

5. The method according to claim 3, characterized in that Determining whether the user is a speaker based on changes in lip information in multiple frames of images includes: Determine the lip opening degree corresponding to each frame of the lip image according to multiple key points of the lip; determining a lip movement state corresponding to each lip image based on a change in the lip opening degree of the lip image relative to the previous lip image; Determine whether the user is a speaker based on lip movement states corresponding to the multiple frames of lip images.

6. The method according to claim 5, characterized in that Determining whether the user is a speaker based on lip movement states corresponding to the multiple frames of lip images includes: Inputting the lip movement states corresponding to the multiple frames of lip images into a machine learning model to determine whether the user is a speaker; Alternatively, the lip movement states corresponding to the multiple frames of lip images are added together to obtain a cumulative change amount, and the cumulative change amount is compared with a preset threshold to determine whether the user is a speaker; The lip movement state of the lip image is specifically the absolute value of the difference between the lip opening degree of the lip image and the lip opening degree of the previous lip image.

7. A speaker detection method, characterized in that: include: Obtaining a real-time video stream of the conference, determining a speaker from the at least one user based on lip information of the at least one user in the real-time video stream, and storing the voiceprint information of the speaker, wherein the lip information indicates the locations of key points of the lip; After storing the voiceprint information, if it is detected that the speaker's lip information is lost, determining the speaker in the video using the stored voiceprint information; wherein the lip information loss is used to indicate that the speaker's lip information cannot be detected; After the speaker is identified through lip information or voiceprint information, the current speaker is marked in the real-time video stream; The speaker determined from the at least one user is a first user; if it is detected that the lip information of the speaker is lost, determining the speaker in the video by using the stored voiceprint information, includes: If it is detected that the lip information of the first user is missing, the first user is tracked based on the stored body information of the speaker, and the stored voiceprint information is used to determine whether the speaker is the first user; if so, the speaker in the video is determined based on the body tracking result; or, When there are multiple current speakers, if it is detected that the lip information of the second user among the current speakers is lost, the voiceprint information corresponding to the speech collected after the lip information is lost is compared with the voiceprint information stored before the loss; If the comparison result is inconsistent, the lip information of the remaining users except the second user in the video is used to determine whether the second user is in a state of continuing to speak during the process of losing the lip information. If the second user is in a state of continuing to speak, the second user stops speaking.

8. A target user detection method, characterized in that: include: Determining at least one user in the video, determining a target user from the at least one user based on image feature information of the at least one user in the video, and storing voice feature information of the target user in the video; If it is detected that the image feature information of the current target user is lost, the target user in the video is determined by the stored sound feature information; The image feature information includes at least one of the following: lip information, face information, action information, and information about items held by the user; the lip information is used to indicate the location of key points of the lip; The method further comprises: During playback of the video, capturing a facial image of each user in the video; Displaying facial images of various users at a preset position on the video playback interface and identifying the facial image of the currently determined speaker; The speaker determined from the at least one user is a first user; if it is detected that the image feature information of the current target user is lost, determining the target user in the video by using the stored sound feature information, includes: If it is detected that the image feature information of the first user is lost, body tracking of the first user is performed based on the stored body information of the speaker, and whether the speaker is the first user is determined by using the stored voice feature information; if so, the speaker in the video is determined based on the body tracking result; or, When there are multiple current speakers, if it is detected that the image feature information of the second user among the current speakers is lost, the sound feature information corresponding to the speech collected after the lip information is lost is compared with the sound feature information stored before the loss; If the comparison result is inconsistent, it is determined whether the second user is in a state of continuing to speak during the process of losing the image feature information based on the image feature information of the remaining users in the video except the second user. If the second user is in a state of continuing to speak, the second user stops speaking.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method according to any one of claims 1 to 8 is implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Voice identification method and system

    CN104409075A

  • No-supervision multi-speaker identification device based on audio and video and method thereof

    CN109410954A

  • Speech processing method, device and system, and storage medium

    CN110808048A

  • Video speaker identification method and device, computer equipment and storage medium

    CN111785279A