Speaker recognition method, device, storage medium, client, and server

By detecting the relevance of voice data on the client and presenting participant identification, combined with the speaker recognition and update function of the server, the problem of poor speaker recognition effect in the prior art is solved, and a higher recognition accuracy is achieved.

CN115440231BActive Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110617973.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-01
Publication Date
2025-06-27
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

The prior art has poor automatic recognition effect due to far-field or alternating sounds in audio speaker recognition.

Method used

By detecting the recognition operation for the target voice data and the recognition text on the client, it is determined whether the voice data is associated with the voice data source device and presents the corresponding participant identification set. The user can select and re-identify the target participant ID, generate a speaker re-identification request and send it to the server. The server recognizes the target voice data and updates the recognition text.

Benefits of technology

It improves the accuracy of speaker recognition, allows users to re-identify existing recognition text, and enhances the accuracy of recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440231B_ABST
    Figure CN115440231B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speaker recognition and presentation method, apparatus, storage medium, and speaker recognition system. By providing an implementation method in which a user at the client re-performs speaker recognition on the speech data corresponding to the first target participant in the target speech data of the associated voice data source device, and then the server re-performs speaker recognition, the speaker recognition accuracy of the speech data from the terminal device corresponding to the first target participant can be specifically improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of audio - video conferencing, and specifically to a speaker recognition method, apparatus, storage medium, client, and server. Background Art

[0002] Currently, when performing speaker recognition on audio, the effect of automatically recognizing speakers is often poor due to reasons such as far - field or alternating voices. Summary of the Invention

[0003] Embodiments of the present disclosure propose a speaker recognition method, apparatus, storage medium, client, and server.

[0004] In a first aspect, embodiments of the present disclosure provide a speaker recognition method, which includes: presenting a target recognition text formed by performing speech recognition and speaker recognition on target speech data;

[0005] In response to detecting a first recognition operation on the target speech data and the target recognition text, determining whether the target speech data is associated with a voice data source device;

[0006] In response to determining that the target speech data is associated with a voice data source device, presenting a set of target participant identifiers corresponding to the target speech data;

[0007] In response to detecting a selection operation on a first target participant identifier in the set of target participant identifiers, obtaining the first target participant identifier;

[0008] In response to detecting a second recognition operation on the first target participant identifier, generating a first speaker re - recognition request for the target speech data, the target recognition text, and the first target participant identifier, and sending the first speaker re - recognition request to a server.

[0009] In some optional embodiments, in response to detecting a second recognition operation on the first target participant identifier, generating a first speaker re - recognition request for the target speech data, the target recognition text, and the first target participant identifier, and sending the first speaker re - recognition request to a server, includes:

[0010] Determining whether a specified speaker number operation on the first target participant identifier is detected;

[0011] In response to determining that it is detected, presenting a first speaker number editing display object;

[0012] In response to detecting a second recognition operation for editing the current speaker count corresponding to the first target participant identifier and the first speaker count display object, a first speaker re-identification request is generated for the target voice data, the target recognition text, the first target participant identifier, and the first target speaker count, where the first target speaker count corresponds to the current speaker count corresponding to the first speaker count display object.

[0013] In some alternative embodiments, the generating a first speaker re-identification request for the target voice data, the target recognition text, and the first target participant identifier in response to detecting a second recognition operation for the first target participant identifier, and sending the first speaker re-identification request to the server further includes:

[0014] In response to determining that no specified speaker count operation for the first target participant identifier is detected and detecting a second recognition operation for the first target participant identifier, a first speaker re-identification request is generated for the target voice data, the target recognition text, and the first target participant identifier.

[0015] In some alternative embodiments, the method further includes:

[0016] In response to determining that the target voice data is not associated with a voice data source device, a second speaker count display object is presented;

[0017] In response to detecting a third recognition operation for the current speaker count corresponding to the second speaker count display object, a second speaker re-identification request is generated for the target voice data, the target recognition text, and the second target speaker count, and the second speaker re-identification request is sent to the server, where the second target speaker count corresponds to the current speaker count corresponding to the second speaker count display object.

[0018] In some alternative embodiments, after sending the second speaker re-identification request to the server, the method further includes:

[0019] Performing the following second progress value presentation operation in real time: calculating a second real-time duration difference between the current duration and the time when the second speaker re-identification request is sent to the server, and determining a ratio of the second real-time duration difference to a second total duration received from the server as a second recognition progress value, and presenting the second recognition progress value.

[0020] In some alternative embodiments, the method further includes:

[0021] In response to receiving the updated target recognition text and the set of target participant identifiers sent by the server, present the updated target recognition text and the set of target participant identifiers.

[0022] In some alternative embodiments, the method further includes:

[0023] In response to receiving the updated target recognition text sent by the server, present the updated target recognition text.

[0024] In some alternative embodiments, after sending the first speaker re-identification request to the server, the method further includes:

[0025] Perform the following first progress value presentation operation in real time: calculate the first real-time duration difference between the current duration and the time when the first speaker re-identification request is sent to the server, and determine the ratio of the first real-time duration difference to the first total duration received from the server as the first recognition progress value, and present the first recognition progress value.

[0026] In some alternative embodiments, determining whether the target voice data is associated with a voice data source device includes:

[0027] Determine whether the target recognition text is in an edited state;

[0028] In response to determining that the target recognition text is not in an edited state, determine whether the target voice data is associated with a voice data source device.

[0029] In some alternative embodiments, the method further includes:

[0030] In response to determining that the target recognition text is in an edited state, present a first prompt message indicating that the target recognition text is in an edited state and speaker re-identification cannot be performed.

[0031] In a second aspect, an embodiment of the present disclosure provides a speaker recognition method, the method includes: in response to receiving a first speaker re-identification request for target voice data, a target recognition text, and a first target participant identifier in a set of target participant identifiers corresponding to the target voice data sent by the client, perform speaker recognition on the voice data in the target voice data originating from the client indicated by the first target participant identifier, and update the target recognition text according to the speaker recognition result.

[0032] In some alternative embodiments, the first speaker re-identification request further includes a first target speaker number; and

[0033] Performing speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier includes:

[0034] Based on the first target number of speakers, performing speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier.

[0035] In some alternative embodiments, the method further includes:

[0036] In response to receiving a second speaker re - recognition request sent by the client for the target speech data, the target recognition text, and the second target number of speakers, performing speaker recognition on the target speech data based on the second target number of speakers, and updating the target recognition text according to the speaker recognition result.

[0037] In some alternative embodiments, after performing speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier, and updating the target recognition text according to the speaker recognition result, the method further includes:

[0038] Updating the target participant identifier set according to the updated target recognition text;

[0039] Sending the updated target recognition text and the target participant identifier set to the client for the client to present the updated target recognition text and the target participant identifier set.

[0040] In some alternative embodiments, after performing speaker recognition on the target speech data based on the second target number of speakers, and updating the target recognition text according to the speaker recognition result, the method further includes:

[0041] Sending the updated target recognition text to the client for the client to present the updated target recognition text.

[0042] In some alternative embodiments, before performing speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier, the method further includes:

[0043] Determining a first total duration of speaker recognition according to the size of the speech data in the target speech data that originates from the client indicated by the first target participant identifier, and sending the first total duration to the client.

[0044] In some alternative embodiments, before performing speaker recognition on the target voice data based on the number of the second target speakers, the method further includes:

[0045] Determining a second total duration of speaker recognition according to the data size of the target voice data, and sending the second total duration to the client.

[0046] In some alternative embodiments, performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier based on the number of the first target speakers includes:

[0047] Segmenting the voice data in the target voice data that originates from the client indicated by the first target participant identifier into a sentence audio sequence according to the speech recognition result and punctuation annotation result corresponding to the target voice data;

[0048] Segmenting the sentence audio in the sentence audio sequence into an audio segment set after audio segment segmentation;

[0049] Extracting voiceprint features from each audio segment to obtain a voiceprint feature set;

[0050] Determining the number of the first target speakers as the number of speaker clustering centers;

[0051] Clustering the voiceprint features in the voiceprint feature set according to the number of speaker clustering centers to obtain the number of voiceprint feature clustering centers equal to the number of speaker clustering centers;

[0052] For the sentence audio in the sentence audio sequence, determining the voiceprint feature clustering center to which the voiceprint feature corresponding to each audio segment included in the sentence video belongs, and determining the speaker identifier corresponding to the sentence audio as the speaker identifier corresponding to the voiceprint feature clustering center with the most occurrences among the voiceprint feature clustering centers to which the voiceprint features corresponding to each audio segment included in the sentence audio belong.

[0053] In some alternative embodiments, updating the target recognition text according to the speaker recognition result includes:

[0054] In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, updating the target recognition text according to the speaker recognition result.

[0055] In some alternative embodiments, updating the target recognition text according to the speaker recognition result includes:

[0056] In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, update the target recognition text according to the speaker recognition result.

[0057] In a third aspect, an embodiment of the present disclosure provides a speaker recognition device, which includes:

[0058] A first presentation unit, configured to present a target recognition text formed by performing speech recognition and speaker recognition on target speech data;

[0059] A first determination unit, configured to determine whether the target speech data is associated with a voice data source device in response to detecting a first recognition operation on the target speech data and the target recognition text;

[0060] A second presentation unit, configured to present a set of target participant identifiers corresponding to the target speech data in response to determining that the target speech data is associated with a voice data source device;

[0061] An acquisition unit, configured to acquire the first target participant identifier in response to detecting a selection operation on a first target participant identifier in the set of target participant identifiers;

[0062] A first sending unit, configured to generate a first speaker re-recognition request for the target speech data, the target recognition text, and the first target participant identifier, and send the first speaker re-recognition request to a server in response to detecting a second recognition operation on the first target participant identifier.

[0063] In some optional embodiments, the first sending unit is further configured to:

[0064] Determine whether a specified speaker number operation on the first target participant identifier is detected;

[0065] In response to determining that it is detected, present a first speaker number editing display object;

[0066] In response to detecting a second recognition operation on the current speaker number corresponding to the first target participant identifier and the first speaker number editing display object, generate the first speaker re-recognition request for the target speech data, the target recognition text, the first target participant identifier, and the first target speaker number, where the first target speaker number corresponds to the current speaker number corresponding to the first speaker number editing display object.

[0067] In some optional embodiments, the first sending unit is further configured to:

[0068] In response to determining that the specified number of speaker operations for the first target participant identifier is not detected, and detecting a second recognition operation for the first target participant identifier, a first speaker re-identification request is generated for the target voice data, the target recognition text, and the first target participant identifier.

[0069] In some alternative embodiments, the apparatus further comprises:

[0070] A third presentation unit, configured to present a second number-of-speakers editing display object in response to determining that the target voice data is not associated with a voice data source device;

[0071] A second sending unit, configured to, in response to detecting a third recognition operation for the current number of speakers corresponding to the second number-of-speakers editing display object, generate a second speaker re-identification request for the target voice data, the target recognition text, and a second target number of speakers, and send the second speaker re-identification request to the server, where the second target number of speakers corresponds to the current number of speakers corresponding to the second number-of-speakers editing display object.

[0072] In some alternative embodiments, the apparatus further comprises:

[0073] A fourth presentation unit, configured to, after sending the second speaker re-identification request to the server, perform the following second progress value presentation operation in real time: calculate a second real-time duration difference between the current duration and the time when the second speaker re-identification request is sent to the server, and determine a ratio of the second real-time duration difference to a second total duration received from the server as a second recognition progress value, and present the second recognition progress value.

[0074] In some alternative embodiments, the apparatus further comprises:

[0075] A fifth presentation unit, configured to present the updated target recognition text and the target participant identifier set in response to receiving the updated target recognition text and the target participant identifier set sent by the server.

[0076] In some alternative embodiments, the apparatus further comprises:

[0077] A sixth presentation unit, configured to present the updated target recognition text in response to receiving the updated target recognition text sent by the server.

[0078] In some alternative embodiments, the apparatus further comprises:

[0079] A seventh presentation unit, configured to perform the following first progress value presentation operation in real time after sending the first speaker re-identification request to the server: calculate a first real-time duration difference between the current duration and the time when the first speaker re-identification request is sent to the server, and determine a ratio between the first real-time duration difference and a first total duration received from the server as a first identification progress value, and present the first identification progress value.

[0080] In some alternative embodiments, the first determination unit is further configured to:

[0081] Determine whether the target recognition text is in an edited state;

[0082] In response to determining that the target recognition text is not in an edited state, determine whether the target voice data is associated with a voice data source device.

[0083] In some alternative embodiments, the apparatus further includes:

[0084] An eighth presentation unit, configured to present a first prompt message for indicating that the target recognition text is in an edited state and cannot perform speaker re-identification in response to determining that the target recognition text is in an edited state.

[0085] In a fourth aspect, an embodiment of the present disclosure provides a speaker recognition apparatus, the apparatus includes:

[0086] A first recognition and update unit, configured to perform speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier in response to receiving a first speaker re-identification request for the target voice data, the target recognition text, and the first target participant identifier in the target participant identifier set corresponding to the target voice data sent by the client, and update the target recognition text according to the speaker recognition result.

[0087] In some alternative embodiments, the first speaker re-identification request further includes a first target speaker number; and

[0088] The performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier includes:

[0089] Based on the first target speaker number, perform speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier.

[0090] In some alternative embodiments, the apparatus further includes:

[0091] A second recognition and update unit, configured to respond to a second speaker re-recognition request sent by the client for the target voice data, the target recognition text, and the second target speaker number, perform speaker recognition on the target voice data based on the second target speaker number, and update the target recognition text according to the speaker recognition result.

[0092] In some alternative embodiments, the apparatus further includes:

[0093] A first update unit, configured to update the target participant identifier set according to the updated target recognition text after performing speaker recognition on the voice data of the client indicated by the first target participant identifier in the target voice data, and updating the target recognition text according to the speaker recognition result;

[0094] A first sending unit, configured to send the updated target recognition text and the target participant identifier set to the client for the client to present the updated target recognition text and the target participant identifier set.

[0095] In some alternative embodiments, the apparatus further includes:

[0096] A second sending unit, configured to send the updated target recognition text to the client after performing speaker recognition on the target voice data based on the second target speaker number, and updating the target recognition text according to the speaker recognition result, for the client to present the updated target recognition text.

[0097] In some alternative embodiments, the apparatus further includes:

[0098] A third sending unit, configured to determine a first total duration of speaker recognition according to the size of the voice data of the client indicated by the first target participant identifier in the target voice data, and send the first total duration to the client before performing speaker recognition on the voice data of the client indicated by the first target participant identifier in the target voice data.

[0099] In some alternative embodiments, the apparatus further includes:

[0100] A fourth sending unit, configured to determine a second total duration of speaker recognition according to the size of the target voice data, and send the second total duration to the client before performing speaker recognition on the target voice data based on the second target speaker number.

[0101] In some alternative embodiments, the speaker recognition of the voice data originating from the client indicated by the first target participant identifier in the target voice data based on the first target speaker number includes:

[0102] Segment the voice data originating from the client indicated by the first target participant identifier in the target voice data according to the speech recognition result and punctuation annotation result corresponding to the target voice data to obtain a sentence audio sequence;

[0103] Segment the sentence audio in the sentence audio sequence to form an audio segment set;

[0104] Extract voiceprint features for each audio segment to obtain a voiceprint feature set;

[0105] Determine the first target speaker number as the number of speaker clustering centers;

[0106] Cluster the voiceprint features in the voiceprint feature set according to the number of speaker clustering centers to obtain the number of speaker clustering centers of voiceprint feature clustering centers;

[0107] For the sentence audio in the sentence audio sequence, determine the voiceprint feature clustering center to which the voiceprint feature corresponding to each audio segment included in the sentence video belongs, and determine the speaker identifier corresponding to the sentence audio as the speaker identifier corresponding to the voiceprint feature clustering center with the most occurrences among the voiceprint feature clustering centers to which the voiceprint features corresponding to each audio segment included in the sentence audio belong.

[0108] In some alternative embodiments, the updating of the target recognition text according to the speaker recognition result includes:

[0109] In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, update the target recognition text according to the speaker recognition result.

[0110] In some alternative embodiments, the updating of the target recognition text according to the speaker recognition result includes:

[0111] In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, update the target recognition text according to the speaker recognition result.

[0112] In a fifth aspect, an embodiment of the present disclosure provides a client, including: one or more processors; a storage device having one or more programs stored thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.

[0113] In a sixth aspect, an embodiment of the present disclosure provides a server, including: one or more processors; a storage device storing one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the second aspect.

[0114] In a seventh aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method described in any implementation manner of the first aspect and / or the method described in any implementation manner of the second aspect.

[0115] In an eighth aspect, an embodiment of the present disclosure provides a speaker recognition system, including a client described in any implementation manner of the fifth aspect and a server described in any implementation manner of the sixth aspect.

[0116] To solve the problem of poor current speaker recognition effect, the speaker recognition method, device, storage medium, client, and server provided by the present disclosure first present the target recognition text formed by performing speech recognition and speaker recognition on the target speech data. When the client detects a first recognition operation on the target speech data and the target recognition text, it determines whether the target speech data is associated with a voice data source device. If it is determined to be the case, the client presents the set of target participant identifiers corresponding to the target speech data. When a selection operation on the first target participant identifier in the set of target participant identifiers is detected, the first target participant identifier targeted by the selection operation is obtained. When a second recognition operation on the first target participant identifier is detected, a first speaker re-recognition request is generated for the target speech data, the target recognition text, and the first target participant identifier and sent to the server. Finally, the server performs speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier, and updates the target recognition text according to the speaker recognition result. Thus, it is possible to perform re-speaker recognition on the speech data corresponding to the first target participant in the target speech data associated with the voice data source device by the user on the client, providing a way for the user to re-recognize the speaker recognition result in the existing recognition text, and improving the speaker recognition accuracy. Optionally, for example, by specifying the number of speakers, the number of speakers specified can be referred to during the speaker recognition process. With the guidance of the corresponding number of speakers, the speaker recognition accuracy can be further improved. For example, the above target speech data can be the audio data of a remote audio-video conference, and further, it is possible to provide the user on the client to re-perform speaker recognition on the audio data of the remote audio-video conference to specify the speech data from which terminal device, improving the speaker recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0117] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings. The drawings are only for the purpose of showing the specific embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0118] Figure 1 is a system architecture diagram of an embodiment of the speaker recognition system according to the present disclosure;

[0119] Figure 2A and Figure 2C is a timing diagram of an embodiment of the speaker recognition system according to the present disclosure;

[0120] Figure 2B is a decomposed flowchart of an embodiment of step 205 according to the present disclosure;

[0121] Figure 3 is a flowchart of an embodiment of a speaker recognition method applied to a client according to the present disclosure;

[0122] Figure 4 is a flowchart of an embodiment of a speaker recognition method applied to a server according to the present disclosure;

[0123] Figure 5 is a schematic structural diagram of an embodiment of a speaker recognition device applied to a client according to the present disclosure;

[0124] Figure 6 is a schematic structural diagram of an embodiment of a speaker recognition device applied to a server according to the present disclosure;

[0125] Figure 7 is a schematic structural diagram of a computer system of a client or a server suitable for implementing an embodiment of the present disclosure. Detailed implementation manners

[0126] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings.

[0127] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0128] Figure 1 An exemplary system architecture 100 to which the speaker recognition method and device and an embodiment of the speaker recognition method and device according to the present disclosure can be applied is shown.

[0129] As Figure 1 shown, the system architecture 100 may include clients 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the clients 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0130] Users can use the clients 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the clients 101, 102, 103, such as audio and video conferencing applications, instant messaging tools, speech recognition applications, web browser applications, shopping applications, search applications, email clients, social platform software, etc.

[0131] The clients 101, 102, and 103 can be either hardware or software. When the clients 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and supporting voice collection and / or video collection, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on. When the clients 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (e.g., for providing distributed services), or as a single software or software module. Specific limitations are not made here.

[0132] The server 105 can be a server providing various services, such as a background server that supports audio and video conferencing applications or voice recognition applications displayed on the clients 101, 102, and 103. The background server can analyze and process data such as speaker recognition requests received, and feedback the processing results (e.g., speaker recognition results, etc.) to the clients.

[0133] It should be noted that the server 105 can be either hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., for providing distributed services), or as a single software or software module. Specific limitations are not made here.

[0134] It should be noted that the speaker recognition method applied to the client generally is executed by the clients 101, 102, and 103. Correspondingly, the speaker recognition device applied to the client generally is set in the clients 101, 102, and 103.

[0135] It should be noted that the speaker recognition method applied to the server generally is executed by the server 105. Correspondingly, the speaker recognition device applied to the server generally is set in the server 105.

[0136] It should be understood that Figure 1 the numbers of clients, networks, and servers in

[0137] Continue to refer to Figure 2A, which shows a time sequence 200 of an embodiment of a speaker recognition system according to the present disclosure. The speaker recognition system in the embodiments of the present disclosure may include a client and a server. The time sequence 200 includes the following steps:

[0138] Step 201, the client presents a target recognition text formed by performing speech recognition and speaker recognition on target speech data.

[0139] In this embodiment, the client (such as Figure 1 the clients 101, 102, 103 shown) may adopt various implementation manners to present the target recognition text formed by performing speech recognition and speaker recognition on the target speech data.

[0140] Here, the target speech data may be various speech data, and the target recognition text may be a text formed by performing speech recognition and speaker recognition on the target speech data. The target recognition text may include a target speech data recognition text obtained by performing speech recognition on the target speech data, and the target speech data recognition text may include characters and symbols. The target recognition text may also include speaker identifiers corresponding to each recognition statement in the target speech data recognition text and obtained by performing speaker recognition based on the target speech data and the target speech data recognition text.

[0141] In some alternative embodiments, the target speech data may be speech data generated by the server in the order of reception time after each terminal device participating in the remote audio and video conference (such as a dedicated remote conference device or a terminal device installed with a remote audio and video conference application) in a physical environment real-time collects the input speech of the terminal device and sends it to the server that supports the remote audio and video conference during the progress of the remote audio and video conference. At this time, the target speech data may include a speech data sequence generated in time sequence and a data source identifier corresponding to each speech data in the speech data sequence, that is, each speech data may be associated with a corresponding speech data source device. Specifically, the speech data can be associated with the speech data source device by the device identifier of the speech data source device or the user identifier of the participating user (i.e., the participant identifier).

[0142] It should be noted that here, the speech data source device associated with the target speech data may be one or more than one. For example, when multiple users respectively use dedicated remote conference devices to conduct a remote audio and video conference in a physical environment, the speech data collected by each dedicated remote conference device used by these users can be sent to the server through the same unified terminal device, and then the target speech data received and formed by the server is only associated with one speech source.

[0143] In some alternative embodiments, the target voice data may also be voice data that is collected in real time by a terminal device participating in a remote audio-video conference (e.g., a dedicated remote conference device or a terminal device installed with a remote audio-video conference application) during the remote audio-video conference, and the input voice (e.g., the voice data of the participant speaking who uses the terminal device to participate in the conference) and output voice (e.g., the voice data of the participant speaking who uses a terminal device other than the terminal device to participate in the conference) of the terminal device are sent to the server that supports the remote audio-video conference. At this time, the target voice data may include voice data generated in sequence, but there is no data source corresponding to different voice segments in the voice data. That is, the target voice data is only voice data and is not associated with a voice data source device.

[0144] In some alternative embodiments, the target voice data may be voice data collected in real time by a terminal device using a sound input device, and the voice data is generated and uploaded to the server after the collection is completed. Correspondingly, the server may perform voice recognition and speaker recognition on the above target voice data to obtain the target recognition text. That is, here, the target voice data is also only voice data and is not associated with a voice data source device.

[0145] Here, the client may, for example, present corresponding display objects for indicating play, pause play, accelerate play, decelerate play, and play progress dragging on a display device, and when detecting operations such as clicking, swiping, and dragging on the above different display objects by the user, use a sound playback device to implement corresponding operations on the target voice data to present the target voice data.

[0146] The client may also present the target recognition text on the display device. For example, different recognition statements and corresponding speaker identifiers may be presented in the order of time corresponding to the target voice data in the target recognition text.

[0147] Step 202, the client determines whether the target voice data is associated with a voice data source device in response to detecting a first recognition operation on the target voice data and the target recognition text.

[0148] In this embodiment, the client may use various implementation methods to detect the first recognition operation on the target voice data and the target recognition text. For example, the client may present a first re-recognize speaker operation display object (e.g., an icon, text, or button with "Re-recognize speaker" or "Rematch speakers" written on it) for triggering the first recognition operation on the page presenting the target voice data and the target recognition text, and determine that the first recognition operation is detected after detecting operations such as single-clicking, double-clicking, hovering, dragging, and swiping on the above first re-recognize speaker operation display object.

[0149] In this embodiment, the client can determine locally or remotely by the server whether the target voice data is associated with a voice data source device. It can be understood that when the client locally stores whether the target voice data is associated with a voice data source device, it can be determined locally by the client; otherwise, the client can generate a determination request for the above determination content and send it to the server for determination, and the server returns the corresponding determination result.

[0150] Here, if the target voice data is associated with a voice data source device, it can indicate that the target voice data is the voice data generated by the server in real time during a remote audio and video conference after the input voice of each terminal device participating in the remote audio and video conference (for example, a dedicated remote conference device or a terminal device installed with a remote audio and video conference application) is collected in real time and sent to the server that supports the remote audio and video conference, and is generated by the server in the order of reception time. Moreover, the terminal devices participating in the remote audio and video conference can respectively correspond to each target participant identifier in the target participant identifier set. Here, the target participant identifier set can be the participant identifier set of the conference corresponding to the above target voice data.

[0151] In this embodiment, when the client determines that the target voice data is associated with a voice data source device, it can proceed to step 203.

[0152] Step 203: The client presents the target participant identifier set corresponding to the target voice data.

[0153] As recorded above, if the client determines in step 202 that the target voice data is associated with a voice data source device, it indicates that the target voice data is the audio of each participating terminal device recorded in real time by the server during the remote audio and video conference. Subsequently, the remote audio and video conference can be associated with a corresponding participant identifier set, which can be considered as the target participant identifier set corresponding to the target voice data. And the client can obtain and present the target participant identifier set before step 203. Optionally, if the client has not obtained the target participant identifier set before step 203, it can generate a participant identifier acquisition request for acquiring the participant identifier set corresponding to the target voice data and send it to the server. The server can then acquire the target participant identifier set corresponding to the target voice data after receiving the above participant identifier acquisition request and feedback it to the client, and then the client can present the above target participant identifier set.

[0154] Here, the client can present the target participant identifier set in various ways.

[0155] For example, the client can present the target participant identifier set in the form of a dropdown list or in the form of icons sorted according to specified rules.

[0156] Step 204: In response to detecting a selection operation on a first target participant identifier in the target participant identifier set, the client obtains the first target participant identifier.

[0157] Here, the selection operation for the first target participant identifier in the target participant identifier set may be, for example, single-click, double-click, hover, drag, slide, or the like.

[0158] Step 205 , in response to detecting the second recognition operation for the first target participant identifier, the client generates a first speaker re-recognition request for the target voice data, the target recognition text and the first target participant identifier, and sends the first speaker re-recognition request to the server.

[0159] In this embodiment, the client may detect the second recognition operation for the first target participant identifier in various implementation methods.

[0160] For example, the client may simultaneously present a second speaker re-identification operation display object for triggering the second recognition operation (e.g., an icon, text, or button with "Are you sure to rematch?" or "Are you sure to rematch?") on the page presenting the target participant identification set, and determine that the second recognition operation has been detected after detecting a single-click, double-click, hover, drag, slide, or other operation on the second speaker re-identification operation display object.

[0161] In this embodiment, the client may also generate a first speaker re-identification request for the target voice data, the target recognition text, and the first target participant identifier in various ways. For example, as can be seen from the records of step 203 and step 204, the target voice data and the target recognition text may correspond to a remote audio and video conference, and the server may store the correspondence between the target voice data, the target recognition text, and the target conference identifier of the corresponding remote audio and video conference. When presenting the target voice data and the target recognition text in step 201, the client may store the above-mentioned target conference identifier, so here the client may generate a first speaker re-identification request based on the above-mentioned target conference identifier and the first target participant identifier.

[0162] Step 206, in response to receiving a first speaker re-identification request sent by the client for the target voice data, the target recognition text, and the first target participant identifier in the target participant identifier set corresponding to the target voice data, the server performs speaker recognition on the voice data in the target voice data originating from the client indicated by the first target participant identifier, and updates the target recognition text according to the speaker recognition result.

[0163] In this embodiment, since the first speaker re-identification request specifies the first target participant identifier in the set of target participant identifiers, that is, the user hopes to re-identify the voice data in the target voice data that originates from the client indicated by the first target participant identifier. Then, the server can perform speaker identification on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and update the target recognition text according to the speaker identification result.

[0164] Here, the server can adopt various speaker identification methods that are currently known or will be developed in the future. The present disclosure does not make specific limitations on this.

[0165] It should be noted that the speaker identification result of performing speaker identification on the voice data in the target voice data that originates from the client indicated by the first target participant identifier can be a sequence of recognition statement information composed of at least one recognition statement information in chronological order. Among them, the recognition statement information can include the recognition statement and the corresponding speaker identifier. Among them, the recognition statement in the recognition statement information can be the statement in the target voice data recognition text in the target recognition text. And updating the target recognition text according to the speaker identification result can be to update the speaker identifier corresponding to the corresponding recognition statement in the target recognition text with the speaker identifier in the obtained recognition statement information.

[0166] In some cases, this embodiment may have the following optional implementation manners:

[0167] Optional implementation manner (1): In step 206, before the server performs speaker identification on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, the server can also determine the first total duration of the speaker identification according to the size of the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and send the first total duration to the client.

[0168] For example, the server can determine the first total duration of the speaker identification according to the size of the voice data in the target voice data that originates from the client indicated by the first target participant identifier according to the complexity of the speaker identification algorithm with respect to the size of the voice data in the target voice data that originates from the client indicated by the first target participant identifier.

[0169] Based on the above optional manner of the server sending the first total duration to the client, the above time sequence 200 may further include performing the following step 205' after step 205:

[0170] Step 205', the client performs the first progress value presentation operation in real time.

[0171] Here, the client can perform the first progress value presentation operation in real time. For example, the first progress value presentation operation can be performed every preset duration (e.g., 0.5 seconds).

[0172] Specifically, the first progress value presentation operation may include:

[0173] First, calculate the first real-time duration difference between the current duration and the time when the first speaker re-identification request is sent to the server. Then, determine the ratio between the first real-time duration difference and the first total duration received from the server as the first recognition progress value. Finally, present the first recognition progress value.

[0174] That is, through the above optional implementation (1), it is possible to present in real time on the client the recognition progress of speaker recognition for the voice data in the target voice data that originates from the client indicated by the first target participant identifier, which is convenient for the user to obtain the specific progress situation.

[0175] Optional implementation (2): Step 205 in the time sequence 200 may include steps 2051 to 2053 as Figure 2B shown:

[0176] Step 2051, determine whether a specified speaker number operation for the first target participant identifier is detected.

[0177] Here, the client can use various implementation methods to detect the specified speaker number operation for the first target participant identifier. For example, the client can simultaneously present on the page showing the target participant identifier set a speaker number specified selection display object (e.g., a radio box, a checkbox, or two icons labeled "Specify Speaker Number" and "Do Not Specify Speaker Number") for triggering the specified speaker number operation. The user can make two selections in the above speaker number specified selection display object, one of which is associated with the specified speaker number operation and the other is not. When the user selects the selection associated with the specified speaker number operation, the specified speaker number operation can be triggered. If it is determined in step 2051 that a specified speaker number operation for the first target participant identifier is detected, proceed to step 2052 for execution.

[0178] Step 2052, present the first speaker number editing display object.

[0179] Here, if the user selects to specify the number of speakers in step 2051, that is, the user wishes to specify how many speakers are speaking at the terminal device corresponding to the first target participant identifier during the remote audio and video conference, the client can present a first speaker number editing display object. Here, the first speaker number editing display object can be various real objects that are convenient for digital editing, such as a text box, a digital increase or decrease button, a digital drop-down list, a digital icon, etc.

[0180] Step 2053, in response to detecting a second recognition operation for the current number of speakers corresponding to the first target participant identifier and the first speaker number editing display object, a first speaker re-recognition request is generated for the target voice data, the target recognition text, the first target participant identifier and the first target speaker number.

[0181] Here, the first target speaker quantity corresponds to the current speaker quantity corresponding to the first speaker quantity editing display object.

[0182] The client may adopt various implementation methods to detect the second recognition operation on the current number of speakers corresponding to the first target participant identifier and the first speaker number editing display object.

[0183] For example, the client may simultaneously present a third speaker re-identification operation display object for triggering the second recognition operation on a page presenting the first speaker number edit display object (for example, an icon, text, or button with the words "Are you sure to rematch speakers according to the number of selected number?" or "Are you sure to rematch speakers according to the number of selected number?"), and determine that a second recognition operation for the first target participant identifier and the current number of speakers corresponding to the first speaker number edit display object has been detected after detecting a single-click, double-click, hover, drag, slide, or other operation on the third speaker re-identification operation display object.

[0184] Here, the client can also generate a first speaker re-identification request for the target voice data, the target recognition text, the first target participant identifier, and the first target number of speakers in various ways. For example, as described in steps 203 and 204, both the target voice data and the target recognition text can correspond to a remote audio-video conference, and the server can store the correspondence between the target voice data, the target recognition text, and the target conference identifier of the corresponding remote audio-video conference. When presenting the target voice data and the target recognition text in step 201, the client can store the above target conference identifier. Therefore, here the client can generate a first speaker re-identification request based on the above target conference identifier, the first target participant identifier, and the first target number of speakers.

[0185] Based on the above steps 2051 to 2053, when the server executes step 206, the first speaker re-identification request received may include the first target number of speakers, and step 206 can be executed as follows:

[0186] In response to receiving the first speaker re-identification request sent by the client for the target voice data, the target recognition text, the first target participant identifier, and the first target number of speakers, the server performs speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data based on the first target number of speakers. That is, speaker recognition is performed according to the first target number of speakers, and the resulting speaker recognition result can include the first target number of speaker identifiers. Performing speaker recognition according to the number of speakers specified by the user corrects the previous speaker recognition result in the target recognition text compared to performing speaker recognition according to the number of speakers specified by the user, which can improve the recognition accuracy of re-identifying speakers compared to the previous speaker recognition.

[0187] Optionally, here, performing speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data based on the first target number of speakers may include the following first to sixth steps:

[0188] First, according to the speech recognition result and punctuation annotation result corresponding to the target voice data, the voice data from the client indicated by the first target participant identifier in the target voice data is segmented to obtain a sentence audio sequence.

[0189] Second, the sentence audio in the sentence audio sequence is segmented into an audio segment set after audio segment segmentation.

[0190] For example, the sentence audio can be continuously and evenly segmented starting from the start time of the sentence audio in chronological order. The time length of each audio segment obtained by segmentation is the first preset time length (for example, 1.5 seconds) until the time length of the last audio segment is less than or equal to the first preset time length. Another example is that the audio segments can also be intercepted from the starting point of the sentence audio according to a preset sliding window. Here, the preset sliding window can include a window length and a sliding step length, and the window length of the preset sliding window can be greater than the sliding step length. For example, the window length can be 1.5 seconds, and the sliding step length can be 0.75 seconds. Among them, the window length is the time length of each intercepted audio segment, and the sliding step length is the time difference between the start times of two adjacent interception operations. Since the window length of the preset sliding window is greater than the sliding step length, there is an overlapping part between two adjacent audio segments in the intercepted audio segments. Furthermore, all the audio data in the sentence audio is reflected in the audio segments, and the information will not be lost.

[0191] In the third step, voiceprint features are extracted from each audio segment to obtain a voiceprint feature set.

[0192] For example, the voiceprint features here can be short-time spectrum features such as MFCC (Mel Frequency Cepstral Coefficient), PLP (Perceptual Linear Prediction), FBank (FilterBanks), i-vector (identity-vector), features extracted based on TDNN (Time Delay Neural Networks) such as x-vector, etc.

[0193] In the fourth step, the number of first target speakers is determined as the number of speaker clustering centers.

[0194] In the fifth step, according to the number of speaker clustering centers, the voiceprint features in the voiceprint feature set are clustered to obtain the number of voiceprint feature clustering centers equal to the number of speaker clustering centers.

[0195] Here, various clustering algorithms can be used, such as distance-based clustering algorithms or density-based clustering algorithms, etc. The present disclosure does not make specific limitations in this regard. For example, Spectral Clustering algorithm, K-Means (K-means) clustering, mean shift clustering, Expectation-Maximization (EM) clustering using Gaussian Mixed Model (GMM), agglomerative hierarchical clustering, Graph Community Detection, etc. can be used.

[0196] Step 6: For the sentence audio in the sentence audio sequence, determine the speaker feature cluster centers to which the voiceprint features corresponding to the audio segments included in the sentence video belong, and determine the speaker identifier corresponding to the sentence audio as the speaker identifier corresponding to the voiceprint feature cluster center with the most occurrences among the voiceprint feature cluster centers to which the voiceprint features corresponding to the audio segments included in the sentence audio belong.

[0197] Alternative Embodiment (3): Based on the above alternative embodiment, in step 205, when it is determined in step 2051 that the specified number of speakers for the first target participant identifier is not detected, that is, the user does not wish to specify the number of speakers corresponding to the first target participant identifier, the following step 2054 can be executed.

[0198] Step 2054: In response to detecting a second recognition operation for the first target participant identifier, generate a first speaker re-identification request for the target voice data, the target recognition text, and the first target participant identifier.

[0199] That is, if the user does not wish to specify the number of speakers, a first speaker re-identification request can be generated for the target voice data, the target recognition text, and the first target participant identifier. Correspondingly, when the server executes step 206, since the first target number of speakers is not included in the first speaker re-identification request, speaker recognition can be performed on the voice data in the target voice data that originates from the client indicated by the first target participant identifier. That is, it may be the same or different from the number before re-identifying the speaker this time.

[0200] Due to page display limitations, continue to refer to Figure 2C , it should be noted that Figure 2C The process of Figure 2C In addition to including the process shown in Figure 2A The various steps shown in

[0201] Alternative Embodiment (4): In the above time sequence 200, when the client determines in step 202 that the target voice data is not associated with the voice data source device, the following step 207 can be executed:

[0202] Step 207: The client presents a second speaker number editing display object in response to determining that the target voice data is not associated with the voice data source device.

[0203] That is, if the target voice data is not the voice data generated by each terminal device participating in the remote audio-visual conference (e.g., a dedicated remote conference device or a terminal device installed with a remote audio-visual conference application) in real time during the remote audio-visual conference, collecting the input voice of the terminal device and sending it to the server that supports the remote audio-visual conference, and the server generates the voice data in the order of receiving time. The target voice data can be the voice data of the input voice (e.g., the voice data of the participant speaking who uses the terminal device to participate in the meeting) and the output voice (e.g., the voice data of the participant speaking who uses other terminal devices other than the terminal device) collected in real time by a certain terminal device participating in the remote audio-visual conference (e.g., a dedicated remote conference device or a terminal device installed with a remote audio-visual conference application) during the remote audio-visual conference and sent to the server that supports the remote audio-visual conference. Or, the target voice data can also be the voice data collected in real time by the terminal device using the sound input device and generated and uploaded to the server after the collection ends.

[0204] To implement re-speaker recognition for the target voice data to improve the accuracy of speaker recognition, here the client can present a second speaker number editing display object for the user to specify how many speakers should participate in the re-speaker recognition of the target voice data. Here, the second speaker number editing display object can be various real objects convenient for digital editing, such as a text box, a digital increase and decrease button, a digital drop-down list, a digital icon, etc.

[0205] Step 208, in response to detecting a third recognition operation for the current speaker number corresponding to the second speaker number editing display object, the client generates a second speaker re-recognition request for the target voice data, the target recognition text, and the second target speaker number, and sends the second speaker re-recognition request to the server.

[0206] Here, the second target speaker number corresponds to the current speaker number corresponding to the second speaker number editing display object.

[0207] The client can detect the third recognition operation for the current speaker number corresponding to the second speaker number editing display object in various implementation manners.

[0208] For example, the client may simultaneously present a fourth speaker re-identification operation display object for triggering a third recognition operation (for example, an icon, text, or button with "Are you sure to rematch speakers for the speech data?" or "Are you sure to rematch speakers for the speech data?") on a page presenting a second speaker number edit display object, and determine that a third recognition operation for the current number of speakers corresponding to the second speaker number edit display object has been detected after detecting a single-click, double-click, hover, drag, or slide operation on the fourth speaker re-identification operation display object.

[0209] Here, the client may also generate a second speaker re-identification request for the target voice data, the target recognition text, and the second target number of speakers in various ways. For example, here, there is a corresponding relationship between the target voice data and the target recognition text, and the two may correspond to a voice data identifier, and the server may store the corresponding relationship between the target voice data, the target recognition text, and the corresponding voice data identifier. When presenting the target voice data and the target recognition text in step 201, the client may also store the above-mentioned voice data identifier, so here the client may generate a second speaker re-identification request based on the above-mentioned voice data identifier and the second target number of speakers.

[0210] Step 209, in response to receiving a second speaker re-identification request for the target voice data, the target recognition text and the second target number of speakers sent by the client, the server performs speaker recognition on the target voice data based on the second target number of speakers, and updates the target recognition text according to the speaker recognition result.

[0211] Here, the server can perform speaker recognition on the target voice data according to the second target speaker number, and the obtained speaker recognition result includes speaker identifiers of the second target speaker number, that is, the target voice data is the voice data corresponding to the second target speaker number speakers. Speaker recognition is performed on the target voice data according to the second target speaker number specified by the user, and the speaker recognition result of the previous time in the target recognition text is corrected, which can improve the recognition accuracy of the re-recognized speaker relative to the previous speaker recognition. Updating the target recognition text according to the above speaker recognition result can make the speaker recognition result in the updated target recognition text more consistent with the actual situation.

[0212] Optionally, here, speaker recognition is performed on the target voice data based on the number of second target speakers, and a similar method to that in step 206 where the server performs speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data can be adopted. For example, when adopting the optional implementation manner including the first to sixth steps shown in step 206 for speaker recognition, the difference between the two is that the recognized data is changed from the voice data from the client indicated by the first target participant identifier in the target voice data to the target voice data itself, and the number of first target speakers is changed to the number of second target speakers.

[0213] Through the above optional implementation manner (four), in the case where the target audio data is not the audio data recorded by the server during the remote audio and video conference, such as when the target audio data is uploaded after the meeting or is a real-time recording, it is also possible to provide speaker recognition after the user specifies the number of speakers, thereby improving the accuracy of speaker recognition.

[0214] Optional implementation manner (five): Before the server performs speaker recognition on the target voice data based on the number of second target speakers in step 209, the second total duration of speaker recognition can also be determined according to the data size of the target voice data, and the second total duration is sent to the client.

[0215] For example, the second total duration of speaker recognition can also be determined according to the data size of the target voice data according to the complexity of the speaker recognition algorithm with respect to the data size of the target voice data.

[0216] Based on the above optional manner in which the server sends the second total duration to the client, the above timing 200 may further include performing the following step 208' after step 209:

[0217] Step 208', the client performs the second progress value presentation operation in real time.

[0218] Here, the client can perform the second progress value presentation operation in real time. For example, it can perform the second progress value presentation operation every preset duration (for example, 0.5 seconds).

[0219] Specifically, the second progress value presentation operation may include:

[0220] First, calculate the second real-time duration difference between the current duration and the time when the second speaker re-recognition request is sent to the server. Then, determine the ratio between the second real-time duration difference and the second total duration received from the server as the second recognition progress value. Finally, present the second recognition progress value.

[0221] That is, through the above optional implementation (5), it is possible to present the recognition progress of speaker recognition of the target voice data in real time on the client side, which is convenient for the user to obtain the specific progress.

[0222] Optional implementation (6): In the above time sequence 200, after step 206, that is, after the server performs speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data and updates the target recognition text according to the speaker recognition result, the following steps 210 to 212 may further be included:

[0223] Step 210: The server updates the target participant identifier set according to the updated target recognition text.

[0224] As can be seen from the above description, since the target voice data corresponds to a participant identifier, when the speaker recognition is re-performed on the voice data from the client indicated by the first target participant identifier in the target voice data, the speaker identifier corresponding to the first target participant identifier may change. For example, in traditional remote audio and video conferences, the voice data from each participating terminal device is mostly distinguished according to the device dimension to form a recognition document. Therefore, each participating terminal device in the target recognition text may correspond to a speaker identifier. Or, even if the distinction is made according to the voice dimension, that is, the voice data corresponding to each participating terminal may correspond to more than one speaker identifier, but if the recognition accuracy of the previous speaker recognition is not ideal, after the server re-recognizes in step 206, the recognition accuracy of the speaker recognition may be improved. Therefore, the participant identifier corresponding to the first target participant in the target recognition text may change. Therefore, the server may update the target participant identifier set according to the updated target recognition text, that is, update the participant identifier corresponding to the first target participant identifier according to the speaker recognition result obtained in step 206.

[0225] Step 211: The server sends the updated target recognition text and the target participant identifier set to the client.

[0226] Since the target recognition text and the target participant identifier set (that is, the participant identifier set corresponding to the target voice data) have been updated, the updated target recognition text and the target participant identifier set may be sent to the client after the update.

[0227] Step 212: The client presents the updated target recognition text and the target participant identifier set in response to receiving the updated target recognition text and the target participant identifier set sent by the server.

[0228] Here, after receiving the updated target recognition text and the set of target participant identifiers sent by the server, the client can present the updated target recognition text and the set of target participant identifiers in various implementation manners. For example, the client can present the target recognition text in the same or different presentation manner and presentation position as that in step 201 for presenting the target recognition text, and can also present the updated set of target participant identifiers in the same or different presentation manner and presentation position as that in step 203.

[0229] Through the above optional implementation manner (six), after the client realizes speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and updates the target recognition text according to the speaker recognition result, the updated target recognition text and the set of target participant identifiers can be presented on the client. That is, the user can use the client to view the impact of the recognition result after re - performing speaker recognition on the target recognition text and the first target participant identifier.

[0230] Optional implementation manner (seven), in the above time sequence 200, after step 209, that is, after the server responds to receiving the second speaker re - recognition request sent by the client for the target voice data, the target recognition text, and the second target speaker quantity, performs speaker recognition on the target voice data based on the second target speaker quantity, and updates the target recognition text according to the speaker recognition result, the following steps 213 and 214 can also be included:

[0231] Step 213, the server sends the updated target recognition text to the client.

[0232] That is, since speaker recognition is re - performed on the target voice data, and as described in steps 207 and 208, the target voice data is voice data originating from the same terminal device. The target voice data may or may not correspond to a set of participant identifiers (when the target voice data is the voice data uploaded by the terminal device after the end of a remote audio - video conference) or not correspond to a set of participant identifiers (when the target voice data is the voice data formed by real - time recording of a certain terminal device). However, regardless of which type of voice data, the voice segments in the target voice data do not correspond to different participant identifiers in the corresponding set of participant identifiers. Therefore, after re - performing speaker recognition according to the second target speaker quantity, the speaker identifiers in the obtained recognition result cannot be corresponded to the participant identifiers, and thus the set of participant identifiers cannot be updated, but only the target recognition text can be updated, and then the updated target recognition text is sent to the client.

[0233] Step 214, the client presents the updated target recognition text in response to receiving the updated target recognition text sent by the server.

[0234] For the client to present the updated target recognition text, reference can be made to the relevant records in step 212, which will not be elaborated here.

[0235] Through the above optional implementation (VII), on the client side, after performing speaker recognition on the target voice data based on the number of second target speakers and updating the target recognition text according to the speaker recognition result, the updated target recognition text can be presented on the client side. That is, the user can use the client to see the impact of the recognition result after re - performing speaker recognition on the target recognition text.

[0236] Optional implementation (VIII), for updating the target recognition text according to the speaker recognition result in step 206 and step 209, can be executed as follows:

[0237] In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to the preset number threshold, update the target recognition text according to the speaker recognition result.

[0238] In practice, it is designed considering the accuracy of the current speaker recognition algorithm. That is, it can be considered that if the number of speaker identifiers included in the result of re - performing speaker recognition is large (greater than the preset number threshold), there may be a problem with the speaker recognition algorithm, and the target recognition text is not updated. If the number of speaker identifiers included in the speaker recognition result is small (less than or equal to the preset number threshold), the speaker recognition result is used to update the target recognition text. Optionally, the preset number threshold here can be 5. That is, if the speaker recognition algorithm recognizes more than 5 people, it is considered that the algorithm is inaccurate, and if the speaker recognition algorithm recognizes less than or equal to 5 people, it is considered that the algorithm is accurate, the speaker recognition result is used to update the target recognition text, and feedback can be sent to the client for presentation to the user.

[0239] The speaker recognition system provided by the above embodiments of the present disclosure can specifically improve the speaker recognition accuracy of the voice data from the terminal device corresponding to the first target participant by providing an implementation method for the user on the client side to re - perform speaker recognition on the voice data corresponding to the first target participant in the target voice data of the associated voice data source device, and then re - performing speaker recognition by the server.

[0240] Continue to refer to Figure 3 , which shows a flow 300 of an embodiment of the speaker recognition method according to the present disclosure. This speaker recognition method is applied to the client and includes the following steps:

[0241] Step 301, present the target recognition text formed by performing speech recognition and speaker recognition on the target voice data.

[0242] Step 302, in response to detecting a first recognition operation on the target voice data and the target recognition text, determine whether the target voice data is associated with a voice data source device.

[0243] Step 303, present the set of target participant identifiers corresponding to the target voice data.

[0244] Step 304, in response to detecting a selection operation on the first target participant identifier in the set of target participant identifiers, obtain the first target participant identifier.

[0245] Step 305, in response to detecting a second recognition operation on the first target participant identifier, generate a first speaker re-identification request for the target voice data, the target recognition text, and the first target participant identifier, and send the first speaker re-identification request to the server.

[0246] In this embodiment, the specific operations of steps 301, 302, 303, 304, and 305 and the technical effects produced by them are Figure 2A substantially the same as the operations and effects of steps 201, 202, 203, 204, and 205 in the embodiment shown, and will not be elaborated here.

[0247] In some alternative embodiments, the above method flow 300 may further include the following steps:

[0248] Step 306, in response to determining that the target voice data is not associated with a voice data source device, present a second speaker number editing display object.

[0249] Here, the specific operation of step 306 and the technical effect produced by it are Figure 2A substantially the same as the operation and effect of step 207 in the embodiment shown, and will not be elaborated here.

[0250] In some alternative embodiments, the above method flow 300 may further include the following steps:

[0251] Step 307, in response to detecting a third recognition operation on the current speaker number corresponding to the second speaker number editing display object, generate a second speaker re-identification request for the target voice data, the target recognition text, and the second target speaker number, and send the second speaker re-identification request to the server.

[0252] Here, the specific operation of step 307 and the technical effect produced by it are Figure 2A substantially the same as the operation and effect of step 208 in the embodiment shown, and will not be elaborated here.

[0253] In some alternative embodiments, the method flow 300 may further include performing the following step 305' after step 305:

[0254] Step 305', perform a first progress value presentation operation in real time.

[0255] Here, the specific operation of step 305' and the technical effects produced thereby are substantially the same as Figure 2A the operations and effects of step 205' in the illustrated embodiment, and will not be elaborated herein.

[0256] In some alternative embodiments, the method flow 300 may further include performing the following step 307' after step 307:

[0257] Step 307', perform a second progress value presentation operation in real time.

[0258] Here, the specific operation of step 307' and the technical effects produced thereby are substantially the same as Figure 2A the operations and effects of step 208' in the illustrated embodiment, and will not be elaborated herein.

[0259] In some alternative embodiments, the method flow 300 may further include the following step 308:

[0260] Step 308, in response to receiving the updated target recognition text and the target attendee identifier set sent by the server, present the updated target recognition text and the target attendee identifier set.

[0261] Here, the specific operation of step 308 and the technical effects produced thereby are substantially the same as Figure 2A the operations and effects of step 212 in the illustrated embodiment, and will not be elaborated herein.

[0262] In some alternative embodiments, the method flow 300 may further include the following step 309:

[0263] Step 309, in response to receiving the updated target recognition text sent by the server, present the updated target recognition text.

[0264] Here, the specific operation of step 309 and the technical effects produced thereby are substantially the same as Figure 2A the operations and effects of step 214 in the illustrated embodiment, and will not be elaborated herein.

[0265] The speaker recognition method provided by the above embodiments of the present disclosure can facilitate the user to specifically improve the speaker recognition accuracy of the voice data from the terminal device corresponding to the first target participant by providing an implementation manner in which the user re-performs speaker recognition on the voice data corresponding to the first target participant in the target voice data for the associated voice data source device on the client side.

[0266] Continuing to refer to Figure 4 , which shows a process 400 of a speaker recognition method according to an embodiment of the present disclosure. The speaker recognition method, applied to a server, includes the following steps:

[0267] Step 401, in response to receiving a first speaker re-recognition request for the target voice data, the target recognition text, and the first target participant identifier in the set of target participant identifiers corresponding to the target voice data sent by the client, perform speaker recognition on the voice data in the target voice data originating from the client indicated by the first target participant identifier, and update the target recognition text according to the speaker recognition result.

[0268] In this embodiment, the specific operations of step 401 and the technical effects thereof are substantially the same as the operations and effects of step 206 in the Figure 2A shown embodiment, and will not be elaborated here.

[0269] In some alternative embodiments, the above method process 400 may further include the following step 402:

[0270] Step 402, in response to receiving a second speaker re-recognition request for the target voice data, the target recognition text, and the second target speaker number sent by the client, perform speaker recognition on the target voice data based on the second target speaker number, and update the target recognition text according to the speaker recognition result.

[0271] Here, the specific operations of step 402 and the technical effects thereof are substantially the same as the operations and effects of step 209 in the Figure 2A shown embodiment, and will not be elaborated here.

[0272] In some alternative embodiments, the above process method 400 may further include performing the following steps 403 and 404 after step 401:

[0273] Step 403, update the set of target participant identifiers according to the updated target recognition text.

[0274] Step 404, send the updated target recognition text and the set of target participant identifiers to the client.

[0275] Here, the specific operations of steps 403 and 404 and the technical effects they produce are basically the same as those of steps 210 and 211 in the Figure 2A embodiment shown, and will not be elaborated here.

[0276] In some alternative embodiments, the above process method 400 may further include the following step 405 after step 402:

[0277] Step 405: Send the updated target recognition text to the client.

[0278] Here, the specific operation of step 405 and the technical effect it produces are basically the same as those of step 213 in the Figure 2A embodiment shown, and will not be elaborated here.

[0279] The speaker recognition method provided by the above embodiments of the present disclosure can support the improvement of the speaker recognition accuracy of the voice data from the terminal device corresponding to the first target participant for the user on the client by processing the request for re - speaker recognition of the voice data corresponding to the first target participant in the target voice data of the associated voice data source device sent by the client to the server.

[0280] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speaker recognition device. This device embodiment corresponds to the Figure 3 method embodiment shown and can be specifically applied to various clients.

[0281] As shown in Figure 5As shown in the figure, the speaker recognition device 500 of this embodiment includes: a first presentation unit 501, a first determination unit 502, a second presentation unit 503, an acquisition unit 504, and a first transmission unit 505. Among them, the first presentation unit 501 is configured to present a target recognition text formed by performing speech recognition and speaker recognition on target speech data; the first determination unit 502 is configured to determine whether the target speech data is associated with a voice data source device in response to detecting a first recognition operation on the target speech data and the target recognition text; the second presentation unit 503 is configured to present a set of target participant identifiers corresponding to the target speech data in response to determining that the target speech data is associated with a voice data source device; the acquisition unit 504 is configured to acquire the first target participant identifier in response to detecting a selection operation on a first target participant identifier in the set of target participant identifiers; the first transmission unit 505 is configured to generate a first speaker re-recognition request for the target speech data, the target recognition text, and the first target participant identifier in response to detecting a second recognition operation on the first target participant identifier, and send the first speaker re-recognition request to the server.

[0282] In this embodiment, for the specific processing of the first presentation unit 501, the first determination unit 502, the second presentation unit 503, the acquisition unit 504, and the first transmission unit 505 of the speaker recognition device 500 and the technical effects brought by them, reference can be made to Figure 3 the relevant descriptions of steps 301, step 302, step 303, step 304, and step 305 in the corresponding embodiments, which will not be elaborated here.

[0283] In some alternative embodiments, the first transmission unit 505 may be further configured to:

[0284] Determine whether a specified speaker number operation for the first target participant identifier is detected;

[0285] In response to determining that it is detected, present a first speaker number editing display object;

[0286] In response to detecting a second recognition operation on the first target participant identifier and the current speaker number corresponding to the first speaker number editing display object, generate the first speaker re-recognition request for the target speech data, the target recognition text, the first target participant identifier, and the first target speaker number, where the first target speaker number corresponds to the current speaker number corresponding to the first speaker number editing display object.

[0287] In some alternative embodiments, the first transmission unit 505 may be further configured to:

[0288] In response to determining that the specified number of speaker operations for the first target participant identifier is not detected, and detecting a second recognition operation for the first target participant identifier, a first speaker re-identification request is generated for the target voice data, the target recognition text, and the first target participant identifier.

[0289] In some alternative embodiments, the apparatus 500 may further include:

[0290] A third presentation unit 506, configured to present a second speaker number editing display object in response to determining that the target voice data is not associated with a voice data source device;

[0291] A second sending unit 507, configured to generate a second speaker re-identification request for the target voice data, the target recognition text, and a second target speaker number in response to detecting a third recognition operation for the current speaker number corresponding to the second speaker number editing display object, and send the second speaker re-identification request to the server, where the second target speaker number corresponds to the current speaker number corresponding to the second speaker number editing display object.

[0292] In some alternative embodiments, the apparatus 500 may further include:

[0293] A fourth presentation unit 508, configured to perform the following second progress value presentation operation in real time after sending the second speaker re-identification request to the server: calculating a second real-time duration difference between the current duration and the time when the second speaker re-identification request is sent to the server, and determining a ratio of the second real-time duration difference to a second total duration received from the server as a second recognition progress value, and presenting the second recognition progress value.

[0294] In some alternative embodiments, the apparatus 500 may further include:

[0295] A fifth presentation unit 509, configured to present the updated target recognition text and the target participant identifier set after receiving the updated target recognition text and the target participant identifier set sent by the server.

[0296] In some alternative embodiments, the apparatus 500 may further include:

[0297] A sixth presentation unit 510, configured to present the updated target recognition text in response to receiving the updated target recognition text sent by the server.

[0298] In some alternative embodiments, the apparatus 500 may further include:

[0299] A seventh presentation unit 511, configured to, after sending the first speaker re-identification request to the server, perform the following first progress value presentation operation in real time: calculate a first real-time duration difference between the current duration and the time when the first speaker re-identification request is sent to the server, and determine a ratio of the first real-time duration difference to a first total duration received from the server as a first identification progress value, and present the first identification progress value.

[0300] In some alternative embodiments, the first determination unit 502 may be further configured to:

[0301] Determine whether the target recognition text is in an edited state;

[0302] In response to determining that the target recognition text is not in an edited state, determine whether the target voice data is associated with a voice data source device.

[0303] In some alternative embodiments, the apparatus 500 may further include:

[0304] An eighth presentation unit 512, configured to, in response to determining that the target recognition text is in an edited state, present a first prompt message indicating that the target recognition text is in an edited state and cannot be used for speaker re-identification.

[0305] It should be noted that the implementation details and technical effects of each unit in the speaker recognition apparatus provided in the embodiments of the present disclosure can be referred to the descriptions of other embodiments in the present disclosure, and will not be elaborated here.

[0306] Further referring to Figure 6 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speaker recognition apparatus. This apparatus embodiment corresponds to the Figure 4 method embodiment shown, and this apparatus can be specifically applied to various servers.

[0307] As shown in Figure 6 , the speaker recognition apparatus 600 of this embodiment includes: a first recognition and update unit 601. Among them, the first recognition and update unit 601 is configured to, in response to receiving a first speaker re-identification request for target voice data, a target recognition text, and a first target participant identifier in a target participant identifier set corresponding to the target voice data sent by the client, perform speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and update the target recognition text according to the speaker recognition result.

[0308] In this embodiment, for the specific processing of the first recognition and update unit 601 of the speaker recognition device 600 and the technical effects brought thereby, reference may be made to Figure 4 the relevant description of step 401 in the corresponding embodiment, which will not be elaborated herein.

[0309] In some alternative embodiments, the first speaker re - recognition request may further include the number of first target speakers; and

[0310] Performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier includes:

[0311] Performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier based on the number of first target speakers.

[0312] In some alternative embodiments, the device 600 may further include:

[0313] A second recognition and update unit 602, configured to, in response to receiving a second speaker re - recognition request sent by the client for the target voice data, the target recognition text, and the number of second target speakers, perform speaker recognition on the target voice data based on the number of second target speakers, and update the target recognition text according to the speaker recognition result.

[0314] In some alternative embodiments, the device 600 may further include:

[0315] A first update unit 603, configured to, after performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier and updating the target recognition text according to the speaker recognition result, update the target participant identifier set according to the updated target recognition text;

[0316] A first sending unit 604, configured to send the updated target recognition text and the target participant identifier set to the client for the client to present the updated target recognition text and the target participant identifier set.

[0317] In some alternative embodiments, the device 600 may further include:

[0318] A second sending unit 605, configured to, after performing speaker recognition on the target voice data based on the second target number of speakers and updating the target recognition text according to the speaker recognition result, send the updated target recognition text to the client for the client to present the updated target recognition text.

[0319] In some alternative embodiments, the apparatus 600 may further include:

[0320] A third sending unit 606, configured to, before performing speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data, determine a first total duration of speaker recognition according to the size of the voice data from the client indicated by the first target participant identifier in the target voice data, and send the first total duration to the client.

[0321] In some alternative embodiments, the apparatus 600 may further include:

[0322] A fourth sending unit 607, configured to, before performing speaker recognition on the target voice data based on the second target number of speakers, determine a second total duration of speaker recognition according to the size of the target voice data, and send the second total duration to the client.

[0323] In some alternative embodiments, performing speaker recognition on the voice data from the client indicated by the first target participant identifier in the target voice data based on the first target number of speakers may include:

[0324] Segmenting the voice data from the client indicated by the first target participant identifier in the target voice data according to the voice recognition result and punctuation annotation result corresponding to the target voice data to obtain a sentence audio sequence;

[0325] Performing audio segment segmentation on the sentence audio in the sentence audio sequence to form an audio segment set;

[0326] Extracting voiceprint features from each audio segment to obtain a voiceprint feature set;

[0327] Determining the first target number of speakers as the number of speaker clustering centers;

[0328] Clustering the voiceprint features in the voiceprint feature set according to the number of speaker clustering centers to obtain the number of speaker clustering centers of voiceprint feature clusters;

[0329] For the sentence audio in the sentence audio sequence, determine the speaker identification corresponding to the sentence audio as the speaker identification corresponding to the acoustic feature clustering center with the most occurrences among the acoustic feature clustering centers to which the acoustic features corresponding to the respective audio segments included in the sentence video belong, and determine the acoustic feature clustering center to which the acoustic features corresponding to the respective audio segments included in the sentence audio belong.

[0330] In some alternative embodiments, the updating of the target recognition text according to the speaker recognition result may include:

[0331] In response to determining that the number of speaker identifications included in the speaker recognition result is less than or equal to a preset number threshold, update the target recognition text according to the speaker recognition result.

[0332] In some alternative embodiments, the updating of the target recognition text according to the speaker recognition result may include:

[0333] In response to determining that the number of speaker identifications included in the speaker recognition result is less than or equal to a preset number threshold, update the target recognition text according to the speaker recognition result.

[0334] It should be noted that the implementation details and technical effects of each unit in the speaker recognition device provided by the embodiments of the present disclosure may refer to the descriptions of other embodiments in the present disclosure, and will not be elaborated herein.

[0335] The following refers to Figure 7 , which shows a schematic structural diagram of a computer system 700 suitable for use in implementing the embodiments of the present disclosure for a client or a server. Figure 7 The illustrated computer system 700 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0336] As Figure 7 shown, the computer system 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0337] Typically, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and a communication device 709. The communication device 709 can allow the computer system 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 a computer system 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0338] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0339] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0340] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0341] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to implement Figure 3 the speaker recognition method shown in the embodiments and their optional embodiments as Figure 4 shown, and / or, the speaker recognition method shown in the embodiments and their optional embodiments as

[0342] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0343] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0344] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the acquisition unit may also be described as "the unit for acquiring the identification of the first target participant".

[0345] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, and should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned disclosure concept. For example, the technical solutions formed by mutually replacing the above-mentioned features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

Claims

1. A speaker recognition method, the method comprising: Presenting a target recognition text formed by performing speech recognition and speaker recognition on target speech data; In response to detecting a first recognition operation on the target speech data and the target recognition text, determining whether the target speech data is associated with a voice data source device; In response to determining that the target speech data is associated with a voice data source device, presenting a set of target participant identifiers corresponding to the target speech data; In response to detecting a selection operation on a first target participant identifier in the set of target participant identifiers, obtaining the first target participant identifier; In response to detecting a second recognition operation on the first target participant identifier, generating a first speaker re-identification request for the target speech data, the target recognition text, and the first target participant identifier, and sending the first speaker re-identification request to a server.

2. The method according to claim 1, wherein In response to detecting a second recognition operation on the first target participant identifier, generating a first speaker re-identification request for the target speech data, the target recognition text, and the first target participant identifier, and sending the first speaker re-identification request to a server, includes: Determining whether a specified speaker number operation on the first target participant identifier is detected; In response to determining that it is detected, presenting a first speaker number editing display object; In response to detecting a second recognition operation on the first target participant identifier and the current speaker number corresponding to the first speaker number editing display object, generating the first speaker re-identification request for the target speech data, the target recognition text, the first target participant identifier, and a first target speaker number, the first target speaker number corresponding to the current speaker number corresponding to the first speaker number editing display object.

3. The method according to claim 2, wherein The step of, in response to detecting a second recognition operation on the first target participant identifier, generating a first speaker re-identification request for the target speech data, the target recognition text, and the first target participant identifier, and sending the first speaker re-identification request to a server, further includes: In response to determining that no specified speaker number operation on the first target participant identifier is detected and detecting a second recognition operation on the first target participant identifier, generating a first speaker re-identification request for the target speech data, the target recognition text, and the first target participant identifier.

4. The method according to claim 1, wherein, The method further includes: In response to determining that the target speech data is not associated with a voice data source device, presenting a second speaker number editing display object; In response to detecting a third recognition operation on the current speaker number corresponding to the second speaker number editing display object, generating a second speaker re-identification request for the target speech data, the target recognition text, and a second target speaker number, and sending the second speaker re-identification request to the server, the second target speaker number corresponding to the current speaker number corresponding to the second speaker number editing display object.

5. The method according to claim 4, wherein After sending the second speaker re-identification request to the server, the method further includes: Performing the following second progress value presentation operation in real time: calculating a second real-time duration difference between the current time and the time when the second speaker re-identification request is sent to the server, and determining a ratio of the second real-time duration difference to a second total duration received from the server as a second identification progress value, and presenting the second identification progress value.

6. The method according to claim 1, wherein, The method further includes: In response to receiving the updated target recognition text and the target attendee identification set sent by the server, presenting the updated target recognition text and the target attendee identification set.

7. The method according to claim 1, wherein, The method further includes: In response to receiving the updated target recognition text sent by the server, presenting the updated target recognition text.

8. According to the method as claimed in any one of claims 1-7, wherein, After sending the first speaker re-identification request to the server, the method further includes: Performing the following first progress value presentation operation in real time: calculating a first real-time duration difference between the current time and the time when the first speaker re-identification request is sent to the server, and determining a ratio of the first real-time duration difference to a first total duration received from the server as a first identification progress value, and presenting the first identification progress value.

9. The method according to claim 1, wherein, The determining whether the target voice data is associated with a voice data source device includes: Determining whether the target recognition text is in an edited state; In response to determining that the target recognition text is not in an edited state, determining whether the target voice data is associated with a voice data source device.

10. The method according to claim 9, wherein, The method further includes: In response to determining that the target recognition text is in an edited state, presenting a first prompt message indicating that the target recognition text is in an edited state and speaker re-identification cannot be performed.

11. A speaker recognition method, including: In response to receiving a first speaker re-identification request for target voice data, a target recognition text, and a first target attendee identification in a target attendee identification set corresponding to the target voice data sent by a client, performing speaker recognition on the voice data in the target voice data originating from the client indicated by the first target attendee identification, and updating the target recognition text according to the speaker recognition result, wherein the target recognition text is formed by performing speech recognition and speaker recognition on the target voice data.

12. The method according to claim 11, wherein, The first speaker re-identification request further includes a first target speaker number; and The performing speaker recognition on the voice data in the target voice data originating from the client indicated by the first target attendee identification includes: Based on the first target speaker number, performing speaker recognition on the voice data in the target voice data originating from the client indicated by the first target attendee identification.

13. The method according to claim 11, wherein, The method further includes: In response to receiving a second speaker re-identification request sent by the client for the target voice data, the target recognition text, and the second target number of speakers, perform speaker recognition on the target voice data based on the second target number of speakers, and update the target recognition text according to the speaker recognition result.

14. According to the method of any one of claims 11-13, wherein After performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and updating the target recognition text according to the speaker recognition result, the method further includes: Update the target participant identifier set according to the updated target recognition text; Send the updated target recognition text and the target participant identifier set to the client for the client to present the updated target recognition text and the target participant identifier set.

15. The method according to claim 13, wherein, After performing speaker recognition on the target voice data based on the second target number of speakers, and updating the target recognition text according to the speaker recognition result, the method further includes: Send the updated target recognition text to the client for the client to present the updated target recognition text.

16. The method according to claim 11, wherein, Before performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier, the method further includes: Determine the first total duration of speaker recognition according to the size of the voice data in the target voice data that originates from the client indicated by the first target participant identifier, and send the first total duration to the client.

17. The method according to claim 13, wherein, Before performing speaker recognition on the target voice data based on the second target number of speakers, the method further includes: Determine the second total duration of speaker recognition according to the size of the target voice data, and send the second total duration to the client.

18. The method according to claim 12, wherein, Performing speaker recognition on the voice data in the target voice data that originates from the client indicated by the first target participant identifier based on the first target number of speakers includes: Segment the voice data in the target voice data that originates from the client indicated by the first target participant identifier according to the voice recognition result and punctuation mark annotation result corresponding to the target voice data to obtain a sentence audio sequence; Perform audio segment segmentation on the sentence audio in the sentence audio sequence to form an audio segment set; Extract voiceprint features from each audio segment to obtain a voiceprint feature set; Determine the first target number of speakers as the number of speaker clustering centers; Cluster the voiceprint features in the voiceprint feature set according to the number of speaker clustering centers to obtain the number of voiceprint feature clustering centers equal to the number of speaker clustering centers; For the sentence audio in the sentence audio sequence, determine the speaker feature cluster center to which the voiceprint feature corresponding to each audio segment included in the sentence audio belongs, and determine the speaker identifier corresponding to the sentence audio as the speaker identifier corresponding to the voiceprint feature cluster center with the most occurrences among the voiceprint feature cluster centers to which the voiceprint features corresponding to each audio segment included in the sentence audio belong.

19. The method according to claim 11, wherein, The updating of the target recognition text according to the speaker recognition result includes: In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, updating the target recognition text according to the speaker recognition result.

20. The method according to claim 13, wherein, The updating of the target recognition text according to the speaker recognition result includes: In response to determining that the number of speaker identifiers included in the speaker recognition result is less than or equal to a preset number threshold, updating the target recognition text according to the speaker recognition result.

21. A speaker recognition device, the device includes: A first presentation unit configured to present a target recognition text formed by performing speech recognition and speaker recognition on target speech data; A first determination unit configured to determine whether the target speech data is associated with a voice data source device in response to detecting a first recognition operation on the target speech data and the target recognition text; A second presentation unit configured to present a set of target participant identifiers corresponding to the target speech data in response to determining that the target speech data is associated with a voice data source device; An acquisition unit configured to acquire the first target participant identifier in response to detecting a selection operation on a first target participant identifier in the set of target participant identifiers; A first sending unit configured to generate a first speaker re-identification request for the target speech data, the target recognition text, and the first target participant identifier in response to detecting a second recognition operation on the first target participant identifier, and send the first speaker re-identification request to a server.

22. A speaker recognition device, including: A recognition and update unit configured to perform speaker recognition on the speech data in the target speech data that originates from the client indicated by the first target participant identifier in response to receiving a first speaker re-identification request sent by the client for the target speech data, the target recognition text, and the set of target participant identifiers corresponding to the target speech data, and update the target recognition text according to the speaker recognition result, where the target recognition text is formed by performing speech recognition and speaker recognition on the target speech data.

23. A client, including: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method according to any one of claims 1-10.

24. A server, including: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 11-20.

25. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by one or more processors, the method according to any one of claims 1-10 and / or the method according to any one of claims 11-20 is implemented.

26. A speaker recognition system, comprising a client according to claim 23 and a server according to claim 24.

Citation Information

Patent Citations

  • Conference record generation method based on voice recognition, device and storage medium

    CN110335612A

  • Identity recognition method and device

    CN110880325A