Speech recognition method and apparatus, readable medium, and electronic device

By extracting the voiceprint features of audio segments and performing overall similarity analysis, the labeling bias caused by background noise and unclear pronunciation in speech recognition is solved, thereby improving the accuracy of speech recognition and the certainty of audio segment attribution.

CN114023315BActive Publication Date: 2025-12-30BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111405339.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-12-30
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

During speech recognition, background noise and unclear pronunciation by the speaker can cause deviations in speaker marking, reducing the accuracy of speech recognition.

Method used

By extracting the voiceprint features of audio segments, the total similarity of each role label is determined, and role labels with a total similarity less than a threshold are marked as target role labels. The specified labels are used to indicate that the affiliation of the audio segment is pending, reducing the impact of background noise interference and unclear pronunciation.

Benefits of technology

It improves the accuracy of speech recognition, reduces the interference of specific audio segments on the accuracy of speech recognition, and enhances the certainty of audio segment attribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114023315B_ABST
    Figure CN114023315B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a speech recognition method and device, readable medium and electronic equipment, and relates to the technical field of electronic information processing. The method comprises: determining at least one audio segment included in the to-be-recognized speech and a role label to which each audio segment belongs according to the obtained to-be-recognized speech; determining a voiceprint feature of each audio segment; determining a total similarity of each role label according to the voiceprint features of the audio segments belonging to the role label; determining a target role label when the total similarity is less than a first similarity threshold; and marking the audio segments belonging to the target role label as specified labels, wherein the specified labels are used to indicate the ownership of the audio segments. The present disclosure extracts the voiceprint features of each audio segment to determine the total similarity of each role label, thereby determining the label of each audio segment, which can effectively improve the accuracy of speech recognition and effectively reduce the interference of specific audio segments on the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to electronic information processing technology, and more specifically, to a speech recognition method, apparatus, readable medium, and electronic device. Background Technology

[0002] With the continuous development of electronic information technology, voice, as an important carrier of information, has been widely used in daily life and work. Voice-related applications typically involve various voice processing methods, especially in scenarios such as telephone conferences and video conferences. These scenarios require the recognition of conference recordings or videos to transcribe the audio signal into text and mark the speaker for each segment of dialogue, allowing users to intuitively identify which speaker is making each conversation. However, audio signals often contain background noise and other interference, and there may be issues such as unclear pronunciation, leading to inaccuracies in speaker marking and reducing the accuracy of voice recognition. Summary of the Invention

[0003] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides a speech recognition method, the method comprising:

[0005] Based on the acquired speech to be recognized, determine at least one audio segment included in the speech to be recognized, and the role label to which each audio segment belongs;

[0006] Determine the voiceprint features of each audio segment;

[0007] The overall similarity of a character tag is determined based on the voiceprint features of the audio segments belonging to each character tag.

[0008] The character tags whose total similarity is less than a first similarity threshold are identified as target character tags, and the audio segments belonging to the target character tags are marked as designated tags, the designated tags being used to indicate that the affiliation of the audio segments is pending.

[0009] Secondly, this disclosure provides a speech recognition device, the device comprising:

[0010] The recognition module is used to determine, based on the acquired speech to be recognized, at least one audio segment included in the speech to be recognized, and the role label to which each audio segment belongs;

[0011] A feature determination module is used to determine the voiceprint features of each audio segment;

[0012] The total similarity determination module is used to determine the total similarity of a role tag based on the voiceprint features of the audio segments belonging to each role tag.

[0013] The processing module is used to determine the role tags whose total similarity is less than a first similarity threshold as target role tags, and to mark the audio segments belonging to the target role tags as designated tags, wherein the designated tags are used to indicate that the affiliation of the audio segments is pending.

[0014] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.

[0015] Fourthly, this disclosure provides an electronic device, comprising:

[0016] A storage device on which computer programs are stored;

[0017] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.

[0018] Through the above technical solution, this disclosure first determines, based on the acquired speech to be recognized, at least one audio segment included in the speech, and the role label to which each audio segment belongs. Then, it extracts the voiceprint features of each audio segment, and determines the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label. Finally, it marks the audio segments belonging to the target role label as designated labels, wherein the total similarity of the target role label is less than a first similarity threshold, and the designated labels are used to indicate that the attribution of the audio segments is pending. This disclosure, by extracting the voiceprint features of each audio segment to determine the total similarity of each role label, thereby determining the label of each audio segment, can effectively improve the accuracy of speech recognition and effectively reduce the interference of specific audio segments on the accuracy of speech recognition.

[0019] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1This is a flowchart illustrating a speech recognition method according to an exemplary embodiment;

[0022] Figure 2 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment;

[0023] Figure 3 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment;

[0024] Figure 4 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment;

[0025] Figure 5 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment;

[0026] Figure 6 This is a block diagram illustrating a speech recognition device according to an exemplary embodiment;

[0027] Figure 7 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment;

[0028] Figure 8 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment;

[0029] Figure 9 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment;

[0030] Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0033] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0037] Figure 1 This is a flowchart illustrating a speech recognition method according to an exemplary embodiment, such as... Figure 1 As shown, the method may include the following steps:

[0038] Step 101: Based on the acquired speech to be recognized, determine at least one audio segment included in the speech to be recognized, and the role label to which each audio segment belongs.

[0039] For example, the process begins by acquiring the speech to be recognized. This speech can be, for instance, speech captured by the terminal device's sound acquisition device (e.g., a microphone), speech stored locally on the terminal device, or speech retrieved from the internet or a database. There can be one or more speech samples, each with an arbitrary duration, such as 30 seconds, 120 seconds, or 2000 seconds; this disclosure does not impose specific limitations on this. After obtaining the speech, it can be recognized using a preset recognition method to determine at least one audio segment and the role label to which each audio segment belongs. An audio segment can be understood as a dialogue segment within the speech, and a role label can be understood as a speaker identifier, indicating that the corresponding audio segment belongs to that role label (i.e., the dialogue's affiliation). Role labels can be used to indicate a specific speaker or to distinguish between different speakers. For example, role labels can include: Zhang San, Li Si, and Wang Wu, used to indicate a specific speaker. Alternatively, role labels can include: A, B, and C, to distinguish three different speakers in the speech.

[0040] The preset recognition method can be any speaker recognition method or speaker segmentation and clustering method; this disclosure does not specifically limit this. For example, a pre-trained speaker recognition model can be used to recognize the speech to be recognized. The speaker recognition model can output at least one audio segment and the role label to which each audio segment belongs. Alternatively, the voiceprint features of each audio frame in the speech to be recognized can be extracted first, and then the audio frames can be clustered according to the voiceprint features. The temporally consecutive audio frames in the clusters are taken as an audio segment, and finally, a role label is assigned to each audio segment in the cluster. It should be noted that, while determining the role label to which each audio segment belongs, each audio segment can also be transcribed into text information, and then the text information corresponding to the audio segment can be labeled as the role label. That is to say, the role label can be used to label both audio segments and the text information corresponding to the audio segment; this disclosure does not specifically limit this.

[0041] Step 102: Determine the voiceprint features of each audio segment.

[0042] For example, features can be extracted from each audio segment to obtain the voiceprint features (also known as voiceprint embedding) for each segment. Voiceprint features can be understood as vectors that can distinguish different speakers and represent the audio segment. Voiceprint features can include multiple dimensions, such as pitch, volume, speech rate, noise level, tone, and loudness. Specifically, audio processing tools such as Sox, Librosa, and Straight can be used to extract the voiceprint features for each audio segment. Alternatively, a pre-trained voiceprint extraction model can be used to extract the voiceprint features for each audio segment. This model can be trained independently or it can be part of a speaker recognition model. In other words, the voiceprint features for each audio segment can be extracted separately after obtaining at least one audio segment and the role label to which each audio segment belongs, or it can be extracted during the process of obtaining at least one audio segment and the role label to which each audio segment belongs. This disclosure does not specifically limit this approach.

[0043] Step 103: Determine the total similarity of the character tag based on the voiceprint features of the audio segments belonging to each character tag.

[0044] Step 104: Character tags with a total similarity less than the first similarity threshold are identified as target character tags, and audio segments belonging to the target character tags are marked as designated tags. The designated tags are used to indicate that the affiliation of the audio segments is pending.

[0045] For example, since voiceprint features can distinguish different speakers, audio segments belonging to the same role label should have very similar voiceprint features, while audio segments belonging to different role labels should have dissimilar voiceprint features. Therefore, the total similarity for each role label can be determined separately to determine whether the labeling of that role is accurate. Specifically, audio segments belonging to that role label (which can be one or more) can be integrated first, and then the total similarity for that role label can be determined based on the voiceprint features of the audio segments belonging to that role label. The total similarity reflects the degree of similarity (or aggregation) between audio segments belonging to that role label, or it can be understood as reflecting the degree of similarity (or aggregation) between the voiceprint features of audio segments belonging to that role label. That is, the higher the total similarity, the more aggregated the audio segments belonging to that role label are; the lower the total similarity, the more discrete the audio segments belonging to that role label are.

[0046] Specifically, the determination of the total similarity can be achieved for example: first, determine the cosine similarity of the voiceprint features of every two audio segments belonging to the role label, and then determine the total similarity based on the variance of multiple cosine similarities. Alternatively, the covariance matrix of the voiceprint features of the audio segments belonging to the role label can be determined first, and then the total similarity can be determined based on the covariance matrix. After obtaining the total similarity for each role label, the total similarity for each role label can be compared with a first similarity threshold. If the total similarity is less than the first similarity threshold, then the corresponding role label can be determined as the target role label. Then, the audio segments belonging to the target role label are marked with a specified label, where the specified label indicates that the attribution of the audio segment is undetermined. That is, the audio segments belonging to the target role label determined in step 101 are marked with a specified label to indicate that the attribution of the audio segment is not the target role label; the audio segment may belong to other role labels, or it may not belong to any role label. In one implementation, the specified label can be in the form of an unknown label or a pending label, used to indicate that the speaker corresponding to the audio segment is undetermined. In another implementation, the designated labels can take the form of error labels and / or unknown labels. Error labels indicate that an audio segment is incorrectly labeled (i.e., the audio segment does not belong to the target role label), while unknown labels indicate that the affiliation of an audio segment is unknown (i.e., the audio segment does not belong to any role label). This allows for further differentiation and adjustment of audio segments labeled with the designated labels, thus avoiding recognition errors caused by background noise interference or unclear speech. The first similarity threshold can be adjusted according to specific needs; it is directly proportional to the accuracy of speech recognition and the number of audio segments labeled with the designated labels in the speech to be recognized.

[0047] In summary, this disclosure first determines at least one audio segment and the role label to which each audio segment belongs based on the acquired speech to be recognized. Then, it extracts the voiceprint features of each audio segment and determines the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label. Finally, it marks the audio segments belonging to the target role label as designated labels, wherein the total similarity of the target role label is less than a first similarity threshold, and the designated labels are used to indicate that the attribution of the audio segments is pending. This disclosure, by extracting the voiceprint features of each audio segment to determine the total similarity of each role label, thereby determining the label of each audio segment, can effectively improve the accuracy of speech recognition and effectively reduce the interference of specific audio segments on the accuracy of speech recognition.

[0048] Figure 2 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment, such as... Figure 2 As shown, the method may further include the following steps:

[0049] Step 105: Determine the total voiceprint features of the character tag based on the voiceprint features of the audio segments belonging to each character tag.

[0050] Step 106: Based on the voiceprint features of the audio segments belonging to the character tag and the total voiceprint features of the character tag, determine the similarity of the audio segments belonging to the character tag.

[0051] Step 107: Mark audio segments with similarity less than the second similarity threshold as specified tags.

[0052] For example, since voiceprint features can distinguish different speakers, the voiceprint features of audio segments belonging to the same role label should be very similar, while the voiceprint features of audio segments belonging to different role labels should be dissimilar. For each role label, if there are multiple audio segments belonging to that role label, then the voiceprint features of these multiple audio segments should be very similar. Therefore, the voiceprint features of multiple audio segments belonging to that role label can also be compared with the total voiceprint features of that role label to determine whether each audio segment belonging to that role label is accurately labeled.

[0053] Specifically, for each character tag, multiple audio segments belonging to that character tag can be integrated first. Then, based on the voiceprint features of these audio segments, the total voiceprint features of that character tag can be determined. The total voiceprint features can reflect the overall characteristics of the multiple audio segments belonging to that character tag. For example, the average of the voiceprint features of multiple audio segments can be used as the total voiceprint features of that character tag. Alternatively, the median of the voiceprint features of multiple audio segments can be used as the total voiceprint features of that character tag. Afterward, the similarity between the voiceprint features of each audio segment belonging to that character tag and the total voiceprint features of that character tag can be determined. Specifically, the cosine similarity (or Pearson correlation coefficient, Jaccard similarity coefficient, Euclidean distance, etc.) between the voiceprint features of each audio segment belonging to that character tag and the total voiceprint features of that character tag can be used as the similarity between each audio segment belonging to that character tag. For example, in the speech to be recognized, there are three audio segments, audio segment a, audio segment b, and audio segment c, which belong to role label A. The similarity of the voiceprint features of audio segment a with the cosine similarity of the total voiceprint features of role label A can be calculated as the similarity of audio segment a. Similarly, the similarity of the voiceprint features of audio segment b with the total voiceprint features of role label A can be calculated as the similarity of audio segment b. Finally, the similarity of the voiceprint features of audio segment c with the total voiceprint features of role label A can be calculated as the similarity of audio segment c. Thus, the similarity of the three audio segments can be obtained.

[0054] For each role label, the above steps are performed to obtain the similarity of each audio segment in the speech to be recognized. Then, the similarity of each audio segment is compared with a second similarity threshold. If the similarity is less than the second similarity threshold, the corresponding audio segment can be labeled with the specified tag, indicating that the attribution of that audio segment's tag is pending. The second similarity threshold can be adjusted according to specific needs; it is directly proportional to the accuracy of speech recognition and the number of audio segments in the speech to be recognized that are labeled with the specified tag.

[0055] Figure 3 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment, such as... Figure 3 As shown, step 103 can be implemented in the following ways:

[0056] Step 1031: Determine the voiceprint covariance matrix based on the voiceprint features of the audio segments belonging to the character tag.

[0057] Step 1032: Determine the total similarity of the character tag based on the voiceprint covariance matrix.

[0058] For example, to determine the overall similarity of each character tag, we can first determine the voiceprint covariance matrix based on the voiceprint features of the audio segments belonging to that character tag. Each element in the voiceprint covariance matrix is ​​the covariance between the voiceprint features of each audio segment belonging to that character tag. Then, we can determine the overall similarity of that character tag based on the trace or determinant of the voiceprint covariance matrix. The overall similarity is negatively correlated with the trace or determinant of the voiceprint covariance matrix; that is, the larger the trace (or determinant) of the voiceprint covariance matrix, the smaller the overall similarity, and vice versa. For example, the negative value of the trace of the voiceprint covariance matrix can be used as the overall similarity of that character tag.

[0059] In one implementation, step 105 can be achieved in the following way:

[0060] The average value of the voiceprint features of the audio segments belonging to each character label is taken as the total voiceprint features of that character label.

[0061] Step 106 can be achieved in the following way:

[0062] The voiceprint features of the audio segments belonging to the character tag are determined, and the cosine similarity between them and the total voiceprint features of the character tag is used as the similarity of the audio segments belonging to the character tag.

[0063] For example, when determining the similarity of each audio segment, we can first determine the total voiceprint features of the character tag to which that audio segment belongs. Specifically, the average of the voiceprint features of all audio segments belonging to that character tag can be used as the total voiceprint features of that character tag. Then, for each audio segment, we can calculate the cosine similarity between the voiceprint features of that audio segment and the total voiceprint features of the character tag to which that audio segment belongs, and use the cosine similarity as the similarity of that audio segment.

[0064] Figure 4 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment, such as... Figure 4 As shown, step 101 may include:

[0065] Step 1011: Cluster the multiple audio frames included in the speech to be recognized to obtain at least one audio segment.

[0066] Step 1012: Determine the text information corresponding to each audio segment and the role tag to which the audio segment belongs, based on each audio segment.

[0067] For example, to determine the audio segments and their corresponding role labels in a speech to be recognized, we can first extract multiple audio frames from the speech. Then, we cluster these audio frames to obtain at least one cluster, with each cluster containing multiple audio frames. For each cluster, we integrate the audio frames based on their temporal continuity within the speech to be recognized, resulting in at least one audio segment. This means that the audio frames within each segment are temporally continuous and belong to the same cluster. Next, we can use a pre-defined conversion model or algorithm to convert each audio segment into corresponding text information. Simultaneously, we can assign a role label to each cluster and mark audio segments belonging to the same cluster with the corresponding role label. Furthermore, we can mark the text information corresponding to audio segments belonging to the same cluster with the corresponding role label.

[0068] Figure 5 This is a flowchart illustrating another speech recognition method according to an exemplary embodiment, such as... Figure 5 As shown, the method may further include:

[0069] Step 108: Mark audio segments with a duration less than a specified duration threshold as specified tags. And / or, mark audio segments with corresponding text information whose text length is less than a specified length threshold as specified tags.

[0070] For example, the shorter the duration of an audio segment, the less accurate the voiceprint information contained in its voiceprint features. This can be understood as a smaller amount of voiceprint information, potentially leading to identification errors. Therefore, audio segments can be filtered based on their duration. If the duration is less than a specified threshold (e.g., 1 second), the corresponding audio segment is marked with a specified label, indicating that its attribution is pending. The duration of an audio segment is also reflected in the length of the corresponding text information. Longer audio segments correspond to longer text information, and vice versa. Therefore, filtering can also be based on the length of the corresponding text information. If the text length is less than a specified length (e.g., 2 characters), the corresponding audio segment is marked with a specified label, indicating that its attribution is pending. This helps avoid identification errors caused by inaccurate voiceprint information contained in the voiceprint features.

[0071] In summary, this disclosure first determines at least one audio segment and the role label to which each audio segment belongs based on the acquired speech to be recognized. Then, it extracts the voiceprint features of each audio segment and determines the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label. Finally, it marks the audio segments belonging to the target role label as designated labels, wherein the total similarity of the target role label is less than a first similarity threshold, and the designated labels are used to indicate that the attribution of the audio segments is pending. This disclosure, by extracting the voiceprint features of each audio segment to determine the total similarity of each role label, thereby determining the label of each audio segment, can effectively improve the accuracy of speech recognition and effectively reduce the interference of specific audio segments on the accuracy of speech recognition.

[0072] Figure 6 This is a block diagram illustrating a speech recognition device according to an exemplary embodiment, such as... Figure 6 As shown, the device 200 includes:

[0073] The recognition module 201 is used to determine, based on the acquired speech to be recognized, at least one audio segment included in the speech to be recognized, and the role label to which each audio segment belongs.

[0074] Feature determination module 202 is used to determine the voiceprint features of each audio segment.

[0075] The total similarity determination module 203 is used to determine the total similarity of the character tag based on the voiceprint features of the audio segments belonging to each character tag.

[0076] The processing module 204 is used to determine the character tags with a total similarity less than the first similarity threshold as target character tags, and to mark the audio segments belonging to the target character tags as specified tags. The specified tags are used to indicate that the affiliation of the audio segments is pending.

[0077] Figure 7 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment, such as... Figure 7 As shown, the device 200 may further include:

[0078] The total feature determination module 205 is used to determine the total voiceprint features of a role tag based on the voiceprint features of the audio segments belonging to each role tag.

[0079] The similarity determination module 206 is used to determine the similarity of audio segments belonging to the character tag based on the voiceprint features of the audio segments belonging to the character tag and the total voiceprint features of the character tag.

[0080] The processing module 204 is also used to mark audio segments with a similarity less than a second similarity threshold as specified tags.

[0081] Figure 8 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment, such as... Figure 8 As shown, the total similarity determination module 203 may include:

[0082] The matrix determination submodule 2031 is used to determine the voiceprint covariance matrix based on the voiceprint features of the audio segments belonging to the character tag.

[0083] The total similarity determination submodule 2032 is used to determine the total similarity of the character's tag based on the voiceprint covariance matrix.

[0084] In one implementation, the total feature determination module 205 can be used to:

[0085] The average value of the voiceprint features of the audio segments belonging to each character label is taken as the total voiceprint features of that character label.

[0086] The similarity determination module 206 can be used for:

[0087] The voiceprint features of the audio segments belonging to the character tag are determined, and the cosine similarity between them and the total voiceprint features of the character tag is used as the similarity of the audio segments belonging to the character tag.

[0088] Figure 9 This is a block diagram illustrating another speech recognition device according to an exemplary embodiment, such as... Figure 9 As shown, the recognition module 201 may include:

[0089] The clustering submodule 2011 is used to cluster multiple audio frames included in the speech to be recognized in order to obtain at least one audio segment.

[0090] The recognition submodule 2012 is used to determine the text information corresponding to each audio segment and the role tag to which the audio segment belongs, based on each audio segment.

[0091] In another implementation, the processing module 204 can also be used for:

[0092] Audio segments whose duration is less than a specified duration threshold are marked with a specified tag. And / or, audio segments whose corresponding text information has a text length less than a specified length threshold are marked with a specified tag.

[0093] In another implementation, the specified labels include: error labels and / or unknown labels. Error labels indicate that the audio segment is incorrectly labeled, and unknown labels indicate that the ownership of the audio segment is unknown.

[0094] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0095] In summary, this disclosure first determines at least one audio segment and the role label to which each audio segment belongs based on the acquired speech to be recognized. Then, it extracts the voiceprint features of each audio segment and determines the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label. Finally, it marks the audio segments belonging to the target role label as designated labels, wherein the total similarity of the target role label is less than a first similarity threshold, and the designated labels are used to indicate that the attribution of the audio segments is pending. This disclosure, by extracting the voiceprint features of each audio segment to determine the total similarity of each role label, thereby determining the label of each audio segment, can effectively improve the accuracy of speech recognition and effectively reduce the interference of specific audio segments on the accuracy of speech recognition.

[0096] The following is for reference. Figure 10 This document illustrates a structural diagram of an electronic device 300 suitable for implementing embodiments of the present disclosure (e.g., the execution entity in the above embodiments, which may be a terminal device or a server). The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0097] like Figure 10 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0098] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0099] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of embodiments of this disclosure.

[0100] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0101] In some implementations, terminal devices and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0102] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0103] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine, based on the acquired speech to be recognized, at least one audio segment included in the speech to be recognized, and a role label to which each audio segment belongs; determine the voiceprint features of each audio segment; determine the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label; determine the role labels whose total similarity is less than a first similarity threshold as target role labels, and mark the audio segments belonging to the target role labels as designated labels, wherein the designated labels are used to indicate that the attribution of the audio segments is pending.

[0104] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0106] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, a feature determination module can also be described as a "module for determining voiceprint features".

[0107] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0108] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0109] According to one or more embodiments of this disclosure, Example 1 provides a speech recognition method, comprising: determining at least one audio segment included in the speech to be recognized, and a role label to which each audio segment belongs, based on the acquired speech to be recognized; determining the voiceprint features of each audio segment; determining the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label; determining the role labels whose total similarity is less than a first similarity threshold as target role labels, and marking the audio segments belonging to the target role labels as designated labels, wherein the designated labels are used to indicate that the affiliation of the audio segments is pending.

[0110] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, the method further comprising: determining the total voiceprint features of the character tag based on the voiceprint features of the audio segments belonging to each of the character tags; determining the similarity between the voiceprint features of the audio segments belonging to the character tag and the total voiceprint features of the character tag; and marking the audio segments with similarity less than a second similarity threshold as the designated tag.

[0111] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein determining the total similarity of a role tag based on the voiceprint features of the audio segments belonging to each role tag includes: determining a voiceprint covariance matrix based on the voiceprint features of the audio segments belonging to the role tag; and determining the total similarity of the role tag based on the voiceprint covariance matrix.

[0112] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein determining the total voiceprint features of a role tag based on the voiceprint features of the audio segments belonging to each of the role tags includes: taking the average value of the voiceprint features of the audio segments belonging to each of the role tags as the total voiceprint features of the role tag; and determining the similarity of the audio segments belonging to the role tag based on the voiceprint features of the audio segments belonging to the role tag and the total voiceprint features of the role tag includes: determining the cosine similarity between the voiceprint features of the audio segments belonging to the role tag and the total voiceprint features of the role tag as the similarity of the audio segments belonging to the role tag.

[0113] According to one or more embodiments of this disclosure, Example 5 provides the methods of Examples 1 to 4, wherein determining at least one audio segment included in the acquired speech to be recognized, and the role label to which each audio segment belongs, includes: clustering multiple audio frames included in the speech to be recognized to obtain at least one audio segment; and determining the text information corresponding to each audio segment and the role label to which the audio segment belongs based on each audio segment.

[0114] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, the method further comprising: marking the audio segment with a duration less than a specified duration threshold as the specified tag; and / or, marking the audio segment with a corresponding text information text length less than a specified length threshold as the specified tag.

[0115] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein the specified label includes: an error label and / or an unknown label; the error label is used to indicate that the audio segment is incorrectly labeled, and the unknown label is used to indicate that the ownership of the audio segment is unknown.

[0116] According to one or more embodiments of this disclosure, Example 8 provides a speech recognition device, comprising: a recognition module, configured to determine at least one audio segment included in the speech to be recognized, and a role label to which each audio segment belongs, based on acquired speech to be recognized; a feature determination module, configured to determine the voiceprint features of each audio segment; a total similarity determination module, configured to determine the total similarity of the role label based on the voiceprint features of the audio segments belonging to each role label; and a processing module, configured to determine the role labels with a total similarity less than a first similarity threshold as target role labels, and mark the audio segments belonging to the target role labels as designated labels, wherein the designated labels are used to indicate that the affiliation of the audio segments is pending.

[0117] According to one or more embodiments of this disclosure, Example 9 provides an apparatus of Example 8, the apparatus further comprising: a total feature determination module, configured to determine the total voiceprint features of the audio segments belonging to each of the role tags based on the voiceprint features of the audio segments belonging to the role tag; a similarity determination module, configured to determine the similarity between the voiceprint features of the audio segments belonging to the role tag and the total voiceprint features of the role tag; and the processing module, further configured to mark the audio segments with similarity less than a second similarity threshold as the designated tag.

[0118] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 8, wherein the total similarity determination module includes: a matrix determination submodule, configured to determine a voiceprint covariance matrix based on the voiceprint features of the audio segment belonging to the role tag; and a total similarity determination submodule, configured to determine the total similarity of the role tag based on the voiceprint covariance matrix.

[0119] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 9, wherein the total feature determination module is configured to: take the average value of the voiceprint features of the audio segments belonging to each of the character tags as the total voiceprint features of the character tag; and the similarity determination module is configured to: determine the cosine similarity between the voiceprint features of the audio segments belonging to the character tag and the total voiceprint features of the character tag as the similarity of the audio segments belonging to the character tag.

[0120] According to one or more embodiments of this disclosure, Example 12 provides an apparatus of Examples 8 to 11, wherein the recognition module includes: a clustering submodule for clustering multiple audio frames included in the speech to be recognized to obtain at least one audio segment; and a recognition submodule for determining, based on each audio segment, the text information corresponding to the audio segment and the role tag to which the audio segment belongs.

[0121] According to one or more embodiments of this disclosure, Example 13 provides the apparatus of Example 12, wherein the processing module is further configured to: mark the audio segment with a duration less than a specified duration threshold as the specified tag; and / or, mark the audio segment with a corresponding text information text length less than a specified length threshold as the specified tag.

[0122] According to one or more embodiments of this disclosure, Example 14 provides the apparatus of Example 8, wherein the designated label includes: an error label and / or an unknown label; the error label is used to indicate that the audio segment is incorrectly labeled, and the unknown label is used to indicate that the ownership of the audio segment is unknown.

[0123] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the methods described in Examples 1 to 7.

[0124] According to one or more embodiments of this disclosure, Example 16 provides an electronic device including: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the methods described in Examples 1 to 7.

[0125] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0126] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0127] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A speech recognition method, characterized by, The method comprises: determining at least one audio segment included in the to-be-recognized speech and a role label to which each audio segment belongs according to the obtained to-be-recognized speech; determining a voiceprint feature of each audio segment; determining a total similarity of each role label according to the voiceprint features of the audio segments belonging to the role label; the total similarity reflects a similarity degree between the audio segments belonging to the role label; determining a target role label from the role labels whose total similarity is less than a first similarity threshold, and marking the audio segments belonging to the target role label as a specified label; the specified label is used to indicate that the attribution of an audio segment is pending; The method further comprises: determining a total voiceprint feature of each role label according to the voiceprint features of the audio segments belonging to the role label; determining a similarity of the audio segments belonging to the role label according to the voiceprint features of the audio segments belonging to the role label and the total voiceprint feature of the role label; marking the audio segments whose similarity is less than a second similarity threshold as the specified label.

2. The method of claim 1, wherein, The determination of the total similarity of each role label according to the voiceprint features of the audio segments belonging to the role label comprises: determining a voiceprint covariance matrix according to the voiceprint features of the audio segments belonging to the role label; determining the total similarity of the role label according to the voiceprint covariance matrix.

3. The method of claim 1, wherein, The determination of the total voiceprint feature of each role label according to the voiceprint features of the audio segments belonging to the role label comprises: taking an average value of the voiceprint features of the audio segments belonging to the role label as the total voiceprint feature of the role label. The determination of the similarity of the audio segments belonging to the role label according to the voiceprint features of the audio segments belonging to the role label and the total voiceprint feature of the role label comprises: determining a cosine similarity of the voiceprint features of the audio segments belonging to the role label and the total voiceprint feature of the role label as the similarity of the audio segments belonging to the role label.

4. The method according to any one of claims 1 to 3, characterized in that, The determination of at least one audio segment included in the to-be-recognized speech and a role label to which each audio segment belongs according to the obtained to-be-recognized speech comprises: clustering a plurality of audio frames included in the to-be-recognized speech to obtain at least one audio segment; determining, according to each audio segment, text information corresponding to the audio segment and the role label to which the audio segment belongs.

5. The method of claim 4, wherein, The method further comprises: marking the audio segments whose duration is less than a specified duration threshold as the specified label; and / or marking the audio segments whose text length of the corresponding text information is less than a specified length threshold as the specified label.

6. The method of claim 1, wherein, The specified label comprises an error label and / or an unknown label; the error label is used to indicate that an audio segment is marked incorrectly, and the unknown label is used to indicate that the attribution of an audio segment is unknown.

7. A speech recognition apparatus characterized by comprising: The device comprises: an identification module configured to determine at least one audio segment included in the to-be-recognized speech and a role label to which each audio segment belongs according to the obtained to-be-recognized speech; a feature determination module configured to determine a voiceprint feature of each audio segment; The total similarity determination module is configured to determine a total similarity of each role label according to the voiceprint features of the audio segments belonging to the role label, and the total similarity reflects a similarity degree between the audio segments belonging to the role label. The processing module is configured to determine the role label with the total similarity less than a first similarity threshold as a target role label, and mark the audio segments belonging to the target role label as specified labels, and the specified labels are used to indicate that the ownership of the audio segments is pending. The device can further include: The total feature determination module is configured to determine a total voiceprint feature of each role label according to the voiceprint features of the audio segments belonging to the role label. The similarity determination module is configured to determine a similarity of the audio segments belonging to the role label according to the voiceprint features of the audio segments belonging to the role label and the total voiceprint feature of the role label. The processing module is further configured to mark the audio segments with the similarity less than a second similarity threshold as specified labels.

8. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by the processing device to implement the steps of the method in any one of claims 1-6.

9. An electronic device, comprising: The device includes: A storage device having one or more computer programs stored thereon; One or more processing devices configured to execute the one or more computer programs in the storage device to implement the steps of the method in any one of claims 1-6.

Citation Information

Patent Citations

  • Audio recognition method and device and computer readable storage medium

    CN109686377A

  • Voiceprint recognition template protection algorithm

    CN112331215A

  • Voice role segmentation method and device, computer equipment and storage medium

    CN113192516A