Speech recognition method and device, electronic equipment and storage medium

By segmenting speech and matching voiceprints, the problems of recognizing multiple speakers and long audio clips are solved, and accurate identification of the main speaker and speech quality are achieved.

CN115410554BActive Publication Date: 2025-11-18MOBVOI (WUHAN) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211057400.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-11-18
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle to effectively recognize long audio clips containing multiple speakers and/or long segments without human voices, leading to recognition errors or failures to recognize audio.

Method used

The speech to be recognized is segmented into speech segments of equal length, and each speech segment is further segmented into overlapping speech frames. Voiceprint recognition features are extracted, and a voiceprint embedding code is generated. The speaker is determined by matching the voiceprint embedding code with the pre-registered voiceprint embedding code, and the main speaker is determined based on the confidence level.

Benefits of technology

It effectively removes the influence of invalid audio segments, accurately identifies the main speaker of the speech, and distinguishes between speech with good voice instructions and speech with poor voice quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410554B_ABST
    Figure CN115410554B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, device, electronic equipment and storage medium. The speech recognition method comprises: dividing the speech to be recognized into a plurality of speech segments of the same length; and obtaining the speaker of each speech segment by: dividing the speech segment into a plurality of speech frames of the same length and overlapping with each other; obtaining a voiceprint recognition feature of each speech frame in the speech segment; obtaining a voiceprint embedding code of the speech segment according to the voiceprint recognition features of all speech frames in the speech segment; and determining the speaker of the speech segment according to the voiceprint embedding code of the speech segment and a pre-registered voiceprint embedding code. The embodiments of the present disclosure can not only effectively remove the influence of invalid audio segments on the entire audio speaker recognition, thereby accurately identifying the speaker of the speech, but also identify the quality of the speech, and identify the speech with better voice instructions and the speech with poor voice quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a speech recognition method, apparatus, electronic device, and storage medium. Background Technology

[0002] Current speech recognition technology is mainly suitable for speech with only one speaker or speech with good default audio quality. For long audio clips containing multiple speakers and / or long segments without voice, existing speech recognition technologies suffer from problems such as failure to recognize or misrecognition. Summary of the Invention

[0003] To address at least one of the aforementioned technical problems, this disclosure provides a speech recognition method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of this disclosure, a speech recognition method is provided, comprising:

[0005] The speech to be recognized is segmented into multiple speech segments of equal length;

[0006] The speaker for each of the aforementioned speech segments is obtained as follows:

[0007] The speech segment is divided into multiple speech frames of the same length that overlap each other;

[0008] Obtain the voiceprint recognition features of each speech frame in the speech segment;

[0009] The voiceprint embedding code of the speech segment is obtained based on the voiceprint recognition features of all speech frames in the speech segment.

[0010] The speaker of the speech segment is determined based on the voiceprint embedding code and the pre-registered voiceprint embedding code of the speech segment.

[0011] Among the possible implementations of the first aspect of this disclosure are:

[0012] Based on the speakers of all speech segments in the speech to be identified, determine the confidence level of each speaker corresponding to the speech to be identified;

[0013] The main speaker of the speech to be identified is determined based on the confidence level of each speaker corresponding to the speech to be identified.

[0014] In some possible implementations of the first aspect of this disclosure, before determining the main speaker of the speech to be identified, the method further includes: determining whether the speech to be identified has a main speaker based on the confidence level of each speaker corresponding to the speech to be identified and a preset proportion threshold.

[0015] In some possible implementations of the first aspect of this disclosure, determining the confidence level of each speaker corresponding to the speech to be recognized based on the speakers of all speech segments in the speech to be recognized includes:

[0016] Based on the speakers of all speech segments in the speech to be identified, determine the number of speech segments for each speaker corresponding to the speech to be identified;

[0017] The confidence level of each speaker is determined based on the number of speech segments for each speaker and the total number of speech segments to be identified.

[0018] In some possible implementations of the first aspect of this disclosure, the confidence level is the ratio of the number of speech segments of the speaker to the total number of speech segments.

[0019] In some possible implementations of the first aspect of this disclosure, determining whether the speech to be recognized has a main speaker based on the confidence level of each speaker corresponding to the speech to be recognized and a preset proportion threshold includes:

[0020] Determine whether there is a speaker among all speakers corresponding to the speech to be recognized whose confidence level is greater than or equal to the preset proportion threshold;

[0021] When there is a speaker whose confidence level is greater than or equal to a preset proportion threshold, it is determined that the speech to be identified has a main speaker;

[0022] If there is no speaker whose confidence level is greater than or equal to a preset percentage threshold, it is determined that the speech to be identified does not have a main speaker.

[0023] In some possible implementations of the first aspect of this disclosure, determining the main speaker of the speech to be identified based on the confidence level of each speaker corresponding to the speech to be identified includes: when it is determined that the speech to be identified has a main speaker, determining the speaker with the highest confidence level as the main speaker of the speech to be identified.

[0024] In some possible implementations of the first aspect of this disclosure, determining the speaker of the speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code includes: obtaining the similarity between the voiceprint embedding code of the speech segment and each pre-registered voiceprint embedding code, and determining the speaker corresponding to the pre-registered voiceprint embedding code whose similarity is greater than a preset similarity threshold as the speaker of the speech segment.

[0025] According to a second aspect of this disclosure, a voice recognition device is provided, comprising:

[0026] The first segmentation unit is used to divide the speech to be recognized into multiple speech segments of the same length;

[0027] The second segmentation unit is used to segment a speech segment into multiple speech frames of the same length that overlap each other;

[0028] A voiceprint recognition unit is used to acquire the voiceprint recognition features of each speech frame in the speech segment;

[0029] An embedding code extraction unit is used to obtain the voiceprint embedding code of the speech segment based on the voiceprint recognition features of all speech frames in the speech segment.

[0030] The first determining unit is used to determine the speaker of the speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code of the speech segment.

[0031] Some possible implementations of the second aspect of this disclosure further include: a confidence unit, configured to determine the confidence level of each speaker corresponding to the speech to be identified based on the speakers of all speech segments in the speech to be identified; and a second determination unit, configured to determine the main speaker of the speech to be identified based on the confidence level of each speaker corresponding to the speech to be identified.

[0032] Some possible implementations of the second aspect of this disclosure further include: a determination unit, configured to determine whether the speech to be identified has a main speaker based on the confidence level of each speaker corresponding to the speech to be identified and a preset proportion threshold.

[0033] In some possible implementations of the second aspect of this disclosure, the confidence unit is specifically used to: determine the number of speech segments for each speaker corresponding to the speech to be identified based on the speakers of all speech segments in the speech to be identified; and determine the confidence level of each speaker based on the number of speech segments for each speaker and the total number of speech segments in the speech to be identified.

[0034] In some possible implementations of the second aspect of this disclosure, the confidence level is the ratio of the number of speech segments of the speaker to the total number of speech segments.

[0035] In some possible implementations of the second aspect of this disclosure, the determination unit is specifically used to: determine whether there is a speaker among all speakers corresponding to the speech to be recognized whose confidence level is greater than or equal to the preset proportion threshold; when there is a speaker whose confidence level is greater than or equal to the preset proportion threshold, determine that the speech to be recognized has a main speaker; when there is no speaker whose confidence level is greater than or equal to the preset proportion threshold, determine that the speech to be recognized does not have a main speaker.

[0036] In some possible implementations of the second aspect of this disclosure, the second determining unit is specifically used to determine the speaker with the highest confidence level as the main speaker of the speech to be identified when the determining unit determines that the speech to be identified has a main speaker.

[0037] In some possible implementations of the second aspect of this disclosure, the first determining unit is specifically used to: obtain the similarity between the voiceprint embedding code of the speech segment and each pre-registered voiceprint embedding code, and determine the speaker corresponding to the pre-registered voiceprint embedding code whose similarity is greater than a preset similarity threshold as the speaker of the speech segment.

[0038] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0039] Memory, the memory storing execution instructions; and

[0040] A processor that executes the execution instructions stored in the memory, causing the processor to perform the above-described speech recognition method.

[0041] According to a fourth aspect of this disclosure, a readable storage medium is provided, wherein executable instructions are stored therein, which, when executed by a processor, are used to implement the above-described speech recognition method.

[0042] The speech recognition method of this disclosure segmentes the speech to be recognized, identifies the speaker in each segmented speech segment, and identifies the main speaker of the speech to be recognized by statistically analyzing the speakers in all speech segments. In this way, not only can the influence of invalid audio segments on the overall audio speaker recognition be effectively removed, thereby accurately identifying the speaker of the speech, but also the speech quality can be identified, distinguishing between speech with good voice instructions and speech with poor voice quality. Attached Figure Description

[0043] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.

[0044] Figure 1 This is a schematic flowchart of a speech recognition method according to one embodiment of the present disclosure.

[0045] Figure 2 This is an example diagram of a speech recognition device employing a hardware implementation of a processing system according to one embodiment of the present disclosure.

[0046] The specific labels in the attached figures are as follows:

[0047] 200 voice recognition devices

[0048] 300 bus

[0049] 400 processor

[0050] 500 memory

[0051] 600 Other circuits. Detailed Implementation

[0052] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.

[0053] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0054] Unless otherwise stated, the exemplary implementations / embodiments shown are to be understood as providing exemplary features of various details that provide ways in which the technical concepts of this disclosure can be implemented in practice. Therefore, unless otherwise stated, the features of various implementations / embodiments may be additionally combined, separated, interchanged and / or rearranged without departing from the technical concepts of this disclosure.

[0055] The use of crosshairs and / or shading in the accompanying drawings is generally used to clarify the boundaries between adjacent components. Thus, unless otherwise stated, the presence or absence of crosshairs or shading does not convey or indicate any preference or requirement for the specific material, material properties, dimensions, proportions, commonalities between the illustrated components, or any other characteristics, properties, etc., of the components. Furthermore, in the accompanying drawings, the dimensions and relative dimensions of components may be exaggerated for clarity and / or descriptive purposes. When exemplary embodiments can be implemented differently, a specific process sequence may be performed in a different order than that described. For example, two consecutively described processes may be performed substantially simultaneously or in the reverse order of their description. Furthermore, the same reference numerals denote the same components.

[0056] When a component is referred to as being "on" or "above" another component, "connected to," or "joined to" another component, the component may be directly on, directly connected to, or directly joined to the other component, or there may be intermediate components. However, when a component is referred to as being "directly on" another component, "directly connected to," or "directly joined to" another component, there are no intermediate components. Therefore, the term "connection" can refer to a physical connection, an electrical connection, etc., and may or may not have intermediate components.

[0057] The terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting. As used herein, unless the context clearly indicates otherwise, the singular forms “a” and “the” are intended to include the plural forms as well. Furthermore, when the terms “comprising” and / or “including” and variations thereof are used in this specification, it indicates the presence of the stated features, integrals, steps, operations, parts, components, and / or groups thereof, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, parts, components, and / or groups thereof. It should also be noted that, as used herein, the terms “substantially,” “about,” and other similar terms are used as approximate terms rather than as terms of degree, thus explaining the inherent biases in measurements, calculated values, and / or provided values ​​that would be recognized by one of ordinary skill in the art.

[0058] Figure 1 This is a schematic flowchart of a speech recognition method according to one embodiment of the present disclosure.

[0059] like Figure 1 As shown, the speech recognition method in this embodiment may include the following steps S12 to S110.

[0060] Step S12: Divide the speech to be recognized into multiple speech segments of the same length;

[0061] Specifically, in step S12, each speech to be recognized can be segmented according to a first preset duration to obtain multiple speech segments for each speech to be recognized, and the length of each speech segment is the first preset duration. The first preset duration can be flexibly set as needed. For example, assuming the first preset duration is 3 seconds and the length of a speech to be recognized is 43 seconds, then the speech to be recognized can be segmented into 14 speech segments, each with a length of 3 seconds.

[0062] After step S12, the speaker of each speech segment is obtained through the following steps S14 to S110.

[0063] Step S14: Divide the speech segment into multiple speech frames of the same length that overlap each other;

[0064] Specifically, in step S14, each speech segment can be divided into multiple speech frames, with the length of a single speech frame being a second preset length and adjacent speech frames overlapping for a third preset duration. Here, both the second and third preset durations can be flexibly set as needed.

[0065] For example, the second preset duration can be set to 25 milliseconds, and the third preset duration can be set to 15 milliseconds. That is, the length of a single speech frame is 25 milliseconds, and adjacent speech frames will overlap for 15 milliseconds. In this case, a speech frame can be segmented every 10 milliseconds.

[0066] Step S16: Obtain the voiceprint recognition features of each speech frame in the speech segment;

[0067] Specifically, in step S16, various applicable voiceprint recognition algorithms or systems can be used to process each speech segment to obtain the voiceprint recognition features of the speech segment. For example, the voiceprint recognition algorithm can be, but is not limited to, fbank (FilterBank), and the voiceprint recognition features can be, but are not limited to, Mel-frequency cepstral coefficients (MFCCs).

[0068] Step S18: Obtain the voiceprint embedding code of the speech segment based on the voiceprint recognition features of all speech frames in the speech segment.

[0069] In some implementations, in step S18, the pre-trained voiceprint model can be used to process the voiceprint recognition features of all speech frames in the speech segment, thereby extracting a fixed-dimensional voiceprint embedding code, which is the voiceprint embedding code of the speech segment. For example, the voiceprint model can be a machine learning model such as a neural network, which can be obtained through pre-training.

[0070] Step S110: Determine the speaker of the speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code of the speech segment.

[0071] In some implementations, step S110 may include: obtaining the similarity between the voiceprint embedding code of the speech segment and each pre-registered voiceprint embedding code, and determining the speaker corresponding to the pre-registered voiceprint embedding code with the highest similarity as the speaker of the speech segment.

[0072] A pre-registered voiceprint embedding code refers to a voiceprint embedding code that has been pre-registered with speaker information. This pre-registered voiceprint embedding code can be obtained in advance through various methods. For example, steps S14 to S18 can be performed on a speech segment sample of a known speaker (the speech segment sample has the same length as the speech segment mentioned earlier, both being a first preset length) to obtain the voiceprint embedding code of the known speaker. Of course, other methods can also be used to obtain the pre-registered voiceprint embedding code, and this disclosure does not limit such methods.

[0073] In some implementations, a pre-trained machine learning model or an algorithm such as cosine distance can be used to obtain the similarity between each pre-registered voiceprint embedding code and the voiceprint embedding code of the speech segment.

[0074] In some implementations, before determining the speaker of a speech segment, step S110 may further include: determining whether the highest similarity corresponding to the voiceprint embedding code of the speech segment is greater than a preset similarity threshold λ (0 < λ < 1). If the highest similarity corresponding to the voiceprint embedding code of the speech segment is greater than the preset similarity threshold, it indicates that the speech segment corresponds to the same speaker, and the speaker corresponding to the pre-registered voiceprint embedding code with the highest similarity can be determined as the speaker of the speech segment. If the highest similarity corresponding to the voiceprint embedding code of the speech segment is less than or equal to the preset similarity threshold, it indicates that the speech segment does not correspond to the same speaker, and the speech segment may be background noise, and it can be determined that the speech segment has no speaker. The similarity threshold can be set according to requirements or taken as an empirical value.

[0075] In some implementations, the speech recognition method may further include:

[0076] Step S112: Determine the confidence level of each speaker corresponding to the speech to be recognized based on the speakers of all speech segments in the speech to be recognized;

[0077] In some implementations, step S112 may include the following steps a1 to a2:

[0078] Step a1: Determine the number of speech segments for each speaker in the speech to be recognized, based on the speakers of all speech segments in the speech to be recognized.

[0079] Step a2: Determine the confidence level of each speaker based on the number of speech segments for each speaker and the total number of speech segments to be recognized.

[0080] In some implementations, the confidence level is the ratio of the number of speech segments by the speaker to the total number of speech segments.

[0081] For example, suppose a speech to be recognized is segmented into N speech segments, corresponding to two speakers: speaker A and speaker B. Speaker A has M1 speech segments, and speaker B has M2 speech segments. The sum of M1 and M2 is less than or equal to N, where N is an integer greater than or equal to 2, and M1 and M2 are integers greater than or equal to 1. Then, the confidence level of speaker A is M1 / N, and the confidence level of speaker B is M2 / N. The " / " signifies "divided by".

[0082] Step S114: Determine the main speaker of the speech to be recognized based on the confidence level of each speaker corresponding to the speech to be recognized.

[0083] In some implementations, before determining the main speaker of the speech to be recognized in step S114, it may further include: determining whether the speech to be recognized has a main speaker based on the confidence level of each speaker corresponding to the speech to be recognized and a preset proportion threshold.

[0084] Specifically, it is determined whether there is a speaker with a confidence level greater than or equal to a preset ratio threshold among all speakers corresponding to the speech to be recognized; when there is a speaker with a confidence level greater than or equal to the preset ratio threshold, it is determined that the speech to be recognized has a main speaker; when there is no speaker with a confidence level greater than or equal to the preset ratio threshold, it is determined that the speech to be recognized does not have a main speaker.

[0085] The preset ratio threshold can take empirical values. For example, the preset ratio threshold can be set to a value greater than 0 and less than 1. For example, the preset ratio threshold can be 0.65, 0.7, 0.75, 0.8 or other values. In specific applications, the specific value of the preset ratio threshold can be flexibly set according to requirements such as recognition accuracy.

[0086] In some embodiments, step S114 may include: when the speech to be recognized has a main speaker, determining the speaker with the highest confidence level as the main speaker of the speech to be recognized.

[0087] In some embodiments, step S114 may include: when the speech to be recognized does not have a main speaker, determining that the speech to be recognized has no corresponding main speaker.

[0088] For example, assume that a certain speech to be recognized is segmented into N speech segments, and these N speech segments correspond to 2 speakers: Speaker A and Speaker B. The confidence level of Speaker A is M1 / N, the confidence level of Speaker B is M2 / N, and the preset ratio threshold is C (0 < C < 1). If either M1 / N or M2 / N is greater than or equal to C, it means that the speech to be recognized has a main speaker, and most of the speech to be recognized is the voice of the same person with good quality. At this time, the speaker corresponding to the higher one of M1 / N and M2 / N can be determined as the main speaker of the recognized speech; if both M1 / N and M2 / N are less than C, it means that the speech to be recognized does not have a main speaker, and most of the speech to be recognized may be background noise with poor quality.

[0089] In this way, while accurately identifying the main speaker of a certain speech to be recognized, it is possible to identify the speech with good voice quality.

[0090] [[ID=1十九]]The speech recognition method of the embodiments of the present disclosure segments the speech to be recognized, respectively performs speaker recognition on the segmented speech segments, and realizes the recognition of the main speaker of the speech to be recognized by counting the number of speakers of all speech segments. In this way, not only can the influence of invalid audio segments on the speaker recognition of the entire audio be effectively removed, so as to accurately identify the main speaker of a long speech, but also at the same time, not only can the influence of invalid audio segments on the speaker recognition of the entire audio be effectively removed, so as to accurately identify the main speaker of a long speech, but also the voice quality can be identified, and the speech with better voice commands and the speech with poor voice quality can be recognized.

[0091] Figure 2 An example diagram of a speech recognition device using a hardware implementation of a processing system is shown.

[0092] The apparatus may include corresponding modules that perform one or more steps in the flowchart above. Therefore, each or more steps in the flowchart above can be performed by a corresponding module, and the apparatus may include one or more of these modules. A module may be one or more hardware modules specifically configured to perform a corresponding step, or implemented by a processor configured to perform a corresponding step, or stored in a computer-readable medium for implementation by a processor, or implemented through some combination thereof.

[0093] This hardware architecture can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 300 connects various circuits, including one or more processors 400, memory 500, and / or hardware modules. Bus 300 can also connect various other circuits 600, such as peripherals, voltage regulators, power management circuits, external antennas, etc.

[0094] Bus 300 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, this diagram uses only one connection line, but this does not imply that there is only one bus or one type of bus.

[0095] Any process or method description in the flowcharts or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain. The processor performs the various methods and processes described above. For example, the method embodiments of this disclosure may be implemented as software programs tangibly contained in a machine-readable medium, such as memory. In some embodiments, part or all of the software program may be loaded and / or installed via memory and / or a communication interface. When the software program is loaded into memory and executed by the processor, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the processor may be configured to perform one of the methods described above by any other suitable means (e.g., by means of firmware).

[0096] The logic and / or steps represented in the flowchart or otherwise described herein may be specifically implemented in any readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0097] For the purposes of this specification, a "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM). Furthermore, a readable storage medium can even be paper or other suitable media on which a program can be printed, since a program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in memory.

[0098] It should be understood that various parts of this disclosure can be implemented in hardware, software, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0099] Those skilled in the art will understand that all or part of the steps of the methods described above can be implemented by a program instructing related hardware, and the program can be stored in a readable storage medium. When executed, the program includes one or a combination of the steps of the method implementation.

[0100] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. The storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0101] like Figure 2 As shown, the speech recognition device 200 of some embodiments of this disclosure may include:

[0102] The first segmentation unit 202 is used to segment the speech to be recognized into multiple speech segments of the same length;

[0103] The second segmentation unit 204 is used to segment a speech segment into multiple speech frames of the same length that overlap each other.

[0104] The voiceprint recognition unit 206 is used to acquire the voiceprint recognition features of each speech frame in the speech segment;

[0105] The embedding code extraction unit 208 is used to obtain the voiceprint embedding code of the speech segment based on the voiceprint recognition features of all speech frames in the speech segment.

[0106] The first determining unit 210 is used to determine the speaker of a speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code of the speech segment.

[0107] In some embodiments, the voice recognition device 200 may further include:

[0108] The confidence unit 212 is used to determine the confidence level of each speaker corresponding to the speech to be recognized based on the speakers of all speech segments in the speech to be recognized.

[0109] The second determining unit 214 is used to determine the main speaker of the speech to be recognized based on the confidence level of each speaker corresponding to the speech to be recognized.

[0110] In some embodiments, the speech recognition device 200 may further include a determination unit 216, used to determine whether the speech to be recognized has a main speaker based on the confidence level of each speaker corresponding to the speech to be recognized and a preset ratio threshold.

[0111] In some implementations, the confidence unit 212 may specifically be used to: determine the number of speech segments for each speaker corresponding to the speech to be recognized, based on the speakers of all speech segments in the speech to be recognized; and determine the confidence level for each speaker based on the number of speech segments for each speaker and the total number of speech segments in the speech to be recognized. For example, the confidence level may be the ratio of the number of speech segments for a speaker to the total number of speech segments.

[0112] In some implementations, the determination unit 216 may specifically be used to: determine whether there is a speaker with a confidence level greater than or equal to a preset proportion threshold among all speakers corresponding to the speech to be recognized; if there is a speaker with a confidence level greater than or equal to the preset proportion threshold, determine that the speech to be recognized has a main speaker; if there is no speaker with a confidence level greater than or equal to the preset proportion threshold, determine that the speech to be recognized does not have a main speaker.

[0113] In some implementations, the second determining unit 214 may specifically be used to determine the speaker with the highest confidence level as the main speaker of the speech to be recognized when the determining unit 216 determines that the speech to be recognized has a main speaker.

[0114] In some implementations, the first determining unit 210 is specifically used to: obtain the similarity between the voiceprint embedding code of the speech segment and each pre-registered voiceprint embedding code, and determine the speaker corresponding to the pre-registered voiceprint embedding code with a similarity greater than a preset similarity threshold as the speaker of the speech segment.

[0115] In some implementations, the first determining unit 210 is specifically used to: determine whether the highest similarity corresponding to the voiceprint embedding code of a speech segment is greater than a preset similarity threshold λ (0 < λ < 1). When the highest similarity corresponding to the voiceprint embedding code of a speech segment is greater than the preset similarity threshold, it indicates that the speech segment corresponds to the same speaker, and the speaker corresponding to the pre-registered voiceprint embedding code with the highest similarity can be determined as the speaker of the speech segment. When the highest similarity corresponding to the voiceprint embedding code of a speech segment is less than or equal to the preset similarity threshold, it indicates that the speech segment does not correspond to the same speaker, and the speech segment may be background noise, and it can be determined that the speech segment has no speaker. The similarity threshold can be set according to requirements or taken as an empirical value.

[0116] This disclosure also provides an electronic device, including: a memory storing execution instructions; and a processor or other hardware module executing the execution instructions stored in the memory, causing the processor or other hardware module to perform the above-described speech recognition method.

[0117] This disclosure also provides a readable storage medium storing execution instructions, which, when executed by a processor, are used to implement the above-described speech recognition method.

[0118] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment / mode or example is included in at least one embodiment / mode or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.

[0119] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0120] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.

Claims

1. A speech recognition method, characterized in that, include: The speech to be recognized is divided into multiple speech segments of equal length according to a first preset duration; wherein, the first preset duration is 3 seconds; The speaker for each speech segment is obtained as follows: the speech segment is divided into multiple speech frames of equal length that overlap each other; the voiceprint recognition features of each speech frame in the speech segment are obtained; the voiceprint embedding code of the speech segment is obtained based on the voiceprint recognition features of all speech frames in the speech segment; the speaker of the speech segment is determined based on the voiceprint embedding code and the pre-registered voiceprint embedding code; wherein, the length of a single speech frame is a second preset duration, and adjacent speech frames overlap for a third preset duration; the second preset duration is 25 milliseconds, and the third preset duration is 15 milliseconds; Determining the confidence level of each speaker corresponding to the speech to be identified based on the speakers of all speech segments in the speech to be identified includes: determining the number of speech segments for each speaker corresponding to the speech to be identified based on the speakers of all speech segments in the speech to be identified; determining the confidence level of each speaker based on the number of speech segments for each speaker and the total number of speech segments in the speech to be identified; wherein the confidence level is the ratio of the number of speech segments for a speaker to the total number of speech segments. The main speaker of the speech to be identified is determined based on the confidence level of each speaker corresponding to the speech to be identified. The step of determining the speaker of the speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code includes: obtaining the similarity between the voiceprint embedding code of the speech segment and each pre-registered voiceprint embedding code, and determining the speaker corresponding to the pre-registered voiceprint embedding code with a similarity greater than a preset similarity threshold as the speaker of the speech segment. Before determining the main speaker of the speech to be identified, the method further includes: determining whether the speech to be identified has a main speaker based on the confidence level and a preset proportion threshold for each speaker corresponding to the speech to be identified; determining whether the speech to be identified has a main speaker based on the confidence level and the preset proportion threshold for each speaker corresponding to the speech to be identified includes: determining whether there is a speaker among all speakers corresponding to the speech to be identified whose confidence level is greater than or equal to the preset proportion threshold; if there is a speaker whose confidence level is greater than or equal to the preset proportion threshold, determining that the speech to be identified has a main speaker; if there is no speaker whose confidence level is greater than or equal to the preset proportion threshold, determining that the speech to be identified does not have a main speaker.

2. The speech recognition method according to claim 1, characterized in that, The step of determining the main speaker of the speech to be identified based on the confidence level of each speaker corresponding to the speech to be identified includes: when it is determined that the speech to be identified has a main speaker, the speaker with the highest confidence level is determined as the main speaker of the speech to be identified.

3. A voice recognition device, characterized in that, include: The first segmentation unit is used to divide the speech to be recognized into multiple speech segments of the same length; The second segmentation unit is used to segment a speech segment into multiple speech frames of the same length that overlap each other; A voiceprint recognition unit is used to acquire the voiceprint recognition features of each speech frame in the speech segment; An embedding code extraction unit is used to obtain the voiceprint embedding code of the speech segment based on the voiceprint recognition features of all speech frames in the speech segment. The first determining unit is used to determine the speaker of the speech segment based on the voiceprint embedding code and the pre-registered voiceprint embedding code of the speech segment. A confidence unit is used to determine the confidence level of each speaker corresponding to the speech to be recognized based on the speakers of all speech segments in the speech to be recognized, including: determining the number of speech segments corresponding to each speaker of the speech to be recognized based on the speakers of all speech segments in the speech to be recognized; determining the confidence level of each speaker based on the number of speech segments of each speaker and the total number of speech segments of the speech to be recognized; wherein the confidence level is the ratio of the number of speech segments of a speaker to the total number of speech segments; The second determining unit is used to determine the main speaker of the speech to be recognized based on the confidence level of each speaker corresponding to the speech to be recognized; Before determining the main speaker of the speech to be identified, the method further includes: determining whether the speech to be identified has a main speaker based on the confidence level and a preset proportion threshold of each speaker corresponding to the speech to be identified; determining whether the speech to be identified has a main speaker based on the confidence level and the preset proportion threshold of each speaker corresponding to the speech to be identified includes: determining whether there is a speaker among all speakers corresponding to the speech to be identified whose confidence level is greater than or equal to the preset proportion threshold; if there is a speaker whose confidence level is greater than or equal to the preset proportion threshold, determining that the speech to be identified has a main speaker; if there is no speaker whose confidence level is greater than or equal to the preset proportion threshold, determining that the speech to be identified does not have a main speaker.

4. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 2.

5. A readable storage medium, characterized in that, The readable storage medium stores execution instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Method and device for identifying speaker in multi-person speaking

    CN108399923A

  • Voice data processing method and device, computer equipment and storage medium

    CN111613231A

  • Voiceprint recognition method and device, electronic equipment and storage medium

    CN114141252A