Off-line speech recognition method and system for mixed languages

By adding calibration links and iterative calling language discrimination engines in the speech recognition process, cutting recording clips by sentences and re-recognizing them, the problem of multilingual recording recognition errors is solved, and the complete and correct recognition of mixed language recordings is achieved.

CN120431918APending Publication Date: 2025-08-05PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359832.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing single-language judgment and single voice recognition engine method cannot correctly recognize multilingual mixed call recordings, resulting in errors in recognition text and affecting recording quality inspection.

Method used

By adding calibration links, time boundaries and punctuation parameters are requested in the first speech recognition, recording clips are cut by sentences, and the language discrimination engine and dedicated speech recognition engine are iteratively identified fragments that do not match the language, and the initial recognition results are corrected.

Benefits of technology

It realizes complete and correct recognition of mixed language recordings, solves the problem of identification errors, and improves the accuracy of recording quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431918A_ABST
    Figure CN120431918A_ABST
Patent Text Reader

Abstract

The invention provides a mixed language offline voice recognition method and system, and the method comprises the steps: carrying out the first overall judgment of an original recording through a language discrimination engine, and obtaining the approximate language of the whole call recording; performing first identification on the original record through a voice identification engine of a corresponding language, setting an identification request to open punctuation marks and time boundary parameters, and returning a first identification result; according to a time boundary and sentence punctuation marks in a first recognition result of a voice recognition engine, original recording is cut into recognized recording segments according to sentences, each recognized recording segment is sent to a language judgment engine for re-judgment, and a complete and correct recognition result is finally obtained according to a re-judged language. According to the method, the language discrimination engine and the special voice recognition engine of each language are iteratively called, and all the recording segments containing multiple languages in the recording are correctly recognized into texts, so that complete and correct recognition of the mixed language recording is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of offline speech quality inspection, and in particular to a mixed-language offline speech recognition method and system. Background Art

[0002] In the call center's call recording quality inspection project, a portion of the call recording is usually cut out and sent to the language discrimination engine to determine the language of the entire call recording. Then it is sent to the language recognition engine for recognition. This method can correctly recognize the text for most single-language voice recordings.

[0003] However, when a customer switches between different languages during a call, for example, speaking Cantonese at the beginning but later finding that their Cantonese skills are not good enough to communicate fluently with the agent and switching back to Mandarin or English, the voice recording saved at this time is a mixed multi-language recording.

[0004] Under the existing conventional method of single-time language determination and execution of a single speech recognition engine, for such mixed multilingual call recordings, there will be some speech recognition errors in the recognized text, making it unreadable and affecting subsequent recording quality inspection. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to propose a mixed-language offline speech recognition method and system. On the basis of the conventional single-time language judgment and docking with a single language recognition engine, a calibration link is added, and time boundary and punctuation parameter requests are added in the first speech recognition call. Then, according to the time boundary and sentence punctuation of the returned result, the original call recording is cut into segments by sentence, and the language of each segment is judged again. The recording segments that do not conform to the first language are sent to the correct speech recognition engine for re-recognition, and the correct recognition results are used to correct the text recognized for the first time. By iteratively calling the language discrimination engine and the dedicated speech recognition engines of each language, all the recording segments containing multiple languages in the recording are correctly recognized as text, thereby achieving complete and correct recognition of mixed-language recordings.

[0006] The present invention provides a mixed language offline speech recognition method, comprising the following steps:

[0007] S1. Perform a first overall judgment on the original recording using a language discrimination engine to obtain an approximate language of the entire call recording; perform a first recognition on the original recording using a speech recognition engine in the corresponding language, with the recognition request set to include punctuation and time boundary parameters, and return the first recognition result;

[0008] For example, the first recognition result returned is: "0.00-1.20 Hello, 1.30-1.50 I 1.50-1.63 Want to 1.70-1.90 Consultation 1.95-2.30 Credit card 2.30-2.60 How to apply 2.70-2.90";

[0009] S2. Based on the time boundaries and sentence punctuation marks in the first recognition result of the speech recognition engine, the original recording is divided into recognized recording segments by sentence, and each recognized recording segment is sent to the language discrimination engine for re-judgment, and a complete and correct recognition result is finally obtained according to the re-judged language.

[0010] Furthermore, the method of finally obtaining a complete and correct recognition result according to the re-determined language in step S2 includes:

[0011] If the language returned by a recognized audio segment does not match the language determined by the language discrimination engine for the first overall recognition, the recognized audio segment is sent to the corresponding speech recognition engine for re-recognition according to the re-determined language to obtain the correct recognition text, and the correct recognition text of each recognized audio segment is used to correct the first recognition result;

[0012] If the language returned by a recognized recording segment matches the language determined by the language identification engine for the first time, the first recognition result is a correct recognition result.

[0013] Furthermore, the method of performing a first overall judgment on the original recording by the language discrimination engine in step S1 includes:

[0014] A portion of the recording is cut from the original recording (for example, from the 5th second to the 20th second of the original recording), and the recording portion of the original recording is sent to the language discrimination engine for judgment.

[0015] Furthermore, the method of performing the first recognition of the original recording by a speech recognition engine of the corresponding language in step S1 includes:

[0016] The audio segment of the original audio recording whose approximate language is obtained after the first overall judgment is sent to the speech recognition engine of the corresponding language for recognition.

[0017] The present invention further provides a mixed-language offline speech recognition system, which executes the mixed-language offline speech recognition method described above, comprising:

[0018] First judgment and recognition module: used to perform a first overall judgment on the original recording using the language discrimination engine to obtain the approximate language of the entire call recording; perform a first recognition on the original recording using the speech recognition engine of the corresponding language, with the recognition request set to open punctuation and time boundary parameters, and return the first recognition result;

[0019] Re-judgment and recognition module: used to cut the original recording into recognized recording segments by sentence based on the time boundaries and sentence punctuation in the first recognition result of the speech recognition engine, and send each recognized recording segment to the language discrimination engine for re-judgment, and finally obtain a complete and correct recognition result according to the re-judged language.

[0020] Specifically, the language identification engine is used to automatically identify the language of the input text or voice data by analyzing specific language features, including vocabulary, grammatical structure, character set and pronunciation pattern, to determine which language the text or voice is in;

[0021] The speech recognition engine is used to convert human speech into machine-understandable text.

[0022] Furthermore, the first judgment and recognition module includes:

[0023] The unit for cutting the audio segment of the original recording is used to cut a portion of the audio segment from the original recording and send the audio segment of the original recording to the language discrimination engine for judgment;

[0024] The first recognition unit is used to send the recording segment of the original recording whose approximate language is obtained after the first overall judgment to the speech recognition engine of the corresponding language for recognition.

[0025] After adopting the technical solution of the present invention, the actual test results of the call center's recording quality inspection project showed that mixed-language recordings all obtained correct recognition texts, and the problem of recognition errors caused by language switching was effectively solved.

[0026] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps of the mixed-language offline speech recognition method as described above are implemented.

[0027] The present invention also provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the mixed-language offline speech recognition method described above are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The mixed-language offline speech recognition method and system provided by the present invention adds a calibration link on the basis of conventional single-time language determination and docking with a single language recognition engine, adds a time boundary and punctuation parameter request in the first speech recognition call, and then cuts the original call recording into segments by sentence based on the time boundary and sentence punctuation of the returned result. The language of each segment is determined again, and the recording segments that do not conform to the first language are sent back to the correct speech recognition engine for re-recognition. The correct recognition results are used to correct the text recognized for the first time. By iteratively calling the language determination engine and the dedicated speech recognition engines for each language, all recording segments containing multiple languages in the recording are correctly recognized as text, effectively achieving complete and correct recognition of mixed-language recordings. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.

[0031] In the attached figure:

[0032] Figure 1 This is an interface diagram of an embodiment of the present invention for cutting an original recording to perform a first overall judgment and determine the approximate language;

[0033] Figure 2 This is an interface diagram of an embodiment of the present invention for re-judging each recognized recording segment and performing language discrimination;

[0034] Figure 3 This is an interface diagram of an embodiment of the present invention for sending a segment whose language was incorrectly determined in the first overall judgment to the correct speech recognition engine for recognition again;

[0035] Figure 4 This is a flow chart of a mixed language offline speech recognition method of the present invention;

[0036] Figure 5 Schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of devices and products consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0038] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0039] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0040] The embodiments of the present invention are described in further detail below.

[0041] The embodiment of the present invention provides a mixed language offline speech recognition method, see Figure 4 As shown, the following steps are included:

[0042] S1. Perform the first overall judgment on the original recording through the language identification engine to obtain the approximate language of the entire call recording (such as Figure 1 As shown); the original recording is recognized for the first time by a speech recognition engine of the corresponding language, the recognition request is set to open punctuation and time boundary parameters, and the first recognition result is returned;

[0043] Methods for making a first overall judgment on the original recording using the language identification engine include:

[0044] A portion of the recording is cut from the original recording. In this embodiment, the recording segment cut is from the 5th second to the 20th second of the front of the original recording. The recording segment of the original recording is sent to the language discrimination engine for judgment.

[0045] The method of performing a first recognition on the original recording by a speech recognition engine of a corresponding language includes:

[0046] The audio segment of the original audio recording whose approximate language is obtained after the first overall judgment is sent to the speech recognition engine of the corresponding language for recognition.

[0047] In this embodiment, the first recognition result returned is as follows:

[0048] "0.00-1.20 Hello, 1.30-1.50 I 1.50-1.63 want 1.70-1.90 consultation 1.95-2.30 credit card 2.30-2.60 how to apply 2.70-2.90";

[0049] S2. Cut the original recording into recognized recording segments by sentence according to the time boundaries and sentence punctuation marks in the first recognition result of the speech recognition engine, and send each recognized recording segment to the language discrimination engine for re-judgment (e.g., Figure 2 As shown in the figure), a complete and correct recognition result is finally obtained according to the re-judged language.

[0050] Methods for obtaining complete and correct recognition results based on the re-determined language include:

[0051] If the language returned by a recognized audio segment does not match the language determined by the language discrimination engine for the first time, the recognized audio segment will be sent to the corresponding speech recognition engine for re-recognition according to the re-determined language (e.g. Figure 3 As shown), obtain the correct recognition text, and then use the correct recognition text of each recognized recording segment to correct the first recognition result;

[0052] If the language returned by a recognized recording segment matches the language determined by the language identification engine for the first time, the first recognition result is a correct recognition result.

[0053] An embodiment of the present invention further provides a mixed-language offline speech recognition system, which executes the mixed-language offline speech recognition method described above, including:

[0054] First judgment and recognition module: used to perform a first overall judgment on the original recording using the language discrimination engine to obtain the approximate language of the entire call recording; perform a first recognition on the original recording using the speech recognition engine of the corresponding language, with the recognition request set to open punctuation and time boundary parameters, and return the first recognition result;

[0055] The first judgment and recognition module includes:

[0056] The unit for cutting the audio segment of the original recording is used to cut a portion of the audio segment from the original recording and send the audio segment of the original recording to the language discrimination engine for judgment;

[0057] The first recognition unit is used to send the recording segment of the original recording whose approximate language is obtained after the first overall judgment to the speech recognition engine of the corresponding language for recognition.

[0058] Re-judgment and recognition module: used to cut the original recording into recognized recording segments by sentence based on the time boundaries and sentence punctuation in the first recognition result of the speech recognition engine, and send each recognized recording segment to the language discrimination engine for re-judgment, and finally obtain a complete and correct recognition result according to the re-judged language.

[0059] The language identification engine is used to automatically identify the language of the input text or voice data by analyzing specific language features, including vocabulary, grammatical structure, character set and pronunciation patterns, to determine which language the text or voice is in;

[0060] The speech recognition engine is used to convert human speech into machine-understandable text.

[0061] The mixed-language offline speech recognition method and system of this embodiment adds a calibration step based on the conventional single-time language determination and docking with a single language recognition engine. It adds time boundary and punctuation parameter requests in the first speech recognition call, and then cuts the original call recording into segments by sentence based on the time boundary and sentence punctuation of the returned result. The language of each segment is determined again, and the recording segments that do not match the first language are sent back to the correct speech recognition engine for re-recognition. The correct recognition results are used to correct the text recognized for the first time. By iteratively calling the language determination engine and the dedicated speech recognition engines for each language, all recording segments containing multiple languages in the recording are correctly recognized as text, thereby achieving complete and correct recognition of mixed-language recordings.

[0062] An embodiment of the present invention further provides a computer device, Figure 5 This is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention; see the accompanying drawings Figure 5 As shown, the computer device includes: an input system 23, an output system 24, a memory 22 and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the mixed language offline speech recognition method provided in the above embodiment; wherein the input system 23, the output system 24, the memory 22 and the processor 21 can be connected by a bus or other means. Figure 5 The bus connection is taken as an example.

[0063] The memory 22 is a readable and writable storage medium of a computing device and can be used to store software programs and computer executable programs, such as the program instructions corresponding to the mixed-language offline speech recognition method described in the embodiment of the present invention. The memory 22 may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created based on the use of the device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 22 may further include a memory remotely located relative to the processor 21, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0064] The input system 23 may be used to receive input digital or character information, and generate key signal input related to user settings and function control of the device; the output system 24 may include display devices such as a display screen.

[0065] The processor 21 executes the software programs, instructions, and modules stored in the memory 22 to perform various functional applications and data processing of the device, that is, to implement the above-mentioned mixed-language offline speech recognition method.

[0066] The computer device provided above can be used to execute the mixed language offline speech recognition method provided in the above embodiment, and has corresponding functions and beneficial effects.

[0067] Embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the mixed-language offline speech recognition method provided in the above embodiments. The storage medium is any of various types of memory devices or storage devices, including: installation media, such as CD-ROMs, floppy disks, or tape systems; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements; the storage medium may also include other types of memory or a combination thereof; in addition, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system, which is connected to the first computer system via a network (such as the Internet); the second computer system may provide program instructions to the first computer for execution. The storage medium includes two or more storage media that may reside in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions (e.g., embodied as a computer program) that can be executed by one or more processors.

[0068] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present invention is not limited to the offline speech recognition method for mixed languages described in the above embodiment, and can also execute related operations in the offline speech recognition method for mixed languages provided in any embodiment of the present invention.

[0069] Thus far, the technical solutions of the present invention have been described in conjunction with preferred embodiments. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is clearly not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0070] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A mixed language offline speech recognition method, characterized in that: The following steps are involved: S1. Perform a first overall judgment on the original recording using a language discrimination engine to obtain an approximate language of the entire call recording; perform a first recognition on the original recording using a speech recognition engine in the corresponding language, with the recognition request set to include punctuation and time boundary parameters, and return the first recognition result; S2. Based on the time boundaries and sentence punctuation marks in the first recognition result of the speech recognition engine, the original recording is divided into recognized recording segments by sentence, and each recognized recording segment is sent to the language discrimination engine for re-judgment, and a complete and correct recognition result is finally obtained according to the re-judged language.

2. The mixed language offline speech recognition method according to claim 1, characterized in that: The method for finally obtaining a complete and correct recognition result according to the re-determined language in step S2 includes: If the language returned by a recognized audio segment does not match the language determined by the language discrimination engine for the first overall recognition, the recognized audio segment is sent to the corresponding speech recognition engine for re-recognition according to the re-determined language to obtain the correct recognition text, and the correct recognition text of each recognized audio segment is used to correct the first recognition result; If the language returned by a recognized recording segment matches the language determined by the language identification engine for the first time, the first recognition result is a correct recognition result.

3. The mixed language offline speech recognition method according to claim 2, characterized in that: The method of performing a first overall judgment on the original recording by the language discrimination engine in step S1 includes: A portion of the recording is cut from the original recording, and the recording portion of the original recording is sent to the language discrimination engine for judgment.

4. The mixed language offline speech recognition method according to claim 3, characterized in that: The method of performing a first recognition of the original recording by a speech recognition engine of a corresponding language in step S1 includes: The audio segment of the original audio recording whose approximate language is obtained after the first overall judgment is sent to the speech recognition engine of the corresponding language for recognition.

5. A mixed language offline speech recognition system, characterized by: Executing the mixed-language offline speech recognition method according to any one of claims 1 to 4 comprises: First judgment and recognition module: used to perform a first overall judgment on the original recording using the language discrimination engine to obtain the approximate language of the entire call recording; perform a first recognition on the original recording using the speech recognition engine of the corresponding language, with the recognition request set to open punctuation and time boundary parameters, and return the first recognition result; Re-judgment and recognition module: used to cut the original recording into recognized recording segments by sentence based on the time boundaries and sentence punctuation in the first recognition result of the speech recognition engine, and send each recognized recording segment to the language discrimination engine for re-judgment, and finally obtain a complete and correct recognition result according to the re-judged language.

6. The mixed language offline speech recognition system according to claim 5, characterized in that: The first judgment and recognition module includes: The unit for cutting the audio segment of the original recording is used to cut a portion of the audio segment from the original recording and send the audio segment of the original recording to the language discrimination engine for judgment; The first recognition unit is used to send the recording segment of the original recording whose approximate language is obtained after the first overall judgment to the speech recognition engine of the corresponding language for recognition.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the mixed-language offline speech recognition method according to any one of claims 1 to 4 are implemented.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the mixed-language offline speech recognition method according to any one of claims 1 to 4 are implemented.