Automatic speech recognition methods, devices, and media based on natural language processing and speech coding.

By using a speech code automatic recognition method based on natural language processing, semantic and syntactic analysis, combined with historical communication records, the method can automatically identify the other party's speech code and language, solving the problem of cumbersome speech code matching and improving the convenience and stability of calls.

CN120692256BActive Publication Date: 2025-10-31WUHAN HUABO COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511049597.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-31
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

In existing technologies, the voice code matching process is cumbersome, leading to call interruptions, especially when the voice code is changed without the caller's handshake or the called party's knowledge, making the call process inconvenient.

Method used

By using an automatic speech coding recognition method based on natural language processing, semantic and syntactic analysis is employed to automatically identify the other party's speech coding and language. Combined with a speech coding priority list from historical communication records, the correct speech coding format is quickly matched.

Benefits of technology

It simplifies the voice code matching process, improves the convenience and stability of calls, reduces connection interruptions caused by handshake failures, and improves call recovery speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692256B_ABST
    Figure CN120692256B_ABST
Patent Text Reader

Abstract

This application relates to an automatic speech code recognition method, device, and medium based on natural language processing, belonging to the field of voice communication technology. The automatic speech code recognition method based on natural language processing includes the following steps: S1, selecting a speech code; S2, restoring the speech information to a standard speech code speech data stream according to the selected speech code format; S3, obtaining the language matching information of the speech data stream; if the speech data stream does not meet the requirements for all languages ​​under the current speech code format, repeating steps S1-S3; if the speech data stream meets the requirements for any language under the current speech code format, proceeding to the next step; S4, obtaining the current speech code format, and outputting the speech information as a speech data stream in the current speech code format. This application can quickly identify speech codes and ensure smooth communication during communication failures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice communication technology, and in particular to an automatic speech recognition method, device and medium based on natural language processing. Background Technology

[0002] With the development of communication technology, people have increasingly rich means of communication, but voice communication, as a backup communication method, has the characteristics of short latency and relatively small data volume, and is still people's most basic communication need.

[0003] To meet the needs of different communication systems, people have invented many different voice coding methods, including but not limited to G.729 and G.711. Before a call, the caller and the called party will shake hands. The caller sends all the audio coding formats it supports to the called party. When the called party rings back, it selects one or more of the coding formats it supports.

[0004] However, communication breakdowns can occur when the calling party fails to initiate a handshake or changes the voice code that was previously matched without the called party's knowledge. In such cases, a new handshake must be initiated after the user discovers the problem, and the voice code format must be matched again before the call can resume. This process is cumbersome and causes significant inconvenience to both parties.

[0005] In view of this, there is an urgent need to provide a simpler and more convenient way to make calls, which can simplify the process of matching voice codes and make calls smoother. Summary of the Invention

[0006] The purpose of this application is to provide a simpler and more convenient way to make calls, which simplifies the process of matching voice codes and makes calls smoother.

[0007] Firstly, the automatic speech recognition method based on natural language processing provided in this application adopts the following technical solution:

[0008] An automatic speech recognition method based on natural language processing coding includes the following steps:

[0009] S1. Select voice encoding;

[0010] S2. Based on the selected speech coding format, restore the speech information to a standard speech coding speech data stream;

[0011] S3. Obtain the language information of the voice data stream that meets the requirements;

[0012] If the speech data stream does not meet the requirements for all languages ​​under the current speech encoding format, repeat steps S1-S3.

[0013] If the speech data stream in any language under the current speech encoding format meets the requirements, proceed to the next step;

[0014] S4. Obtain the current speech encoding format, and output the speech data stream in the current speech encoding format.

[0015] Further, S3 includes:

[0016] S31. Select a language from the language database;

[0017] S32. Based on the selected language, perform semantic and syntactic analysis on the speech data stream based on the processing of natural language data;

[0018] S33. Detect whether the speech continuity of the speech data stream reaches the first threshold under the current language condition;

[0019] If the speech continuity of the speech data stream reaches the threshold, lock the currently selected language;

[0020] If the speech continuity of the speech data stream does not reach the threshold, repeat steps S31-S33 until all languages ​​in the language library are selected.

[0021] Further, S32 includes:

[0022] Identify whether the speech data stream contains time, location, and personal pronouns;

[0023] Recognize verbs in a speech data stream;

[0024] Analyze the word order and grammatical structure of the speech data stream.

[0025] Furthermore, the threshold is set to 85-100%.

[0026] Furthermore, the language library includes Chinese, Japanese, and English.

[0027] Furthermore, in S1, the voice encoding is selected sequentially from high to low priority according to the voice encoding priority list in the historical communication records.

[0028] Furthermore, the speech coding priority list includes G.711, G.729, G.726, G.722, and G.723.1.

[0029] Secondly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned automatic speech code recognition method.

[0030] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the above-described automatic speech code recognition method.

[0031] In summary, this application includes at least one of the following beneficial technical effects:

[0032] 1. This application analyzes the voice encoding and language during a call to determine the other party's voice encoding and language. This effectively solves the problem of manually adapting different voice encoding formats when accessing different voice systems, making calls simpler and more convenient.

[0033] 2. The voice encoding in this application is selected in descending order of priority from the voice encoding priority list in historical communication records, so as to quickly match the correct voice encoding format when the other party changes the voice encoding, and ensure that the call can be quickly restored to normal.

[0034] 3. This application achieves "blind matching" through localized NLP analysis, which reduces the time required for communication establishment compared to the SIP protocol, and can reduce the occurrence of connection interruption due to handshake failure. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the automatic speech coding recognition method in this application;

[0036] Figure 2 yes Figure 1 A flowchart of the S3 process. Detailed Implementation

[0037] The following will be combined with the appendix Figures 1-2 The technical solution of this application is clearly and completely described. The following embodiments are exemplary and are only used to explain this application, and should not be construed as limiting this application. In the following description, the same reference numerals are used to denote the same or equivalent elements, and repeated descriptions are omitted.

[0038] In the description of this application, it should be understood that the terms "upper", "lower", "inner", "outer", "left", "right", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this application is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used to facilitate the description of this application and to simplify the description, and are not intended to indicate or imply that the equipment or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0039] Furthermore, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0040] It should also be further understood that the term "and / or" as used in this application specification and the corresponding claims refers to any combination of one or more of the listed items, as well as all possible combinations.

[0041] Firstly, the automatic speech recognition method based on natural language processing provided in this application adopts the following technical solution:

[0042] An automatic speech recognition method based on natural language processing coding, referring to Figure 1 This includes the following steps:

[0043] S1. Select voice encoding;

[0044] The selection order of voice codes is set according to the actual situation. In this embodiment, voice codes are selected in descending order of priority based on the voice code priority list in historical communication records.

[0045] The voice code priority list records the frequency of use of various voice codes during past calls and sorts them according to usage frequency. During a call, the most frequently used voice codes are prioritized for faster matching.

[0046] It should be noted that in other embodiments, the speech code can also be selected by random sampling from a speech code library.

[0047] Furthermore, referring to the mainstream speech coding types, in this embodiment, the speech coding priority list includes G.711, G.729, G.726, G.722, and G.723.1.

[0048] Based on the voice coding used during a call with a particular party, the five voice codings listed above are ranked by priority as follows: G.729 > G.711 > G.726 > G.722 > G.723.1. Under this condition, voice coding G.729 is selected first for processing voice information. It should be noted that the order of the various voice codings obtained from the voice coding priority list may be adjusted accordingly for different parties in the call.

[0049] S2. Based on the selected speech coding format, restore the speech information to a standard speech coding speech data stream;

[0050] S3. Obtain the language information of the voice data stream that meets the requirements;

[0051] Reference Figure 2 In this embodiment, S3 specifically includes the following steps:

[0052] S31. Select a language from the language database.

[0053] In this embodiment, the language library includes Chinese, Japanese, and English. In other embodiments, the language library may be further expanded according to usage, for example, to include Korean, German, Russian, etc.

[0054] S32. Based on the selected language, perform semantic and syntactic analysis on the speech data stream based on the processing of natural language data.

[0055] Once the language is determined, the voice data stream is analyzed to check whether the voice data stream actually matches the selected language, in order to determine whether the selected language meets the requirements.

[0056] In this embodiment, S32 specifically includes the following steps:

[0057] S321. Identify whether the speech data stream contains time, location, or personal pronouns.

[0058] S322, Identify verbs in the speech data stream.

[0059] S323. Analyze the word order and grammatical structure of the speech data stream.

[0060] This involves analyzing whether the speech data stream contains textual information corresponding to the selected language, namely, the aforementioned time, location, personal pronouns, and verbs. Based on this, it further analyzes whether the grammatical structure is the same as the selected language to determine whether the selected language is correctly identified under the current speech encoding.

[0061] S33. Detect whether the speech continuity of the speech data stream reaches the threshold under the current language condition;

[0062] The threshold setting is used to detect whether the current language selection is appropriate. If the speech continuity of the speech data stream does not reach the threshold under the current language condition, the current language is determined to be unacceptable.

[0063] In this embodiment, the threshold is set to 85-100%.

[0064] Based on the specific steps of S32 above, using Chinese, Japanese, and English as the languages ​​in the language database, and taking Chinese as the selected language as an example, the above steps will be explained.

[0065] The specific value of the threshold is determined three times, including a first threshold, a second threshold, and a third threshold. Based on step S321, the first threshold is determined; based on step S322, the second threshold is determined; and based on step S323, the third threshold is determined.

[0066] First, after translating the speech data stream into Chinese, it is determined whether the speech data stream contains relevant information under Chinese conditions. If time, location, and personal pronouns are all present, a first threshold is output. Next, verbs in the speech data stream are identified, coefficient 'a' is determined, and a second threshold is output, where the second threshold equals coefficient 'a' multiplied by the first threshold. Based on this, the word order and grammatical structure of the speech data stream are further analyzed to determine coefficient 'b', and a third threshold is output, where the third threshold equals coefficient 'b' multiplied by the second threshold.

[0067] It should be noted that when the first threshold is not met, it indicates that the current language does not meet the requirements, and subsequent verb analysis, word order analysis, and grammatical structure analysis will stop. When the first threshold is met but the second threshold is not met, subsequent word order and grammatical structure analysis will stop.

[0068] Furthermore, time, place, personal pronouns, and verbs provide the foundation for word order and grammatical structure analysis, which in turn is a decisive factor in determining whether a language meets the requirements. Specifically, the coefficient b is b1*b2, and the coefficients b, b1, and b2 are all 0 or 1.

[0069] The word order in Chinese and English is typically subject + verb + object, while in Japanese it is typically subject + object + verb. If the current word order is actually subject + verb + object, the current language is determined to be either Chinese or English, which meets the Chinese requirement, so b1=1. If, based on the output meeting the Chinese requirement, a grammatical structure of subject + copula + predicate or subject + verb + conjunction + clause is found, the current language is determined to be English; otherwise, it is determined to be Chinese. If the current grammatical structure is subject + copula + predicate, it does not meet the Chinese requirement, so b2=0. Therefore, b=0, and the output speech data stream does not meet the requirements.

[0070] When the speech continuity of the speech data stream reaches a threshold, the currently selected language is locked. For example, when the selected language is English and the speech continuity of the speech data stream reaches the threshold, the speech encoding and language are confirmed. The speech information is then restored from the determined speech encoding to the speech data stream and output in the confirmed language.

[0071] If the speech continuity of the speech data stream does not reach the threshold, repeat steps S31-S33 until all languages ​​in the language library are selected.

[0072] This involves selecting other languages ​​and re-analyzing the speech data stream until all languages ​​have been selected. At this point, it can be determined that the speech encoding was selected incorrectly. It should be noted that in another embodiment, multiple languages ​​in the language library can be analyzed simultaneously to improve the speed of speech encoding and language identification.

[0073] If the speech data streams for all languages ​​do not meet the requirements under the current speech encoding format, repeat steps S1-S3. That is, according to the speech encoding priority list, select other speech encodings and further analyze the various speech languages ​​in the language library.

[0074] If the speech data stream in any language under the current speech encoding format meets the requirements, proceed to the next step.

[0075] S4. Obtain the current speech encoding format, and output the speech data stream in the current speech encoding format.

[0076] Secondly, this application provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, it implements the above-mentioned automatic speech code recognition method.

[0077] Memory can be used to store instructions, programs, code, code sets, or instruction sets.

[0078] In one embodiment, the memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing the above-described automatic speech code recognition method, and the data storage area may store data involved in the above-described automatic speech code recognition method, such as a speech code priority list, a language library, etc.

[0079] The processor may include one or more processing cores. The processor executes or runs instructions, programs, code sets, or instruction sets stored in memory, and calls data stored in memory to perform various functions and process data as described in this application.

[0080] The processor can be at least one of the following: application-specific integrated circuit, digital signal processor, digital signal processing device, programmable logic device, field-programmable gate array, central processing unit, controller, microcontroller, and microprocessor. It is understood that, for different devices, the electronic device used to implement the above-described processor functions can also be other types, and the embodiments of this application do not specifically limit this.

[0081] In addition, this application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and execute the above-described automatic speech encoding recognition method.

[0082] Computer-readable storage media are configured according to actual needs, and may include, for example, USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0083] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Identical components are represented by the same reference numerals. Therefore, all equivalent changes made to the structure, shape, and principle of this application should be covered within the scope of protection of this application.

Claims

1. An automatic speech recognition method based on natural language processing coding, characterized in that, Includes the following steps: S1. Select a voice code. The voice code is selected in descending order of priority according to the voice code priority list in the historical communication records. S2. Based on the selected speech coding format, restore the speech information to a standard speech coding speech data stream; S3. Obtain the language information of the voice data stream that meets the requirements, specifically including: S31. Select a language from the language database; S32. Based on the selected language, perform semantic and syntactic analysis on the speech data stream based on the processing of natural language data; S33. Detect whether the speech continuity of the speech data stream reaches the threshold under the current language condition; If the speech continuity of the speech data stream reaches the threshold, lock the currently selected language; If the speech continuity of the speech data stream does not reach the threshold, repeat steps S31-S33 until all languages ​​in the language library have been selected. If the speech data stream does not meet the requirements for all languages ​​under the current speech encoding format, repeat steps S1-S3. If the speech data stream in any language under the current speech encoding format meets the requirements, proceed to the next step; S4. Obtain the current speech encoding format, and output the speech data stream in the current speech encoding format.

2. The automatic speech recognition method based on natural language processing coding according to claim 1, characterized in that, S32 includes: Identify whether the speech data stream contains time, location, and personal pronouns; Recognize verbs in a speech data stream; Analyze the word order and grammatical structure of the speech data stream.

3. The automatic speech recognition method based on natural language processing according to claim 1, characterized in that, The threshold is set to 85-100%.

4. The automatic speech recognition method based on natural language processing coding according to claim 1, characterized in that, The language library includes Chinese, Japanese, and English.

5. The automatic speech recognition method based on natural language processing coding according to claim 1, characterized in that, The speech coding priority list includes G.711, G.729, G.726, G.722, and G.723.

1.

6. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the speech coding automatic recognition method as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the speech encoding automatic recognition method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Voice code conversion method and device

    CN102377755A

  • Subtitle correction method and display equipment

    CN112580302A