Speech code automatic identification method and device based on natural language processing, and medium
By performing semantic and grammatical analysis on the voice data stream and combining it with a priority list of historical communication records, the system automatically identifies voice codes and languages, thus resolving the problem of call aberrations caused by code mismatches in voice communications, achieving fast and convenient voice code matching, and improving call quality.
Patent Information
- Application Number
- CN202511049597.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In the prior art, during voice communication, call failures due to a lack of handshake or the other party changing the voice codec require a tedious rematching process, causing inconvenience to the user.
Through the automatic voice coding recognition method based on natural language processing, the semantic analysis and grammatical analysis of the voice data stream are used to automatically identify the other party's voice coding and language. Combined with the priority list of historical communication records, the correct voice coding format is quickly matched.
It simplifies the voice codec matching process, improves call fluency, reduces connection interruptions due to handshake failures, and reduces communication establishment time.
Smart Images

Figure CN120692256A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice communication technology, and in particular to a method, device and medium for automatic voice coding recognition based on natural language processing. Background Art
[0002] With the development of communication technology, people's communication means are becoming increasingly diverse, but voice communication, as a guaranteed communication means, has the characteristics of short delay and relatively small data volume, and is still people's most basic communication need.
[0003] To meet the needs of different communication systems, people have invented many different voice coding methods, including but not limited to G.729 and G.711. Before a call, the caller and the called party will handshake. The caller sends the called party all the audio coding formats it supports. The called party selects one or more of the coding formats it supports when the call is ringing back.
[0004] However, if the caller fails to perform a handshake, or if the caller changes the voice codec that was previously matched without the called party's knowledge, communication can become disrupted. In this case, the user must discover the problem, perform a handshake again, match the voice codec, and then resume the call. This is a cumbersome process and causes significant inconvenience to both parties.
[0005] In view of this, there is an urgent need to provide a simpler and more convenient way to make calls, which can simplify the process of matching voice codes and make calls smoother. Summary of the Invention
[0006] The purpose of this application is to provide a simpler and more convenient way to make calls, which can simplify the process of matching voice codes and make calls smoother.
[0007] In the first aspect, the present application provides a method for automatic speech coding recognition based on processing natural language, which adopts the following technical solutions: A method for automatic speech coding recognition based on processing natural language comprises the following steps: S1, select speech coding; S2, according to the selected voice coding format, the voice information is restored to a voice data stream of standard voice coding; S3. Obtaining whether the voice data stream meets the required language; If the voice data streams in all languages do not meet the requirements under the current voice coding format, repeat steps S1-S3; If the voice data stream in any language meets the requirements under the current voice coding format, proceed to the next step; S4. Obtain the current voice coding format, and output the voice information as a voice data stream in the current voice coding format.
[0008] Furthermore, the S3 includes: S31. Select a language from the language library; S32. Performing semantic analysis and grammatical analysis on the voice data stream based on processing natural language data according to the selected language; S33, detecting whether the voice continuity of the voice data stream reaches a first threshold in the current language; If the voice continuity of the voice data stream reaches the threshold, the currently selected language is locked; If the voice continuity of the voice data stream does not reach the threshold, S31 to S33 are repeated until all languages in the language library are selected.
[0009] Furthermore, the S32 includes: Identify whether the voice data stream contains time, place, and personal pronouns; Recognize verbs in speech data streams; Analyze the word order and grammatical structure of the voice data stream.
[0010] Furthermore, the threshold is set to 85-100%.
[0011] Furthermore, the language library includes Chinese, Japanese and English.
[0012] Furthermore, in S1, the voice codec is selected in descending order of priority according to the voice codec priority list in the historical communication record.
[0013] Furthermore, the voice coding priority list includes G.711, G.729, G.726, G.722, and G.723.1.
[0014] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned method for automatic speech coding recognition is implemented.
[0015] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned method for automatic speech coding recognition is implemented.
[0016] In summary, this application includes at least one of the following beneficial technical effects: 1. This application analyzes the voice coding and language during a call to determine the other party's voice coding and language. This effectively solves the problem of manually adapting to different voice coding formats when accessing different voice systems. This method makes calls simpler and more convenient.
[0017] 2. The voice coding in this application is selected in descending order according to the voice coding priority list in the historical communication records, so that when the other party changes the voice coding, the correct voice coding format can be quickly matched to ensure that the call can be restored to normal quickly.
[0018] 3. This application implements "blind matching" through localized NLP analysis. Compared with the SIP protocol which requires multiple signaling interactions, it reduces the time required to establish communication and can reduce the occurrence of connection interruptions due to handshake failures. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of the automatic speech coding recognition method in this application; Figure 2 yes Figure 1 Schematic diagram of the S3 process. DETAILED DESCRIPTION
[0020] The following will be combined with the Figure 1-2 The technical solution of the present application is described clearly and completely. The following embodiments are exemplary and are only used to explain the present application, and should not be construed as limiting the present application. In the following description, the same reference numerals are used to represent the same or equivalent elements, and repeated descriptions are omitted.
[0021] In the description of this application, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the products of this application are conventionally placed when in use, or are the orientations or positional relationships conventionally understood by those skilled in the art. These are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they should not be understood as limitations on this application.
[0022] In addition, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this application based on the specific circumstances.
[0023] It should be further understood that the term “and / or” used in this specification and the corresponding claims refers to any and all possible combinations of one or more of the listed items.
[0024] In the first aspect, the present application provides a method for automatic speech coding recognition based on processing natural language, which adopts the following technical solutions: A method for automatic speech coding recognition based on processing natural language, referring to Figure 1 , including the following steps: S1, select speech coding; The selection order of the voice codes is specifically set according to the actual situation. In this embodiment, the voice codes are selected in descending order of priority based on the voice code priority list in the historical communication records.
[0025] The voice codec priority list records the number of times each voice codec was used during past calls with the other party and sorts the different voice codecs by the number of times they were used. During a call, the voice codecs with the highest number of uses are prioritized for faster matching.
[0026] It should be noted that, in other embodiments, the speech code may be selected by random sampling from a speech code library.
[0027] Furthermore, referring to mainstream speech coding types, in this embodiment, the speech coding priority list includes G.711, G.729, G.726, G.722, and G.723.1.
[0028] Based on the voice codec usage during a call with a particular party, the five voice codecs listed above are prioritized as G.729 > G.711 > G.726 > G.722 > G.723.1. Under these conditions, G.729 is prioritized for voice information processing. It should be noted that the order of the various voice codecs in the voice codec priority list may be adjusted accordingly for different call parties.
[0029] S2, according to the selected voice coding format, the voice information is restored to a voice data stream of standard voice coding; S3. Obtaining whether the voice data stream meets the required language; Reference Figure 2 In this embodiment, S3 specifically includes the following steps: S31. Select a language from the language library.
[0030] In this embodiment, the language library includes Chinese, Japanese and English. In other embodiments, the language library can also be further increased according to usage, for example: Korean, German, Russian, etc.
[0031] S32. According to the selected language, the speech data stream is subjected to semantic analysis and grammatical analysis based on processing natural language data.
[0032] When the language is determined, the voice data stream is analyzed to detect whether the voice data stream actually matches the selected language, so as to determine whether the selected language meets the requirements.
[0033] In this embodiment, S32 specifically includes the following steps: S321. Identify whether the voice data stream contains time, place, and personal pronouns.
[0034] S322: Identify verbs in the voice data stream.
[0035] S323: Analyze the word order and grammatical structure of the voice data stream.
[0036] Specifically, the voice data stream is analyzed to see if it contains text information corresponding to the selected language, namely the time, place, personal pronouns, and verbs mentioned above. Based on this, further analysis is performed to see if the grammatical structure is the same as that of the selected language, thereby determining whether the selected language is correct under the current voice encoding.
[0037] S33, detecting whether the voice continuity of the voice data stream reaches a threshold in the current language; The threshold is set to detect whether the current language selection is appropriate. If the voice continuity of the voice data stream under the current language condition does not reach the threshold, it is determined that the current language does not meet the requirements.
[0038] In this embodiment, the threshold is set to 85-100%.
[0039] In combination with the specific steps of S32 above, the above steps are explained by taking Chinese, Japanese, and English as the languages of the language library and taking Chinese as the selected language as an example.
[0040] The specific value of the threshold is determined three times, and the thresholds specifically include a first threshold, a second threshold, and a third threshold. Based on step S321, the first threshold is determined; based on step S322, the second threshold is determined; and based on step S323, the third threshold is determined.
[0041] First, after translating the voice data stream into Chinese, the system determines whether the data stream contains relevant information under Chinese conditions. If so, the system outputs a first threshold. Next, the system identifies verbs in the data stream, determines coefficient a, and outputs a second threshold: coefficient a × first threshold. Furthermore, the system analyzes the word order and grammatical structure of the data stream, determines coefficient b, and outputs a third threshold: coefficient b × second threshold.
[0042] It should be noted that if the first threshold does not meet the threshold range requirement, it indicates that the current language no longer meets the requirements, and subsequent verb analysis, word order, and grammatical structure analysis will be stopped. If the first threshold meets the threshold range requirement but the second threshold is not met, subsequent word order and grammatical structure analysis will be stopped.
[0043] Furthermore, time, place, personal pronouns, and verbs provide the basis for word order and grammatical structure analysis, which are crucial factors in determining whether a language meets the requirements. Specifically, coefficient b is b1*b2, where coefficients b, b1, and b2 are all 0 or 1.
[0044] The word order in Chinese and English is typically subject + predicate + object, while Japanese is typically subject + object + predicate. If the current word order is subject + predicate + object, the current language is determined to be Chinese or English, thus meeting the Chinese language requirements, and b1 = 1. If the output meets the Chinese language requirements and the grammatical structure is further determined to be subject + copula + predicate or subject + predicate + introductory word + clause, the current language is determined to be English. Otherwise, it is determined to be Chinese. If the current grammatical structure is subject + copula + predicate, it does not meet the Chinese language requirements, and b2 = 0. Therefore, if b = 0, the output voice data stream does not meet the requirements.
[0045] When the voice continuity of the voice data stream reaches a threshold, the currently selected language is locked. For example, if the selected language is English and the voice data stream's voice continuity reaches the threshold, the voice codec and language are confirmed. The voice information, after the confirmed voice codec, is restored to the voice data stream and output in the confirmed language.
[0046] If the voice continuity of the voice data stream does not reach the threshold, S31 to S33 are repeated until all languages in the language library are selected.
[0047] That is, another language is selected and the voice data stream is reanalyzed until all languages are selected. At this point, it can be determined that the voice codec selection was incorrect. It should be noted that in another embodiment, multiple languages in the language library can be analyzed simultaneously to increase the speed of voice codec and language confirmation.
[0048] If the voice data streams in all languages under the current voice coding format do not meet the requirements, repeat steps S1-S3. That is, reselect another voice coding format according to the voice coding priority list and further analyze the various voices in the language library.
[0049] If the voice data stream in any language meets the requirements under the current voice coding format, proceed to the next step.
[0050] S4. Obtain the current voice coding format, and output the voice information as a voice data stream in the current voice coding format.
[0051] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned method for automatic speech coding recognition is implemented.
[0052] The memory may be used to store instructions, programs, codes, code sets, or instruction sets.
[0053] In one embodiment, the memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing the above-mentioned automatic speech coding recognition method, and the data storage area may store data involved in the above-mentioned automatic speech coding recognition method, such as: a speech coding priority list, a language library, etc.
[0054] The processor may include one or more processing cores. The processor executes or processes instructions, programs, code sets, or instruction sets stored in the memory, calls data stored in the memory, and performs various functions of the present application and processes data. The processor may be at least one of an application-specific integrated circuit, a digital signal processor, a digital signal processing device, a programmable logic device, a field programmable gate array, a central processing unit, a controller, a microcontroller, and a microprocessor. It is understood that for different devices, the electronic components used to implement the above-mentioned processor functions may also be other, and the embodiments of the present application are not specifically limited thereto. In addition, the present application also provides a computer-readable storage medium, which stores a computer program that can be loaded by a processor and execute the above-mentioned automatic speech coding recognition method.
[0055] The computer-readable storage medium is specifically configured according to actual needs, and includes, for example, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0056] The examples of this specific embodiment are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Identical components are represented by the same reference numerals. Therefore, any equivalent changes made based on the structure, shape, and principle of this application should be included in the scope of protection of this application.
Claims
1. A method for automatic speech coding recognition based on processing natural language, characterized in that: The following steps are involved: S1, select speech coding; S2, according to the selected voice coding format, the voice information is restored to a voice data stream of standard voice coding; S3. Obtaining whether the voice data stream meets the required language; If the voice data streams in all languages do not meet the requirements under the current voice coding format, repeat steps S1-S3; If the voice data stream in any language meets the requirements under the current voice coding format, proceed to the next step; S4. Obtain the current voice coding format, and output the voice information as a voice data stream in the current voice coding format.
2. The method for automatic speech coding recognition based on processing natural language according to claim 1, characterized in that: The S3 includes: S31. Select a language from the language library; S32. Performing semantic analysis and grammatical analysis on the voice data stream based on processing natural language data according to the selected language; S33, detecting whether the voice continuity of the voice data stream reaches a threshold in the current language; If the voice continuity of the voice data stream reaches the threshold, the currently selected language is locked; If the voice continuity of the voice data stream does not reach the threshold, S31 to S33 are repeated until all languages in the language library are selected.
3. The method for automatic speech coding recognition based on processing natural language according to claim 2, characterized in that: The S32 includes: Identify whether the voice data stream contains time, place, and personal pronouns; Recognize verbs in speech data streams; Analyze the word order and grammatical structure of the voice data stream.
4. The method for automatic speech coding recognition based on processing natural language according to claim 2, characterized in that: The threshold is set to 85-100%.
5. The method for automatic speech coding recognition based on processing natural language according to claim 2, characterized in that: The language library includes Chinese, Japanese and English.
6. The method for automatic speech coding recognition based on processing natural language according to claim 1, characterized in that: In S1, the voice codec is selected in descending order of priority according to the voice codec priority list in the historical communication record.
7. The method for automatic speech coding recognition based on processing natural language according to claim 6, characterized in that: The voice coding priority list includes G.711, G.729, G.726, G.722, and G.723.
1.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method for automatic speech coding recognition according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method for automatic speech coding recognition according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice code conversion method and device
CN102377755A
Chinese and English hybrid speech recognition model training method and device
CN111816169A
Subtitle correction method and display equipment
CN112580302A
Speech recognition method and device and computer readable storage medium
CN114283786A
Encoding device
US3953846A