Speech recognition method suitable for multiple languages and intelligent customer service robot

By detecting the dominant language at the voice transmitter and receiver and selecting an appropriate semantic conversion mode, the problem of semantic understanding bias in multilingual voice interaction is solved, and accurate semantic transmission is achieved in a multilingual environment.

CN121747576APending Publication Date: 2026-03-27SHENZHEN YUEGANG TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In multilingual voice interaction, differences in language and region can lead to a lack of shared semantic understanding among the interacting parties, creating communication barriers.

Method used

By acquiring the main languages ​​of the voice transmitter and receiver, performing conflict detection, and selecting an appropriate semantic conversion mode, the voice data is translated into an output message that the receiver understands, including direct semantic conversion, cross-language and cross-regional conversion.

Benefits of technology

It effectively avoids semantic deviations caused by direct translation, achieves accurate semantic transmission in a multilingual environment, and reduces interaction barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747576A_ABST
    Figure CN121747576A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method suitable for multiple languages and an intelligent customer service robot, and relates to the technical field of language processing, and the voice recognition method comprises the steps: obtaining voice data transmitted by a voice transmitting end to a voice receiving end; according to a sending main language corresponding to the sent voice data and a receiving main language corresponding to the voice receiving end, conflict detection of direct semantic conversion is carried out on the voice data, and a conflict detection result is obtained; according to the conflict detection result, obtaining a semantic conversion mode of the sent voice data; according to the semantic conversion mode, translating the sent voice data into an output message which is in semantic fit with the receiving main language; and sending the output message to a voice receiving end corresponding to the receiving main language. According to the method and the intelligent customer service robot provided by the invention, interaction obstacles occurring during interaction between different language voice data can be effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of language processing, in particular to a voice recognition method suitable for multiple languages and an intelligent customer service robot. BACKGROUND

[0002] In the process of voice interaction in multiple languages, the input voice of different voice input terminals can be recognized and translated by using an intelligent customer service robot, so as to ensure that multiple parties can complete interactive communication in different languages.

[0003] In the process of voice interaction in multiple languages, if one party inputs some sentence patterns, due to different languages and different regions of the parties in voice interaction, the other party cannot obtain the same semantic understanding based on the text content of the same sentence pattern, so that there is a certain semantic deviation in understanding of the parties in voice interaction, and further communication barriers are formed. SUMMARY

[0004] The present application provides a voice recognition method suitable for multiple languages and an intelligent customer service robot, to solve the technical problem that in the prior art, due to different languages and different regions of the parties in voice interaction, the other party cannot obtain the same semantic understanding based on the text content of the same sentence pattern, so that there is a certain semantic deviation in understanding of the parties in voice interaction, and further communication barriers are formed.

[0005] To achieve the above object and other related objects, the present application provides a voice recognition method suitable for multiple languages, comprising: obtaining the voice data sent by a voice sending terminal to a voice receiving terminal; detecting the conflict of direct semantic conversion of the voice data according to the sending main language corresponding to the voice data and the receiving main language corresponding to the voice receiving terminal, to obtain a conflict detection result; obtaining the semantic conversion mode of the voice data according to the conflict detection result; translating the voice data into an output message with semantic close to the receiving main language according to the semantic conversion mode; and sending the output message to the voice receiving terminal corresponding to the receiving main language.

[0006] In an embodiment of the present application, the method further comprises: obtaining historical dialogue data before the voice data is sent, the historical dialogue data comprising first historical output data corresponding to the voice sending end and second historical output data corresponding to the voice receiving end; performing usage language analysis on the voice sending end and the voice receiving end according to the historical dialogue data to determine a first skilled language set corresponding to the voice sending end and a second skilled language set corresponding to the voice receiving end, the first skilled language set comprising a plurality of first skilled languages and the second skilled language set comprising a plurality of second skilled languages; performing usage language analysis on the voice data to obtain a corresponding sending language; determining whether the sending language is in the first skilled language set; if yes, taking a first skilled language corresponding to a maximum proficiency in the first skilled language set as the sending main language of the voice sending end; if no, taking a first skilled language corresponding to a maximum proficiency in the first skilled language set as the sending main language of the voice sending end; determining whether there is a second skilled language corresponding to the sending main language in the second skilled language set; if yes, taking the second skilled language corresponding to the sending main language as the receiving main language; if no, taking a second skilled language corresponding to a maximum proficiency in the second skilled language set as the receiving main language of the voice sending end.

[0007] In an embodiment of the present application, the conflict detection of the direct semantic conversion of the voice data is performed according to the sending main language corresponding to the voice data and the receiving main language corresponding to the voice receiving end to obtain a conflict detection result, comprising: determining whether the sending main language and the receiving main language are consistent; if yes, taking the direct semantic conversion of the voice data as a first conflict detection result; if no, taking a language difference between the sending main language and the receiving main language as a second conflict detection result, the language difference comprising at least one of a language difference and a regional difference; wherein the determination of whether the sending main language and the receiving main language are consistent comprises: determining whether the sending main language and the receiving main language satisfy the following conditions simultaneously: the language of the sending main language is the same as that of the receiving main language; the region of the sending main language is the same as that of the receiving main language; if yes, it is determined that the sending main language and the receiving main language are consistent; if no, it is determined that the sending main language and the receiving main language are inconsistent.

[0008] In an embodiment of the present application, the semantic conversion mode for the outgoing voice data is obtained according to the conflict detection result, including: when the conflict detection result is that there is a language difference between the outgoing main language and the receiving main language, then the semantic conversion of the outgoing voice data is performed according to the outgoing main language and the first standard language, as a first semantic conversion mode for the outgoing voice data, and the first standard language is a standard language corresponding to the language of the receiving main language; when the conflict detection result is that there is a regional difference between the outgoing main language and the receiving main language, then the semantic conversion of the outgoing voice data is performed according to the outgoing main language, the second standard language and the receiving main language, as a second semantic conversion mode for the outgoing voice data, and the second standard language is a standard language corresponding to the language of the outgoing main language and the receiving main language; when the conflict detection result is that there is a language difference and a regional difference between the outgoing main language and the receiving main language, then the semantic conversion of the outgoing voice data is performed according to the outgoing main language, the third standard language, the fourth standard language and the receiving main language, as a third semantic conversion mode for the outgoing voice data, and the third standard language is a standard language corresponding to the language of the outgoing main language, and the fourth standard language is a standard language corresponding to the language of the receiving main language.

[0009] In an embodiment of the present application, the outgoing voice data is translated into an output message that is semantically close to the receiving main language according to the semantic conversion mode, including: converting the outgoing voice data into an outgoing semantic text corresponding to the outgoing main language; performing semantic conversion across languages on the outgoing semantic text according to the semantic conversion mode to obtain a detailed semantic text corresponding to the receiving main language; and performing semantic simplification on the detailed semantic text to translate the outgoing voice data into an output message that is semantically close to the receiving main language.

[0010] In an embodiment of the present application, the semantic conversion mode includes a first semantic conversion mode of performing semantic conversion on the outgoing voice data according to the outgoing main language and the first standard language; and the semantic conversion across languages on the outgoing semantic text to obtain a detailed semantic text corresponding to the receiving main language according to the semantic conversion mode and the previous dialogue data of the outgoing voice data, including: when the semantic conversion mode is the first semantic conversion mode, a first semantic difference text library between the outgoing main language and the first standard language is called; the outgoing semantic text is converted and difference monitored according to the first semantic difference text library; when it is monitored that there is an outgoing text sentence of the first difference text corresponding to the first semantic difference text library in the outgoing semantic text, the outgoing text sentence in the outgoing semantic text is removed to obtain a first remaining outgoing text; and the first remaining outgoing text and the first difference text are translated into a detailed semantic text corresponding to the receiving main language through a first large language model, and the first large language model is trained through language data corresponding to the first standard language.

[0011] In an embodiment of the present application, the semantic conversion mode includes a second semantic conversion mode of converting the outgoing semantic text according to the outgoing primary language, the second standard language and the receiving primary language; and converting the outgoing semantic text according to the semantic conversion mode to obtain the detailed semantic text corresponding to the receiving primary language, including: when the semantic conversion mode is the second semantic conversion mode, calling the second semantic difference text library between the outgoing primary language and the second standard language; converting the outgoing semantic text according to the second semantic difference text library; when the outgoing semantic text is monitored to have the outgoing text sentence corresponding to the second difference text in the second semantic difference text library, the outgoing text sentence in the outgoing semantic text is removed to obtain the second remaining outgoing text; the second remaining outgoing text and the second difference text are translated into the intermediate semantic long text corresponding to the second standard language through the second large language model corresponding to the second standard language, and the second large language model is trained through the language data corresponding to the second standard language; calling the third semantic difference text library between the second standard language and the receiving primary language; converting the intermediate semantic long text according to the third semantic difference text library; when the intermediate semantic long text is monitored to have the semantic text sentence corresponding to the third difference text in the third semantic difference text library, the semantic text sentence in the intermediate semantic long text is removed to obtain the third remaining outgoing text; and the third remaining outgoing text and the third difference text are translated into the detailed semantic text corresponding to the receiving primary language through the third large language model corresponding to the receiving primary language, and the third large language model is trained through the language data corresponding to the receiving primary language.

[0012] In an embodiment of the present application, the semantic conversion mode includes a third semantic conversion mode of converting the outgoing semantic text according to the outgoing primary language, the third standard language, the fourth standard language and the receiving primary language; and converting the outgoing semantic text according to the semantic conversion mode to obtain the detailed semantic text corresponding to the receiving primary language, including: when the semantic conversion mode is the third semantic conversion mode, converting the outgoing semantic text according to the third standard language, and translating the outgoing semantic text into the first intermediate semantic long text corresponding to the third standard language; converting the first intermediate semantic long text according to the fourth standard language, and translating the first intermediate semantic long text into the second intermediate semantic long text corresponding to the fourth standard language; converting the second intermediate semantic long text according to the receiving primary language, and translating the second intermediate semantic long text into the detailed semantic text corresponding to the receiving primary language.

[0013] In an embodiment of the present application, the outgoing semantic text is converted and detected according to a third standard language, and the outgoing semantic text is translated into a first intermediate semantic long text corresponding to the third standard language; the first intermediate semantic long text is converted and detected according to a fourth standard language, and the first intermediate semantic long text is translated into a second intermediate semantic long text corresponding to the fourth standard language; the second intermediate semantic long text is converted and detected according to a receiving main language, and the second intermediate semantic long text is translated into a detailed semantic text corresponding to the receiving main language, including: calling a fourth semantic difference text library between the outgoing main language and the third standard language; converting and monitoring the outgoing semantic text according to the fourth semantic difference text library; when it is monitored that there is an outgoing text sentence of a fourth difference text corresponding to the fourth semantic difference text library in the outgoing semantic text, the outgoing text sentence in the outgoing semantic text is eliminated to obtain a fourth remaining outgoing text, and the fourth remaining outgoing text and the fourth difference text are translated into the first intermediate semantic long text corresponding to the third standard language through a fourth large language model corresponding to the third standard language, and the fourth large language model is obtained by training language data corresponding to the third standard language; calling a fifth semantic difference text library between the third standard language and the fourth standard language; converting and monitoring the first intermediate semantic long text according to the fifth semantic difference text library; when it is monitored that there is a first semantic text sentence of a fifth difference text corresponding to the fifth semantic difference text library in the first intermediate semantic long text, the first semantic text sentence in the first intermediate semantic long text is eliminated to obtain a fifth remaining outgoing text, and the fifth remaining outgoing text and the fifth difference text are translated into the second intermediate semantic long text corresponding to the fourth standard language through a fifth large language model corresponding to the fourth standard language, and the fifth large language model is obtained by training language data corresponding to the fourth standard language; calling a sixth semantic difference text library between the fourth standard language and the receiving main language; converting and monitoring the second intermediate semantic long text according to the sixth semantic difference text library; when it is monitored that there is a second semantic text sentence of a sixth difference text corresponding to the sixth semantic difference text library in the second intermediate semantic long text, the second semantic text sentence in the second intermediate semantic long text is eliminated to obtain a sixth remaining outgoing text, and the sixth remaining outgoing text and the sixth difference text are translated into the detailed semantic text corresponding to the receiving main language through a sixth large language model corresponding to the receiving main language, and the sixth large language model is obtained by training language data corresponding to the receiving main language.

[0014] To achieve the above object and other related objects, the present application also provides a voice recognition intelligent customer service robot suitable for multiple languages, comprising: a data acquisition unit, configured to acquire voice data sent by a voice sending end to a voice receiving end; a conflict detection unit, configured to detect conflicts of direct semantic conversion of the voice data according to a sending main language corresponding to the voice data and a receiving main language corresponding to the voice receiving end, and obtain a conflict detection result; a mode acquisition unit, configured to acquire a semantic conversion mode of the voice data according to the conflict detection result; a message conversion unit, configured to translate the voice data into output messages in line with the semantic of the receiving main language according to the semantic conversion mode; and a message sending unit, configured to send the output messages to the voice receiving end corresponding to the receiving main language.

[0015] The voice recognition method and the intelligent customer service robot suitable for multiple languages provided by the present application can accurately translate the voice data sent by the voice sending end into the message content that the voice sending end really wants to express, and the semantic of the message content can also be in line with the understanding habit of the voice receiving end, so that the message content can be sent to the receiving main language, which can facilitate the conversion of the voice data into semantic content in line with the language habit of the voice receiving end in the multiple language and multiple habit voice interaction, and can effectively avoid the problem of interaction obstacle between different language voice data when the voice data sent by the voice sending end is directly translated. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application. It is to be understood that the drawings are only schematic, and that they do not necessarily represent a limiting case of the application. For the purpose of explanation and clearness, elements and structures shown in the drawings have not necessarily been drawn on scale. Moreover, in the drawings, like reference numerals designate corresponding parts throughout the different views, and similar components of the application bear the same reference numerals.

[0017] In the drawings: Figure 1 The flowchart of the voice recognition method suitable for multiple languages provided by the embodiments of the present application is shown.

[0018] Figure 2A flow chart showing a detailed semantic text obtaining process in a first semantic conversion mode provided by an embodiment of the present application.

[0019] Figure 3 A flow chart showing a detailed semantic text obtaining process in a second semantic conversion mode provided by an embodiment of the present application.

[0020] Figure 4 A flow chart showing a detailed semantic text obtaining process in a third semantic conversion mode provided by an embodiment of the present application.

[0021] Figure 5 A structural block diagram of a voice recognition intelligent customer service robot suitable for multiple languages provided by an embodiment of the present application.

[0022] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0023] Reference signs are as follows: Electronic device 1; voice recognition intelligent customer service robot suitable for multiple languages 11; memory 12; processor 13; data obtaining unit 111; mode detection unit 112; mode obtaining unit 113; message conversion unit 114; message sending unit 115. DETAILED DESCRIPTION

[0024] The present application is described in more detail by the following specific examples. Other advantages and effects of the present application can be easily understood by those skilled in the art from the contents disclosed in this specification. The present application can also be implemented or applied by other different specific embodiments, and each detail in the specification can be modified or changed based on different views and applications without departing from the spirit of the present application. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0025] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present application, and the drawings only show the components related to the present application, not the number, shape and size of the components when actually implemented. The type, number and ratio of each component when actually implemented can be arbitrarily changed, and the layout type of the components can also be more complex.

[0026] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application, however, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details, and in other embodiments, the known structures and devices are shown in the form of block diagrams rather than in the form of details, to avoid making the embodiments of the present application difficult to understand.

[0027] The application provides a speech recognition method suitable for multiple languages, which can determine whether there is a semantic conflict with the speech receiving end when the outgoing speech data of the speech sending end is translated into multiple languages, by determining whether the main language of the speech sending end and the speech receiving end is consistent, and determining whether to directly convert the semantic of the outgoing speech data based on the conflict detection result. If the semantic cannot be directly converted, the corresponding semantic conversion mode can be selected based on the difference between the outgoing main language and the receiving main language, so as to accurately translate the outgoing speech data of the speech sending end into the message content that the speech sending end really wants to express, and the semantic of the message content can also conform to the understanding habit of the speech receiving end. In this way, the message content is sent to the receiving main language, which can facilitate the conversion of the outgoing speech data into semantic content conforming to the language habit of the speech receiving end in multiple language and multiple habit voice interactions, and can effectively avoid the problem of interaction obstacles between different language voice data when the speech sending end directly translates the outgoing speech data, resulting in understanding deviation when the speech receiving end understands the translated text of the outgoing speech data.

[0028] Figure 1 A flowchart of a speech recognition method suitable for multiple languages in an exemplary embodiment of the application is shown, which is applied to a speech recognition system between a speech sending end and a speech receiving end, for receiving the outgoing speech data sent by the speech sending end, and sending it to the corresponding speech receiving end after translation processing, including steps S10-S50. The technical solutions of the application will be described in detail below. Figure 1

[0029] Firstly, step S10 is performed to obtain the outgoing speech data sent by the speech sending end to the speech receiving end.

[0030] When the speech sending end interacts with other user ends, the outgoing speech data sent by the speech sending end is pre-acquired by the speech recognition system, and is sent to each speech receiving end after being processed by the speech recognition system to conform to the understanding habit of each speech receiving end, so as to ensure that the speech receiving end can accurately understand the intention of the speech sending end.

[0031] After step S10, it further includes: The historical dialogue data before the outgoing speech data is obtained, including the first historical output data corresponding to the speech sending end and the second historical output data corresponding to the speech receiving end; ​According to historical dialogue data, the language use of the voice sending end and the voice receiving end is analyzed to determine a first set of proficient languages corresponding to the voice sending end and a second set of proficient languages corresponding to the voice receiving end, the first set of proficient languages including a plurality of first proficient languages, and the second set of proficient languages including a plurality of second proficient languages; The voice sending data is analyzed for language use to obtain a corresponding sending language; It is determined whether the sending language is in the first set of proficient languages; If yes, the first proficient language corresponding to the first set of proficient languages is taken as the main sending language of the voice sending end; If no, the first proficient language corresponding to the maximum proficiency in the first set of proficient languages is taken as the main sending language of the voice sending end; It is determined whether there is a second proficient language corresponding to the main sending language in the second set of proficient languages; If yes, the second proficient language corresponding to the main sending language is taken as the main receiving language; If no, the second proficient language corresponding to the maximum proficiency in the second set of proficient languages is taken as the main receiving language of the voice sending end.

[0032] After the voice recognition system obtains the voice sending data, it first collects historical dialogue data corresponding to the voice sending end and the voice receiving end before the voice sending data, which can be close to the time node of the voice sending data. After obtaining the historical dialogue data, the voice sending end can be analyzed for language use according to the first historical output data, so as to determine the first set of proficient languages corresponding to the voice sending end. Specifically, the language types involved in the first historical output data are first searched to obtain all language types used by the voice sending end for communication. Then, the proficiency of each language type corresponding to the voice sending end is determined according to the word amount proportion, word syntax accuracy, and speech speed fluency in the voice in the first historical output data. The word amount proportion can be the word amount corresponding to each language type in the total word amount of the first historical output data , which can be expressed by the formula . The word syntax accuracy can be the first proportion of the error word amount corresponding to each language type in the total sentence word amount of the same sentence , which can be expressed by the formula ; and the second proportion of the sentence amount corresponding to each language type in the total amount of all sentences of the language type , the formula can be expressed as ; and according to the first proportion weight corresponding to the designed first proportion of the first proportion weight , the second proportion weight corresponding to the second proportion of the second proportion weight , the word grammar accuracy of the corresponding language type is calculated . The speech speed fluency in the speech can be determined by calculating the reading speed of each word when reading compared with the regular speed corresponding to the corresponding language type , combined with the preset fluency conversion coefficient , to determine the speech speed fluency . Finally, based on the preset first proficiency conversion coefficient corresponding to the word amount proportion , the second proficiency conversion coefficient corresponding to the word grammar accuracy and the third proficiency conversion coefficient corresponding to the speech speed fluency , the word amount proportion, the word grammar accuracy, the speech speed fluency in the speech are integrated to determine the proficiency of each language type .

[0033] After calculating the proficiency of each language type of the speech sending end, the language type whose proficiency exceeds the set threshold value can be selected as the first fluent language, and the first fluent language set can be formed based on the first fluent language. Similarly, the proficiency of each language type of the speech receiving end can be determined based on the second historical output data corresponding to the speech receiving end, and the second fluent language can be selected based on the comparison with the set threshold value, so as to form the second fluent language set.

[0034] After determining the first fluent language set corresponding to the speech sending end and the second fluent language set corresponding to the speech receiving end, the use language of the current speech sending end, that is, the sending language, can be determined by using language analysis on the sending speech data. Then, it is determined whether the sending language of the current speech sending end is fluent by judging whether the sending language is in the first fluent language set. When it is not fluent, that is, the speech sending end cannot truly understand the deep semantics of the sending language corresponding to the sending speech data, therefore, the first fluent language corresponding to the maximum proficiency of the speech sending end is also used as the sending main language of the speech sending end.

[0035] When determining the main language, the speech receiving end first judges whether the second fluent language corresponding to the sending main language exists in the second fluent language set of the speech receiving end. If it exists, it means that the speech receiving end can understand the sending speech data of the speech sending end. Otherwise, the second fluent language corresponding to the maximum proficiency of the second fluent language set can also be used as the receiving main language of the speech sending end to ensure the voice interaction between the speech sending end and the speech receiving end.

[0036] Then, step S20 is performed to detect conflicts of direct semantic conversion of the voice data according to the sending main language corresponding to the sending voice data and the receiving main language corresponding to the voice receiving end, to obtain a conflict detection result.

[0037] After the voice recognition system determines the sending main language corresponding to the sending voice data and the receiving main language corresponding to the voice receiving end, it can determine whether the voice data can be directly semantically converted based on the difference between the sending main language and the receiving main language, and when the conflict detection result is direct semantic conversion, only direct conversion can be completed. If direct semantic conversion is not possible, the corresponding translation processing mode is selected based on the corresponding conflict situation to ensure that the semantics of the voice sending end can be accurately understood by the voice receiving end in accordance with the understanding mode of the voice receiving end.

[0038] In step S20, conflicts of direct semantic conversion of the sending voice data are detected according to the sending main language corresponding to the sending voice data and the receiving main language corresponding to the voice receiving end, to obtain a conflict detection result, including: determining whether the sending main language and the receiving main language are consistent; if yes, direct semantic conversion of the sending voice data is taken as the first conflict detection result; if no, the language difference between the sending main language and the receiving main language is taken as the second conflict detection result, and the language difference includes at least one of a language difference and a regional difference; wherein determining whether the sending main language and the receiving main language are consistent includes: determining whether the sending main language and the receiving main language simultaneously satisfy the following conditions: the language of the sending main language and the receiving main language is the same; the region of the sending main language and the receiving main language is the same; if yes, it means that the sending main language and the receiving main language are consistent; if no, it means that the sending main language and the receiving main language are not consistent.

[0039] When performing conflict detection, it is necessary to determine whether the sending main language and the receiving main language are consistent, that is, to determine whether the sending main language and the receiving main language are the same language and the same region. If yes, it means that the sending main language and the receiving main language are consistent. If at least one item is not the same, it means that the sending main language and the receiving main language are not consistent. Therefore, when consistent, direct semantic conversion of the voice data can be taken as the first conflict detection result, and the first conflict detection result can be used to directly complete translation of the sending voice data corresponding to the sending main language, that is, corresponding to the receiving main language.

[0040] Then, step S30 is performed to obtain a semantic conversion mode for the outgoing voice data according to the conflict detection result. When the conflict detection result is the second conflict detection result, it indicates that there is a language difference or a region difference between the outgoing primary language and the receiving primary language, or both.

[0041] In step S30, the semantic conversion mode for the outgoing voice data is obtained according to the conflict detection result, including: When the conflict detection result is that there is a language difference between the outgoing primary language and the receiving primary language, the semantic conversion of the outgoing voice data is performed according to the outgoing primary language and a first standard language, which is a standard language of the corresponding language of the receiving primary language, as a first semantic conversion mode for the outgoing voice data. When the conflict detection result is that there is a region difference between the outgoing primary language and the receiving primary language, the semantic conversion of the outgoing voice data is performed according to the outgoing primary language, a second standard language, and the receiving primary language, which is a standard language of the corresponding language of the outgoing primary language and the receiving primary language, as a second semantic conversion mode for the outgoing voice data. When the conflict detection result is that there is both a language difference and a region difference between the outgoing primary language and the receiving primary language, the semantic conversion of the outgoing voice data is performed according to the outgoing primary language, a third standard language, a fourth standard language, and the receiving primary language, which is a standard language of the corresponding language of the outgoing primary language and a standard language of the corresponding language of the receiving primary language, as a third semantic conversion mode for the outgoing voice data.

[0042] The language difference of the conflict detection result includes a language difference or a region difference, or both. When the conflict detection result is that there is a language difference between the outgoing primary language and the receiving primary language, it indicates that the proficient language of the language receiving end is a standard language of the corresponding language, and at this time, there is no need to perform a region language conversion. Therefore, only through the standard language of the corresponding language of the outgoing primary language and the receiving primary language, the semantic conversion of the outgoing voice data can be completed as a first semantic conversion mode for the outgoing voice data.

[0043] When the conflict detection result is that there is a region difference between the outgoing primary language and the receiving primary language, it indicates that the outgoing primary language is a standard language of the corresponding language, and there is no need to perform a conversion of the corresponding language, and the receiving primary language is not a standard language, and needs to be converted according to the corresponding standard language, that is, the semantic conversion of the outgoing voice data is performed through the outgoing primary language, a second standard language of the corresponding language of the receiving primary language, and the receiving primary language, as a second semantic conversion mode for the outgoing voice data.

[0044] In addition, when the conflict detection result is that there are both language differences and regional differences between the sending primary language and the receiving primary language, it indicates that neither the sending primary language nor the receiving primary language is the standard language of the corresponding language, and both need to be converted into the standard language, that is, the semantic conversion is performed on the sending voice data by using the sending primary language, a third standard language corresponding to the language of the sending primary language, a fourth standard language corresponding to the language of the receiving primary language, and the receiving primary language, as a third semantic conversion mode of the sending voice data.

[0045] Through the above three semantic conversion modes, the sending voice data can be effectively translated into output messages that can be accurately understood according to the language and regional conditions of different primary language interaction parties.

[0046] Then, step S40 is performed, and the sending voice data is translated into output messages that are semantically close to the receiving primary language according to the semantic conversion mode.

[0047] After determining the above several semantic conversion modes, the sending voice data can be translated into output messages that are semantically close to the receiving primary language according to different conditions of the sending primary language and the receiving primary language, so as to ensure that the voice receiving end can accurately understand the intention of the voice sender after viewing the messages.

[0048] In step S40, the sending voice data is translated into output messages that are semantically close to the receiving primary language according to the semantic conversion mode, including: The sending voice data is converted into sending semantic text corresponding to the sending primary language; The sending semantic text is subjected to cross-language semantic conversion according to the semantic conversion mode, to obtain detailed semantic text corresponding to the receiving primary language; The detailed semantic text is subjected to semantic simplification, and the sending voice data is translated into output messages that are semantically close to the receiving primary language.

[0049] When the outgoing voice data is translated, the outgoing voice data can be converted into the outgoing semantic text corresponding to the outgoing master language in advance. If the outgoing language corresponding to the outgoing voice data is different from the outgoing master language, the outgoing master language corresponding to the understanding habit is directly translated. For example, the English "good good study, day day up" is directly translated into "good good study, day day up" when the outgoing master language is Chinese and the English of the voice outgoing end is not a fluent language. Then, the semantic text is obtained by detailed conversion of the semantic text, that is, the outgoing semantic text can be further expanded based on the translated text by the large language model corresponding to the outgoing master language to accurately express the outgoing semantic of the translated text. For example, after expansion, the outgoing semantic text of "good good study, day day up" is "for the learning spirit, no matter at which stage of life, an open mind should be maintained, and learning should be pursued to constantly improve and progress". After obtaining the outgoing semantic text, the semantic conversion mode corresponding to the semantic conversion mode is further used to realize the semantic conversion of the outgoing semantic text across languages to obtain the detailed semantic text corresponding to the receiving master language. Finally, in order to ensure the simplification of translation, the detailed semantic text needs to be processed by semantic simplification to obtain the output message of the receiving master language semantic, and sent to the voice receiving end to ensure that the voice outgoing end can accurately express its own intention, and the voice receiving end can also accurately understand the intention of the voice outgoing end. That is, if the outgoing master language is Chinese and the English "good good study, day day up" is spoken, when the receiving master language of the voice receiving end is English, "study hard and make progress every day" or "study well and progress every day" can be further translated. Of course, the translated semantic can be further simplified.

[0050] For example, when Chinese says "Skate on thin ice" to express the intended meaning, due to the difference in context, "Skate on thin ice" in Chinese is understood as "extremely cautious in dangerous situations", and if this "Skate on thin ice" is directly sent to the voice receiving end whose main language is English, it will be misinterpreted as "when in danger, can handle it skillfully". If the translated English of "extremely cautious in dangerous situations" is the standard semantic, and the regional language of the receiving end is English, the translated English needs to be further translated into the regional language of the voice receiving end to ensure that the voice receiving end can accurately understand the accurate semantics of Chinese "Skate on thin ice" when speaking in English "Skate on thin ice". Of course, there can be other different main language types corresponding to different semantics.

[0051] For the semantic conversion mode, the first semantic conversion mode for converting the outgoing semantic text according to the outgoing main language and the first standard language can be included.

[0052] When the semantic conversion mode is the first semantic conversion mode, the outgoing semantic text is converted according to the semantic conversion mode, and the detailed semantic text corresponding to the receiving main language is obtained, including: When the semantic conversion mode is the first semantic conversion mode, the first semantic difference text library between the outgoing main language and the first standard language is called; According to the first semantic difference text library, the conversion difference of the outgoing semantic text is monitored; When it is monitored that the outgoing text sentence of the outgoing semantic text exists in the first difference text corresponding to the first semantic difference text library, the outgoing text sentence in the outgoing semantic text is eliminated to obtain the first remaining outgoing text; the first remaining outgoing text and the first difference text are translated into the detailed semantic text corresponding to the receiving main language through the first large language model, and the first large language model is trained through the language data corresponding to the first standard language.

[0053] Please refer to Figure 2In the cross-lingual semantic conversion of the issued semantic text by using the first semantic conversion mode, the first semantic difference text library between the issued main language and the first standard language can be called first. The first semantic difference text library includes each text data corresponding to the issued main language and the standard text data corresponding to the first standard language having a corresponding conversion relationship. Therefore, the conversion difference text existing in the issued semantic text of the voice issuing end can be monitored first. When the issued text sentence with differences is found, the first difference text corresponding to the issued text sentence is found based on the first semantic difference text library, and the issued text sentence is replaced to ensure the semantic accuracy of the entire issued semantic text. Then, in order to ensure the semantic fluency of the first remaining issued text after the first difference text is added, the first large language model trained by the language data corresponding to the first standard language can be used to translate the first remaining issued text and the first difference text into the detailed semantic text corresponding to the receiving main language.

[0054] For the semantic conversion mode, the second semantic conversion mode for converting the issued voice data according to the issued main language, the second standard language and the receiving main language can also be included.

[0055] When the semantic conversion mode is the second semantic conversion mode, the cross-lingual semantic conversion of the issued semantic text according to the semantic conversion mode is performed to obtain the detailed semantic text corresponding to the receiving main language, including: When the semantic conversion mode is the second semantic conversion mode, the second semantic difference text library between the issued main language and the second standard language is called; According to the second semantic difference text library, the conversion difference of the issued semantic text is monitored; When the issued text sentence with the second difference text corresponding to the second semantic difference text library is monitored in the issued semantic text, the issued text sentence in the issued semantic text is removed to obtain the second remaining issued text. The second remaining issued text and the second difference text are translated into the intermediate semantic long text corresponding to the second standard language by the second large language model corresponding to the second standard language. The second large language model is trained by the language data corresponding to the second standard language; The third semantic difference text library between the second standard language and the receiving main language is called; According to the third semantic difference text library, the conversion difference of the intermediate semantic long text is monitored; When it is monitored that there is a third difference text corresponding to the third semantic difference text library in the intermediate semantic long text, the semantic text sentence in the intermediate semantic long text is removed, and a third remaining sending text is obtained; the third remaining sending text and the third difference text are translated into detailed semantic text corresponding to the receiving main language by the third large language model corresponding to the receiving main language, and the third large language model is trained by language data corresponding to the receiving main language.

[0056] Please refer to Figure 3 When the second semantic conversion mode is used for cross-language semantic conversion of the sending semantic text, the second semantic difference text library between the sending main language and the second standard language can be called first, and the second semantic difference text library includes text data corresponding to each sending main language and standard text data corresponding to the second standard language having a corresponding conversion relationship. Therefore, the conversion difference text existing in the sending semantic text of the voice sending end can be monitored first, and when the difference exists, the second difference text corresponding to the sending text sentence is found based on the second semantic difference text library to replace the sending text sentence, so as to ensure the semantic accuracy of the entire sending semantic text. Then, in order to ensure the semantic fluency of the second remaining sending text after adding the second difference text, the second large language model trained by the language data corresponding to the second standard language can be used to translate the second remaining sending text and the second difference text into the intermediate semantic long text corresponding to the second standard language.

[0057] Then, the third semantic difference text library between the second standard language and the receiving main language can be called again, and the third semantic difference text library includes text data corresponding to each receiving main language and standard text data corresponding to the second standard language having a corresponding conversion relationship. Therefore, the conversion difference text existing in the intermediate semantic long text can be monitored first, and when the difference exists, the third difference text corresponding to the semantic text sentence is found based on the third semantic difference text library to replace the semantic text sentence, so as to ensure the semantic accuracy of the entire intermediate semantic long text. Then, in order to ensure the semantic fluency of the third remaining sending text after adding the third difference text, the third large language model trained by the language data corresponding to the receiving main language can be used to translate the third remaining sending text and the third difference text into the detailed semantic text corresponding to the receiving main language.

[0058] For the semantic conversion mode, the third semantic conversion mode for converting the semantic voice data according to the sending main language, the third standard language, the fourth standard language and the receiving main language is also included.

[0059] When the semantic conversion mode is the third semantic conversion mode, the outgoing semantic text is converted and detected according to the third standard language, and the outgoing semantic text is translated into the first intermediate semantic long text corresponding to the third standard language; When the semantic conversion mode is the third semantic conversion mode, the outgoing semantic text is converted and detected according to the third standard language, and the outgoing semantic text is translated into the first intermediate semantic long text corresponding to the third standard language; The first intermediate semantic long text is converted and detected according to the fourth standard language, and the first intermediate semantic long text is translated into the second intermediate semantic long text corresponding to the fourth standard language; The second intermediate semantic long text is converted and detected according to the receiving main language, and the second intermediate semantic long text is translated into the detailed semantic text corresponding to the receiving main language.

[0060] In the detailed semantic text conversion using the third semantic conversion mode, in order to ensure the accuracy of the detailed semantic text corresponding to the receiving main language, the outgoing semantic text needs to be converted into the first intermediate semantic long text corresponding to the third standard language, then the first intermediate semantic long text is converted into the second intermediate semantic long text corresponding to the fourth standard language, and finally the second intermediate semantic long text is continued to be converted and translated into the detailed semantic text corresponding to the receiving main language based on the corresponding receiving main language under the standard language, so that the outgoing semantic text can be accurately understood by the voice receiving end when it is sent to the voice receiving end.

[0061] Among them, for converting and detecting the outgoing semantic text according to the third standard language, and translating the outgoing semantic text into the first intermediate semantic long text corresponding to the third standard language; converting and detecting the first intermediate semantic long text according to the fourth standard language, and translating the first intermediate semantic long text into the second intermediate semantic long text corresponding to the fourth standard language; converting and detecting the second intermediate semantic long text according to the receiving main language, and translating the second intermediate semantic long text into the detailed semantic text corresponding to the receiving main language, including: Call the fourth semantic difference text library between the outgoing main language and the third standard language; According to the fourth semantic difference text library, the conversion difference of the outgoing semantic text is monitored; When it is monitored that the outgoing text sentence of the outgoing semantic text exists in the fourth difference text corresponding to the fourth semantic difference text library, the outgoing text sentence in the outgoing semantic text is removed to obtain the fourth remaining outgoing text, and the fourth remaining outgoing text and the fourth difference text are translated into the first intermediate semantic long text corresponding to the third standard language through the fourth large language model corresponding to the third standard language, and the fourth large language model is obtained by training the language data corresponding to the third standard language; retrieve a fifth semantic difference text library between the third standard language and the fourth standard language; According to the fifth semantic difference text library, the first intermediate semantic long text is monitored for conversion difference; When it is monitored that the first semantic text statement corresponding to the fifth difference text in the first intermediate semantic long text exists in the fifth semantic difference text library, the first semantic text statement in the first intermediate semantic long text is removed to obtain a fifth remaining sending text. The fifth remaining sending text and the fifth difference text are translated into the second intermediate semantic long text corresponding to the third standard language by a fifth large language model corresponding to the fourth standard language. The fifth large language model is trained by language data corresponding to the fourth standard language. retrieve a sixth semantic difference text library between the fourth standard language and the receiving main language; According to the sixth semantic difference text library, the second intermediate semantic long text is monitored for conversion difference; When it is monitored that the second semantic text statement corresponding to the sixth difference text in the second intermediate semantic long text exists in the sixth semantic difference text library, the second semantic text statement in the second intermediate semantic long text is removed to obtain a sixth remaining sending text. The sixth remaining sending text and the sixth difference text are translated into the detailed semantic text corresponding to the receiving main language by a sixth large language model corresponding to the receiving main language. The sixth large language model is trained by language data corresponding to the receiving main language.

[0062] Please refer to Figure 4 In the semantic conversion of the sending semantic text across languages by the third semantic conversion mode, the fourth semantic difference text library between the sending main language and the third standard language can be retrieved first. The fourth semantic difference text library includes text data corresponding to each sending main language and standard text data corresponding to the fourth standard language having a corresponding conversion relationship. Therefore, the conversion difference text existing in the sending semantic text of the voice sending end can be monitored first. When the sending text statement with difference is found, the fourth difference text corresponding to the sending text statement is found based on the fourth semantic difference text library to replace the sending text statement, so as to ensure the semantic accuracy of the entire sending semantic text. Then, in order to ensure the semantic fluency of the fourth remaining sending text after the fourth difference text is added, the fourth large language model trained by the language data corresponding to the third standard language can be used to translate the fourth remaining sending text and the fourth difference text into the first intermediate semantic long text corresponding to the third standard language.

[0063] Then, the fifth semantic difference text library between the third standard language and the fourth standard language can be called again, and the fifth semantic difference text library includes the standard text data corresponding to each third standard language and the standard text data corresponding to the fourth standard language with a corresponding conversion relationship. Therefore, the conversion difference text in the first intermediate semantic long text can be monitored first, and when the first semantic text statement with a difference is found, the fifth difference text corresponding to the first semantic text statement can be found based on the fifth semantic difference text library to replace the first semantic text statement, so as to ensure the semantic accuracy of the entire first intermediate semantic long text. Then, in order to ensure the semantic fluency of the fifth remaining sending text after the fifth difference text is added, the fifth large language model trained by the language data corresponding to the fourth standard language can be used to translate the fifth remaining sending text and the fifth difference text into the second intermediate semantic long text corresponding to the fourth standard language.

[0064] Subsequently, the sixth semantic difference text library between the fourth standard language and the receiving main language can be called again, and the sixth semantic difference text library includes the text data corresponding to each receiving main language and the standard text data corresponding to the fourth standard language with a corresponding conversion relationship. Therefore, the conversion difference text in the second intermediate semantic long text can be monitored first, and when the second semantic text statement with a difference is found, the sixth difference text corresponding to the second semantic text statement can be found based on the sixth semantic difference text library to replace the second semantic text statement, so as to ensure the semantic accuracy of the entire second intermediate semantic long text. Then, in order to ensure the semantic fluency of the sixth remaining sending text after the sixth difference text is added, the sixth large language model trained by the language data corresponding to the receiving main language can be used to translate the sixth remaining sending text and the sixth difference text into the detailed semantic text corresponding to the receiving main language.

[0065] Next, step S50 is performed to send the output message to the voice receiving end corresponding to the receiving main language. When the user at the voice receiving end receives the output message, the semantic intention of the other user sending the message can be accurately identified based on the output message, and the problem of semantic understanding deviation is solved.

[0066] Please refer to Figure 5The application further provides a multi-language voice recognition intelligent customer service robot 11, comprising: a data acquisition unit 111, configured to acquire voice data sent by a voice sending end to a voice receiving end; a conflict detection unit 112, configured to detect conflicts of direct semantic conversion of the voice data according to a sending main language corresponding to the voice data and a receiving main language corresponding to the voice receiving end, to obtain a conflict detection result; a mode acquisition unit 113, configured to acquire a semantic conversion mode of the voice data according to the conflict detection result; a message conversion unit 114, configured to translate the voice data into output messages in line with semantics of the receiving main language according to the semantic conversion mode; and a message sending unit 115, configured to send the output messages to the voice receiving end corresponding to the receiving main language.

[0067] It should be noted that the multi-language voice recognition intelligent customer service robot 11 provided in the above embodiment and the multi-language voice recognition method provided in the above embodiment belong to the same concept, and the specific manner in which each module and unit performs operations has been described in detail in the method embodiment, which will not be described here. In actual application, the multi-language voice recognition intelligent customer service robot 11 provided in the above embodiment can allocate the above functions to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above, and this is not limited here.

[0068] Please refer to Figure 6 The electronic device 1 can include a memory 12, a processor 13 and a bus, and can further include a computer program, such as a multi-language voice recognition program, stored in the memory 12 and executable on the processor 13.

[0069] The memory 12 includes at least one type of readable storage medium, such as a flash memory, a mobile hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 12 can include both an internal storage unit and an external storage device of the electronic device 1. The memory 12 can be used not only to store application software and various data installed in the electronic device 1, such as codes for multi-language voice recognition, but also to temporarily store data that has been output or will be output.

[0070] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the electronic device 1, connecting various components of the electronic device 1 through various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., speech recognition programs suitable for multiple languages) and calls data stored in the memory 12 to perform various functions of the electronic device 1 and process data.

[0071] The processor 13 executes the operating system of the electronic device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-described speech recognition method applicable to multiple languages.

[0072] For example, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into units within a multilingual voice recognition intelligent customer service robot.

[0073] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The software functional module, stored in the storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute some functions of the speech recognition method applicable to multiple languages ​​described in the various embodiments of this application.

[0074] In summary, the speech recognition method and intelligent customer service robot disclosed in this invention, applicable to multiple languages, first determines whether there is a semantic conflict between the main languages ​​of the speaker and receiver when the speaker sends voice data to the receiver. Based on the conflict detection results, it determines whether to directly perform semantic conversion on the sent voice data. If direct semantic conversion is not possible, an appropriate semantic conversion mode can be selected based on the differences between the main languages ​​of the speaker and receiver to accurately translate the sent voice data into the message content that the speaker actually wants to express. The semantics of this message content should also conform to the understanding habits of the receiver. This allows the message content to be sent to the receiver's main language. This facilitates the conversion of sent voice data into semantic content that conforms to the language habits of the receiver in multilingual and multi-habit multi-segment voice interactions. It effectively avoids the problem of misunderstanding when different receivers with different language spans understand the translated voice data due to direct translation of the sent voice data, which can cause interaction barriers when interacting with voice data in different languages. Therefore, this invention effectively overcomes the various shortcomings of the prior art and has high industrial application value.

[0075] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A speech recognition method applicable to multiple languages, characterized in that, include: Acquire the voice data sent from the voice transmitter to the voice receiver; Based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiving end, a conflict detection is performed on the voice data through direct semantic conversion to obtain a conflict detection result. Based on the conflict detection results, obtain the semantic conversion pattern of the emitted voice data; According to the semantic conversion mode, the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language; The output message is sent to the voice receiving terminal corresponding to the receiving host language.

2. The speech recognition method applicable to multiple languages ​​according to claim 1, characterized in that, Also includes: Acquire historical dialogue data prior to the transmission of voice data, wherein the historical dialogue data includes first historical output data corresponding to the voice transmitting end and second historical output data corresponding to the voice receiving end; Based on the historical dialogue data, language analysis is performed on the voice transmitter and the voice receiver to determine a first set of proficient languages ​​corresponding to the voice transmitter and a second set of proficient languages ​​corresponding to the voice receiver. The first set of proficient languages ​​includes multiple first proficient languages, and the second set of proficient languages ​​includes multiple second proficient languages. The emitted voice data is subjected to language analysis to obtain the corresponding emitted language; Determine whether the language spoken is within the first set of proficient languages; If so, the first proficient language corresponding to the first proficient language set shall be the primary language of the speech output end; If not, then the first proficient language corresponding to the highest proficiency in the first proficient language set shall be the primary language of the speech output end. Determine whether a second proficient language corresponding to the spoken language exists in the second proficient language set. If so, the second proficient language corresponding to the sending main language shall be used as the receiving main language; If not, then the second proficiency language corresponding to the maximum proficiency level in the second proficiency language set shall be the receiving primary language of the speech transmitter.

3. The speech recognition method applicable to multiple languages ​​according to claim 1, characterized in that, Based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiver, conflict detection is performed on the transmitted voice data through direct semantic conversion to obtain conflict detection results, including: Determine whether the sending language and the receiving language are consistent; If so, the transmitted voice data will be directly semantically converted and used as the first conflict detection result; If not, the language difference between the sending language and the receiving language is taken as the second conflict detection result, and the language difference includes at least one of language difference and regional difference; The determination of whether the sending language and the receiving language are consistent includes: Determine whether the sending language and the receiving language simultaneously satisfy the following conditions: The languages ​​of the sending and receiving main languages ​​are the same; The sending and receiving languages ​​are from the same region; If so, it means that the sending language and the receiving language are the same; If not, it means that the sending language and the receiving language are inconsistent.

4. The speech recognition method applicable to multiple languages ​​according to claim 1, characterized in that, Based on the conflict detection results, the semantic conversion pattern of the transmitted voice data is obtained, including: When the conflict detection result indicates that there is a language difference between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language and the first standard language as the first semantic conversion mode for the sent speech data. The first standard language is the standard language of the language corresponding to the receiving language. When the conflict detection result indicates that there is a regional difference between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language, the second standard language, and the receiving language as the second semantic conversion mode for the sent speech data. The second standard language is the standard language of the language that both the sending language and the receiving language correspond to. When the conflict detection result indicates that there are both language differences and regional differences between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language, the third standard language, the fourth standard language, and the receiving language as the third semantic conversion mode for the sent speech data. The third standard language is the standard language of the language corresponding to the sending language, and the fourth standard language is the standard language of the language corresponding to the receiving language.

5. The speech recognition method applicable to multiple languages ​​according to claim 1, characterized in that, According to the semantic conversion mode, translating the transmitted voice data into an output message that is semantically appropriate to the receiving host language includes: Convert the emitted voice data into emitted semantic text corresponding to the emitting main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language; The detailed semantic text is semantically simplified, and the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language.

6. The speech recognition method applicable to multiple languages ​​according to claim 5, characterized in that, The semantic conversion mode includes a first semantic conversion mode that performs semantic conversion on the emitted speech data based on the emitting main language and the first standard language; Based on the semantic conversion mode and the pre-existing dialogue data prior to the transmitted voice data, cross-language semantic conversion is performed on the transmitted semantic text to obtain detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the first semantic conversion mode, the first semantic difference text library between the issuing host language and the first standard language is retrieved; Based on the first semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the first difference text in the first semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the first remaining sent text; the first remaining sent text and the first difference text are translated into the detailed semantic text corresponding to the receiving host language by the first large language model, the first large language model being trained by language data corresponding to the first standard language.

7. The speech recognition method applicable to multiple languages ​​according to claim 5, characterized in that, The semantic conversion mode includes a second semantic conversion mode that performs semantic conversion on the transmitted speech data based on the transmitting main language, the second standard language, and the receiving main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the second semantic conversion mode, the second semantic difference text library between the issuing main language and the second standard language is retrieved; Based on the second semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the second difference text in the second semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the second remaining sent text; the second remaining sent text and the second difference text are translated into the second standard language corresponding to the second standard language through the second large language model, and the second large language model is trained through the language data corresponding to the second standard language. Retrieve a third semantic difference text library between the second standard language and the receiving host language; Based on the third semantic difference text library, the conversion difference monitoring is performed on the intermediate semantic long text; When a semantic text statement corresponding to the third difference text in the third semantic difference text library is detected in the intermediate semantic long text, the semantic text statement in the intermediate semantic long text is removed to obtain the third remaining transmitted text; the third remaining transmitted text and the third difference text are translated into the detailed semantic text corresponding to the receiving host language through the third major language model corresponding to the receiving host language, and the third major language model is trained through the language data corresponding to the receiving host language.

8. The speech recognition method applicable to multiple languages ​​according to claim 5, characterized in that, The semantic conversion mode includes a third semantic conversion mode that performs semantic conversion on the transmitted voice data based on the transmitting main language, the third standard language, the fourth standard language, and the receiving main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the third semantic conversion mode, the transmitted semantic text is converted and detected according to the third standard language, and the transmitted semantic text is translated into the first intermediate semantic long text corresponding to the third standard language; The first intermediate semantic long text is converted and detected according to the fourth standard language, and the first intermediate semantic long text is translated into the second intermediate semantic long text corresponding to the fourth standard language. The second intermediate semantic long text is converted and detected according to the receiving host language, and then translated into the detailed semantic text corresponding to the receiving host language.

9. The speech recognition method applicable to multiple languages ​​according to claim 8, characterized in that, The transmitted semantic text is converted and detected according to the third standard language, and translated into a first intermediate semantic long text corresponding to the third standard language; the first intermediate semantic long text is converted and detected according to the fourth standard language, and translated into a second intermediate semantic long text corresponding to the fourth standard language; the second intermediate semantic long text is converted and detected according to the receiving host language, and translated into the detailed semantic text corresponding to the receiving host language, including: Retrieve a fourth semantic difference text library between the main language and the third standard language; Based on the fourth semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the fourth difference text in the fourth semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the fourth remaining sent text. The fourth remaining sent text and the fourth difference text are translated into the first intermediate semantic long text corresponding to the third standard language through the fourth major language model corresponding to the third standard language. The fourth major language model is trained through the language data corresponding to the third standard language. Retrieve the fifth semantic difference text library between the third standard language and the fourth standard language; Based on the fifth semantic difference text library, the first intermediate semantic long text is subjected to conversion difference monitoring. When a first semantic text statement corresponding to the fifth difference text in the fifth semantic difference text library is detected in the first intermediate semantic long text, the first semantic text statement in the first intermediate semantic long text is removed to obtain the fifth remaining text. The fifth remaining text and the fifth difference text are then translated into the second intermediate semantic long text corresponding to the fourth standard language through the fifth major language model corresponding to the fourth standard language. The fifth major language model is trained using the language data corresponding to the fourth standard language. Retrieve the sixth semantic difference text library between the fourth standard language and the receiving host language; Based on the sixth semantic difference text library, the conversion difference monitoring is performed on the second intermediate semantic long text; When a second semantic text statement corresponding to the sixth difference text in the sixth semantic difference text library is detected in the second intermediate semantic long text, the second semantic text statement in the second intermediate semantic long text is removed to obtain the sixth remaining transmitted text. The sixth remaining transmitted text and the sixth difference text are then translated into the detailed semantic text corresponding to the receiving host language through the sixth major language model corresponding to the receiving host language. The sixth major language model is trained using the language data corresponding to the receiving host language.

10. A speech recognition intelligent customer service robot applicable to multiple languages, characterized in that, include: The data acquisition unit is used to acquire the voice data sent from the voice transmitter to the voice receiver. The conflict detection unit is used to perform direct semantic conversion on the voice data based on the sending language corresponding to the voice data and the receiving language corresponding to the voice receiver, and obtain the conflict detection result. The pattern acquisition unit is used to acquire the semantic conversion pattern of the transmitted voice data based on the conflict detection result. The message conversion unit is used to translate the transmitted voice data into an output message that is semantically appropriate to the receiving host language according to the semantic conversion mode. The message sending unit is used to send the output message to the voice receiving terminal corresponding to the receiving main language.

Citation Information

Patent Citations

  • Information processing method and device for voice conversion

    CN110442881A

  • Text translation method and device

    CN112749569A

  • Text semantic understanding method and device, equipment and storage medium

    CN114970541A

  • Cross-language voice communication system integrating AI voice cloning and real-time translation

    CN119580703A

  • Deep learning-based sign language-to-multilingual text speech mutual conversion method and application thereof

    CN120580985A