A multi-language voice recognition method and intelligent customer service robot
Patent Information
- Application Number
- CN202512031508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-12-30
AI Technical Summary
[0004]本发明提供一种适用于多语言的语音识别方法和智能客服机器人,以解决现有技术中由于语音交互各方的语种不同、地域不同等,交互另一方无法基于该同一句式的文本内容来获取到相同的习惯理解的语义,从而导致语音交互各方在理解时存在一定的语义偏差,进而形成交流障碍的技术问题
[0015]本发明的有益效果:本发明提出的一种适用于多语言的语音识别方法和智能客服机器人,在语音发出端向语音接收端发出的发出语音数据时,先通过确定语音发出端和语音接收端的主语言是否一致,来确定发出语音数据进行多语言翻译时是否与语音接收端存在语义冲突,并基于冲突检测结果来确定是否直接对发出语音数据进行语义转换,若不可以直接语义转换时,则还可以基于发出主语言与接收主语言之间的差异,选择相应的语义转换模式,以将语音发出端的发出语音数据准确地翻译为语音发出端真实想表达的消息内容,且该消息内容的语义还能够贴合语音接收端理解习惯,以此将该消息内容发送给接收主语言,可以便于在多语言多习惯的多段语音交互下,能够实现将发出语音数据转换为符合语音接收端的语言习惯的语义内容,可以有效地避免因直接翻译语音发出端的发出语音数据,而导致存在语言跨度的不同语音接收端在对发出语音数据的翻译译文进行理解时,出现理解偏差,造成不同语言语音数据之间进行交互时出现交互障碍的问题。
Smart Images

Figure CN121747576B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language processing technology, and in particular to a speech recognition method and intelligent customer service robot applicable to multiple languages. Background Technology
[0002] When conducting multilingual voice interaction, intelligent customer service robots can be used to perform speech recognition and translation processing on the input speech from different voice input terminals, so as to ensure that the multiple parties involved can complete interactive communication in different languages.
[0003] In the process of multilingual voice interaction, existing intelligent customer service robots may encounter semantic discrepancies when one party inputs certain sentence structures. Due to differences in language and region among the parties involved, the other party may not be able to obtain the same semantic understanding based on the text content of the same sentence structure. This can lead to certain semantic discrepancies in understanding among the parties involved, thus creating communication barriers. Summary of the Invention
[0004] This invention provides a speech recognition method and intelligent customer service robot applicable to multiple languages, in order to solve the technical problem in the prior art that, due to the different languages and regions of the parties involved in the voice interaction, the other party cannot obtain the same habitual understanding of the semantics based on the same sentence text content, resulting in certain semantic deviations in the understanding of the parties involved in the voice interaction, thus forming a communication barrier.
[0005] To achieve the above and other related objectives, the present invention provides a speech recognition method applicable to multiple languages, comprising: acquiring transmitted speech data sent from a speech transmitter to a speech receiver; performing direct semantic conversion conflict detection on the speech data based on the transmitting host language corresponding to the transmitted speech data and the receiving host language corresponding to the speech receiver, and obtaining a conflict detection result; acquiring a semantic conversion pattern for the transmitted speech data based on the conflict detection result; translating the transmitted speech data into an output message semantically appropriate to the receiving host language based on the semantic conversion pattern; and sending the output message to the speech receiver corresponding to the receiving host language.
[0006] In one embodiment of the present invention, the method further includes: acquiring historical dialogue data prior to the transmission of voice data, the historical dialogue data including first historical output data corresponding to the voice transmitting end and second historical output data corresponding to the voice receiving end; performing language analysis on the voice transmitting end and the voice receiving end based on the historical dialogue data to determine a first set of proficient languages corresponding to the voice transmitting end and a second set of proficient languages corresponding to the voice receiving end, the first set of proficient languages including multiple first proficient languages and the second set of proficient languages including multiple second proficient languages; performing language analysis on the transmitted voice data to obtain the corresponding transmitted language; determining whether the transmitted language is within the first set of proficient languages; if yes, then using the first proficient language corresponding to the first set of proficient languages as the primary language of the voice transmitting end; if no, then using the first proficient language corresponding to the maximum proficiency in the first set of proficient languages as the primary language of the voice transmitting end; determining whether there is a second proficient language corresponding to the primary language in the second set of proficient languages; if yes, then using the second proficient language corresponding to the primary language as the primary language of the receiving end; if no, then using the second proficient language corresponding to the maximum proficiency in the second set of proficient languages as the primary language of the receiving end.
[0007] In one embodiment of the present invention, conflict detection is performed on the transmitted voice data by direct semantic conversion based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiving end, to obtain a conflict detection result. This includes: determining whether the transmitting language and the receiving language are consistent; if yes, then the direct semantic conversion of the transmitted voice data is taken as the first conflict detection result; if no, then the language difference between the transmitting language and the receiving language is taken as the second conflict detection result, where the language difference includes at least one of language difference and regional difference; wherein, determining whether the transmitting language and the receiving language are consistent includes: determining whether the transmitting language and the receiving language simultaneously satisfy the following conditions: the transmitting language and the receiving language are of the same language; the transmitting language and the receiving language are of the same region; if yes, then the transmitting language and the receiving language are consistent; if no, then the transmitting language and the receiving language are inconsistent.
[0008] In one embodiment of the present invention, obtaining a semantic conversion mode for transmitted voice data based on conflict detection results includes: when the conflict detection result indicates a language difference between the transmitting and receiving languages, semantic conversion is performed on the transmitted voice data based on the transmitting language and a first standard language, serving as a first semantic conversion mode for the transmitted voice data, where the first standard language is the standard language corresponding to the language of the receiving language; when the conflict detection result indicates a regional difference between the transmitting and receiving languages, semantic conversion is performed on the transmitted voice data based on the transmitting language, a second standard language, and the receiving language, serving as a second semantic conversion mode for the transmitted voice data, where the second standard language is the standard language corresponding to the language of both the transmitting and receiving languages; when the conflict detection result indicates both language and regional differences exist between the transmitting and receiving languages, semantic conversion is performed on the transmitted voice data based on the transmitting language, a third standard language, a fourth standard language, and the receiving language, serving as a third semantic conversion mode for the transmitted voice data, where the third standard language is the standard language corresponding to the language of the transmitting language, and the fourth standard language is the standard language corresponding to the language of the receiving language.
[0009] In one embodiment of the present invention, translating transmitted voice data into an output message that is semantically appropriate to the receiving host language according to a semantic conversion mode includes: converting transmitted voice data into transmitted semantic text corresponding to the transmitting host language; performing cross-language semantic conversion on the transmitted semantic text according to the semantic conversion mode to obtain detailed semantic text corresponding to the receiving host language; and performing semantic simplification on the detailed semantic text to translate the transmitted voice data into an output message that is semantically appropriate to the receiving host language.
[0010] In one embodiment of the present invention, the semantic conversion mode includes a first semantic conversion mode for semantic conversion of transmitted speech data based on the transmitting subject language and a first standard language; and cross-language semantic conversion of transmitted semantic text based on the semantic conversion mode and prior dialogue data before the transmitted speech data to obtain detailed semantic text corresponding to the receiving subject language, including: when the semantic conversion mode is the first semantic conversion mode, retrieving a first semantic difference text library between the transmitting subject language and the first standard language; performing conversion difference monitoring on the transmitted semantic text based on the first semantic difference text library; when a transmitted text statement corresponding to the first difference text library is detected in the transmitted semantic text, the transmitted text statement in the transmitted semantic text is removed to obtain a first remaining transmitted text; and translating the first remaining transmitted text and the first difference text into detailed semantic text corresponding to the receiving subject language through a first large language model, wherein the first large language model is trained using language data corresponding to the first standard language.
[0011] In one embodiment of the present invention, the semantic conversion mode includes a second semantic conversion mode that performs semantic conversion on the transmitted speech data according to the transmitting subject language, a second standard language, and a receiving subject language; according to the semantic conversion mode, cross-language semantic conversion is performed on the transmitted semantic text to obtain detailed semantic text corresponding to the receiving subject language, including: when the semantic conversion mode is the second semantic conversion mode, retrieving a second semantic difference text library between the transmitting subject language and the second standard language; monitoring conversion differences in the transmitted semantic text according to the second semantic difference text library; when a transmitted text statement corresponding to the second difference text library is detected in the transmitted semantic text, the transmitted text statement in the transmitted semantic text is removed to obtain a second remaining transmitted text; the second remaining transmitted text and the second difference text are then compared and contrasted. The second major language model corresponding to the second standard language is used to translate the intermediate semantic long text into the second standard language. The second major language model is trained using language data corresponding to the second standard language. A third semantic difference text library between the second standard language and the receiving host language is retrieved. Based on the third semantic difference text library, the intermediate semantic long text is subjected to conversion difference monitoring. When a semantic text statement corresponding to the third difference text in the third semantic difference text library is detected in the intermediate semantic long text, the semantic text statement in the intermediate semantic long text is removed to obtain the third remaining transmitted text. The third remaining transmitted text and the third difference text are then translated into detailed semantic text corresponding to the receiving host language using the third major language model corresponding to the receiving host language. The third major language model is trained using language data corresponding to the receiving host language.
[0012] In one embodiment of the present invention, the semantic conversion mode includes a third semantic conversion mode that performs semantic conversion on the transmitted speech data according to the transmitting subject language, a third standard language, a fourth standard language, and the receiving subject language; according to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain detailed semantic text corresponding to the receiving subject language, including: when the semantic conversion mode is the third semantic conversion mode, performing conversion detection on the transmitted semantic text according to the third standard language, and translating the transmitted semantic text into a first intermediate semantic long text corresponding to the third standard language; performing conversion detection on the first intermediate semantic long text according to the fourth standard language, and translating the first intermediate semantic long text into a second intermediate semantic long text corresponding to the fourth standard language; performing conversion detection on the second intermediate semantic long text according to the receiving subject language, and translating the second intermediate semantic long text into detailed semantic text corresponding to the receiving subject language.
[0013] In one embodiment of the present invention, the transmitted semantic text is converted and detected according to a third standard language, and translated into a first intermediate semantic long text corresponding to the third standard language; the first intermediate semantic long text is converted and detected according to a fourth standard language, and translated into a second intermediate semantic long text corresponding to the fourth standard language; the second intermediate semantic long text is converted and detected according to the receiving main language, and translated into detailed semantic text corresponding to the receiving main language, including: retrieving a fourth semantic difference text library between the transmitting main language and the third standard language; performing conversion difference monitoring on the transmitted semantic text according to the fourth semantic difference text library; when a transmitted text statement corresponding to the fourth difference text library is detected in the transmitted semantic text, the transmitted text statement in the transmitted semantic text is removed to obtain a fourth remaining transmitted text; the fourth remaining transmitted text and the fourth difference text are translated into a first intermediate semantic long text corresponding to the third standard language through a fourth major language model corresponding to the third standard language, the fourth major language model being trained using language data corresponding to the third standard language; and retrieving a fifth semantic difference text library between the third standard language and the fourth standard language. Based on the fifth semantic difference text library, conversion difference monitoring is performed on the first intermediate semantic long text. When a first semantic text statement corresponding to the fifth difference text in the fifth semantic difference text library is detected in the first intermediate semantic long text, the first semantic text statement in the first intermediate semantic long text is removed to obtain the fifth remaining transmitted text. The fifth remaining transmitted text and the fifth difference text are then translated into the second intermediate semantic long text corresponding to the fourth standard language through the fifth major language model corresponding to the fourth standard language. The fifth major language model is trained using the language data corresponding to the fourth standard language. The sixth semantic difference text library between the fourth standard language and the receiving host language is retrieved. Based on the sixth semantic difference text library, conversion difference monitoring is performed on the second intermediate semantic long text. When a second semantic text statement corresponding to the sixth difference text in the second intermediate semantic long text is detected, the second semantic text statement in the second intermediate semantic long text is removed to obtain the sixth remaining transmitted text. The sixth remaining transmitted text and the sixth difference text are then translated into the detailed semantic text corresponding to the receiving host language through the sixth major language model corresponding to the receiving host language. The sixth major language model is trained using the language data corresponding to the receiving host language.
[0014] To achieve the above and other related objectives, the present invention also provides a speech recognition intelligent customer service robot applicable to multiple languages, comprising: a data acquisition unit for acquiring voice data sent from a voice transmitter to a voice receiver; a conflict detection unit for performing direct semantic conversion conflict detection on the voice data based on the transmitting language corresponding to the transmitting voice data and the receiving language corresponding to the receiving language of the voice receiver, and obtaining a conflict detection result; a pattern acquisition unit for acquiring a semantic conversion pattern of the transmitted voice data based on the conflict detection result; a message conversion unit for translating the transmitted voice data into an output message semantically appropriate to the receiving language based on the semantic conversion pattern; and a message sending unit for sending the output message to the voice receiver corresponding to the receiving language.
[0015] The beneficial effects of this invention are as follows: The speech recognition method and intelligent customer service robot proposed in this invention, applicable to multiple languages, first determine whether the primary languages of the speech transmitter and receiver are consistent to determine whether there is a semantic conflict when translating the speech data into multiple languages. Based on the conflict detection result, it is determined whether to directly perform semantic conversion on the speech data. If direct semantic conversion is not possible, an appropriate semantic conversion mode can be selected based on the differences between the primary languages of the transmitter and receiver to accurately translate the speech data into the message content that the transmitter actually wants to express. The semantics of this message content can also conform to the understanding habits of the receiver. This allows the message content to be sent to the receiver's primary language. This facilitates the conversion of speech data into semantic content that conforms to the language habits of the receiver in multilingual and multi-habit multi-segment speech interactions. It can effectively avoid the problem of misunderstanding when different receivers with different language spans understand the translated speech data due to direct translation of the speech data, which can cause interaction barriers when interacting with speech data in different languages. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0017] In the attached diagram: Figure 1 This is a flowchart illustrating a speech recognition method applicable to multiple languages provided in an embodiment of the present invention.
[0018] Figure 2This is a flowchart illustrating the detailed semantic text acquisition process under a first semantic conversion mode provided in an embodiment of the present invention.
[0019] Figure 3 This is a flowchart illustrating the detailed semantic text acquisition process under the second semantic conversion mode provided in an embodiment of the present invention.
[0020] Figure 4 This is a flowchart illustrating the detailed semantic text acquisition process under the third semantic conversion mode provided in an embodiment of the present invention.
[0021] Figure 5 The diagram shown is a structural block diagram of a speech recognition intelligent customer service robot applicable to multiple languages, provided in an embodiment of the present invention.
[0022] Figure 6 The diagram shown is a structural schematic of an electronic device according to an embodiment of the present invention.
[0023] The attached figures are labeled as follows: Electronic device 1; a multilingual voice recognition intelligent customer service robot 11; a memory 12; a processor 13; a data acquisition unit 111; a burst detection unit 112; a pattern acquisition unit 113; a message conversion unit 114; and a message sending unit 115. Detailed Implementation
[0024] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0025] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0026] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0027] This invention provides a speech recognition method applicable to multiple languages. When a speech data is sent from a speaker to a receiver, the method first determines whether the primary languages of the speaker and receiver are consistent. This helps to identify semantic conflicts when translating the speech data into multiple languages. Based on the conflict detection results, it determines whether to directly perform semantic conversion on the speech data. If direct semantic conversion is not possible, an appropriate semantic conversion mode can be selected based on the differences between the primary languages of the speaker and receiver. This ensures that the speech data is accurately translated into the message content that the speaker intends to express, and that the semantics of this message content aligns with the understanding habits of the receiver. This allows for the conversion of speech data into semantic content that conforms to the language habits of the receiver in multilingual and multi-segment speech interactions. This effectively avoids the problem of misunderstandings and interaction barriers that can occur when different receivers with different language gaps interpret the translated speech data due to direct translation.
[0028] Figure 1 A flowchart of a speech recognition method applicable to multiple languages according to an exemplary embodiment of this application is shown. Applied to a speech recognition system, the speech recognition system is located between a speech transmitter and a speech receiver, and is used to receive the spoken speech data emitted by the speech transmitter, and send it to the corresponding speech receiver after translation processing, including steps S10-S50. The following will be combined with… Figure 1 The technical solution of this application will be described in detail below.
[0029] First, execute step S1 0 to obtain the voice data sent from the voice transmitter to the voice receiver.
[0030] When the voice transmitter interacts with other user terminals, the voice data emitted by the voice transmitter is pre-acquired by the voice recognition system. After the voice recognition system performs corresponding translation processing to match the understanding habits of each voice receiver, it is sent to each voice receiver to ensure that the voice receiver can accurately understand the intention that the voice transmitter wants to express.
[0031] Following step S10, the method further includes: Acquire historical dialogue data prior to the transmission of voice data. The historical dialogue data includes the first historical output data corresponding to the voice transmitting end and the second historical output data corresponding to the voice receiving end. Based on historical dialogue data, language analysis is performed on the voice transmitter and the voice receiver to determine the first set of proficient languages corresponding to the voice transmitter and the second set of proficient languages corresponding to the voice receiver. The first set of proficient languages includes multiple first proficient languages, and the second set of proficient languages includes multiple second proficient languages. The spoken voice data is analyzed using language analysis to obtain the corresponding spoken language. Determine whether the language spoken is within the first set of proficient languages; If so, then the first proficient language corresponding to the first proficient language set shall be the primary language of the speech source. If not, then the first proficient language corresponding to the highest proficiency in the first proficient language set will be taken as the primary language of speech production. Determine whether a second proficient language exists in the set of second proficient languages that corresponds to the primary language spoken. If so, the second proficient language corresponding to the sending language will be used as the receiving language. If not, then the second proficiency language corresponding to the highest proficiency level in the second proficiency language set will be used as the primary receiving language of the speech transmitter.
[0032] After acquiring the transmitted speech data, the speech recognition system first collects historical dialogue data between the speaker and receiver prior to the transmission of the speech data. This historical dialogue data can be from a point in time close to the transmission of the speech data. After acquiring the historical dialogue data, the system can perform language analysis on the speaker based on the first historical output data to determine the speaker's first set of proficient languages. Specifically, it first searches for the language types involved in the first historical output data to obtain all the language types used by the speaker in communication. Then, based on the proportion of words of each language type in the first historical output data, the accuracy of word and grammar, and the fluency of speech, the proficiency of the speaker for each language type is determined. The proportion of words can be the proportion of words corresponding to each language type. Total number of words corresponding to the first historical output data The proportion below The formula can be expressed as Vocabulary grammatical accuracy can be achieved by monitoring the vocabulary grammar for each language type, and identifying the number of erroneous words when errors are detected. Total number of words in the same sentence The first proportion below The formula can be expressed as Then identify the number of sentences containing grammatical errors. Total number of statements corresponding to the language type The second percentage The formula can be expressed as ; and according to the first proportion weight corresponding to the designed first proportion. The second proportion corresponds to the second proportional weight. To calculate the grammatical accuracy of words and phrases in the corresponding language type. Fluency in speech can be assessed by calculating the reading speed of each word. Compared to the standard speed for the corresponding language type The difference, combined with the pre-set smoothness conversion coefficient To determine the fluency of speech rate Finally, based on the preset vocabulary proportions, the first proficiency conversion coefficient is used. The conversion coefficient of second proficiency corresponding to the accuracy of vocabulary and grammar And the third proficiency conversion coefficient corresponding to speech speed and fluency. By integrating the proportion of vocabulary, the accuracy of vocabulary and grammar, and the fluency of speech, the proficiency level of each language type is determined. .
[0033] After calculating the proficiency of each language type at the speech transmitter, the language type with a proficiency exceeding a set threshold can be selected as the first proficient language, and a first proficient language set can be constructed based on the first proficient language. Similarly, the proficiency of each language type at the speech receiver can be determined based on the corresponding second historical output data, and a second proficient language can be selected based on a comparison with a set threshold, thereby forming a second proficient language set.
[0034] After determining the first proficient language set corresponding to the voice transmitter and the second proficient language set corresponding to the voice receiver, the language used by the current voice transmitter, i.e., the transmitting language, can be determined by performing language analysis on the transmitted voice data. Then, by judging whether the transmitting language is within the first proficient language set, the proficiency of the current voice transmitter in the transmitting language is determined. If the language is not proficient, meaning that the voice transmitter cannot truly understand the deep semantics of the transmitting language corresponding to the transmitted voice data, the first proficient language corresponding to the voice transmitter's maximum proficiency will be used as the primary transmitting language of the voice transmitter.
[0035] When determining the primary language, the voice receiver first checks if a second proficient language corresponding to the primary language exists in its second proficient language set. If it does, it means the voice receiver can understand the voice data transmitted by the voice transmitter. Otherwise, similarly, the second proficient language corresponding to the highest proficiency level in the second proficient language set can be used as the primary language for receiving the voice data to ensure voice interaction with the voice transmitter.
[0036] Next, step S20 is executed, and conflict detection is performed on the voice data by direct semantic conversion based on the sending language corresponding to the voice data and the receiving language corresponding to the voice receiver, so as to obtain the conflict detection result.
[0037] After determining the issuing language of the spoken data and the receiving language of the receiving end, the speech recognition system can perform difference detection between the two languages to determine whether direct semantic translation of the speech data is possible. If the conflict detection result indicates that direct semantic translation is possible, the translation can be performed directly. If direct semantic translation is not possible, the system selects the appropriate translation processing method based on the corresponding conflict situation to ensure that the semantics of the spoken data can be accurately understood by the receiving end in accordance with its understanding pattern.
[0038] In step S20, conflict detection is performed on the transmitted voice data based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiver, to obtain the conflict detection result, including: Determine whether the sending and receiving languages are consistent. If so, the transmitted voice data will be directly semantically converted and used as the first conflict detection result; If not, the language difference between the sending and receiving languages will be used as the second conflict detection result. The language difference includes at least one of language difference and regional difference. The determination of whether the sending and receiving languages are consistent includes: Determine whether the sending and receiving languages simultaneously satisfy the following conditions: The languages of the main language being spoken and the main language being received are the same; The regions where the main language is sent and the regions where the main language is received are the same; If so, it means that the sending and receiving languages are the same; If not, it means that the sending language and the receiving language are inconsistent.
[0039] During conflict detection, it is necessary to determine whether the speaking and receiving languages are consistent. This means determining whether the speaking and receiving languages are the same language and geographical location. If so, the speaking and receiving languages are consistent. If at least one of these differs, the speaking and receiving languages are inconsistent. Therefore, when they are consistent, direct semantic conversion of the speech data can be used as the first conflict detection result. This first conflict detection result can then be used to directly translate the spoken speech data corresponding to the speaking language, which is also the receiving language.
[0040] Next, step S30 is executed to obtain the semantic conversion pattern of the transmitted speech data based on the conflict detection result. When the conflict detection result is the second conflict detection result, it indicates that there is a language difference or regional difference between the transmitting and receiving languages, or both language difference and regional difference exist simultaneously.
[0041] In step S30, based on the conflict detection result, the semantic conversion pattern of the transmitted speech data is obtained, including: When the conflict detection result indicates that there is a language difference between the sending and receiving languages, the sent speech data will be semantically converted according to the sending language and the first standard language as the first semantic conversion mode for the sent speech data. The first standard language is the standard language of the language corresponding to the receiving language. When the conflict detection result indicates that there is a regional difference between the transmitting and receiving languages, the transmitted speech data will be semantically converted based on the transmitting language, the second standard language, and the receiving language. This will serve as the second semantic conversion mode for the transmitted speech data. The second standard language is the standard language of the language that both the transmitting and receiving languages correspond to. When the conflict detection result indicates that there are both language and regional differences between the transmitting and receiving languages, the transmitted speech data will be semantically converted based on the transmitting language, the third standard language, the fourth standard language, and the receiving language. This will serve as the third semantic conversion mode for the transmitted speech data. The third standard language is the standard language of the language corresponding to the transmitting language, and the fourth standard language is the standard language of the language corresponding to the receiving language.
[0042] The language differences in conflict detection results can include language type differences, regional differences, or both. When the conflict detection result indicates a language difference between the speaking and receiving languages, it means that the language the receiver is fluent in is the standard language of the corresponding language, and regional language conversion is not necessary. Therefore, semantic conversion of the transmitted speech data can be completed solely through the standard languages of the corresponding languages of the speaking and receiving languages, serving as the first semantic conversion mode for the transmitted speech data.
[0043] When the conflict detection result indicates a regional difference between the transmitting and receiving languages, it means that the transmitting language is the standard language of the corresponding language and no further conversion is needed. However, the receiving language is not the standard language and needs to be converted according to its corresponding standard language. In other words, semantic conversion of the transmitted speech data is performed using the transmitting language, the second standard language corresponding to the receiving language, and the receiving language. This semantic conversion method is then used as the second semantic conversion mode for the transmitted speech data.
[0044] In addition, when the conflict detection result shows that there are both language differences and regional differences between the sending and receiving languages, it means that neither the sending nor the receiving language is a standard language of the corresponding language, and both need to be converted to standard languages. That is, semantic conversion of the sent speech data is performed by using the sending language, the third standard language of the language corresponding to the sending language, the fourth standard language of the language corresponding to the receiving language, and the receiving language as the third semantic conversion mode for the sent speech data.
[0045] Through the three semantic conversion modes mentioned above, it is possible to effectively translate the transmitted voice data into an output message that can be accurately understood, taking into account the language and regional conditions of different main language interactors.
[0046] Next, step S4 0 is executed, in which the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language according to the semantic conversion mode.
[0047] After determining the above semantic conversion modes, we can adapt to different situations of the sending and receiving languages to translate the sent voice data into an output message that is semantically appropriate to the receiving language, so as to ensure that the voice receiver can accurately understand the meaning of the speaker after viewing the message.
[0048] In step S40, according to the semantic conversion mode, the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language, including: Convert the emitted voice data into semantic text corresponding to the emitting host language; Based on the semantic conversion mode, the transmitted semantic text is converted across languages to obtain the detailed semantic text corresponding to the receiving host language; The detailed semantic text is semantically simplified, and the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language.
[0049] When translating outgoing voice data, the outgoing voice data can be converted into an outgoing semantic text corresponding to the main outgoing language in advance. Wherein, when the outgoing language corresponding to the outgoing voice data is different from the main outgoing language, the translation is directly performed according to the understanding habits corresponding to the main outgoing language. For example, the English phrase "good good study, day day up", when the main outgoing language is Chinese and English is not a proficient language for the voice transmitter, can be directly translated into "好好学习,天天向上", and then detailed semantic conversion is performed to obtain the outgoing semantic text, that is, the large language model corresponding to the main outgoing language can further expand based on the translated text to accurately express the outgoing semantics of the translated text. For example, after "好好学习,天天向上" is expanded, the obtained outgoing semantic text is "In terms of learning spirit, one should maintain an open mind at any stage of life, be willing to learn, and pursue continuous self-improvement and progress". After the outgoing semantic text is obtained, cross-lingual semantic conversion is further implemented on the outgoing semantic text according to the corresponding semantic conversion mode, to obtain a detailed semantic text corresponding to the main receiving language. Finally, in order to simplify the translation, it is also necessary to perform semantic simplification processing on the detailed semantic text to obtain an output message with appropriate semantics in the main receiving language, and send the output message to the voice receiving end, so as to ensure that the voice transmitting end can accurately express its intention, and the voice receiving end can also accurately understand the intention of the voice transmitting end. That is, when the main outgoing language is Chinese, the English phrase "good good study, day day up" is spoken, and the main receiving language of the voice receiving end is English, further translation will be performed to obtain "study hard and make progress every day" or "study well and progress every day ", and of course, further simplification processing can also be performed based on the translated semantics.
[0050] For example, when someone says "Skate on thin ice" in English, meaning "walking on thin ice" in Chinese, due to differences in context, "walking on thin ice" in Chinese means "being extremely cautious in dangerous situations." However, if this "Skate on thin ice" is directly sent to a voice receiver whose primary language is English, it might be mistranslated as "being able to handle a dangerous situation skillfully." If the English translation of "being extremely cautious in dangerous situations" is standard, but the receiving language is a regional language, then the English translation needs to be further translated into the receiving language to ensure the receiver accurately understands the meaning of "walking on thin ice" when expressed in English. Of course, there could also be other situations where different primary languages correspond to different meanings.
[0051] The semantic conversion mode may include a first semantic conversion mode that performs semantic conversion on the spoken speech data based on the uttering host language and a first standard language.
[0052] When the semantic conversion mode is the first semantic conversion mode, cross-language semantic conversion is performed on the transmitted semantic text according to the semantic conversion mode to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the first semantic conversion mode, retrieve the first semantic difference text library between the issuing host language and the first standard language; Based on the first semantic difference text library, the conversion difference of the emitted semantic text is monitored. When a sent text statement that corresponds to the first difference text in the first semantic difference text library is detected, the sent text statement in the sent semantic text is removed to obtain the first remaining sent text. The first remaining sent text and the first difference text are then translated into detailed semantic text corresponding to the receiving host language by the first major language model. The first major language model is trained using language data corresponding to the first standard language.
[0053] Please see Figure 2When performing cross-language semantic conversion on the transmitted semantic text using the first semantic conversion model, we can first retrieve the first semantic difference text library between the transmitting host language and the first standard language. This library includes text data corresponding to each transmitting host language with corresponding conversion relationships, as well as standard text data corresponding to the first standard language. Therefore, we can first monitor the conversion difference texts present in the transmitted semantic text at the speech transmitter. When a difference is found, the corresponding first difference text is found based on the first semantic difference text library and replaced with that text to ensure the semantic accuracy of the entire transmitted semantic text. Then, to ensure the semantic fluency of the first remaining transmitted text after adding the first difference text, we can use the first large language model trained on the language data corresponding to the first standard language to translate the first remaining transmitted text and the first difference text into detailed semantic text corresponding to the receiving host language.
[0054] The semantic conversion mode may also include a second semantic conversion mode that performs semantic conversion on the transmitted speech data based on the transmitting main language, the second standard language, and the receiving main language.
[0055] When the semantic conversion mode is the second semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion according to the semantic conversion mode to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the second semantic conversion mode, retrieve the second semantic difference text library between the main language and the second standard language; Based on the second semantic difference text library, the conversion difference of the emitted semantic text is monitored. When a sent text statement that corresponds to the second difference text in the second semantic difference text library is detected in the sent semantic text, the sent text statement in the sent semantic text is removed to obtain the second remaining sent text; the second remaining sent text and the second difference text are translated into the second standard language corresponding to the second standard language through the second large language model, and the second large language model is trained through the language data corresponding to the second standard language. Retrieve a third semantic difference text library between the second standard language and the receiving host language; Based on the third semantic difference text library, the conversion difference of intermediate semantic long text is monitored; When a semantic text statement corresponding to the third difference text in the third semantic difference text library is detected in the intermediate semantic long text, the semantic text statement in the intermediate semantic long text is removed to obtain the third remaining transmitted text. The third remaining transmitted text and the third difference text are then translated into detailed semantic text corresponding to the receiving host language by the third major language model corresponding to the receiving host language. The third major language model is trained by the language data corresponding to the receiving host language.
[0056] Please see Figure 3 When performing cross-language semantic conversion on transmitted semantic text using the second semantic conversion model, we can first retrieve a second semantic difference text library between the transmitting subject language and the second standard language. This library includes text data corresponding to each transmitting subject language with corresponding conversion relationships, as well as standard text data corresponding to the second standard language. Therefore, we can first monitor the conversion difference texts present in the transmitted semantic text at the speech transmitter. When a difference is found, the corresponding second difference text is found based on the second semantic difference text library and replaced with that text to ensure the semantic accuracy of the entire transmitted semantic text. Then, to ensure the semantic fluency of the second remaining transmitted text after adding the second difference text, we can use a second language model trained on the language data corresponding to the second standard language to translate the second remaining transmitted text and the second difference text into an intermediate long semantic text corresponding to the second standard language.
[0057] Then, a third semantic difference text library can be retrieved between the second standard language and the receiving host language. This library includes text data corresponding to each receiving host language with corresponding conversion relationships, along with standard text data corresponding to the second standard language. Therefore, the conversion difference texts existing in the intermediate semantic long text can be monitored first. When a semantic text statement with a difference is found, the corresponding third difference text is found based on the third semantic difference text library and replaced with that statement to ensure the semantic accuracy of the entire intermediate semantic long text. Then, to ensure the semantic fluency of the third remaining transmitted text after adding the third difference text, a third language model trained on the language data corresponding to the receiving host language can be used to translate the third remaining transmitted text and the third difference text into detailed semantic text corresponding to the receiving host language.
[0058] The semantic conversion mode also includes a third semantic conversion mode that performs semantic conversion on the transmitted speech data based on the transmitting host language, the third standard language, the fourth standard language, and the receiving host language.
[0059] When the semantic conversion mode is the third semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion according to the semantic conversion mode to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the third semantic conversion mode, the semantic text is converted and detected according to the third standard language, and the semantic text is translated into the first intermediate semantic long text corresponding to the third standard language. The first intermediate semantic long text is converted and detected according to the fourth standard language, and then the first intermediate semantic long text is translated into the second intermediate semantic long text corresponding to the fourth standard language. The second intermediate semantic long text is converted and detected based on the receiving main language, and then translated into the detailed semantic text corresponding to the receiving main language.
[0060] When using the third semantic conversion mode for detailed semantic text conversion, in order to ensure the accuracy of the detailed semantic text corresponding to the receiving host language, the transmitted semantic text needs to be converted into the first intermediate semantic long text corresponding to the third standard language first. Then, the first intermediate semantic long text is converted into the second intermediate semantic long text corresponding to the fourth standard language. Finally, based on the receiving host language corresponding to the standard language, the second intermediate semantic long text is further converted and translated into the detailed semantic text corresponding to the receiving host language. This ensures that when the transmitted semantic text is transmitted to the voice receiver, it can be accurately understood by the voice receiver to express its true meaning.
[0061] Specifically, the process includes: performing conversion detection on the transmitted semantic text according to the third standard language and translating the transmitted semantic text into a first intermediate semantic long text corresponding to the third standard language; performing conversion detection on the first intermediate semantic long text according to the fourth standard language and translating the first intermediate semantic long text into a second intermediate semantic long text corresponding to the fourth standard language; and performing conversion detection on the second intermediate semantic long text according to the received host language and translating the second intermediate semantic long text into detailed semantic text corresponding to the received host language, including: Retrieve the fourth semantic difference text library between the main language and the third standard language; Based on the fourth semantic difference text library, the conversion difference of the emitted semantic text is monitored. When it is detected that there is a sent text statement in the sent semantic text that corresponds to the fourth difference text in the fourth semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the fourth remaining sent text. The fourth remaining sent text and the fourth difference text are translated into the first intermediate semantic long text corresponding to the third standard language through the fourth major language model corresponding to the third standard language. The fourth major language model is trained through the language data corresponding to the third standard language. Retrieve the fifth semantic difference text library between the third and fourth standard languages; Based on the fifth semantic difference text library, the conversion difference monitoring is performed on the first intermediate semantic long text; When the first semantic text statement in the first intermediate semantic long text is detected to be the first semantic text statement corresponding to the fifth semantic difference text library, the first semantic text statement in the first intermediate semantic long text is removed to obtain the fifth remaining text. The fifth remaining text and the fifth difference text are then translated into the second intermediate semantic long text corresponding to the fourth standard language through the fifth major language model corresponding to the fourth standard language. The fifth major language model is trained using the language data corresponding to the fourth standard language. Retrieve the sixth semantic difference text library between the fourth standard language and the receiving host language; Based on the sixth semantic difference text library, the conversion difference monitoring is performed on the second intermediate semantic long text; When a second semantic text statement corresponding to the sixth difference text in the sixth semantic difference text library is detected in the second intermediate semantic long text, the second semantic text statement in the second intermediate semantic long text is removed to obtain the sixth remaining transmitted text. The sixth remaining transmitted text and the sixth difference text are then translated into detailed semantic text corresponding to the receiving host language by the sixth major language model corresponding to the receiving host language. The sixth major language model is trained by the language data corresponding to the receiving host language.
[0062] Please see Figure 4 When performing cross-language semantic conversion on transmitted semantic text using the third semantic conversion model, a fourth semantic difference text library can be retrieved between the transmitting host language and the third standard language. This library includes text data corresponding to each transmitting host language with corresponding conversion relationships, as well as standard text data corresponding to the fourth standard language. Therefore, the conversion difference texts present in the transmitted semantic text can be monitored first. When a difference is found, the corresponding fourth difference text is found based on the fourth semantic difference text library and replaced to ensure the semantic accuracy of the entire transmitted semantic text. Then, to ensure the semantic fluency of the remaining transmitted text after adding the fourth difference text, a fourth language model trained on the language data corresponding to the third standard language can be used to translate the remaining transmitted text and the fourth difference text into a first intermediate semantic long text corresponding to the third standard language.
[0063] Then, a fifth semantic difference text library can be retrieved between the third and fourth standard languages. This library includes standard text data corresponding to each of the third and fourth standard languages with corresponding conversion relationships. Therefore, the conversion difference texts existing in the first intermediate semantic long text can be monitored first. When a difference in the first semantic text statement is found, the fifth difference text corresponding to the first semantic text statement is found based on the fifth semantic difference text library and replaced to ensure the semantic accuracy of the entire first intermediate semantic long text. Then, to ensure the semantic fluency of the fifth remaining text after adding the fifth difference text, a fifth language model trained on the language data corresponding to the fourth standard language can be used to translate the fifth remaining text and the fifth difference text into a second intermediate semantic long text corresponding to the fourth standard language.
[0064] Subsequently, a sixth semantic difference text library can be retrieved between the fourth standard language and the receiving host language. This library includes text data corresponding to each receiving host language with corresponding conversion relationships, along with standard text data corresponding to the fourth standard language. Therefore, the conversion difference texts existing in the second intermediate semantic long text can be monitored first. When a difference in the second semantic text statement is found, the sixth difference text corresponding to the second semantic text statement is found based on the sixth semantic difference text library and replaced to ensure the semantic accuracy of the entire second intermediate semantic long text. Then, to ensure the semantic fluency of the sixth remaining transmitted text after adding the sixth difference text, a sixth major language model trained on the language data corresponding to the receiving host language can be used to translate the sixth remaining transmitted text and the sixth difference text into detailed semantic text corresponding to the receiving host language.
[0065] Next, step S50 is executed, sending the output message to the voice receiver corresponding to the receiving language. When the user at the voice receiver receives the output message, they can accurately identify the semantic intent of the other user who sent the message based on the output message, thus solving the problem of semantic understanding bias.
[0066] Please see Figure 5The present invention also provides a speech recognition intelligent customer service robot 11 applicable to multiple languages, comprising: a data acquisition unit 111, used to acquire voice data sent from a voice transmitter to a voice receiver; a conflict detection unit 112, used to perform direct semantic conversion conflict detection on the voice data according to the transmitting host language corresponding to the voice data and the receiving host language corresponding to the voice receiver, and obtain a conflict detection result; a pattern acquisition unit 113, used to acquire a semantic conversion pattern of the voice data according to the conflict detection result; a message conversion unit 114, used to translate the voice data into an output message that is semantically appropriate to the receiving host language according to the semantic conversion pattern; and a message sending unit 115, used to send the output message to the voice receiver corresponding to the receiving host language.
[0067] It should be noted that the multilingual speech recognition intelligent customer service robot 11 provided in the above embodiments and the multilingual speech recognition method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the multilingual speech recognition intelligent customer service robot 11 provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation here.
[0068] Please see Figure 6 The electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a speech recognition program suitable for multiple languages.
[0069] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 12 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 1. Furthermore, the memory 12 can include both internal and external storage units of the electronic device 1. The memory 12 can be used not only to store application software and various types of data installed on the electronic device 1, such as code for multilingual speech recognition, but also to temporarily store data that has been output or will be output.
[0070] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the electronic device 1, connecting various components of the electronic device 1 through various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., speech recognition programs suitable for multiple languages) and calls data stored in the memory 12 to perform various functions of the electronic device 1 and process data.
[0071] The processor 13 executes the operating system of the electronic device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-described speech recognition method applicable to multiple languages.
[0072] For example, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into units within a multilingual voice recognition intelligent customer service robot.
[0073] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The software functional module, stored in the storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute some functions of the speech recognition method applicable to multiple languages described in the various embodiments of this application.
[0074] In summary, the speech recognition method and intelligent customer service robot disclosed in this invention, applicable to multiple languages, first determines whether there is a semantic conflict between the main languages of the speaker and receiver when the speaker sends voice data to the receiver. Based on the conflict detection results, it determines whether to directly perform semantic conversion on the sent voice data. If direct semantic conversion is not possible, an appropriate semantic conversion mode can be selected based on the differences between the main languages of the speaker and receiver to accurately translate the sent voice data into the message content that the speaker actually wants to express. The semantics of this message content should also conform to the understanding habits of the receiver. This allows the message content to be sent to the receiver's main language. This facilitates the conversion of sent voice data into semantic content that conforms to the language habits of the receiver in multilingual and multi-habit multi-segment voice interactions. It effectively avoids the problem of misunderstanding when different receivers with different language spans understand the translated voice data due to direct translation of the sent voice data, which can cause interaction barriers when interacting with voice data in different languages. Therefore, this invention effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0075] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A speech recognition method applicable to multiple languages, characterized in that, include: Acquire the voice data sent from the voice transmitter to the voice receiver; Based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiving end, a conflict detection is performed on the voice data through direct semantic conversion to obtain a conflict detection result. Based on the conflict detection results, obtain the semantic conversion pattern of the emitted voice data; According to the semantic conversion mode, the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language; The output message is sent to the voice receiving terminal corresponding to the receiving main language; Based on the conflict detection results, the semantic conversion pattern of the transmitted voice data is obtained, including: When the conflict detection result indicates that there is a language difference between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language and the first standard language as the first semantic conversion mode for the sent speech data. The first standard language is the standard language of the language corresponding to the receiving language. When the conflict detection result indicates that there is a regional difference between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language, the second standard language, and the receiving language as the second semantic conversion mode for the sent speech data. The second standard language is the standard language of the language that both the sending language and the receiving language correspond to. When the conflict detection result indicates that there are both language differences and regional differences between the sending language and the receiving language, the sent speech data will be semantically converted according to the sending language, the third standard language, the fourth standard language, and the receiving language as the third semantic conversion mode for the sent speech data. The third standard language is the standard language of the language corresponding to the sending language, and the fourth standard language is the standard language of the language corresponding to the receiving language.
2. The speech recognition method applicable to multiple languages according to claim 1, characterized in that, Also includes: Acquire historical dialogue data prior to the transmission of voice data, wherein the historical dialogue data includes first historical output data corresponding to the voice transmitting end and second historical output data corresponding to the voice receiving end; Based on the historical dialogue data, language analysis is performed on the voice transmitter and the voice receiver to determine a first set of proficient languages corresponding to the voice transmitter and a second set of proficient languages corresponding to the voice receiver. The first set of proficient languages includes multiple first proficient languages, and the second set of proficient languages includes multiple second proficient languages. The emitted voice data is subjected to language analysis to obtain the corresponding emitted language; Determine whether the language spoken is within the first set of proficient languages; If so, the first proficient language corresponding to the first proficient language set shall be the primary language of the speech output end; If not, then the first proficient language corresponding to the highest proficiency in the first proficient language set shall be the primary language of the speech output end. Determine whether a second proficient language corresponding to the spoken language exists in the second proficient language set. If so, the second proficient language corresponding to the sending main language shall be used as the receiving main language; If not, then the second proficiency language corresponding to the maximum proficiency level in the second proficiency language set shall be the receiving primary language of the speech transmitter.
3. The speech recognition method applicable to multiple languages according to claim 1, characterized in that, Based on the transmitting language corresponding to the transmitted voice data and the receiving language corresponding to the voice receiver, conflict detection is performed on the transmitted voice data through direct semantic conversion to obtain conflict detection results, including: Determine whether the sending language and the receiving language are consistent; If so, the transmitted voice data will be directly semantically converted and used as the first conflict detection result; If not, the language difference between the sending language and the receiving language is taken as the second conflict detection result, and the language difference includes at least one of language difference and regional difference; The determination of whether the sending language and the receiving language are consistent includes: Determine whether the sending language and the receiving language simultaneously satisfy the following conditions: The languages of the sending and receiving main languages are the same; The sending and receiving languages are from the same region; If so, it means that the sending language and the receiving language are the same; If not, it means that the sending language and the receiving language are inconsistent.
4. The speech recognition method applicable to multiple languages according to claim 1, characterized in that, According to the semantic conversion mode, translating the transmitted voice data into an output message that is semantically appropriate to the receiving host language includes: Convert the emitted voice data into emitted semantic text corresponding to the emitting main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language; The detailed semantic text is semantically simplified, and the transmitted voice data is translated into an output message that is semantically appropriate to the receiving host language.
5. The speech recognition method applicable to multiple languages according to claim 4, characterized in that, The semantic conversion mode includes a first semantic conversion mode that performs semantic conversion on the emitted speech data based on the emitting main language and the first standard language; Based on the semantic conversion mode and the pre-existing dialogue data prior to the transmitted voice data, cross-language semantic conversion is performed on the transmitted semantic text to obtain detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the first semantic conversion mode, the first semantic difference text library between the issuing host language and the first standard language is retrieved; Based on the first semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the first difference text in the first semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the first remaining sent text; the first remaining sent text and the first difference text are translated into the detailed semantic text corresponding to the receiving host language by the first large language model, the first large language model being trained by language data corresponding to the first standard language.
6. The speech recognition method applicable to multiple languages according to claim 4, characterized in that, The semantic conversion mode includes a second semantic conversion mode that performs semantic conversion on the transmitted speech data based on the transmitting main language, the second standard language, and the receiving main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the second semantic conversion mode, the second semantic difference text library between the issuing main language and the second standard language is retrieved; Based on the second semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the second difference text in the second semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the second remaining sent text; the second remaining sent text and the second difference text are translated into the second standard language corresponding to the second standard language through the second large language model, and the second large language model is trained through the language data corresponding to the second standard language. Retrieve a third semantic difference text library between the second standard language and the receiving host language; Based on the third semantic difference text library, the conversion difference monitoring is performed on the intermediate semantic long text; When a semantic text statement corresponding to the third difference text in the third semantic difference text library is detected in the intermediate semantic long text, the semantic text statement in the intermediate semantic long text is removed to obtain the third remaining transmitted text; the third remaining transmitted text and the third difference text are translated into the detailed semantic text corresponding to the receiving host language through the third major language model corresponding to the receiving host language, and the third major language model is trained through the language data corresponding to the receiving host language.
7. The speech recognition method applicable to multiple languages according to claim 4, characterized in that, The semantic conversion mode includes a third semantic conversion mode that performs semantic conversion on the transmitted voice data based on the transmitting main language, the third standard language, the fourth standard language, and the receiving main language; According to the semantic conversion mode, the transmitted semantic text is subjected to cross-language semantic conversion to obtain the detailed semantic text corresponding to the receiving host language, including: When the semantic conversion mode is the third semantic conversion mode, the emitted semantic text is converted and detected according to the third standard language, and the emitted semantic text is translated into the first intermediate semantic long text corresponding to the third standard language; The first intermediate semantic long text is converted and detected according to the fourth standard language, and the first intermediate semantic long text is translated into the second intermediate semantic long text corresponding to the fourth standard language. The second intermediate semantic long text is converted and detected according to the receiving host language, and then translated into the detailed semantic text corresponding to the receiving host language.
8. The speech recognition method applicable to multiple languages according to claim 7, characterized in that, The transmitted semantic text is converted and detected according to the third standard language, and translated into a first intermediate semantic long text corresponding to the third standard language; the first intermediate semantic long text is converted and detected according to the fourth standard language, and translated into a second intermediate semantic long text corresponding to the fourth standard language; the second intermediate semantic long text is converted and detected according to the receiving host language, and translated into the detailed semantic text corresponding to the receiving host language, including: Retrieve a fourth semantic difference text library between the main language and the third standard language; Based on the fourth semantic difference text library, the conversion difference of the emitted semantic text is monitored; When it is detected that there is a sent text statement in the sent semantic text that corresponds to the fourth difference text in the fourth semantic difference text library, the sent text statement in the sent semantic text is removed to obtain the fourth remaining sent text. The fourth remaining sent text and the fourth difference text are translated into the first intermediate semantic long text corresponding to the third standard language through the fourth major language model corresponding to the third standard language. The fourth major language model is trained through the language data corresponding to the third standard language. Retrieve the fifth semantic difference text library between the third standard language and the fourth standard language; Based on the fifth semantic difference text library, the first intermediate semantic long text is subjected to conversion difference monitoring. When a first semantic text statement corresponding to the fifth difference text in the fifth semantic difference text library is detected in the first intermediate semantic long text, the first semantic text statement in the first intermediate semantic long text is removed to obtain the fifth remaining text. The fifth remaining text and the fifth difference text are then translated into the second intermediate semantic long text corresponding to the fourth standard language through the fifth major language model corresponding to the fourth standard language. The fifth major language model is trained using the language data corresponding to the fourth standard language. Retrieve the sixth semantic difference text library between the fourth standard language and the receiving host language; Based on the sixth semantic difference text library, the conversion difference monitoring is performed on the second intermediate semantic long text; When a second semantic text statement corresponding to the sixth difference text in the sixth semantic difference text library is detected in the second intermediate semantic long text, the second semantic text statement in the second intermediate semantic long text is removed to obtain the sixth remaining transmitted text. The sixth remaining transmitted text and the sixth difference text are then translated into the detailed semantic text corresponding to the receiving host language through the sixth major language model corresponding to the receiving host language. The sixth major language model is trained using the language data corresponding to the receiving host language.
9. An intelligent customer service robot applying the speech recognition method applicable to multiple languages as described in any one of claims 1-8, characterized in that, include: The data acquisition unit is used to acquire the voice data sent from the voice transmitter to the voice receiver. The conflict detection unit is used to perform direct semantic conversion on the voice data based on the sending language corresponding to the voice data and the receiving language corresponding to the voice receiver, and obtain the conflict detection result. The pattern acquisition unit is used to acquire the semantic conversion pattern of the transmitted voice data based on the conflict detection result. The message conversion unit is used to translate the transmitted voice data into an output message that is semantically appropriate to the receiving host language according to the semantic conversion mode. The message sending unit is used to send the output message to the voice receiving terminal corresponding to the receiving main language.
Citation Information
Patent Citations
Information processing method and device for voice conversion
CN110442881A