Customer service method, system and device and readable storage medium

By querying customer call history and acquiring voice data in real time, and utilizing a multilingual interactive platform and a third-party voice service engine for instant transcription and translation, the system solves the problems of poor real-time communication and untimely recognition of common languages ​​in existing technologies, thus achieving efficient, convenient, and instant feedback in the customer service system.

CN121397145APending Publication Date: 2026-01-23CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511543565.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

The existing customer service system suffers from poor real-time communication and an inability to promptly recognize customers' common language during the translation process, resulting in a poor customer experience.

Method used

By querying customers' historical call information to determine their preferred language, the system acquires customers' voices in real time and uses a multilingual interaction platform and a third-party voice service engine to perform instant transcription, automatic sentence segmentation, and multilingual translation, thereby achieving real-time voice feedback.

Benefits of technology

It improves the real-time nature of communication and customer experience, ensuring that customers can receive immediate feedback during the conversation, and significantly enhances the convenience and efficiency of the service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397145A_ABST
    Figure CN121397145A_ABST
Patent Text Reader

Abstract

The invention provides a customer service method, system and device and a readable storage medium, and the method comprises the steps: querying the historical incoming call information of an incoming call number corresponding to a call request according to the call request of a client, so as to determine the habitual language of a customer; further confirming whether the customer service language currently selected by the customer is non-Chinese or not in response to the condition that the queried habitual language is non-Chinese or the habitual language of the customer is not queried; and if the customer service language currently selected by the customer is non-Chinese, acquiring customer voice in real time in a voice conversation process between the client and the corresponding first seat end, and forwarding the customer voice to a third-party voice service engine through the multi-language interaction platform. According to the method, the system, the device and the readable storage medium, the problems of poor communication real-time performance, poor customer experience and poor customer experience caused by incapability of timely identifying whether the conventional language of the customer is Chinese or not due to the fact that translation can be carried out after the customer finishes speaking in the existing customer service method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of communication technology, and in particular to a customer service method, system, device and readable storage medium. BACKGROUND

[0002] With the acceleration of globalization, international exchanges are becoming more frequent, and various problems encountered by foreign friends in their activities such as life, work and study in China are also increasing. In order to better serve foreign friends and improve the international level of service, it is particularly important to establish an efficient, convenient and foreign language supported service hotline system.

[0003] To achieve the above-mentioned efficient, convenient and foreign language supported service hotline system, there are mainly two kinds of existing customer service solutions: one is that the customer uses translation equipment, translation software or translation personnel to realize it; the other is to add a translation server between the customer and the agent end. However, these two solutions at least have the following technical defects:

[0004] (1) The existing technologies all need to wait for the customer to finish speaking (generate the first audio and video data) before translation, and if the user speaks for a long time (such as more than 30 seconds), the subsequent process of translation, agent reply and translation again will cause a long delay, resulting in poor real-time communication and poor customer experience.

[0005] (2) The existing technologies cannot timely identify whether the customer's habitual language is Chinese, and often need the customer service terminal (agent end) to hear the customer's voice and confirm that the customer uses non-Chinese before initiating the barrier-free call request, which will significantly reduce the customer's experience. SUMMARY

[0006] The present disclosure provides a customer service method, system, device and readable storage medium.

[0007] In a first aspect, the present disclosure provides a customer service method applied to a traffic platform, the method comprising:

[0008] According to the call request of the client, the historical call information of the call request corresponding incoming number is queried to determine the habitual language of the customer;

[0009] In response to the habitual language being non-Chinese or the habitual language of the customer not being queried, it is further confirmed whether the customer service language currently selected by the customer is non-Chinese;

[0010] If the current selected service language of the customer is not Chinese, during the voice conversation between the client and the first agent terminal, the customer voice is acquired in real time, and the customer voice is forwarded to the third-party voice service engine through the multi-language interaction platform, so that the third-party voice service engine transcribes the customer voice into text in the corresponding first language in real time, and performs automatic sentence segmentation during the transcription process, and after each sentence is segmented, the corresponding target text is translated into text in the second language which is the habitual language of the first agent terminal, and is further synthesized into voice in the second language;

[0011] The synthesized voice in the first language is sent to the client, wherein the synthesized voice in the first language is obtained by translating and synthesizing the voice or text replied by the first agent terminal after the third-party voice service engine forwards the text and / or voice in the second language to the first agent terminal through the multi-language interaction platform;

[0012] The synthesized voice in the first language is sent to the client, wherein the synthesized voice in the first language is obtained by translating and synthesizing the voice or text replied by the first agent terminal after the third-party voice service engine forwards the text and / or voice in the second language to the first agent terminal through the multi-language interaction platform;

[0013] Further, in response to the query that the habitual language is not Chinese or the habitual language of the customer is not queried, it is further confirmed whether the current selected service language of the customer is not Chinese, and specifically includes:

[0014] In response to the query that the habitual language is not Chinese or the habitual language of the customer is not queried, the customer is prompted to select the service language by voice;

[0015] If the selected service language of the customer is Chinese, it is confirmed that the current selected service language of the customer is Chinese;

[0016] If the selected service language of the customer is other, it is confirmed that the current selected service language of the customer is not Chinese;

[0017] Before the client and the first agent terminal are matched and the call is established, the method further includes:

[0018] The client is matched with the first agent terminal and the call is established.

[0019] Further, the method further includes:

[0020] In response to the query that the habitual language is Chinese or the current selected service language of the customer is Chinese, the client is matched with the second agent terminal and the call is established.

[0021] Further, the method further includes at least one of the following:

[0022] The voice activity detection (VAD) is used to detect whether the voice of the client or the first agent terminal is complete;

[0023] In response to the end of the voice dialogue, the habitual language of the customer is obtained from the third-party voice service engine through the multi-language interaction platform and stored, wherein the habitual language of the customer is the first language;

[0024] During waiting for the synthesized voice in the first language, the original voice or preset voice prompt of the first agent end is played to the customer.

[0025] In the second aspect, the disclosure provides a customer service method applied to a multi-language interaction platform, and the method comprises:

[0026] The customer voice sent by the traffic platform is received, wherein the customer voice is obtained by the traffic platform according to the call request of the client, querying the historical call information of the incoming call number corresponding to the call request to determine the habitual language of the customer, and in response to the fact that the habitual language obtained by the query is not Chinese or the habitual language of the customer is not queried, further confirming that the customer currently selected service language is not Chinese, and the customer voice is obtained and sent in real time during the voice dialogue process between the client and the corresponding first agent end.

[0027] The customer voice is forwarded to the third-party voice service engine, so that the third-party voice service engine transcribes the customer voice into text in the corresponding first language in real time, and performs automatic sentence segmentation during the transcription process, translates the corresponding target text into text in the second language which is the habitual language of the first agent end after each sentence is segmented, and further synthesizes the text into voice in the second language;

[0028] The text and / or voice in the second language sent by the third-party voice service engine is received;

[0029] The text and / or voice in the second language is sent to the first agent end, and the voice or text replied by the first agent end is forwarded to the third-party voice service engine, so that the third-party voice service engine translates and synthesizes the voice or text replied by the first agent end to obtain synthesized voice in the first language;

[0030] The synthesized voice in the first language sent by the third-party voice service engine is received, and the synthesized voice in the first language is sent to the client through the traffic platform.

[0031] Further, the method further comprises:

[0032] The real-time voice stream of the client and the first agent end is sent to the third-party voice service engine, so that the third-party voice service engine transcribes the real-time voice stream into text in the corresponding language, translates the text in the corresponding language into text in the third language which is the habitual language of the listener, and further synthesizes the text into voice in the third language;

[0033] receive the synthesized third language voice sent by the third party voice service engine, and send the synthesized third language voice to the listening end.

[0034] In a third aspect, the disclosure provides a customer service method applied to a third party voice service engine, the method comprising:

[0035] receiving a customer voice forwarded by a multi-language interaction platform, the customer voice being obtained by a traffic platform in response to a call request from a client, querying historical call information of an incoming call number corresponding to the call request to determine the habitual language of the customer, and further confirming that the customer currently selected service language is non-Chinese when the habitual language is non-Chinese or the habitual language of the customer is not queried;

[0036] transcribing the customer voice into text in a corresponding first language, and performing automatic sentence segmentation during the transcription process, translating the corresponding target text into text in a second language habitually used by a first agent end after each sentence is segmented, and further synthesizing the text in the second language into voice in the second language;

[0037] forwarding the text and / or voice in the second language to the first agent end through the multi-language interaction platform, receiving voice or text replied by the first agent end sent by the multi-language interaction platform, and translating and synthesizing the voice or text replied by the first agent end to obtain synthesized voice in the first language;

[0038] sending the synthesized voice in the first language to the client through the multi-language interaction platform and the traffic platform.

[0039] Further, the transcribing the customer voice into text in a corresponding first language and performing automatic sentence segmentation during the transcription process specifically comprises:

[0040] identifying the language corresponding to the customer voice through a pre-set large model, and taking the language as the first language;

[0041] transcribing the customer voice into text in the first language, and extracting voice features and text features during the transcription process, inputting the extracted voice features and text features into a trained sentence segmentation model for sentence segmentation, wherein the trained sentence segmentation model performs sentence segmentation according to voice pause rules, tone change rules, grammar rules and punctuation rules.

[0042] Further, the translating and synthesizing the voice or text replied by the first agent end to obtain synthesized voice in the first language specifically comprises:

[0043] transcribe the voice of the reply into text in a corresponding language, and translate the text in the corresponding language into text in the first language;

[0044] translate the text of the reply into text in the first language;

[0045] synthesize the translated text in the first language, and model the prosodic features of the voice, introduce emotion labels and style parameters in the process of synthesis, and adopt a preset vocoder technology.

[0046] In a fourth aspect, the present disclosure provides a customer service system, comprising a traffic platform, a multi-language interaction platform and a third-party voice service engine;

[0047] The traffic platform is configured to perform the customer service method of the first aspect;

[0048] The multi-language interaction platform is configured to perform the customer service method of the second aspect;

[0049] The third-party voice service engine is configured to perform the customer service method of the third aspect.

[0050] In a fifth aspect, the present disclosure provides a customer service device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the customer service method of the first aspect or the second aspect or the third aspect.

[0051] In a sixth aspect, the present disclosure provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the customer service method of the first aspect or the second aspect or the third aspect.

[0052] The embodiments provided by the present disclosure determine the habitual language or further confirm the customer's current selected customer service language by querying the customer's historical incoming information, accurately identify the customer's language demand in advance, and ensure smooth communication. At the same time, by acquiring the customer's voice in real time, with the cooperation of the multi-language interaction platform and the third-party voice service engine, the customer's voice is transcribed, automatically segmented, and quickly translated and synthesized between multiple languages, so that the customer can obtain feedback in real time during the conversation, effectively ensuring the real-time performance of voice interaction, and significantly improving the customer experience. The existing customer service method needs to wait for the customer to finish speaking before translating, which leads to poor real-time communication and poor customer experience, and cannot identify whether the customer's habitual language is Chinese in time, which leads to poor customer experience.

[0053] It should be understood that nothing in this section is intended to limit the scope of the embodiments of the present disclosure. Other aspects of the present disclosure will become apparent to those skilled in the art upon reading the following specification and appended claims. BRIEF DESCRIPTION OF DRAWINGS

[0054] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and serve to explain the principles of the present disclosure, and are not intended to limit the disclosure. The above and other features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which:

[0055] Figure 1 A flow chart of a customer service method of Embodiment 1 of the present disclosure;

[0056] Figure 2 An architecture diagram of a customer service system of the present disclosure;

[0057] Figure 3 A flow chart of automatic punctuating of the present disclosure;

[0058] Figure 4 A flow chart of monitoring of the present disclosure;

[0059] Figure 5 A flow chart of a customer service method of Embodiment 2 of the present disclosure;

[0060] Figure 6 A flow chart of a customer service method of Embodiment 3 of the present disclosure;

[0061] Figure 7 A structure diagram of a customer service system of Embodiment 4 of the present disclosure;

[0062] Figure 8 A structure diagram of a customer service device of Embodiment 5 of the present disclosure. DETAILED DESCRIPTION

[0063] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered merely as exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0064] The embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0065] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0066] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. "Coupled" or "connected" or similar terms are not restricted to physical or mechanical connections or associations, but can also include electrical connections, whether direct or indirect.

[0067] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.

[0068] Embodiment 1:

[0069] The embodiment provides a customer service method applied to a traffic platform, as shown in the accompanying drawings, the method comprises: Figure 1 The method comprises:

[0070] Step S101: querying historical call-in information of an incoming call number corresponding to a call request of a client according to the call request to determine a habitual language of the client.

[0071] In the embodiment, a client calls a traffic platform through a client, the traffic platform receives a call request of the client, and queries historical call-in information of an incoming call number, the historical call-in information comprising a habitual language.

[0072] Step S102: in response to the habitual language being non-Chinese or the habitual language of the client not being queried, further confirming whether a customer service language currently selected by the client is non-Chinese.

[0073] In the embodiment, if the habitual language is non-Chinese or the habitual language of the client is not queried, the client is further prompted to confirm whether the customer service language currently selected is non-Chinese.

[0074] Further, the further confirming whether the customer service language currently selected by the client is non-Chinese in response to the habitual language being non-Chinese or the habitual language of the client not being queried comprises:

[0075] In response to the query that the habitual language is not Chinese or the habitual language of the customer is not queried, the voice prompts the customer to select the service language;

[0076] If the customer selects the service language as Chinese, it is confirmed that the service language currently selected by the customer is Chinese.

[0077] If the customer selects the service language as other, it is confirmed that the service language currently selected by the customer is not Chinese.

[0078] In the embodiment, the service language currently selected by the customer can be confirmed by voice prompt. For example, the call platform can play English and Chinese voice to prompt the customer to select the service language, such as "Please select the language of customer service, Chinese please press 1, other please press 2; Please select the language of customer service, Chinese please press 1, other please press 2". Then the service language currently selected by the customer is confirmed according to the key result of the customer.

[0079] Step S103: If the service language currently selected by the customer is not Chinese, the customer voice is acquired in real time during the voice dialogue process between the client and the corresponding first agent end, and the customer voice is forwarded to the third-party voice service engine through the multi-language interaction platform, so that the third-party voice service engine transcribes the customer voice into text in the corresponding first language in real time, and performs automatic sentence segmentation during the transcription process. After completing a sentence, the corresponding target text is translated into text in the second language which is the habitual language of the first agent end, and is further synthesized into voice in the second language.

[0080] In the embodiment, if the service language currently selected by the customer is not Chinese, the first agent end is matched for the client and the call is established, and then the customer voice is acquired in real time during the voice dialogue process between the client and the corresponding first agent end, and the customer voice is forwarded to the third-party voice service engine through the multi-language interaction platform.

[0081] It should be noted that the call platform is a core component of the system, and the functional expansion of the call platform has high implementation complexity and may affect the original stability. Therefore, in order to avoid interference and ensure stability, a multi-language interaction platform for calling the third-party voice service engine is specially set up, so that the original call platform does not need to be newly added or modified in function.

[0082] In the embodiment, after the third-party voice service engine receives the customer voice forwarded by the multi-language interaction platform, the customer voice is transcribed into text in a corresponding first language in real time, and automatic sentence segmentation is performed during the transcription process. After each sentence is segmented, the corresponding target text is translated into text in a second language commonly used by the first agent, and is further synthesized into voice in the second language.

[0083] Optionally, the transcription of the customer voice into text in the first language and the automatic sentence segmentation during the transcription process specifically include:

[0084] The language corresponding to the customer voice is identified by a preset large model, and is taken as the first language;

[0085] The customer voice is transcribed into text in the first language, and voice features and text features are extracted during the transcription process. The extracted voice features and text features are input into a trained sentence segmentation model for sentence segmentation. The trained sentence segmentation model performs sentence segmentation according to voice pause rules, tone change rules, grammar rules, and punctuation rules.

[0086] In the embodiment, the third-party voice service engine identifies the language corresponding to the customer voice (i.e., the first language) by a large model, and transcribes it into corresponding text. During the transcription process, a trained sentence segmentation model is used for sentence segmentation. The trained sentence segmentation model is trained in advance using annotated training data according to certain voice pause rules, tone change rules, grammar rules, and punctuation rules. It should be noted that the traditional automatic sentence segmentation model training only considers the rules on the text, i.e., the grammar rules and punctuation rules. In order to make the sentence segmentation more reasonable, the present disclosure also adds the rules on the voice, i.e., the voice pause rules and the tone change rules.

[0087] The sentence segmentation model uses VAD (Voice activity detection) to detect the voice pause rules, which include natural pauses and pauses corresponding to punctuation marks. Tone changes are usually related to features such as pitch, volume, and duration of the voice signal. The corresponding tone change rules include questioning tone (when the speaker asks a question, the tone usually rises), emphasis tone (when the speaker emphasizes, the speaker may raise the volume, speed up the speech, or prolong certain syllables), and emotional tone (different emotional expressions are also accompanied by changes in tone).

[0088] Specifically, the third-party voice service engine extracts voice features and text features from the customer voice and the transcribed text respectively during the transcription process, the voice features are, for example, pause position, pause duration, tone change features, and the text features are, for example, punctuation marks, words, and parts of speech, and then inputs the extracted voice features and text features into the trained punctuation model to punctuate the text.

[0089] Optionally, the method further comprises:

[0090] In response to the fact that the habitual language is Chinese or the fact that the customer currently selects the customer service language as Chinese, the second agent terminal is matched for the client terminal, and a call is established.

[0091] In this embodiment, if the habitual language is Chinese or the customer currently selects the customer service language as Chinese, the second agent terminal is matched for the client terminal, and the second agent terminal establishes a call with the client terminal to directly perform voice conversation.

[0092] Step S104: receiving the synthesized first language voice sent by the multilingual interaction platform, wherein the synthesized first language voice is obtained by translating and synthesizing the voice or text replied by the first agent terminal after the third-party voice service engine forwards the second language text and / or voice to the first agent terminal through the multilingual interaction platform.

[0093] In this embodiment, the third-party voice service engine sends the second language text and / or voice to the first agent terminal through the multilingual interaction platform. After the first agent terminal receives the second language text and / or voice, the first agent terminal replies through text or voice and sends the reply to the third-party voice service engine through the multilingual interaction platform. The third-party voice service engine receives the voice or text replied by the first agent terminal sent by the multilingual interaction platform, and translates and synthesizes the voice or text replied by the first agent terminal to obtain the synthesized first language voice.

[0094] Optionally, the translation and synthesis of the voice or text replied by the first agent terminal to obtain the synthesized first language voice specifically comprises:

[0095] For the voice replied by the first agent terminal, the replied voice is transcribed into text in the corresponding language, and the text in the corresponding language is translated into first language text;

[0096] For the text replied by the first agent terminal, the replied text is translated into first language text;

[0097] The translated first language text is synthesized, and the prosodic features of the voice are modeled, emotion labels and style parameters are introduced during the synthesis, and a preset vocoder technology is used.

[0098] In this embodiment, in order to improve the quality and expressiveness of synthesized speech, fine prosody modeling, emotion and style control, high-quality vocoder and other technologies can be used in the synthesis process to provide users with a better voice interaction experience.

[0099] (1) Fine prosody modeling: fine modeling of prosodic features of speech, such as predicting the fundamental frequency and energy at the phoneme level, correlating duration prediction with prosodic components such as fundamental frequency and energy, and using an autoregressive structure for fine-grained modeling to make the synthesized speech more natural and smooth.

[0100] (2) Emotion and style control: by introducing emotion labels, style parameters and other methods, the model can synthesize speech with specific emotions and styles according to the needs. For example, the emotion CSS model based on heterogeneous graph context modeling can learn emotional cues in the dialogue context and infer the emotional style of the current speech through the emotion rendering module to realize emotion speech synthesis.

[0101] (3) High-quality vocoder: advanced vocoder technology such as deep learning-based vocoder can more accurately convert acoustic features into natural speech waveforms, improving the quality and expressiveness of synthesized speech. The acoustic features are extracted from the original speech of the customer, including the spectral features, fundamental frequency, duration, and energy of the speech. Based on the spectral features, the "tone" details are restored, combined with the fundamental frequency and duration to adjust the "pitch" and "rhythm", and the energy corresponds to the "volume" restoration.

[0102] It should be noted that the conventional speech synthesis method used by the media server in the prior art often results in insufficient quality and expressiveness of the speech received by the customer. However, by using fine prosody modeling, targeted emotion and style control, and high-quality vocoder, the embodiment can significantly improve the quality and expressiveness of synthesized speech, thereby effectively enhancing the customer's interaction experience.

[0103] Step S105: sending the synthesized speech of the first language to the client.

[0104] In this embodiment, the third-party voice service engine sends the synthesized speech of the first language to the client through the multi-language interaction platform and the traffic platform for playback and the next round of interaction.

[0105] Optionally, the method further comprises at least one of the following:

[0106] detecting whether the speech of the client or the first agent end is complete based on voice activity detection (VAD);

[0107] In response to the end of the voice dialogue, the habitual language of the customer is obtained from the third-party voice service engine through the multi-language interaction platform, and is stored, wherein the habitual language of the customer is the first language;

[0108] During the waiting for the synthesized voice in the first language, the original voice or preset voice prompt of the first agent end is played to the customer.

[0109] In the embodiment, the voice activity detection (VAD) is used to detect whether the voice of the customer end / agent (the voice of the first agent end or the voice of the second agent end) is finished, and the next step is jumped to after the voice is finished. The basic principle of VAD is to analyze various characteristics of the voice signal to determine whether the current audio frame contains human voice.

[0110] In the embodiment, the habitual language of the customer is obtained from the third-party voice service engine through the multi-language interaction platform after the call is ended, and is stored in the database, so that the habitual language of the customer can be quickly obtained when the customer calls in next time.

[0111] In the embodiment, when the customer finishes speaking, the customer cannot hear the sound feedback before the reply of the agent is transcribed, translated and synthesized (i.e. during the waiting for the synthesized voice in the first language), which causes the service to be not smooth, and the customer may think that the service is interrupted. To solve the problem, the original voice of the first agent end (i.e. the voice in the second language) can be directly played to the customer during the waiting for the reply of the agent to be transcribed, translated and synthesized, and then the synthesized voice is switched to and played to the customer after the reply of the agent is transcribed, translated and synthesized.

[0112] Optionally, in the occasion of call monitoring and call quality inspection, the real-time voice streams of the customer end and the first agent end are sent to the third-party voice service engine through the multi-language interaction platform, so that the third-party voice service engine transcribes the real-time voice streams into texts in the corresponding language, translates the texts in the corresponding language into texts in the third language which is habitual to the monitoring end, and further synthesizes the texts in the third language into voice in the third language; meanwhile, the multi-language interaction platform receives the synthesized voice in the third language sent by the third-party voice service engine, and sends the synthesized voice in the third language to the monitoring end.

[0113] It should be noted that, for the occasion of call monitoring and call quality inspection, the prior art can only be used to perform the monitoring and quality inspection through recording after the call, and cannot be used to perform real-time monitoring, while the embodiment can uniformly translate and synthesize the real-time voice streams of the customer end and the first agent end into voice in the third language which is habitual to the monitoring personnel, so as to facilitate the real-time monitoring and quality inspection.

[0114] In a specific embodiment, the customer service method is applied to a customer service system, and an architecture diagram of the system is as followsFigure 2 As shown, it includes a client (such as a user-side mobile terminal), an agent terminal (referred to as an agent end), a traffic platform, a multi-language interaction platform, and a third-party voice service engine. The traffic platform is inherent to the customer service system itself and is used to efficiently manage and process a large amount of telephone traffic. The multi-language interaction platform is deployed locally, and the third-party voice service engine is provided by a third party and can be accessed through the multi-language interaction platform to obtain related services. The third-party voice service engine has many options, such as Microsoft Azure Cognitive Services, Google Cloud Speech, Baidu Voice Open Platform, and iFLYER.

[0115] Based on Figure 2 As shown in the architecture diagram, the customer service method supporting multi-language interaction provided by the embodiment specifically includes the following steps:

[0116] S1: The customer calls the traffic platform through the client, and the traffic platform judges the habitual language of the customer, specifically including:

[0117] S1.1: The traffic platform queries the historical incoming information of the incoming number, and the historical incoming information includes the habitual language;

[0118] If the habitual language is Chinese, go to step S2.

[0119] If the habitual language is not Chinese or the habitual language is not queried, go to step S1.2.

[0120] S1.2: The traffic platform plays an English and Chinese voice prompt for the customer to select the service language, such as "Please select the language of customer service, Chinese please press 1, other please press 2; Please select the language of customer service, Chinese please press 1, other please press 2".

[0121] S1.3: The traffic platform obtains the customer's input service language selection, and if the customer's selected service language is Chinese, go to step S2. If the customer's selected service language is other, go to step S3.

[0122] S2: The traffic platform matches the agent end (i.e., the second agent end) for the client, the agent end establishes a call with the client, the agent end and the client directly conduct a voice dialogue, and the method ends.

[0123] S3: The traffic platform matches the agent end (i.e., the first agent end) for the client, the agent end establishes a call with the client, the traffic platform obtains the customer's voice in real time, and sends the customer's voice to the third-party voice service engine through the multi-language interaction platform;

[0124] It should be noted that the traffic platform is a core component of the system, and its function expansion not only has high implementation complexity, but also may affect the original stability. Therefore, in order to avoid interference and ensure stability, a multilingual interaction platform (mainly responsible for forwarding function) is specially set up for calling third-party voice service engine, so that the original traffic platform does not need to be added or modified.

[0125] S4: The third-party voice service engine identifies the language of the customer's voice through a large model, and the language is recorded as language A, and the text is transcribed into the corresponding text;

[0126] S5: The third-party voice service engine translates the text into the language (such as Chinese) used by the agent end through a large model, and the language is recorded as language B, and the text of language B is synthesized into voice;

[0127] S6: The text and / or voice of language B is sent to the agent end through the multilingual interaction platform;

[0128] S7: After receiving the text and voice of language B, the agent end replies through text or voice, and sends it to the third-party voice service engine through the multilingual interaction platform;

[0129] S8: For voice reply, the third-party voice service engine first transcribes it into the corresponding text; for text reply, directly jump to step S9;

[0130] S9: The third-party voice service engine translates the text into language A through a large model, and synthesizes the text of language A into voice;

[0131] S10: The voice of language A is sent to the client end through the multilingual interaction platform and the traffic platform.

[0132] Repeat steps S3-S10 to realize multi-round interaction between the client end and the agent end.

[0133] It should be noted that after the first identification of the language of the client, the language of the subsequent conversation can be defaulted to be unchanged, and it does not need to be re-identified.

[0134] Further, after the method ends, the user call-in information (such as user number, habitual language, call-in time) of this time is stored in the database, so that the habitual language of the client can be quickly obtained when the client calls in next time.

[0135] It should be noted that after the call ends, the business platform obtains the habitual language of the client from the third-party voice service engine through the multilingual interaction platform, and stores it in the database.

[0136] Further, the traffic platform also includes voice activity detection (VAD) to detect whether the customer / agent voice is complete, and after completion, jump to the next step. The basic principle of VAD is to analyze various characteristics of the voice signal to determine whether the current audio frame contains human voice. A typical VAD system usually includes the following steps: preprocessing: pre-emphasis, framing, windowing, etc. on the input audio signal. Feature extraction: extract energy, zero-crossing rate, spectrum, cepstrum, etc. Feature parameters. Speech / non-speech decision: use statistical models or machine learning algorithms to make decisions based on extracted features. Smoothing processing: smooth the decision results to avoid frequent state switching.

[0137] Further, the agent interface is provided with a common language window, which supports the configuration of various common language sentences and displays them in the form of quick buttons. The agent clicks the common language button, and the system automatically translates and synthesizes the voice of language A and sends it to the client.

[0138] Specifically, when the agent replies to the customer, for common language, the agent can directly click the corresponding content in the common language window of the operation interface without oral expression. The content is forwarded to the third-party voice service engine through the multilingual interaction platform, translated and synthesized, and then sent to the client through the multilingual interaction platform and the traffic platform in turn.

[0139] It should be noted that the common language in the common language window needs to be updated regularly to ensure that it meets the latest situation. At the same time, it can be optimized according to feedback from others to improve communication effectiveness.

[0140] Further, to shorten the delay caused by language conversion, the embodiment deploys an automatic sentence breaking module in the multilingual interaction platform or the third-party voice service engine.

[0141] The automatic sentence breaking process of the automatic sentence breaking module is as shown in Figure 3 As shown in step S4, it is not necessary to wait for the entire text to be transcribed before performing automatic sentence breaking during the transcription process into the corresponding text. The specific execution steps of the automatic sentence breaking module are as follows:

[0142] S4.1: Sentence breaking model training: using annotated training data, according to certain voice pause rules, tone change rules, grammar rules and punctuation rules to train the sentence breaking model. Traditional automatic sentence breaking model training only considers the rules on the text, i.e. grammar rules and punctuation rules. Considering that the embodiment is to break the text after voice recognition, the rules on the voice are also added, i.e. voice pause rules, tone change rules, so that the sentence breaking is more reasonable. The following is an example:

[0143] Voice activity detection (VAD) can detect voice pauses, and voice pause rules are as follows:

[0144] Natural pauses: In natural language communication, people naturally pause between sentences, between clauses, and where emphasis is needed when speaking. For example, when listing things, people will pause slightly after each item, such as "I like fruits, apples, bananas, and oranges." In speech, people may have short pauses after "apples" and "bananas," and these pause positions can be used as a reference for sentence breaks.

[0145] Punctuation corresponding pauses: Punctuation marks such as periods, commas, and semicolons represent different pause lengths. In text segmentation after speech recognition, the pause time in the speech can be used to infer the punctuation marks that should be inserted. For example, a longer pause may correspond to a period, indicating the end of a sentence; a shorter pause may correspond to a comma, indicating an internal division of a sentence.

[0146] Pitch changes are usually related to the pitch, intensity, and duration of the speech signal. The speech signal can be preprocessed to extract these features. For example, Fourier transform and other methods can be used to obtain the frequency spectrum information of the speech signal, and then analyze the pitch change; or directly calculate the amplitude of the speech signal to analyze the intensity change. Pitch change rules are as follows:

[0147] Questioning intonation: When the speaker asks a question, the intonation usually rises. For example, in the Chinese sentence "Can you get me subscription A" or the English sentence "Can you get me subscription A", the intonation of the last word rises, which can be an important feature for identifying a question, so that the question mark can be correctly added when segmenting sentences.

[0148] Emphasis intonation: When the speaker expresses emphasis, the speaker may increase the volume, speed up the speech, or prolong certain syllables. For example, in the Chinese sentence "I really liked your answer" or the English sentence "I really liked your answer", the word "really" may be emphasized, and the intonation may change, which can help identify the key part of the sentence and affect the way of segmentation, such as pausing slightly after "really" to enhance the emphasis effect.

[0149] Emotional intonation: Different emotional expressions are also accompanied by changes in intonation. For example, when excited, the intonation may be higher and more fluctuating, and when sad, the intonation may be lower and more stable. In text segmentation after speech recognition, the change in intonation can be used to judge the emotional color of the sentence, so as to more accurately segment the sentence to meet the speaker's emotional expression.

[0150] It should be noted that the voice pause detection and the intonation detection are performed simultaneously, and then the punctuation is performed based on the voice pause and the intonation.

[0151] S4.2: Text preprocessing: During the process of real-time transcription of the voice stream into text, the input text is cleaned and standardized in real time, such as uniform encoding format (referring to character encoding format, such as UTF-8 format), removal of redundant spaces and line breaks, etc.

[0152] S4.3: Voice feature extraction: Extracting features in the voice, such as pause position, pause duration, intonation change feature, as the basis for subsequent punctuation.

[0153] S4.4: Text feature extraction: Extracting features in the text, such as punctuation marks, words, parts of speech, etc., as the basis for subsequent punctuation.

[0154] It should be noted that the features extracted for different languages will be different, for example, Chinese does not have a clear word separator, and can extract lexical, grammatical and semantic features for punctuation. For English punctuation, there are spaces between words in English, and punctuation is mainly based on punctuation marks and grammatical structure. By analogy, different languages have different punctuation habits and rules, and need to be set and processed according to the characteristics of the specific language.

[0155] S4.5: Automatic punctuation: input the extracted voice features and text features into the trained punctuation model to punctuate the text.

[0156] It should be noted that the automatic punctuation module processes the text and punctuates at appropriate positions, dividing a long paragraph into several segments, and after each punctuation is completed, subsequent processing (including synthesis, sending) is performed in time. In this way, subsequent processing can be performed without waiting for the entire text to be transcribed, which shortens the delay caused by language conversion and improves customer experience.

[0157] Further, after the customer's speech is finished, the agent cannot hear the voice feedback before the transcription, translation and synthesis are completed, which may cause the service to be not smooth, and the customer may think that the service has been interrupted. To solve this problem, during the waiting for the agent's reply to the transcription, translation and synthesis process, the agent's original voice can be directly played to the customer (i.e. the Chinese original voice directly sent by the agent, such as "You have 5GB of remaining traffic"), and then switched to the synthesized voice after the agent's reply to the transcription, translation and synthesis is completed, and played to the customer.

[0158] In step S9, to improve the sound quality and expressiveness of the synthesized voice, the following can also be included:

[0159] Fine prosody modeling: Fine modeling of prosodic features of speech, such as predicting the fundamental frequency and energy at the phoneme level, associating duration prediction with fundamental frequency, energy, and other prosodic components, and using an autoregressive structure for fine-grained modeling to make the synthesized speech more natural and smooth.

[0160] Emotion and style control: By introducing emotion labels, style parameters, etc., the model can synthesize speech with specific emotions and styles according to requirements. For example, the emotion CSS model based on heterogeneous graph context modeling can learn emotional cues in the dialogue context and infer the emotional style of the current speech through the emotion rendering module to realize emotional speech synthesis.

[0161] High-quality vocoder: Advanced vocoder technology, such as deep learning-based vocoder, can more accurately convert acoustic features into natural speech waveforms, improving the quality and expressiveness of synthesized speech. Acoustic features are extracted from the original speech of the customer service end, including spectral features, fundamental frequency, duration, and energy. Based on spectral features, "tone" details are restored, combined with fundamental frequency and duration to adjust "pitch" and "rhythm", and energy corresponds to restore "volume".

[0162] It should be noted that in the process of synthesizing speech from text, fine prosody modeling, emotion and style control, high-quality vocoder, etc. are preferably used. The application of these technologies significantly improves the naturalness and expressiveness of synthesized speech, providing users with a better voice interaction experience.

[0163] It should be noted that during the waiting for the agent to reply to the transcription, translation, and synthesis process, no audio can be played, or a pre-set audio can be played directly to the customer, such as playing the voice prompt "waiting for the agent to reply, please wait".

[0164] Further, the method further includes a listening function, and the listening process of the listening function is as shown in Figure 4 , which specifically includes:

[0165] T1: The multi-language interaction platform acquires real-time speech streams of the client and the agent end and sends them to the third-party speech service engine;

[0166] T2: The third-party speech service engine transcribes the real-time speech into corresponding text and translates the text into the language commonly used by the listening end, denoted as C language, and then synthesizes the text of C language into speech;

[0167] T3: The listening end receives and plays the synthesized speech.

[0168] For example, when the customer service center has visitors from different countries, the monitoring function can convert the content of the conversation between the customer and the agent (which may involve multiple languages) into the language familiar to the visitor in real time according to the visitor's habitual language, making it easier for the visitor to understand the conversation process.

[0169] It should be noted that the customer service method provided by the present disclosure has the following beneficial effects:

[0170] a) The text is processed by the automatic sentence breaking module, and is broken at appropriate positions to divide a long paragraph into several segments. After each sentence breaking, subsequent processing is performed in time. In this way, subsequent processing can be performed without waiting for the entire transcription to be completed, thereby shortening the delay caused by language conversion and improving the customer experience.

[0171] b) Considering that the present disclosure performs text breaking after speech recognition, the breaking rules include not only text-level rules but also speech-level rules, i.e. speech pause rules and intonation change rules, making the breaking more reasonable.

[0172] c) When the customer calls in, the present disclosure queries the historical call information of the incoming number through the traffic platform to obtain the habitual language of the customer. When the customer's habitual language is not Chinese, it can automatically switch to a multi-language interaction mode, improving the customer experience.

[0173] d) The present disclosure improves the sound quality and expressiveness of the synthesized speech through fine prosody modeling, emotion and style control, and high-quality vocoders, thereby improving the customer experience.

[0174] e) During the waiting for the agent to reply, the present disclosure can directly play the original voice or voice prompt of the agent to the user, avoiding a smooth service connection, and improving the customer experience.

[0175] f) For occasions that require call monitoring and call quality inspection, the present disclosure can translate and synthesize the A language of the client and the B language of the agent into the habitual C language of the monitoring personnel, making it easier to monitor and inspect.

[0176] The customer service method provided by the embodiments of the present disclosure first determines the habitual language of the customer according to the historical call-in information of the call-in number corresponding to the call request of the client by the traffic platform; then, in response to the habitual language being non-Chinese or the habitual language of the customer not being queried, further confirms whether the currently selected customer service language of the customer is non-Chinese; if the currently selected customer service language of the customer is non-Chinese, in the process of voice dialogue between the client and the corresponding first agent terminal, the customer voice is acquired in real time, and the customer voice is forwarded to the third-party voice service engine through the multilingual interaction platform, so that the third-party voice service engine transcribes the customer voice into text in the corresponding first language in real time, and performs automatic sentence segmentation in the transcription process; after each sentence is segmented, the corresponding target text is translated into text in the second language which is habitual to the first agent terminal, and is further synthesized into voice in the second language; the synthesized voice in the first language sent by the multilingual interaction platform is received, the synthesized voice in the first language is obtained by translating and synthesizing the voice or text replied by the first agent terminal after the third-party voice service engine forwards the text and / or voice in the second language to the first agent terminal through the multilingual interaction platform; and finally, the synthesized voice in the first language is sent to the client. By querying the historical call-in information of the customer to determine the habitual language or further confirm the currently selected customer service language of the customer, the language demand of the customer is accurately identified in advance, and smooth communication is ensured. At the same time, by acquiring the customer voice in real time, with the cooperation of the multilingual interaction platform and the third-party voice service engine, the customer voice is transcribed, automatically segmented, and quickly translated and synthesized between multiple languages, so that the customer can obtain feedback in real time during the dialogue process, the real-time performance of voice interaction is effectively ensured, and the customer experience is significantly improved. The existing customer service method needs to wait for the customer to finish speaking before translation, which leads to poor real-time performance of communication and poor customer experience, and the habitual language of the customer cannot be identified in time, which leads to poor customer experience.

[0177] Embodiment 2:

[0178] As shown in Figure 5 , the present embodiment provides a customer service method applied to a multilingual interaction platform, the method comprising:

[0179] Step S201: receiving the customer voice sent by the traffic platform, the customer voice being acquired and sent in real time by the traffic platform during the voice dialogue between the client and the corresponding first agent terminal in response to the habitual language being non-Chinese or the habitual language of the customer not being queried, after the traffic platform determines the habitual language of the customer according to the historical call-in information of the call-in number corresponding to the call request of the client;

[0180] In the embodiment, the historical incoming call information includes a habitual language.

[0181] Step S202: forwarding the customer voice to a third-party voice service engine to make the third-party voice service engine transcribe the customer voice into text in a corresponding first language in real time, and performing automatic sentence segmentation during the transcription process, translating the corresponding target text into text in a second language habitual to the first agent end after each sentence is segmented, and further synthesizing the text into voice in the second language;

[0182] In the embodiment, the third-party voice service engine identifies the language corresponding to the customer voice (i.e., the first language) through a large model, and transcribes it into corresponding text, and performs sentence segmentation using a trained sentence segmentation model during the transcription process. The trained sentence segmentation model is trained in advance using annotated training data according to certain voice pause rules, tone change rules, grammar rules, and punctuation rules.

[0183] Step S203: receiving the text and / or voice in the second language sent by the third-party voice service engine;

[0184] Step S204: sending the text and / or voice in the second language to the first agent end, and forwarding the voice or text replied by the first agent end to the third-party voice service engine to make the third-party voice service engine translate and synthesize the voice or text replied by the first agent end to obtain synthesized voice in the first language;

[0185] Step S205: receiving the synthesized voice in the first language sent by the third-party voice service engine, and sending the synthesized voice in the first language to the client through the traffic platform.

[0186] In the embodiment, in order to improve the sound quality and expressiveness of the synthesized voice, the third-party voice service engine can use fine prosody modeling, emotion and style control, high-quality vocoder and other technologies during the synthesis process to provide users with a better voice interaction experience.

[0187] Further, the method further comprises:

[0188] sending the real-time voice stream of the client and the first agent end to the third-party voice service engine to make the third-party voice service engine transcribe the real-time voice stream into text in a corresponding language, and translate the text in the corresponding language into text in a third language habitual to the listener end, and further synthesize the text into voice in the third language;

[0189] receiving the synthesized voice in the third language sent by the third-party voice service engine, and sending the synthesized voice in the third language to the listener end.

[0190] In the embodiment, in the case where call monitoring and call quality inspection are required, the multi-language interactive platform can unify and translate the real-time voice streams of the client and the first agent end into the third language voice which is used by the monitoring personnel, so as to facilitate real-time monitoring and quality inspection.

[0191] Embodiment 3

[0192] As shown in Figure 6 the embodiment, a customer service method is provided, which is applied to a third-party voice service engine, and the method comprises the following steps:

[0193] Step S301: receiving the customer voice forwarded by the multi-language interactive platform, wherein the customer voice is obtained in real time during the voice dialogue between the client and the corresponding first agent end, and is sent in response to the fact that the habitual language of the customer is non-Chinese or the habitual language of the customer is not queried after the history call information of the incoming call number corresponding to the call request of the client is queried by the call platform to determine the habitual language of the customer.

[0194] In the embodiment, the history call information comprises the habitual language.

[0195] Step S302: transcribing the customer voice into the text of the corresponding first language, and performing automatic sentence segmentation during the transcription process, wherein after each sentence is segmented, the corresponding target text is translated into the text of the second language which is used by the first agent end, and is further synthesized into the voice of the second language.

[0196] Optionally, the step of transcribing the customer voice into the text of the corresponding first language and performing automatic sentence segmentation during the transcription process specifically comprises the following steps:

[0197] recognizing the language corresponding to the customer voice by using a preset large model, and taking the language as the first language;

[0198] transcribing the customer voice into the text of the first language, and extracting the voice features and text features during the transcription process, and inputting the extracted voice features and text features into a trained sentence segmentation model for sentence segmentation, wherein the trained sentence segmentation model performs sentence segmentation according to the voice pause rules, the intonation change rules, the grammar rules and the punctuation rules.

[0199] In the embodiment, the trained sentence segmentation model is obtained by using the annotated training data in advance according to certain voice pause rules, intonation change rules, grammar rules and punctuation rules. The voice pause rules comprise natural pauses and pauses corresponding to punctuation marks. The intonation change rules comprise question intonation, emphasis intonation and emotion intonation.

[0200] Step S303: forwarding the text and / or voice in the second language to the first agent end through the multi-language interaction platform, receiving the voice or text replied by the first agent end, and obtaining the synthesized voice in the first language after translating and synthesizing the voice or text replied by the first agent end.

[0201] Optionally, the step of obtaining the synthesized voice in the first language after translating and synthesizing the voice or text replied by the first agent end comprises:

[0202] for the voice replied by the first agent end, transcribing the replied voice into text in the corresponding language, and translating the text in the corresponding language into text in the first language;

[0203] for the text replied by the first agent end, translating the replied text into text in the first language;

[0204] synthesizing the translated text in the first language, modeling the prosodic features of the voice, introducing emotion labels and style parameters in the process of synthesis, and using a preset vocoder technology.

[0205] In this embodiment, in order to improve the sound quality and expressiveness of the synthesized voice, the third-party voice service engine can use multiple technologies such as fine prosodic modeling, emotion and style control, and high-quality vocoder in the process of synthesis, to provide users with a better voice interaction experience.

[0206] Step S304: sending the synthesized voice in the first language to the client through the multi-language interaction platform and the traffic platform.

[0207] In this embodiment, the third-party voice service engine sends the synthesized voice in the first language to the client through the multi-language interaction platform and the traffic platform for playing, and performs the next round of interaction.

[0208] Embodiment 4:

[0209] With reference to Figure 7 , this embodiment provides a customer service system, comprising: a traffic platform 11, a multi-language interaction platform 12, and a third-party voice service engine 13.

[0210] The traffic platform 11 is configured to perform the customer service method in Embodiment 1.

[0211] The multi-language interaction platform 12 is configured to perform the customer service method in Embodiment 2.

[0212] The third-party voice service engine 13 is configured to perform the customer service method in Embodiment 3.

[0213] Embodiment 5:

[0214] Reference Figure 8 The embodiment provides a customer service device, comprising a memory 21 and a processor 22, the memory 21 stores a computer program, and the processor 22 is configured to execute the computer program to perform the customer service method in the embodiment 1 or the embodiment 2 or the embodiment 3.

[0215] The memory 21 is connected with the processor 22, the memory 21 can adopt a flash memory or a read-only memory or other memories, and the processor 22 can adopt a central processing unit or a single-chip microcomputer.

[0216] Embodiment 6:

[0217] The embodiment provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the customer service method in the above embodiment 1 or the embodiment 2 or the embodiment 3.

[0218] The computer readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, computer program modules or other data. The computer readable storage medium includes but is not limited to RAM (Random Access Memory, Random Access Memory), ROM (Read-Only Memory, Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory, Electrically Erasable Programmable Read-Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory, Compact Disc Read-Only Memory), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, magnetic tape, magnetic disc storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer.

[0219] In summary, the customer service method, system, device, and readable storage medium provided in this disclosure firstly, based on the client's call request, query the historical call information of the corresponding incoming number to determine the customer's preferred language; then, in response to the query finding that the preferred language is not Chinese or that the customer's preferred language is not found, further confirmation is made as to whether the customer's currently selected customer service language is not Chinese; if the customer's currently selected customer service language is not Chinese, then during the voice dialogue between the client and the corresponding first agent, the customer's voice is acquired in real time and forwarded to a third-party voice service engine through a multilingual interaction platform, so that the third-party language... The voice service engine transcribes the client's voice into text in the corresponding first language in real time, and performs automatic sentence segmentation during the transcription process. After each sentence segmentation is completed, the corresponding target text is translated into text in the second language commonly used by the first agent, and further synthesized into speech in the second language. Then, it receives the synthesized speech in the first language sent by the multilingual interaction platform. The synthesized speech in the first language is obtained by the third-party voice service engine after forwarding the text and / or speech in the second language to the first agent through the multilingual interaction platform, and then translating and synthesizing the speech or text replied by the first agent. Finally, the synthesized speech in the first language is sent to the client. This disclosure determines a customer's preferred language or confirms their current language selection by querying their historical inbound call information, accurately identifying their language needs in advance and ensuring smooth communication. Simultaneously, by acquiring customer voice recordings in real time and leveraging the collaboration of a multilingual interaction platform and a third-party voice service engine, it achieves instant transcription, automatic sentence segmentation, and rapid translation and speech synthesis between multiple languages. This allows customers to receive immediate feedback during the conversation, effectively ensuring the real-time nature of voice interaction and significantly improving the customer experience. It solves the problems of existing customer service methods that require waiting for the customer to finish speaking before translation, resulting in poor real-time communication and a subpar customer experience, as well as the inability to promptly identify whether a customer's preferred language is Chinese, leading to a poor customer experience.

[0220] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A customer service method characterized by, The method applied to a traffic platform comprises: According to a call request of a client, historical incoming information of an incoming number corresponding to the call request is queried to determine a habitual language of a customer; In response to the habitual language being non-Chinese or the habitual language of the customer not being queried, it is further determined whether a currently selected service language of the customer is non-Chinese; If the currently selected service language of the customer is non-Chinese, during a voice dialogue process between the client and a first agent terminal, customer voice is acquired in real time, and the customer voice is forwarded to a third-party voice service engine through a multi-language interaction platform, so that the third-party voice service engine transcribes the customer voice into text in a first language in real time, and performs automatic sentence segmentation during the transcription, and after each sentence is segmented, the corresponding target text is translated into text in a second language which is habitual to the first agent terminal, and is further synthesized into voice in the second language; The synthesized voice in the first language sent by the multi-language interaction platform is received, the synthesized voice in the first language being obtained by translating and synthesizing voice or text replied by the first agent terminal through the multi-language interaction platform after the second language text and / or voice is forwarded to the first agent terminal by the third-party voice service engine; The synthesized voice in the first language is sent to the client.

2. The method of claim 1, wherein, In response to the habitual language being non-Chinese or the habitual language of the customer not being queried, the currently selected service language of the customer is further determined to be non-Chinese, specifically comprising: In response to the habitual language being non-Chinese or the habitual language of the customer not being queried, the customer is prompted to select a service language by voice; If the selected service language of the customer is Chinese, it is determined that the currently selected service language of the customer is Chinese; If the selected service language of the customer is other, it is determined that the currently selected service language of the customer is non-Chinese; Before the customer voice is acquired in real time during the voice dialogue process between the client and the first agent terminal, the method further comprises: The first agent terminal is matched for the client, and a call is established.

3. The method of claim 2, wherein, The method further comprises: In response to the habitual language being Chinese or the currently selected service language of the customer being Chinese, a second agent terminal is matched for the client, and a call is established.

4. The method of claim 1, wherein, The method further comprises at least one of the following: Voice activity detection (VAD) is performed to detect whether the voice of the client or the first agent terminal is complete; In response to the voice dialogue being ended, the habitual language of the customer is acquired from the third-party voice service engine through the multi-language interaction platform, and is stored, wherein the habitual language of the customer is the first language; During waiting for the synthesized voice in the first language, the original voice of the first agent terminal or a preset voice prompt is played to the customer.

5. A customer service method characterized by, The method applied to a multi-language interaction platform comprises: receive customer voice sent by a traffic platform, the customer voice being obtained by the traffic platform in real time during a voice dialogue process between a client and a corresponding first agent terminal in response to a call request of the client, the traffic platform querying historical call information of an incoming number corresponding to the call request to determine a habitual language of the client, and the traffic platform further confirming that a currently selected service language of the client is non-Chinese in response to the habitual language being non-Chinese or the habitual language of the client not being queried; forward the customer voice to a third-party voice service engine, so that the third-party voice service engine transcribes the customer voice into text in a corresponding first language in real time, and performs automatic sentence segmentation during the transcription, translates corresponding target text into text in a second language which is habitual to the first agent terminal after each sentence is segmented, and further synthesizes the text into voice in the second language; receive the text and / or voice in the second language sent by the third-party voice service engine; send the text and / or voice in the second language to the first agent terminal, and forward voice or text replied by the first agent terminal to the third-party voice service engine, so that the third-party voice service engine translates and synthesizes the voice or text replied by the first agent terminal to obtain synthesized voice in the first language; receive the synthesized voice in the first language sent by the third-party voice service engine, and send the synthesized voice in the first language to the client through the traffic platform.

6. The method of claim 5, wherein, The method further comprises: send real-time voice streams of the client and the first agent terminal to the third-party voice service engine, so that the third-party voice service engine transcribes the real-time voice streams into text in a corresponding language, translates the text in the corresponding language into text in a third language which is habitual to a listening terminal, and further synthesizes the text in the third language into voice in the third language; receive the synthesized voice in the third language sent by the third-party voice service engine, and send the synthesized voice in the third language to the listening terminal.

7. A customer service method characterized by, The method applied to a third-party voice service engine comprises: receive customer voice forwarded by a multi-language interaction platform, the customer voice being obtained by a traffic platform in real time during a voice dialogue process between a client and a corresponding first agent terminal in response to a call request of the client, the traffic platform querying historical call information of an incoming number corresponding to the call request to determine a habitual language of the client, and the traffic platform further confirming that a currently selected service language of the client is non-Chinese in response to the habitual language being non-Chinese or the habitual language of the client not being queried; transcribe the customer voice into text in a corresponding first language, and perform automatic sentence segmentation during the transcription, translate corresponding target text into text in a second language which is habitual to the first agent terminal after each sentence is segmented, and further synthesize the text into voice in the second language; forward the text and / or voice in the second language to the first agent terminal through the multi-language interaction platform, receive voice or text replied by the first agent terminal sent by the multi-language interaction platform, and translate and synthesize the voice or text replied by the first agent terminal to obtain synthesized voice in the first language; The synthesized speech in the first language is sent to the client through a multilingual interaction platform and a call center platform.

8. The method of claim 7, wherein, The step of transcribing the customer's speech into text in the corresponding first language, and performing automatic sentence segmentation during the transcription process, specifically includes: The language corresponding to the customer's voice is identified by a pre-set large model and used as the first language; The customer's speech is transcribed into text in the first language, and speech features and text features are extracted during the transcription process. The extracted speech features and text features are then input into a trained sentence segmentation model for sentence segmentation. The trained sentence segmentation model segments sentences according to speech pause rules, intonation change rules, grammatical rules, and punctuation rules.

9. The method of claim 7, wherein, The process of translating and synthesizing the speech or text replied by the first agent terminal to obtain synthesized speech in the first language specifically includes: For the voice reply from the first agent terminal, the voice reply is transcribed into text in the corresponding language, and the text in the corresponding language is translated into text in the first language. For the text reply from the first agent terminal, translate the text reply into text in the first language; The translated text in the first language is synthesized, and during the synthesis process, the prosodic features of the speech are modeled, emotion tags and style parameters are introduced, and a pre-set vocoder technology is used.

10. A customer service system, characterized by This includes a call center platform, a multilingual interaction platform, and a third-party voice service engine; The call center platform is used to perform the customer service method according to any one of claims 1-4; The multilingual interaction platform is used to perform the customer service method as described in claim 5 or 6; The third-party voice service engine is used to perform the customer service method as described in any one of claims 7-9.

11. A customer service device, characterized by It includes a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the customer service method as described in any one of claims 1-4, or the customer service method as described in claims 5 or 6, or the customer service method as described in any one of claims 7-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the customer service method as described in any one of claims 1-4, or the customer service method as described in claim 5 or 6, or the customer service method as described in any one of claims 7-9.