Method for configuring multilingual language interaction skills based on large model assistance

Through the multilingual language interaction skills configuration method assisted by large-model, existing Chinese skills are converted into multilingual skills as a whole, and combined with automatic translation and manual correction, the problems of computing power and human resources consumption in the existing technology are solved, and efficient and low-cost multilingual voice interaction is achieved.

CN120255841APending Publication Date: 2025-07-04AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308035.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, multilingual voice skills configuration methods have problems such as huge computing power consumption, waste of server resources and serious human resources consumption, especially in real-time translation models and multi-skill configuration solutions.

Method used

The overall conversion and local maintenance methods based on large models are adopted, and the existing Chinese skills are converted into multilingual skills with one click through the translation capabilities of large models, and combined with automatic translation and manual correction, the precise adjustment of multilingual reply text is achieved.

Benefits of technology

It greatly reduces the labor costs caused by multi-skill maintenance, reduces the server computing burden, and improves the operation efficiency and user experience of the voice interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255841A_ABST
    Figure CN120255841A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction, in particular to a method for configuring multilingual language interaction skills based on large model assistance, which comprises the following steps of: integrally converting existing Chinese skills, creating skills, turning on an internationalized broadcast switch, selecting languages needing to be supported, and configuring Chinese skill contents. Converting Chinese to multilingual configuration based on large model capability, and issuing and using skills; for the maintenance of existing multilingual skills, the skills are opened, dialogue conditions needing to be maintained are selected, a multilingual voice reply configuration interface is entered, a target language is selected, if the language is a Chinese language, single sentence translation is automatically carried out, and if the language is other languages, manual correction and skill release and use are carried out. According to the method and the device, the labor cost caused by multi-skill maintenance can be greatly reduced, and the overall operation efficiency of the voice interaction system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and in particular to a method for configuring multi-language language interaction skills assisted by a large model. Background Art

[0002] With the rapid development of intelligent voice interaction technology, voice skills have been widely applied in multiple fields such as in-vehicle systems, smart homes, and customer service robots. Against the backdrop of the growing globalization and multi-language communication needs, how to efficiently configure voice skills that support multiple languages has become a hot topic in the industry. Currently, the mainstream voice skill configuration methods mainly focus on how to quickly support multiple languages while ensuring natural and fluent voice interaction, providing users with richer and more convenient language services.

[0003] Currently, there are mainly two solutions for the mainstream multi-language voice skill configuration methods in the industry. One is the real-time translation model solution, which intervenes during the use of voice skills. Specifically, the voice language skills are still configured in the form of Chinese voice interaction skills on the DUI platform, and there is no change during the skill production process. During actual use, for the output of the voice skill, the Chinese reply is converted into the text of the target language through a real-time translation model, thereby achieving multi-language support. The other is the multi-skill configuration solution, which intervenes during the voice skill configuration stage and directly configures an independent skill for each supported language, such as configuring an English skill, a French skill, etc. Each language skill not only includes the reply text but also involves a separate configuration of a large amount of dialogue logic. During use, multiple skills work together to achieve multi-language interaction.

[0004] However, both of the above two solutions have obvious defects. In the real-time translation model solution, since each voice reply must be converted into the target language through a real-time translation model, even if many replies are fixed texts, real-time translation calculation is still required, resulting in huge computing power consumption, waste of server resources, and possible reply delay problems. In the multi-skill configuration solution, to support multiple languages, this solution needs to separately configure and maintain voice skills for each language. Since the voice skills on the DUI platform not only involve reply texts but also include complex dialogue logic, the maintenance workload of each language skill is extremely large, seriously consuming human resources. Traditionally, to solve this problem, a multi-skill maintenance method is often adopted, supporting multiple languages by increasing human resources. However, as the number of supported languages increases and the labor cost continues to rise, the cost of skill maintenance becomes increasingly significant. Summary of the Invention

[0005] This application provides a method for configuring multi-language language interaction skills assisted by a large model, which can significantly reduce the labor cost generated by multi-skill maintenance and significantly improve the overall operation efficiency of the voice interaction system. This application provides the following technical solutions:

[0006] In a first aspect, the present application provides a method for configuring multi - language language interaction skills assisted by a large - model, and the method includes:

[0007] For the overall conversion of existing Chinese skills, create a skill, turn on the international broadcast switch and select the languages to be supported, configure the Chinese skill content, perform the configuration conversion from Chinese to multiple languages based on the large - model capabilities, and release and use the skill;

[0008] For the maintenance of existing multi - language skills, open the skill and select the dialogue conditions to be maintained, enter the multi - language voice response configuration interface and select the target language. If the target language is Chinese, automatic single - sentence translation is performed; if it is other languages, manual correction is carried out, and then the skill is released and used.

[0009] In a specific feasible implementation, for the overall conversion of existing Chinese skills, creating a skill includes:

[0010] On the skill customization page of the DUI platform, click the customize skill button, and configure the skill type and skill name according to requirements to create a new voice skill.

[0011] In a specific feasible implementation, the configuration of the Chinese skill content includes:

[0012] According to the Chinese voice skill configuration process of the DUI platform, configure the Chinese semantic content and dialogue logic of the skill.

[0013] In a specific feasible implementation, the configuration conversion from Chinese to multiple languages based on the large - model capabilities includes:

[0014] After the international broadcast function is enabled, the international broadcast – NLG full - volume translation button is displayed on the skill home page; click the button to call the rule - based translation ability of the large - model to perform an overall conversion of all Chinese response texts and related semantic information in the skill, and generate multi - language response content in the corresponding target languages.

[0015] In a specific feasible implementation, entering the multi - language voice response configuration interface and selecting the target language includes:

[0016] Enter the corresponding response text configuration interface, and in the response configuration interface, select the target language to be maintained according to actual needs.

[0017] In a specific feasible implementation, if the target language is Chinese, automatic single - sentence translation includes:

[0018] After selecting the Chinese language, the translation button is automatically displayed in the voice reply text configuration area; after clicking this button, the rule-based translation ability of the large model is called to automatically translate the Chinese reply text of the current sentence into a single sentence, generating the reply text corresponding to the target language.

[0019] In a specific implementable embodiment, the manual correction for other languages includes:

[0020] For non-Chinese languages selected by the operator, manual modification of the translation result is allowed in the text configuration area.

[0021] In a second aspect, the present application provides a configuration system for multi-language interaction skills assisted by a large model, adopting the following technical solutions:

[0022] A configuration system for multi-language interaction skills assisted by a large model, including:

[0023] An overall conversion module for existing Chinese skills, used for the overall conversion of existing Chinese skills, creating skills, turning on the international broadcast switch and selecting the languages to be supported, configuring Chinese skill content, converting Chinese to multi-language configuration based on the large model ability, and releasing and using skills;

[0024] A maintenance module for existing multi-language skills, used for the maintenance of existing multi-language skills, opening the skills and selecting the dialogue conditions to be maintained, entering the multi-language voice reply configuration interface and selecting the target language, automatically performing single-sentence translation if it is the Chinese language, and performing manual correction if it is other languages, and releasing and using skills.

[0025] In a third aspect, the present application provides an electronic device, the device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a configuration method for multi-language interaction skills assisted by a large model as described in the first aspect.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, a program is stored in the storage medium, and when the program is executed by a processor, it is used to implement a configuration method for multi-language interaction skills assisted by a large model as described in the first aspect.

[0027] In summary, the beneficial effects of the present application at least include:

[0028] 1) By using the rule translation ability of the large model to achieve overall multi - language conversion of existing Chinese skills, a single skill can generate response texts in multiple languages, eliminating the need to separately configure and maintain independent skills for each language. This integrated configuration significantly reduces the repetitive labor of manual work in the process of skill development, debugging, updating, and maintenance, simplifies the operation process of the entire voice interaction system, and improves the development and operation efficiency. In addition, unified skill management also helps reduce the error risk caused by multi - skill configuration, further improving the overall human efficiency.

[0029] 2) Adopt the pre - translation method. Before skill release, use the large model to complete the conversion of all fixed response texts and related semantic information, rather than performing real - time translation during each user voice interaction. In this way, the system avoids frequently calling the real - time translation model during actual interactions, significantly reducing the computing burden of the online server and the risk of response delay. By completing multi - language conversion in advance during the production process, not only valuable computing resources are saved, but also users can enjoy a fast and stable voice interaction experience during actual use.

[0030] By using the rule translation ability of the large model through the overall conversion method, the existing Chinese skills are converted into multi - language skills with one click, and response texts in all target languages are pre - generated, thus eliminating the repetitive calculations of real - time translation. At the same time, by adopting a maintenance method that combines single - sentence - level automatic translation and manual correction for the existing multi - language skills, precise adjustment of the voice response texts is achieved. Thereby, both the waste of server computing power is effectively reduced, and the labor cost generated by multi - skill maintenance is greatly reduced, significantly improving the overall operation efficiency of the voice interaction system.

[0031] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application and implement it in accordance with the content of the specification, the following takes the preferred embodiments of this application and combines with the attached drawings to elaborate in detail as follows. Brief Description of the Drawings

[0032] Figure 1 It is a schematic flowchart of the multi - language voice skill configuration method based on the overall conversion of existing Chinese skills in this application.

[0033] Figure 2 It is a schematic flowchart of the multi - language voice skill configuration method based on the maintenance of existing multi - language skills in this application.

[0034] Figure 3 It is a structural block diagram of the multi - language language interaction skill configuration system based on the large - model assistance in the embodiments of this application.

[0035] Figure 4It is a block diagram of an electronic device configured with multi - language language interaction skills assisted by a large model in an embodiment of the present application. Detailed implementation manners

[0036] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0037] Optionally, in this application, the configuration method of multi - language language interaction skills assisted by a large model provided in each embodiment is taken as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.

[0038] An embodiment of the present application provides a configuration method of multi - language language interaction skills assisted by a large model. It should be noted that there are two intervention methods for the multi - language skill configuration scheme: the overall conversion of existing Chinese skills and the maintenance of existing multi - language skills. Correspondingly, at the beginning, ordinary Chinese skills are converted into multi - language supported skills with one key by using the large - model rule translation ability; and for the maintenance of existing multi - language skills, single - sentence modification of the skill response text is performed.

[0039] Refer to Figure 1 , which is a schematic flowchart of the multi - language voice skill configuration method based on the overall conversion of existing Chinese skills in the present application. This method uses the large - model rule translation ability to achieve the overall multi - language conversion of the original Chinese skills, thereby reducing the consumption of real - time translation computing resources and simplifying the skill maintenance work. The specific implementation steps are as follows:

[0040] Step S101: Create a skill.

[0041] In step S101, the operator clicks the "Customize Skill" button on the skill customization page of the DUI platform and configures the skill type and skill name according to the requirements, thereby creating a new voice skill.

[0042] Step S102: Turn on the international broadcast switch and select the languages to be supported.

[0043] In step S102, on the skill creation interface and the skill editing interface after creation, the operator selects to turn on the "International Broadcast" function in the advanced configuration options of the skill. After turning it on, the multi - language options supported by the platform will be automatically displayed below the page, and the operator selects the target language according to actual needs to achieve the preset of multi - language capabilities.

[0044] Step S103: Configure the Chinese skill content.

[0045] In step S103, after completing the above basic settings, the operator configures the Chinese semantic content and dialogue logic of the skill in detail according to the conventional Chinese voice skill configuration process of the DUI platform, which serves as the basis for subsequent multilingual conversion.

[0046] Step S104: Chinese to multilingual configuration conversion based on the capabilities of the large model.

[0047] In step S104, after confirming that the skill already has a Chinese configuration and the "Internationalized Announcement" function is enabled, the "Internationalized Announcement - NLG Full Translation" button will be displayed on the skill home page. After the operator clicks this button, the "Rule Translation" ability of the large model is called to perform an overall conversion of all Chinese response texts and related semantic information within the skill, generating multilingual response content in the corresponding target languages. This conversion process is automatically executed and, although it takes some time, it achieves a seamless conversion from Chinese to multiple languages.

[0048] Optionally, the large model in this application is the Sibyl - Dongfeng large model, and other large models can also be selected. This application does not limit the specific type of the large model.

[0049] Step S105: Skill release and use.

[0050] In step S105, after the overall conversion is completed, the operator publishes the converted skill to the production environment. When the released skill is actually used, it realizes the function of a single skill supporting multiple language responses, which not only avoids the waste of server computing power caused by real - time translation but also reduces the labor cost generated by maintaining multiple skills.

[0051] Refer to Figure 2 , which is the flow schematic diagram of the multilingual voice skill configuration method based on the maintenance of existing multilingual skills in this application. By performing local maintenance on the single - sentence response texts in the voice skill, accurate adjustment of the multilingual response texts is achieved, thereby improving the accuracy of voice interaction and the user experience. Based on the original skill maintenance process of the DUI platform, this solution performs targeted maintenance on the skills with the activated internationalized announcement function through a combination of automation and manual intervention. The specific steps are as follows:

[0052] Step S201: Open the skill and select the dialogue conditions that need to be maintained.

[0053] In step S201, the operator opens the skill that needs to be maintained in the skill management interface of the DUI platform and selects the corresponding dialogue conditions for maintenance according to the requirements. This step is consistent with the traditional DUI platform skill maintenance process.

[0054] Step S202: Enter the multilingual voice response configuration interface and select the target language.

[0055] In step S202, among the skills that activate the "internationalized broadcast" function, the voice reply configuration is no longer limited to Chinese, but is extended to support multi-language replies. The operator enters the corresponding reply text configuration interface to prepare for subsequent language selection and text adjustment. In the reply configuration interface, the operator selects the target language to be maintained according to actual needs. The processing methods for different languages are different. Among them, for the Chinese language, the automatic single-sentence translation method is adopted, while for other languages, manual text adjustment is supported to meet the expression habits and quality requirements of different languages.

[0056] Step S203: If it is the Chinese language, automatic single-sentence translation is performed; if it is other languages, manual correction is performed.

[0057] In step S203, after the operator selects the Chinese language, in the voice reply text configuration area (such as a text input box), the system automatically displays the "Translate" button. After the operator clicks this button, the "rule translation" ability of the Sibyl-Dongfeng large model is called to perform automatic single-sentence translation on the Chinese reply text of the current sentence, generating the reply text corresponding to the target language and realizing efficient automatic processing. For non-Chinese languages selected by the operator, the system allows manual modification of the translation result in the text configuration area. This method is used to correct inaccurate or context-inconsistent expressions that may occur during the automatic translation process to ensure that the reply texts in each language meet the expected quality and expression effects.

[0058] Step S204: Skill release and use.

[0059] In step S204, after completing the text maintenance at the single-sentence level, the operator publishes the adjusted skill to the production environment according to the regular process of the DUI platform. After the skill is published and in actual use, it can support multi-language voice replies, providing users with accurate and timely interaction experiences, while ensuring the flexibility and efficiency of skill maintenance.

[0060] In summary, the present application has deduced and optimized the inherent defects of the real-time translation model solution and the multi-skill configuration solution: in the traditional real-time translation solution, even if most of the responses are fixed texts, the translation model needs to be called in real time for each interaction, resulting in a large consumption of computing resources and potential response delays; while in the multi-skill configuration solution, in order to support each language, complex dialogue logics must be configured and maintained separately, resulting in a sharp increase in labor costs. Therefore, the present application proposes a multi-language language interaction skill configuration method assisted by a large model. Among them, by using the rule translation ability of the large model through an overall conversion method, the existing Chinese skills are converted into multi-language skills with one click, and the response texts in all target languages are generated in advance, thus eliminating the repeated calculations of real-time translation; at the same time, by adopting a maintenance method that combines single-sentence automatic translation and manual correction for the existing multi-language skills, the precise adjustment of the voice response text is realized. Thereby, both the waste of server computing power is effectively reduced, and the labor costs generated by multi-skill maintenance are also greatly reduced, significantly improving the overall operation efficiency of the voice interaction system.

[0061] Figure 3 FIG. is a structural block diagram of a multi-language language interaction skill configuration system assisted by a large model provided by an embodiment of the present application. The device at least includes the following several modules:

[0062] An overall conversion module for existing Chinese skills, which is used for the overall conversion of existing Chinese skills, creating skills, turning on the international broadcast switch and selecting the languages to be supported, configuring the Chinese skill content, performing the configuration conversion from Chinese to multi-languages based on the large model capabilities, and releasing and using the skills;

[0063] A maintenance module for existing multi-language skills, which is used for the maintenance of existing multi-language skills, turning on the skills and selecting the dialogue conditions to be maintained, entering the multi-language voice response configuration interface and selecting the target language, automatically performing single-sentence translation if it is the Chinese language, and performing manual correction if it is other languages, and releasing and using the skills.

[0064] For relevant details, refer to the above method embodiment.

[0065] Figure 4 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.

[0066] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0067] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the configuration method of the multi-lingual language interaction skill based on large model assistance provided in the method embodiments of the present application.

[0068] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.

[0069] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.

[0070] Optionally, the present application also provides a computer-readable storage medium, and a program is stored in the computer-readable storage medium. The program is loaded and executed by the processor to implement the configuration method of the multi-lingual language interaction skill based on large model assistance in the above method embodiments.

[0071] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium. A program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the configuration method of the multi-language language interaction skill assisted by a large model in the above method embodiment.

[0072] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0073] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A configuration method for multi - language interaction skills assisted by large models, characterized in that, The method includes: For the overall conversion of existing Chinese skills, create a skill, turn on the international broadcast switch and select the languages to be supported, configure the Chinese skill content, perform the conversion of Chinese to multiple languages based on the capabilities of the large model, and release and use the skill; For the maintenance of existing multi-language skills, open the skill and select the dialogue conditions to be maintained, enter the multi-language voice response configuration interface and select the target language. If it is the Chinese language, automatic single-sentence translation will be performed. If it is other languages, manual correction will be carried out, and then release and use the skill.

2. The configuration method of the multi-language language interaction skill assisted by the large model according to claim 1, wherein, For the overall conversion of existing Chinese skills, creating a skill includes: On the skill customization page of the DUI platform, click the customize skill button, and configure the skill type and skill name according to requirements to create a new voice skill.

3. The configuration method of the multi-language language interaction skill assisted by the large model according to claim 1, wherein The configuration of the Chinese skill content includes: According to the Chinese voice skill configuration process of the DUI platform, configure the Chinese semantic content and dialogue logic of the skill.

4. The configuration method of the multi - language language interaction skill assisted by the large - model according to claim 1, characterized in that, The conversion of Chinese to multiple languages based on the capabilities of the large model includes: After the international broadcast function is enabled, the international broadcast – NLG full translation button is displayed on the skill home page; click the button to call the rule translation ability of the large model to perform an overall conversion of all Chinese response texts and related semantic information in the skill, and generate multi-language response content in the corresponding target language.

5. The configuration method of multi - language interaction skills assisted by large models according to claim 1, wherein, Entering the multi-language voice response configuration interface and selecting the target language includes: Enter the corresponding response text configuration interface, and in the response configuration interface, select the target language to be maintained according to actual needs.

6. The configuration method of the multi - language language interaction skill assisted by the large model according to claim 1, wherein, If it is the Chinese language, automatic single-sentence translation will be performed, including: After selecting the Chinese language, in the voice response text configuration area, the translation button will be automatically displayed; click the button to call the rule translation ability of the large model to perform automatic single-sentence translation of the Chinese response text of the current sentence, and generate the corresponding response text in the target language.

7. The configuration method of the multi - language language interaction skill assisted by a large - model according to claim 1, wherein, If it is other languages, manual correction will be carried out, including: For non-Chinese languages selected by the operator, manual modification of the translation result is allowed in the text configuration area.

8. A configuration system for multi - language interaction skills assisted by large models, characterized in that, including: An overall conversion module for existing Chinese skills, which is used for the overall conversion of existing Chinese skills, creating a skill, turning on the international broadcast switch and selecting the languages to be supported, configuring the Chinese skill content, performing the conversion of Chinese to multiple languages based on the capabilities of the large model, and releasing and using the skill; A maintenance module for existing multi-language skills, which is used for the maintenance of existing multi-language skills, opening the skill and selecting the dialogue conditions to be maintained, entering the multi-language voice response configuration interface and selecting the target language. If it is the Chinese language, automatic single-sentence translation will be performed. If it is other languages, manual correction will be carried out, and then release and use the skill.

9. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a configuration method for a multi-language language interaction skill assisted by a large model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and when the program is executed by the processor, it is used to implement a configuration method for a multi-language language interaction skill assisted by a large model as described in any one of claims 1 to 7.