Spoken language learning method and related device
By training the target large language model to infer multi-role dialogue simulation data and mixed language data in a thinking chain, combined with self-distillation supervision fine-tuning, the problem of insufficient accuracy and coherence of dialogue reply in the existing oral learning methods is solved, and a more efficient oral learning experience is achieved.
Patent Information
- Application Number
- CN202510462039.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-08
AI Technical Summary
Among the existing oral learning methods, the accuracy and coherence of dialogue responses between language learners and large language models are poor, especially when facing mixed language dialogue data, output errors occur frequently and the language level of language learners are not considered, resulting in inconsistent dialogue responses with the expected language level.
The multi-role dialogue simulation data inferred by the thinking chain method is trained in the target large language model, combined with the mixed language dialogue simulation data for training, and fine-tuning the optimization model through self-distillation supervision to generate dialogue data for thinking chain thinking, and provide learner dialogue prompts during the oral dialogue process to improve the consistency and accuracy of the dialogue.
It improves the dialogue context coherence and consistency between language learners and large language models, adapts to learning needs of different language difficulties, provides a realistic oral learning experience, solves the accuracy problem under mixed language dialogue data, and enhances the oral practice effect of language learners.
Smart Images

Figure CN120279791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a spoken language learning method and related devices. Background Art
[0002] With the acceleration of the globalization process and the increasing frequency of cross-cultural exchanges, the importance of multilingual ability has become more prominent. At the same time, in the past two years, large language models have been widely used in various dialogue systems, translation systems, and education systems due to their excellent performance in natural language generation and understanding. In this context, how to have a coherent and accurate dialogue with language learners through large language models to help language learners efficiently improve their spoken language dialogue ability has become one of the hot research directions in the education field. Summary of the Invention
[0003] In view of the above problems, this application provides a spoken language learning method and related devices to achieve the purpose of having a coherent and accurate dialogue with language learners through large language models. The specific solutions are as follows:
[0004] The first aspect of this application provides a spoken language learning method, including:
[0005] Obtain the target spoken language learning scenario and target language difficulty input by the language learner;
[0006] Invoke the target large language model to have at least one round of spoken language dialogue with the language learner based on the target spoken language learning scenario and the target language difficulty;
[0007] Among them, the training data of the target large language model includes multi-role dialogue simulation data inferred in a chain-of-thought manner, and the learner dialogue simulation data in the multi-role dialogue simulation data includes mixed-language dialogue simulation data.
[0008] In a possible implementation, the process of obtaining the multi-role dialogue simulation data includes:
[0009] Obtain a first prompt instruction, which is used to instruct the first large language model to perform chain-of-thought dialogue reasoning according to a preset first language difficulty and spoken language learning scenario to obtain teacher dialogue simulation data that meets the set dialogue goal;
[0010] Obtain a second prompt instruction, which is used to instruct the second large language model to adopt a chain-of-thought dialogue reasoning strategy to obtain learner dialogue simulation data for the teacher dialogue simulation data;
[0011] Input the first prompt instruction into the first large language model, and input the second prompt instruction into the second large language model, so that the first large language model and the second large language model conduct at least one round of spoken dialogue simulation to obtain the multi-role dialogue simulation data.
[0012] In one possible implementation, the obtaining of the first prompt instruction includes:
[0013] Obtain the first language difficulty and the spoken language learning scenario;
[0014] Obtain a first prompt instruction template, where the first prompt instruction template includes a first language difficulty slot and a learning scenario slot;
[0015] Fill the first language difficulty into the first language difficulty slot, and fill the spoken language learning scenario into the learning scenario slot to obtain the first prompt instruction.
[0016] In one possible implementation, the obtaining of the second prompt instruction includes:
[0017] Obtain the preset role information and language style;
[0018] Obtain a second prompt instruction template, where the second prompt instruction template includes a role slot and a language style slot;
[0019] Fill the role information into the role slot, and fill the language style into the language style slot to obtain the second prompt instruction.
[0020] In one possible implementation, the training data of the target large language model further includes: learner dialogue prompt data inferred in the way of chain of thought;
[0021] The process of obtaining the learner dialogue prompt data includes:
[0022] Obtain the preset second language difficulty, chat background and the dialogue content under the chat background;
[0023] Obtain a third prompt instruction template, where the third prompt instruction template includes a chat background slot, a dialogue history slot and a second language difficulty slot, and the third prompt instruction template is used to instruct the third large language model to conduct chain-of-thought dialogue reasoning according to the chat background in the chat background slot, the dialogue content in the dialogue history slot and the second language difficulty in the second language difficulty slot to obtain learner dialogue prompt data that meets the set dialogue requirements;
[0024] Fill the chat background into the chat background slot, fill the dialogue content under the chat background into the dialogue history slot, and fill the second language difficulty into the second language difficulty slot to obtain the third prompt instruction;
[0025] Input the third prompt instruction into the third large language model to obtain the learner dialogue prompt data.
[0026] In a possible implementation, the oral learning method further includes:
[0027] During the oral dialogue with the language learner, if a dialogue prompt request is detected, obtain the target oral learning scenario, the target language difficulty, and the historical dialogue data in the oral dialogue process;
[0028] Call the target large language model to obtain target learner dialogue prompt data according to the target oral learning scenario, the target language difficulty, and the historical dialogue data;
[0029] Output and display the target learner dialogue prompt data.
[0030] In a possible implementation, it further includes:
[0031] Eliminate the data that does not meet the data quality requirements in the multi-role dialogue simulation data and the learner dialogue prompt data, and the remaining data is used as the training data.
[0032] In a possible implementation, the target large language model is obtained by performing self-distillation supervised fine-tuning on a pre-trained Decode-Only large language model using the training data, and the cross-entropy loss and KL divergence regularization loss are used in the process of self-distillation supervised fine-tuning.
[0033] A second aspect of the present application provides an oral learning device, including:
[0034] A dialogue environment acquisition module, configured to acquire a target oral learning scenario and a target language difficulty input by a language learner;
[0035] A human-machine dialogue module, configured to call a target large language model to perform at least one round of oral dialogue with the language learner based on the target oral learning scenario and the target language difficulty, where the training data of the target large language model includes multi-role dialogue simulation data inferred in a chain-of-thought manner, and the learner dialogue simulation data in the multi-role dialogue simulation data includes mixed-language dialogue simulation data.
[0036] A third aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, cause the electronic device to implement the oral learning method in the first aspect or any implementation manner of the first aspect.
[0037] A fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:
[0038] The memory is used to store a computer program;
[0039] The processor is used to execute the computer program so that the electronic device can implement the oral language learning method in the above first aspect or any implementation manner of the first aspect.
[0040] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the oral language learning method in the above first aspect or any implementation manner of the first aspect.
[0041] By means of the above technical solution, for the oral language learning method provided by the present application, the target oral language learning scenario and the target language difficulty input by the language learner are obtained, and the target large language model is called to have at least one round of oral conversation with the language learner based on the target oral language learning scenario and the target language difficulty. Since the training data used by the target large language model in the training process is the data inferred in the way of chain of thought, when the target large language model has an oral conversation with the language learner, it can also think in the way of chain of thought, improving the coherence and consistency of the dialogue context.
[0042] Furthermore, when the target large language model conducts chain-of-thought thinking, in addition to considering the dialogue data of the language learner, it only needs to consider the information in two dimensions of the target oral language learning scenario and the target language difficulty. The amount of information is small, reducing the task difficulty of chain-of-thought thinking and improving the accuracy of the dialogue output by the target large language model. At the same time, the present application supports the language learner to input different language difficulties, which can adapt to the oral language learning of the language learner with different language difficulties and improve the oral language learning experience of the language learner.
[0043] Even further, considering that the language learner often outputs mixed-language dialogue data when their language level is insufficient, in order to enable the target large language model to still have an accurate dialogue with the language learner in the face of mixed-language dialogue data, the present application uses mixed-language dialogue simulation data as the learner dialogue simulation data to train the target large language model, further improving the accuracy of the dialogue output by the target large language model and providing a more practical oral language learning experience for the language learner. Description of the Drawings
[0044] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original and elements are not necessarily drawn to scale.
[0045] Figure 1 A schematic diagram of a system architecture provided by this application;
[0046] Figure 2 An optional hardware structure diagram of the terminal 100 provided by this application;
[0047] Figure 3 A structural diagram of a server 200 provided by this application;
[0048] Figure 4 A flowchart of a spoken language learning method provided by this application;
[0049] Figure 5 A flowchart of a spoken language conversation between a target large language model and a language learner provided by this application;
[0050] Figure 6 A flowchart of a conversation prompting process provided by this application;
[0051] Figure 7 A structural diagram of a spoken language learning device provided by this application;
[0052] Figure 8 A structural diagram of an electronic device provided by this application. Specific Embodiments
[0053] The embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0054] The embodiments of the present application will be described below in conjunction with the accompanying drawings. Those of ordinary skill in the art will appreciate that as technology develops and new scenarios emerge, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0055] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0056] Existing oral learning methods mainly include an oral learning method based on intent recognition and template matching and an oral learning method based on prompt engineering to call a large language model.
[0057] The oral learning method based on intent recognition and template matching is as follows: collect a large amount of monolingual dialogue data, train an intent recognition model to identify the intent category input by the language learner, then retrieve the corresponding template response from a preset dialogue template library according to the intent category, and finally perform sentence generation and optimization on the retrieved template response to obtain a dialogue response and output it to the language learner.
[0058] The oral learning method based on prompt engineering to call a large language model is as follows: construct prompts containing multi-dimensional information such as dialogue scenarios, role settings, and task descriptions, then combine the user input with the prompts to form an input for calling the large language model, and directly generate the corresponding dialogue response by the large language model.
[0059] However, the above-mentioned oral learning methods based on intent recognition and template matching and oral learning methods based on prompt engineering to call a large language model have the following problems: (1) The accuracy and coherence of the dialogue responses output to the language learner are poor. Especially for the oral learning method based on intent recognition and template matching, when facing the mixed-language dialogue data output by the language learner, it often outputs incorrect dialogue responses. For the oral learning method based on prompt engineering to call a large language model, due to the multi-dimensional information constraints in the prompts, the large language model often shows a gradual weakening of the instruction-following ability in multi-round dialogues; (2) The language level of the language learner is not considered when generating dialogue responses, resulting in the language level of the output dialogue responses being inconsistent with the language level expected by the language learner.
[0060] In view of the above problems, the present application provides a spoken language learning method, which can be applicable to the scenario where a language learner practices spoken language with a spoken language dialogue device. A target large language model is installed on the spoken language dialogue device, and the spoken language dialogue device can use the target large language model to conduct at least one round of spoken language dialogue with the language learner. The spoken language learning method of the present application can be applied to, for example, Figure 1 the system architecture shown in the figure. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 illustrated by taking one server as an example), and the terminal 100 is the above-mentioned spoken language dialogue device.
[0061] The terminal 100 can be used alone to execute the spoken language learning method provided by the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the spoken language learning method provided by the embodiments of the present application.
[0062] Next, the product form of the terminal 100 will be described. Figure 1 in the figure;
[0063] The terminal 100 in the embodiments of the present application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present application do not make any restrictions on this.
[0064] Figure 2 shows an optional schematic diagram of the hardware structure of the terminal 100.
[0065] Referring to Figure 2 shown in the figure, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160, a speaker 161, a microphone 162, a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art can understand that Figure 2 is only an example of a terminal or a multifunctional device, and does not constitute a limitation on the terminal or the multifunctional device. It may include more or fewer components than those shown in the figure, or combine certain components, or different components.
[0066] The input unit 130 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function control of the portable multifunctional device. Specifically, the input unit 130 can include a touch screen 131 and / or other input devices 132. The touch screen 131 can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object such as a finger, a joint, a stylus, etc. on or near the touch screen), and drive corresponding connection devices according to a preset program. The touch screen can detect the touch action of the user on the touch screen, convert the touch action into a touch signal and send it to the processor 170, and can receive and execute commands sent by the processor 170; the touch signal at least includes contact coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch screen. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, the other input devices 132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0067] Among them, the other input devices 132 can receive input data and the like.
[0068] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, an interactive interface, file display, and / or the playing of any multimedia file. In the embodiment of the present application, the display unit 140 can be used to display each interactive interface, processing result, etc. in the oral language learning method.
[0069] The memory 120 can be used to store instructions and data. The memory 120 mainly includes a storage instruction area and a storage data area. The storage data area can store various data, such as multimedia files, texts, etc.; the storage instruction area can store software units such as an operating system, an application, instructions required for at least one function, or their subsets or extended sets. It can also include a non-volatile random access memory; it provides the processor 170 with management of the hardware, software, and data resources in the computing processing device, supports control software and applications. It is also used for the storage of multimedia files and the storage of running programs and applications.
[0070] The processor 170 is the control center of the terminal 100, connecting various parts of the entire terminal 100 through various interfaces and circuits. By running or executing instructions stored in the memory 120 and calling data stored in the memory 120, it executes various functions of the terminal 100 and processes data, thereby exercising overall control over the terminal device. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 170 either. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be separately implemented on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing and processing device, read and process data in the software, especially read and process the data and programs in the memory 120, so that each functional module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0071] Among them, the memory 120 can be used to store software codes related to the oral learning method, and the processor 170 can execute the steps of the oral learning method or schedule other units (such as the above input unit 130 and display unit 140) to implement corresponding functions.
[0072] The radio frequency unit 110 (optional) can be used for receiving and transmitting information or signals during a call. For example, after receiving the downlink information from the base station, it is sent to the processor 170 for processing. Additionally, the data designed for uplink is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0073] Among them, in the embodiment of the present application, the radio frequency unit 110 can send data to the server 200 and receive the processing result sent by the server 200.
[0074] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network interface.
[0075] The terminal 100 also includes a power supply 190 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0076] The terminal 100 also includes an external interface 180. This external interface can be a standard Micro USB interface or a multi-pin connector, and can be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100.
[0077] Although not shown, the terminal 100 may also include a flashlight, a Wireless Fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here. Some or all of the methods described below can be applied to the terminal 100 as Figure 2 shown.
[0078] Next, the product form of the server 200 will be described. Figure 1 The product form of the server 200 in Figure 1
[0079] Figure 3 A schematic structural diagram of a server 200 is provided, as Figure 3 shown. The server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other through the bus 201.
[0080] The bus 201 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 3 only a thick line is used to represent it in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0081] The processor 202 can be any one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP), etc.
[0082] The memory 204 can include a volatile memory, such as a Random Access Memory (RAM). The memory 204 can also include a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD).
[0083] Among them, the memory 204 can be used to store software codes related to the oral learning method, and the processor 202 can execute the steps of the oral learning method of the chip, or can also schedule other units to implement corresponding functions.
[0084] It should be understood that the above terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), general-purpose processor, DSP, microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware system without the function of executing instructions and the hardware system with the function of executing instructions.
[0085] Next, the oral learning method provided by the embodiments of the present application will be introduced. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 1 the terminal 100 in Figure 4 , or a system composed of the terminal 100 and the server 200. Referring to
[0086] Step S401: Obtain the target oral learning scenario and target language difficulty input by the language learner.
[0087] Optionally, this embodiment can provide a user interface, and set an oral learning scenario input box and a language difficulty input box on the user interface, so that the language learner can input the target oral learning scenario in the oral learning scenario input box and input the target language difficulty in the language difficulty input box.
[0088] Here, the target oral learning scenario refers to the conversation topic that the language learner wants to practice this time; the target language difficulty refers to the conversation difficulty that the language learner wants to practice this time. This conversation difficulty can be equivalent to the language learner's own language level, or higher or lower than the language learner's own language level, and can specifically be set by the language learner according to their own needs.
[0089] To avoid the language learner from inputting invalid information, optionally, the oral learning scenario input box can be provided with drop-down options so that the language learner can input the preset conversation topics; optionally, the drop-down options can include a custom option so that the user can freely set the conversation topics.
[0090] Similarly, optionally, the language difficulty input box can be provided with drop-down options for language learners to input preset language difficulties, such as coarse-grained language difficulties like primary school, junior high school, high school, etc., or fine-grained language difficulties like beginner, intermediate, advanced, etc., and so on. The present application does not make specific limitations. For example, if a language learner is not clear about their own language level, they can input a coarse-grained language difficulty; otherwise, they can input a fine-grained language difficulty to practice speaking more targeted.
[0091] In another possible implementation, the drop-down options of the language difficulty input box can also include a language proficiency test option. When the user clicks on the language proficiency test option, the present application can detect the language proficiency of the language learner based on the spoken language input by the language learner and set the target language difficulty according to the detection result.
[0092] Optionally, the process of "detecting the language proficiency of the language learner based on the spoken language input by the language learner" can include: scoring the spoken language input by the language learner from multiple perspectives, and performing weighted summation based on the scores and weights from multiple perspectives to obtain the language proficiency of the language learner.
[0093] Optionally, the multiple perspectives include at least two of the following perspectives: pronunciation and intonation, vocabulary and grammar, fluency, and content depth. Among them, pronunciation and intonation include: pronunciation accuracy (such as the accuracy of vowels, consonants, liaison, and weak reading), intonation naturalness (such as the naturalness of rising tone, falling tone, and stress), and rhythm (such as the rhythm of pauses and morpheme control); vocabulary and grammar include: vocabulary size (such as basic vocabulary, high-frequency vocabulary, academic vocabulary, etc.), grammar accuracy (such as the accuracy of tense, sentence pattern, and clause usage), and expression richness (such as synonym replacement, idiom usage, etc.); fluency includes: pause frequency (such as the number of self-corrections), speech rate stability (such as being too fast, too slow, or stuttering), and thinking coherence (such as logical jumps and repeated expressions); content depth includes: topic expansion ability (such as detailed description, example illustration, etc.), opinion expression (such as critical thinking, dialectical analysis, etc.), and cultural adaptability (such as the understanding of idioms, slang, and metaphors).
[0094] It should be noted that the above way of setting drop-down options is only an example. For example, the spoken language learning scenario input box and the language difficulty input box can also be voice input boxes or handwriting input boxes, etc., and are not limited to the present application.
[0095] Step S402: Invoke the target large language model to have at least one round of spoken language conversation with the language learner based on the target spoken language learning scenario and the target language difficulty.
[0096] It can be understood that in some spoken dialogue scenarios, the target large language model can initiate the dialogue first. Then, in this embodiment, a first target prompt instruction can be generated based on the target spoken language learning scenario and the target language difficulty, and the first target prompt instruction can be input into the target large language model so that the target large language model can have a round of spoken dialogue with the language learner to obtain the first dialogue data input by the language learner. Further, a second target prompt instruction can be generated based on the target spoken language learning scenario, the target language difficulty, and the first dialogue data, and the second target prompt instruction can be input into the target large language model so that the target large language model can have a second round of spoken dialogue with the language learner to obtain the second dialogue data input by the language learner. A third target prompt instruction can be generated based on the target spoken language learning scenario, the target language difficulty, and the second dialogue data, and the third target prompt instruction can be input into the target large language model so that the target large language model can have a third round of spoken dialogue with the language learner; and so on until the dialogue ends.
[0097] In some other spoken dialogue scenarios, the language learner can initiate the dialogue first. Then, in this embodiment, the dialogue data input by the language learner can be obtained, and a target prompt instruction can be generated based on the target spoken language learning scenario, the target language difficulty, and the dialogue data input by the language learner so that the target large language model can have a dialogue with the language learner. This process is similar to the related process of "the target large language model initiates the dialogue first" above. For details, reference can be made to the above introduction and will not be elaborated here.
[0098] In this embodiment, the training data of the target large language model includes multi-role dialogue simulation data inferred in a chain-of-thought manner, and the learner dialogue simulation data in the multi-role dialogue simulation data can include mixed-language dialogue simulation data. Here, the mixed-language dialogue simulation data refers to that in the learner dialogue simulation data of a round of dialogue, multiple languages are included.
[0099] For example, the mixed-language dialogue simulation data with in-sentence switching is as follows:
[0100] "English-Chinese mixed (English-dominant): Where is my schoolbag?", or, "Chinese-English mixed (Chinese-dominant): Where is my backpack?".
[0101] The mixed-language dialogue simulation data with between-sentence switching is as follows:
[0102] "Where is my schoolbag? Do you know?"
[0103] Optionally, in order to improve the dialogue effect of the target large language model, the learner dialogue simulation data in the multi-role dialogue simulation data can also include: single-language dialogue simulation data.
[0104] For example, the single-language dialogue simulation data is as follows:
[0105] "Fully in Chinese: Where is my schoolbag?", or, "Fully in English: Where is my backpack?".
[0106] Considering that language learners may input normal or abnormal dialogue data during oral practice, examples of normal dialogue data are as shown in the above mixed-language dialogue simulation data and single-language dialogue simulation data, and examples of abnormal dialogue data are like "Where is my my my backpack?". To enable the target large language model to accurately identify abnormal dialogue data, optionally, the learner dialogue simulation data in the multi-role dialogue simulation data can include normal dialogue data and abnormal dialogue data.
[0107] Optionally, the above multi-role dialogue simulation data can include: teacher dialogue simulation data and learner dialogue simulation data. Taking the scenario of a language learner practicing English oral skills as an example, the teacher dialogue simulation data are all fully English dialogue simulation data, while the learner dialogue simulation data can be the above mixed-language dialogue simulation data or the above single-language dialogue simulation data.
[0108] Of course, the multi-role dialogue simulation data can also include dialogue simulation data of other roles, which is not specifically limited in this application.
[0109] In summary, the oral learning method provided in this application obtains the target oral learning scenario and target language difficulty input by the language learner, and calls the target large language model to have at least one round of oral dialogue with the language learner based on the target oral learning scenario and target language difficulty. Since the training data used by the target large language model during training is data inferred in a chain-of-thought manner, when the target large language model has an oral dialogue with the language learner, it can also perform chain-of-thought thinking, improving the coherence and consistency of the dialogue context.
[0110] Furthermore, when the target large language model performs chain-of-thought thinking, in addition to considering the dialogue data of the language learner, it only needs to consider the information in two dimensions: the target oral learning scenario and the target language difficulty. The amount of information is small, reducing the task difficulty of chain-of-thought thinking and improving the accuracy of the dialogue output by the target large language model. At the same time, this application supports the language learner to input different language difficulties, can adapt to the oral learning of the language learner with different language difficulties, and improves the oral learning experience of the language learner.
[0111] Furthermore, considering that language learners often output mixed-language dialogue data when their language proficiency is insufficient, in order to enable the target large language model to accurately converse with language learners when faced with mixed-language dialogue data, this application uses mixed-language dialogue simulation data as the learner dialogue simulation data to train the target large language model, further improving the accuracy of the dialogue output by the target large language model and providing a more practical oral learning experience for language learners.
[0112] In some embodiments of this application, taking the example that the multi-role dialogue simulation data may include teacher role dialogue simulation data and learner dialogue simulation data, the process of obtaining the multi-role dialogue simulation data will be introduced.
[0113] Considering that there is currently less mixed-language dialogue simulation data, in order to obtain a sufficient amount of mixed-language dialogue simulation data, this embodiment provides a method that can automatically generate a large batch of mixed-language dialogue simulation data.
[0114] Optionally, two large language models with strong instruction-following capabilities can be used to role-play and converse with each other to obtain a sufficient amount of mixed-language dialogue simulation data. Optionally, one of the two large language models is Role 1: an English speaking teacher. The other large language model is Role 2: students with different portraits, such as primary school students, junior high school students, etc., or more fine-grained background information can be specified, which is not specifically limited in this application.
[0115] In this embodiment, different prompt instructions can be set for the above two large language models respectively, so that under the prompt of the prompt instructions, the two large language models can simulate humans to converse with each other and obtain a sufficient amount of mixed-language dialogue simulation data.
[0116] Based on this, optionally, the process of obtaining the multi-role dialogue simulation data may include: obtaining a first prompt instruction, which is used to instruct the first large language model to perform a thought-chain dialogue reasoning according to a preset first language difficulty and oral learning scenario to obtain teacher dialogue simulation data that meets the set dialogue goal; obtaining a second prompt instruction, which is used to instruct the second large language model to adopt a thought-chain dialogue reasoning strategy to obtain learner dialogue simulation data for the teacher dialogue simulation data; inputting the first prompt instruction into the first large language model and inputting the second prompt instruction into the second large language model, so that the first large language model and the second large language model perform at least one round of oral dialogue simulation to obtain multi-role dialogue simulation data.
[0117] In a possible implementation, the process of "obtaining the first prompt instruction" may include: obtaining the first language difficulty, the oral learning scenario, and the first prompt instruction template, where the first prompt instruction template includes a first language difficulty slot and a learning scenario slot. The first prompt instruction template is used to instruct the first large language model to perform a thought chain dialogue reasoning based on the first language difficulty in the first language difficulty slot and the oral learning scenario in the learning scenario slot, so as to obtain teacher dialogue simulation data that meets the set dialogue goal; filling the first language difficulty into the first language difficulty slot and filling the oral learning scenario into the learning scenario slot to obtain the first prompt instruction.
[0118] Optionally, the process of obtaining the oral learning scenario may include: randomly or sequentially obtaining the oral learning scenario from a preset set of oral learning scenarios; or obtaining a custom oral learning scenario.
[0119] It should be noted that in addition to the first language difficulty slot and the learning scenario slot, the above first prompt instruction template may further include other information slots, such as a dialogue goal slot, a role information slot, etc., which are not specifically limited in this application.
[0120] It should also be noted that the above process of obtaining the first prompt instruction is only an example. In addition, there may be other implementation manners. For example, multiple first prompt instructions are pre-generated, and when it is necessary to obtain the first prompt instruction, one first prompt instruction is randomly or selected as required from the multiple first prompt instructions.
[0121] In a possible implementation, the process of "obtaining the second prompt instruction" may include: obtaining the preset role information and language style; obtaining the second prompt instruction template, where the second prompt instruction template includes a role slot and a language style slot. The second prompt instruction template is used to instruct the second large language model to perform a thought chain dialogue reasoning based on the role information in the role slot and the language style in the language style slot, so as to obtain learner dialogue simulation data for the teacher dialogue simulation data; filling the role information into the role slot and filling the language style into the language style slot to obtain the second prompt instruction.
[0122] Optionally, the process of obtaining the role information may include: randomly or sequentially obtaining the role information from a preset set of role information; or obtaining custom role information; similarly, the process of obtaining the language style may include: randomly or sequentially obtaining the language style from a preset set of language styles; or obtaining a custom language style.
[0123] The above language style may be, for example, a single language style, a bilingual style, etc., which are not specifically limited in this application.
[0124] It should be noted that in addition to the role slot and language style slot, the above second prompt instruction template may also include other information slots, which are not specifically limited in this application.
[0125] It should also be noted that the above process of obtaining the second prompt instruction is only an example. In addition, there may be other implementation methods. For example, multiple second prompt instructions are pre-generated, and when a second prompt instruction needs to be obtained, one second prompt instruction is randomly or selected as required from the multiple second prompt instructions.
[0126] It should further be noted that the prompt instruction following capabilities of both the first large language model and the second large language model meet the expectations, so that both the first large language model and the second large language model can accurately respond and output based on the prompt instructions they receive.
[0127] Taking the scenario of Chinese students practicing English speaking as an example, an optional first prompt instruction template and second prompt instruction template are as follows respectively.
[0128] First prompt instruction template: "Role setting: A patient and professional English speaking teacher to help students improve their English speaking ability. Language style: Pure English. Conversation goal: Guide students to express in English, correct grammar mistakes, provide encouragement and feedback. Language difficulty: {{Language_level}}. The conversation topic is: {{Topic}}. Output the reasoning process before outputting
Reply
[0129] Second prompt instruction template: "Role setting: {{User_profile}}. Language style: {{Mix_style}}".
[0130] Among them, the slot contents in the first language difficulty slot {{Language_level}}, learning scenario slot {{Topic}}, role slot {{User_profile}}, and language style slot {{Mix_style}} can all be customized, which are not limited in this application.
[0131] By customizing a variety of contents in the slots of the above two prompt instruction templates, sufficient teacher conversation simulation data and learner conversation simulation data can be obtained under the interaction between the first large language model and the second large language model. Since the second large language model simulates a Chinese student having a conversation with an English speaking teacher, and Chinese students usually output a large amount of mixed language conversation simulation data due to their limited English speaking ability, the learner conversation simulation data output by the second large language model can include a large amount of mixed language conversation simulation data.
[0132] Considering that some of the multi-role dialogue simulation data generated based on the first language model and the second language model may not meet the data quality requirements, for example, in a scenario where Chinese students practice oral English, the teacher dialogue simulation data contains non-English data, then it can be considered that the multi-role dialogue simulation data containing the teacher dialogue simulation data does not meet the data quality requirements. In order to train a target large language model with better dialogue effect, optionally, this embodiment can remove the data that does not meet the data quality requirements from the multi-role dialogue simulation data, and use the remaining multi-role dialogue simulation data as training data.
[0133] Optionally, the first language model and the second language model can both be Spark-4.0-Ultra (the flagship version in the Spark large model series, focusing on improving office productivity and complex task processing capabilities) language model. Of course, the first language model and the second language model can also be others, such as GPT-4, etc., which are not specifically limited in this application.
[0134] It should also be noted that the first large language model and the second large language model can be the same large language model or two different large language models, which is not limited in this application. In addition, the first prompt instruction template and the second prompt instruction template are only examples and are not intended to limit the embodiments of this application.
[0135] In summary, this embodiment sets roles for the large language model so that the large language model can simulate human thinking for conversation. At the same time, in order to avoid problems such as context incoherence, irrelevant answers, and inconsistencies in the process of simulating human conversations by the large language model, this embodiment can enable the large language model to generate chain-of-thought data with reasoning while generating conversations, thereby improving the coherence and consistency of the conversation context.
[0136] In other embodiments of the present application, after careful study of the oral dialogue process of language learners, it is found that during the oral dialogue process, language learners sometimes cannot output dialogues due to lack of expression, or the output cannot form a complete dialogue due to stumbling and other issues.
[0137] In order to improve the fluency and completeness of the language learner's dialogue, this embodiment can add learner dialogue prompt data inferred by using the thought chain method to the training data of the target large language model when training the target large language model.
[0138] Since there is little existing learner dialogue prompt data, in order to obtain sufficient learner dialogue prompt data, this embodiment provides a method for automatically generating learner dialogue prompt data in large quantities.
[0139] Optionally, the process of obtaining the learner dialogue prompt data may include: obtaining a third prompt instruction, which is used to instruct a third large language model to perform a chain-of-thought dialogue reasoning based on a preset chat background, the dialogue content in the chat background, and a second language difficulty, so as to obtain learner dialogue prompt data that meets the set dialogue requirements; inputting the third prompt instruction into the third large language model to obtain the learner dialogue prompt data.
[0140] Optionally, the process of "obtaining the third prompt instruction" may include: obtaining a second language difficulty, a chat background, the dialogue content in the chat background, and a third prompt instruction template. The third prompt instruction template includes a chat background slot, a dialogue history slot, and a second language difficulty slot, and is used to instruct the third large language model to perform a chain-of-thought dialogue reasoning based on the chat background in the chat background slot, the dialogue content in the dialogue history slot, and the second language difficulty in the second language difficulty slot, so as to obtain learner dialogue prompt data that meets the set dialogue requirements; filling the chat background into the chat background slot, filling the dialogue content in the chat background into the dialogue history slot, and filling the second language difficulty into the second language difficulty slot to obtain the third prompt instruction.
[0141] Optionally, the above preset chat background may be a high-frequency background, such as topic discussion, picture-based dialogue, etc. Of course, the chat background may also be a low-frequency background or a medium-frequency background, etc., which is not limited in this application.
[0142] Here, the prompt instruction following ability of the third large language model meets the expectation. For example, the third large language model may be the Spark-4.0-Ultra language model, GPT-4, etc.
[0143] It should be noted that the third large language model may be the same large language model as the first large language model or the second large language model mentioned above, or may be different from the first large language model and the second large language model mentioned above. This application does not make specific limitations.
[0144] An optional third prompt instruction template is as follows.
[0145] "Based on the
Chat Background
ConversationHistory
reply
[0146] In this embodiment, the slot contents in the chat background slot {{Chat Background}}, the conversation history slot {{Conversation History}}, and the second language difficulty slot {{Language_level}} can all be customized based on existing real conversation scene data, and this application does not limit this.
[0147] By customizing a variety of content in the slots of the third prompt instruction template as described above, it is possible to utilize existing real dialogue scene data and construct sufficient learner dialogue prompt data (i.e., the next round of USER dialogue) through the third language model with powerful instruction following capabilities.
[0148] After obtaining a sufficient amount of learner dialogue prompt data, this embodiment may use the learner dialogue prompt data as training data to train the target large language model.
[0149] Considering that some of the learner dialogue prompt data generated based on the third language model may not meet the data quality requirements, for example, the second language difficulty in the third prompt instruction is simple (generally, the length of simple learner dialogue data is very short), but the length of the generated learner dialogue prompt data is very long, then it can be considered that the learner dialogue prompt data does not meet the data quality requirements. In order to train a target large language model with better dialogue effect, optionally, this embodiment can remove the data that does not meet the data quality requirements from the learner dialogue prompt data, and the remaining learner dialogue prompt data is used as training data to train the target large language model.
[0150] Since the embodiment of the present application enables the third large language model to generate learner dialogue prompt data while generating reasoning process data, the output of the learner dialogue prompt data is data derived through chain thinking. Using the learner dialogue prompt data as training data to train the target large language model can enhance the ability of the target large language model to think first and then answer, and effectively solve problems such as inconsistencies and repeated dialogues.
[0151] Furthermore, after the target large language model is trained, the present embodiment can detect whether there is a dialogue prompt request during the oral dialogue with the language learner in the previous step S402. If a dialogue prompt request is detected, the target oral learning scenario, the target language difficulty, and the historical dialogue data during the oral dialogue are obtained, the target large language model is called to obtain the target learner dialogue prompt data according to the target oral learning scenario, the target language difficulty, and the historical dialogue data, and the target learner dialogue prompt data is output and displayed.
[0152] Optionally, in this embodiment, a context manager may be preset to store the historical conversation data during the spoken conversation. That is, when the language learner has a conversation with the target large language model, the conversation data of each round is stored as historical conversation data in the context manager. When this embodiment detects a conversation prompt request, it can call the historical conversation data during the spoken conversation from the context manager.
[0153] Optionally, this embodiment may provide a user interface and set a conversation prompt button on the user interface, so that the language learner can click the conversation prompt button, and then this embodiment can respond to the operation of the language learner clicking the conversation prompt button to generate a conversation prompt request.
[0154] Of course, the above-mentioned conversation prompt button is only an example. In addition, there may be other implementation manners, such as setting a conversation prompt voice input box on the user interface, etc. This application does not make a limitation.
[0155] Optionally, the process of "outputting and displaying the target learner's conversation prompt data" may include: outputting and displaying the target learner's conversation prompt data in text form, and at the same time outputting the target learner's conversation prompt data as voice, so that the language learner can learn the pronunciation.
[0156] For example, if the target learner's conversation prompt data is the target learner's conversation prompt text, then the target learner's conversation prompt text is displayed to the language learner, and at the same time it is converted into voice data, and the voice data is output, so that the language learner can learn the pronunciation in the voice data.
[0157] In summary, this embodiment can dynamically generate a reference reply according to the current conversation environment, so that when the language learner gets stuck during the spoken conversation, it can promptly give a prompt to the language learner, so that the language learner can refer to the target learner's conversation prompt data to reply to the conversation, improving the spoken language learning experience of the language learner.
[0158] In a possible implementation, the target large language model is obtained by performing self-distillation supervised fine-tuning on a pre-trained Decode-Only large language model using the training data in the foregoing, and the cross-entropy loss and KL divergence regularization loss are used in the process of self-distillation supervised fine-tuning.
[0159] Here, self-distillation is a model optimization technique that allows the model to learn from its own output to further improve the performance of the model, which can alleviate catastrophic forgetting and improve the generalization ability of the model.
[0160] Specifically, first, a pre-trained Decode-Only large language model is selected as the base model. Here, the Decode-Only large language model is a model optimized specifically for text generation tasks, such as the GPT (Generative Pre-trained Transformer) series of large language models. Compared with traditional autoregressive generation models, it has advantages such as efficient decoding, low resource occupancy, high flexibility, and good controllability.
[0161] It should be noted that the parameter scale of the Decode-Only large language model is not limited in this embodiment. For example, parameter scales such as 7B (7B means the model contains 7 billion trainable parameters), 13B (13B means the model contains 13 billion trainable parameters), 70B (70B means the model contains 70 billion trainable parameters), etc. are all acceptable.
[0162] In this embodiment, the pre-trained base model can be loaded as the teacher model, and the weights of the teacher model are copied as the initial parameters of the student model (the student model has the same network structure and size as the teacher model). Then, the training data (multi-role dialogue simulation data and / or learner dialogue prompt data) described above is used to perform supervised fine-tuning (SFT) training on the student model. The standard cross-entropy loss is used in the training process to minimize the difference between the output of the student model and the corresponding dialogue data in the training data. At the same time, the KL (Kullback-Leibler) divergence regularization loss is added as the self-distillation loss to measure the difference between the output of the teacher model and the output of the student model, so as to prevent the student model from deviating too much from the teacher model.
[0163] To further improve the performance of the student model, optionally, this embodiment can perform iterative training on the student model multiple times (for example, 3 times) using the training data described above.
[0164] In summary, this embodiment performs supervised fine-tuning in a self-distillation manner, which can effectively improve the understanding ability of the target large language model for multi-language mixed input, enhance the coherence and consistency of the target large language model in multi-turn conversations, and optimize the ability of the target large language model to generate user dialogue prompts.
[0165] To make those skilled in the art better understand this application, the following will introduce this application in detail in combination with an optional application scenario.
[0166] See Figure 5 , which is a flowchart of an oral dialogue between a target large language model and a language learner provided by this application.
[0167] Step S501: Obtain the target oral learning scenario and target language difficulty input by the language learner.
[0168] Here, the target oral learning scenario and target language difficulty can be stored in the context manager.
[0169] In the scenario where the target large language model starts the conversation first, the target large language model can be called to obtain model text data according to the target oral learning scenario and target language difficulty, and the model text data can be converted into model speech data through text-to-speech technology, and the model speech data can be output for oral conversation with the language learner.
[0170] Step S502: Determine whether oral speech data input by the language learner is received.
[0171] Specifically, in the scenario where the language learner starts the conversation first, in the first round of conversation, the language learner can directly start the conversation to obtain oral speech data; in non-first-round conversations, or in the scenario where the target large language model starts the conversation first, if the language learner hears the model speech data, a conversation reply can be made to the model speech data to obtain oral speech data.
[0172] The embodiments of the present application can continuously monitor whether oral speech data input by the language learner is received. If oral speech data is detected, step S503 is executed. If no oral speech data is detected within a preset duration (for example, within the preset duration after the output of the model speech data, or within the preset duration after detecting the target oral learning scenario and target language difficulty), this conversation is ended.
[0173] Step S503: Convert the oral speech data input by the language learner into oral text data through speech-to-text technology.
[0174] Optionally, the speech-to-text technology can be a multilingual speech-to-text ASR (Automatic Speech Recognition) model.
[0175] Of course, the speech-to-text technology can also be other, which is not limited in this application.
[0176] Step S504: Call the target large language model to obtain model text data according to the target oral learning scenario, target language difficulty, and oral text data.
[0177] Here, the target large language model is obtained by self-distillation supervised fine-tuning of the pre-trained Decode-Only large language model using multi-role dialogue simulation data, and the cross-entropy loss and KL divergence regularization loss are used in the process of self-distillation supervised fine-tuning.
[0178] The multi-role dialogue model data is the data inferred by the chain of thought method, and the learner dialogue simulation data in the multi-role dialogue simulation data includes mixed language dialogue simulation data.
[0179] The process of obtaining the multi-role dialogue simulation data can be referred to the previous introduction and will not be elaborated here.
[0180] Step S505: Convert the model text data into model voice data through text-to-speech technology, output the model voice data, and return to step S502.
[0181] Optionally, the text-to-speech technology can be a text-to-speech (TTS) model for speech synthesis.
[0182] Of course, the text-to-speech technology can also be other, and this application does not make a limitation.
[0183] In the above process, the oral text data and the model text data can be stored as historical dialogue data in the context manager.
[0184] It can be understood that in the above process of having at least one round of oral dialogue between the target large language model and the language learner, there may be a situation where the user needs a prompt. At this time, it can be referred to Figure 6 Execute.
[0185] Such as Figure 6 , which is a flowchart of a dialogue prompt process provided by this application.
[0186] Step S601: Continuously monitor the dialogue prompt request.
[0187] As introduced before, the language learner can click the dialogue prompt button on the user interface to generate a dialogue prompt request.
[0188] Step S602: Determine whether a dialogue prompt request is detected.
[0189] If yes, execute step S603; if no, continue to monitor through step S601.
[0190] Step S603: Obtain the target oral learning scenario, the target language difficulty, and the historical dialogue data during the oral dialogue process from the context manager.
[0191] It can be understood that the context manager can store multiple conversations of data, and the historical dialogue data in this embodiment refers to the historical dialogue data during this oral dialogue process.
[0192] Here, an oral conversation process refers to the entire process from the time when the language learner inputs the target oral learning scenario and the target language difficulty, and the language learner or the target large language model starts the conversation, until the language learner no longer inputs oral voice data.
[0193] The historical dialogue data refers to the dialogue data generated in the current spoken dialogue process before the step S603 of acquiring data from the context manager.
[0194] Step S604: calling the target large language model to obtain target learner dialogue prompt data according to the target oral learning scenario, target language difficulty and historical dialogue data.
[0195] Here, the target large language model is obtained by self-distillation supervised fine-tuning the pre-trained Decode-Only large language model (or the target large language model in step S504) using the learner's dialogue prompt data. The self-distillation supervised fine-tuning process uses cross entropy loss and KL divergence regularization loss.
[0196] The learner dialogue prompt data is the data inferred by the thinking chain method. The specific acquisition process can be referred to the previous introduction and will not be repeated here.
[0197] Step S605: Output the target learner's dialogue prompt data in text form and output it as voice.
[0198] In summary, this embodiment can construct multi-role dialogue simulation data through role-playing dialogue through a large language model, and can construct learner dialogue prompt data through a large language model. By replacing traditional manual annotations with self-generated data from the model, the data construction cost is greatly reduced and the data diversity is improved; the large language model is fine-tuned based on the constructed data so that it can understand multi-language mixed input at the same time and generate mixed language responses that conform to the context, thereby achieving seamless support for multi-language mixed input and being more in line with the oral learning needs of language learners; by introducing a multi-round dialogue dynamic context management mechanism and constructing reasoning chain data, the coherence and consistency of oral dialogue can be achieved; by designing language learner dialogue prompts, reference responses can be dynamically generated according to the language level and dialogue environment of the language learner, thereby solving the problem of language learner dialogue jamming.
[0199] The spoken language learning device provided in the embodiment of the present application is described below. The spoken language learning device described below and the spoken language learning method described above can be referenced to each other.
[0200] See also Figure 7 , Figure 7 A schematic diagram of the structure of a spoken language learning device provided in an embodiment of the present application.
[0201] likeFigure 7 As shown, the device may include:
[0202] A dialogue environment acquisition module 701, configured to acquire a target oral learning scenario and a target language difficulty input by a language learner;
[0203] A human-machine dialogue module 702, configured to call a target large language model to conduct at least one round of oral dialogue with the language learner based on the target oral learning scenario and the target language difficulty. Among them, the training data of the target large language model includes multi-role dialogue simulation data inferred in a chain-of-thought manner, and the learner dialogue simulation data in the multi-role dialogue simulation data includes mixed-language dialogue simulation data.
[0204] In a possible implementation, when the human-machine dialogue module acquires the multi-role dialogue simulation data, it may specifically be used for:
[0205] Acquire a first prompt instruction, where the first prompt instruction is used to instruct a first large language model to perform chain-of-thought dialogue reasoning according to a preset first language difficulty and an oral learning scenario, so as to obtain teacher dialogue simulation data that meets the set dialogue goal;
[0206] Acquire a second prompt instruction, where the second prompt instruction is used to instruct a second large language model to obtain learner dialogue simulation data for the teacher dialogue simulation data by adopting a chain-of-thought dialogue reasoning strategy;
[0207] Input the first prompt instruction into the first large language model, and input the second prompt instruction into the second large language model, so that the first large language model and the second large language model conduct at least one round of oral dialogue simulation to obtain multi-role dialogue simulation data.
[0208] In a possible implementation, when the human-machine dialogue module acquires the first prompt instruction, it may specifically be used for:
[0209] Acquire the first language difficulty and the oral learning scenario;
[0210] Acquire a first prompt instruction template, where the first prompt instruction template includes a first language difficulty slot and a learning scenario slot;
[0211] Fill the first language difficulty into the first language difficulty slot, and fill the oral learning scenario into the learning scenario slot to obtain the first prompt instruction.
[0212] In a possible implementation, when the human-machine dialogue module acquires the second prompt instruction, it may specifically be used for:
[0213] Acquire preset role information and language styles;
[0214] Acquire a second prompt instruction template, where the second prompt instruction template includes a role slot and a language style slot;
[0215] Fill the role information into the role slot and fill the language style into the language style slot to obtain the second prompt instruction.
[0216] In one possible implementation, the training data of the above-mentioned target large language model may further include: learner dialogue prompt data inferred in a chain-of-thought manner.
[0217] Based on this, the above-mentioned human-computer dialogue module can also be used for: during the oral dialogue with the language learner, if a dialogue prompt request is detected, obtain the target oral learning scenario, target language difficulty, and historical dialogue data during the oral dialogue; call the target large language model to obtain target learner dialogue prompt data according to the target oral learning scenario, target language difficulty, and historical dialogue data; output and display the target learner dialogue prompt data.
[0218] In one possible implementation, when the above-mentioned human-computer dialogue module obtains learner dialogue prompt data, it can specifically be used for:
[0219] Obtain the preset second language difficulty, chat background, and conversation content under the chat background;
[0220] Obtain the third prompt instruction template. The third prompt instruction template includes a chat background slot, a dialogue history slot, and a second language difficulty slot. The third prompt instruction template is used to instruct the third large language model to perform chain-of-thought dialogue reasoning based on the chat background in the chat background slot, the conversation content in the dialogue history slot, and the second language difficulty in the second language difficulty slot to obtain learner dialogue prompt data that meets the set dialogue requirements;
[0221] Fill the chat background into the chat background slot, fill the conversation content under the chat background into the dialogue history slot, and fill the second language difficulty into the second language difficulty slot to obtain the third prompt instruction;
[0222] Input the third prompt instruction into the third large language model to obtain learner dialogue prompt data.
[0223] In one possible implementation, after the above-mentioned human-computer dialogue module obtains multi-role dialogue simulation data and learner dialogue prompt data, it can also be used for: removing the data that does not meet the data quality requirements in the multi-role dialogue simulation data and learner dialogue prompt data, and using the remaining data as training data.
[0224] In one possible implementation, the above-mentioned target large language model is obtained by performing self-distillation supervised fine-tuning on a pre-trained Decode-Only large language model using training data, and the cross-entropy loss and KL divergence regularization loss are used in the process of self-distillation supervised fine-tuning.
[0225] The oral learning device provided by this application obtains the target oral learning scenario and target language difficulty input by the language learner, and calls the target large language model to conduct at least one round of oral conversation with the language learner based on the target oral learning scenario and target language difficulty. Since the training data used by the target large language model during training is data inferred in a chain-of-thought manner, the target large language model can also perform chain-of-thought thinking when conducting an oral conversation with the language learner, improving the coherence and consistency of the conversation context.
[0226] Furthermore, when the target large language model conducts chain-of-thought thinking, in addition to considering the conversation data of the language learner, it only needs to consider the information in two dimensions: the target oral learning scenario and the target language difficulty. The amount of information is small, reducing the task difficulty of chain-of-thought thinking and improving the accuracy of the dialogue output by the target large language model. At the same time, this application supports the language learner to input different language difficulties, which can adapt to the oral learning of language learners with different language difficulties and improve the oral learning experience of language learners.
[0227] Even further, considering that language learners often output mixed-language conversation data when their language level is insufficient, in order to enable the target large language model to still conduct accurate conversations with language learners in the face of mixed-language conversation data, this application uses mixed-language conversation simulation data as the learner conversation simulation data to train the target large language model, further improving the accuracy of the dialogue output by the target large language model and providing a more realistic oral learning experience for language learners.
[0228] This application embodiment also provides an electronic device. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in this application embodiment. The electronic device in this application embodiment may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 8 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of this application embodiment.
[0229] As Figure 8As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage device 808 into the random access memory (RAM) 803. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0230] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a memory card, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0231] In an embodiment of the present application, a computer program product is further provided, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the oral learning methods provided in the embodiments of the present application.
[0232] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the oral learning methods provided in the embodiments of the present application.
[0233] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0235] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0236] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A method for oral language learning, characterized in that, including: obtaining the target oral learning scenario and target language difficulty input by the language learner; invoking the target large language model to conduct at least one round of oral dialogue with the language learner based on the target oral learning scenario and the target language difficulty; wherein, the training data of the target large language model includes multi-role dialogue simulation data inferred in a chain of thought manner, and the learner dialogue simulation data in the multi-role dialogue simulation data includes mixed-language dialogue simulation data.
2. The oral learning method according to claim 1, wherein The process of obtaining the multi-role dialogue simulation data includes: obtaining a first prompt instruction, which is used to instruct the first large language model to conduct chain-of-thought dialogue reasoning according to a preset first language difficulty and oral learning scenario, so as to obtain teacher dialogue simulation data that meets the set dialogue goal; obtaining a second prompt instruction, which is used to instruct the second large language model to adopt a chain-of-thought dialogue reasoning strategy to obtain learner dialogue simulation data for the teacher dialogue simulation data; inputting the first prompt instruction into the first large language model, and inputting the second prompt instruction into the second large language model, so that the first large language model and the second large language model conduct at least one round of oral dialogue simulation to obtain the multi-role dialogue simulation data.
3. The oral learning method according to claim 2, characterized in that The obtaining of the first prompt instruction includes: obtaining the first language difficulty and the oral learning scenario; obtaining a first prompt instruction template, which includes a first language difficulty slot and a learning scenario slot; filling the first language difficulty into the first language difficulty slot, and filling the oral learning scenario into the learning scenario slot to obtain the first prompt instruction.
4. The oral learning method according to claim 2, wherein The obtaining of the second prompt instruction includes: obtaining preset role information and language style; obtaining a second prompt instruction template, which includes a role slot and a language style slot; filling the role information into the role slot, and filling the language style into the language style slot to obtain the second prompt instruction.
5. The oral learning method according to claim 2, wherein The training data of the target large language model further includes: learner dialogue prompt data inferred in the chain-of-thought manner; The process of obtaining the learner dialogue prompt data includes: obtaining a preset second language difficulty, chat background, and dialogue content under the chat background; obtaining a third prompt instruction template, which includes a chat background slot, a dialogue history slot, and a second language difficulty slot, and the third prompt instruction template is used to instruct the third large language model to conduct chain-of-thought dialogue reasoning according to the chat background in the chat background slot, the dialogue content in the dialogue history slot, and the second language difficulty in the second language difficulty slot, so as to obtain learner dialogue prompt data that meets the set dialogue requirements; filling the chat background into the chat background slot, filling the dialogue content under the chat background into the dialogue history slot, and filling the second language difficulty into the second language difficulty slot to obtain a third prompt instruction; inputting the third prompt instruction into the third large language model to obtain the learner dialogue prompt data.
6. The oral learning method according to claim 5, wherein The oral learning method further includes: During the oral conversation with the language learner, if a conversation prompt request is detected, obtain the target oral learning scenario, the target language difficulty, and the historical conversation data during the oral conversation; Call the target large language model to obtain target learner conversation prompt data based on the target oral learning scenario, the target language difficulty, and the historical conversation data; Output and display the target learner conversation prompt data.
7. The oral learning method according to claim 5, wherein It further includes: Eliminate the data that does not meet the data quality requirements in the multi-role conversation simulation data and the learner conversation prompt data, and the remaining data is used as the training data.
8. The oral learning method according to claim 1 or 6, characterized in that, The target large language model is obtained by performing self-distillation supervised fine-tuning on a pre-trained Decode-Only large language model using the training data, and the cross-entropy loss and KL divergence regularization loss are used in the process of self-distillation supervised fine-tuning.
9. A spoken language learning device, characterized in that, It includes: A conversation environment acquisition module for obtaining the target oral learning scenario and the target language difficulty input by the language learner; A human-computer conversation module for calling a target large language model to conduct at least one round of oral conversation with the language learner based on the target oral learning scenario and the target language difficulty. Among them, the training data of the target large language model includes multi-role conversation simulation data inferred in a chain-of-thought manner, and the learner conversation simulation data in the multi-role conversation simulation data includes mixed-language conversation simulation data.
10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, where: The memory is used to store computer programs; The processor is used to execute the computer programs so that the electronic device can implement the oral learning method described in any one of claims 1 to 8.