Method for obtaining text instruction generation model and method and device for obtaining text instruction

By performing supervised optimization on lightweight large language models, a text instruction generation model was developed. This solves the problems of high resource consumption, low accuracy, and low security of large language models in industrial scenarios, and achieves efficient conversion and accurate execution of natural language into text instructions in a specific format.

CN120633828APending Publication Date: 2025-09-12CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510605945.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-12

Smart Images

  • Figure CN120633828A_ABST
    Figure CN120633828A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for obtaining a text instruction generation model and a method and device for obtaining a text instruction. By utilizing the efficient understanding ability and generation ability of a lightweight large language model for human natural languages, two-stage innovation improvement is performed on fine tuning training of the lightweight large language model, a lightweight text instruction generation model for converting natural languages into text instructions is developed, the number of parameters is small, and the generation efficiency is high. And the occupied space is small. Through two-stage supervision and optimization, intelligent tasks of a lightweight large language model are automatically arranged, so that the problems of illusion, low safety, high resource consumption and the like of an original large language model are solved; therefore, the text instruction generation model generates the text instruction of the specific format which can be executed by the system according to the voice instruction of the natural language input by the user, and the accuracy of the generated text instruction of the specific format is high, so that the electronic equipment can execute the generated text instruction of the specific format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for obtaining a text instruction generation model, and a method and device for obtaining a text instruction. Background Art

[0002] With the rapid development of technology, large language models rely on unsupervised learning of massive data and have superior language understanding and generation capabilities, becoming a hot research direction.

[0003] Gradually, large language models at home and abroad have been released like mushrooms after rain, driving productivity to leap towards machine intelligence. Large language models seem to be becoming a new mainstream production tool in the economic society.

[0004] However, the inventors found that while large language models bring about changes in productivity, they are also difficult to transform and apply, and difficult to implement in industrial scenarios. Summary of the Invention

[0005] The present application discloses a method and device for obtaining a text instruction generation model, and a method and device for obtaining text instructions.

[0006] In a first aspect, the present application provides a method for obtaining a text instruction generation model, which is applied to an electronic device. The method includes:

[0007] Acquire multiple different dialogue training data sets; each dialogue training data set includes: at least one round of dialogue text, wherein a round of dialogue text includes: historical text input by the user to the intelligent dialogue model during the historical process, and response text answered by the intelligent dialogue model in response to the historical text input by the user;

[0008] Perform supervised optimization on the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model.

[0009] Acquire multiple different instruction training data sets, the instruction training data sets including: text instructions in natural language and text instructions in a specific format; the text instructions in natural language and text instructions in a specific format are both used to implement specific types of functions; the text instructions in the specific format are used to be directly executed by the electronic device, while the text instructions in natural language cannot be directly executed by the electronic device;

[0010] Based on multiple different instruction training data sets, the intermediate large language model is optimized in a supervised manner to obtain a text instruction generation model.

[0011] In a second aspect, the present application provides a method for obtaining a text instruction, which is applied to an electronic device. The method includes:

[0012] Acquiring a natural language voice command input by a user to an electronic device, where the natural language voice command is used to control the electronic device to implement a specific type of function;

[0013] Perform semantic recognition on the natural language voice command to obtain the natural language text command corresponding to the natural language voice command;

[0014] Generate text instructions in a specific format based on the text instructions in natural language based on the text instruction generation model;

[0015] Execute text instructions in a specific format to control electronic devices to perform specific functions;

[0016] Among them, the text instruction generation model is trained based on the method of the first aspect.

[0017] In a third aspect, the present application provides a device for obtaining a text instruction generation model, which is applied to an electronic device. The device includes:

[0018] A first acquisition module is configured to acquire a plurality of different dialogue training data sets; each dialogue training data set includes: at least one round of dialogue text, wherein a round of dialogue text includes: historical text input by a user to the intelligent dialogue model during a historical process, and response text answered by the intelligent dialogue model in response to the historical text input by the user;

[0019] The first optimization module is used to perform supervised optimization on the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model;

[0020] a second acquisition module, configured to acquire a plurality of different instruction training data sets, the instruction training data sets including: text instructions in natural language and text instructions in a specific format; the text instructions in natural language and text instructions in a specific format are both used to implement specific functions; the text instructions in the specific format are used to be directly executed by the electronic device, while the text instructions in natural language cannot be directly executed by the electronic device;

[0021] The second optimization module is used to perform supervised optimization on the intermediate large language model based on multiple different instruction training data sets to obtain a text instruction generation model.

[0022] In a fourth aspect, the present application provides a device for obtaining a text instruction, which is applied to an electronic device, and includes:

[0023] a third acquisition module, configured to acquire a natural language voice instruction input by a user to the electronic device, wherein the natural language voice instruction is used to control the electronic device to implement a specific type of function;

[0024] A recognition module is used to perform semantic recognition on the natural language voice command to obtain the natural language text command corresponding to the natural language voice command;

[0025] A generation module, configured to generate text instructions in a specific format according to the text instructions in the natural language based on the text instruction generation model;

[0026] An execution module, configured to execute text instructions in a specific format to control the electronic device to perform specific functions;

[0027] The text instruction generation model is obtained by training based on the method of the first aspect, or the text instruction generation model is obtained by training based on the device of the third aspect.

[0028] In a fifth aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.

[0029] In a sixth aspect, the present application shows a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.

[0030] In a seventh aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any one of the above aspects.

[0031] The technical solution provided by this application may have the following beneficial effects:

[0032] This application utilizes the lightweight large language model's efficient understanding and generation capabilities of human natural language, and conducts a two-stage innovative improvement to the fine-tuning training of the lightweight large language model, and develops a lightweight large language model for converting natural language into text instructions, for example, a text instruction generation model. The text instruction generation model has a small number of parameters and occupies a small space.

[0033] This application uses two-stage supervised optimization to automatically orchestrate the intelligent tasks of a lightweight large language model, solving the problems of hallucinations, low security, and high resource consumption of the original large language model. It enables the text instruction generation model to generate text instructions in a specific format that can be executed by the system based on the natural language voice instructions input by the user, and the generated text instructions in a specific format have a high accuracy rate, enabling electronic devices to execute the generated text instructions in a specific format to meet the user's needs to control electronic devices based on natural language voice instructions.

[0034] This application breaks the tradition and broadens the application channels of large language models, so that the application of large language models is no longer in the form of a stereotyped intelligent question-answering assistant. It innovatively applies it to instruction arrangement tasks, allowing it to be deeply applied to daily life and work scenarios, bringing convenience to the development of science and technology to ordinary mass users. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a scenario diagram of this application.

[0036] Figure 2 This is a schematic diagram of the input data and output results of a text instruction generation model of the present application.

[0037] Figure 3 This is a flow chart of a solution of this application.

[0038] Figure 4 This is a schematic diagram of the process of producing a dialogue training dataset and an instruction training dataset in this application.

[0039] Figure 5 This is a flowchart of the steps of a method for obtaining a text instruction generation model in the present application.

[0040] Figure 6 It is a schematic diagram of a prompt word of this application.

[0041] Figure 7 This is a schematic diagram of a data format of this application.

[0042] Figure 8 This is a schematic diagram of an instruction training data set of this application.

[0043] Figure 9 This is a schematic diagram of data of a nested function of this application.

[0044] Figure 10 This is a schematic diagram of a text instruction of a nested function of this application.

[0045] Figure 11 This is a flowchart of the steps of a method for obtaining text instructions in this application.

[0046] Figure 12 This is a structural block diagram of a device for obtaining a text instruction generation model in the present application.

[0047] Figure 13 This is a structural block diagram of a device for obtaining text instructions in the present application.

[0048] Figure 14 This is a block diagram of an electronic device of the present application.

[0049] Figure 15 This is a block diagram of an electronic device of the present application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] The difficulties of large language models in implementing industrial scenarios can be manifested in the following aspects:

[0052] Insufficient computing resources. The primary challenge facing the training and application of large language models is a severe shortage of computing resources. These models require enormous computing power, massive amounts of data, and a large number of parameters, resulting in significant resource demands and high time costs for training and inference. This restricts the participation of small and medium-sized enterprises in R&D, hindering the market and the general public from fully benefiting from technological advancements. Furthermore, the long training cycles lead to a lag in knowledge updates for large language models, preventing them from gaining real-time access to the latest information.

[0053] Lack of industry expertise. Although large language models have a broad knowledge base that can cover multiple fields, they often lack in-depth expertise in specific verticals, such as agriculture, biotechnology, and other niche markets. This makes them unable to provide accurate knowledge services in these specialized fields.

[0054] Poor user experience. The Large Language Model is a probabilistic generative model. Its core algorithmic framework makes it prone to hallucinations when generating responses to questions, namely, "giving wrong answers in a serious manner." It can generate some seemingly authentic erroneous information, misleading users.

[0055] The current application model for large language models is relatively simple. It primarily involves intelligent question-and-answer assistants that interact with users through dialog boxes, replacing the traditional question-and-answer robots with low accuracy and blunt responses. However, this application approach fails to fully unleash the high-performance potential of large language models. Consequently, large language models cannot be effectively applied in diverse industrial environments, failing to maximize their advantages.

[0056] As can be seen, current challenges include high resource consumption and high barriers to entry during large language model training and inference, as well as low accuracy and security of output results. Furthermore, while large language models offer broad knowledge coverage, specialized applications require deeper and more precise knowledge.

[0057] The industry's goal is to enable large language models to meet the refined needs of specialized fields while maintaining broad knowledge and generalization capabilities. This can be achieved by reducing resource consumption, for example by developing lightweight large language models. These models require far fewer parameters during training than traditional large language models without sacrificing output accuracy, achieving true speed and efficiency.

[0058] To this end, the technical solution of this application is proposed. In order to expand the application scenarios of large language models, a lightweight large language model is developed by utilizing the large language model's profound understanding and generation capabilities of human natural language. The lightweight large language model has excellent language understanding capabilities, can understand and convert users' natural language intentions into instructions, can safely and accurately execute task requirements, output system instructions, and realize automated task orchestration, overcoming the limitations of traditional methods, breaking through the application difficulties of large language models, and improving the intelligence level of the system.

[0059] The number of parameters of the lightweight large language model does not exceed 1.1 billion, and the size of the lightweight large language model does not exceed 800MB. It can be easily deployed on end devices such as mobile phones and can be expanded to various business systems to achieve fast and accurate task automation orchestration.

[0060] This application not only makes the task orchestration process intelligent, but also can avoid data transmission over the network by being deployed on end devices such as mobile phones, ensuring data and privacy security while avoiding the problem of model response delay.

[0061] Secondly, supervised fine-tuning of the parameters of the lightweight large language model in specific application fields reduces the risk of erroneous information in the output results generated by the lightweight large language model, thereby improving the accuracy of the output results of the lightweight large language model.

[0062] Before introducing the technical solution of this application, the technologies that may be involved in the technical solution of this application are first explained.

[0063] LLM (Large Language Model): A data-driven deep learning model with billions of parameters after unsupervised learning on massive amounts of data. It can process and generate text in multiple languages ​​and is widely used in translation, writing, and dialogue systems.

[0064] Light-LLM (Lightweight Large Language Models): Through methods such as transfer learning and knowledge distillation, large language models can retain massive amounts of knowledge while having smaller parameters and being more lightweight, making them easier to deploy on the client side and implement in applications.

[0065] AIGC (Artificial Intelligence Generated Content): Generative AI technology is an emerging form of AI application that analyzes large amounts of data to learn language patterns and then generates new content. This technology can be trained on specific topics or writing styles to produce customized content.

[0066] SFT (Supervised Fine-Tune): Generally refers to a technique in machine learning and deep learning, particularly when training neural network models. This process involves taking a pre-trained model and further training it on a specific task to improve the model's performance on that task.

[0067] Retrieval-Augmented Generation (RAG): An AI technology that combines retrieval and generative models. It augments the output of the Large Language Model (LLM) by retrieving information from private or proprietary data sources to assist in text generation, improving the relevance of the search experience without retraining the model.

[0068] ASR (Automatic Speech Recognition): A technology that converts human speech into text information that can be understood by computers.

[0069] Optical Character Recognition (OCR): A technology that recognizes and extracts text from image files. OCR software processes scanned documents, photos, or digital images and converts them into editable, searchable text data.

[0070] API (Application Programming Interface): A set of predefined functions, protocols, and tools used to build software applications. An API defines how different software components communicate with each other. It allows developers to access functionality or data from an application or service without exposing the system's inner workings.

[0071] ADB (Android Debug Bridge) is a command-line tool used to communicate with connected Android devices for debugging and development work. Through adb commands, you can perform various operations, such as installing and debugging applications, accessing the device's shell, and managing the device's file system.

[0072] NLP (Natural Language Processing): It is a technology in the field of AI used to understand and generate human language, enabling machines to perform tasks such as language translation and sentiment analysis.

[0073] The Instruction-Tuned Large Language Model (LLM) is a fine-tuned version of the BaseLLM. It is trained on task-specific datasets to better understand and follow human instructions. This fine-tuning process helps the model respond more accurately to instructions, improving its performance in real-world scenarios.

[0074] In this application, a text instruction generation model is obtained by optimizing a lightweight large language model.

[0075] See also Figure 1 As shown, the function of the text instruction generation model is to convert the natural language voice instructions input by the user to the electronic device into text instructions in a specific format.

[0076] The text instruction generation model has the ability to understand and analyze the intentions in user sentences and generate text instructions in a specific format, such as generating instructions with input parameters and nested tasks to directly control end-side devices or different digital platforms.

[0077] See also Figure 2 , showing the input data (Input) and output results (Output) of the text instruction generation model.

[0078] Among them, "Ring up John at home, number is +12345678900" is the input data of the text instruction generation model, which is natural language text.

[0079] The text instruction generation model understands from the input data that the user needs to make a phone call and has directly given a phone number. Therefore, the output result of the text instruction generation model is a text instruction in a specific format:<api_0> ('+12345678900')<api_end> ”.

[0080] “<api_0> " is an identifier for an API that specifies a specific type of functionality (e.g., making a phone call).

[0081] "('+12345678900')" is an input parameter that needs to be input to the API of a specific type of function.

[0082] “<api_end> " is the command terminator for text commands in a specific format.

[0083] In addition, the flowchart of this application can be found in Figure 3 shown.

[0084] The process of this application involves data level, model level, and application level.

[0085] At the data level, it involves sorting out the functional APIs and fine-tuning data production.

[0086] Functional analysis. An application can implement a single function on a mobile phone. This application analyzes six single functions, one nested function, and one unrelated function. Single functions include: calling, contacts, browser, photo taking, weather, and email. When user input does not involve any of the above six functions, the unrelated function is returned.

[0087] The nested function includes: after searching for relevant content using a browser, the searched content can be sent via email.

[0088] Fine-tuning data production. This involves collecting and organizing high-quality conversational data and creating natural language-text instruction data pairs. The former relies on open-source datasets, while the latter involves designing API function descriptions based on eight functions, input parameters, and output parameters, using prompt words to enable data synthesis for a general model.

[0089] At the model level, it involves fine-tuning of dialogue capabilities and target tasks.

[0090] Fine-tune conversational capabilities. Using a lightweight large language model as the base model, perform supervised fine-tuning of all parameters to enhance the model's contextual understanding and conversational capabilities, resulting in an intermediate large language model.

[0091] Fine-tuning the target task. Align the intermediate large language model with task data of a specific type of function, learn the fixed format and output mode of the fine-tuning data, and obtain a text instruction generation model.

[0092] At the application level, it involves Android engineering simulation.

[0093] Android engineering simulation. This project verifies the superior ability of the trained text command generation model to convert natural language text into multi-task orchestration commands. This capability is coordinated and invoked within a computer development environment, allowing mobile phones to perform tasks through commands. Furthermore, the application and deployment of this text command generation model consumes significantly fewer resources than traditional natural language processing algorithms and general-purpose large language models, enabling widespread application. This project is applicable not only to mobile phones but also to other digital platforms, simplifying the operation of complex systems.

[0094] In this application, fine-tuning the lightweight large language model is mainly divided into two steps:

[0095] 1. Perform supervised optimization on a lightweight large language model (for example, HARE-1.1B) based on multiple different conversation training datasets to obtain an intermediate large language model (for example, HARE-chat). This model is then given improved conversational and natural language understanding capabilities.

[0096] 2. Use multiple different high-quality instruction training datasets (such as natural language-text instruction data pairs) to perform supervised fine-tuning of all parameters of the intermediate large language model (such as HARE-chat) to ensure that it has the input and output capabilities required by the project tasks and obtain a text instruction generation model.

[0097] The process of making dialogue training dataset and instruction training dataset can be found in Figure 4 shown.

[0098] Among them, reference Figure 5 , shows a flowchart of the steps of a method for obtaining a text instruction generation model of the present application, the method is applied to an electronic device, and the method includes:

[0099] In step S101, multiple different dialogue training datasets are obtained. Each dialogue training dataset includes at least one round of dialogue text. Each round of dialogue text includes the historical text input by the user to the intelligent dialogue model during the historical process, and the response text of the intelligent dialogue model in response to the historical text input by the user.

[0100] Select multiple dialogue training datasets from open-source datasets, such as the Cornell Movie Dialogs Corpus, Persona-Chat, and UltraChat. Intelligent dialogue models include commercially available AI chat models, such as Kimi or Wenxin Yiyan.

[0101] Dialogue structuring: Convert the dialogues of different dialogue data into a standard input-output pair format, that is, each dialogue turn consists of a user input and the corresponding model response.

[0102] In step S102 , supervised optimization is performed on the lightweight large language model based on multiple different dialogue training data sets to obtain an intermediate large language model.

[0103] In one embodiment of the present application, historical texts and incorrect punctuation marks, HTML (HyperText Markup Language) tags, and / or meaningless characters (such as filler words like "um", "ah", and "ya") in response texts in multiple different dialogue training datasets can be removed to obtain multiple different first dialogue training datasets. Duplicate removal processing for the dialogue texts in the multiple different first dialogue training datasets is performed on the historical texts or response texts to obtain multiple different second dialogue training datasets (it can be determined whether texts are the same based on hash encoding of the texts. If the historical texts or response texts in two dialogue texts are the same, any one of the two dialogue texts can be removed). The letters in the historical texts and response texts in the multiple different second dialogue training datasets are uniformly converted to lowercase letters to obtain multiple different third dialogue training datasets. The full English names in the historical texts and response texts in the multiple different third dialogue training datasets are converted to English abbreviations to obtain multiple different fourth dialogue training datasets. Through the foregoing processing, it is ensured that the texts in the fourth dialogue training datasets are clean texts. Subsequently, the lightweight large language model can be optimized in a supervised manner based on the multiple different fourth dialogue training datasets to obtain an intermediate large language model, avoiding affecting the semantic understanding ability of the intermediate large language model.

[0104] In another embodiment of the present application, at least one character in the real names in the historical texts and response texts in multiple different dialogue training datasets can be replaced with a character different from the at least one character in the real name to obtain multiple different fifth dialogue training datasets. At least one character in the real addresses in the historical texts and response texts in the multiple different fifth dialogue training datasets is replaced with a character different from the at least one character in the real address to obtain multiple different sixth dialogue training datasets. At least one character in the real phone numbers in the historical texts and response texts in the multiple different sixth dialogue training datasets is replaced with a character different from the at least one character in the real phone number to obtain multiple different seventh dialogue training datasets. The lightweight large language model is optimized in a supervised manner based on the multiple different seventh dialogue training datasets to obtain an intermediate large language model. In this way, privacy information can be protected and the complete exposure of privacy information can be avoided.

[0105] The character different from the at least one character in the real name can be a character preset by technicians, etc. What specific characters are preset can be determined according to the actual situation, and the present application does not limit this.

[0106] The characters different from the at least one character in the real address may be characters pre-set by a technician, etc. The specific characters of the pre-set characters may be determined according to actual conditions, and this application does not impose any limitation on this.

[0107] The characters different from the at least one character in the real telephone number may be characters pre-set by a technician, etc. The specific characters of the pre-set characters may be determined according to actual conditions, and this application does not impose any limitation on this.

[0108] In another embodiment of the present application, the conversation text can be diversified and expanded to obtain an extended conversation text, wherein the historical text in the extended conversation text is different from the historical text in the conversation text, and / or the response text in the extended conversation text is different from the response text in the conversation text. Based on the conversation texts in multiple different conversation training datasets and the extended conversation text obtained by the diversified expansion, a lightweight large language model is subjected to supervised optimization to obtain an intermediate large language model.

[0109] When performing diversity expansion on the dialogue text, characters may be randomly inserted into the historical text and / or the response text in the dialogue text, and / or characters may be randomly deleted from the historical text and / or the response text in the dialogue text, and / or characters may be randomly replaced from the historical text and / or the response text in the dialogue text, and / or the historical text in the dialogue text may be replaced with a first replacement text, the semantics of the first replacement text being the same as that of the historical text replaced by the first replacement text, and / or the response text in the dialogue text may be replaced with a second replacement text, the semantics of the second replacement text being the same as that of the response text replaced by the second replacement text, thereby obtaining the dialogue expanded text.

[0110] In this way, we can continue to add diverse training texts based on the dialogue texts in the dialogue training dataset to increase the diversity of the dataset used for supervised optimization of the lightweight large language model, thereby improving the robustness of the optimized intermediate large language model.

[0111] It should be noted that any two of the aforementioned three embodiments can be comprehensively and organically combined for use, or the aforementioned three embodiments can be comprehensively and organically combined for use.

[0112] In addition, manual review can be performed on a sample basis or in its entirety to ensure that the text is semantically natural and fluent, and contains no obvious errors or inappropriate content.

[0113] In step S103, multiple different instruction training data sets are obtained. The instruction training data sets include: text instructions in natural language and text instructions in a specific format. Both the text instructions in natural language and the text instructions in a specific format are used to implement specific functions. The text instructions in the specific format are used to be directly executed by the electronic device.

[0114] Text instructions in natural language cannot be directly executed by electronic devices.

[0115] In one embodiment of the present application, this step can be implemented through the following process, including:

[0116] 1031. Obtain a prompt word. The prompt word is used to prompt the task details of the text instruction generation task of the large language model for instruction tuning. The task details at least include the format of the text instruction generated by the large language model for instruction tuning.

[0117] The instruction-tuned large language model has the ability to generate text, for example, the ability to generate text instructions, including a convolutional neural network or a recurrent neural network, or a model based on a convolutional neural network or a model based on a recurrent neural network.

[0118] The prompt words are clear and specific, and require structured output. For example, the task details must at least include the format of the text instructions generated by the instruction-tuned large language model, and the instruction-tuned large language model must be required to check whether the conditions are met, and provide few-shot prompts.

[0119] Secondly, the prompt word needs to give the instruction-tuned large language model time to think, such as specifying the steps required to complete the task. Since the instruction-tuned large language model sometimes falsifies parameters, the prompt word should remind the instruction-tuned large language model not to falsify any parameters.

[0120] See also Figure 6 , showing the prompt words designed for phone call instructions, the task details of the text instruction generation task of the large language model with prompt instructions tuned in the prompt words, etc.

[0121] 1032. Obtain at least one sample example data set, where the sample example data set includes: sample example text instructions in a natural language and sample example text instructions in a specific format. The sample example text instructions in the natural language and the sample example text instructions in the specific format are both used to implement a specific type of function. The sample example text instructions in the specific format are directly executable by the electronic device.

[0122] The sample example text instructions of natural language cannot be directly executed by the electronic device.

[0123] The purpose of at least one sample example data set is to allow the instruction-tuned large language model to imitate, so that the instruction-tuned large language model can output text instructions in a specific format based on natural language text instructions.

[0124] For example, Figure 6 Two sample example data sets are given in the paper for the instruction-tuned large language model to imitate, so that the instruction-tuned large language model can output text instructions in a specific format based on natural language text instructions.

[0125] The format of the text instruction generated by the large language model for instruction tuning includes: an identifier of an API of a specific type of function, input parameters that need to be input to the API of the specific type of function, and an instruction terminator of the text instruction.

[0126] Specific categories of functions include: call function, contact function, browser function, camera function, weather function, email function, and nesting function.

[0127] Nested functions are: two or more functions are used serially in the call function, contact function, browser function, camera function, weather function and email function, and the output result of the adjacent previous function is the input data of the adjacent next function.

[0128] Specific types of functions also include: irrelevant functions, which are functions other than the calling function, the contact function, the browser function, the camera function, the weather function, the email function, and the nested function.

[0129] Based on the principle of functional universality, we collected data on the functions that most terminals on the market (such as mobile phones) have when they leave the factory. Each function corresponds to an API. We then built a functional description for each function and defined the input and output parameters of the API.

[0130] This application has sorted out a total of 6 functions, 1 unrelated function and 1 nested function, the details of which are as follows:

[0131]

[0132] 1033. Obtain function description information of a specific type of function.

[0133] The functional description information of a specific type of function may enable a large language model for instruction tuning to fully understand the specific type of function.

[0134] Among them, see Figure 7 , the prompt word, at least one sample example data set, and description information of a specific type of function can be reproduced into a required data format.

[0135] 1034. With the aid of the instruction-tuned large language model, multiple different instruction training data sets are generated based on prompt words, at least one sample example data set, and description information of a specific type of function.

[0136] The instruction training dataset includes both natural language text instructions and text instructions in a specific format. Both natural language text instructions and text instructions in a specific format are used to implement specific functions. Text instructions in a specific format are intended to be directly executed by electronic devices, while natural language text instructions cannot be directly executed by electronic devices.

[0137] The large language model for instruction tuning includes: Instruction Tuned LLM, etc.

[0138] Through this embodiment, the diversity of the instruction training data set can be improved.

[0139] Through the design of diverse prompts, data synthesis constraints are provided to the large language model for instruction tuning in the form of prompts, guiding the large language model for instruction tuning to generate data that meets task requirements.

[0140] In another embodiment of the present application, for any type of function, multiple instruction training data sets corresponding to that type of function can be obtained, and the same is true for each other type of function, thereby obtaining multiple instruction training data sets corresponding to each type of function respectively.

[0141] See also Figure 8 ,Each instruction training dataset is generated in json format and contains two fields.

[0142] “human” refers to text instructions in natural language.

[0143] "Assistant" is a text instruction in a specific format, which includes a token representing a special function of the API, input parameters required by the API, and a fixed end token.

[0144] In step S104, supervised optimization is performed on the intermediate large language model based on a plurality of different instruction training data sets to obtain a text instruction generation model.

[0145] Among them, the natural language text of browser functions and the natural language text of unrelated functions can easily confuse the large language model. The intermediate large language model's judgment on whether the user wants to obtain a certain type of information and the user's requirements are unknown or irrelevant is relatively vague. Therefore, when synthesizing data, we focus on enhancing and distinguishing the natural language data synthesized by these two APIs.

[0146] Also, see Figure 9 , you can intuitively see the key data content and structure of nested functions.

[0147] The first line is Figure 7 The "system" field in , which is mainly used to tell Light-LLM the task and require it to follow the task context during operation.

[0148] The second line "Query" is the natural language input by the user (corresponding to Figure 7 The "human" field content in .

[0149] The third line "Response" is the text content output by the text instruction generation model.<api_5> ('stu@yahoo.com','Interest Rate in the US',<api_3> ('Interest rate in the US'))<api_end> " is a text instruction.

[0150] <api_3> The search results for the input data 'Interest rate in the US' are<api_5> The third input parameter (the text instructions of the nested function can be seen in Figure 10 ).

[0151] The data construction here enables the large language model for instruction tuning to learn to parse the order of multiple tasks from the user's natural language and automatically nest them, thereby realizing the automated orchestration of multiple tasks.

[0152] "Function description" is the<api_3> 、<api_5> The corresponding functional description information is used to help the large language model understand the instruction tuning during training. However, during actual application reasoning, the model does not output this part of the content.

[0153] Furthermore, through manual review, the data set is sampled and checked to ensure that the content of the generated text instructions correctly corresponds to the requirements expressed in natural language, and that the input parameters and output parameters are extracted correctly.

[0154] This application uses a data-driven approach to achieve language understanding and generation through pre-training and fine-tuning, and can achieve excellent performance for different tasks through only fine-tuning.

[0155] The base model of the large language model is one-way, which can continue to write text based on given words or sentences, but lacks context understanding and dialogue capabilities.

[0156] Therefore, in this application, a two-stage fine-tuning is used, and a high-quality dialogue training dataset is used to enhance the lightweight large language model's perception of the context before and after the dialogue, so that the lightweight large language model can better handle the conversation dynamics, coherence and user engagement. The lightweight large language model is further supervised and optimized to obtain an intermediate large language model; then, the intermediate large language model is supervised and optimized using an instruction training dataset (natural language-text instruction data pairs) to obtain a text instruction generation model, so that the text instruction generation model can accurately complete tasks of specific types of functions.

[0157] Full parameter supervised fine-tuning is based on the intermediate large language model obtained after pre-training, and is obtained by fine-tuning it with labeled data to adapt it to downstream tasks of specific types of functions. The goal is to achieve better performance on specific tasks by fine-tuning the parameters of the intermediate large language model obtained after pre-training. This application uses two-stage fine-tuning to guide model learning by using different data, allowing the base model to gradually achieve the target performance capability. The specific training settings are the same.

[0158] The data contains the "human" and "assistant" fields. "Assistant" is the text instruction that the text instruction generation model is expected to output. Therefore, during training, the text instruction generation model will calculate the loss of the "assistant" part to participate in the weight update of the text instruction generation model.

[0159] The loss function used in the supervised optimization process may include a cross entropy loss function, for example, nn.CrossEntropyLoss.

[0160] The cross entropy loss function combines the operations of nn.LogSoftmax and nn.NLLLoss (negative log-likelihood loss) without explicitly applying a softmax function to the input.

[0161] Assume that the training sample set is x i is the input, y i is the corresponding label, then the cross entropy loss function is: N is the total number of samples; C is the total number of categories; y ij is an indicator variable. If sample i enters category j, then y ij =1, otherwise y ij =0; P(y j |x i ; θ) is the probability that the model predicts that sample i belongs to category j. This probability is given by the softmax layer of the model; θ represents the parameters of the model.

[0162] nn.CrossEntropyLoss has a slightly different data representation for multi-classification problems because it operates directly on the raw scores (logits) instead of probability distributions.

[0163] For example, for each sample i and category C, the loss function can be expressed as zj is the raw score (logits) of sample i for category j; yi is the true category label of sample i; C is the total number of categories, and for the entire dataset, the loss is the average of all samples

[0164] If you set the labels in non-assistant positions to ignore_index, the loss function will ignore all labels specified by ignore_index, and the corresponding terms in the mathematical expression will be omitted. In this way, for these ignored labels, their loss contribution is 0 and will not affect gradient calculation and model training.

[0165] For the settings of the optimization process:

[0166] The learning rate scheduler uses cosine annealing. The initial value of the learning rate is 1 times (10 to the power of 06), which changes according to the cosine function and gradually decreases during training.

[0167] Gradient accumulation uses a batch of data for forward and backward propagation each time, saves the gradient but does not clear it, and does not update the model parameters. After the model accumulates a certain number of times (for example, 8 or 9 times), the parameters are updated according to the accumulated gradient, and then the gradient is cleared to proceed to the next cycle.

[0168] Gradient checkpoints selectively save activation values. During backpropagation, gradients are calculated by recalculating unsaved activation values, trading time for space and optimizing memory utilization. Adamw is selected as the optimizer for gradient descent, with weight decay.

[0169] For the second-stage training, after data preparation, fine-tuning is performed using conversational data to obtain an intermediate large language model. Before fine-tuning for the target task, the tokenizer of the intermediate large language model is expanded to add identifiers for specific types of functions. The model's embedding and head layers are resized, and supervised fine-tuning of all parameters is performed using the Firefly framework. Ultimately, the final text command generation model with automatic task orchestration capabilities is obtained. The size of the text command generation model is approximately 2GB, but after int4 quantization, it is reduced to only 800MB.

[0170] This application utilizes the lightweight large language model's efficient understanding and generation capabilities of human natural language, and conducts a two-stage innovative improvement to the fine-tuning training of the lightweight large language model, and develops a lightweight large language model for converting natural language into text instructions, for example, a text instruction generation model. The text instruction generation model has a small number of parameters and occupies a small space.

[0171] This application uses two-stage supervised optimization to automatically orchestrate the intelligent tasks of a lightweight large language model, solving the problems of hallucinations, low security, and high resource consumption of the original large language model. It enables the text instruction generation model to generate text instructions in a specific format that can be executed by the system based on the natural language voice instructions input by the user, and the generated text instructions in a specific format have a high accuracy rate, enabling electronic devices to execute the generated text instructions in a specific format to meet the user's needs to control electronic devices based on natural language voice instructions.

[0172] It greatly simplifies the manual operation process of electronic devices, fully demonstrates the bridging ability of the text command generation model between users and electronic devices, can also be expanded to more other electronic devices and is compatible with multimodal input, and is a stable and reliable application method.

[0173] This application breaks the tradition and broadens the application channels of large language models, so that the application of large language models is no longer in the form of a stereotyped intelligent question-answering assistant. It innovatively applies it to instruction arrangement tasks, allowing it to be deeply applied to daily life and work scenarios, bringing convenience to the development of science and technology to ordinary mass users.

[0174] After obtaining the text instruction generation model, the text instruction generation model can be put into use. Figure 11 , shows a flowchart of the steps of a method for obtaining text instructions of the present application, which is applied to an electronic device and includes:

[0175] In step S201 , a natural language voice instruction input by a user to an electronic device is obtained, where the natural language voice instruction is used to control the electronic device to implement a specific type of function.

[0176] The specific category of functions includes one of the eight functions mentioned in the preceding table.

[0177] In step S202, semantic recognition is performed on the natural language voice instruction to obtain a natural language text instruction corresponding to the natural language voice instruction.

[0178] Any existing semantic recognition algorithm can be used to perform semantic recognition on natural language voice commands.

[0179] In step S203 , a text instruction in a specific format is generated according to the text instruction in the natural language based on the text instruction generation model.

[0180] The text instruction generation model is trained based on the training method shown in the above embodiment.

[0181] In step S204, a text instruction in a specific format is executed to control the electronic device to implement a specific type of function.

[0182] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions involved are not necessarily required by this application.

[0183] Reference Figure 12 , shows a device for obtaining a text instruction generation model of the present application, which is applied to an electronic device, and the device includes:

[0184] A first acquisition module 11 is configured to acquire a plurality of different dialogue training data sets; each dialogue training data set includes at least one round of dialogue text, wherein a round of dialogue text includes historical text input by a user to the intelligent dialogue model during a historical process, and response text provided by the intelligent dialogue model in response to the historical text input by the user;

[0185] A first optimization module 12 is configured to perform supervised optimization on the lightweight large language model based on multiple different dialogue training data sets to obtain an intermediate large language model;

[0186] A second acquisition module 13 is configured to acquire a plurality of different instruction training data sets, wherein the instruction training data sets include: text instructions in natural language and text instructions in a specific format; both the text instructions in natural language and the text instructions in a specific format are used to implement specific functions; the text instructions in the specific format are used to be directly executed by the electronic device, while the text instructions in natural language cannot be directly executed by the electronic device;

[0187] The second optimization module 14 is used to perform supervised optimization on the intermediate large language model based on multiple different instruction training data sets to obtain a text instruction generation model.

[0188] In an optional implementation, the first optimization module includes:

[0189] a removal unit, configured to remove incorrect punctuation marks, HTML tags, and / or meaningless characters from historical texts and response texts in the plurality of different dialogue training data sets, to obtain a plurality of different first dialogue training data sets;

[0190] a deduplication unit, configured to perform deduplication processing on the conversation texts in the plurality of different first conversation training data sets with respect to the historical texts or the response texts, to obtain a plurality of different second conversation training data sets;

[0191] a first conversion unit, configured to uniformly convert letters in historical texts and response texts in the plurality of different second dialogue training data sets into lowercase letters to obtain a plurality of different third dialogue training data sets;

[0192] a second conversion unit, configured to convert the full English names in the historical texts and the response texts in the plurality of different third dialogue training data sets into English abbreviations to obtain a plurality of different fourth dialogue training data sets;

[0193] The first optimization unit is configured to perform supervised optimization on the lightweight large language model based on multiple different fourth dialogue training data sets to obtain an intermediate large language model.

[0194] In an optional implementation, the first optimization module includes:

[0195] a first replacing unit, configured to replace at least one character of a real name in a history text and a response text in a plurality of different dialogue training data sets with a character different from the at least one character in the real name, to obtain a plurality of different fifth dialogue training data sets;

[0196] a second replacing unit, configured to replace at least one character in the real address in the historical text and the response text in the plurality of different fifth dialogue training data sets with a character different from the at least one character in the real address, to obtain a plurality of different sixth dialogue training data sets;

[0197] a third replacing unit, configured to replace at least one character of a real phone number in the historical text and the response text in the plurality of different sixth dialogue training data sets with a character different from the at least one character in the real phone number, to obtain a plurality of different seventh dialogue training data sets;

[0198] The second optimization unit is used to perform supervised optimization on the lightweight large language model based on multiple different seventh dialogue training data sets to obtain an intermediate large language model.

[0199] In an optional implementation, the first optimization module includes:

[0200] an expansion unit configured to perform diversity expansion on the dialogue text to obtain a dialogue expansion text, wherein the history text in the dialogue expansion text is different from the history text in the dialogue text, and / or the response text in the dialogue expansion text is different from the response text in the dialogue text;

[0201] The third optimization unit is used to perform supervised optimization on the lightweight large language model based on the dialogue texts in multiple different dialogue training data sets and the dialogue extension texts obtained by diversity expansion to obtain an intermediate large language model.

[0202] In an optional implementation, the extension unit includes:

[0203] An insert subunit is used to randomly insert characters into the historical text and / or response text in the dialogue text, and / or a delete subunit is used to randomly delete some characters in the historical text and / or response text in the dialogue text, and / or a first replacement subunit is used to randomly replace some characters in the historical text and / or response text in the dialogue text, and / or a second replacement subunit is used to replace the historical text in the dialogue text with a first replacement text, the semantics of the first replacement text being the same as those of the historical text replaced by the first replacement text; and / or a third replacement subunit is used to replace the response text in the dialogue text with a second replacement text, the semantics of the second replacement text being the same as those of the response text replaced by the second replacement text, thereby obtaining a dialogue extended text.

[0204] In an optional implementation, the second acquisition module includes:

[0205] The first acquisition unit is configured to acquire a prompt word; the prompt word is used to prompt the task details of the text instruction generation task of the large language model for instruction tuning, and the task details at least include the format of the text instruction generated by the large language model for instruction tuning;

[0206] a second acquiring unit, configured to acquire at least one sample example data set, the sample example data set including: sample example text instructions in a natural language and sample example text instructions in a specific format; the sample example text instructions in the natural language and the sample example text instructions in the specific format are both used to implement specific types of functions; the sample example text instructions in the specific format are used to be directly executed by the electronic device, while the sample example text instructions in the natural language cannot be directly executed by the electronic device;

[0207] a third acquiring unit, configured to acquire function description information of a specific type of function;

[0208] The generation unit is configured to generate a plurality of different instruction training data sets based on the prompt words, at least one sample example data set, and description information of a specific type of function by using the instruction-tuned large language model.

[0209] In an optional implementation, the format of the text instruction generated by the large language model for instruction tuning includes: an identifier of the application programming interface API of a specific type of function, input parameters that need to be input to the API of the specific type of function, and an instruction terminator of the text instruction.

[0210] In an optional implementation, the specific categories of functions include: a call function, a contact function, a browser function, a camera function, a weather function, an email function, and a nested function;

[0211] Nested functions are defined as the use of two or more functions serially in the call function, contacts function, browser function, camera function, weather function, and email function, where the output of the previous function is used as the input data for the next function.

[0212] Specific types of functions also include: irrelevant functions, which are functions other than the calling function, the contact function, the browser function, the camera function, the weather function, the email function, and the nested function.

[0213] This application utilizes the lightweight large language model's efficient understanding and generation capabilities of human natural language, and conducts a two-stage innovative improvement to the fine-tuning training of the lightweight large language model, and develops a lightweight large language model for converting natural language into text instructions, for example, a text instruction generation model. The text instruction generation model has a small number of parameters and occupies a small space.

[0214] This application uses two-stage supervised optimization to automatically orchestrate the intelligent tasks of a lightweight large language model, solving the problems of hallucinations, low security, and high resource consumption of the original large language model. It enables the text instruction generation model to generate text instructions in a specific format that can be executed by the system based on the natural language voice instructions input by the user, and the generated text instructions in a specific format have a high accuracy rate, enabling electronic devices to execute the generated text instructions in a specific format to meet the user's needs to control electronic devices based on natural language voice instructions.

[0215] It greatly simplifies the manual operation process of electronic devices, fully demonstrates the bridging ability of the text command generation model between users and electronic devices, can also be expanded to more other electronic devices and is compatible with multimodal input, and is a stable and reliable application method.

[0216] This application breaks the tradition and broadens the application channels of large language models, so that the application of large language models is no longer in the form of a stereotyped intelligent question-answering assistant. It innovatively applies it to instruction arrangement tasks, allowing it to be deeply applied to daily life and work scenarios, bringing convenience to the development of science and technology to ordinary mass users.

[0217] Reference Figure 13 , shows a device for obtaining text instructions of the present application, which is applied to an electronic device, and the device includes:

[0218] a third acquisition module 21 for acquiring a natural language voice command inputted by a user into the electronic device, wherein the natural language voice command is used to control the electronic device to implement a specific type of function;

[0219] The recognition module 22 is used to perform semantic recognition on the natural language voice instruction to obtain the natural language text instruction corresponding to the natural language voice instruction;

[0220] A generating module 23, configured to generate a text instruction in a specific format according to the text instruction in the natural language based on the text instruction generation model;

[0221] An execution module 24 is used to execute text instructions in a specific format to control the electronic device to perform specific functions;

[0222] The text instruction generation model is trained based on any one of claims 1-8.

[0223] Optionally, an embodiment of the present application also provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0224] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the various processes of the above-described method embodiments are implemented and the same technical effects are achieved. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0225] Figure 148 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0226] Reference Figure 14 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0227] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0228] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0229] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0230] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also monitor the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0231] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0232] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0233] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can monitor the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also monitor the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to monitor the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0234] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0235] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0236] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0237] Figure 15 1 is a block diagram of an electronic device 1900 shown in the present application. For example, the electronic device 1900 can be provided as a server.

[0238] Reference Figure 15 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0239] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0240] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0241] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0242] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0243] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0244] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0245] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0246] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0247] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0248] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0249] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for obtaining a text instruction generation model, characterized in that: Applied to electronic equipment, the method includes: Acquire multiple different dialogue training data sets; each dialogue training data set includes: at least one round of dialogue text, wherein a round of dialogue text includes: historical text input by the user to the intelligent dialogue model during the historical process, and response text answered by the intelligent dialogue model in response to the historical text input by the user; Perform supervised optimization on the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model. Acquire multiple different instruction training data sets, the instruction training data sets including: text instructions in natural language and text instructions in a specific format; the text instructions in natural language and text instructions in the specific format are both used to implement specific types of functions; the text instructions in the specific format are used to be directly executed by the electronic device; Based on multiple different instruction training data sets, the intermediate large language model is optimized in a supervised manner to obtain a text instruction generation model.

2. The method according to claim 1, characterized in that The supervised optimization of the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model includes: removing incorrect punctuation marks, Hypertext Markup Language (HTML) tags, and / or meaningless characters from historical texts and response texts in the multiple different dialogue training data sets to obtain multiple different first dialogue training data sets; performing deduplication processing on the conversation texts in the plurality of different first conversation training data sets with respect to the historical texts or the response texts to obtain a plurality of different second conversation training data sets; Converting letters in history texts and response texts in multiple different second dialogue training data sets to lowercase letters to obtain multiple different third dialogue training data sets; Converting the full English names in the historical texts and response texts in the multiple different third dialogue training data sets into English abbreviations to obtain multiple different fourth dialogue training data sets; Based on multiple different fourth dialogue training datasets, the lightweight large language model is optimized in a supervised manner to obtain an intermediate large language model.

3. The method according to claim 1, characterized in that The supervised optimization of the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model includes: replacing at least one character of a real name in history texts and response texts in a plurality of different dialogue training data sets with a character different from the at least one character in the real name, to obtain a plurality of different fifth dialogue training data sets; replacing at least one character in the real address in the historical text and the response text in the plurality of different fifth dialogue training data sets with a character different from the at least one character in the real address, to obtain a plurality of different sixth dialogue training data sets; replacing at least one character of a real phone number in the historical text and the response text in the plurality of different sixth conversation training data sets with a character different from the at least one character in the real phone number, to obtain a plurality of different seventh conversation training data sets; Based on multiple different seventh dialogue training datasets, the lightweight large language model is optimized in a supervised manner to obtain an intermediate large language model.

4. The method according to claim 1, wherein The supervised optimization of the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model includes: Performing diversity expansion on the dialogue text to obtain a dialogue extension text, wherein the history text in the dialogue extension text is different from the history text in the dialogue text, and / or the response text in the dialogue extension text is different from the response text in the dialogue text; Based on the conversation texts from multiple different conversation training datasets and the conversation extension texts obtained through diversity expansion, the lightweight large language model is supervisedly optimized to obtain an intermediate large language model.

5. The method according to claim 4, characterized in that The step of performing diversity expansion on the dialogue text to obtain the dialogue extension text includes: Randomly insert characters into the historical text and / or response text in the dialogue text, and / or randomly delete some characters in the historical text and / or response text in the dialogue text, and / or randomly replace some characters in the historical text and / or response text in the dialogue text, and / or replace the historical text in the dialogue text with a first replacement text, the semantics of the first replacement text being the same as those of the historical text replaced by the first replacement text; and / or replace the response text in the dialogue text with a second replacement text, the semantics of the second replacement text being the same as those of the response text replaced by the second replacement text, thereby obtaining a dialogue extended text.

6. The method according to claim 1, characterized in that The obtaining of a plurality of different instruction training data sets, wherein the instruction training data sets include text instructions in natural language and text instructions in a specific format, includes: Obtaining a prompt word; the prompt word is used to prompt the task details of the text instruction generation task of the large language model for instruction tuning, and the task details at least include the format of the text instruction generated by the large language model for instruction tuning; Obtaining at least one sample example data set, the sample example data set including: sample example text instructions in a natural language and sample example text instructions in a specific format; the sample example text instructions in the natural language and the sample example text instructions in the specific format are both used to implement a specific type of function; the sample example text instructions in the specific format are used to be directly executed by the electronic device; Get the function description information of a specific type of function; With the help of a large language model tuned for instructions, multiple different instruction training data sets are generated based on prompt words, at least one sample example data set, and description information of specific types of functions.

7. The method according to claim 1 or 6, characterized in that The format of the text instruction generated by the large language model for instruction tuning includes: an identifier of an application programming interface (API) of a specific type of function, input parameters that need to be input to the API of the specific type of function, and an instruction terminator of the text instruction.

8. The method according to claim 1, characterized in that Specific categories of functionality include: call functionality, contact functionality, browser functionality, camera functionality, weather functionality, email functionality, and nesting functionality; Nested functions are defined as the use of two or more functions serially in the call function, contacts function, browser function, camera function, weather function, and email function, where the output of the previous function is used as the input data for the next function. Specific types of functions also include: irrelevant functions, which are functions other than the calling function, the contact function, the browser function, the camera function, the weather function, the email function, and the nested function.

9. A method for obtaining text instructions, characterized in that: Applied to electronic equipment, the method includes: Acquiring a natural language voice command input by a user to an electronic device, where the natural language voice command is used to control the electronic device to implement a specific type of function; Perform semantic recognition on the natural language voice command to obtain the natural language text command corresponding to the natural language voice command; Generate text instructions in a specific format based on the text instructions in natural language based on the text instruction generation model; Execute text instructions in a specific format to control electronic devices to perform specific functions; The text instruction generation model is trained based on any one of claims 1-8.

10. A device for obtaining a text instruction generation model, characterized in that: Applied to electronic equipment, the device comprises: A first acquisition module is configured to acquire a plurality of different dialogue training data sets; each dialogue training data set includes: at least one round of dialogue text, wherein a round of dialogue text includes: historical text input by a user to the intelligent dialogue model during a historical process, and response text answered by the intelligent dialogue model in response to the historical text input by the user; The first optimization module is used to perform supervised optimization on the lightweight large language model based on multiple different dialogue training datasets to obtain an intermediate large language model; a second acquisition module, configured to acquire a plurality of different instruction training data sets, the instruction training data sets including: text instructions in natural language and text instructions in a specific format; the text instructions in natural language and text instructions in a specific format are both used to implement specific functions; the text instructions in the specific format are used to be directly executed by the electronic device, while the text instructions in natural language cannot be directly executed by the electronic device; The second optimization module is used to perform supervised optimization on the intermediate large language model based on multiple different instruction training data sets to obtain a text instruction generation model.

11. A device for obtaining text instructions, characterized in that: Applied to electronic equipment, the device comprises: a third acquisition module, configured to acquire a natural language voice instruction input by a user to the electronic device, wherein the natural language voice instruction is used to control the electronic device to implement a specific type of function; A recognition module is used to perform semantic recognition on the natural language voice command to obtain the natural language text command corresponding to the natural language voice command; A generation module, configured to generate text instructions in a specific format according to the text instructions in the natural language based on the text instruction generation model; An execution module, configured to execute text instructions in a specific format to control the electronic device to perform specific functions; The text instruction generation model is trained based on any one of claims 1-8.

12. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by the processor.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the method according to any one of claims 1 to 9 when executed by a processor.