Instruction generation method and device, computing system and computing equipment
By using a feedback mechanism to generate and evaluate instruction sets in a large language model, the problem of low efficiency in constructing instruction sets in existing technologies is solved, and high-quality instruction sets are generated efficiently.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for applying Large Language Models (LLMs) to specialized fields rely on a large number of labeled instructions to construct instruction sets, resulting in complex implementation and low efficiency.
By inputting prompt words into a large language model to generate an instruction set, and adjusting the prompt words based on the instruction evaluation results, a feedback relationship between instruction generation and evaluation is established to improve the quality of the instruction set.
It achieves efficient generation of instruction sets while improving the quality of the instruction sets, without relying on labeled instructions to train a dedicated instruction generation model, thus improving efficiency and effectiveness.
Smart Images

Figure CN121742902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence (AI), and in particular relates to an instruction generation method and device, a computing system and a computing device. BACKGROUND
[0002] With the rapid development of large language model (LLM) technology, there are more and more demands for applying LLM to professional fields. At present, when applying LLM to professional fields, it is usually necessary to fine-tune LLM based on the instruction set of the field. The instruction set is a data set constructed for a specific task, which contains explicit questions and corresponding expected outputs, and can help LLM learn the specific task of the professional field, so that LLM can accurately generate outputs meeting the requirements according to the given instructions.
[0003] In related technologies, when applying LLM to professional fields, the instruction set is constructed in the following way: a large number of labeled instructions are used to train an instruction generation model, the trained instruction generation model is used to generate instructions in the professional field, and at the same time, a training instruction scoring model is used to score and evaluate the generated instructions, and then instructions meeting the conditions are selected.
[0004] However, the above method needs to rely on a large number of labeled instructions for model training, which is complex and inefficient. SUMMARY
[0005] The present application provides an instruction generation method, device, computing system and computing device, which utilizes the generation capability of the large language model, and can improve the quality of the instruction set on the basis of efficiently generating the instruction set.
[0006] In a first aspect, the present application provides an instruction generation method applied to a scenario of constructing an instruction set according to a specified field, wherein the present application does not limit the field related to the instruction set, for example, it can be the field of history, literature and art, music, medicine, communication technology, programming technology, etc. The method comprises:
[0007] inputting a first prompt word into a large language model to generate a first instruction set related to a document segment;
[0008] evaluating the first instruction set to obtain a first evaluation result, the first evaluation result indicating the quality of each instruction in the first instruction set;
[0009] when the first evaluation result does not meet a preset condition, adjusting the first prompt word based on the first evaluation result to obtain a second prompt word, so as to determine a target instruction set related to the document segment according to the second prompt word.
[0010] In the above method, the generation capability of the large language model is utilized, the prompt word is input into the large language model to generate an instruction set related to the document segment, then the instruction set is evaluated, and when the evaluation result does not meet the preset condition, the prompt word is adjusted in a timely manner according to the evaluation result to determine a target instruction set related to the document segment. In this way, a feedback relationship between instruction generation and instruction evaluation is established, which can adjust the prompt word in a timely manner according to the instruction evaluation result in the process of efficiently generating the instruction set, that is, adjust the strategy of the large language model to generate instructions, thereby improving the quality of the finally obtained instruction set. In addition, the above method can input the prompt word into the existing large language model, so that the large language model quickly generates a corresponding instruction set according to the indication of the prompt word, and the quality of the generated instruction set is ensured by combining the instruction evaluation result, thereby not relying on labeled instructions to train a special instruction generation model, achieving convenience and high efficiency. In contrast, related technologies need to rely on a large number of labeled instructions for model training when applying LLM to a professional field, which is complex and inefficient.
[0011] In some embodiments, the first prompt word includes the document segment and / or at least one labeled instruction. The first prompt word includes the document segment and / or at least one labeled instruction means that the first prompt word can include the document segment, or the document segment and at least one labeled instruction. The at least one labeled instruction can be a labeled instruction randomly selected from the labeled instruction set, or a labeled instruction selected by a user, which is not limited in the present application. Wherein, in the case that the first prompt word includes the document segment and at least one labeled instruction, since the first prompt word includes both the document segment of the specified field and the labeled instruction, the large language model can have both the knowledge of the specified field and the examples of the instructions when generating the instructions, and generate the instructions of the specified field that meet the examples, based on which the generated instructions can be directly applied to the specified field, improving the application effect, for example, using the generated instructions to train the large language model, so that the large language model quickly learns the knowledge of the specified field, etc.
[0012] In some embodiments, the first prompt word further includes a constraint condition, the constraint condition being used to describe a feature that the instruction needs to have; the first prompt word is adjusted based on the first evaluation result to obtain a second prompt word, including: at least one of the at least one labeled instruction and the constraint condition is adjusted based on the first evaluation result to obtain the second prompt word.
[0013] When the first evaluation result does not satisfy the preset condition, it indicates that the overall quality of the instructions in the first instruction set does not meet the requirements (or the number of instructions with qualified quality is small). Since the first evaluation result can indicate the quality of each instruction in the first instruction set, the first prompt can be adjusted based on the first evaluation result, at least one of the at least one labeled instruction and the constraint condition is adjusted to obtain a second prompt. For example, the instructions with qualified quality indicated by the first evaluation result are used as new labeled instructions. In this way, the second prompt can guide the large language model to generate instructions based on the new labeled instructions, thereby increasing the number of instructions with qualified quality. For another example, the constraint condition is adjusted based on the instructions with unqualified quality indicated by the first evaluation result. In this way, the second prompt can guide the large language model to generate instructions based on the adjusted constraint condition, thereby improving the quality of the instructions.
[0014] In some embodiments, based on the first evaluation result, at least one of the at least one labeled instruction and the constraint condition is adjusted to obtain a second prompt, including:
[0015] Based on the first evaluation result, at least one first instruction that meets the evaluation rule is determined from the first instruction set, and the evaluation rule is used to evaluate the characteristics of the instructions to measure the quality of the instructions.
[0016] Based on the at least one first instruction and the constraint condition, a second prompt is obtained.
[0017] In the above manner, when the first evaluation result does not satisfy the preset condition, since the first evaluation result can indicate the quality of each instruction in the first instruction set, the labeled instructions in the first prompt can be adjusted based on the first evaluation result to obtain a second prompt. For example, the first instructions that meet the evaluation rule are used as labeled instructions in the second prompt, or the first instructions are added to the first prompt to obtain the second prompt. In this way, the second prompt can guide the large language model to generate a new instruction set that meets the evaluation rule based on the adjusted labeled instructions, thereby increasing the number of instructions with qualified quality in the instruction set and improving the quality of the instruction set.
[0018] In some embodiments, based on the first evaluation result, at least one of the at least one labeled instruction and the constraint condition is adjusted to obtain a second prompt, including:
[0019] At least one second instruction that does not meet the evaluation rule is determined from the first instruction set;
[0020] Based on the evaluation rule that is not met by the at least one second instruction, the constraint condition is adjusted to obtain an adjusted constraint condition;
[0021] Based on the adjusted constraint condition and the at least one labeled instruction, a second prompt is obtained.
[0022] In this way, when the first evaluation result does not satisfy the preset condition, since the first evaluation result can indicate the quality of each instruction in the first instruction set, the constraint condition in the first prompt word can be adjusted in a timely manner according to the first evaluation result to obtain a second prompt word, for example, a new feature required by the instruction is determined according to the evaluation rule that the second instruction does not meet, and the new feature is added to the constraint condition to obtain the second prompt word, so that the second prompt word can guide the large language model to generate the instruction with the adjusted constraint condition, and the instruction that meets the feature requirement is regenerated, thereby improving the quality of the instruction set.
[0023] In some embodiments, the preset condition is that the qualified rate of the instructions included in the first instruction set is greater than a threshold, and the qualified rate indicates the proportion of the instructions in the first instruction set that meet the evaluation rule.
[0024] In some embodiments, the first instruction set is evaluated to obtain the first evaluation result, including: evaluating each instruction in the first instruction set by a large language model or an instruction scoring model to obtain the first evaluation result.
[0025] In this way, the quality of each instruction in the first instruction set is quickly evaluated by using the capability of the large language model, and the efficiency of obtaining the first evaluation result is improved on the basis of intuitively reflecting the quality of each instruction. Alternatively, the quality of each instruction in the first instruction set is quantitatively scored by using the instruction scoring model, which can intuitively reflect the quality of the instructions in the first instruction set, and facilitate subsequent screening of the instructions in the first instruction set.
[0026] In some embodiments, the target instruction set is used for at least one of the following: training a large language model; storing to a knowledge base to realize retrieval augmented generation (RAG); and storing to a question and answer database to realize question and answer pair matching. Since the instruction generation method provided by the present application can generate a high-quality instruction set, when the target instruction set is applied to the above scenarios, the accuracy of the answer can be improved. Moreover, the application scenarios of the target instruction set are not limited to the above few scenarios, which are not limited by the present application.
[0027] In a second aspect, the present application provides an instruction generation device, the device comprising at least one functional module, and the at least one functional module is configured to implement the instruction generation method provided by the first aspect or any one of the possible implementation manners of the first aspect.
[0028] In a third aspect, the present application provides a computing system, the system comprising a host and an acceleration card, the host being configured to control the acceleration card, and the acceleration card being configured to run a large language model, and the system being configured to implement the instruction generation method provided by the first aspect or any one of the possible implementation manners of the first aspect.
[0029] In a fourth aspect, the present application provides a computing device, comprising a processor and a memory, wherein the processor is configured to execute at least one piece of program code stored in the memory, so as to enable the computing device to implement the instruction generation method provided in the first aspect or any possible implementation manner of the first aspect.
[0030] In a fifth aspect, the present application provides a computer program product, which, when executed on a computing device, enables the computing device to implement the instruction generation method provided in the first aspect or any possible implementation manner of the first aspect. The computer program product can be a software package, which can be downloaded and executed on the computing device when the method provided in the first aspect is needed to be implemented.
[0031] In a sixth aspect, the present application provides a computer readable storage medium, which is configured to store at least one piece of program code, and the at least one piece of program code is configured to implement the instruction generation method provided in the first aspect or any possible implementation manner of the first aspect. The storage medium can include, but is not limited to, a volatile memory such as a random access memory (RAM), a non-volatile memory such as a flash memory, a hard disk drive (HDD) and a solid state drive (SSD). BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0033] Figure 2 FIG. 2 is a schematic diagram of a server provided by an embodiment of the present application;
[0034] Figure 3 FIG. 3 is a structural schematic diagram of a computing device provided by an embodiment of the present application;
[0035] Figure 4 FIG. 4 is a principle schematic diagram of an instruction generation method provided by an embodiment of the present application;
[0036] Figure 5 FIG. 5 is an architecture schematic diagram of a computing system provided by an embodiment of the present application;
[0037] Figure 6 FIG. 6 is a flowchart of an instruction generation method provided by an embodiment of the present application;
[0038] Figure 7 FIG. 7 is a schematic diagram of an instruction generation method provided by an embodiment of the present application;
[0039] Figure 8is a schematic diagram of another instruction generation method provided by an embodiment of the present application.
[0040] Figure 9 is a structural schematic diagram of an instruction generation apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION
[0041] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the prompt words, document fragments, annotation instructions and the like involved in the present application are obtained under the condition of full authorization.
[0042] For the convenience of understanding, the key terms and key concepts involved in the present application will be described first.
[0043] An artificial intelligence (AI) model is a mathematical algorithm model that uses machine learning ideas to solve practical problems. Generally, an AI model includes a large number of parameters and calculation formulas (or calculation rules).
[0044] An acceleration card, also known as an acceleration device, an accelerator or an acceleration chip, is a special hardware device or computer system designed to accelerate the calculation process in an AI scenario. In the embodiments of the present application, the acceleration device is, for example, a graphics processing unit (GPU), a neural network processing unit (XPU), an intelligent processing unit (IPU), a tensor processing unit (TPU), a domain specific architecture (DSA) chip, etc., without being limited thereto.
[0045] Large language model (LLM) is an AI language processing model trained on large-scale text data. LLM is usually trained on massive amounts of text data, which comes from a wide range of sources, including internet web pages, books, news articles, academic papers, social media posts, etc. LLM can understand the questions raised by users and generate accurate answers. Through learning a large amount of knowledge, LLM can answer questions in various fields, including history, science, technology, culture, etc.
[0046] Supervised fine-tuning (SFT), also known as instruction tuning, refers to further fine-tuning of an already trained large language model by using labeled task-specific data to make the model follow instructions. When applying LLM to a certain field, in order to make the model have the ability to understand and respond to human instructions, it is also necessary to fine-tune the model using the instruction set of that field.
[0047] Instruction set is a data set built for a specific task, containing explicit questions and corresponding expected outputs, which can help LLM learn specific tasks in a professional field, so that LLM can accurately generate outputs that meet the requirements according to the given instructions.
[0048] Labeled instruction is an instruction obtained by labeling documents by machine or manually, which usually has accuracy and reliability. In some scenarios, labeled instruction is also called seed instruction. For example, if a document contains "XXX is a practical programming tool", the labeled instruction obtained by labeling this part of the content is "question: what is XXX; answer: XXX is a programming tool".
[0049] Prompt is an input provided to AI model to guide and stimulate AI model to generate corresponding response or content. For example, prompt is a complete text content (or a sentence), which can be used as input of LLM, so that LLM can output corresponding reasoning result according to the prompt.
[0050] The prompt word template is a template text including variables and fixed text content, and a prompt word can be dynamically generated by inputting specific values of the variables. For example, in actual application, by adjusting the specific values of the variables in the prompt word template, multiple different prompt words can be dynamically generated. Illustratively, taking that the LLM is used to generate an instruction set as an example, the prompt word template is:
Please generate instructions based on the content of the document {variable}
Please generate instructions based on the content of the document {XXX is a practical programming tool}
[0051] Retrieval augmented generation (RAG) is a method that combines retrieval technology with generative models to improve the accuracy, relevance, and reliability of generated text. For example, when the LLM answers a question about a historical event, the system will retrieve relevant historical records, biographies of historical figures, and other text fragments from a historical literature knowledge base, and then combine these text fragments with the LLM to generate the final answer.
[0052] The application scenarios and implementation environments of the present application are introduced as follows.
[0053] The present application is applied to the scenario of constructing an instruction set according to a specified field, wherein the present application does not limit the field related to the instruction set, for example, it can be a field of history, literature, music, medicine, communication technology, programming technology, etc.
[0054] In the related art, when the LLM is applied to a professional field, the way to construct an instruction set is, for example, to train an instruction generation model using a large number of labeled instructions, to generate instructions in the professional field using the trained instruction generation model, and to train an instruction scoring model to score and evaluate the generated instructions, and then to select instructions that meet the conditions. However, this method requires a large number of labeled instructions for model training, and is complex to implement and low in efficiency.
[0055] Therefore, the present application provides an instruction generation method that utilizes the generation capability of a large language model (LLM) and combines the evaluation results of the instruction set during the generation of the instruction set, thereby improving the quality of the instruction set on the basis of efficiently generating the instruction set. The type of LLM is not limited in the present application, and any LLM that can provide instruction generation function is applicable to the present application. For example, the LLM is an open source LLM.
[0056] The implementation environment of the present application is introduced as follows. Figures 1 to 3
[0057] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. For example... Figure 1 As shown, the implementation environment includes a computing system 100, which includes a host 101 and an accelerator card 102, and the host 101 and the accelerator card 102 are communicatively connected. In some embodiments, the accelerator card 102 is also referred to as an acceleration device, accelerator, etc., but this application is not limited thereto.
[0058] In this embodiment, the computing system 100 has instruction generation capabilities, enabling it to determine prompt words based on documents in a specified domain, and input the prompt words into a large language model to obtain a set of instructions related to the document. Schematic, the host 101 provides AI services and can control the accelerator card 102 to execute various computational tasks involved in the instruction generation process. For example, in response to an instruction generation request for a document fragment, the host 101 determines the prompt words corresponding to the document fragment, sends the prompt words to the accelerator card 102, and controls the accelerator card 102 to run a large language model to process the prompt words and generate a set of instructions related to the document fragment.
[0059] The number of hosts 101 can be one or more, and this application does not limit this. Accelerator cards 102 can be, for example, GPUs, XPUs, IPUs, TPUs, DSA chips, etc., and this application is not limited to these. Furthermore, Figure 1 The number of accelerator cards 102 shown is for illustrative purposes only. The number of accelerator cards 102 may be more or less, and this application is not limited thereto.
[0060] In the computing system 100, each accelerator card 102 can be connected via a high-speed interconnect link (or chip bus), enabling different accelerator cards 102 to quickly access each other's memory and achieve efficient data transfer between them. Examples of high-speed interconnect links include NVIDIA Link, Compute Express Link (CXL), Universal Chiplet Interconnect Express (UCIe), Huawei Cache Coherent System (HCCS), Cache Coherent Interconnect for Accelerators (CCIX), etc., and this application is not limited to these.
[0061] Furthermore, in the computing system 100, the host 101 and the accelerator card 102 can be integrated into a single server or configured separately; this application does not impose any limitations on this. Illustratively, taking the integration of the host 101 and the accelerator card 102 into a single server as an example, refer to... Figure 2, Figure 2 is a schematic diagram of a server provided by an embodiment of the present application, as shown in Figure 2 The host 101 and the acceleration card 102 are connected in communication through a peripheral component interconnect express (PCIe) link, and the host 101 and the acceleration card 102 exchange data through the PCIe link. For example, the server is a standalone physical server.
[0062] In some embodiments, the computing system 100 can also be deployed on a server cluster composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and big data and artificial intelligence platform. For example, the cloud server is also called a cloud platform (i.e., an abbreviation of a cloud computing platform), which is a service based on hardware resources and software resources, providing computing, network, and storage capabilities. Through the network "cloud", huge data computing is processed and analyzed in the far end and returned to the user, with characteristics such as large-scale, distributed, virtualization, high availability, scalability, on-demand service, and security. The cloud platform can realize the rapid delivery and release of configurable computing resources with small management cost or low interaction complexity between the user and the service provider.
[0063] In some embodiments, the wireless network or the wired network described above uses standard communication technologies and / or protocols. The network is usually a transmission control protocol / internet protocol (TCP / IP) network and an RDMA network in a data center network, such as a RDMA over converged Ethernet (RoCE) network, an InfiniBand (IB) network, etc., without limitation. In some other embodiments, custom and / or dedicated data communication technologies can be used instead of or in addition to the above data communication technologies.
[0064] The hardware structure of the host 101 in the computing system 100 described above will be introduced below. The present application provides a computing device that can be configured as the host 101 described above, and the hardware structure of the host 101 is shown in Figure 3 , Figure 3 is a structural schematic diagram of a computing device provided by an embodiment of the present application. As shown in Figure 3As shown, the computing device 300 includes a memory 301, a processor 302, a communication interface 303, and a bus 304. The memory 301, the processor 302, and the communication interface 303 are communicatively connected to each other through the bus 304.
[0065] The memory 301 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magneto-optical disk, a magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto. Illustratively, the memory 301 is used to store at least one program code, and when the program code stored in the memory 301 is executed by the processor 302, the processor 302 is used to execute the following instruction generation method.
[0066] The processor 302 can be a network processor (NP), a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or an integrated circuit for controlling program execution of the scheme of the present application. The processor 302 can be a single-CPU processor or a multi-CPU processor. The number of processors 302 can be one or more.
[0067] The communication interface 303 uses a transceiving module such as a transceiver to realize communication between the computing device 300 and other devices or communication networks. For example, data can be obtained through the communication interface 303.
[0068] The memory 301 and the processor 302 can be separately arranged or integrated together.
[0069] Bus 304 can include a pathway that transmits information between various components of computing device 300 (e.g., memory 301, processor 302, communication interface 303).
[0070] Based on the above introduction of the application scenario and implementation environment of the present application, the instruction generation method provided by the present application will be introduced below.
[0071] For the convenience of understanding, the principle of the present application will be introduced below with reference to Figure 4 and Figure 5 .
[0072] Figure 4 is a schematic diagram of the principle of an instruction generation method provided by an embodiment of the present application. As shown in Figure 4 , the instruction generation method provided by the present application involves the following parts: documents of a specified field, a large language model LLM, instruction generation, and instruction evaluation. Illustratively, after the computing system obtains the documents of the specified field, the documents are segmented to obtain a plurality of document segments (also referred to as domain knowledge), and the labeled instructions obtained based on the documents (also referred to as seed instructions) are obtained. In the process of generating instructions based on the LLM, the prompt words are generated in combination with the domain knowledge and / or the labeled instructions, the prompt words are input into the LLM, the instruction set of the field is generated, and it should be understood that the content of the prompt words can also be understood as an instruction generation strategy, which can guide the LLM to generate a corresponding instruction set according to the prompt words. Then, the instruction set generated by the LLM is evaluated to obtain an evaluation result to indicate the quality of each instruction, and the evaluation result is used to provide feedback for the instruction generation strategy, for example, adjusting the content of the prompt words, so that the LLM regenerates the instruction set of the field according to the adjusted prompt words. For another example, the evaluation result indicates that the quality of the currently generated instruction set meets the requirements, and the currently generated instruction set is taken as the final instruction set (the currently generated instruction set can also be further screened). In this way, a feedback relationship between instruction generation and instruction evaluation is established, which can adjust the instruction generation strategy in time according to the instruction evaluation result in the instruction generation process, and thus improve the quality of the finally generated instruction set.
[0073] Based on the principle shown in Figure 4 , the functions possessed by the computing system provided by the present application will be introduced below in combination with the computing system involved in the above implementation environment. Figure 5 is a schematic diagram of the architecture of a computing system provided by an embodiment of the present application. As shown in Figure 5 , the computing system is used to provide instruction generation function 501 and instruction evaluation function 502.
[0074] The instruction generation function 501 includes prompt word determination, instruction generation, instruction analysis, and instruction filtering, etc. For example, the computing system determines a prompt word according to a document segment and a labeled instruction of a specified field, inputs the prompt word into a large language model to generate a plurality of instructions related to the document segment, and obtains an instruction set related to the document segment after analyzing (such as unifying the instruction format, etc.) and filtering (such as screening instructions containing the same question) the plurality of instructions. Illustratively, the instruction generation function 501 is implemented by the host and the accelerator card of the computing system in cooperation, for example, the host determines the prompt word and sends the prompt word to the accelerator card to control the accelerator card to run the large language model to generate a plurality of instructions related to the document segment.
[0075] The instruction evaluation function 502 includes instruction evaluation, result feedback, result screening, etc. For example, the computing system evaluates the instruction set generated by the instruction generation function 501 based on evaluation rules to obtain an evaluation result to indicate the quality of each instruction. According to the indication of the evaluation result, feedback is provided for the instruction generation strategy, such as adjusting the content of the prompt word; or further screening the current generated instruction set. Through the above process, the target instruction set related to the document segment is obtained. Illustratively, the instruction evaluation function 502 is implemented by the host and the accelerator card of the computing system in cooperation, for example, the host controls the accelerator card to run the large language model to evaluate the instruction set to obtain the evaluation result, and the host adjusts the content of the prompt word or performs the instruction set screening step according to the indication of the evaluation result.
[0076] In addition, the function division of the computing system is not limited to the above Figure 5 The content shown, in actual application, more functions can be set according to user needs, for example, the computing system is also used to provide storage function, for providing storage space for document segments, labeled instructions, large language models, instruction sets, etc. In addition, the functions provided by the computing system can be allocated to the host and / or the accelerator card according to the needs, which is not limited in the present application.
[0077] The following refers to the embodiment shown in Figure 6 The process of the instruction generation method provided by the computing system of the present application is introduced.
[0078] Figure 6 is a flowchart of an instruction generation method provided by an embodiment of the present application. As shown in the method executed by the computing system, for example, the following steps 601 to 605 are included. Figure 6
[0079] 601、The computing system obtains a document segment and a labeled instruction.
[0080] In the embodiments of the present application, the document segments and the labeling instructions belong to the same field, wherein the document segments refer to the texts obtained by segmenting the documents of the specified field, and the labeling instructions refer to the instructions obtained by labeling the documents of the specified field. In this step, the present application does not limit the number and source of the document segments obtained by the computing system, and the number and source of the labeling instructions. For example, the document segments can be the document segments obtained by segmenting the documents after the computing system obtains the documents, or the document segments uploaded by the user. For another example, the labeling instructions can be the authorized public labeling instructions, or the labeling instructions uploaded by the user, etc. Illustratively, when the number of the labeling instructions obtained by the computing system is more than one, the multiple labeling instructions are also called a labeling instruction set. In addition, the present application does not limit the field to which the document segments and the labeling instructions belong, for example, the fields of history, literature and art, music, medicine, communication technology, programming technology, etc.
[0081] In some embodiments, taking the example of segmenting the obtained documents by the computing system to obtain the document segments, after obtaining the documents, the computing system performs document parsing to convert the documents into a format convenient for subsequent processing by the computing system, and then the computing system segments the obtained documents into multiple document segments in the manner of sentence segmentation or fixed-length segmentation. It should be understood that since the document formats are various (such as.txt,.csv,.doc,.docx,.xls,.xlsx,.ppt,.pptx, PDF file, web file, etc.), the sources of the documents are also various, and may contain noise, impurities and inconsistent formats, therefore, by performing document parsing, text cleaning and standardization processing, etc., the documents are converted into a unified format for subsequent processing by the computing system. Among them, the text cleaning is, for example, removing special characters, stop words, redundant spaces, etc.; the standardization processing is, for example, converting the text into a unified case, stem extraction, part-of-speech tagging, etc.
[0082] 602、The computing system inputs the first prompt word into the large language model to generate a first instruction set related to the document segment.
[0083] In the embodiments of the present application, the first prompt word includes a document segment and / or at least one annotation instruction. After obtaining the document segment and the annotation instruction through the foregoing step 601, for any one document segment, the computing system fills the document segment and / or the at least one annotation instruction into the first prompt word template of the large language model to obtain the first prompt word, inputs the first prompt word into the large language model, and generates a first instruction set related to the document segment, wherein the first instruction set includes multiple instructions. For example, the host in the computing system sends the first prompt word template to the accelerator card, and sends the document segment and / or the at least one annotation instruction to the accelerator card, the accelerator card fills the document segment and / or the at least one annotation instruction into the first prompt word template of the large language model to obtain the first prompt word, inputs the first prompt word into the large language model, and runs the large language model to generate the instruction set. In some embodiments, the first prompt word is uploaded by a user, which is not limited in the present application.
[0084] In addition, the first prompt word includes a document segment and / or at least one annotation instruction means that the first prompt word can include a document segment, or can include a document segment and at least one annotation instruction. The at least one annotation instruction can be a randomly selected annotation instruction from the annotation instruction set, or can be an annotation instruction selected by a user, which is not limited in the present application. Wherein, in the case that the first prompt word includes a document segment and at least one annotation instruction, since the first prompt word includes both a document segment of a specified field and an annotation instruction, the large language model can have both knowledge of the specified field and examples of instructions when generating instructions, so as to generate instructions in the specified field that conform to the examples, based on which the generated instructions can be directly applied to the specified field, improving the application effect, for example, using the generated instructions to train the large language model, so that the large language model quickly learns the knowledge of the specified field, etc.
[0085] In some embodiments, the first prompt word further includes a constraint condition for describing a feature that the instruction generated by the large language model needs to have, or in other words, by adding the constraint condition in the first prompt word, the large language model is guided to generate instructions with specified features. Wherein, the features that the instruction needs to have include subjective features of the instruction and / or objective features of the instruction, the subjective features refer to features that are affected by human factors, and the objective features refer to features that can be quantified and are not easily affected by human factors. Illustratively, the subjective features of the instruction include but are not limited to at least one of the following: accuracy, used to represent the degree of consistency of the instruction with objective facts; logicality, used to represent the clarity of the language logic of the instruction; completeness, used to represent the comprehensiveness of the content of the instruction; language quality, used to represent the quality of the textual expression of the instruction; and the like. The objective features of the instruction include but are not limited to at least one of the following: instruction word number, instruction number, instruction sentence structure, instruction language, and the like.
[0086] Correspondingly, the constraint condition includes at least one of the following: the instruction needs to be accurate, the instruction needs to be logical, the instruction needs to be complete, the number of instruction words is greater than XX, the number of instructions is greater than YY, the sentence structure of the instruction is ZZ structure, the language of the instruction is WW language, etc. The present application does not limit this.
[0087] It should be understood that the content of the first prompt word can also be understood as an instruction generation strategy that can guide the large language model to generate a corresponding first instruction set according to the first prompt word. The content of the first prompt word is illustrated below. Taking the first prompt word template as an example:
Please generate instructions based on the following content: document segment {variable 1}; label instruction {variable 2}; constraint condition {variable 3}
[0088] For example, the document segment is "XXX is a practical programming tool", the label instruction is "question: what is XXX; answer: XXX is a programming tool", and the constraint condition is "generate 3 instructions; the number of words in each instruction question is not more than 10; the number of words in each instruction answer is greater than 5; the instruction is accurate". Based on this, the first prompt word is
Please generate instructions based on the following content: document segment {XXX is a practical programming tool}; label instruction {question: what is XXX; answer: XXX is a programming tool}; constraint condition {generate 3 instructions; the number of words in each instruction question is not more than 10; the number of words in each instruction answer is greater than 5}
[0089] 603、The computing system evaluates the first instruction set to obtain a first evaluation result, the first evaluation result indicating the quality of each instruction in the first instruction set.
[0090] In the embodiments of the present application, for each instruction in the first instruction set, the computing system evaluates the instruction based on the evaluation rule to obtain a first evaluation result. The evaluation rule can be a default evaluation rule configured by the computing system or an evaluation rule uploaded by a user, which is not limited in the present application. In addition, the number of evaluation rules can be one or more, which is not limited in the present application. It should be understood that the evaluation rule is used to evaluate the characteristics of the instruction to measure the quality of the instruction, that is, for the instruction generated by the large language model, the computing system can evaluate the characteristics currently possessed by the instruction according to the evaluation rule to obtain the quality of the instruction. Illustratively, the characteristics of the instruction evaluated by the evaluation rule include objective characteristics of the instruction and / or subjective characteristics of the instruction.
[0091] For example, the instruction characteristic evaluated by the evaluation rule is the subjective characteristic-accuracy, the evaluation rule can be "whether the answer in the instruction contains incorrect or misleading expressions"; the instruction characteristic evaluated by the evaluation rule is the subjective characteristic-logic, the evaluation rule can be "whether the answer in the instruction has self-contradiction or unreasonable jump"; the instruction characteristic evaluated by the evaluation rule is the subjective characteristic-completeness, the evaluation rule can be "whether the answer in the instruction gives a comprehensive response to the question"; the instruction characteristic evaluated by the evaluation rule is the subjective characteristic-language quality, the evaluation rule can be "whether the grammar of the question and answer in the instruction is correct, whether the words are appropriate and accurate, etc.". That is, the content of the evaluation rule can be refined according to the subjective characteristics of the instruction to be evaluated, the subjective characteristics of the instruction are measured by specific evaluation indexes, and then the quality of the evaluated instruction is obtained.
[0092] For example, the instruction characteristic evaluated by the evaluation rule is the objective characteristic-instruction word number, the evaluation rule can be "whether the word number of the instruction is greater than XX"; the instruction characteristic evaluated by the evaluation rule is the objective characteristic-instruction quantity, the evaluation rule can be "whether the instruction quantity is greater than YY"; the instruction characteristic evaluated by the evaluation rule is the objective characteristic-sentence structure of the instruction, the evaluation rule can be "whether the sentence structure of the instruction is ZZ structure"; the instruction characteristic evaluated by the evaluation rule is the objective characteristic-language of the instruction, the evaluation rule can be "whether the language of the instruction is WW language". That is, the content of the evaluation rule can be set according to the objective characteristics of the instruction to be evaluated, the objective characteristics of the instruction are accurately measured by specific evaluation indexes, and then the quality of the evaluated instruction is obtained.
[0093] It should be noted that the disclosure does not limit the content of the evaluation rule, and in actual application, a suitable evaluation rule can be selected according to the needs to improve the comprehensiveness of the evaluation result. In addition, based on the foregoing introduction of the first prompt word, in some scenarios, the first prompt word can include a constraint condition, and the constraint condition is used to describe the characteristics that the instruction generated by the large language model needs to have. The characteristics evaluated by the evaluation rule can also not include the characteristics described by the constraint condition, or more than the characteristics described by the constraint condition. It should be understood that the constraint condition in the first prompt word and the evaluation rule in this step are two types of information with different functions. The constraint condition is to guide the large language model to generate an instruction with a specified feature, and the evaluation rule is to reflect the quality of the instruction by evaluating the characteristics of the instruction. In some embodiments, the characteristics described by the constraint condition are different from the characteristics evaluated by the evaluation rule, for example, the constraint condition is used to describe the objective characteristics of the instruction, and the evaluation rule is used to evaluate the subjective characteristics of the instruction. In other embodiments, the characteristics described by the constraint condition and the characteristics evaluated by the evaluation rule have overlapping parts, for example, the constraint condition is used to describe the objective characteristics and subjective characteristics of the instruction, and the evaluation rule is used to evaluate the subjective characteristics of the instruction, and the like, which are not limited by the present application.
[0094] Illustratively, the computing system evaluates the instructions includes: evaluating each instruction in the first instruction set by the large language model or the instruction scoring model to obtain the first evaluation result. The following will be introduced respectively:
[0095] Method one, through the large language model, whether each instruction in the first instruction set conforms to the evaluation rule is evaluated to obtain the first evaluation result. Wherein, the computing system fills the first instruction set and the evaluation rule to the second prompt word template to obtain the third prompt word, inputs the third prompt word into the large language model, and evaluates whether each instruction in the first instruction set conforms to the evaluation rule through the large language model to obtain the first evaluation result.
[0096] The content of the third prompt word will be illustrated below. Taking the second prompt word template as an example:
Please evaluate the instructions according to the following evaluation rules: evaluation rule {variable 4}; instruction set {variable 5}
[0097] For example, the evaluation rules include rule A "whether the question of the instruction has a syntax problem", rule B "whether the answer of the instruction is complete and correct", and rule C "whether the answer of the instruction has self-contradiction or unreasonable jumps", hereinafter referred to as "rule A, rule B, rule C", and the first instruction set includes instruction 1, instruction 2, and instruction 3. Based on this, the third prompt word is
Please evaluate the instructions according to the following evaluation rules: evaluation rules {rule A, rule B, rule C}; instruction set {instruction 1, instruction 2, instruction 3}
[0098] In the above manner, the ability of the large language model is utilized to quickly evaluate each instruction in the first instruction set, and on the basis of intuitively reflecting the quality of each instruction, the efficiency of obtaining the first evaluation result is improved.
[0099] Method two, the quality of each instruction in the first instruction set is scored by an instruction scoring model to obtain the first evaluation result. The instruction scoring model is an AI model trained based on evaluation rules and sample instructions, which can quantitatively score the quality of the instruction. The instruction scoring model can be an instruction scoring model uploaded by a user, or a large language model, and the type and structure of the instruction scoring model are not limited in the present application.
[0100] For example, the first instruction set includes instruction 4, instruction 5, and instruction 6, and the instruction scoring model can score the quality of each instruction according to evaluation rules such as rule D "the number of syntax problems in the question of the instruction", rule E "the accuracy of the answer in the instruction", and rule F "the number of syntax problems in the answer of the instruction", to obtain the quality score (for example, the score interval is 0-10) of each instruction. Correspondingly, the first evaluation result is, for example: instruction 4 is 9 points, instruction 5 is 7.5 points, and instruction 6 is 5 points.
[0101] In the above manner, the quality of each instruction in the first instruction set is quantitatively scored by the instruction scoring model, which can intuitively reflect the quality of the instructions in the first instruction set, and facilitate subsequent screening of the instructions in the first instruction set.
[0102] It should be noted that the above several evaluation methods are only for illustration, and in actual application, a suitable evaluation method can be selected to obtain the first evaluation result according to the needs, which is not limited in the present application.
[0103] 604、The computing system adjusts the first prompt word based on the first evaluation result to obtain a second prompt word when the first evaluation result does not satisfy a preset condition.
[0104] In the embodiments of the present application, the preset condition refers to that the qualified rate of the instructions included in the first instruction set is greater than a threshold value, wherein the qualified rate indicates the proportion of the instructions in the first instruction set that meet the evaluation rules. The threshold value is a preset threshold value, which can be set according to the requirements in actual application, for example, the threshold value is 60%, that is, when the qualified rate of the instructions included in the first instruction set is greater than 60%, it is determined that the first instruction set satisfies the preset condition. Based on the foregoing step 603, it can be known that the first evaluation result can be obtained by the large language model evaluating whether the quality of each instruction in the first instruction set meets the evaluation rules, or can be obtained by the instruction scoring model scoring the quality of each instruction in the first instruction set. Correspondingly, different types of first evaluation results can correspond to different qualified rate determination methods.
[0105] In some embodiments, the first evaluation result is obtained by the large language model evaluating whether each instruction in the first instruction set meets the evaluation rules, and correspondingly, the first evaluation result includes the conclusion of whether each instruction in the first instruction set meets the evaluation rules. Based on this, the computing system can determine the qualified rate of the instructions in the first instruction set according to the conclusion of whether each instruction meets the evaluation rules. For example, the first evaluation result includes: instruction 1 meets rule A and rule B, and does not meet rule C; instruction 2 meets rule A, rule B and rule C; instruction 3 meets rule A, and does not meet rule B and rule C. Based on this, the computing system regards the instructions meeting all evaluation rules as the instructions meeting the evaluation rules (or instructions with qualified quality), that is, instruction 2 is the instruction meeting the evaluation rules, and the qualified rate is 1 / 3. It should be noted that this is only an example for illustration, and the computing system can also regard the instructions meeting part of the evaluation rules as the instructions with qualified quality, which is not limited in the present application.
[0106] In other embodiments, the first evaluation result is obtained by the instruction scoring model scoring the quality of each instruction in the first instruction set, and correspondingly, the first evaluation result includes the quality score of each instruction in the first instruction set. Based on this, the computing system can determine the qualified rate of the instructions in the first instruction set according to the quality score of each instruction. For example, the first evaluation result includes: instruction 4 is 9 points, instruction 5 is 7.5 points, and instruction 6 is 5 points. Based on this, the computing system can regard the instructions with scores greater than 7 as the instructions meeting the evaluation rules (or instructions with qualified quality), that is, instruction 4 and instruction 5 are the instructions meeting the evaluation rules, and the qualified rate is 2 / 3. It should be noted that this is only an example for illustration, and the computing system can also set other score thresholds to determine the instructions with qualified quality, which is not limited in the present application.
[0107] In this step, when the first evaluation result indicates that the first instruction set does not meet the preset condition, it means that the overall quality of the instructions in the first instruction set does not meet the requirements (or the number of instructions with qualified quality is small). Since the first evaluation result can indicate the quality of each instruction in the first instruction set, the first prompt can be used as a reference to adjust at least one of the at least one labeled instruction and the constraint condition in the first prompt to obtain a second prompt. For example, the instructions with qualified quality indicated by the first evaluation result can be used as new labeled instructions. In this way, the second prompt can guide the large language model to generate instructions using the new labeled instructions as examples, thereby increasing the number of instructions with qualified quality. For another example, according to the instructions with unqualified quality indicated by the first evaluation result, the constraint condition is adjusted. In this way, the second prompt can guide the large language model to generate instructions using the adjusted constraint condition, thereby improving the quality of the instructions.
[0108] It should be understood that based on the foregoing introduction, the content of the first prompt can be understood as an instruction generation strategy that can guide the large language model to generate a corresponding instruction set according to the first prompt. Accordingly, adjusting the first prompt to obtain the second prompt can also be understood as adjusting the instruction generation strategy, thereby guiding the large language model to generate a corresponding instruction set according to the second prompt.
[0109] Illustratively, since the first prompt includes at least one of the at least one labeled instruction and the constraint condition, in this step, the computing system can adjust at least one of the at least one labeled instruction and the constraint condition based on the first evaluation result to obtain the second prompt. The ways of adjusting the at least one labeled instruction and adjusting the constraint condition are introduced as follows:
[0110] Method one, adjusting the at least one labeled instruction in the first prompt based on the first evaluation result to obtain the second prompt.
[0111] In this method, the computing system determines at least one first instruction that meets the evaluation rule from the first instruction set based on the first evaluation result; and obtains the second prompt based on the at least one first instruction and the constraint condition. Illustratively, the computing system can replace the at least one labeled instruction in the first prompt with the at least one first instruction as the labeled instruction; or add the at least one first instruction to the labeled instruction set to determine at least one labeled instruction from the labeled instruction set again; or add the at least one first instruction as the labeled instruction to the first prompt to obtain the second prompt; the present application does not limit the way in which the computing system obtains the second prompt based on the at least one first instruction.
[0112] In some embodiments, the first evaluation result is obtained by the large language model evaluating whether each instruction in the first instruction set meets the evaluation rule, and accordingly, the first evaluation result includes the conclusion of whether each instruction in the first instruction set meets the evaluation rule. Based on this, the computing system can determine at least one first instruction in the first instruction set that meets the evaluation rule according to the conclusion of whether each instruction meets the evaluation rule. For example, the first evaluation result includes: instruction 1 meets rule A and rule B, and does not meet rule C; instruction 2 meets rule A, rule B, and rule C; and instruction 3 meets rule A, and does not meet rule B and rule C. Based on this, the computing system takes the instruction that meets all the evaluation rules as the first instruction, that is, takes instruction 2 as the first instruction. It should be noted that this is only an example, and the computing system can also take the instruction that meets part of the evaluation rules as the first instruction, which can be set according to the needs in actual application, and the present application does not limit this.
[0113] In other embodiments, the first evaluation result is obtained by the instruction scoring model, and the first evaluation result is obtained by the instruction scoring model scoring the quality of each instruction in the first instruction set. Accordingly, the first evaluation result includes the quality score of each instruction in the first instruction set. Based on this, the computing system can determine the pass rate of the instructions in the first instruction set according to the quality score of each instruction. For example, the first evaluation result includes: instruction 4 is 9 points, instruction 5 is 7.5 points, and instruction 6 is 5 points. Based on this, the computing system can take the instruction with a score greater than 7 as the first instruction, or sort the instructions in the first instruction set according to the quality score, and select the top k instructions as the first instruction (k is a positive integer), and the present application does not limit this.
[0114] In the above manner, when the first evaluation result does not meet the preset condition, since the first evaluation result can indicate the quality of each instruction in the first instruction set, the marked instruction in the first prompt word can be adjusted based on the first evaluation result to obtain a second prompt word, for example, the first instruction that meets the evaluation rule is taken as the marked instruction in the second prompt word, or the first instruction is added to the first prompt word to obtain the second prompt word, so that the second prompt word can guide the large language model to regenerate an instruction set that meets the evaluation rule better with the adjusted marked instruction as an example, thereby improving the number of instructions in the instruction set that meet the quality requirements, and further improving the quality of the instruction set.
[0115] Method two: adjusting the constraint condition in the first prompt word based on the first evaluation result to obtain a second prompt word.
[0116] The computing system determines at least one second instruction that does not conform to the evaluation rule from the first instruction set; adjusts the constraint condition based on the evaluation rule that the at least one second instruction does not conform to, to obtain an adjusted constraint condition; and obtains a second prompt word based on the adjusted constraint condition and the at least one labeled instruction. The computing system can determine a new feature that the instruction generated by the large language model needs to have according to the evaluation rule that the second instruction does not conform to, add the new feature to the constraint condition to obtain the adjusted constraint condition, and then obtain the second prompt word. This process is to adjust the features described by the constraint condition using the features evaluated by the evaluation rule, so that the adjusted constraint condition adds new features to the features described by the original constraint condition, and then obtains the second prompt word. In this way, the large language model can generate instructions with higher quality according to the second prompt word.
[0117] For example, the constraint condition includes "the instruction needs to have accuracy", the evaluation rule that the second instruction does not conform to is used to evaluate this subjective feature of the instruction, the content of the evaluation rule is "whether the answer in the instruction contains incorrect or misleading expressions", and the adjusted constraint condition includes "the instruction needs to have accuracy and the answer in the instruction does not contain incorrect or misleading expressions".
[0118] For another example, the constraint condition includes "the number of words in the instruction is greater than XX", the evaluation rule that the second instruction does not conform to is used to evaluate this objective feature of the instruction, the content of the evaluation rule is "whether the number of words in the instruction is greater than XX", and since the number of words in the second instruction is equal to XX, the adjusted constraint condition is "the number of words in the instruction is greater than XX and cannot be equal to XX".
[0119] For another example, the constraint condition includes "the number of words in the instruction is greater than XX", the evaluation rule that the second instruction does not conform to is used to evaluate the accuracy of the instruction, the content of the evaluation rule is "the answer in the instruction does not contain incorrect or misleading expressions", and the adjusted constraint condition includes "the number of words in the instruction is greater than XX; the answer in the instruction does not contain incorrect or misleading expressions".
[0120] It should be noted that this is only an example for illustration and does not constitute a limitation on the present application. Based on the foregoing step 603, the features described by the constraint condition and the features evaluated by the evaluation rule can be different or can have overlapping parts. In some scenarios, the way of adjusting the constraint condition can be flexibly set according to business needs.
[0121] The following will take different types of first evaluation results as examples to introduce the way of adjusting the constraint condition by the computing system.
[0122] In some embodiments, the first evaluation result is obtained by the large language model evaluating whether each instruction in the first instruction set meets the evaluation rule, and accordingly, the first evaluation result includes the conclusion of whether each instruction in the first instruction set meets the evaluation rule. Based on this, the computing system can determine at least one second instruction that does not meet the evaluation rule from the first instruction set according to the conclusion of whether each instruction meets the evaluation rule. For example, the first evaluation result includes: instruction 1 meets rule A and rule B, and does not meet rule C; instruction 2 meets rule A, rule B, and rule C; and instruction 3 meets rule A, and does not meet rule B and rule C. Based on this, the computing system takes the instruction that does not meet all the evaluation rules as the second instruction, that is, takes instruction 1 and instruction 3 as the second instruction, and since instruction 1 and instruction 3 do not meet rule B and rule C, the constraint condition is adjusted based on rule B and rule C. For example, the constraint condition of the first prompt word is “generate 3 instructions; the number of words of the question in each instruction is not more than 10; and the number of words of the answer in each instruction is greater than 5”, rule B is “whether the answer of the instruction is complete and correct”, and rule C is “whether the answer of the instruction is self-contradictory or unreasonable in jumping”, and the adjusted constraint condition is “generate 3 instructions; the number of words of the question in each instruction is not more than 10; the number of words of the answer in each instruction is greater than 5; the answer of the instruction is complete and correct; and the answer of the instruction is not self-contradictory or unreasonable in jumping”.
[0123] In other embodiments, the first evaluation result is obtained by the instruction scoring model scoring the quality of each instruction in the first instruction set, and accordingly, the first evaluation result includes the quality score of each instruction in the first instruction set. Based on this, the computing system can determine at least one second instruction that does not meet the evaluation rule from the first instruction set according to the quality score of each instruction. For example, the first evaluation result includes: instruction 4 is 9 points, instruction 5 is 7.5 points, and instruction 6 is 5 points. Based on this, the computing system can take the instruction with a score less than or equal to 7 as the second instruction, that is, take instruction 6 as the second instruction, and adjust the constraint condition according to whether instruction 6 meets the evaluation rule. It should be noted that the adjustment of the instruction requirement according to the evaluation rule can be set according to business needs, which is not limited in the present application.
[0124] In the above manner, when the first evaluation result does not meet the preset condition, since the first evaluation result can indicate the quality of each instruction in the first instruction set, the constraint condition in the first prompt word can be adjusted in time according to the first evaluation result to obtain the second prompt word, for example, the new characteristics required by the instruction are determined according to the evaluation rule that the second instruction does not meet, and are added to the constraint condition to obtain the second prompt word, so that the second prompt word can guide the large language model to generate the instruction with the adjusted constraint condition, and the instruction set with better quality can be generated again.
[0125] It should be understood that the above-mentioned manner one and manner two can also be used in combination, that is, the computing system adjusts at least one annotation instruction and constraint condition in the first prompt word based on the first evaluation result to obtain a second prompt word. The principle of this process is the same as the aforementioned manner one and manner two, and thus will not be described again. By using manner one and manner two in combination, when the first evaluation result indicates that the first instruction set does not meet the preset condition, the annotation instruction and the constraint condition in the first prompt word are adjusted in a timely manner according to the first evaluation result to obtain the second prompt word, which can enable the large language model to generate a new instruction set that meets the evaluation rules and the feature requirements more in accordance with the adjusted annotation instruction and the adjusted constraint condition, greatly improving the quality of the instruction set.
[0126] 605、The computing system determines the target instruction set related to the document segment according to the second prompt word.
[0127] In the embodiments of the present application, the computing system inputs the second prompt word into the large language model to generate the target instruction set related to the document segment. Since the second prompt word is the prompt word adjusted on the basis of the first prompt word and based on the first evaluation result of the first instruction set, the second prompt word can better guide the large language model to generate an instruction set with higher quality than the first prompt word, and obtain the target instruction set.
[0128] In some embodiments, the computing system inputs the second prompt word into the large language model to generate a second instruction set related to the document segment, evaluates the second instruction set to obtain a second evaluation result, the second evaluation result indicating the quality of each instruction in the second instruction set, and adjusts the second prompt word based on the second evaluation result to obtain a fourth prompt word when the second evaluation result indicates that the second instruction set does not meet the preset condition, so as to determine the target instruction set related to the document segment according to the fourth prompt word. This process is also that the computing system can regenerate the instruction set related to the document segment according to the second prompt word, and perform the process same as steps 603 to 605 until the instruction set that meets the preset condition is obtained.
[0129] In addition, in steps 604 and 605, the first instruction set is taken as an example to introduce that the first evaluation result indicates that the first instruction set does not meet the preset condition. In some embodiments, when the first evaluation result indicates that the first instruction set meets the preset condition, the computing system filters the target instruction set from the first instruction set based on the first evaluation result. For example, the first evaluation result includes: instruction 1 meets rule A and rule B, and does not meet rule C; instruction 2 meets rule A, rule B, and rule C; instruction 3 meets rule A, and does not meet rule B and rule C, and the like. Based on this, the computing system filters the instructions that meet all evaluation rules to obtain the target instruction set, and the target instruction set includes instruction 2. For another example, the first evaluation result includes: instruction 4 is 9 points, instruction 5 is 7.5 points, instruction 6 is 5 points, and the like. Based on this, the computing system can filter the instructions greater than 7 points to obtain the target instruction set, and the target instruction set includes instruction 4 and instruction 5. It should be noted that this is only an example, and the computing system can also use other ways to filter the target instruction set, which is not limited in the present application. Of course, the computing system can also take the first instruction set as the target instruction set, and the like.
[0130] After obtaining the target instruction set related to the document segment through steps 601 to 605, the target instruction set can be used for at least one of the following:
[0131] (1) Training a large language model. That is, using the target instruction set to train the large language model, so that the large language model learns the knowledge of the field to which the target instruction set belongs, thereby applying the large language model to the field, so that the large language model can accurately generate a required answer according to a question in the field.
[0132] (2) Storing to a knowledge base to realize retrieval augmented generation (RAG). That is, storing the target instruction set to the knowledge base, so that the large language model can retrieve related instructions from the knowledge base when answering a question in the field to which the target instruction set belongs, and combining the retrieved instructions with the large language model to generate a final answer.
[0133] (3) Storing to a question and answer database to realize question and answer pair matching. That is, storing the target instruction set to the question and answer database, and in some question and answer scenarios, the instructions retrieved from the question and answer database can be fed back to the user according to the question input by the user.
[0134] It should be noted that since the instruction generation method provided by the present application can generate a high-quality instruction set, the accuracy of the answer can be improved when the target instruction set is applied to the above scenarios. Moreover, the application scenarios of the target instruction set are not limited to the above few scenarios, which are not limited in the present application.
[0135] It can be seen that, in the instruction generation method provided in the application, the generation capability of the large language model is utilized, the prompt word is input into the large language model to generate an instruction set related to the document segment, then the instruction set is evaluated, and when the evaluation result indicates that the instruction set does not meet the preset condition, the prompt word is adjusted in a timely manner according to the evaluation result to determine a target instruction set related to the document segment. In this way, a feedback relationship between instruction generation and instruction evaluation is established, and the prompt word can be adjusted in a timely manner according to the instruction evaluation result in the process of efficiently generating the instruction set, that is, the strategy of the large language model for generating instructions is adjusted, so that the quality of the finally obtained instruction set is improved.
[0136] The method embodiment shown in the above Figure 7 and Figure 8 is exemplarily explained below with reference to the embodiments shown in the above Figure 6 .
[0137] Figure 7 is a schematic diagram of an instruction generation method provided in the application. As shown in the figure, the instruction generation method includes two parts of instruction generation and instruction evaluation. Figure 7
[0138] In the instruction generation part, the computing system fills the document segment "XXX is a practical programming tool..." and the labeled instruction "question: what is XXX; answer: XXX is a programming tool" into the first prompt word template of the large language model to obtain the first prompt word 701, wherein the first prompt word 701 further includes the constraint condition "YYY". Then, the computing system inputs the first prompt word 701 into the large language model to generate a plurality of instructions related to the document segment, and after analyzing (such as unifying the instruction format) and filtering (such as screening instructions containing the same question) the plurality of instructions, obtains a first instruction set related to the document segment.
[0139] In the instruction evaluation part, the computing system fills the first instruction set and the evaluation rule into the second prompt word template to obtain the third prompt word 702, inputs the third prompt word 702 into the large language model, and through the large language model, evaluates whether the quality of each instruction in the first instruction set meets the evaluation rule (such as rule A, rule B and rule C in the figure) to obtain the first evaluation result 703. Then, the computing system judges whether the first instruction set indicated by the first evaluation result meets the preset condition, that is, the qualified rate r of the instructions included in the first instruction set is greater than the threshold θ (such as 60%). The qualified rate r indicates the proportion of the instructions in the first instruction set that meet the evaluation rule. For example, r is calculated by the following formula (1):
[0140]
[0141] In formula (1) above, I(·) represents the indicator function; n represents the number of instructions in the first instruction set; d i ∈R indicates that the i-th instruction is of acceptable quality, where i is a positive integer.
[0142] Furthermore, if the first evaluation result indicates that the first instruction set meets the preset conditions, then the target instruction set is selected from the first instruction set. If the first evaluation result indicates that the first instruction set does not meet the preset conditions, then the constraints in the first prompt word are adjusted according to the first evaluation result to obtain the second prompt word, and the corresponding instruction set is regenerated based on the second prompt word. This process refers to steps 604 and 605 above, so it will not be described again.
[0143] visible, Figure 7 In the instruction generation method shown, a feedback relationship is established between instruction generation and instruction evaluation. In the process of efficiently generating instruction sets, the constraints in the prompt words can be adjusted in a timely manner according to the instruction evaluation results, thereby improving the quality of the final instruction set. This process can also be understood as an adaptive iterative instruction generation process based on the constraints in the prompt words, which can be expressed as the following formula (2):
[0144] D n =LLM(p gen r n-1 );r n =LLM(p eval D n (2)
[0145] In the above formula (2), D n r represents the instruction set generated in the nth iteration; n p represents the evaluation result generated in the nth iteration. gen This indicates the prompt word generated based on the first prompt word template, i.e., the prompt word used to generate the instruction set; p eval This indicates the prompt words generated based on the second prompt word template, which are also the prompt words used to generate the evaluation results.
[0146] Figure 8 This is a schematic diagram of another instruction generation method provided in an embodiment of this application. For example... Figure 8 As shown, the instruction generation method includes two parts: instruction generation and instruction evaluation. The instruction generation part is similar to the previously mentioned... Figure 7 The same principle applies to the embodiments shown, so they will not be repeated here. The instruction evaluation part will be introduced below.
[0147] In the instruction evaluation part, the computing system scores the quality of each instruction in the first instruction set by the instruction scoring model, to obtain a first evaluation result 801. Then, the computing system determines whether the first instruction set indicated by the first evaluation result satisfies a preset condition, i.e., the qualified rate r of the instructions included in the first instruction set is greater than a threshold value a (such as 60%, it should be noted that the threshold value can be the same as or different from that in the embodiment shown in the figure, which is not limited in the present application), wherein the qualified rate r indicates the proportion of the instructions in the first instruction set that meet the evaluation rules. For example, r is calculated by the following formula (3): Figure 7
[0148]
[0149] In the above formula (3), I(·) represents an indicator function; n represents the number of instructions in the first instruction set, f(x) represents the instruction scoring model, the input x is an instruction, and the output is a quality score; f(di) > σ represents that the ith instruction is a qualified instruction, i is a positive integer; and σ represents a score threshold (such as 7 points) of the instructions meeting the evaluation rules. i
[0150] Further, if the first evaluation result indicates that the first instruction set satisfies the preset condition, a target instruction set is selected from the first instruction set. If the first evaluation result indicates that the first instruction set does not satisfy the preset condition, the labeled instructions in the first prompt word are adjusted according to the first evaluation result, such as selecting the instructions with the top k scores as new labeled instructions, replacing the labeled instructions in the first prompt word to obtain a second prompt word, and generating a corresponding instruction set based on the second prompt word. This process is described in the foregoing steps 604 and 605, and will not be described again.
[0151] As can be seen, Figure 8 In the instruction generation method shown in the figure, a feedback relationship between instruction generation and instruction evaluation is established, which can timely adjust the labeled instructions in the prompt word according to the instruction evaluation result in the process of efficiently generating the instruction set, thereby improving the quality of the finally obtained instruction set. This process can also be understood as a self-adaptive iterative instruction generation process based on the labeled instructions in the prompt word, which can be expressed as the following formula (4):
[0152]
[0153] In the above formula (4), D n represents the instruction set generated in the nth iteration; represents the labeled instructions selected in the nth iteration; p gen represents the prompt word generated based on the first prompt word template, i.e., the prompt word used to generate the instruction set; and k represents the number of selected labeled instructions.
[0154] Based on the above Figures 4 to 8 According to the method embodiment shown in the method embodiment, the application further provides an instruction generation device. Illustratively, referring to Figure 9 , Figure 9 is a structural schematic diagram of an instruction generation device provided by an embodiment of the application. As shown in the figure, the device comprises a generation module 901, an evaluation module 902 and an adjustment module 903: Figure 9 The generation module 901 is configured to input a first prompt word into a large language model to generate a first instruction set related to a document segment.
[0155] The evaluation module 902 is configured to evaluate the first instruction set to obtain a first evaluation result, the first evaluation result indicating the quality of each instruction in the first instruction set.
[0156] The adjustment module 903 is configured to, when the first evaluation result does not satisfy a preset condition, adjust the first prompt word based on the first evaluation result to obtain a second prompt word, so as to determine a target instruction set related to the document segment according to the second prompt word.
[0157] In some embodiments, the first prompt word comprises the document segment and / or at least one labeled instruction.
[0158] In some embodiments, the first prompt word further comprises a constraint condition, the constraint condition being used to describe a feature required by the instruction.
[0159] The adjustment module 903 is configured to adjust at least one of the at least one labeled instruction and the constraint condition based on the first evaluation result to obtain the second prompt word.
[0160] In some embodiments, the adjustment module 903 is configured to:
[0161] based on the first evaluation result, determine at least one first instruction from the first instruction set that meets an evaluation rule, the evaluation rule being used to evaluate the feature of the instruction to measure the quality of the instruction;
[0162] based on the at least one first instruction and the constraint condition, obtain the second prompt word.
[0163] In some embodiments, the adjustment module 903 is configured to:
[0164] determine at least one second instruction from the first instruction set that does not meet the evaluation rule;
[0165] based on the evaluation rule that the at least one second instruction does not meet, adjust the constraint condition to obtain an adjusted constraint condition;
[0166] based on the adjusted constraint condition and the at least one labeled instruction, obtain the second prompt word.
[0167]
[0168] In some embodiments, the preset condition refers to a qualified rate of instructions included in the first instruction set being greater than a threshold value, the qualified rate indicating a proportion of instructions in the first instruction set that meet the evaluation rule.
[0169] In some embodiments, the evaluation module 902 is configured to evaluate each instruction in the first instruction set by using the large language model or the instruction scoring model to obtain a first evaluation result.
[0170] In some embodiments, the target instruction set is used for at least one of the following:
[0171] training the large language model;
[0172] storing to a knowledge base to implement retrieval augmented generation (RAG);
[0173] storing to a question and answer database to implement question and answer pair matching.
[0174] By using the above device, the generation capability of the large language model is utilized, the prompt word is input into the large language model to generate an instruction set related to the document segment, then the instruction set is evaluated, and when the evaluation result indicates that the instruction set does not meet the preset condition, the prompt word is adjusted in a timely manner to determine a target instruction set related to the document segment. In this way, a feedback relationship between instruction generation and instruction evaluation is established, and the prompt word can be adjusted in a timely manner according to the instruction evaluation result in the process of efficiently generating the instruction set, that is, the strategy of generating instructions by the large language model is adjusted, so as to improve the quality of the finally obtained instruction set.
[0175] Of course, the device can also include other functional units to implement the functions involved in the computing system in the above method embodiments. In actual applications, the above functions can be completed by different functional units as needed, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above. In some possible implementation manners, each functional unit can be located on the same computing device, or can be located in different computing devices to cooperatively complete all or part of the functions described above. In addition, the instruction generation device and the instruction generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be described here.
[0176] The present application also provides a computer readable storage medium, which is used to store at least one program code, when the at least one program code is executed by a computing device, the computing device implements the above instruction generation method.
[0177] The present application also provides a computer program product, when the computer program product is run on a computing device, the computing device implements the above instruction generation method.
[0178] The terms "first", "second", and the like in the present application are used to distinguish between similar or identical items or elements having substantially the same function and should be understood as not having a logical or chronological dependency between "first", "second", "n-th", and the like. It should also be understood that although the following description uses the terms first, second, and the like to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various described examples, a first prompt can be referred to as a second prompt, and similarly, a second prompt can be referred to as a first prompt. The first prompt and the second prompt can both be prompts, and in some cases, can be separate and distinct prompts.
[0179] The term "at least one" in the present application means one or more, and the term "a plurality" in the present application means two or more, for example, a plurality of prompts means two or more prompts.
[0180] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0181] In the above-described embodiments, all or part of the steps can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the program structure information can be implemented. The program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, all or part of the processes or functions in the embodiments of the present application are generated.
[0182] A person of ordinary skill in the art can understand that all or part of the steps of the above-described embodiments can be completed by hardware, or by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0183] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, a person of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating instructions, characterized in that, The method includes: Input the first prompt word into the large language model to generate the first instruction set related to the document fragment; The first instruction set is evaluated to obtain a first evaluation result, which indicates the quality of each instruction in the first instruction set. If the first evaluation result does not meet the preset conditions, the first prompt word is adjusted based on the first evaluation result to obtain a second prompt word, so as to determine the target instruction set related to the document fragment according to the second prompt word.
2. The method according to claim 1, characterized in that, The first prompt word includes the document fragment and / or at least one annotation instruction.
3. The method according to claim 2, characterized in that, The first prompt word also includes constraints, which are used to describe the characteristics that the instruction must possess; Based on the first evaluation result, the first prompt word is adjusted to obtain the second prompt word, including: Based on the first evaluation result, adjust at least one of the at least one annotation instruction and at least one of the constraints to obtain the second prompt word.
4. The method according to claim 3, characterized in that, The step of adjusting at least one of the at least one annotation instruction and at least one of the constraints based on the first evaluation result to obtain the second prompt word includes: Based on the first evaluation result, at least one first instruction that conforms to the evaluation rules is determined from the first instruction set. The evaluation rules are used to evaluate the characteristics of the instruction to measure the quality of the instruction. The second prompt word is obtained based on the at least one first instruction and the constraints.
5. The method according to claim 3 or 4, characterized in that, The step of adjusting at least one of the at least one annotation instruction and at least one of the constraints based on the first evaluation result to obtain the second prompt word includes: Identify at least one second instruction from the first instruction set that does not conform to the evaluation rules; Based on the evaluation rule that at least one of the second instructions is not met, the constraints are adjusted to obtain the adjusted constraints. Based on the adjusted constraints and the at least one annotation instruction, the second prompt word is obtained.
6. The method according to claim 1, characterized in that, The preset condition refers to the fact that the pass rate of instructions included in the first instruction set is greater than a threshold, and the pass rate indicates the proportion of instructions in the first instruction set that meet the evaluation rules.
7. The method according to claim 1, characterized in that, The evaluation of the first instruction set to obtain a first evaluation result includes: The first evaluation result is obtained by evaluating each instruction in the first instruction set using the large language model or instruction scoring model.
8. An instruction generation device based on a large language model, characterized in that, The device includes: The generation module is used to input the first prompt word into the large language model and generate the first instruction set related to the document fragment; An evaluation module is used to evaluate the first instruction set and obtain a first evaluation result, wherein the first evaluation result indicates the quality of each instruction in the first instruction set; An adjustment module is used to adjust the first prompt word to obtain a second prompt word based on the first evaluation result when the first evaluation result does not meet the preset conditions, so as to determine the target instruction set related to the document fragment according to the second prompt word.
9. The apparatus according to claim 8, characterized in that, The first prompt word includes the document fragment and / or at least one annotation instruction.
10. The apparatus according to claim 9, characterized in that, The first prompt word also includes constraints, which are used to describe the characteristics that the instruction must possess; The adjustment module is used to adjust at least one of the at least one annotation instruction and the constraint condition based on the first evaluation result to obtain the second prompt word.
11. The apparatus according to claim 10, characterized in that, The adjustment module is used for: Based on the first evaluation result, at least one first instruction that conforms to the evaluation rules is determined from the first instruction set. The evaluation rules are used to evaluate the characteristics of the instruction to measure the quality of the instruction. The second prompt word is obtained based on the at least one first instruction and the constraints.
12. The apparatus according to claim 10 or 11, characterized in that, The adjustment module is used for: Identify at least one second instruction from the first instruction set that does not conform to the evaluation rules; Based on the evaluation rule that at least one of the second instructions is not met, the constraints are adjusted to obtain the adjusted constraints. Based on the adjusted constraints and the at least one annotation instruction, the second prompt word is obtained.
13. The apparatus according to claim 8, characterized in that, The preset condition refers to the fact that the pass rate of the instructions included in the first instruction set is greater than a threshold, and the pass rate indicates the proportion of instructions in the first instruction set that meet the evaluation rules.
14. The apparatus according to claim 8, characterized in that, The evaluation module is used to evaluate each instruction in the first instruction set using the large language model or instruction scoring model to obtain the first evaluation result.
15. A computing system, characterized in that, The computing system includes a host and an accelerator card. The host is used to control the accelerator card, which is used to run a large language model. The system is used to implement the instruction generation method as described in any one of claims 1 to 7.
16. A computing device, characterized in that, The computing device includes a processor and a memory, the processor being configured to execute at least one piece of program code stored in the memory to enable the computing device to implement the instruction generation method as described in any one of claims 1 to 7.
17. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device implements the instruction generation method as described in any one of claims 1 to 7.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which is used to implement the instruction generation method as described in any one of claims 1 to 7.