Large language model training method and information processing method
By first entering the initial knowledge information in the large language model, and then building a training sample set containing the reply information for supervision and training, the problem of low reply accuracy in the existing technology is solved, and higher reply accuracy is achieved.
Patent Information
- Application Number
- CN202510143097.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, by inputting relevant knowledge into large language models for training, the response accuracy of large language models is low.
A training method of a large language model is adopted. First, the first knowledge information is input into the large language model for preliminary learning. Then, the training sample set is constructed based on the second knowledge information, including knowledge samples, first reply information and second reply information. Through this information, the learned large language model is supervised and trained to obtain the target large language model.
Through this method, the reply accuracy of the large language model is improved, so that it can better identify and provide more appropriate reply information.
Smart Images

Figure CN119990325A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method and an information processing method for a large language model. Background Art
[0002] With the emergence of large language models, these models can handle a variety of complex and diverse problems with their excellent understanding ability. For example, intelligent customer service in the e-commerce field can be realized through large language models. Large language model technology can process large amounts of text data, automatically learn and extract language patterns and knowledge graphs, so as to achieve goals such as automatically answering customer questions and providing personalized services. However, in the existing technology, the training of large language models is often achieved by inputting relevant knowledge into the large language model. This method has the problem of relatively low accuracy of the large language model's responses.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present application provide a training method and an information processing method for a large language model, so as to at least solve the technical problem in the related art that the training of the large language model is achieved by inputting relevant knowledge into the large language model, resulting in relatively low accuracy of the large language model's responses.
[0005] According to one aspect of an embodiment of the present application, a training method for a large language model is provided, comprising: acquiring first knowledge information; inputting the first knowledge information into a large language model so that the large language model learns the first knowledge information to obtain a learned large language model; constructing a training sample set based on second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; and performing supervised training on the learned large language model through the training sample set to obtain a target large language model.
[0006] Furthermore, constructing a training sample set based on the second knowledge information includes: constructing multiple knowledge samples based on the second knowledge information; obtaining first reply information and second reply information corresponding to the knowledge samples; and constructing the training sample set based on the multiple knowledge samples, the first reply information and the second reply information.
[0007] Furthermore, obtaining the first reply information and the second reply information corresponding to the knowledge sample includes: for the target knowledge sample, constructing first prompt information based on the target knowledge sample, wherein the target knowledge sample is any one of the multiple knowledge samples, and the first prompt information is used to indicate that the target model outputs reply information consistent with the target knowledge sample; constructing second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs reply information inconsistent with the target knowledge sample; processing the first prompt information through the target model to obtain the first reply information; processing the second prompt information through the large language model to obtain the second reply information.
[0008] Furthermore, the first prompt information is processed through the target model to obtain the first reply information, including: processing the first prompt information through the target model to obtain multiple initial reply information; calculating the matching degree between the multiple initial reply information and the target knowledge sample to obtain a matching value; and screening the multiple initial reply information according to the matching value to obtain the first reply information.
[0009] Furthermore, the learned large language model is supervisedly trained through a training sample set to obtain a target large language model, including: processing the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; obtaining a target loss function based on the predicted reply information, the first reply information and the second reply information; and supervised training is performed on the learned large language model based on the target loss function to obtain the target large language model.
[0010] Furthermore, based on the predicted reply information, the first reply information and the second reply information, obtaining the target loss function includes: performing cross entropy calculation based on the predicted reply information and the first reply information to obtain a first loss function; performing cross entropy calculation based on the predicted reply information and the second reply information to obtain a second loss function; performing calculation based on the first loss function and the second loss function to obtain the target loss function.
[0011] Furthermore, calculating the degree of matching between the multiple initial reply information and the target knowledge sample to obtain a matching value includes: vectorizing the multiple initial reply information to obtain a first sparse vector and a first dense vector; vectorizing the target knowledge sample to obtain a second sparse vector and a second dense vector; and calculating based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain the matching value.
[0012] According to another aspect of an embodiment of the present application, an information processing method is also provided, including: obtaining question information input by a target object; processing the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained using any of the large language model training methods described above; and returning the target answer information to the target object.
[0013] According to another aspect of an embodiment of the present application, an information processing method is also provided, including: obtaining question information input by a client; processing the question information through a target large language model in a cloud server to obtain target answer information, wherein the target large language model is obtained using any of the large language model training methods described above; and returning the target answer information to the client.
[0014] According to another aspect of an embodiment of the present application, a language model training device is also provided, including: a first acquisition unit, used to acquire first knowledge information; an input unit, used to input the first knowledge information into a large language model, so that the large language model learns the first knowledge information to obtain a learned large language model; a construction unit, used to construct a training sample set based on second knowledge information, wherein the training sample set includes at least multiple knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; an adjustment unit, used to perform supervised training on the learned large language model through the training sample set to obtain a target large language model.
[0015] Furthermore, the construction unit includes: a first construction subunit, used to construct multiple knowledge samples based on the second knowledge information; an acquisition subunit, used to obtain first reply information and second reply information corresponding to the knowledge samples; and a second construction subunit, used to construct the training sample set based on the multiple knowledge samples, the first reply information and the second reply information.
[0016] Furthermore, the acquisition subunit includes: a first construction module, used to construct first prompt information for a target knowledge sample based on the target knowledge sample, wherein the target knowledge sample is any one of the multiple knowledge samples, and the first prompt information is used to indicate that the target model outputs reply information that is consistent with the target knowledge sample; a second construction module, used to construct second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs reply information that is inconsistent with the target knowledge sample; a first processing module, used to process the first prompt information through the target model to obtain the first reply information; and a second processing module, used to process the second prompt information through the large language model to obtain the second reply information.
[0017] Furthermore, the first processing module includes: a first processing sub-module, used to process the first prompt information through the target model to obtain multiple initial reply information; a calculation sub-module, used to calculate the matching degree between the multiple initial reply information and the target knowledge sample to obtain a matching value; and a screening sub-module, used to screen the multiple initial reply information according to the matching value to obtain the first reply information.
[0018] Furthermore, the adjustment unit includes: a first processing subunit, used to process the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; a second processing subunit, used to obtain a target loss function based on the predicted reply information, the first reply information and the second reply information; an adjustment subunit, used to perform supervised training on the learned large language model based on the target loss function to obtain the target large language model.
[0019] Furthermore, the second processing sub-unit includes: a first calculation module, used to perform cross entropy calculation based on the predicted reply information and the first reply information to obtain a first loss function; a second calculation module, used to perform cross entropy calculation based on the predicted reply information and the second reply information to obtain a second loss function; a third calculation module, used to perform calculation based on the first loss function and the second loss function to obtain the target loss function.
[0020] Furthermore, the calculation submodule includes: a first processing submodule, used to vectorize the multiple initial reply information to obtain a first sparse vector and a first dense vector; a second processing submodule, used to vectorize the target knowledge sample to obtain a second sparse vector and a second dense vector; a calculation submodule, used to calculate based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain the matching value.
[0021] According to another aspect of an embodiment of the present application, an information processing device is also provided, including: a second acquisition unit, used to acquire question information input by a target object; a processing unit, used to process the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained using any of the large language model training methods described above; and a return unit, used to return the target answer information to the target object.
[0022] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute any one of the above-mentioned large language model training methods or the information processing methods.
[0023] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, comprising a memory storing an executable program; and a processor for running the program, wherein when the program is running, the training method for a large language model described in any one of the above, or the information processing method, is executed.
[0024] According to another aspect of an embodiment of the present invention, a computer program product is also provided, which includes a stored computer program, and when the computer program is executed by a processor, any one of the above-mentioned methods for training a large language model or the above-mentioned information processing method is implemented.
[0025] In an embodiment of the present application, the following steps are adopted: obtaining first knowledge information; inputting the first knowledge information into a large language model so that the large language model learns second knowledge information to obtain a learned large language model; constructing a training sample set based on the second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; and supervising the learned large language model through the training sample set to obtain a target large language model, thereby solving the technical problem in the related art that the large language model is trained by inputting relevant knowledge into the large language model, resulting in relatively low accuracy of the large language model responses. In this solution, the first knowledge information is first input into the large language model so that the large model can preliminarily learn the second knowledge information. Then, based on the second knowledge information, multiple knowledge samples are constructed, as well as first reply information and second reply information corresponding to the knowledge samples. The first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample. The output of the learned large language model is intervened by the first reply information and the second reply information, so that the large language model can better identify more appropriate reply information, thereby achieving the technical effect of improving the accuracy of the large language model's replies. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0027] Figure 1 is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of the present application;
[0028] Figure 2 This is the process of the training method of the large language model provided in Example 1 of the present application Figure 1 ;
[0029] Figure 3 This is the process of the training method of the large language model provided in Example 1 of the present application Figure 2 ;
[0030] Figure 4 is a flowchart of an information processing method provided according to Embodiment 2 of the present application;
[0031] Figure 5 is a flowchart of an information processing method provided according to Embodiment 3 of the present application;
[0032] Figure 6 is a schematic diagram of a large language model training device provided according to Embodiment 4 of the present application;
[0033] Figure 7 is a schematic diagram of an information processing device provided according to Embodiment 5 of the present application;
[0034] Figure 8 This is a structural block diagram of an electronic device provided according to Embodiment 6 of the present application. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0038] Example 1
[0039] According to an embodiment of the present application, a training method for a large language model is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a training method for a large language model. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the processor set 102 may include a processor set, Figure 1 102a, 102b, ..., 102n are used to illustrate), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB, Universal Ser ia l Bus) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0041] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0042] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the training method of the large language model in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the training method of the large language model mentioned above. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0044] The display may be, for example, a touch screen liquid crystal display that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0045] Under the above operating environment, this application provides Figure 2 The training method of the large language model shown. Figure 2 : is a flow chart of a method for training a large language model according to the first embodiment of the present application. The training method comprises:
[0046] Step S201, obtaining first knowledge information.
[0047] Optionally, the relevant first knowledge information can be collected according to the required large language model. For example, if the user wants to implement e-commerce intelligent customer service through the large language model, the corresponding first knowledge information can be knowledge information related to e-commerce intelligent customer service, for example, relevant knowledge information of various products in the prior art. In an optional embodiment, the above-mentioned first knowledge information can be collected from the Internet or a related knowledge base by crawling.
[0048] Step S202: input the first knowledge information into the large language model so that the large language model learns the first knowledge information to obtain a learned large language model.
[0049] Optionally, the first knowledge information is input into a large language model so that the large language model learns the first knowledge information, thereby obtaining the learned large language model.
[0050] Step S203, constructing a training sample set based on the second knowledge information, wherein the training sample set includes at least multiple knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge samples, and the second reply information is inconsistent with the knowledge samples.
[0051] Optionally, after obtaining the above-mentioned learned large language model, a training sample set is constructed according to the second knowledge information. It should be noted that the second knowledge information may overlap with the first knowledge information or may not overlap. The second knowledge information is the information that the user prefers to be output by the large language model. For example, the user wants to use the large language model to achieve e-commerce intelligent customer service, and the second knowledge information is the knowledge information of the relevant products in the e-commerce platform corresponding to the user.
[0052] In an optional embodiment, the training sample set includes at least a plurality of knowledge samples, and the knowledge samples are obtained from the second knowledge information. For example, the knowledge of the corresponding product is extracted according to the identifier corresponding to the product in the e-commerce platform corresponding to the user to form a plurality of knowledge samples. The training sample set also includes the first reply information and the second reply information corresponding to the knowledge sample.
[0053] It should be noted that the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample, wherein the first reply information is consistent with the knowledge sample means that the content expressed by the first reply information is consistent with the meaning expressed by the corresponding knowledge sample, and the second reply information is inconsistent with the knowledge sample means that the content expressed by the second reply information is different from or contradictory to the meaning expressed by the corresponding knowledge sample.
[0054] For example, the knowledge corresponding to the knowledge sample is "a certain locator has an ultra-long standby function", and the corresponding first reply information can be "a certain locator does not need to be charged frequently"; the second reply information can be "a certain locator requires frequent charging and has limited battery life."
[0055] Step S204, performing supervised training on the learned large language model through the training sample set to obtain a target large language model.
[0056] Optionally, the learned large language model is supervised and trained using a training sample set so that the large language model can better identify more appropriate response information, thereby obtaining the above-mentioned target large language model.
[0057] To summarize, the first knowledge information is first input into the large language model so that the large model can preliminarily learn the second knowledge information, and then based on the second knowledge information, multiple knowledge samples are constructed, as well as the first reply information and the second reply information corresponding to the knowledge samples. The first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample. The output of the learned large language model is intervened by the first reply information and the second reply information, so that the large language model can better identify more appropriate reply information, thereby achieving the technical effect of improving the accuracy of the large language model's replies.
[0058] In order to improve the quality of the training sample set, in the training method of the large language model provided in Example 1 of the present application, constructing a training sample set based on the second knowledge information includes: constructing multiple knowledge samples based on the second knowledge information; obtaining first reply information and second reply information corresponding to the knowledge samples; and constructing a training sample set based on the multiple knowledge samples, the first reply information and the second reply information.
[0059] Optionally, the second knowledge information is split, for example, the second knowledge information is split according to the product identification to obtain multiple knowledge samples, or the second knowledge information is split according to the function of the product to obtain multiple knowledge samples, and the second knowledge information can be split according to user needs, which is not limited here.
[0060] After obtaining the above-mentioned multiple knowledge samples, the first reply information and the second reply information corresponding to the knowledge samples are obtained. In an optional embodiment, the first reply information and the second reply information can be generated manually, or the content expressed by the knowledge samples can be summarized through a language learning model, and then the summarized content is refined into the first reply information, and then the second reply information opposite to the first reply information is generated through the first reply information. Finally, a training sample set is constructed based on the multiple knowledge samples, the first reply information and the second reply information.
[0061] By constructing multiple knowledge samples, first reply information and second reply information to supervise and fine-tune the large language model, the large language model has the ability to distinguish knowledge priorities, thereby improving the accuracy of the subsequent reply information output by the large language model.
[0062] In order to improve the efficiency of obtaining the first reply information and the second reply information, in the training method of the large language model provided in Example 1 of the present application, obtaining the first reply information and the second reply information corresponding to the knowledge sample includes: for the target knowledge sample, constructing first prompt information based on the target knowledge sample, wherein the target knowledge sample is any one of a plurality of knowledge samples, and the first prompt information is used to indicate that the target model outputs reply information consistent with the target knowledge sample; constructing second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs reply information inconsistent with the target knowledge sample; processing the first prompt information through the target model to obtain the first reply information; processing the second prompt information through the large language model to obtain the second reply information.
[0063] Optionally, for any one of the multiple knowledge samples, i.e., the target knowledge sample, the target knowledge sample is used to construct a first prompt information, i.e., a first prompt. It should be noted that the first prompt is used to instruct the target model to output a reply information that is consistent with the target knowledge sample, and the target model may also be a large language model. The target knowledge sample is used to construct a second prompt information, i.e., a second prompt. The second prompt is used to instruct the target model to output a reply information that is inconsistent with the target knowledge sample.
[0064] For example, the first prompt could be:
[0065] “##Task: Please generate a product description that matches the following content. Do not add any irrelevant descriptions other than the following content;
[0066] The answer obtained based on the product attributes: The locator has an ultra-long standby function.
[0067] The answer obtained from the product knowledge base: The locator does not need to be charged for use.
[0068] ## Output requirements
[0069] Simply output the partially matching product descriptions.
[0070] Output results within 30 words."
[0071] For example, the second prompt could be:
[0072] “##Task: Please generate a product description that contradicts the following content. Do not add irrelevant descriptions other than the following content;
[0073] The answer obtained based on the product attributes: The locator has an ultra-long standby function.
[0074] The answer obtained from the product knowledge base: The locator does not need to be charged for use.
[0075] ## Output requirements
[0076] Simply output the partially contradictory product descriptions.
[0077] Output results within 30 words."
[0078] After obtaining the first prompt information and the second prompt information, the prompt information is input into the target model, the first prompt information is processed by the target model to obtain the first reply information, and the second prompt information is processed to obtain the second reply information.
[0079] The target model can be used to quickly and accurately obtain the first reply information and the second reply information corresponding to the knowledge sample, thereby improving the efficiency of fine-tuning the large language model.
[0080] In order to improve the accuracy of the first reply information, in the training method of the large language model provided in Example 1 of the present application, the first prompt information is processed by the target model to obtain the first reply information, including: processing the first prompt information by the target model to obtain multiple initial reply information; calculating the matching degree of the multiple initial reply information with the target knowledge sample to obtain the matching value; screening the multiple initial reply information according to the matching value to obtain the first reply information.
[0081] Optionally, the first prompt information is first processed by the target model to obtain multiple initial reply information. For example, the target model may output a preset number of initial reply information. Alternatively, when the first prompt information is processed by the target model, a preset number of multiple initial reply information are obtained according to the generation probability. For example, the reply information ranked in the first few places may be used as the above-mentioned multiple initial reply information.
[0082] After obtaining the above-mentioned multiple initial reply information, in order to more accurately select the first reply information, the matching degree between the multiple initial reply information and the target knowledge sample can be calculated, that is, the content consistency between the initial reply information and the target knowledge sample can be evaluated, and finally, the multiple initial reply information can be screened according to the matching value to obtain the first reply information. It should be noted that in order to avoid the deviation of the large language model, it is not necessary to select only the initial reply information with a high matching degree as the first reply information, and the first reply information can also be selected comprehensively based on the content of the reply information.
[0083] In an optional embodiment, the second reply information can be obtained in the same manner, that is, the first prompt information is first processed through the target model to obtain multiple initial reply information, the matching degree between the multiple initial reply information and the target knowledge sample is calculated, that is, the content consistency between the initial reply information and the target knowledge sample is evaluated, and finally, the multiple initial reply information is screened according to the matching value to obtain the second reply information. It should be noted that when selecting the first reply information according to the matching value, it is preferred to select the one with a higher matching degree, while when selecting the second reply information, it is preferred to select the one with a lower matching degree.
[0084] Multiple initial response information is output through the target model, and then the initial response information is screened according to the content consistency between the initial response information and the target knowledge sample to obtain the first response information, thereby improving the accuracy of determining the first response information.
[0085] How to supervise the learned large language model for training is also crucial. Therefore, in the large language model training method provided in Example 1 of the present application, the learned large language model is supervised and trained through a training sample set to obtain a target large language model, including: processing the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; obtaining a target loss function based on the predicted reply information, the first reply information and the second reply information; and supervising and training the learned large language model based on the target loss function to obtain a target large language model.
[0086] Optionally, the training sample set is input into the learned large language model, and then the knowledge samples in the training sample set are processed by the learned large language model. For example, a prompt is constructed according to the knowledge samples to prompt the learned large language model to output predicted reply information. Then, the target loss function is determined according to the predicted reply information, the first reply information, and the second reply information. For example, the target loss function is determined according to the similarity between the predicted reply information and the first reply information and the second reply information, that is, the similarity calculation function is determined as the target loss function. Finally, supervised training is performed on the learned large language model according to the target loss function to obtain a target large language model, that is, supervised training is performed on the learned large language model so that the large model is more inclined to output the first reply information.
[0087] Through the supervision and fine-tuning of multiple knowledge samples, the first reply information and the second reply information, the large language model has the ability to distinguish knowledge priorities, thereby improving the accuracy of the reply information output by the subsequent large language model.
[0088] In order to improve the accuracy of the target loss function, in the training method of the large language model provided in Example 1 of the present application, the target loss function is obtained based on the predicted reply information, the first reply information and the second reply information, including: performing cross entropy calculation based on the predicted reply information and the first reply information to obtain the first loss function; performing cross entropy calculation based on the predicted reply information and the second reply information to obtain the second loss function; and performing calculation based on the first loss function and the second loss function to obtain the target loss function.
[0089] Optionally, a first loss function is obtained by calculating the cross entropy between the predicted reply information and the first reply information, and a second loss function is obtained by calculating the cross entropy between the predicted reply information and the second reply information. Finally, the first loss function and the second loss function are determined as the target loss function. For example, the target loss function can be as follows:
[0090] L=L1+L2
[0091] Among them, L1 is the first loss function, L2 is the second loss function, and L is the target loss function.
[0092] In an optional embodiment, weights corresponding to the first loss function and the second loss function may also be set, and then the target loss function may be determined based on the weights, the first loss function, and the second loss function.
[0093] By iterating the large language model through the cross entropy between the predicted reply information and the first reply information and the cross entropy between the predicted reply information and the second reply information, the large language model can be enabled to have the ability to distinguish knowledge priorities, thereby improving the accuracy of the reply information subsequently output by the large language model.
[0094] In order to improve the accuracy of calculating the matching degree between multiple initial reply information and the target knowledge sample, in the training method of the large language model provided in Example 1 of the present application, the matching degree between multiple initial reply information and the target knowledge sample is calculated, and the matching value is obtained, including: vectorizing the multiple initial reply information to obtain a first sparse vector and a first dense vector; vectorizing the target knowledge sample to obtain a second sparse vector and a second dense vector; calculating based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain the matching value.
[0095] Optionally, multiple initial reply information are vectorized to obtain first sparse vectors and first dense vectors corresponding to the multiple initial reply information. For example, keyword extraction is performed on the initial reply information, and the corresponding keywords are converted into vectors to obtain the above-mentioned first sparse vector. For example, the initial reply information is decomposed to obtain the corresponding token sequence, and then a vector representation of each token is generated through a multi-layer Transformer encoder. Finally, the above-mentioned first dense vector is obtained based on the vector representation of each token. The second sparse vector and the second dense vector can be obtained in the same way.
[0096] After obtaining the first sparse vector, the first dense vector, the second sparse vector, and the second dense vector, a matching value is obtained by calculating based on the first sparse vector, the first dense vector, the second sparse vector, and the second dense vector. For example, the similarity between the first sparse vector and the second sparse vector is calculated, the similarity between the first dense vector and the second dense vector is calculated, and finally, the similarity is determined as the above matching value.
[0097] The sparse vector and dense vector respectively capture the keyword information and semantic information of the text, which can effectively improve the accuracy of calculating the matching degree.
[0098] In an optional embodiment, the Figure 3The flowchart shown implements the training of the large language model, specifically including: inputting general knowledge into the large language model so that the large language model learns to have the ability to answer prompts, and then constructing aligned samples and non-aligned samples through the knowledge information in the current platform. It should be noted that the aligned samples are the knowledge information and the first answer information corresponding to the knowledge information, and the non-aligned samples are the knowledge information and the second answer information corresponding to the knowledge information. Finally, the large language model is supervised and fine-tuned based on the aligned samples and the non-aligned samples to obtain the final large language model.
[0099] In the training method of the large language model provided in the first embodiment of the present application, first knowledge information is obtained; the first knowledge information is input into the large language model so that the large language model learns the first knowledge information to obtain the learned large language model; based on the second knowledge information, a training sample set is constructed, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; supervised training is performed on the learned large language model through the training sample set to obtain a target large language model, thereby solving the technical problem in the related art that the large language model is trained by inputting relevant knowledge into the large language model, resulting in relatively low accuracy of the large language model responses. In this solution, the first knowledge information is first input into the large language model so that the large model can preliminarily learn the knowledge information, and then based on the second knowledge information, multiple knowledge samples are constructed, as well as first reply information and second reply information corresponding to the knowledge samples. The first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample. The output of the learned large language model is intervened by the first reply information and the second reply information, so that the large language model can better identify more appropriate reply information, thereby achieving the technical effect of improving the accuracy of the large language model's replies.
[0100] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0102] Example 2
[0103] According to an embodiment of the present application, an information processing method is also provided, such as Figure 4 As shown, the information processing method includes:
[0104] Step S401, obtaining question information input by the target object.
[0105] Optionally, after the target large language model is obtained through the first embodiment, the target large language model can be deployed on a required platform, such as an e-commerce platform. The e-commerce platform receives question information input by the user (ie, the target object mentioned above).
[0106] Step S402, processing the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method in Embodiment 1;
[0107] Optionally, the acquired question information is input into a target large language model, and the question information is processed by the target large language model. For example, the target large language model divides the question text into words, phrases, or sub-words, and then converts them into key-value pairs, and finally infers the target answer information based on the key-value pairs.
[0108] Step S403, returning the target reply information to the target object.
[0109] In summary, the target large language model obtained by supervised fine-tuning through multiple knowledge samples, first reply information and second reply information has the ability to distinguish knowledge priorities, thereby improving the accuracy of the target reply information.
[0110] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0112] Example 3
[0113] According to an embodiment of the present application, an information processing method is also provided, such as Figure 5 As shown, the information processing method includes:
[0114] Step S501, obtaining question information input by the client;
[0115] Step S502: Process the question information in the cloud server by using the target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method in the first embodiment;
[0116] Step S503: Return the target response information to the client.
[0117] It should be noted that the method for processing information in the cloud server is the same as that in the second embodiment, and will not be repeated here.
[0118] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0120] Example 4
[0121] According to an embodiment of the present application, a large language model training device for implementing the above-mentioned large language model training method is also provided, such as Figure 6 As shown, the device includes: a first acquisition unit 601, an input unit 602, a construction unit 603 and an adjustment unit 604.
[0122] A first acquisition unit 601 is used to acquire first knowledge information;
[0123] An input unit 602 is used to input the first knowledge information into the large language model so that the large language model learns the first knowledge information to obtain a learned large language model;
[0124] A construction unit 603 is used to construct a training sample set according to the second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge samples, and the second reply information is inconsistent with the knowledge samples;
[0125] The adjustment unit 604 is used to perform supervised training on the learned large language model through the training sample set to obtain a target large language model.
[0126] In the training device of the large language model provided in the fourth embodiment of the present application, the first knowledge information is acquired by the first acquisition unit 601; the input unit 602 inputs the first knowledge information into the large language model so that the large language model learns the first knowledge information and obtains the learned large language model; the construction unit 603 constructs a training sample set based on the second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; the adjustment unit 604 performs supervised training on the learned large language model through the training sample set to obtain the target large language model, thereby solving the technical problem in the related art that the large language model is trained by inputting relevant knowledge into the large language model, resulting in relatively low accuracy of the large language model response. In this solution, the first knowledge information is first input into the large language model so that the large model can preliminarily learn the knowledge information, and then based on the second knowledge information, multiple knowledge samples are constructed, as well as first reply information and second reply information corresponding to the knowledge samples. The first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample. The output of the learned large language model is intervened by the first reply information and the second reply information, so that the large language model can better identify more appropriate reply information, thereby achieving the technical effect of improving the accuracy of the large language model's replies.
[0127] Optionally, in the training device of the large language model provided in Example 4 of the present application, the construction unit 603 includes: a first construction subunit, used to construct multiple knowledge samples based on the second knowledge information; an acquisition subunit, used to obtain first reply information and second reply information corresponding to the knowledge samples; and a second construction subunit, used to construct a training sample set based on multiple knowledge samples, the first reply information and the second reply information.
[0128] Optionally, in the training device of the large language model provided in Example 4 of the present application, the acquisition subunit includes: a first construction module, used to construct first prompt information for a target knowledge sample based on the target knowledge sample, wherein the target knowledge sample is any one of a plurality of knowledge samples, and the first prompt information is used to indicate that the target model outputs reply information that is consistent with the target knowledge sample; a second construction module, used to construct second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs reply information that is inconsistent with the target knowledge sample; a first processing module, used to process the first prompt information through the target model to obtain first reply information; and a second processing module, used to process the second prompt information through the large language model to obtain second reply information.
[0129] Optionally, in the training device of the large language model provided in Example 4 of the present application, the first processing module includes: a first processing sub-module, used to process the first prompt information through the target model to obtain multiple initial reply information; a calculation sub-module, used to calculate the matching degree between the multiple initial reply information and the target knowledge sample to obtain a matching value; a screening sub-module, used to screen the multiple initial reply information according to the matching value to obtain the first reply information.
[0130] Optionally, in the training device of the large language model provided in Example 4 of the present application, the adjustment unit 604 includes: a first processing subunit, used to process the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; a second processing subunit, used to obtain a target loss function based on the predicted reply information, the first reply information and the second reply information; an adjustment subunit, used to perform supervised training on the learned large language model based on the target loss function to obtain a target large language model.
[0131] Optionally, in the training device of the large language model provided in Example 4 of the present application, the second processing subunit includes: a first calculation module, used to perform cross entropy calculation based on the predicted reply information and the first reply information to obtain a first loss function; a second calculation module, used to perform cross entropy calculation based on the predicted reply information and the second reply information to obtain a second loss function; a third calculation module, used to perform calculations based on the first loss function and the second loss function to obtain a target loss function.
[0132] Optionally, in the training device of the large language model provided in Example 4 of the present application, the calculation sub-module includes: a first processing sub-module, used to vectorize multiple initial reply information to obtain a first sparse vector and a first dense vector; a second processing sub-module, used to vectorize the target knowledge sample to obtain a second sparse vector and a second dense vector; a calculation sub-module, used to calculate based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain a matching value.
[0133] It should be noted that the first acquisition unit 601, input unit 602, construction unit 603 and adjustment unit 604 described above correspond to steps S201 to S204 in Embodiment 1, and the examples and application scenarios implemented by the four units and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1 described above. It should be noted that the above modules as part of the device can be run in the computer terminal 10 provided in Embodiment 1.
[0134] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0135] Example 5
[0136] According to an embodiment of the present application, there is also provided an information processing device for implementing the information processing method, such as Figure 7 As shown, the device includes: a second acquisition unit 701, a processing unit 702 and a returning unit 703.
[0137] The second acquisition unit 701 is used to acquire question information input by the target object;
[0138] A processing unit 702 is used to process the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method in the first embodiment;
[0139] The returning unit 703 is used to return the target reply information to the target object.
[0140] It should be noted that the second acquisition unit 701, the processing unit 702 and the returning unit 703 correspond to steps S401 to S403 in the second embodiment, and the three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the first embodiment. It should be noted that the above modules as part of the device can be run in the computer terminal 10 provided in the first embodiment.
[0141] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 2 as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 2.
[0142] Example 6
[0143] The embodiment of the present application may provide an electronic device, which may be any electronic device in an electronic device terminal group. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.
[0144] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0145] In this embodiment, the above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: obtaining first knowledge information; inputting the first knowledge information into the large language model so that the large language model learns the first knowledge information to obtain the learned large language model; constructing a training sample set based on the second knowledge information, wherein the training sample set includes at least multiple knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; and performing supervised training on the learned large language model through the training sample set to obtain the target large language model.
[0146] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: constructing a training sample set based on the second knowledge information includes: constructing multiple knowledge samples based on the second knowledge information; obtaining the first reply information and the second reply information corresponding to the knowledge samples; constructing a training sample set based on the multiple knowledge samples, the first reply information and the second reply information.
[0147] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: obtaining the first reply information and the second reply information corresponding to the knowledge sample includes: for the target knowledge sample, constructing the first prompt information based on the target knowledge sample, wherein the target knowledge sample is any one of the multiple knowledge samples, and the first prompt information is used to indicate that the target model outputs the reply information consistent with the target knowledge sample; constructing the second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs the reply information inconsistent with the target knowledge sample; processing the first prompt information through the target model to obtain the first reply information; processing the second prompt information through the large language model to obtain the second reply information.
[0148] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: processing the first prompt information through the target model to obtain the first reply information includes: processing the first prompt information through the target model to obtain multiple initial reply information; calculating the matching degree between the multiple initial reply information and the target knowledge sample to obtain the matching value; screening the multiple initial reply information according to the matching value to obtain the first reply information.
[0149] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: supervised training is performed on the learned large language model through the training sample set to obtain the target large language model, including: processing the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; obtaining the target loss function based on the predicted reply information, the first reply information and the second reply information; supervised training is performed on the learned large language model based on the target loss function to obtain the target large language model.
[0150] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: obtaining the target loss function based on the predicted reply information, the first reply information and the second reply information includes: performing cross entropy calculation based on the predicted reply information and the first reply information to obtain the first loss function; performing cross entropy calculation based on the predicted reply information and the second reply information to obtain the second loss function; calculating based on the first loss function and the second loss function to obtain the target loss function.
[0151] The above-mentioned electronic device can execute the program code of the following steps in the training method of the large language model: calculating the matching degree between multiple initial response information and the target knowledge sample to obtain the matching value including: vectorizing the multiple initial response information to obtain a first sparse vector and a first dense vector; vectorizing the target knowledge sample to obtain a second sparse vector and a second dense vector; calculating based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain the matching value.
[0152] The above-mentioned electronic device can execute the program code of the following steps in the information processing method: obtaining question information input by the target object; processing the question information through the target large language model to obtain target answer information, wherein the target large language model is obtained using any of the above-mentioned large language model training methods; and returning the target answer information to the target object.
[0153] Optionally, Figure 8 is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 8 As shown, the electronic device 80 may include: one or more ( Figure 8 Only one is shown in the figure) processor 802 and memory 804. The electronic device 80 may also include a storage controller, through which the memory 804 is controlled and managed; the electronic device 80 may also include a peripheral interface, through which the radio frequency module, audio module and display screen are connected.
[0154] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method and device of the large language model in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned training method of the large language model. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the electronic device 80 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0155] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain the first knowledge information; input the first knowledge information into the large language model so that the large language model learns the first knowledge information to obtain the learned large language model; construct a training sample set based on the second knowledge information, wherein the training sample set includes at least multiple knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; supervise the learned large language model through the training sample set to obtain the target large language model.
[0156] Optionally, the processor may also execute program code of the following steps: constructing a training sample set based on the second knowledge information, including: constructing multiple knowledge samples based on the second knowledge information; obtaining first reply information and second reply information corresponding to the knowledge samples; constructing a training sample set based on multiple knowledge samples, the first reply information and the second reply information.
[0157] Optionally, the processor may also execute program code for the following steps: obtaining first reply information and second reply information corresponding to the knowledge sample includes: for a target knowledge sample, constructing first prompt information based on the target knowledge sample, wherein the target knowledge sample is any one of a plurality of knowledge samples, and the first prompt information is used to indicate that the target model outputs reply information that is consistent with the target knowledge sample; constructing second prompt information based on the target knowledge sample, wherein the second prompt information is used to indicate that the target model outputs reply information that is inconsistent with the target knowledge sample; processing the first prompt information through the target model to obtain first reply information; and processing the second prompt information through the large language model to obtain second reply information.
[0158] Optionally, the processor may also execute program code of the following steps: processing the first prompt information through the target model to obtain first reply information, including: processing the first prompt information through the target model to obtain multiple initial reply information; calculating the matching degree between the multiple initial reply information and the target knowledge sample to obtain a matching value; screening the multiple initial reply information according to the matching value to obtain the first reply information.
[0159] Optionally, the processor may also execute the program code of the following steps: performing supervised training on the learned large language model through a training sample set to obtain a target large language model, including: processing the knowledge samples in the training sample set through the learned large language model to obtain predicted reply information; obtaining a target loss function based on the predicted reply information, the first reply information and the second reply information; performing supervised training on the learned large language model based on the target loss function to obtain a target large language model.
[0160] Optionally, the processor may also execute program code for the following steps: obtaining a target loss function based on the predicted reply information, the first reply information and the second reply information, including: performing cross entropy calculation based on the predicted reply information and the first reply information to obtain a first loss function; performing cross entropy calculation based on the predicted reply information and the second reply information to obtain a second loss function; performing calculation based on the first loss function and the second loss function to obtain a target loss function.
[0161] Optionally, the processor may also execute program code for the following steps: calculating the degree of matching between multiple initial reply information and the target knowledge sample to obtain a matching value, including: vectorizing the multiple initial reply information to obtain a first sparse vector and a first dense vector; vectorizing the target knowledge sample to obtain a second sparse vector and a second dense vector; and calculating based on the first sparse vector, the first dense vector, the second sparse vector and the second dense vector to obtain a matching value.
[0162] Optionally, the processor may also execute the program code of the following steps: obtaining question information input by the target object; processing the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained using any of the large language model training methods mentioned above; and returning the target answer information to the target object.
[0163] It can be understood by those skilled in the art that Figure 8 The structure shown is for illustration only, and the electronic device 80 may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 8The structure of the electronic device is not limited. Figure 8 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 8 Different configurations are shown.
[0164] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0165] Example 7
[0166] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product can be used to store the program code executed by the large language model training method provided in the first embodiment.
[0167] Optionally, in this embodiment, the computer program product may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0168] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0169] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0170] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0171] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0172] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or CD-ROM and other media that can store program codes.
[0174] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for training a large language model, characterized in that: include: Acquire first knowledge information; Inputting the first knowledge information into a large language model so that the large language model learns the first knowledge information to obtain a learned large language model; Constructing a training sample set according to the second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; The learned large language model is supervisedly trained using the training sample set to obtain a target large language model.
2. The method according to claim 1, characterized in that According to the second knowledge information, constructing a training sample set includes: constructing a plurality of knowledge samples according to the second knowledge information; Obtaining first reply information and second reply information corresponding to the knowledge sample; The training sample set is constructed according to the multiple knowledge samples, the first reply information and the second reply information.
3. The method according to claim 2, characterized in that Acquiring the first reply information and the second reply information corresponding to the knowledge sample includes: For a target knowledge sample, constructing first prompt information according to the target knowledge sample, wherein the target knowledge sample is any one of the multiple knowledge samples, and the first prompt information is used to instruct the target model to output answer information consistent with the target knowledge sample; Constructing second prompt information according to the target knowledge sample, wherein the second prompt information is used to instruct the target model to output answer information that is contrary to the target knowledge sample; Processing the first prompt information through the target model to obtain the first reply information; The second prompt information is processed by the large language model to obtain the second reply information.
4. The method according to claim 3, characterized in that Processing the first prompt information by the target model to obtain the first reply information includes: Processing the first prompt information through the target model to obtain a plurality of initial response information; Calculating the matching degree between the plurality of initial reply information and the target knowledge sample to obtain a matching value; The multiple initial reply information are screened according to the matching value to obtain the first reply information.
5. The method according to claim 1, characterized in that The learned large language model is supervised and trained by the training sample set to obtain a target large language model, including: Processing the knowledge samples in the training sample set by the learned large language model to obtain predicted answer information; Obtaining a target loss function according to the predicted reply information, the first reply information and the second reply information; The learned large language model is supervised and trained according to the target loss function to obtain the target large language model.
6. The method according to claim 5, characterized in that Obtaining a target loss function based on the predicted reply information, the first reply information, and the second reply information includes: Performing cross entropy calculation based on the predicted reply information and the first reply information to obtain a first loss function; Performing cross entropy calculation based on the predicted reply information and the second reply information to obtain a second loss function; The target loss function is obtained by performing calculations based on the first loss function and the second loss function.
7. The method according to claim 4, characterized in that Calculating the matching degree between the plurality of initial reply information and the target knowledge sample to obtain the matching value includes: Performing vectorization processing on the multiple initial reply information to obtain a first sparse vector and a first dense vector; Performing vectorization processing on the target knowledge sample to obtain a second sparse vector and a second dense vector; The matching value is obtained by performing calculation according to the first sparse vector, the first dense vector, the second sparse vector and the second dense vector.
8. An information processing method, characterized in that: include: Get the question information input by the target object; Processing the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method described in any one of claims 1 to 7; The target reply information is returned to the target object.
9. An information processing method, characterized in that: include: Get the question information entered by the client; In the cloud server, the question information is processed by a target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method described in any one of claims 1 to 7; The target reply information is returned to the client.
10. A large language model training device, characterized in that: include: A first acquisition unit, used to acquire first knowledge information; An input unit, used for inputting the first knowledge information into a large language model, so that the large language model learns the first knowledge information, and obtains a learned large language model; A construction unit, configured to construct a training sample set according to the second knowledge information, wherein the training sample set includes at least a plurality of knowledge samples, and first reply information and second reply information corresponding to the knowledge samples, wherein the first reply information is consistent with the knowledge sample, and the second reply information is inconsistent with the knowledge sample; The adjustment unit is used to perform supervised training on the learned large language model through the training sample set to obtain a target large language model.
11. An information processing device, characterized in that: include: A second acquisition unit is used to acquire question information input by a target object; a processing unit, configured to process the question information through a target large language model to obtain target answer information, wherein the target large language model is obtained by using the large language model training method according to any one of claims 1 to 7; A returning unit is used to return the target reply information to the target object.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is run, the device where the storage medium is located is controlled to execute the large language model training method described in any one of claims 1 to 7, or the information processing method described in claim 8.
13. An electronic device, characterized in that: include: A memory storing an executable program; A processor, used to run the program, wherein when the program is run, the training method of a large language model described in any one of claims 1 to 7, or the information processing method described in claim 8 is executed.
14. A computer program product, characterized in that The computer program product includes a stored computer program, and when the computer program is executed by a processor, the large language model training method described in any one of claims 1 to 7, or the information processing method described in claim 8 is implemented.