Large language model training method, correlation determination method and related device
By combining relevance and general language tasks to train a large language model, and using samples of different quality and task instructions to optimize the training process, the accuracy and interpretability problems of existing models in search relevance discrimination are solved, achieving higher discrimination accuracy and robustness.
Patent Information
- Application Number
- CN202411775093.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-04
AI Technical Summary
Existing search relevance discrimination models are limited in application scenarios and lack interpretability, and the discrimination accuracy of large language models trained solely on relevance tasks still has room for improvement.
By combining relevance tasks and general language tasks to train a large language model, alternate training is performed using sample sets of different quality, task instructions and target characters are introduced to improve the model's understanding and discrimination capabilities, and a loss function is used to optimize the model training process.
It improves the accuracy and interpretability of search relevance judgment and enhances the model's ability to transfer and robustness between different tasks.
Smart Images

Figure CN119782524B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as intelligent search, intelligent recommendation, deep learning, and large models, and can be used in application scenarios such as generative search, intelligent document editing, intelligent assistants, virtual assistants, and intelligent e-commerce. Background Art
[0002] Search relevance, in the field of information retrieval, measures the degree of correlation between user queries and search results, thereby providing search results that best match user intent. Search relevance models aim to quantify the degree of match between search results and user intent and filter and sort search results based on relevance probability scores.
[0003] Search relevance determination is a crucial part of search engines, which can directly affect the quality of search results and user experience. Summary of the Invention
[0004] The present disclosure provides a large language model training method, a relevance determination method, and related devices.
[0005] According to one aspect of the present disclosure, a method for training a large language model is provided, comprising:
[0006] Obtaining a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a relevance task, and the relevance task is used to determine search relevance; the task type of the second type of task sample is a general language task;
[0007] The task instructions prompt the large language model to perform the task type, so as to train the large language model based on samples of different task types.
[0008] According to one aspect of the present disclosure, a method for determining relevance is provided, comprising:
[0009] Get user queries and search results;
[0010] Constructing a prompt word based on the user query and the search result; the prompt word includes a task instruction that prompts the large language model to perform a relevance task;
[0011] The prompt word is input into the large language model to obtain a prediction result of the large language model for the search relevance between the user query and the search result, and a generation reason.
[0012] According to another aspect of the present disclosure, a large language model training apparatus is provided, comprising:
[0013] A first acquisition module is configured to acquire a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a relevance task, which is used to determine search relevance; and the task type of the second type of task sample is a general language task;
[0014] The training module is used to prompt the large language model with the task type to be performed based on the task instructions, so as to train the large language model based on samples of different task types.
[0015] According to another aspect of the present disclosure, a correlation determination device is provided, comprising:
[0016] The second acquisition module is used to obtain user queries and search results;
[0017] A construction module, configured to construct a prompt word based on the user query and the search results; the prompt word includes a task instruction prompting the large language model to perform a relevance task;
[0018] A processing module is configured to input the prompt word into the large language model, obtain a prediction result of the large language model for the search relevance between the user query and the search result, and generate a reason.
[0019] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0020] at least one processor; and
[0021] a memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0023] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute any method according to the embodiments of the present disclosure.
[0024] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0025] In the disclosed embodiment, by learning general language tasks, the general capabilities of the large language model in content understanding, generalization, reasoning, etc. can be fully explored, and then these general capabilities can be transferred to the search relevance judgment task, which can improve the accuracy of search relevance judgment.
[0026] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0028] Figure 1 1 is a flow chart of a method for training a large language model according to an embodiment of the present disclosure;
[0029] Figure 2 is a schematic diagram of a process for determining dirty samples in a first sample set according to an embodiment of the present disclosure;
[0030] Figure 3 is a schematic diagram of search results provided according to an embodiment of the present disclosure;
[0031] Figure 4 is a schematic diagram of a process for determining clean samples in the second sample set according to an embodiment of the present disclosure;
[0032] Figure 5 1 is a flow chart of training a large language model according to an embodiment of the present disclosure;
[0033] Figure 6 1 is a schematic diagram of the architecture of a large language model provided according to an embodiment of the present disclosure;
[0034] Figure 7 is a schematic diagram of prompt information of a large language model provided according to an embodiment of the present disclosure;
[0035] Figure 8 is a flowchart of a correlation determination method provided according to an embodiment of the present disclosure;
[0036] Figure 9 2 is a schematic diagram of a large language model training device according to an embodiment of the present disclosure;
[0037] Figure 10 is a structural diagram of a correlation determination device provided according to an embodiment of the present disclosure;
[0038] Figure 11 It is a block diagram of an electronic device used to implement the large language model training method and / or relevance determination method of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0039] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0040] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0041] In search engines, information relevance discriminant networks can be used to identify the search relevance between user queries and search results. For example, the BERT model (Bidirectional Encoder Representation from Transformers, a bidirectional pre-trained language model based on Transformers) can be used to map query terms and documents into vectors, and then calculate the inner product to obtain a relevance score.
[0042] However, the search relevance obtained based on traditional information relevance discrimination networks has limited application scenarios and lacks interpretability. Therefore, a large language model can be used to analyze the correlation between information and provide reasons to increase interpretability.
[0043] Large language models (LLMs) are a specific type of large model designed specifically for processing text data. These models are neural network-based natural language processing models that can be used to generate, understand, and process text data. Large language models can have tens of billions of parameters, generate high-quality text, and are used for a variety of natural language processing tasks such as question answering, text generation, and dialogue systems.
[0044] Large language models have excellent reasoning capabilities and the ability to learn from small samples. Large language models can be used to understand large models or train on a large number of samples. Based on these large language models, accurate semantic understanding can be achieved.
[0045] However, if we simply use the correlation task for determining search relevance to train and optimize the large language model, its discrimination accuracy still needs to be improved. Therefore, a training method for a large language model is proposed in the embodiment of the present disclosure. By adding training for general language tasks, this method can fully tap the ability of the large language model in general language tasks to increase the discrimination accuracy of correlation tasks. Figure 1 FIG. 1 is a flow chart of a method for training a large language model according to an embodiment of the present disclosure, including the following contents:
[0046] S101, obtaining a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a relevance task, which is used to determine search relevance; the task type of the second type of task sample is a general language task;
[0047] Among them, in the field of information retrieval, relevance tasks are used to measure the degree of relevance between user queries and search results, aiming to provide search results that match user intentions.
[0048] General language tasks are natural language processing tasks that are not specific to a particular field or industry. They typically involve broad understanding and generation of language. For example, general language tasks can include natural language generation tasks such as semantic understanding, paragraph summarization, information reasoning, and text rewriting.
[0049] S102 , prompting the large language model with the task type to be performed based on the task instruction, so as to train the large language model based on samples of different task types.
[0050] Among them, by learning correlation tasks and general language tasks, the model capabilities learned by the large language model based on the general language task are transferred to the correlation task, so as to make full use of the general language task analysis capabilities of the large language model to improve the processing accuracy of the correlation task.
[0051] The task instruction may be an instruction instruction, which is used to distinguish different tasks.
[0052] In the disclosed embodiment, based on different task instructions, the large language model can learn the ability of general language tasks while learning to distinguish search relevance. Through the learning of general language tasks, the general capabilities of the large language model in content understanding, generalization, reasoning, etc. can be fully explored, and then these general capabilities can be transferred to the intent understanding, key point extraction, and cause summarization capabilities of the search relevance discrimination task. Ultimately, the accuracy of search relevance discrimination can be further improved. At the same time, using general language tasks to train the large language model can enable the large language model to have the ability to handle different tasks, so as to achieve the ability transfer between multiple tasks.
[0053] In some embodiments, the first type of task samples includes a first sample set and a second sample set; the sample quality of the first sample set is lower than the sample quality of the second sample set, and the first sample set includes dirty samples.
[0054] The samples in the first sample set may be low-quality samples, such as dirty samples. Dirty samples generally refer to data that does not meet requirements and cannot be directly analyzed. In common data mining work, dirty samples refer to data examples that are incorrectly labeled.
[0055] The samples in the second sample set can be high-quality samples, exemplified by clean samples. Clean samples refer to data instances in the dataset that are not mislabeled. They are crucial for model training because the accuracy and robustness of the model largely depend on the quality of the training data.
[0056] During specific training, the large language model can be trained alternately using the first sample set and the second sample set. For example, the first sample set can be used in the i-th batch, and the second sample set can be used in the (i+1)th batch. Furthermore, the first sample set and the second sample set can also be included in the same batch, which is not limited in the present embodiment.
[0057] In the disclosed embodiment, the large language model is trained using samples of different qualities, so that the large language model can learn the patterns of clean samples and dirty samples. Even in a low-quality sample environment, inference learning can be performed to obtain accurate search relevance judgment results. Better knowledge can also be learned from low-quality samples to further improve the robustness of the large language model.
[0058] In some embodiments, dirty samples in the first sample set may be determined based on the following method: Figure 2 As shown, including:
[0059] S201: Obtain a first sample search request, a first sample search result recommended for the first sample search request, and a network address of the first sample search result.
[0060] A search request, or user query, typically refers to a query a user enters into a search engine, database, directory, or other information retrieval system to find specific information or data. This request can be a simple keyword or a complex query statement, depending on the user's needs and the capabilities of the search system.
[0061] During implementation, the first sample search request serves as the training sample search request. The recommendation engine can filter the search results for the first sample search request as the first sample search results. Typically, the first sample search results correspond to a network address, such as a landing page, which is required when constructing dirty samples.
[0062] S202: Compress the effective information of the first sample search result to obtain the first core content of the first sample search result.
[0063] Among them, effective information compression refers to reducing data through specific algorithms and technologies while retaining the core information and value of the data.
[0064] The first core content is used to represent the important and core content of the first sample search result.
[0065] During implementation, when the first sample search result is text, a large language model can be used to extract keywords to obtain the first core content.
[0066] If the first sample search results include an image, optical character recognition (OCR) can be used to identify the image and obtain text information. The large language model can then extract keywords from the text information to obtain the first core content. Furthermore, a multimodal large model can be used to understand the content of the image to generate text content that can describe the image, and then the first core content can be extracted from this text content.
[0067] It is understandable that corresponding core content can be extracted from information of different modalities to construct the first core content of the first sample search results. The embodiment of the present disclosure does not limit how to obtain the first core content.
[0068] S203: If labeling is required, obtain the second sample search result from the network address of the first sample search result.
[0069] The first sample search result and the second sample search result may be the same or different. For example, in a search advertising scenario, Figure 3 As shown, the first sample search request recommends the first sample search result at time point A. When constructing the training sample, when obtaining the first sample search result, the first core content is generated based on the first sample search result of its network address at time point A. This is then stored for subsequent annotation.
[0070] There's a time difference between time point B and time point A. During this time, the content of the network address may or may not have changed. At this point, the annotator clicks on the network address to obtain a second sample search result. The second sample search result may or may not differ from the first sample search result.
[0071] For example, for web pages influenced by current events, the content update cycle is short, which may cause the content of the same network address to change. In this case, the first sample search result and the second sample search result may be different. In this case, the annotator will compare the second sample search result with the first sample search request to determine the search relevance and thus obtain the first label.
[0072] S204: Labeling is performed based on the first sample search request and the second sample search results to obtain a dirty sample consisting of the first sample search request, the first core content, and the first label.
[0073] Continuing with the above example, the label obtained by annotating the second sample search result may be different from the first core content. Therefore, in this case, the labeling may not be reasonable, thereby obtaining dirty data.
[0074] In the disclosed embodiment, dirty data obtained in this manner is used to make the quality of samples different, thereby providing powerful sample data for subsequent large language models to learn.
[0075] In some embodiments, the clean samples in the second sample set may be determined based on the following method: Figure 4 As shown, including:
[0076] S401: Obtain a second sample search request and a snapshot of a third sample search result recommended for the second sample search request.
[0077] A snapshot is a record of the state of data storage at a certain moment.
[0078] During implementation, the snapshot of the third sample search result is the third sample search result obtained by searching based on the second sample search request, and the third sample search result and the current time point are stored in the form of a snapshot.
[0079] S402: Compress the effective information of the third sample search result to obtain the second core content of the third sample search result.
[0080] The specific method for performing effective information compression has been described above and will not be repeated here in the embodiment of the present disclosure.
[0081] S403: If labeling is required, labeling is performed based on the snapshot of the second sample search request and the third sample search result to obtain a clean sample consisting of the second sample search request, the second core content, and the second label.
[0082] During implementation, since a snapshot of the third sample search results is stored, the annotator can annotate the snapshot of the third sample search results to obtain a label. This approach can avoid the aforementioned problem of time differences causing the label and core content to differ, thereby improving the quality of the training sample and obtaining a clean sample.
[0083] In some embodiments, when the first type of task samples includes a first sample set and a second sample set, different task instructions are used for different sample sets to improve learning quality. During implementation, the task instructions for the first sample set not only prompt the large language model to learn the relevant task, but also indicate that the samples being learned by the large language model are dirty samples.
[0084] Similarly, the task instructions of the second sample set not only prompt the large language model to learn the relevant task, but also prompt the large language model to learn clean samples.
[0085] For example, the task instruction of the first sample set can be a, and the task instruction of the second sample set can be b. During training, they are trained with sample instructions corresponding to other common second-category task samples for learning of the large language model.
[0086] Specifically, the task instruction for the correlation task can be set to R, the task instruction for the first sample set can be Ra, the task instruction for the second sample set can be Rb, and the task instruction for the universal language task can be set to P, with the task instruction for universal language task 1 being P1, the task instruction for universal language task 2 being P2, and so on. It should be noted that the specific form of the task instruction corresponding to each task is not limited in the embodiments of the present disclosure; any task instruction that can be recognized by the large language model is applicable to the embodiments of the present disclosure.
[0087] In the disclosed embodiment, task instructions are used to distinguish low-quality samples from high-quality samples, so that the large language model can learn the patterns and knowledge of samples of different qualities during the inference process, thereby improving the adaptability of the large language model and further improving the accuracy of search relevance.
[0088] After obtaining the samples, the large language model can be trained. In some embodiments, the large language model is trained based on the first type of task samples, such as Figure 5 As shown, it can be implemented as:
[0089] S501, constructing the first type of prompt information based on the first type of task samples; the first type of prompt information includes the third sample search request in the first type of task samples, the third core content corresponding to the third sample search request, the search intention of the third sample search request, and the task instructions corresponding to the related tasks.
[0090] The third sample search request, the third core content, the search intent, and the task instructions of the related tasks can be spliced together according to a preset format to obtain the first type of prompt information. The embodiment of this disclosure does not limit the form of the first type of prompt information, as long as it meets the construction requirements of the prompt project of the large language model.
[0091] S502: Training a large language model based on the first type of prompt information.
[0092] During implementation, the first type of prompt information is input into the large language model to obtain a probability score and a generated reason. This is then compared with the label in the input information to obtain a loss value. If the convergence conditions are met, the large language model training is completed. If not, the model parameters are adjusted until the convergence conditions are met.
[0093] In the embodiment of the present disclosure, a search intent that can explain the user's intention is introduced into the relevance task, so that the large language model can more accurately understand the user's search intent, thereby optimizing the discrimination quality of the search relevance. For example, in a search scenario (such as advertising), the user's query words are often affected by current hot topics, and the use of a large language model as a discrimination method is subject to the influence of the training corpus and has a certain lag constraint, which makes it difficult for the model to understand new words and hot words that have not appeared, thereby affecting the discrimination. Therefore, by introducing a subdomain that can explain the user's intention, the large language model can more accurately understand the user's search intent, thereby optimizing the discrimination quality of the search relevance.
[0094] Similarly, for general language tasks, it can be implemented as follows:
[0095] Step A1: construct a second type of prompt information based on the second type of task samples; the second type of prompt information includes the second type of task samples and task instructions corresponding to the general language task.
[0096] During implementation, the second type of prompt information is input into the large language model to obtain the result generated by the generative network of the large language model, which can be compared with the generated label to optimize the model parameters.
[0097] In summary, the large language model provided by the embodiment of the present disclosure is exemplified as follows: Figure 6The system may include a large language processing module (constructed by LLM), which performs reasoning analysis on the input prompt information and outputs a discriminant vector and a generative vector. The discriminant vector is input into the discriminant network to obtain a discriminant result of search relevance. The generative vector is input into the generative network, which outputs the generation reason to improve interpretability.
[0098] During implementation, in order to better learn the correlation between user queries and search results, it is necessary to fully integrate the features of both to obtain a high-quality discriminant vector. Therefore, in the disclosed embodiment, for the first type of task sample, the prompt information of the first type of task sample also includes a first target character; the first target character is inferred by the large language model to obtain a discriminant vector; this discriminant vector is used to input into the discriminant network of the large language model to determine the search relevance between the user query and the search results.
[0099] Among them, the first target character can be exemplarily "[CLS]", which can be understood as a special token. The character obtains a discriminant vector through reasoning of the large language model. The discriminant vector can be used to prompt the output of the search relevance score of the third sample search request and the corresponding third core content.
[0100] In addition, the first prompt information may also include a generation reason, which can be marked in the form of "[gMASK]", which can be understood as a special token. The generation reason is used to prompt the output of the descriptive information of the relevance reason between the third sample search request and the corresponding third core content.
[0101] In the embodiment of the present disclosure, using the target character as the content of the first type of prompt information can enable the user query and the search results to be fully integrated and reflected in the discriminant vector corresponding to the target character, so as to provide high-quality feature information for the discriminant network, thereby improving the accuracy of the search relevance score output by the discriminant network.
[0102] For example, for the relevance task, it is necessary to output the relevance score and the reason for the score. The first type of prompt information for the clean data in the first type of task sample can be:
[0103] Input: Task instruction: Relevance Task - Clean Data, third sample search request: ```Cough with phlegm```, search intent: ```The user may be looking for effective medicine for cough and phlegm```. Third core content: ```Wuji Bufei Pills, Xifeng Pharmaceutical, Bufei Yiqi Zhike```. The first target character of the third sample search request and the corresponding search result: [CLS], generation reason: [gMASK].
[0104] -Output-Probability score: 0~1
[0105] -Output-Generation reason: The third sample search request: "cough and phlegm" is a respiratory disease. The network address of the search result corresponding to the third sample search request is displayed as the [***] Bufei Pills product page, which can cure cough and phlegm and meet general needs.
[0106] For general language tasks, such as text summarization, it is necessary to output summary results. For example, the second type of prompt information can be:
[0107] Input: "Ma Liang, the Magic Brush" is a story in Chinese folklore. Ma Liang is a poor but kind little boy who dreams of owning a magic brush that can draw pictures of life. One day, he really gets such a brush, and whatever he draws becomes reality. When the greedy emperor learns of this, he wants to use Ma Liang's magic brush to satisfy his own desires. But Ma Liang is resourceful and brave. He uses his wisdom and magic brush to fight against the emperor, and ultimately protects the people and becomes a folk hero." Task instructions for the general language task: Please generate a summary for the above content.
[0108] Output - Generate Summary: Ma Liang, a kindhearted poor boy, obtains a magic brush that makes his drawings come true. He uses his wisdom to fight against the greedy emperor, protect the people, and become a hero.
[0109] In some embodiments, for the first type of task samples, the loss used to train the large language model includes a first sub-loss and a second sub-loss;
[0110] The first sub-loss is a discriminant loss determined based on the first type of task samples; the discriminant loss is used to measure the gap between the search relevance prediction value output by the discriminant network of the large language model and the actual search relevance value;
[0111] The second sub-loss is a first generation loss determined based on the first type of task samples; the first generation loss is used to measure the gap between the first generation result and the first generation label output by the generation network of the large language model for the first type of task samples.
[0112] Among them, the expression of the loss L(U) of the large language model is shown in formula (1)
[0113] L(U)=mask*αL cls (U)+(1-α)L gen (U) (1)
[0114] Wherein, when mask is 1, it is the loss corresponding to the first type of task samples; α is a hyperparameter, which can be set based on actual conditions and is not limited in the present disclosure; L cls (U) is the first child loss; L gen (U) is the second sub-loss, i.e., generation loss.
[0115] L cls The specific expression of (U) is shown in formula (2):
[0116] L cls (U)=-log(softmax(F cls W cls )) (2)
[0117] Among them, F cls Denotes the discriminant vector, W cls Represents the parameters of the discriminant network.
[0118] The first type of task sample can be exemplarily U={u1,...u n ,u n+1 ,...u m}, where u1-u n is the input information, u n+1 -u m is the generated result, L gen The specific expression of (U) is shown in (3):
[0119]
[0120] Where θ is the model parameter, n and m are positive integers greater than 1, where n is less than m, and u i Generate the result for the i-th one.
[0121] In the disclosed embodiment, since the correlation tasks include generation tasks and discrimination tasks, appropriate loss functions are designed for the loss functions of the first type of task samples to determine the corresponding generation task loss and discrimination task loss, and the two loss relationships are combined to obtain the target loss relationship for training the preset large language model, which can effectively improve the model training efficiency and training effect.
[0122] In some embodiments, for the second type of task samples, the loss used to train the large language model is a second generation loss;
[0123] The second generation loss is used to measure the gap between the second generation result output by the generation network of the large language model for the second type of task samples and the second generation label.
[0124] When implementing, in the case of the second type of task samples, the mask is 0, that is, the second generation loss is obtained, and the specific expression is shown in formula (4)
[0125] L(U)′=(1-α)L gen (U) (4)
[0126] Among them, when mask is 0, it is the loss corresponding to the second type of task samples. The meanings of other parameters are similar to the above, and the embodiments of this disclosure will not be repeated here.
[0127] In the embodiment of the present disclosure, a corresponding loss function calculation method is provided for the second type of task samples. At the same time, a mask is used to quickly perform the calculation of the corresponding loss function based on different task instructions. This method can effectively improve the model training efficiency and training effect.
[0128] In summary, the architecture diagram of the large language model in the embodiment of the present disclosure is as follows: Figure 7 As shown, corresponding prompt information is constructed based on the first type of task sample and / or the second type of task sample. For the first type of task sample, the prompt information includes the search request, the core content corresponding to the search request, the search intent, and the task instruction corresponding to the related task. The task instruction is also distinguished as indicating whether the sample is a dirty sample or a clean sample.
[0129] For the second type of task samples, the prompt information may include the text content to be processed and the task instructions corresponding to the general language task. The task instructions are used to indicate the task to be performed, and exemplary examples may include semantic understanding, paragraph summarization, and information reasoning. After the above construction is completed, the large language model is input to obtain the generated result. The discriminant vector corresponding to the generated result is obtained and input into the discriminant network to obtain the final result.
[0130] Based on the same technical concept, a correlation discrimination method is also proposed in the embodiment of the present disclosure, which uses the large language model obtained by the above training process, such as Figure 8 As shown, including:
[0131] S801, obtaining user query and search results.
[0132] The user query and the search results are two pieces of information for which search relevance needs to be calculated. The search results are the search results obtained by the recommendation engine for the user query.
[0133] S802, constructing prompt words based on the user query and search results; the prompt words include task instructions that prompt the large language model to perform a related task;
[0134] The specific method of constructing the prompt words is similar to the training process, and will not be described in detail in this embodiment of the present disclosure.
[0135] S803: Input the prompt word into the large language model to obtain the prediction result of the large language model for the search relevance between the user query and the search result, as well as the generation reason.
[0136] For the same user query, multiple search results may be obtained. Each search result can be used to determine its search relevance to the user query based on the method provided by the embodiments of the present disclosure. Subsequently, the multiple search results can be sorted based on the search relevance to obtain high-quality search results for recommendation to the user.
[0137] The generated reason, i.e., the training result generated based on the generative network, can be used to explain the reason for the provided search relevance. For example, in the aforementioned example, the reason for providing Bufei Pills in AI medical treatment.
[0138] In the disclosed embodiment, search relevance is provided based on the large language model, and reasons are generated, so that the decision-making principle of the large language model can be more easily understood, and the accuracy and rationality of the recommendation engine's recommendations can be improved.
[0139] In the embodiment of the present disclosure, when the first type of task samples for training the large language model include a first sample set and a second sample set, the task instructions are instructions corresponding to the second sample set;
[0140] The sample quality of the first sample set is lower than that of the second sample set, and the first sample set includes dirty samples; the second sample set includes clean samples.
[0141] In the embodiment of the present disclosure, different types of first sample sets and second sample sets are used for training. In the inference stage, the task instructions of the second sample set can be used for inference, so that the large language model can accurately understand the relationship between the input user query and the search results, and output accurate generation reasons and search relevance scores.
[0142] Based on the same technical concept, the present disclosure also proposes a large language model training device 900, such as Figure 9 As shown, including:
[0143] The first acquisition module 901 is used to acquire a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a correlation task, which is used to determine search correlation; the task type of the second type of task sample is a general language task;
[0144] The training module 902 is used to prompt the large language model with the task type to be performed based on the task instruction, so as to train the large language model based on samples of different task types.
[0145] In some embodiments, the first type of task samples includes a first sample set and a second sample set;
[0146] The sample quality of the first sample set is lower than the sample quality of the second sample set, and the first sample set includes dirty samples.
[0147] In some embodiments, the system further includes a first determining module configured to:
[0148] Obtaining a first sample search request, a first sample search result recommended for the first sample search request, and a network address of the first sample search result;
[0149] Performing effective information compression on the first sample search results to obtain a first core content of the first sample search results;
[0150] If labeling is required, obtaining a second sample search result from the network address of the first sample search result;
[0151] The dirty sample consisting of the first sample search request, the first core content, and the first label is obtained by labeling based on the first sample search request and the second sample search result.
[0152] In some embodiments, the system further includes a second determining module configured to:
[0153] Obtaining a second sample search request and a snapshot of a third sample search result recommended for the second sample search request;
[0154] Performing effective information compression on the third sample search results to obtain a second core content of the third sample search results;
[0155] If labeling is required, labeling is performed based on the snapshot of the second sample search request and the third sample search result to obtain the clean sample consisting of the second sample search request, the second core content and the second label.
[0156] In some embodiments, the task instruction of the first sample set is further used to indicate that the sample learned by the large language model is a dirty sample;
[0157] The task instructions of the second sample set are also used to prompt that the samples learned by the large language model are clean samples.
[0158] In some embodiments, the training module includes:
[0159] a construction unit, configured to construct a first type of prompt information based on the first type of task sample; the first type of prompt information includes a third sample search request in the first type of task sample, a third core content corresponding to the third sample search request, a search intent of the third sample search request, and a task instruction corresponding to the related task;
[0160] A training unit is configured to train the large language model based on the first type of prompt information.
[0161] In some embodiments, for the first type of task samples, the loss used to train the large language model includes a first sub-loss and a second sub-loss;
[0162] The first sub-loss is a discriminant loss determined based on the first type of task samples; the discriminant loss is used to measure the gap between the search relevance prediction value output by the discriminant network of the large language model and the actual search relevance value;
[0163] The second sub-loss is a first generation loss determined based on the first type of task samples; the first generation loss is used to measure the gap between the first generation result and the first generation label output by the generation network of the large language model for the first type of task samples.
[0164] In some embodiments, for the second type of task samples, the loss used to train the large language model is a second generation loss;
[0165] The second generation loss is used to measure the gap between the second generation result output by the generation network of the large language model for the second type of task samples and the second generation label.
[0166] In some embodiments, for the first type of task sample, the prompt information of the first type of task sample further includes a first target character;
[0167] The first target character is subjected to inference by the large language model to obtain a discriminant vector;
[0168] The discriminant vector is used to input into the discriminant network of the large language model to discriminate the search relevance between the user query and the search results.
[0169] Based on the same technical concept, the embodiment of the present disclosure also proposes a correlation determination device 1000, such as Figure 10 As shown, the large language model obtained based on the above device is applied, including:
[0170] The second acquisition module 1001 is used to obtain user queries and search results;
[0171] A construction module 1002 is configured to construct a prompt word based on the user query and the search results; the prompt word includes a task instruction prompting the large language model to perform a relevance task;
[0172] The processing module 1003 is configured to input the prompt word into the large language model, obtain a prediction result of the large language model for the search relevance between the user query and the search result, and generate a reason.
[0173] In some embodiments, when the first type of task samples for training the large language model include a first sample set and a second sample set, the task instructions are instructions corresponding to the second sample set;
[0174] The sample quality of the first sample set is lower than that of the second sample set, and the first sample set includes dirty samples; the second sample set includes clean samples.
[0175] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0176] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0177] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0178] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0179] like Figure 11 As shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the device 1100 can also be stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0180] Various components in device 1100 are connected to I / O interface 1105, including an input unit 1106, such as a keyboard and mouse; an output unit 1107, such as various types of displays and speakers; a storage unit 1108, such as a magnetic disk and optical disk; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0181] The computing unit 1101 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1101 performs the various methods and processes described above, such as the training method and / or the relevance determination method of the large language model. For example, in some embodiments, the training method and / or the relevance determination method of the large language model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the training method and / or the relevance determination method of the large language model described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the large language model training method and / or the relevance determination method in any other appropriate manner (for example, by means of firmware).
[0182] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0183] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0184] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0185] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0186] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0187] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0188] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0189] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a large language model, comprising: Obtaining a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a correlation task, and the correlation task is used to determine search correlation; The task type of the second type of task samples is general language task; Based on the task instructions, the large language model is prompted with the task type to be performed, so as to train the large language model based on samples of different task types; The first type of task samples includes a first sample set and a second sample set; The sample quality of the first sample set is lower than the sample quality of the second sample set, and the first sample set includes dirty samples, where the dirty samples refer to data instances that are incorrectly labeled, and the second sample set includes clean samples, where the clean samples refer to data instances that are not incorrectly labeled; The method further includes determining dirty samples in the first sample set based on the following method: Obtaining a first sample search request, a first sample search result recommended for the first sample search request, and a network address of the first sample search result; Performing effective information compression on the first sample search results to obtain a first core content of the first sample search results; If labeling is required, obtaining a second sample search result from the content corresponding to the network address of the first sample search result; Based on the first sample search request and the second sample search results, the dirty sample consisting of the first sample search request, the first core content and the first label is obtained, wherein the first label represents the search relevance between the second sample search results and the first sample search request.
2. The method according to claim 1, wherein Determine the clean samples in the second sample set based on the following method: Obtaining a second sample search request and a snapshot of a third sample search result recommended for the second sample search request; Performing effective information compression on the third sample search results to obtain a second core content of the third sample search results; If labeling is required, labeling is performed based on the snapshot of the second sample search request and the third sample search result to obtain the clean sample consisting of the second sample search request, the second core content and the second label.
3. The method according to claim 1, wherein The task instruction of the first sample set is further used to prompt that the samples learned by the large language model are dirty samples; The task instructions of the second sample set are also used to prompt that the samples learned by the large language model are clean samples.
4. The method according to any one of claims 1 to 3, wherein Training the large language model based on the first type of task samples includes: Constructing a first type of prompt information based on the first type of task sample; the first type of prompt information includes a third sample search request in the first type of task sample, a third core content corresponding to the third sample search request, a search intent of the third sample search request, and a task instruction corresponding to the related task; The large language model is trained based on the first type of prompt information.
5. The method according to any one of claims 1 to 3, wherein For the first type of task samples, the loss used to train the large language model includes a first sub-loss and a second sub-loss; The first sub-loss is a discriminant loss determined based on the first type of task samples; the discriminant loss is used to measure the gap between the search relevance prediction value output by the discriminant network of the large language model and the actual search relevance value; The second sub-loss is a first generation loss determined based on the first type of task samples; the first generation loss is used to measure the gap between the first generation result and the first generation label output by the generation network of the large language model for the first type of task samples.
6. The method according to any one of claims 1 to 3, wherein For the second type of task samples, the loss used to train the large language model is the second generation loss; The second generation loss is used to measure the gap between the second generation result output by the generation network of the large language model for the second type of task samples and the second generation label.
7. The method according to any one of claims 1 to 3, wherein For the first type of task sample, the prompt information of the first type of task sample further includes a first target character; The first target character is subjected to inference by the large language model to obtain a discriminant vector; The discriminant vector is used to input into the discriminant network of the large language model to discriminate the search relevance between the user query and the search results.
8. A method for determining relevance, using a large language model obtained by the method according to any one of claims 1 to 7, comprising: Get user queries and search results; Constructing prompt words based on the user query and the search results; The prompt words include task instructions that prompt the large language model to perform a related task; The prompt word is input into the large language model to obtain a prediction result of the large language model for the search relevance between the user query and the search result, and a generation reason.
9. The method according to claim 8, wherein, when the first type of task samples for training the large language model includes a first sample set and a second sample set, the task instructions are instructions corresponding to the second sample set; in, The sample quality of the first sample set is lower than that of the second sample set, and the first sample set includes dirty samples; the second sample set includes clean samples.
10. A large language model training device, comprising: A first acquisition module is used to acquire a first type of task sample and a second type of task sample, wherein the task type of the first type of task sample is a correlation task, and the correlation task is used to determine search correlation; The task type of the second type of task samples is general language task; A training module is used to prompt the large language model with the task type to be performed based on the task instructions, so as to train the large language model based on samples of different task types; The first type of task samples includes a first sample set and a second sample set; The sample quality of the first sample set is lower than the sample quality of the second sample set, and the first sample set includes dirty samples, where the dirty samples refer to data instances that are incorrectly labeled, and the second sample set includes clean samples, where the clean samples refer to data instances that are not incorrectly labeled; The system further includes a first determining module, configured to: Obtaining a first sample search request, a first sample search result recommended for the first sample search request, and a network address of the first sample search result; Performing effective information compression on the first sample search results to obtain a first core content of the first sample search results; If labeling is required, obtaining a second sample search result from the content corresponding to the network address of the first sample search result; Based on the first sample search request and the second sample search results, the dirty sample consisting of the first sample search request, the first core content and the first label is obtained, wherein the first label represents the search relevance between the second sample search results and the first sample search request.
11. The apparatus according to claim 10, further comprising a second determining module, configured to: Obtaining a second sample search request and a snapshot of a third sample search result recommended for the second sample search request; Performing effective information compression on the third sample search results to obtain a second core content of the third sample search results; If labeling is required, labeling is performed based on the snapshot of the second sample search request and the third sample search result to obtain a clean sample consisting of the second sample search request, the second core content and the second label.
12. The device according to claim 10, wherein The task instruction of the first sample set is further used to prompt that the sample learned by the large language model is a dirty sample; The task instructions of the second sample set are also used to prompt that the samples learned by the large language model are clean samples.
13. The device according to any one of claims 10 to 12, wherein: The training module includes: a construction unit, configured to construct a first type of prompt information based on the first type of task sample; the first type of prompt information includes a third sample search request in the first type of task sample, a third core content corresponding to the third sample search request, a search intent of the third sample search request, and a task instruction corresponding to the related task; A training unit is configured to train the large language model based on the first type of prompt information.
14. The device according to any one of claims 10 to 12, wherein: For the first type of task samples, the loss used to train the large language model includes a first sub-loss and a second sub-loss; The first sub-loss is a discriminant loss determined based on the first type of task samples; the discriminant loss is used to measure the gap between the search relevance prediction value output by the discriminant network of the large language model and the actual search relevance value; The second sub-loss is a first generation loss determined based on the first type of task samples; the first generation loss is used to measure the gap between the first generation result and the first generation label output by the generation network of the large language model for the first type of task samples.
15. The device according to any one of claims 10 to 12, wherein: For the second type of task samples, the loss used to train the large language model is the second generation loss; The second generation loss is used to measure the gap between the second generation result output by the generation network of the large language model for the second type of task samples and the second generation label.
16. The device according to any one of claims 10 to 12, wherein: For the first type of task sample, the prompt information of the first type of task sample further includes a first target character; The first target character is subjected to inference by the large language model to obtain a discriminant vector; The discriminant vector is used to input into the discriminant network of the large language model to discriminate the search relevance between the user query and the search results.
17. A relevance determination device, using a large language model obtained by the device according to any one of claims 10 to 16, comprising: The second acquisition module is used to obtain user queries and search results; A construction module, configured to construct prompt words based on the user query and the search results; The prompt words include task instructions that prompt the large language model to perform a related task; A processing module is configured to input the prompt word into the large language model, obtain a prediction result of the large language model for the search relevance between the user query and the search result, and generate a reason.
18. The apparatus according to claim 17, wherein, when the first type of task samples for training the large language model includes a first sample set and a second sample set, the task instructions are instructions corresponding to the second sample set; in, The sample quality of the first sample set is lower than that of the second sample set, and the first sample set includes dirty samples; the second sample set includes clean samples.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Information search method and search engine
CN103729374A
Search request recommendation method and device, electronic equipment and storage medium
CN116628253A