Model training method, correlation determination method, device, equipment and storage medium
By adding a full connection layer to the large language model, a correlation model with both discrimination and generation capabilities is built, the problem of lack of interpretability of existing models is solved, and more efficient search results optimization and user experience improvement is achieved.
Patent Information
- Application Number
- CN202311340353.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-10-16
AI Technical Summary
The existing correlation model lacks interpretability in search engines, making it difficult to understand the decision-making principles of the model, resulting in difficulty in optimization and debugging, and being unable to accurately provide appropriate search results to users.
Add a full connection layer to the large language model to build a correlation model with both discriminant and generational capabilities, and realize knowledge transfer between discriminant and generational tasks through the training process, and output correlation scores and reasons.
It improves the interpretability and optimization efficiency of the model, can better understand the principle of model decisions, and improves the accuracy and user experience of search results.
Smart Images

Figure CN117573817B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to the fields of intelligent search, deep learning, natural language processing, large language models, etc. Background Art
[0002] With the development of Internet technology, more and more users begin to use terminal devices to search for various information on the Internet. When searching, users usually input query terms, and the search engine can use a relevance model to determine the relevance scores between the query terms and each candidate search result, and then feedback appropriate search results to the user according to the ranking of the relevance scores. Summary of the Invention
[0003] The present disclosure provides a model training method, a relevance determination method, an apparatus, a device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided a method for training a relevance model, including:
[0005] Obtain sample input data, where the sample input data includes sample query terms, sample search results, and standard relevance reasons, the sample input data corresponds to a sample label, the sample label is used to represent the standard relevance score between the sample query terms and the sample search results, and the standard relevance reason is the relevance reason corresponding to the standard relevance score;
[0006] Input the sample input data into a preset relevance model, where the preset relevance model includes a large language model and a fully connected layer connected to the large language model;
[0007] Determine a sample relevance score according to the output of the fully connected layer, and determine a sample relevance reason according to the output of the large language model;
[0008] Determine a target loss relationship according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason, and train the preset relevance model based on the target loss relationship.
[0009] According to another aspect of the present disclosure, there is provided a method for determining relevance, including:
[0010] Construct input data according to the user's query terms, where the input data includes query terms and candidate search results;
[0011] Input the input data into a target relevance model, where the target relevance model includes a target large language model and a target fully connected layer connected to the target large language model;
[0012] Determine a target relevance score between the query term and the candidate search result according to the output of the target fully connected layer, and determine a target relevance reason between the query term and the candidate search result according to the output of the target large language model.
[0013] According to another aspect of the present disclosure, there is provided a training device for a relevance model, including:
[0014] A sample data acquisition module, configured to acquire sample input data, where the sample input data includes a sample query term, a sample search result, and a standard relevance reason, the sample input data corresponds to a sample label, and the sample label is used to represent a standard relevance score between the sample query term and the sample search result, and the standard relevance reason is a relevance reason corresponding to the standard relevance score;
[0015] A sample data input module, configured to input the sample input data into a preset relevance model, where the preset relevance model includes a large language model and a fully connected layer connected to the large language model;
[0016] A sample score determination module, configured to determine a sample relevance score according to the output of the fully connected layer;
[0017] A sample reason determination module, configured to determine a sample relevance reason according to the output of the large language model;
[0018] A loss relationship determination module, configured to determine a target loss relationship according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason;
[0019] A training module, configured to train the preset relevance model based on the target loss relationship.
[0020] According to another aspect of the present disclosure, there is provided a relevance determination device, including:
[0021] An input data construction module, configured to construct input data according to a user's query term, where the input data includes the query term and candidate search results;
[0022] A data input module, configured to input the input data into a target relevance model, where the target relevance model includes a target large language model and a target fully connected layer connected to the target large language model;
[0023] A score determination module, configured to determine a target relevance score between the query term and the candidate search result according to the output of the target fully connected layer;
[0024] A reason determination module, configured to determine an objective relevance reason between the query term and the candidate search results according to the output of the target large language model.
[0025] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the embodiments of the present disclosure.
[0029] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the embodiments of the present disclosure.
[0030] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the method described in any embodiment of the present disclosure when executed by a processor.
[0031] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0033] Figure 1 is a flowchart of a method for training a relevance model provided by an embodiment of the present disclosure;
[0034] Figure 2 is a flowchart of another method for training a relevance model provided by an embodiment of the present disclosure;
[0035] Figure 3 is a flowchart of a method for determining relevance provided by an embodiment of the present disclosure;
[0036] Figure 4 is a schematic diagram of an objective relevance model provided by an embodiment of the present disclosure;
[0037] Figure 5 is a schematic structural diagram of a device for training a relevance model provided by an embodiment of the present disclosure;
[0038] Figure 6 It is a schematic structural diagram of a relevance determination device provided according to an embodiment of the present disclosure;
[0039] Figure 7 It is a block diagram of an electronic device for implementing the method of the embodiment of the present disclosure. Specific Embodiments
[0040] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0041] To better understand the technical solutions of the embodiments of the present disclosure, the related technologies are introduced below. The relevance in the present disclosure can be understood as search relevance. Search relevance is generally used in the field of information retrieval to measure the degree of relevance between a user's search query term (hereinafter referred to as the query term) and the search results. Search relevance discrimination is a crucial part of a search engine, which directly affects the quality of search results and the user experience. The relevance discrimination method is usually implemented based on a search relevance model (which can be abbreviated as a relevance model). The relevance model can quantify the matching degree between the search results and the user's search intention, and then the search results can be screened and sorted according to the relevance score (generally a probability score). Existing relevance models are usually machine learning models. Although these models can achieve a certain prediction accuracy, they lack interpretability, which makes it difficult to understand the decision principle of the model, is not conducive to problem diagnosis, makes it difficult to optimize and debug the model, and is also difficult to optimize the quality of the candidate search results themselves.
[0042] Figure 1 It is a flowchart of a method for training a relevance model provided according to an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of training a relevance model for an information search scenario. This method can be executed by a training device for a relevance model, and this device can be implemented in a hardware and / or software manner and can be configured in an electronic device. Refer to Figure 1 and the method specifically includes the following:
[0043] S101. Obtain sample input data, where the sample input data includes a sample query term, sample search results, and a standard relevance reason. The sample input data corresponds to a sample label, and the sample label is used to represent the standard relevance score of the sample query term and the sample search results. The standard relevance reason is the relevance reason corresponding to the standard relevance score;
[0044] S102. Input the sample input data into a preset correlation model, where the preset correlation model includes a large language model and a fully connected layer connected to the large language model;
[0045] S103. Determine a sample correlation score according to the output of the fully connected layer, and determine a sample correlation reason according to the output of the large language model;
[0046] S104. Determine a target loss relationship according to the sample correlation score, the sample label corresponding to the sample input data, the sample correlation reason, and the standard correlation reason, and train the preset correlation model based on the target loss relationship.
[0047] Among them, the sample data volume and model parameters of the large language model (Large Language Model, LLM) can both be at the massive level. For example, it can have billions or even hundreds of billions of parameters. It is trained on large-scale text data or other modal data and can be used for various natural language processing tasks. It has good text generation ability and can be applied to language translation, question answering systems, etc. The large language model can, for example, adopt an N-layer Transformer network structure with an encoder and a decoder, or a Unified pre-trained Language Model (UniLM) network structure. It can be understood that the large language model can also be other neural network models based on the Transformer network structure, which is not limited here. The input of the large language model is generally composed of tokens (also known as tokens or markers, etc.). Each token can correspond to a single character, character, word, or special symbol, etc.
[0048] In the embodiments of the present disclosure, the good generation ability of the large language model can be used to generate the correlation reason between the query term and the candidate search results. However, the existing large language models do not have a discrimination ability and cannot directly output the correlation score between the query term and the candidate search results. Therefore, it is impossible to perform fine-grained sorting between the candidate search results and accurately provide appropriate search results to users.
[0049] In the embodiments of the present disclosure, the model structure at the top layer of the large language model is modified, and a fully connected layer connected thereto is added to the large language model. That is, the relevance model in the embodiments of the present disclosure includes the large language model and the fully connected layer connected to the large language model. The fully connected layer is used to output the relevance score. Through one training process, the relevance model in the embodiments of the present disclosure can have both discriminative ability and generative ability, that is, it can output the relevance reason and the relevance score at the same time, and the discriminative task and the generative task share the model parameters to achieve knowledge transfer between the discriminative task and the generative task and improve the performance of the model on each task. Among them, the specific connection manner between the large language model and the fully connected layer is not limited.
[0050] Exemplarily, the sample input data can be understood as labeled data. For each pair of sample query words (query words for model training) and sample search results (search results for model training), manual or machine labeling can be performed, that is, the corresponding sample label is labeled. The sample label is used to represent the relevance score determined by manual or machine means for the sample query word and the sample search result, denoted as the standard relevance score, and can also be understood as the expected relevance score. In addition, manual or machine means can be used to add a reason description for why the labeled sample label is such, that is, the reason description for the value of the relevance score of the sample query word and the sample search result being the standard relevance score, denoted as the standard relevance reason.
[0051] Exemplarily, the preset relevance model can be understood as the relevance model in the training process of the embodiments of the present disclosure. The sample input data is input into the preset relevance model, the sample relevance score is determined according to the output of the fully connected layer in the preset relevance model, and the sample relevance reason is determined according to the output of the large language model in the preset relevance model. The target loss relationship is determined according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason, and the preset relevance model is trained based on the target loss relationship to optimize the model parameters in the preset relevance model. Among them, the target loss relationship can be represented by a loss value. During the training process, the target can be to minimize the target loss relationship, and training means such as stochastic gradient descent are used to continuously optimize the model parameters in the preset relevance model until the preset training termination condition is met. The specific training termination condition can be set according to actual needs and is not limited in the embodiments of the present disclosure. For example, it can be set based on the number of iterations, the degree of loss value convergence, or the model accuracy, etc.
[0052] In the technical solution provided by the embodiments of the present disclosure, the preset relevance model includes a large language model and a fully connected layer connected to the large language model. The sample relevance score is determined according to the output from the fully connected layer, the sample relevance reason is determined according to the output of the large language model, the target loss relationship is determined according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason, and the preset relevance model is trained based on the target loss relationship. Through one training process, the relevance model in the embodiments of the present disclosure can have both discriminative ability and generative ability, and the discriminative task and the generative task share model parameters, realizing knowledge transfer between the discriminative task and the generative task, improving the performance of the model on each task, enabling the trained preset relevance model to have good interpretability while ensuring accurate prediction of the relevance score, being able to better understand the decision principle of the model, being beneficial to problem diagnosis, making the optimization and debugging of the model more targeted, improving the optimization and debugging efficiency, and also being beneficial to efficiently optimizing the quality of the candidate search results themselves. In addition, compared with the multi-model solution of separately training a discriminative model and a generative model, it can effectively improve the training efficiency, save the resources required for training and deployment, improve the flexibility of the model, and avoid the problem that the consistency of the results of the two models for the same sample cannot be guaranteed.
[0053] In an alternative embodiment, the sample input data further includes a sample classification tag and a sample mask tag. The sample classification tag is associated with description information for prompting the output of the relevance score between the sample query term and the sample search result, and the sample mask tag is associated with description information for prompting the output of the relevance reason between the sample query term and the sample search result; the fully connected layer is used to receive the first sample output vector corresponding to the sample classification tag output by the large language model; wherein, the determining the sample relevance reason according to the output of the large language model includes: determining the sample relevance reason according to the second sample output vector corresponding to the sample mask tag output by the large language model. Thus, the relevance score and the relevance reason can be determined more accurately.
[0054] Exemplarily, the description information associated with the sample classification tag for prompting the output of the relevance score between the sample query term and the sample search result is denoted as the first description information. The first description information may be located before the sample classification tag in the sample input data and adjacent to the sample classification tag. The description information associated with the sample mask tag for prompting the output of the relevance reason between the sample query term and the sample search result may be denoted as the second description information. The second description information may be located before the sample mask tag in the sample input data and adjacent to the sample mask tag.
[0055] As an example, the sample input data can be specifically: Query term: "XX Spoken Language". Search result: "Study abroad planning, the study abroad planning team will conduct a study abroad assessment for you and help you select an ideal study abroad institution". The relevance score between the query term and the search result is [CLS], and the reason is [gMASK]. The query term "XX Spoken Language" involves the spoken language part of the XX exam, while the search result mainly focuses on study abroad planning and the selection of study abroad institutions, without mentioning content related to "XX Spoken Language", so the relevance is relatively low.
[0056] Among them, "XX" refers to the name of a spoken language exam, which can be filled according to the actual situation.
[0057] In the above example, the sample query term is "XX Spoken Language"; the sample search result is "Study abroad planning; the study abroad planning team will conduct a study abroad assessment for you and help you select an ideal study abroad institution"; the sample classification label is "[CLS]", which can be understood as a special token. The description information associated with the sample classification label for prompting the output of the relevance score between the sample query term and the sample search result is "The relevance score between the query term and the search result is"; the sample mask label is "[gMASK]", which can be understood as a special token. The description information associated with the sample mask label for prompting the reason for the relevance between the sample query term and the sample search result; the standard relevance reason is: The query term "XX Spoken Language" involves the spoken language part of the XX exam, while the search result mainly focuses on study abroad planning and the selection of study abroad institutions, without mentioning content related to "XX Spoken Language", so the relevance is relatively low.
[0058] Figure 2 It is a flowchart of another method for training a relevance model provided according to an embodiment of the present disclosure. This embodiment is optimized based on the above-mentioned alternative solutions in the foregoing embodiment. The method may include:
[0059] S201. Obtain sample input data, where the sample input data includes a sample query term, a sample search result, a sample classification label, a sample mask label, and a standard relevance reason. The sample classification label is associated with description information for prompting the output of the relevance score between the sample query term and the sample search result, and the sample mask label is associated with description information for prompting the reason for the relevance between the sample query term and the sample search result.
[0060] S202. Input the sample input data into a preset relevance model, where the preset relevance model includes a large language model and a fully connected layer connected to the large language model. The fully connected layer is used to receive the first sample output vector corresponding to the sample classification label output by the large language model.
[0061] S203. Determine the sample correlation score based on the output of the fully connected layer, and determine the sample correlation reason based on the second sample output vector corresponding to the sample mask token output by the large language model.
[0062] S204. Determine the first loss relationship according to the sample correlation score and the sample label corresponding to the sample input data.
[0063] In an alternative embodiment, the first loss relationship is determined based on the cross-entropy loss function of the normalized exponential function (softmax). The first loss relationship can also be referred to as the discriminative loss.
[0064] Exemplarily, assume a sample data is denoted as U, U = {u1,...u n , u n+1 ,...u m}, where u1 - u n are the input data other than the standard correlation reason, and u n+1 - u m are the standard correlation reasons. The first loss relationship can be expressed by the following expression:
[0065] L cls (U) = -log(softmax(F cls W cls ))
[0066] where L cls (U) represents the first loss relationship, F cls represents the first sample output vector, and W cls represents the parameters of the fully connected layer. Specifically, by adding softmax to obtain multiple nodes, the purpose of the first loss relationship is to make the softmax value of the node corresponding to the sample label the largest.
[0067] S205. Determine the second loss relationship according to the sample correlation reason and the standard correlation reason.
[0068] In an alternative embodiment, the second loss function is determined based on the standard language modeling objective loss function, and the standard language modeling loss function is essentially a cross-entropy loss function. The second loss relationship can also be referred to as the generation loss.
[0069] As in the above example, the first loss relationship can be expressed by the following expression:
[0070]
[0071] where L gen(U) represents the second loss relationship, and θ represents the parameters of the large language model.
[0072] S206. Determine the target loss relationship according to the first loss relationship and the second loss relationship.
[0073] In an alternative implementation, the target loss relationship can be determined according to the weighted sum of the first loss relationship and the second loss relationship. For example, the product of the first weight coefficient and the first loss relationship, plus the product of the second weight coefficient and the second loss relationship, obtains the target loss relationship, and the sum of the first weight coefficient and the second weight coefficient is 1.
[0074] As in the above example, the target loss relationship can be expressed by the following expression:
[0075] L(U) = αL cls (U) + (1 - α)L gen (U)
[0076] Among them, L(U) represents the target loss relationship, and α is a hyperparameter used to control the distribution of the weight coefficients of the first loss relationship and the second loss relationship.
[0077] S207. Train the preset relevance model based on the target loss relationship.
[0078] Exemplarily, through stochastic gradient descent, with the goal of minimizing L(U), the preset relevance model is trained to update the model parameters.
[0079] The training method of the relevance model provided by the embodiments of the present disclosure designs appropriate loss functions for the discrimination task and the generation task respectively to determine the corresponding loss relationships, and combines the two loss relationships to obtain the target loss relationship for training the preset relevance model, which can effectively improve the model training efficiency and training effect, so that the trained model can more accurately output the relevance score and the relevance reason at the same time.
[0080] In an alternative implementation, it may further include: determining a first relevance model according to the training result of the preset relevance model; removing the fully connected layer in the first relevance model to obtain a second relevance model, where the second relevance model is used to determine the relevance reason between the query word input by the user and the candidate search result. Thus, by pruning the trained first relevance model, it can be adapted to application scenarios that do not require relevance reasons, and through the joint training of the discrimination task and the generation task, the knowledge of the generation task can be used to help the discrimination task, which can make the second relevance model have higher accuracy and better performance than the model obtained by training the discrimination task alone.
[0081] In an alternative embodiment, it may further include: determining a first relevance model according to the training result of the preset relevance model; removing the model structure corresponding to the sample mask token in the first relevance model to obtain a third relevance model, where the third relevance model is used to determine the relevance score between the query term input by the user and the candidate search results. Thus, by pruning the trained first relevance model, it can be adapted to application scenarios that do not require relevance scores. And through the joint training of the discrimination task and the generation task, the knowledge of the discrimination task can be used to assist the generation task, enabling the third relevance model to have higher accuracy and better performance compared to the model obtained by training the generation task alone.
[0082] Optionally, for the first relevance model, a backup can be made, and after the above-mentioned removal process respectively, a second relevance model and a third relevance model are obtained respectively, so that the first relevance model, the second relevance model and the third relevance model can be respectively applicable to different application scenarios.
[0083] In an alternative embodiment, the candidate search results include online placement information. Thus, it can better adapt to search scenarios that include online placement information.
[0084] Exemplarily, the online placement information may include advertisement information. In the field of search advertising, interpretability can improve decision-making transparency and facilitate problem diagnosis. For example, it can help advertisers understand the effect of their advertisement placements, why the relevance is poor, so as to provide a decision-making basis and optimize the advertisement quality.
[0085] Figure 3 It is a flowchart of a relevance determination method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of predicting the relevance between a query term of a user and candidate search results in an information search scenario. This method can be executed by a relevance determination device, which can be implemented in a hardware and / or software manner and can be configured in an electronic device. Refer to Figure 3 , the method specifically includes the following:
[0086] S301. Construct input data according to the query term of the user, where the input data includes the query term and candidate search results;
[0087] S302. Input the input data into a target relevance model, where the target relevance model includes a target large language model and a target fully connected layer connected to the target large language model;
[0088] S303. Determine the target relevance score between the query term and the candidate search results based on the output of the target fully connected layer, and determine the target relevance reason between the query term and the candidate search results based on the output of the target large language model.
[0089] Exemplarily, the target relevance model can be trained using the training method of the relevance model in any of the foregoing embodiments.
[0090] Exemplarily, the user can input the content they want to search for, such as keywords, etc., through the input control (such as an input box) corresponding to the search engine, and use the user's input content as the query term. There can be multiple candidate search results corresponding to the query term, forming a candidate search result set. When constructing the input data, for each candidate search result in the candidate search result set, the corresponding input data can be constructed respectively to determine the target relevance score and the target relevance reason corresponding to each candidate search result.
[0091] Optionally, the candidate search result set corresponding to the query term input by the current user can be determined according to the specific search scenario to better adapt to different search scenarios. For example, for the online promotion information search scenario, the candidate search results in the candidate search result set can include online promotion information, such as advertisement information, etc.; for another example, for the information search scenario, the candidate search results in the candidate search result set can include information, such as news, etc.
[0092] In the technical solution provided by the embodiments of the present disclosure, the target relevance model includes a target large language model and a target fully connected layer connected to the target large language model. The target relevance score is determined based on the output from the target fully connected layer, and the target relevance reason is determined based on the output of the target large language model, so as to realize using one model to output accurate relevance scores and relevance reasons simultaneously, making the decision principle of the model easier to understand, which can help users select the search results to be viewed more quickly, is beneficial to problem diagnosis, and is also beneficial to efficiently optimizing the quality of the candidate search results themselves.
[0093] Exemplarily, after obtaining the target relevance scores and target relevance reasons corresponding to each candidate search result in the candidate search result set, the candidate search results can be sorted according to the target relevance scores, and all or part of the sorted results can be displayed. The candidate search results in the displayed sorted results can be recorded as target search results, and each target search result can be associated with and displayed with the corresponding target relevance reason, helping the user intuitively understand why the target search result is in the current ranking based on the target relevance reason, and also enabling the user to quickly select the target search result they want to further understand based on the reason, improving the browsing efficiency of the search results and enhancing the user experience. Optionally, the candidate search results may include the topic information of the information object, and this topic information may be, for example, the title or abstract of the information object. When displaying the sorted results, the topic information and the corresponding target relevance reason may be displayed first, and in response to the user's trigger operation on the target topic information, the information object corresponding to the target topic information is displayed.
[0094] In an alternative embodiment, the input data further includes a classification token and a masking token. The classification token is associated with description information for prompting the output of the relevance score of the query term and the candidate search results, and the masking token is associated with description information for prompting the output of the relevance reason of the query term and the candidate search results; the target fully connected layer is used to receive the first output vector corresponding to the classification token output by the target large language model; wherein, determining the target relevance reason of the query term and the candidate search results according to the output of the target large language model includes: determining the target relevance reason of the query term and the candidate search results according to the second output vector corresponding to the masking token output by the target large language model. Thus, the relevance score and the relevance reason can be determined more accurately.
[0095] Exemplarily, the description information associated with the classification token for prompting the output of the relevance score of the query term and the candidate search results is denoted as the third description information. The third description information may be located before the classification token in the input data and adjacent to the classification token. The description information associated with the masking token for prompting the output of the relevance reason of the query term and the candidate search results may be denoted as the fourth description information. The fourth description information may be located before the masking token in the input data and adjacent to the masking token.
[0096] Figure 4 is a schematic diagram of a target relevance model provided according to an embodiment of the present disclosure. As Figure 4As shown, T represents a token. From T0 to gMASK represents the input data, where CLS represents the classification token and gMASK represents the masked token. E represents the embedding vector, and F represents the output vector. CLS is input into the target large language model (LLM) and converted into the embedding vector E CLS After that, through the processing of the internal model structure of the LLM, the first output vector F is obtained CLS and is input into the target fully connected layer. The probability score, that is, the target relevance score, is determined according to the output of the target fully connected layer. gMASK is input into the LLM and converted into the embedding vector E gMASK After that, through the processing of the internal model structure of the LLM, the second output vector F is obtained gMASK According to F gMASK the generation reason G is determined n+1 to G m that is, the target relevance reason. For example, the target relevance reason is obtained by decoding F gMASK
[0097] Figure 5 is a schematic structural diagram of a training device for a relevance model provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of training a relevance model for an information search scenario. This device can be implemented in a hardware and / or software manner and can be configured in an electronic device. Refer to Figure 5 The training device 500 for the relevance model includes:
[0098] A sample data acquisition module 501, configured to acquire sample input data, where the sample input data includes a sample query term, sample search results, and a standard relevance reason. The sample input data corresponds to a sample label, and the sample label is used to represent the standard relevance score of the sample query term and the sample search results. The standard relevance reason is the relevance reason corresponding to the standard relevance score;
[0099] A sample data input module 502, configured to input the sample input data into a preset relevance model, where the preset relevance model includes a large language model and a fully connected layer connected to the large language model;
[0100] A sample score determination module 503, configured to determine a sample relevance score according to the output of the fully connected layer;
[0101] A sample reason determination module 504, configured to determine a sample relevance reason according to the output of the large language model;
[0102] A loss relationship determination module 505, configured to determine a target loss relationship according to the sample correlation score, the sample label corresponding to the sample input data, the sample correlation reason, and the standard correlation reason;
[0103] A training module 506, configured to train the preset correlation model based on the target loss relationship.
[0104] In the technical solution provided by the embodiments of the present disclosure, the preset correlation model includes a large language model and a fully connected layer connected to the large language model. The sample correlation score is determined according to the output from the fully connected layer, the sample correlation reason is determined according to the output of the large language model, the target loss relationship is determined according to the sample correlation score, the sample label corresponding to the sample input data, the sample correlation reason, and the standard correlation reason, and the preset correlation model is trained based on the target loss relationship. Through one training process, the correlation model in the embodiments of the present disclosure can have both discriminative ability and generative ability, and the discriminative task and the generative task share model parameters, realizing knowledge transfer between the discriminative task and the generative task, improving the performance of the model on each task, enabling the preset correlation model after training to have good interpretability while ensuring accurate prediction of the correlation score, being able to better understand the decision principle of the model, being beneficial to problem diagnosis, making the optimization and debugging of the model more targeted, improving the optimization and debugging efficiency, and also being beneficial to efficiently optimizing the quality of the candidate search results themselves. In addition, compared with the multi-model solution of separately training a discriminative model and a generative model, it can effectively improve the training efficiency, save the resources required for training and deployment, improve the flexibility of the model, and avoid the problem that the consistency of the results of the two models for the same sample cannot be guaranteed.
[0105] In an alternative embodiment, the sample input data further includes a sample classification label and a sample mask label. The sample classification label is associated with description information for prompting the output of the correlation score between the sample query term and the sample search result, and the sample mask label is associated with description information for prompting the output of the correlation reason between the sample query term and the sample search result; the fully connected layer is used to receive the first sample output vector corresponding to the sample classification label output by the large language model;
[0106] Wherein, the sample reason determination module is specifically configured to:
[0107] Determine the sample correlation reason according to the second sample output vector corresponding to the sample mask label output by the large language model.
[0108] In an alternative embodiment, the loss relationship determination module includes:
[0109] A first loss relationship determination unit, configured to determine a first loss relationship according to the sample correlation score and the sample label corresponding to the sample input data;
[0110] A second loss relationship determination unit, configured to determine a second loss relationship according to the sample correlation reason and the standard correlation reason;
[0111] A third loss relationship determination unit, configured to determine a target loss relationship according to the first loss relationship and the second loss relationship.
[0112] In an alternative embodiment, the first loss relationship is determined based on a cross-entropy loss function of a normalized exponential function; and / or, the second loss function is determined based on a standard language modeling loss function.
[0113] In an alternative embodiment, it further includes:
[0114] A first model determination module, configured to determine a first correlation model according to the training result of the preset correlation model;
[0115] A second model determination module, configured to remove the fully connected layer in the first correlation model to obtain a second correlation model, where the second correlation model is used to determine the correlation reason between the query word input by the user and the candidate search result; or, configured to remove the model structure corresponding to the sample mask tag in the first correlation model to obtain a third correlation model, where the third correlation model is used to determine the correlation score between the query word input by the user and the candidate search result.
[0116] In an alternative embodiment, the sample search result includes online placement information.
[0117] Figure 6 It is a schematic structural diagram of a correlation determination device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of predicting the correlation between the query word of the user and the candidate search result in an information search scenario. This device can be implemented in a hardware and / or software manner and can be configured in an electronic device. Refer to Figure 6 , the correlation determination device 600 includes:
[0118] An input data construction module 601, configured to construct input data according to the query word of the user, where the input data includes the query word and the candidate search result;
[0119] A data input module 602, configured to input the input data into a target correlation model, where the target correlation model includes a target large language model and a target fully connected layer connected to the target large language model;
[0120] A score determination module 603, configured to determine a target relevance score between the query term and the candidate search result according to the output of the target fully-connected layer;
[0121] A reason determination module 604, configured to determine a target relevance reason between the query term and the candidate search result according to the output of the target large language model.
[0122] In the technical solution provided by the embodiments of the present disclosure, the target relevance model includes a target large language model and a target fully-connected layer connected to the target large language model. The target relevance score is determined according to the output from the target fully-connected layer, and the target relevance reason is determined according to the output of the target large language model, so as to realize simultaneously outputting an accurate relevance score and a relevance reason by using one model, making the decision principle of the model easier to understand, which is beneficial to helping users more quickly select search results to view, beneficial to problem diagnosis, and also beneficial to efficiently optimizing the quality of the candidate search results themselves.
[0123] In an alternative embodiment, the input data further includes a classification token and a mask token. The classification token is associated with description information for prompting the output of the relevance score between the query term and the candidate search result, and the mask token is associated with description information for prompting the output of the relevance reason between the query term and the candidate search result; the target fully-connected layer is used to receive the first output vector corresponding to the classification token output by the target large language model;
[0124] Wherein, the reason determination module is specifically configured to:
[0125] Determine the target relevance reason between the query term and the candidate search result according to the second output vector corresponding to the mask token output by the target large language model.
[0126] In an alternative embodiment, the candidate search result includes online placement information.
[0127] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0128] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0129] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0131] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 701 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the training method and / or the correlation determination method of the correlation model. For example, in some embodiments, the training method and / or the correlation determination method of the correlation model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method and / or the correlation determination method of the correlation model described above may be executed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the training method and / or the correlation determination method of the correlation model in any other suitable manner (e.g., by means of firmware).
[0133] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs, the one or more computer programs may be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor, and may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes may be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0137] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0138] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system or a server combined with blockchain.
[0139] Artificial intelligence is a discipline that studies the simulation of certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) by computers, including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0140] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0141] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitations are imposed herein.
[0142] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for training a relevance model, comprising: Obtaining sample input data, wherein the sample input data includes a sample query term, sample search results, and a standard relevance reason, the sample input data corresponds to a sample label, and the sample label is used to represent the standard relevance score of the sample query term and the sample search results, and the standard relevance reason is the relevance reason corresponding to the standard relevance score; Inputting the sample input data into a preset relevance model, wherein the preset relevance model includes a large language model and a fully connected layer connected to the large language model; Determining a sample relevance score according to the output of the fully connected layer, and determining a sample relevance reason according to the output of the large language model; Determining a target loss relationship according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason, and training the preset relevance model based on the target loss relationship; The sample input data further includes a sample classification marker and a sample mask marker, the sample classification marker is associated with first description information, the first description information is used to prompt the output of the relevance score of the sample query term and the sample search results, the first description information is located before the sample classification marker in the sample input data and is adjacent to the sample classification marker; the sample mask marker is associated with second description information, the second description information is used to prompt the output of the relevance reason of the sample query term and the sample search results, the second description information is located before the sample mask marker in the sample input data and is adjacent to the sample mask marker; the fully connected layer is used to receive the first sample output vector corresponding to the sample classification marker output by the large language model; Wherein, the determining the sample relevance reason according to the output of the large language model includes: Determining the sample relevance reason according to the second sample output vector corresponding to the sample mask marker output by the large language model.
2. The method according to claim 1, wherein The determining the target loss relationship according to the sample relevance score, the sample label corresponding to the sample input data, the sample relevance reason, and the standard relevance reason includes: Determining a first loss relationship according to the sample relevance score and the sample label corresponding to the sample input data; Determining a second loss relationship according to the sample relevance reason and the standard relevance reason; Determining the target loss relationship according to the first loss relationship and the second loss relationship.
3. The method according to claim 2, wherein The first loss relationship is determined based on a cross-entropy loss function of a normalized exponential function; and / or, the second loss relationship is determined based on a standard language modeling loss function.
4. The method according to any one of claims 1-3, further comprising: Determining a first relevance model according to the training result of the preset relevance model; Remove the fully connected layer in the first relevance model to obtain a second relevance model, where the second relevance model is used to determine the relevance reason between the query term input by the user and the candidate search results; or, remove the model structure corresponding to the sample mask token in the first relevance model to obtain a third relevance model, where the third relevance model is used to determine the relevance score between the query term input by the user and the candidate search results.
5. The method according to any one of claims 1-3, wherein The sample search results include online placement information.
6. A relevance determination method, comprising: Construct input data according to the query term of the user, where the input data includes the query term and candidate search results; Input the input data into a target relevance model, where the target relevance model includes a target large language model and a target fully connected layer connected to the target large language model; Determine the target relevance score between the query term and the candidate search results according to the output of the target fully connected layer, and determine the target relevance reason between the query term and the candidate search results according to the output of the target large language model; Wherein, the input data further includes a classification token and a mask token, the classification token is associated with third description information, the third description information is used to prompt the output of the relevance score between the query term and the candidate search results, the third description information is located before the classification token in the input data and is adjacent to the classification token; the mask token is associated with fourth description information, the fourth description information is used to prompt the output of the relevance reason between the query term and the candidate search results, the fourth description information is located before the mask token in the input data and is adjacent to the mask token; the target fully connected layer is used to receive the first output vector corresponding to the classification token output by the target large language model; Wherein, the determining the target relevance reason between the query term and the candidate search results according to the output of the target large language model includes: Determine the target relevance reason between the query term and the candidate search results according to the second output vector corresponding to the mask token output by the target large language model.
7. The method according to claim 6, wherein The candidate search results include online placement information.
8. A training device for a relevance model, comprising: A sample data acquisition module, configured to acquire sample input data, where the sample input data includes a sample query term, sample search results, and a standard relevance reason, the sample input data corresponds to a sample label, and the sample label is used to represent the standard relevance score between the sample query term and the sample search results, and the standard relevance reason is the relevance reason corresponding to the standard relevance score; A sample data input module, configured to input the sample input data into a preset relevance model, where the preset relevance model includes a large language model and a fully connected layer connected to the large language model; A sample score determination module, configured to determine a sample relevance score according to the output of the fully connected layer; A sample reason determination module, configured to determine a sample relevance reason according to the output of the large language model; A loss relationship determination module, configured to determine a target loss relationship according to the sample correlation score, the sample label corresponding to the sample input data, the sample correlation reason, and the standard correlation reason; A training module, configured to train the preset correlation model based on the target loss relationship; Wherein, the sample input data further includes a sample classification tag and a sample mask tag, the sample classification tag is associated with first description information, the first description information is used to prompt the output of the correlation score between the sample query term and the sample search result, and the first description information is located before the sample classification tag in the sample input data and is adjacent to the sample classification tag; the sample mask tag is associated with second description information, the second description information is used to prompt the output of the correlation reason between the sample query term and the sample search result, and the second description information is located before the sample mask tag in the sample input data and is adjacent to the sample mask tag; the fully connected layer is used to receive the first sample output vector corresponding to the sample classification tag output by the large language model; Wherein, the sample reason determination module is specifically configured to: determine the sample correlation reason according to the second sample output vector corresponding to the sample mask tag output by the large language model.
9. The device according to claim 8, wherein The loss relationship determination module includes: A first loss relationship determination unit, configured to determine a first loss relationship according to the sample correlation score and the sample label corresponding to the sample input data; A second loss relationship determination unit, configured to determine a second loss relationship according to the sample correlation reason and the standard correlation reason; A third loss relationship determination unit, configured to determine a target loss relationship according to the first loss relationship and the second loss relationship.
10. The apparatus according to claim 9, wherein, The first loss relationship is determined based on a cross-entropy loss function of a normalized exponential function; and / or, the second loss relationship is determined based on a standard language modeling loss function.
11. The device according to any one of claims 8-10, further comprising: A first model determination module, configured to determine a first correlation model according to the training result of the preset correlation model; A second model determination module, configured to remove the fully connected layer in the first correlation model to obtain a second correlation model, wherein the second correlation model is used to determine the correlation reason between the query term input by the user and the candidate search result; or, configured to remove the model structure corresponding to the sample mask tag in the first correlation model to obtain a third correlation model, wherein the third correlation model is used to determine the correlation score between the query term input by the user and the candidate search result.
12. The device according to any one of claims 8-10, wherein, The sample search result includes online placement information.
13. A correlation determination device, comprising: An input data construction module, configured to construct input data according to a query term of a user, wherein the input data includes the query term and candidate search results; A data input module, configured to input the input data into a target correlation model, wherein the target correlation model includes a target large language model and a target fully connected layer connected to the target large language model; A score determination module, configured to determine a target relevance score between the query term and the candidate search result according to the output of the target fully-connected layer; A reason determination module, configured to determine a target relevance reason between the query term and the candidate search result according to the output of the target large language model; Wherein, the input data further includes a classification token and a masking token, the classification token is associated with third description information, the third description information is used to prompt the output of the relevance score between the query term and the candidate search result, the third description information is located before the classification token in the input data and is adjacent to the classification token; the masking token is associated with fourth description information, the fourth description information is used to prompt the output of the relevance reason between the query term and the candidate search result, the fourth description information is located before the masking token in the input data and is adjacent to the masking token; the target fully-connected layer is used to receive the first output vector corresponding to the classification token output by the target large language model; Wherein, the reason determination module is specifically configured to: determine the target relevance reason between the query term and the candidate search result according to the second output vector corresponding to the masking token output by the target large language model.
14. The device according to claim 13, wherein, The candidate search results include online placement information.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Search result recommendation reason generation method and device, equipment and storage medium
CN113434763A
Sorting model training method and device of information retrieval system, medium and equipment
CN114780846A