Training method of keyword generation model, keyword generation method and related device
By cross-training large models on click and relevance tasks, the accuracy and relevance of keywords generated by large language models in commercial search scenarios are solved, thereby improving cost-effectiveness and user experience.
Patent Information
- Application Number
- CN202410940604.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-16
AI Technical Summary
Existing large-scale language models struggle to simultaneously generate keywords with both accuracy and relevance in commercial search scenarios, and traditional fine-tuning training methods are costly and cannot share information.
A multi-task fine-tuning instruction cross-training method is adopted. By constructing first-class and second-class samples, the click task and relevance task of the large model are trained respectively. Combined with the preset association relationship, the generation ability of the model is improved.
It improves the accuracy and relevance of keyword generation, reduces data acquisition and annotation costs, and enhances model diversity and user satisfaction.
Smart Images

Figure CN121350604A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the technical fields of large models and keyword generation. Background Technology
[0002] With the advancement of artificial intelligence technology, especially the emergence of large-scale language models, the application of AI (Artificial Intelligence) is becoming increasingly widespread in various scenarios. For example, in recommender systems, how to provide services using large models still faces challenges. Summary of the Invention
[0003] This disclosure provides a training method for a keyword generation model, a keyword generation method, and related apparatus.
[0004] According to one aspect of this disclosure, a method for training a keyword generation model is provided, including:
[0005] A first fine-tuning instruction is constructed based on a first type of sample; the first type of sample includes a first search sample and a first type of word set containing at least one first keyword; wherein, when searching for resources based on the first search sample, the resources described by the first keyword at least satisfy part of the search requirements of the first search sample;
[0006] The second fine-tuning instruction is constructed based on the second type of samples; the second type of samples includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship;
[0007] The initial large model is trained by cross-training the first and second fine-tuning instructions to obtain the keyword generation model.
[0008] According to one aspect of this disclosure, a keyword generation method is provided, including:
[0009] Based on the target search information, construct the prompt information;
[0010] Input the prompt information into the keyword generation model to obtain at least one keyword output by the keyword generation model.
[0011] According to another aspect of this disclosure, a training apparatus for a keyword generation model is provided, comprising:
[0012] The first construction module is used to construct a first fine-tuning instruction based on a first type of sample; the first type of sample includes a first search sample and a first type of word set containing at least one first keyword; wherein, when searching for resources based on the first search sample, the resources described by the first keyword at least satisfy part of the search requirements of the first search sample;
[0013] The second construction module is used to construct a second fine-tuning instruction based on the second type of samples; the second type of samples includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship;
[0014] The training module is used to cross-train the initial large model based on the first and second fine-tuning instructions to obtain the keyword generation model.
[0015] According to another aspect of this disclosure, a keyword generation apparatus is provided, comprising:
[0016] The third module is used to construct prompt information based on the target search information;
[0017] The generation module is used to input the prompt information into the keyword generation model and obtain at least one keyword output by the keyword generation model.
[0018] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0019] At least one processor; and
[0020] The memory is communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0022] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0023] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0024] In this embodiment of the disclosure, the model is trained by constructing multi-task fine-tuning instructions to make the keywords generated by the model more diverse, thereby further improving the quality of the model output, increasing user satisfaction, and reducing the cost of data acquisition and annotation.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0026] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0027] Figure 1 This is a flowchart illustrating a training method for a keyword generation model according to an embodiment of the present disclosure;
[0028] Figure 2 This is a schematic flowchart illustrating the process of constructing a first type of sample according to an embodiment of the present disclosure;
[0029] Figure 3 This is a flowchart illustrating the construction of a first fine-tuning instruction based on a first type of sample according to an embodiment of the present disclosure;
[0030] Figure 4 This is a schematic diagram of the process of training an initial large model based on a first fine-tuning instruction according to an embodiment of the present disclosure;
[0031] Figure 5 This is a schematic flowchart of obtaining a second type of sample according to an embodiment of the present disclosure;
[0032] Figure 6 This is a schematic diagram of the process of training an initial large model based on a second fine-tuning instruction according to an embodiment of the present disclosure;
[0033] Figure 7 This is a schematic flowchart of a keyword generation method according to an embodiment of the present disclosure;
[0034] Figure 8 This is a schematic diagram of the structure of a keyword generation model training device according to an embodiment of the present disclosure;
[0035] Figure 9 This is a schematic diagram of the structure of a keyword generation apparatus according to an embodiment of the present disclosure;
[0036] Figure 10 This is a block diagram of an electronic device used to implement the keyword generation model training method and / or keyword generation method of the embodiments of this disclosure. Detailed Implementation
[0037] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0038] The terms “first,” “second,” etc., used in this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0039] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0040] Artificial intelligence technology is gradually playing a role in various industries. In commercial search scenarios, large-scale models are also slowly coming into play.
[0041] There are three roles in a commercial search scenario: users, advertisers, and commercial search engines. Users refer to the end-users of the commercial search engine system. Users enter a query into the commercial search engine system to obtain information about relevant goods or services. Advertisers refer to businesses or other commercial operators who place advertisements through commercial search engines to promote their products or services. Commercial search engines act as a bridge between users and advertisers, providing information retrieval services to users while simultaneously offering commercial promotion services to advertisers. Matching queries with advertising keywords (also known as keywords) is a crucial task for commercial search engines.
[0042] In related technologies, large AI models generate corresponding responses based on user input. This technology is increasingly being used in various search engine systems to generate advertising keywords based on user-input queries. In this process, the relevance of the generated keywords is crucial. Higher relevance leads to more accurate resource recommendations to the user, better meeting their needs.
[0043] However, due to model capacity limitations, related technologies struggle to achieve general natural language understanding and reasoning capabilities, and the relevance and accuracy of generated keywords still need improvement. Furthermore, traditional fine-tuning training uses single downstream task data, with each downstream task trained independently, preventing information sharing. Maintaining multiple models for different downstream tasks inevitably increases costs.
[0044] In view of this, embodiments of this disclosure provide a training method for a keyword generation model. This method employs a model with a large number of parameters to achieve self-improvement in model capacity, resulting in a well-performing model across different tasks. Simultaneously, a large model is trained alternately based on fine-tuning instructions for different tasks, with a relevance task used to assist the model in generating keywords, thereby ensuring the relevance and accuracy of the generated keywords. This method is not only applicable to keyword generation in commercial scenarios but also to other fields requiring keyword generation; embodiments of this disclosure do not limit its application to these applications.
[0045] like Figure 1 The diagram shown is a flowchart illustrating the training method of the keyword generation model proposed in this embodiment, including:
[0046] S101, construct a first fine-tuning instruction based on the first type of samples; the first type of samples includes a first search sample and a first type of word set containing at least one first keyword; wherein, when searching for resources based on the first search sample, the resources described by the first keyword at least satisfy part of the search requirements of the first search sample.
[0047] In other words, based on the resource set recommended by the first search sample, at least one recommended resource described by the first keyword has been visited. Only such a first keyword can be combined with the corresponding first search sample to form the first class of samples. The first class of samples constructed in this way is confident and can be used to model the click task, so that the large model can learn good examples and learn to generate accurate keywords.
[0048] The first search sample can be the search query entered when performing a web search. For example, the search query might be "wireless headphones". The first term set contains at least one primary keyword that is related to the entered search query, and the corresponding resource should be easily clickable by the user, i.e., easily accessible to the user. For example, the first term set might include "Bluetooth headphones, noise-canceling headphones, wireless earbuds".
[0049] Among them, a first fine-tuning instruction is constructed based on the first type of samples. This instruction aims to optimize the ability of the large model to recognize and understand the first search samples, so that the large model can generate accurate keywords.
[0050] S102, construct a second fine-tuning instruction based on the second type of samples; the second type of samples includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship.
[0051] Similar to the first search sample, the second search sample can also be a search statement entered during a web search.
[0052] The preset association relationship requires that the second search sample and the second keyword be used to model the relevance task, so that the large model can learn how to generate keywords that are related to the search statement.
[0053] For example, the search term is "healthy eating plan". The second set of terms contains secondary keywords that are closely related to the user's search intent. For example, the secondary keywords could be "vegetable recipes, low-sugar foods, protein intake".
[0054] A second fine-tuning instruction is constructed based on the second type of samples, which focuses on enhancing the large model's learning and understanding of the relevance between search samples and keywords.
[0055] It is understood that the embodiments disclosed herein do not limit the generation order of the first fine-tuning instruction and the second fine-tuning instruction.
[0056] S103, based on the cross-training of the first and second fine-tuning instructions, the initial large model is trained to obtain the keyword generation model.
[0057] During training, the initial large model is trained by alternately using the first and second fine-tuning instructions. This ensures that the large model learns both the click task of the first fine-tuning instruction and the relevance task of the second fine-tuning instruction.
[0058] In this embodiment, by constructing a first fine-tuning instruction based on a first type of sample, the initial large model can be trained to generate keywords that meet click requirements, thus ensuring the accuracy of the generated keywords. Constructing fine-tuning instructions based on a second type of sample allows the large model to learn the correlation between keywords and search samples, enabling it to reason about and understand correlations. Therefore, through cross-training of the first and second fine-tuning instructions, the large model can learn both correlation modeling and reasoning, thereby assisting it in generating accurate keywords. The final trained keyword generation model can generate both relevant and accurate keywords.
[0059] Furthermore, the diversity of generated keywords is also very important. Therefore, in this embodiment of the disclosure, for the first search sample, multiple first keywords in the first category of words meet the diversity requirements.
[0060] In other words, for the same first search sample, its corresponding first-category term set not only includes the primary keywords that users are most likely to click, but these primary keywords are also diverse, adapting to the search needs of the first search sample from different angles, levels, or expressions. For example, different primary keywords in the first-category term set cover different core entities, but all can meet the search needs of the first search sample to a certain extent. For instance, if the search term in the first search sample is "berry," then the first-category term set includes multiple different core entities such as "berry tea," "strawberry," and "wild berry tea," thus satisfying the diversity of keywords.
[0061] In this embodiment, a keyword set that meets diversity requirements can improve the language understanding ability of a large model, enabling it to infer keywords that satisfy diversity needs. Therefore, when recommending resources, it can better serve users and improve user experience.
[0062] In this embodiment of the disclosure, in order to quickly and accurately collect a first type of sample from a large amount of data, this embodiment of the disclosure provides an implementation method for constructing the first type of sample, such as... Figure 2 The following content is shown:
[0063] S201, sample the set of search statements for the target object to obtain the first search sample.
[0064] During implementation, provided it complies with relevant laws and regulations, does not violate public order and good morals, and with the permission of the target audience, the search queries of different target audiences and the recommended resources accessed for each search query can be recorded. To avoid sensitive information, the identifiers of the target audiences will be anonymized. Each target audience can be distinguished by a unique identifier.
[0065] The first search sample can be obtained by randomly sampling a set of search statements for different target objects. Taking a commercial search engine as an example, it contains a large number of search records. For instance, m search statements can be randomly selected from the search records of the most recent month as the first search sample. The specific number of the first search samples required can be determined based on the required sample size, such as tens of millions or millions of samples.
[0066] Since each first search sample has recommended corresponding resources, and some of these resources have been accessed, in S202, the keywords corresponding to each accessed resource in the set of accessed resources of the first search sample can be queried to obtain the first preliminary selection set.
[0067] Continuing with the commercial search engine example, examine the m recommended resources in the first search sample and identify the resource identifiers that were actually clicked and viewed. For each accessed resource, extract its corresponding keywords. For example, if the search query is "summer breathable sports shoes," and the target user clicks on a product after executing the search query, extract keywords such as "summer sports shoes," "breathable sports shoes," etc. This yields a first preliminary set containing multiple keywords.
[0068] S203, From the first preliminary selection set, select the first candidate keyword set. That is, the first preliminary selection set needs to be filtered to find high-quality keywords. The filtering of the first candidate keyword set can be implemented as follows:
[0069] Step A1: Determine the access frequency of each initial term in the first initial selection set, and the degree of relevance between the initial term and the first search sample.
[0070] In this context, the access frequency of each initial term in the first initial selection set refers to its popularity or fulfillment of current needs within a certain timeframe. Relevance is used to assess the correlation between the initial term and the first search sample. In this embodiment, the natural language understanding capability of a neural network can be used to analyze the correlation between the first search sample and the initial term. A high relevance indicates that the initial term and the first search sample are closely related.
[0071] Step A2: Select the first initial words that meet the preset requirements in terms of access frequency and relevance to construct the first candidate word set.
[0072] During implementation, the preset requirements can be tailored to specific application scenarios and needs, setting different filtering criteria. For example, the threshold for the number of visits can be set to a specific value, or the relevance level can need to reach a certain level. Another example is sorting by visit count and relevance level, filtering the top-ranked initial keywords. Understandably, only the top-ranked initial keywords that meet the visit count and relevance requirements will be selected into the first candidate keyword set to improve the quality of the first type of sample construction.
[0073] In this embodiment of the disclosure, by determining the access frequency of each first preliminary word in the first preliminary selection set and its relevance to the first search sample, the words most likely related to the search intent can be selected from a large number of first preliminary words, thus providing more accurate and useful first-type samples. By combining access frequency and relevance, the quality and effectiveness of the first-type samples used for training can be improved, thereby enhancing the generation quality of the keyword generation model.
[0074] S204, Construct the first type of word set based on the first candidate word set.
[0075] In this embodiment of the disclosure, the first candidate word set can be used as the first type of word set. When there are multiple target objects, each target object may be sampled from the same first search sample, thus resulting in a separate first candidate word set for each target object. In this case, the union of these multiple first candidate word sets can be used as the first type of word set.
[0076] In cases with a large amount of data, in order to remove redundant information and allow the large model to focus on learning useful knowledge, in this embodiment of the disclosure, when a first search sample is sampled from a set of search statements for multiple target objects, a clustering operation is performed on the first candidate word set obtained for each target object to obtain a first class of word set.
[0077] Clustering can group candidate words related to the first search sample together, extracting a representative set of keywords from search statements of multiple target objects, thereby improving the quality of the constructed first type of sample and thus improving the training efficiency of the model.
[0078] During implementation, clustering can be performed based on the following methods:
[0079] Step B1 involves merging the first candidate word sets obtained from each target object to obtain the union set to be processed.
[0080] In practice, all the first candidate words are gathered into a single set to form a union, which may contain duplicate or very similar keywords.
[0081] Step B2: From the set to be processed, multiple first candidate words that meet the merging conditions are merged into a first keyword to obtain the first type of word set. The merging conditions include: identical entities and semantic feature similarity greater than a preset similarity threshold.
[0082] Among these requirements, "entity similarity" means that the first candidate words for merging must refer to the same entity. For example, the keywords "strawberry" and "small strawberry" both refer to "strawberry," thus satisfying the entity similarity condition. Semantic feature similarity greater than a preset similarity threshold means that the first candidate words for merging must be semantically close enough. For example, "strawberry price" and "strawberry cost" are semantically identical. However, "strawberry cultivation knowledge" and "strawberry price" have completely different semantics and do not need to be merged.
[0083] During implementation, the keywords to be processed and clustered are analyzed and compared according to the above merging conditions. For the first candidate words that meet the merging conditions, they are merged into a specific word set, thus completing the clustering operation. The first keyword obtained by merging multiple first candidate words that meet the merging conditions can be the center of that category, or a first candidate word randomly selected from that category.
[0084] In this embodiment of the disclosure, by merging candidate words with the same entity and high semantic similarity into a single keyword, data redundancy can be reduced and sample quality improved, thereby enhancing the training effect of the model.
[0085] In summary, in this embodiment of the present disclosure, by sampling the set of search statements for the target object to obtain the first search sample, and by filtering and constructing a word set based on the keywords in the accessed resource set, it can be ensured that the keywords in the training data are closely related to the first search sample, so that the first keyword can not only complete click modeling, but also improve the sample quality, thereby improving the efficiency of model training.
[0086] In this embodiment of the disclosure, the implementation of constructing the first fine-tuning instruction based on the first type of sample can be as follows: Figure 3 As shown, it includes:
[0087] S301, sort the first keywords in the first category of words according to the preset rules to obtain the first word sequence; the preset rules require that the first keywords with higher access probability should be sorted first.
[0088] In this embodiment of the disclosure, each keyword in the first category of word sets is sorted to obtain an ordered list of keywords, i.e., the first word sequence. For example, if the first search sample is "berry", the keywords in the first category of word sets are sorted as follows: "small strawberry, berry tea, how much does berry tea cost per pound, wild berry tea...". The sorting is based on a preset rule, which is based on the access probability of the resource corresponding to the keyword. This access probability can be represented by the click-through rate. Access probability can also refer to the probability that a user selects a certain keyword during the search or browsing process. Usually, this probability can be obtained through historical data analysis, such as statistical indicators such as user click-through rate and consumption frequency.
[0089] S302, Generate a first fine-tuning instruction containing the first search sample, the first word sequence, and the first keyword generation requirements.
[0090] One of the key requirements for keyword generation is to sort the keywords according to their affordability. Taking the first search sample as "berry" as an example, the keyword list generated according to the affordability of the keywords might be "how much does berry tea cost per pound, wild berry tea, berry tea, small strawberries...".
[0091] In this embodiment, the first keywords in the first category of words are sorted according to preset rules, which allows keywords with higher access probability to be ranked higher. This makes the model pay more attention to these important keywords during training, thereby improving the model's understanding and prediction accuracy of search samples. By generating a first search sample, a first word sequence, and a first fine-tuning instruction that meets the keyword generation requirements, strong support is provided for further training of the model.
[0092] In this embodiment of the disclosure, the specific implementation of training the initial large model based on the first fine-tuning instruction is as follows: Figure 4 As shown:
[0093] S401, Obtain the first set of predicted words generated by the initial large model for the first search sample.
[0094] The initial large model is used to process the first search sample, generating a set of predicted terms. This first set of predicted terms contains a series of keywords that the large model believes are most likely to be accessed based on the search sample.
[0095] S402, adjust the model parameters of the initial large model based on the loss between the first predicted word set and the first type of word set.
[0096] By calculating the loss value between the first predicted word set and the first type of word set, the model parameters are adjusted using an optimization algorithm based on the loss value to better match the first type of word set.
[0097] In this embodiment of the disclosure, by calculating the loss between the first prediction word set and the first type of word set, the accuracy of the model prediction can be quantitatively evaluated, and the model parameters can be adjusted to minimize the loss, improve the generation capability of large models, and reduce the fluctuation and uncertainty of the model when facing new data.
[0098] In addition to the keyword generation model trained using the first fine-tuning instructions based on the first type of samples as described above, the method also includes training a keyword generation model using the second type of samples based on the second type of samples. The implementation of obtaining the second type of samples is as follows: Figure 5 As shown:
[0099] S501, based on the recall method, recalls multiple recall terms of the second search sample in the second type of sample.
[0100] S502, based on the preset association relationship, select at least one recall term from multiple recall terms as the second keyword.
[0101] For example, using a recall method may result in the recall phase retrieving a massive number of keywords. Some of these keywords are relevant to the second search sample, while others are not.
[0102] To better train the model, suitable recall terms need to be selected based on the pre-defined association requirements to construct the second type of samples. In practice, second keywords that meet the relevance requirements of the second search samples need to be selected. This selection process can be done manually or by leveraging the understanding capabilities of a large language model.
[0103] S503, construct the second type of sample based on the second keyword and the second search sample.
[0104] In this embodiment, a massive number of recall terms are retrieved using a recall method, and keywords with a preset association relationship with the second search sample are selected based on preset association relationships. This approach can provide more valuable data for model training and enrich the information content of the dataset. Constructing a second type of sample using multiple recall terms and the second search sample not only provides a learning pattern for association relationships but also increases the diversity of training data, helping the model learn more comprehensive features and patterns, thereby improving the model's generalization ability and enabling it to better adapt to different search scenarios and user needs. Filtering relevant recall terms through preset association relationships can enhance the model's semantic understanding of the search sample, enabling the model to more accurately grasp the relevance relationship between the second keyword and the second search sample, and improving the model's semantic representation ability.
[0105] Since training a large model typically requires a large number of data samples, in this embodiment of the disclosure, the preset association relationship requires that the second type of training sample set used to train the initial large model includes positive sample samples that satisfy the preset association relationship and negative sample samples that do not satisfy the preset association relationship, for training different aspects of the model.
[0106] Positive examples that satisfy the preset association relationship refer to the second search sample and the second keyword in the second type of samples. For example, the search sample is "a pair of professional running shoes" and the keywords are "cushioning, breathability, suitable for long-distance running".
[0107] Negative examples that do not satisfy the preset association relationship refer to samples in the second category where there is no relevance between the second search sample and the second keyword. For example, the search sample is "a pair of professional running shoes," and the keywords are "business style, formal occasion, leather material." Similarly, "food" and "dog food" are not related.
[0108] In this embodiment of the disclosure, positive and negative examples enhance the model's discrimination and understanding capabilities, preventing overfitting. By combining samples that satisfy and do not satisfy preset correlations, the model can ensure that it captures useful relevant knowledge while also identifying irrelevant noise or anomalies, improving its generalization ability and enabling it to generate more accurate relevant keywords.
[0109] In this embodiment of the disclosure, the specific implementation of training the initial large model based on the second fine-tuning instruction is as follows: Figure 6 As shown:
[0110] S601, Obtain the correlation analysis results between the second search sample predicted by the initial large model and the second keyword.
[0111] During implementation, the initial model evaluates the relationship between the search sample and the second keyword, providing a relevance score. This score reflects whether the model considers the second search sample and the second keyword to be the same entity or to be semantically similar.
[0112] S602, adjust the model parameters of the initial large model based on the loss between the correlation analysis results and the true value.
[0113] The ground truth value indicates whether the second type of search sample and its corresponding second keyword are related. The loss value is calculated by comparing the relevance score predicted by the model with the ground truth value. Based on the loss value, an optimization algorithm is used to adjust the model parameters so that the model can learn the meaning and paradigm of the relevance between keywords.
[0114] In this embodiment, the accuracy of the model prediction can be verified by obtaining the relevance analysis results of the initial large model prediction and comparing them with the true values. Based on this, adjusting the model parameters helps improve the model's understanding of the relevance between the second search sample and the second keyword. Adjusting the model parameters according to the loss between the relevance analysis results and the true values makes the model more adaptable to the keyword generation task, thereby improving the model's ability to generate relevant keywords.
[0115] To further improve the model's ability to understand natural language and its reasoning and expression capabilities in complex situations, the keyword prediction model in this embodiment can employ a large model with hundreds of billions of parameters. For example, a GLM (General Language Model) pre-trained model. This large model with hundreds of billions of parameters, leveraging its robust parameter base, can reason and express more complex logic and knowledge, thus possessing better natural language understanding and reasoning capabilities and improving the quality of the keywords generated by the model.
[0116] Based on the same technical concept, this disclosure also provides a keyword generation method, which is applicable to the aforementioned trained keyword generation model. For example... Figure 7 As shown, the following steps may be included:
[0117] S701, constructs prompt information based on target search information.
[0118] For example, based on a given target search information, the advertising system generates different auction terms in order of purchasing power. Target search information: "berry", matching auction terms are: "How much is berry tea per pound, wild berry tea, berry tea, small strawberries...".
[0119] S702, input the prompt information into the keyword generation model to obtain at least one keyword output by the keyword generation model.
[0120] In this embodiment of the disclosure, by constructing prompt information based on target search information and inputting this prompt information into a keyword generation model to obtain keywords, the search engine can understand the search intent more quickly, thereby improving the accuracy and relevance of keyword generation.
[0121] This embodiment of the keyword generation model can generate multiple keywords with good relevance and accuracy in batches. When the model is trained based on training samples that meet the diversity requirements, the keyword generation model relies on its learned knowledge and reasoning ability to generate accurate keywords that meet the requirements of diversity and relevance.
[0122] Based on the same technical concept, this disclosure also provides a training device 800 for a keyword generation model, such as... Figure 8 As shown, it includes:
[0123] The first construction module 801 is used to construct a first fine-tuning instruction based on a first type of sample; the first type of sample includes a first search sample and a first type of word set containing at least one first keyword; wherein, when searching for resources based on the first search sample, the resources described by the first keyword at least satisfy part of the search requirements of the first search sample;
[0124] The second construction module 802 is used to construct a second fine-tuning instruction based on the second type of samples; the second type of samples includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship;
[0125] Training module 803 is used to cross-train the initial large model based on the first fine-tuning instruction and the second fine-tuning instruction to obtain the keyword generation model.
[0126] In some embodiments, for a first search sample, multiple first keywords in a first set of terms meet the diversity requirement.
[0127] In some embodiments, a first sampling module is further included, comprising:
[0128] The sampling unit is used to sample the set of search statements for the target object to obtain the first search sample;
[0129] The query unit is used to query the keywords corresponding to each accessed resource in the accessed resource set of the first search sample to obtain the first preliminary selection set;
[0130] The filtering unit is used to filter out the first set of candidate words from the first initial selection set;
[0131] The building unit is used to construct the first type of word set based on the first candidate word set.
[0132] In some embodiments, the query unit is specifically used for:
[0133] Determine the access frequency of each initial term in the first initial selection set, and the degree of relevance between the initial terms and the first search sample;
[0134] Select the first initial words that meet the preset requirements in terms of access frequency and relevance to construct the first candidate word set.
[0135] In some embodiments, the building unit is specifically used for:
[0136] When a first search sample is obtained from a set of search statements for multiple target objects, a clustering operation is performed on the first candidate word set obtained for each target object to obtain a first class of word set.
[0137] In some embodiments, the building unit is specifically used for:
[0138] The first candidate word sets obtained from each target object are merged to obtain the union set to be processed;
[0139] From the unprocessed and collected sets, multiple first candidate words that meet the merging conditions are merged into a first keyword to obtain the first type of word set;
[0140] The merging conditions include: the entities are the same and the semantic feature similarity is greater than the preset similarity threshold.
[0141] In some embodiments, the first building module is specifically used for:
[0142] The first keywords in the first category of words are sorted according to a preset rule to obtain the first word sequence; the preset rule requires that the first keywords with higher access probability be sorted first.
[0143] Generate a first fine-tuning instruction that includes the first search sample, the first word sequence, and the keyword generation requirements.
[0144] In some embodiments, the training module is specifically used for:
[0145] Obtain the first set of predicted words generated by the initial large model for the first search sample;
[0146] The model parameters of the initial large model are adjusted based on the loss between the first predicted word set and the first type of word set.
[0147] In some embodiments, a second sampling module is further included, for:
[0148] Based on the recall method, multiple recall words are recalled from the second search sample in the second type of sample;
[0149] Based on the preset association relationship, at least one recall word that satisfies the preset association relationship with the second search sample is selected from multiple recall words as the second keyword;
[0150] The second type of sample is constructed based on the second keyword and the second search sample.
[0151] In some embodiments, the preset association relationship requires that the second type of training sample set for training the initial large model includes positive samples that satisfy the preset association relationship and negative samples that do not satisfy the preset association relationship.
[0152] In some embodiments, the training module is specifically used for:
[0153] Obtain the correlation analysis results between the second search sample predicted by the initial large model and the second keyword;
[0154] Based on the loss between the correlation analysis results and the true values, the model parameters of the initial large model are adjusted.
[0155] In some embodiments, the keyword prediction model is a large model containing tens of billions of parameters.
[0156] Based on the same technical concept, this disclosure also provides a keyword generation device 900, such as... Figure 9 As shown, it includes:
[0157] The third construction module 901 is used to construct prompt information based on the target search information;
[0158] The generation module 902 is used to input the prompt information into the keyword generation model and obtain at least one keyword output by the keyword generation model.
[0159] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0160] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0161] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0162] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0163] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0164] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as keyword generation model training methods and / or keyword generation methods. For example, in some embodiments, the keyword generation model training methods and / or keyword generation methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the keyword generation model training methods and / or keyword generation methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured in any other suitable manner (e.g., by means of firmware) as a keyword generation model training method and / or a keyword generation method.
[0165] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0166] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0167] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0169] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0170] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0171] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0172] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a keyword generation model, comprising: constructing a first fine-tuning instruction based on a first type of sample; the first type of sample includes a first search sample and a first type of word set containing at least one first keyword; wherein, in the case of searching for resources based on the first search sample, the resources expressed by the first keyword at least meet part of the search demand of the first search sample; constructing a second fine-tuning instruction based on a second type of sample; the second type of sample includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship; cross-training an initial large model based on the first fine-tuning instruction and the second fine-tuning instruction to obtain a keyword generation model.
2. The method of claim 1, wherein, For the first search sample, multiple first keywords in the first type of word set meet the diversity requirement.
3. The method of claim 1 or 2, further comprising obtaining the first type of sample based on the following method: sampling a search sentence set of a target object to obtain the first search sample; querying the first search sample in the accessed resource set, the keywords corresponding to each accessed resource, to obtain a first preliminary selection set; screening a first candidate word set from the first preliminary selection set; constructing the first type of word set based on the first candidate word set.
4. The method of claim 3, wherein, The first candidate word set is screened from the first preliminary selection set, comprising: determining the number of accesses of each first preliminary word in the first preliminary selection set and the relevance degree of the first preliminary word to the first search sample; screening first preliminary words that meet the preset requirements in terms of access frequency and relevance degree to construct the first candidate word set.
5. The method of claim 3, wherein, The first type of word set is constructed based on the first candidate word set, comprising: In the case of sampling the first search sample from the search sentence set of multiple target objects, clustering the first candidate word set obtained for each target object to obtain the first type of word set.
6. The method of claim 5, wherein, The first candidate word set obtained for each target object is clustered to obtain the first type of word set, comprising: merge the first candidate word set obtained from each target object to obtain a to-be-processed union set; merge multiple first candidate words that meet the merging condition from the to-be-processed union set into one first keyword to obtain the first type of word set; The merging condition includes: the same entity and the semantic feature similarity is greater than a preset similarity threshold.
7. The method of claim 1, wherein, The first fine-tuning instruction is constructed based on the first type of sample, comprising: sorting each first keyword in the first type of word set according to a preset rule to obtain a first word sequence; the preset rule requires that the first keyword with higher access possibility is sorted earlier; generating a first fine-tuning instruction containing the first search sample, the first word sequence, and the first keyword generation requirement.
8. The method of any one of claims 1-7, wherein, Training the initial large model based on the first fine-tuning instruction, comprising: obtaining a first predicted word set generated by the initial large model for the first search sample; adjusting the model parameters of the initial large model based on the loss between the first predicted word set and the first type of word set.
9. The method of any one of claims 1-8, further comprising obtaining the second type of sample based on the following method: based on a recall method, recalling a plurality of recall words of the second search sample in the second type of sample; based on the preset association relationship, screening at least one recall word from the plurality of recall words as the second keyword; based on the second keyword and the second search sample, constructing the second type of sample.
10. The method of claim 1, wherein, The preset association relationship requires that the second type of training sample for training the initial large model includes positive example samples satisfying the preset association relationship and negative example samples not satisfying the preset association relationship.
11. The method of any one of claims 1-10, wherein, Training the initial large model based on the second fine-tuning instruction, comprising: obtaining a correlation analysis result between the second search sample and the second keyword predicted by the initial large model; based on the loss between the correlation analysis result and the true value, adjusting the model parameters of the initial large model.
12. The method of any one of claims 1-11, wherein, The keyword prediction model is a large model containing hundreds of millions of parameters.
13. A keyword generation method applied to a keyword generation model trained by the method of any one of claims 1-12, comprising: based on target search information, constructing prompt information; inputting the prompt information into the keyword generation model to obtain at least one keyword output by the keyword generation model.
14. A training device of a keyword generation model, comprising: a first construction module configured to construct a first fine-tuning instruction based on a first type of sample; The first type of sample includes a first search sample and a first type of word set containing at least one first keyword; wherein, under the condition of searching resources based on the first search sample, the resources represented by the first keyword at least satisfy part of the search demand of the first search sample; a second construction module configured to construct a second fine-tuning instruction based on a second type of sample; the second type of sample includes a second search sample and a second type of word set containing at least one second keyword, and the second search sample and the second keyword are constructed based on a preset association relationship; a training module configured to cross-train an initial large model based on the first fine-tuning instruction and the second fine-tuning instruction to obtain a keyword generation model.
15. The apparatus of claim 14, wherein, For the first search sample, a plurality of first keywords in the first type of word set satisfy the diversity requirement.
16. The device of claim 14 or 15, further comprising a first sampling module, comprising: a sampling unit configured to sample a search sentence set of a target object to obtain the first search sample; a query unit configured to query keywords corresponding to each accessed resource in an accessed resource set of the first search sample to obtain a first preliminary selection set; a screening unit configured to screen a first candidate word set from the first preliminary selection set; a construction unit configured to construct the first type of word set based on the first candidate word set.
17. The apparatus of claim 16, wherein, The query unit is specifically configured to: determine the number of accesses of each first preliminary selection word in the first preliminary selection set and the degree of correlation between the first preliminary selection word and the first search sample; The first preliminary words meeting preset requirements of the number of access times and the degree of relevance are screened to construct the first candidate word set.
18. The apparatus of claim 16, wherein, The constructing unit is specifically configured to: In a case where the first search samples are sampled from the search sentence set of the plurality of target objects, the first candidate word set obtained for each target object is subjected to a clustering operation to obtain the first category word set.
19. The apparatus of claim 18, wherein, The constructing unit is specifically configured to: The first candidate word sets obtained from the target objects are merged to obtain a to-be-processed union set; From the to-be-processed union set, a plurality of first candidate words meeting a merging condition are merged into one first keyword to obtain the first category word set; The merging condition includes that the entities are the same and the semantic feature similarity is greater than a preset similarity threshold.
20. The apparatus of claim 14, wherein, The first constructing module is specifically configured to: The first keywords in the first category word set are sorted according to a preset rule to obtain a first word sequence; the preset rule requires that the higher the access possibility of a first keyword, the higher the sorting position of the first keyword; A first fine-tuning instruction meeting a keyword generation requirement is generated, and the first fine-tuning instruction includes the first search sample, the first word sequence, and the first keyword.
21. The apparatus of any of claims 14-20, wherein, The training module is specifically configured to: Obtain a first prediction word set generated by the initial large model for the first search sample; Based on a loss between the first prediction word set and the first category word set, adjust the model parameters of the initial large model.
22. The apparatus according to any one of claims 14-21, further comprising a second sampling module configured to: Based on a recall method, recall a plurality of recall words of a second search sample in the second category sample; Based on the preset association relationship, screen at least one recall word meeting the preset association relationship with the second search sample from the plurality of recall words as the second keyword; Based on the second keyword and the second search sample, construct the second category sample.
23. The apparatus of claim 14, wherein, The preset association relationship requires that the second category training sample set for training the initial large model includes a positive example sample meeting the preset association relationship and a negative example sample not meeting the preset association relationship.
24. The apparatus of any of claims 14-23, wherein, The training module is specifically configured to: Obtain a relevance analysis result between the second search sample and the second keyword predicted by the initial large model; Based on a loss between the relevance analysis result and a true value, adjust the model parameters of the initial large model.
25. The apparatus of any of claims 14-24, wherein, The keyword prediction model is a large model including hundreds of millions of parameters.
26. A keyword generation apparatus, applied to a keyword generation model trained by the apparatus of any one of claims 14-25, comprising: A third constructing module configured to construct prompt information based on target search information; A generating module configured to input the prompt information into the keyword generation model to obtain at least one keyword output by the keyword generation model.
27. An electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
28. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1-13.
29. A computer program product comprising computer program which, when executed by a processor, implements the method according to any one of claims 1-13.
Citation Information
Patent Citations
Search keyword recommendation model generation method and keyword recommendation method and device
CN111324804A
Text classification model training sample generation method and device and electronic equipment
CN111831821A
Vertical search method and device, electronic equipment and storage medium
CN113821711A
Semantic retrieval network training method and device, electronic equipment and storage medium
CN113988157A
Keyword combination generation model training method and device
CN114003706A
Cited By
A search keyword generation model training method
CN122412963A