Method, System, Storage Medium and Computer Device for Judging Similar Search Terms
By building a similar search term judgment system, using co-click data and deep learning models, the problem of how to obtain similar search terms is solved, and the accuracy of advertising recall and platform revenue are improved.
Patent Information
- Application Number
- CN202010654923.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-07-09
AI Technical Summary
In the prior art, how to better obtain search terms with high similarity to user search terms to improve the accuracy of advertising recall and platform revenue is a challenge.
By obtaining the co-click data, using the Bert classification model and Bilstm network structure, combining two rounds of fine-tuning training strategies, a similar search term judgment system is built to judge the similarity of search terms and form a training set, and the final model is trained to judge the similarity of search terms.
It realizes more accurately displaying ads with similar search terms, improving advertising promotion performance and platform revenue.
Smart Images

Figure CN113918840B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prediction of similar search terms, and particularly to a method, a system, a storage medium and a computer device for judging similar search terms. Background Art
[0002] In the prior art, a platform will display advertisements to a user according to the user's search term. In order to reasonably display advertisements to the user, before finally displaying an advertisement to the user, advertisement recall can be performed according to the user's search term; so that the advertisement finally displayed to the user is not only an advertisement that exactly matches the user's search term, but also an advertisement of a similar search term with a high similarity to the search term can be displayed to the user.
[0003] However, how to better obtain similar search terms in advertisement recall is a technical problem to be solved.
[0004] In summary, the prior art obviously has inconveniences and defects in actual use, so it is necessary to improve. Summary of the Invention
[0005] Aiming at the above defects, the purpose of the present invention is to provide a method, a system, a storage medium and a computer device for judging similar search terms, which can obtain search terms with a high similarity to the user's search term, promote the advertisements of advertisers, and can also increase the income of the platform.
[0006] To achieve the above purpose, the present invention provides a method for judging similar search terms, including:
[0007] Obtain a plurality of co-click data, each of the co-click data including a search term, a url clicked under the search term, and the number of clicks on the url under the search term;
[0008] According to the url clicked under the search term and the number of clicks on the url under the search term in two of the co-click data, judge the similarity of the two search terms, and classify the two search terms with high similarity as a positive sample; classify the two search terms with low similarity as a negative sample; a plurality of the positive samples and a plurality of the negative samples form a training set;
[0009] Use the training set as input to train a semantic pre-training model to obtain a final model.
[0010] According to the method for judging similar search terms, the step of judging the similarity of two search terms based on the URLs clicked under the search terms and the number of clicks on the URLs under the search terms in the two co-click data, and classifying two search terms with high similarity as a positive sample; and classifying two search terms with low similarity as a negative sample includes:
[0011] Score the search terms in a number of the co-click data according to the URLs clicked under the search terms and the number of clicks on the URLs under the search terms.
[0012] The search term with a score greater than the first threshold is the first search term, and the search term with a score less than or equal to the first threshold is the second search term.
[0013] Classify two first search terms with the same URL in a number of the co-click data as a positive sample to obtain a number of positive samples; classify two second search terms with different URLs as a negative sample to obtain a number of negative samples.
[0014] According to the method for judging similar search terms, the calculation formula for scoring is:
[0015] S u1q1 =C u1q1 -avg(C u1 )
[0016] Where, assuming the search term is query1, the URL clicked under the search term is url1, S u1q1 represents the score of query1, C u1q1 represents the number of clicks on url1 when searching for query1, where n refers to the number of times url1 has been clicked when searching for n search terms.
[0017] According to the method for judging similar search terms, the step of obtaining a number of co-click data includes: obtaining the number of co-click data from the click records of users.
[0018] According to the method for judging similar search terms, the semantic pre-training model is a Bert classification model.
[0019] According to the method for judging similar search terms, the step of taking a number of the training sets as inputs, training the semantic pre-training model, and obtaining the final model includes:
[0020] Input the training set into the Bert classification model for training to obtain the first model;
[0021] Add a layer of Bilstm network structure to the first model to obtain a second model;
[0022] Obtain a data set, use the data set as input in the second model for training, and adopt a training strategy of two rounds of fine-tuning to obtain the final model.
[0023] According to the above-mentioned method for judging similar search terms, the steps of obtaining a data set, using the data set as input in the second model for training, and adopting a training strategy of two rounds of fine-tuning to obtain the final model include:
[0024] Input the first initial data sets of the first quantity into the first model respectively, and the first model outputs, label the first initial data sets as the first positive samples or the first negative samples to obtain the first data set;
[0025] Manually label the second initial data sets of the second quantity, and manually label the second initial data sets as the second positive samples or the second negative samples to obtain the second data set;
[0026] Use the first data set as input in the second model for training and perform the first round of fine-tuning;
[0027] Use the second data set as input in the second model for training and perform the second round of fine-tuning to obtain the final model.
[0028] To achieve the above object, the present invention also provides a system for judging similar search terms, including:
[0029] A co-click data acquisition module, configured to acquire a number of co-click data, each of the co-click data including a search term, a url clicked under the search term, and the number of clicks on the url under the search term;
[0030] A training set acquisition module, configured to judge the similarity of two search terms according to the url clicked under the search term and the number of clicks on the url under the search term in two of the co-click data, and classify two search terms with high similarity as a positive sample; classify two search terms with low similarity as a negative sample; a number of the positive samples and a number of the negative samples form a training set;
[0031] A final model acquisition module, configured to use the training set as input to train a semantic pre-training model to obtain a final model.
[0032] To achieve the above object, the present invention also provides a storage medium for storing a computer program for executing the judgment method of any one of the above similar search terms.
[0033] To achieve the above object, the present invention also provides a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the judgment method of the similar search terms described in any one of the above is implemented.
[0034] The present invention obtains a number of co-click data, each of the co-click data including a search term, a URL clicked under the search term, and the number of clicks on the URL under the search term; determines the similarity of two search terms according to the URL clicked under the search term and the number of clicks on the URL under the search term in two of the co-click data, and classifies two search terms with high similarity as a positive sample; classifies two search terms with low similarity as a negative sample; a number of the positive samples and a number of the negative samples form a training set; uses the training set as an input to train a semantic pre-training model to obtain a final model. Inputs the user's search term and the search term provided by the advertiser into the final model, and the final model can determine whether the above two search terms are similar. If they are similar, the advertisement related to the search term provided by the advertiser is also displayed to the user, thereby realizing the promotion of the advertiser's advertisement and also increasing the revenue of the platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is one of the schematic diagrams of the judgment system for similar search terms in the preferred embodiment of the present invention;
[0036] Figure 2 is the second schematic diagram of the judgment system for similar search terms in the preferred embodiment of the present invention;
[0037] Figure 3 is the flowchart of the judgment method for similar search terms in the preferred embodiment of the present invention;
[0038] Figure 4 is the schematic structural diagram of the computer device provided by the present invention.
[0039] Figure 5 is the schematic diagram of the final model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] It should be noted that the references to "one embodiment", "embodiment", "exemplary embodiment", etc. in this specification mean that the described embodiment may include specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures, or characteristics with an embodiment, whether or not explicitly described, it has been shown that combining such features, structures, or characteristics with other embodiments is within the knowledge of those skilled in the art.
[0042] In addition, in the specification and subsequent claims, certain terms are used to refer to specific components or parts. Those of ordinary skill in the art should understand that a manufacturer may use different nouns or terms to refer to the same component or part. This specification and subsequent claims do not use the difference in name as a way to distinguish components or parts, but use the difference in function of components or parts as the criterion for distinction. The terms "comprising" and "including" mentioned throughout the specification and subsequent claims are open-ended terms and should be interpreted as "including but not limited to". In addition, the term "connected" herein includes any direct and indirect electrical connection means. Indirect electrical connection means include connection through other devices.
[0043] See Figures 1 to 2 , in the first embodiment of the present invention, a judgment system 100 for similar search terms is provided, including:
[0044] A co-click data acquisition module 10, configured to acquire a plurality of co-click data, each of the co-click data including a search term, a URL clicked under the search term, and the number of clicks on the URL under the search term;
[0045] A training set acquisition module 20, configured to judge the similarity of two search terms according to the URL clicked under the search term and the number of clicks on the URL under the search term in two of the co-click data, classify two search terms with high similarity as a positive sample; classify two search terms with low similarity as a negative sample; a plurality of the positive samples and a plurality of the negative samples form a training set;
[0046] A final model acquisition module 30, configured to use the training set as an input to train a semantic pre-training model to obtain a final model.
[0047] In this embodiment, the user searches for the desired information on the platform through search terms. During the advertisement recall process, the final model of the system 100 can determine whether the search terms of the user are similar to the search terms related to the advertisements provided by the advertisers. If they are similar, the advertisements provided by the advertiser can be displayed to the user, thereby increasing the revenue of the platform. The acquisition of the final model requires predetermined data to train the relevant model. Specifically, a training set for training the semantic pre-training model is obtained by acquiring a number of co-click data; preferably, the co-click data acquisition module 10 acquires the number of co-click data from the user's click records. The co-click data can be obtained from the logs of the pc browser. The logs will record the search terms (queries) entered by the user in the browser and return the links (urls) clicked by the user in the returned search results. Training the semantic pre-training model requires positive samples and negative samples. If the similarity between two search terms is high, the two search terms are considered a positive sample; if the similarity between two search terms is low, the two search terms are considered a negative sample. The similarity between two search terms is judged by the urls and click counts of the two search terms in the co-click data. It can be considered that the higher the similarity between two search terms with the same url and high click counts in the co-click data.
[0048] Examples of a number of the co-click data:
[0049]
[0050] Among them, query represents the search term, and Ctr represents the click count.
[0051] See Figure 2 , in the second embodiment of the present invention, the training set acquisition module 20 includes:
[0052] The scoring sub-module 21 is used to score the search terms in a number of the co-click data according to the urls clicked under the search terms and the click counts of clicking the urls under the search terms; the search terms with a score greater than the first threshold are the first search terms, and the search terms with a score less than or equal to the first threshold are the second search terms;
[0053] The classification sub-module 22 is used to classify two first search terms with the same url in a number of the co-click data into a positive sample to obtain a number of the positive samples; classify two second search terms with different urls into a negative sample to obtain a number of the negative samples.
[0054] In this embodiment, the scoring sub-module 21 scores the search terms in the co-click data, and divides the search terms in several pieces of the co-click data into first search terms or second search terms according to the scores. If the URLs of two first search terms in the co-click data are the same, the above two first search terms form a positive sample. That is, under the same URL, two first search terms with high scores have high URL similarity and are used as a positive sample. Since the search terms with the same URL in the co-click data may have a certain degree of similarity, and the search terms with different URLs in the co-click data are dissimilar, therefore, two search terms with low scores and different URLs in the co-click data are used as negative samples.
[0055] In the third embodiment of the present invention, the calculation formula for the scoring sub-module 21 to perform scoring is:
[0056] S u1q1 = C u1q1 - avg(C u1 )
[0057] Wherein, assuming the search term is query1, the URL clicked under the search term is url1, and S u1q1 represents the score of query1, and C u1q1 represents the number of clicks on url1 when searching for query1. Wherein, n refers to that url1 has been clicked when searching for n search terms.
[0058] In this embodiment, generally, a URL will be clicked during the search for multiple search terms, and when searching for one search term, multiple URLs may be clicked. It is difficult to accurately determine whether a search term belongs to a positive sample or a negative sample only by the number of clicks. Through the above calculation formula, not only the relationship between the current URL and the search term is considered, but also the performance of the current search term among all URLs is included.
[0059] In the fourth embodiment of the present invention, the semantic pre-training model is a Bert classification model.
[0060] In this embodiment, the deep Transformer encoder of the Bert classification model has strong semantic representation ability. In this embodiment, the Bert classification model has better effects than using the DSSM model.
[0061] See Figure 2 , in the fifth embodiment of the present invention, the final model acquisition module 30 includes:
[0062] The first model acquisition sub-module 31 is used to input the training set into the Bert classification model for training to obtain a first model;
[0063] The second model acquisition sub-module 32 is used to add a layer of Bilstm network structure to the first model to obtain a second model;
[0064] The data set acquisition sub-module 33 is used to acquire a data set, input the data set into the second model for training, and adopt a training strategy of two rounds of fine-tuning to obtain the final model.
[0065] In this embodiment, Bilstm (Bi-directional Long Short-Term Memory) is a type of deep network that can further abstract the features extracted by bert. After obtaining the second model, continue to train the second model, and adopt a training strategy of two rounds of fine-tuning. When training the model, the parameters of each layer in the network will be trained. Before training, an initial value needs to be set for the parameters. When using the fine-turning training strategy, a trained parameter is used as the initial value to solve the problem of insufficient training samples. For the schematic diagram of the final model, see Figure 5 , the Bilstm layer can further extract the features extracted by bertvec; cos-sim is a matching layer that can calculate the distance between two semantic vectors through cosine similarity, and softmax can be used to calculate the probability of classification problems.
[0066] See Figure 2 , in the sixth embodiment of the present invention, the data set acquisition sub-module 33 includes:
[0067] The first data set acquisition unit 331 is used to input the first initial data sets of the first quantity into the first model respectively. The first model outputs and labels the first initial data sets as the first positive samples or the first negative samples to obtain a first data set;
[0068] The second data set acquisition unit 332 is used to manually label the second initial data sets of the second quantity, and manually label the second initial data sets as the second positive samples or the second negative samples to obtain a second data set;
[0069] The first-round fine-tuning unit 333 is used to input the first data set into the second model for training and perform the first round of fine-tuning;
[0070] The second-round fine-tuning unit 334 is used to train the second dataset as input in the second model and perform a second round of fine-tuning to obtain the final model.
[0071] In this embodiment, manually annotating the second quantity of the second initial dataset can enhance the quality of the data, making the positive and negative samples more reasonable. Set the initial parameter values before training the second dataset as input in the second model, that is, perform the second round of fine-tuning; use the model parameters after training the first dataset as input in the second model as the initial parameter values when performing the second round of fine-tuning.
[0072] In the seventh embodiment of the present invention, the second dataset acquisition unit 332 divides the second negative samples into first similarity negative samples and second similarity negative samples; the second positive samples are divided into third similarity positive samples and fourth similarity positive samples;
[0073] The similarity between two search terms in the first similarity negative samples is lower than that in the second similarity negative samples;
[0074] The similarity between two search terms in the fourth similarity positive samples is higher than that in the third similarity positive samples.
[0075] In this embodiment, by further quantifying the positive and negative samples, it is more accurate than directly labeling them as positive or negative samples.
[0076] Examples of the format of manual annotation:
[0077] 1-Cotton and linen brand women's clothing-Flowered chiffon shirt, short-sleeved women's summer clothing;
[0078] 0-School education culture wall-Northern Education School;
[0079] 2-150 kW generator-How much is a 20 kW generator?
[0080] 3-Intermediate accounting registration process-Intermediate accountant registration time;
[0081] The above marked as 0 are the first similarity negative samples, marked as 1 are the second similarity negative samples, marked as 2 are the third similarity positive samples, and marked as 3 are the fourth similarity positive samples.
[0082] Preferably, the ratio of the first quantity to the second quantity is 5:1. For example, the first quantity is 1 million and the second quantity is 200,000.
[0083] Figure 3It is a flowchart of a method for judging similar search terms in an embodiment of the present invention. The method for judging similar search terms includes:
[0084] Step S301: Obtain a number of co-click data. Each piece of co-click data includes a search term, the URL clicked under the search term, and the number of clicks on the URL under the search term. This is achieved through the co-click data acquisition module 10.
[0085] Step S302: According to the URL clicked under the search term and the number of clicks on the URL under the search term in two pieces of co-click data, judge the similarity of the two search terms. Classify the two search terms with high similarity as a positive sample, and classify the two search terms with low similarity as a negative sample. A number of positive samples and a number of negative samples form a training set. This is achieved through the training set acquisition module 20.
[0086] Step S303: Use the training set as input to train a semantic pre-training model to obtain a final model. This is achieved through the final model acquisition module 30.
[0087] In this embodiment, the user searches for the desired information through search terms on the platform. During the advertisement recall process, the final model in the judgment method can judge whether the user's search term is similar to the search terms related to the advertisements provided by the advertisers. If they are similar, the advertisements provided by the advertisers can be displayed to the user, thereby increasing the revenue of the platform. The method for judging similar search terms can be implemented through the judgment system 100 in the above various embodiments. For the specific implementation process, refer to the above various embodiments and will not be elaborated here.
[0088] In an embodiment of the present invention, step S302 includes:
[0089] Score the search terms in a number of co-click data according to the URL clicked under the search term and the number of clicks on the URL under the search term. The search term with a score greater than the first threshold is the first search term, and the search term with a score less than or equal to the first threshold is the second search term. This is achieved through the scoring sub-module 21.
[0090] Classify two first search terms with the same URL in a number of co-click data as a positive sample to obtain a number of positive samples. Classify two second search terms with different URLs as a negative sample to obtain a number of negative samples. This is achieved through the classification sub-module 22.
[0091] In an embodiment of the present invention, the calculation formula for scoring is:
[0092] S u1q1=C u1q1 -avg(C u1 )
[0093] Wherein, assuming the search term is query1, the url clicked under the search term is url1, S u1q1 represents the score of query1, and C u1q1 represents the number of clicks on url1 when searching for query1. Wherein, n refers to the situation where url1 has been clicked when searching for n search terms. It is implemented by the scoring sub-module 21.
[0094] In an embodiment of the present invention, the step S301 includes: obtaining the plurality of co-click data from the user's click records. It is implemented by the co-click data acquisition module 10.
[0095] In an embodiment of the present invention, the semantic pre-training model is a Bert classification model (a semantic relevance model).
[0096] In an embodiment of the present invention, the step of taking the plurality of training sets as inputs, training the semantic pre-training model, and obtaining the final model includes:
[0097] Inputting the training set into the Bert classification model for training to obtain a first model; it is implemented by the first model acquisition sub-module 31;
[0098] Adding a layer of Bilstm network structure to the first model to obtain a second model; it is implemented by the second model acquisition sub-module 32;
[0099] Obtaining a data set, using the data set as an input in the second model for training, and adopting a training strategy of two rounds of fine-tuning to obtain the final model, which is implemented by the data set acquisition sub-module 33.
[0100] In an embodiment of the present invention, the step of obtaining a data set, using the data set as an input in the second model for training, and adopting a training strategy of two rounds of fine-tuning to obtain the final model includes:
[0101] Inputting the first quantity of first initial data sets into the first model respectively, and the first model outputs, marking the first initial data sets as first positive samples or first negative samples to obtain a first data set; it is implemented by the first data set acquisition unit 331;
[0102] Manually annotating the second quantity of second initial data sets, and manually marking the second initial data sets as second positive samples or second negative samples to obtain a second data set; it is implemented by the second data set acquisition unit 332;
[0103] The first data set is used as input for training in the second model, and the first round of fine-tuning is performed; this is achieved through the first-round fine-tuning unit 333;
[0104] The second data set is used as input for training in the second model, and the second round of fine-tuning is performed to obtain the final model; this is achieved through the second-round fine-tuning unit 334.
[0105] In one embodiment of the present invention, the second negative samples are divided into first similarity negative samples and second similarity negative samples; the second positive samples are divided into third similarity positive samples and fourth similarity positive samples;
[0106] The similarity between two search terms in the first similarity negative samples is lower than that in the second similarity negative samples;
[0107] The similarity between two search terms in the fourth similarity positive samples is higher than that in the third similarity positive samples.
[0108] Preferably, the ratio of the first quantity to the second quantity is 5:1.
[0109] The present invention also provides a storage medium for storing a computer program for executing any one of the above task scheduling methods. For example, computer program instructions, when executed by a computer, can call or provide the methods and / or technical solutions according to the present application through the operations of the computer. The program instructions for calling the methods of the present application may be stored in a fixed or removable storage medium, and / or transmitted and / or stored in the storage medium of a computer device running according to the program instructions through a data stream in a broadcast or other signal-bearing medium. Here, in one embodiment of the present application, there is a computer device 400 as Figure 4 shown. The computer device 400 preferably includes a storage medium 200 for storing a computer program and a processor 300 for executing the computer program. When the computer program is executed by the processor 300, the computer device 400 is triggered to execute the methods and / or technical solutions based on the foregoing multiple embodiments.
[0110] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and the like. Additionally, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0111] The method according to the present invention can be implemented on a computer as a computer-implemented method, or in dedicated hardware, or in a combination of both. The executable code or a part thereof for the method according to the present invention can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-temporary program code components stored on a computer-readable medium for executing the method according to the present invention when the program product is executed on a computer.
[0112] In a preferred embodiment, the computer program includes computer program code components suitable for executing all steps of the method according to the present invention when the computer program runs on a computer. Preferably, the computer program is embodied on a computer-readable medium.
[0113] In summary, the present invention obtains a number of co-click data, each of the co-click data including a search term, a url clicked under the search term, and the number of clicks on the url under the search term; according to the url clicked under the search term and the number of clicks on the url under the search term in two of the co-click data, determines the similarity of the two search terms, and classifies two search terms with high similarity as a positive sample; classifies two search terms with low similarity as a negative sample; a number of the positive samples and a number of the negative samples form a training set; uses the training set as an input to train a semantic pre-training model to obtain a final model. Inputs the user's search term and the search term provided by the advertiser into the final model, and the final model can determine whether the above two search terms are similar. If they are similar, also displays the advertisement related to the search term provided by the advertiser to the user, thereby realizing the promotion of the advertiser's advertisement and also increasing the revenue of the platform.
[0114] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for judging similar search terms, characterized in that, Including: Obtain a number of co-click data, each of the co-click data including a search term, a URL clicked under the search term, and the number of clicks on the URL under the search term; According to the URL clicked under the search term and the number of clicks on the URL under the search term in two of the co-click data, determine the similarity of the two search terms, and classify the two search terms with high similarity as a positive sample; Classify the two search terms with low similarity as a negative sample; a number of the positive samples and a number of the negative samples form a training set; Use the training set as an input to train a semantic pre-training model to obtain a final model; The step of determining the similarity of the two search terms according to the URL clicked under the search term and the number of clicks on the URL under the search term in two of the co-click data, and classifying the two search terms with high similarity as a positive sample; The step of classifying the two search terms with low similarity as a negative sample includes: Score the search terms in a number of the co-click data according to the URL clicked under the search term and the number of clicks on the URL under the search term; The search term with a score greater than a first threshold is a first search term, and the search term with a score less than or equal to the first threshold is a second search term; Classify two of the first search terms with the same URL in a number of the co-click data as a positive sample to obtain a number of the positive samples; classify two of the second search terms with different URLs as a negative sample to obtain a number of the negative samples.
2. The method for judging similar search terms according to claim 1, wherein The calculation formula for the scoring is: S u1q1 = C u1q1 - avg(C u1 ) Among them, assume that the search term is query1, and the url clicked under the search term is url1, S u1q1 represents the score of query1, C u1q1 represents the number of clicks on url1 when searching for query1, where n means that url1 has been clicked when searching for n search terms.
3. The method for judging similar search terms according to claim 1, characterized in that, The step of obtaining a number of co-click data includes: obtaining the number of co-click data from the click records of users.
4. The method for judging similar search terms according to claim 1, wherein the semantic pre-training model is a Bert classification model.
5. The method for determining similar search terms according to claim 4, characterized in that, The step of using the training set as an input to train a semantic pre-training model to obtain a final model includes: Input the training set into the Bert classification model for training to obtain a first model; Add a layer of Bilstm network structure to the first model to obtain a second model; Obtain a data set, use the data set as an input in the second model for training, and perform a training strategy of two rounds of fine-tuning to obtain the final model.
6. The method for judging similar search terms according to claim 5, characterized in that The step of obtaining a data set, using the data set as an input in the second model for training, and performing a training strategy of two rounds of fine-tuning to obtain the final model includes: Input a first number of first initial data sets into the first model respectively, the first model outputs, and label the first initial data sets as first positive samples or first negative samples to obtain a first data set; Manually label a second number of second initial data sets, and manually label the second initial data sets as second positive samples or second negative samples to obtain a second data set; Use the first data set as an input in the second model for training and perform the first round of fine-tuning; In the second model, the second data set is used as input for training, and a second round of fine-tuning is performed to obtain the final model.
7. The method for judging similar search terms according to claim 6, wherein, The second negative samples are divided into first similarity negative samples and second similarity negative samples; the second positive samples are divided into third similarity positive samples and fourth similarity positive samples; The similarity between two search terms in the first similarity negative samples is lower than that in the second similarity negative samples; The similarity between two search terms in the fourth similarity positive samples is higher than that in the third similarity positive samples.
8. The method for judging similar search terms according to claim 6, characterized in that, The ratio of the first quantity to the second quantity is 5:
1.
9. A judgment system for similar search terms, characterized in that, Including: A co-click data acquisition module, configured to acquire a plurality of co-click data, each co-click data including a search term, a url clicked under the search term, and the number of clicks on the url under the search term; A training set acquisition module, configured to determine the similarity between two search terms according to the url clicked under the search term and the number of clicks on the url under the search term in two co-click data, and classify the two search terms with high similarity as a positive sample; Classify the two search terms with low similarity as a negative sample; a plurality of the positive samples and a plurality of the negative samples form a training set; A final model acquisition module, configured to use the training set as input to train a semantic pre-training model to obtain a final model; The training set acquisition module includes: A scoring sub-module, configured to score the search terms in a plurality of the co-click data according to the url clicked under the search term and the number of clicks on the url under the search term; the search term with a score greater than a first threshold is a first search term, and the search term with a score less than or equal to the first threshold is a second search term; A classification sub-module, configured to classify two first search terms with the same url in a plurality of the co-click data as a positive sample to obtain a plurality of positive samples; classify two second search terms with different urls as a negative sample to obtain a plurality of negative samples.
10. The determination system for similar search terms according to claim 9, characterized in that, The calculation formula for the scoring sub-module to perform scoring is: S u1q1 = C u1q1 - avg(C u1 ) Among them, assume that the search term is query1, the URL clicked under the search term is url1, S u1q1 represents the score of query1, C u1q1 represents the number of clicks on url1 when searching for query1, where n means that url1 has been clicked when searching for n search terms.
11. The determination system for similar search terms according to claim 9, characterized in that, The co-click data acquisition module acquires the plurality of co-click data from the click records of the user.
12. The determination system for similar search terms according to claim 9, characterized in that, The semantic pre-training model is a Bert classification model.
13. The determination system for similar search terms according to claim 12, wherein The final model acquisition module includes: A first model acquisition sub-module, configured to input the training set into the Bert classification model for training to obtain a first model; A second model acquisition sub-module, configured to add a layer of Bilstm network structure to the first model to obtain a second model; A data set acquisition sub-module, configured to acquire a data set, use the data set as input in the second model for training, and adopt a training strategy of two rounds of fine-tuning to obtain the final model.
14. The determination system for similar search terms according to claim 13, wherein The data set acquisition sub-module includes: A first data set acquisition unit, configured to respectively input a first quantity of first initial data sets into the first model, and the first model outputs, and label the first initial data sets as first positive samples or first negative samples to obtain a first data set; A second data set acquisition unit, configured to manually label a second number of second initial data sets, and manually label the second initial data sets as second positive samples or second negative samples to obtain a second data set; A first round of fine-tuning unit, configured to use the first data set as an input in the second model for training and perform a first round of fine-tuning; A second round of fine-tuning unit, configured to use the second data set as an input in the second model for training and perform a second round of fine-tuning to obtain the final model.
15. The determination system for similar search terms according to claim 14, wherein The second data set acquisition unit divides the second negative samples into first similarity negative samples and second similarity negative samples; the second positive samples are divided into third similarity positive samples and fourth similarity positive samples; The similarity between two search terms in the first similarity negative samples is lower than that in the second similarity negative samples; The similarity between two search terms in the fourth similarity positive samples is higher than that in the third similarity positive samples.
16. The determination system for similar search terms according to claim 14, characterized in that, The ratio of the first number to the second number is 5:
1.
17. A storage medium, characterized in that, For storing a computer program for executing a method for judging any one of the similar search terms in claims 1 to 8.
18. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for judging similar search terms according to any one of claims 1 to 8.
Citation Information
Patent Citations
Similarity prediction model training method, apparatus, and computer-readable storage medium
CN109284399A
Search behavior identification method and device and identification device for search behavior
CN110147479A