Article recall method and related device
Through a model based on paragraph label training, paragraph labels are determined for article paragraphs, and the correlation scores between articles and query phrases are adjusted based on the matching degree of paragraph labels and query phrases, the problem of poor article recall in the existing technology is solved, and effective recall of related articles and content expansion is improved.
Patent Information
- Application Number
- CN202311553216.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-23
AI Technical Summary
The existing article recall method is difficult to effectively recall articles that meet the needs of the object when there is insufficient ecological resources, resulting in the position of the relevant articles in the recommendation list.
Through a model trained based on paragraphs and corresponding tags, paragraph labels are determined for article paragraphs, and the correlation scores between articles and query phrases are adjusted based on the matching degree of paragraph labels and query phrases, thereby improving the order of articles in the recommended list.
It improves the sorting of articles in the recommendation list, realizes effective recall of related articles, and improves the content expansion of recalled articles.
Smart Images

Figure CN120030138A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to an article recall method and related devices. Background Art
[0002] With the rapid development of the Internet, people can quickly browse the information on the Internet. Especially with the development of search engines, as long as people enter a query string in the search box of the search engine, the search engine can search for pages on the Internet that match the query term for the user to access, greatly facilitating the user's information acquisition. However, the expressions of a user's intention vary widely, resulting in the diversity of queries. And the existing articles lack a certain degree of generalization, such as the hypernyms of the article body, synonyms of article keywords, information about article carriers, etc. As a result, even if the article content meets the user's needs, it is difficult to recall.
[0003] The current solution is to expand on the query side, that is, to rewrite the currently input query to obtain its expanded vocabulary, and then recall relevant articles based on the relationship between the query and the global information of the article. However, in the case of insufficient ecological resources, only the local content of a small number of articles may meet the user's needs. Currently, only the global information of these articles is concerned, and the scores of these articles may be relatively low, resulting in these articles being ranked relatively low in the retrieval and recommendation list. Summary of the Invention
[0004] Embodiments of this application provide an article recall method and related devices, which are used to determine paragraph labels for article paragraphs based on a model trained with paragraphs and corresponding labels, and then adjust the relevance score between the article and the query phrase based on the matching degree between the paragraph label and the query phrase, improving the ranking of the article in the recommendation list to achieve the effect of recalling relevant articles. And if the template phrases in the training data of the second model are rare but frequently used phrases, the content extensibility of the recalled articles can be improved.
[0005] In view of this, on the one hand, this application provides an article recall method, including:
[0006] Input the paragraphs of a first article into a first model to output a first paragraph label. The first article is an article outside the preset sorting range in the recommendation list. The first model is trained and generated using first training data, and the first training data includes first data pairs of article paragraphs and first labels;
[0007] Inputting a paragraph of the first article into a second model to obtain a second paragraph label, the second model is trained and generated using second training data, the second training data includes a second data pair of the article paragraph and a matching phrase, the matching phrase includes an intersection of the first label and a template phrase in a template phrase library, the template phrase is a phrase with a preset period flow less than a threshold, and the phrase is repeated at least a first preset number of times;
[0008] Determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively;
[0009] increasing the relevance score of the first article and the first query phrase whose matching degree is greater than a preset threshold;
[0010] The ranking of the first article in the recommendation list is adjusted according to the relevance score between the first article and the first query phrase.
[0011] On the other hand, the present application provides an article recall device, comprising:
[0012] An input unit is used to input a paragraph of a first article into a first model to output a first paragraph label, wherein the first article is an article outside a preset sorting range of a recommendation list, and the first model is trained and generated using first training data, wherein the first training data includes a first data pair of an article paragraph and a first label; input a paragraph of the first article into a second model to obtain a second paragraph label, and the second model is trained and generated using second training data, wherein the second training data includes a second data pair of an article paragraph and a matching phrase, wherein the matching phrase includes an intersection of the first label and a template phrase in a template phrase library, wherein the template phrase is a word that appears at least a first preset number of times in a phrase whose preset period flow is less than a threshold value;
[0013] A determination unit, used to determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively;
[0014] The adjustment unit is used to increase the relevance score between the first article and the first query phrase whose matching degree is greater than a preset threshold; and adjust the ranking of the first article in the recommendation list according to the relevance score between the first article and the first query phrase.
[0015] In a possible design, in another implementation of another aspect of the embodiment of the present application, the device further includes an acquisition unit, and the acquisition unit is specifically used to:
[0016] Acquire a fourth data pair, the fourth data pair being a data pair in which the number of clicks on the target object exceeds the second preset number and the relevance score is greater than the first preset score, and the fourth data pair includes the third article and the third query phrase;
[0017] Get public datasets, which include public articles and phrases.
[0018] In a possible design, in another implementation of another aspect of the embodiment of the present application, the apparatus further includes a labeling unit, and the labeling unit is specifically used to:
[0019] Label the article paragraphs to obtain the first label.
[0020] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0021] The device also includes a filtering unit, which is specifically used for
[0022] Data pairs whose matching degree between the article paragraph and the first label in the first training data is outside a first preset range are filtered.
[0023] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0024] The device also includes an acquisition unit, which is specifically used to:
[0025] Obtain unpopular phrases with a preset period flow rate less than a threshold;
[0026] Clustering the unpopular phrases to select high-frequency words that appear at least a first preset number of times;
[0027] Phrases whose matching degree with the corresponding second article in the high-frequency word segmentation is outside a second preset range are filtered to obtain template phrases.
[0028] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0029] The device also includes a filtering unit, which is specifically used for:
[0030] Filter entity words in template phrases.
[0031] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0032] The device also includes a training unit, which is specifically used for:
[0033] The article paragraphs and matching phrases with relevance scores greater than a second preset score are taken as positive samples;
[0034] Randomly extract template phrases and corresponding articles from the template phrase library as negative samples;
[0035] The second model is trained based on the positive and negative samples.
[0036] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0037] Input cells are also used to:
[0038] Inputting the paragraph of the first article into the second model to output a plurality of candidate tags whose cosine similarity between the paragraph of the first article and the template phrase meets the requirement;
[0039] The determination unit is also used to:
[0040] Determine the relevance scores of the plurality of candidate tags and the paragraph of the first article, the second paragraph tag being a candidate tag whose relevance score satisfies the first condition.
[0041] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0042] The device also includes a trigger unit, which is specifically used to:
[0043] generating a recommendation list according to a third query phrase;
[0044] Obtaining a first query phrase;
[0045] When the first query phrase is the same as the third query phrase, a step of inputting the paragraph of the first article into the first model to output a first paragraph tag is triggered.
[0046] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0047] The adjustment unit is also used for:
[0048] Adjusting the rough ranking score of the first article according to the first query phrase, the first paragraph tag, and the second paragraph tag;
[0049] Add a mark to the first article whose rough ranking score meets the second condition;
[0050] Merge the first article that adds the tag into the recommendation list.
[0051] In one possible design, in another implementation of another aspect of the embodiment of the present application,
[0052] The adjustment unit is specifically used for:
[0053] The first article in the recommendation list that is marked is moved forward to be inserted before the fourth article, and the relevance score of the fourth article to the first query phrase is less than the relevance score of the first article that is marked to the first query phrase.
[0054] Another aspect of the present application provides a computer device, comprising:
[0055] memories, transceivers, processors, and bus systems;
[0056] Wherein, the memory is used to store programs;
[0057] The processor is used to execute the program in the memory, including executing the above-mentioned methods;
[0058] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
[0059] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer is enabled to execute the above-mentioned methods.
[0060] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided by the above aspects.
[0061] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0062] In the embodiment of the present application, a paragraph of the first article outside the preset sorting range of the recommendation list is input into the first model, and the first model is trained and generated using the first training data, and the first training data includes a first data pair of the article paragraph and the first label, and the first model outputs a first paragraph label related to the paragraph of the first article, and then the paragraph of the first article is input into the second model, and the second model is trained and generated using the second training data, and the second paragraph label is output, and the second training data includes a second data pair of the article paragraph and the matching phrase, and the matching phrase is determined by the intersection of the first label and the template phrase library, and the template phrase library includes multiple template phrases that are phrases that are repeated at least a first preset number of times in the phrases whose preset period flow is less than the threshold. Then, the matching degree of the first query phrase to be queried and the first paragraph label, as well as the second paragraph label, can be determined, and when the matching degree is greater than the preset threshold, the relevance score of the first article and the first query phrase is increased, and based on the increased relevance score, the position of the first article on the recommendation list is also improved accordingly. Through the above method, the model trained based on paragraphs and corresponding labels determines paragraph labels for article paragraphs, and then adjusts the relevance score between the article and the query phrase based on the matching degree between the paragraph label and the query phrase, thereby improving the ranking of the article in the recommendation list and achieving the effect of recalling related articles. Moreover, if the template phrases in the training data of the second model are unpopular and high-frequency phrases, the content extensibility of the recalled articles can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a schematic diagram of the architecture of the search system in the embodiment of the present application;
[0064] Figure 2 A schematic diagram of a process of an article recall method in an embodiment of the present application;
[0065] Figure 3 This is a schematic diagram of the GPT2 model in the embodiment of this application;
[0066] Figure 4 This is a schematic diagram of pre-similarity calculation in an embodiment of the present application;
[0067] Figure 5 This is a schematic diagram of the refined sorting optimization in the embodiment of this application;
[0068] Figure 6 This is a schematic diagram of the search architecture in the embodiment of the present application;
[0069] Figure 7 This is a schematic diagram of the structure of an article recall device in an embodiment of the present application;
[0070] Figure 8 It is a structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The embodiment of the present application provides an article recall method and related devices, which are used to determine paragraph labels for article paragraphs based on a model trained on paragraphs and corresponding labels, and then adjust the relevance score of the article and the query phrase based on the matching degree of the paragraph label and the query phrase, so as to improve the ranking of the article in the recommendation list and achieve the effect of recalling related articles. Moreover, if the template phrases in the training data of the second model are unpopular and high-frequency phrases, the content extensibility of the recalled articles can be improved.
[0072] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0073] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0074] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0075] In addition, in order to better illustrate the present application, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present application can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present application.
[0076] In order to facilitate understanding of the above-mentioned technical solutions and the technical effects produced by the embodiments of the present application, the embodiments of the present application first explain the relevant professional terms:
[0077] Query: refers to the short query text entered by the object, such as "sweet and delicious apples".
[0078] Term: that is, word segmentation. For example, the word segmentation result of "fragrant and delicious apple" is "fragrant and sweet, delicious, apple", which contains 3 terms, namely "fragrant and sweet", "delicious", and "apple".
[0079] Item: The permutation and combination of Term. For example, the combinations of three Term: "sweet", "delicious", "apple" are: 1 Term combination: "sweet", "delicious", "apple"; 2 Term combinations: "sweet and delicious", "sweet apple", "delicious apple"; 3 Term combinations: "sweet and delicious apple".
[0080] Word weight: Identify the weight of each word in the query. The higher the score, the more important the corresponding word. Word weight can be represented by word weight score, such as a score between 0 and 1. Word weight can also be represented by word weight level, such as 1-5.
[0081] Bert: Bidirectional Encoder Representation from Transformers, a bidirectional encoder representation technology based on transformers, is a pre-training technology for natural language processing.
[0082] With the rapid development of the Internet, people can quickly browse information on the Internet. Especially with the development of search engines, as long as people enter a query in the search box of the search engine, the search engine can search for pages on the Internet that match the search term based on the search term for the object to access, which greatly facilitates the object's information acquisition. However, the expression of a certain intention by the object is ever-changing, resulting in the diversity of queries, while the existing article side lacks certain generalization, such as the hypernyms of the article body, synonyms of the article keywords, article carriers and other information, resulting in difficulty in recalling even if the article content meets the needs of the object.
[0083] The current solution is to expand on the query side, that is, to rewrite the currently input query to obtain its expanded vocabulary after expansion, and then recall relevant articles based on the relationship between the query and the global information of the article. However, in the case of insufficient ecological resources, only the local content of a small number of articles may meet the needs of the object. Currently, only the global information of these articles is concerned, and the scores of these articles may be low, resulting in the position of these articles at the back of the search recommendation list.
[0084] Based on this, the embodiment of the present application provides an article recall method, wherein a paragraph of a first article outside the preset sorting range of a recommendation list is input into a first model, the first model is generated by training using first training data, the first training data includes a first data pair of an article paragraph and a first label, and the first model outputs a first paragraph label related to the paragraph of the first article, and then the paragraph of the first article is input into a second model, the second model is generated by training using second training data, the second training data includes a second data pair of an article paragraph and a matching phrase, the matching phrase is determined by the intersection of the first label and a template phrase library, and the template phrase library includes multiple template phrases that are phrases that are repeated at least a first preset number of times in a phrase with a preset period flow less than a threshold. Then, the matching degree between the first query phrase to be queried and the first paragraph label and the second paragraph label can be determined, and when the matching degree is greater than a preset threshold, the relevance score of the first article and the first query phrase is increased, and based on the increased relevance score, the position of the first article on the recommendation list is also improved accordingly. Through the above method, the model trained based on paragraphs and corresponding labels determines paragraph labels for article paragraphs, and then adjusts the relevance score between the article and the query phrase based on the matching degree between the paragraph label and the query phrase, thereby improving the ranking of the article in the recommendation list and achieving the effect of recalling related articles. Moreover, if the template phrases in the training data of the second model are unpopular and high-frequency phrases, the content extensibility of the recalled articles can be improved.
[0085] The embodiments of the present application are applied to the field of artificial intelligence (AI). Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0086] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0087] Specifically, the article recall method described in the embodiment of the present application relates to natural language processing (NLP) and machine learning (ML) in the field of artificial intelligence. Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language used by people in daily life, so it is closely related to the study of linguistics. Machine learning is a multi-field interdisciplinary subject involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by formula.
[0088] The article recall method of the search engine provided in the embodiment of the present application can be implemented by various electronic devices, for example, it can be implemented by a terminal device alone, or it can be implemented by a server and a terminal device in collaboration. For example, the terminal device alone executes the article recall method described below, or the terminal device and the server jointly execute the article recall method described below, for example, the terminal sends a message carrying a first query phrase to the server, and the server determines the execution order of each word according to the query phrase to search, obtains a recommendation list, and then inputs the paragraph of the first article outside the preset sorting range of the recommendation list into the first model to generate a first paragraph label, the first model is trained and generated using the first training data, the first training data includes a first data pair of the article paragraph and the first label, and then inputs the paragraph of the first article into the second model to generate a second paragraph label, the second model is trained and generated using the second training data, the second training data includes a second data pair of the article paragraph and the matching phrase, the matching phrase includes the intersection of the first label and the template phrase in the template phrase library, and the template phrase is a phrase with a preset period flow less than a threshold value and is repeated at least a first preset number of times. The server can determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively, increase the relevance score between the first article with a matching degree of a preset threshold and the first query phrase, and then adjust the ranking of the first article in the recommendation list according to the relevance score, that is, improve the ranking of the first article accordingly for the increased relevance score, and finally return the adjusted recommendation list to the terminal device.
[0089] The electronic device for article recall provided in the embodiments of the present application may be various types of terminal devices or servers, wherein the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the terminal may be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0090] Taking the server as an example, it can be a server cluster deployed in the cloud, opening artificial intelligence cloud services (AIaaS, AI as a Service) to objects. The AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to an AI theme mall. All objects can access one or more artificial intelligence services provided by the AIaaS platform through an application programming interface.
[0091] For example, one of the artificial intelligence cloud services may be a search engine article recall service, that is, the cloud server is encapsulated with the search engine article recall program provided by the embodiment of the present application. The terminal device responds to the object's search operation on the search engine by calling the search engine article recall service in the cloud service, so that the server deployed in the cloud calls the encapsulated search engine article recall program, determines the execution order of each word to search, obtains a recommendation list, and then adjusts the order of the first article in the recommendation list based on the relevance score between the paragraph of the first article and the first query phrase, and finally returns the adjusted recommendation list to the terminal device.
[0092] The following is an example of an article recall method provided by an embodiment of the present application being implemented in collaboration between a server and a terminal. Figure 1 , Figure 1 The terminal device 11 is connected to the server 13 via the network 12, and the network 12 can be a wide area network or a local area network, or a combination of the two.
[0093] In some embodiments, the search in the terminal device 11 responds to a search operation on a search engine and sends a first query phrase to the server 13. The server 13 searches the database based on the first query phrase to obtain a recommendation list, and then adjusts the order of the first article in the recommendation list based on the relevance score between the paragraphs of the first article and the first query phrase. Finally, the adjusted recommendation list is returned to the terminal device 11 to be displayed in the terminal device 11.
[0094] The article recall method provided in the embodiment of the present application will be described below in conjunction with the accompanying drawings. The executor of the following article recall method takes a terminal device as an example, and can be specifically implemented by the terminal device by running the various computer programs mentioned above; of course, based on the understanding of the following text, it is not difficult to see that the article recall method provided in the embodiment of the present application can also be implemented collaboratively by the terminal device and the server.
[0095] See also Figure 2 , Figure 2 The figure is a flow chart of a method for recalling an article provided in an embodiment of the present application, the method comprising:
[0096] Step 201. Input a paragraph of a first article into a first model to output a first paragraph label, wherein the first article is an article outside a preset sorting range of a recommendation list, and the first model is generated by training using first training data, wherein the first training data includes a first data pair of an article paragraph and a first label.
[0097] In one or more embodiments, when the terminal device user wants to obtain the target content, it can be achieved through the application that provides the search service in the terminal device client. Specifically, the terminal user enters the query text (Query) in the application that provides the search service, and triggers the search control to enable the application background server that provides the search service to process the query text. Optionally, the application that provides the search service includes but is not limited to content acquisition applications, shopping applications, life service applications, social applications, audio and video applications, etc. The embodiment of this application takes an article as an example.
[0098] After the user inputs the first query phrase, the terminal device can segment the first query phrase, and then search for the segmentation, and determine the recommendation list according to the relevance of the global features of each article and the first query phrase. The embodiment of the present application can recall articles that are outside the recommendation list or outside the preset sorting range of the recommendation list to the front of the recommendation list. Among them, the terminal device can select an article outside the preset sorting range of the recommendation list as the first article, for example, an article ranked after top250 as the first article, and input the paragraph of the first article into the first model, and the first model is pre-trained and generated using the first training data, wherein the first training data includes a first data pair of an article paragraph and a first label, that is, each article paragraph in the first training data has a corresponding first label, then the first model generated based on the first training data can output a corresponding label after inputting the paragraph of the first article, and the corresponding label can be called a first paragraph label.
[0099] The first model can use the generative pre-trained transformer 2 (GPT2) model proposed by OpenAI as the backbone. After pre-training, the model is fine-tuned on the above data in a supervised manner. The GPT2 model is only an example and can be replaced by a generative model with a larger parameter magnitude. Figure 3The GPT2 model diagram shown in the figure uses a transformer decoder for the language model, which contains 12 identical transformer blocks. The input of the transformer block can be the embedding of the text & position, where the position indicates which part of the text is instructing the model to work at the corresponding position. After the text & position embedding passes through the masked multi head self-attention layer of the GPT2 model, it is residually connected with the original text & position embedding to retain the original information. The residual connection result enters the normalization layer (layer norm) for processing, that is, the input value of each layer of neurons is normed, because when passing through the neural network, it is passed through one sample at a time, and the input value of each neuron is the value of a sample with different dimensions. The normalized result is then input into the feedforward layer, and the output of the feedforward layer is residually connected with the normalized result, and then processed by the normalization layer. The normalized output result is used to form a text prediction through a layer of full connection and softmax (not shown in the figure) to predict the probability of the next token. The normalized output result is combined with a layer of full connection and softmax (not shown in the figure) to form a task classifier to predict the probability of each category.
[0100] Step 202. Input the paragraph of the first article into the second model to obtain a second paragraph label. The second model is trained and generated using second training data. The second training data includes second data pairs of article paragraphs and matching phrases. The matching phrases include the intersection of the first label and a template phrase in the template phrase library. The template phrase is a word that appears at least a first preset number of times in a phrase with a preset period flow less than a threshold.
[0101] In one or more embodiments, the second model is pre-trained, and the second model is trained using the second training data during the training process, wherein the embodiment of the present application may pre-save a template phrase library, and the template phrases in the template phrase library may be obtained by selecting multiple phrases with a preset period flow less than a threshold, and determining the words (one or more participles) that appear repeatedly at least a first preset number of times in the multiple phrases as template phrases. The terminal device may intersect the first label in the first training data with the template phrase, that is, map the first label to the template phrase in the template phrase library, and the mapping condition is that the first label completely hits the template phrase, or the template phrase completely hits the first label, wherein the complete hit indicates that one party is completely mentioned in the other party, for example, if the first label is "emoji package" and the template phrase is "cute emoji package", the first label completely hits the template phrase, otherwise, the template phrase completely hits the first label. The template phrase that meets the mapping condition may be called a matching phrase, and accordingly, if the matching phrase corresponds to an article paragraph, the data pair of the matching phrase and the article paragraph may be called the second training data, and after the paragraph of the first article is input into the second model, the corresponding second paragraph label may be obtained.
[0102] Step 203: Determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively.
[0103] In one or more embodiments, when the object determines the first query phrase as the phrase to be queried, the terminal device may respectively match the first paragraph tag and the second paragraph tag of the paragraph of the first article to obtain a matching degree, which may be a literal matching degree.
[0104] In the embodiment of the present application, the first paragraph tag and the second paragraph tag can be incorporated into the construction of the article index chain to solve the problem of intersection failure caused by the missing of some words.
[0105] Step 204: Increase the relevance score between the first article and the first query phrase whose matching degree is greater than a preset threshold.
[0106] In one or more embodiments, when the degree of match is greater than a preset threshold, it can be considered that the paragraph of the first article meets the query requirements of the first query phrase. Since the first article was originally ranked later, in order to enable the object to view the first article faster, the relevance score assessed between the first article and the first query phrase can be increased, that is, the relevance between the first article and the first query phrase can be improved.
[0107] Step 205: Adjust the ranking of the first article in the recommendation list according to the relevance score between the first article and the first query phrase.
[0108] In one or more embodiments, the recommendation list may be a list that is filtered and ranked by the terminal device according to normal filtering conditions. When the first article is in the recommendation list, the embodiment of the present application may raise the ranking position of the first article in the recommendation list based on the relevance score between the first article and the first query phrase. When the first article is not in the recommendation list, the embodiment of the present application may insert the first article into the recommendation list based on the relevance score between the first article and the first query phrase.
[0109] In the embodiment of the present application, a paragraph of the first article outside the preset sorting range of the recommendation list is input into the first model, and the first model is trained and generated using the first training data, and the first training data includes a first data pair of the article paragraph and the first label, and the first model outputs a first paragraph label related to the paragraph of the first article, and then the paragraph of the first article is input into the second model, and the second model is trained and generated using the second training data, and the second paragraph label is output, and the second training data includes a second data pair of the article paragraph and the matching phrase, and the matching phrase is determined by the intersection of the first label and the template phrase library, and the template phrase library includes multiple template phrases that are phrases that are repeated at least a first preset number of times in the phrases whose preset period flow is less than the threshold. Then, the matching degree of the first query phrase to be queried and the first paragraph label, as well as the second paragraph label, can be determined, and when the matching degree is greater than the preset threshold, the relevance score of the first article and the first query phrase is increased, and based on the increased relevance score, the position of the first article on the recommendation list is also improved accordingly. Through the above method, the model trained based on paragraphs and corresponding labels determines paragraph labels for article paragraphs, and then adjusts the relevance score between the article and the query phrase based on the matching degree between the paragraph label and the query phrase, thereby improving the ranking of the article in the recommendation list and achieving the effect of recalling related articles. Moreover, if the template phrases in the training data of the second model are unpopular and high-frequency phrases, the content extensibility of the recalled articles can be improved.
[0110] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the first training data further includes a fourth data pair and a public data set, and before inputting the paragraph of the first article into the first model to output the first paragraph label, the method further includes:
[0111] Acquire a fourth data pair, the fourth data pair being a data pair in which the number of clicks on the target object exceeds the second preset number and the relevance score is greater than the first preset score, and the fourth data pair includes the third article and the third query phrase;
[0112] Get public datasets, which include public articles and phrases.
[0113] In one or more embodiments, a method for obtaining first training data is introduced. The click object doc of the target object for the query being queried will constitute a query-doc pair, and the corresponding relationship of the query-doc pair will be saved in the click log of the target object. The terminal device can filter the query-doc pairs with more than a second preset number of clicks from the click log, and then determine the relevance score for the query-doc pair, and then select the query-doc pair with a relevance score greater than the first preset score as the fourth data pair. Correspondingly, the query in the fourth data pair is the third query phrase, and the doc in the fourth data pair is the third article. Exemplarily, query-doc pairs with more than 20 clicks and a relevance score higher than 3 levels can be filtered from the object click log, where the relevance score level can be set to 5 levels. The public data set is a collection of data pairs that can be freely collected on the Internet. The data pair can be a data pair of a public article and a phrase, or a data pair of a paragraph and a phrase, that is, a query-doc pair or a query-para pair, which is not limited here.
[0114] Secondly, in the embodiment of the present application, a method for obtaining the first training data is provided. Through the above method, the data pairs with the number of clicks exceeding the second preset number and the relevance score greater than the first preset score, and the public data set are involved in the training of the first model, so that the first model can combine the information of the fourth data pair and the public data set to improve the accuracy of the first model.
[0115] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes:
[0116] Label the article paragraphs to obtain the first label.
[0117] In one or more embodiments, a method of marking article paragraphs is introduced. The first data pair may be obtained by pre-marking the article paragraphs, and the randomly selected article paragraphs are marked based on the preset marking conditions. Exemplarily, the embodiment of the present application may use ChatGPT to tag the paragraphs, and ChatGPT may be required to generate the first tag under the following two prompts: 1. "Please generate no more than 5 search engine query phrases based on the paragraph content, meeting the following requirements: 1) The query phrases are as concise as possible, and each phrase is no more than 15 words; 2) The query phrases contain expressions that are not in the paragraph; 3) Please focus on the time, place, person, and organization information in the paragraph. The paragraph content is as follows: \n{}"; 2. Generate tags based on query and paragraph: "Please generate multiple search engine query phrases similar to the given query based on the paragraph content, and the query phrases are required to be as concise as possible. The query and paragraph content are as follows: \n query:{}\n paragraph content:{}".
[0118] Secondly, in the embodiment of the present application, a method for marking article paragraphs is provided. Through the above method, by marking the article paragraphs, the specification of the first data pair meets the model requirements, and the accuracy of article recall is improved.
[0119] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes:
[0120] Data pairs whose matching degree between the article paragraph and the first label in the first training data is outside a first preset range are filtered.
[0121] In one or more embodiments, a method of filtering the first training data is introduced. The terminal device can select data pairs whose matching degree between the first label and the article paragraph is within a first preset range as training data for the first model. A too high matching degree is likely to lead to low scalability of the model, while a too low matching degree is likely to affect the accuracy of the model. Exemplarily, the filtered data pairs can be data pairs in which some first labels in the first training data completely hit the article paragraphs, or data pairs in which some first labels completely miss the article paragraphs, wherein the correlation of the completely hit data pairs is too strong, resulting in limitations in article recall, and the data pairs that completely miss are obviously less relevant.
[0122] Secondly, in the embodiment of the present application, a method for filtering the first training data is provided. By filtering the data pairs whose matching degree between the article paragraph and the first label is outside the first preset range, the richness of the model content is improved, and the expansibility of article recall is ensured.
[0123] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the second model to obtain the second paragraph label, the method further includes:
[0124] Obtain unpopular phrases with a preset period flow rate less than a threshold;
[0125] Clustering the unpopular phrases to select high-frequency words that appear at least a first preset number of times;
[0126] Phrases whose matching degree with the corresponding second article in the high-frequency word segmentation is outside a second preset range are filtered to obtain template phrases.
[0127] In one or more embodiments, a method for obtaining template phrases is introduced. For some unpopular phrases that include high-frequency participles, the selection range of article recall can be expanded. In the embodiment of the present application, a threshold of a preset periodic flow rate can be set, and phrases less than the threshold are called unpopular phrases. For example, phrases with a weekly flow rate less than 3 are called unpopular phrases. Then the unpopular phrases are clustered, for example, by the degree of literal similarity. If multiple unpopular phrases are clustered in each cluster, then the high-frequency common subsequences in each cluster can be screened. The high-frequency common subsequences can be called high-frequency participles, where the high frequency in the embodiment of the present application can be the degree of repetition of more than a first preset number of times, for example, repetition of more than 3 times. For example, assuming that the phrases in a cluster are (pictures of people, pictures of cats, pictures of scenery, types of cats), then the high-frequency participle is a picture. After obtaining the high-frequency participles, phrases in the high-frequency participles whose matching degree with the corresponding second article is outside a second preset range can be filtered. The second preset range can be that the high-frequency participle appears in the second article. Outside the second preset range is that the second article does not include the template phrase, that is, when the corresponding high-frequency participle does not appear in the content of the second article, the high-frequency participle can be called a template phrase.
[0128] For example, for obtaining the template phrase library, the embodiment of the present application can perform a word-granularity intersection operation on the second article that is finally exposed under the high-frequency segmentation according to the high-frequency segmentation, and select the high-frequency segmentation in the high-frequency segmentation but not in the second article (doc) as the template phrase (pattern). For example, the high-frequency segmentation is "something emoticon package", and the second article does not mention the "something emoticon package", but the second article has been clicked under the entry.
[0129] Secondly, in the embodiment of the present application, a method for obtaining template phrases is provided. Through the above method, unpopular phrases are screened by thresholds, high-frequency word segmentation is obtained by clustering, and phrases whose matching degree with the second article is outside the second preset range are filtered to generate template phrases, and unpopular high-frequency phrases are used to participate in the training of the second model to improve the scalability of the model, and the unpopular high-frequency phrases further filter phrases that do not meet the conditions, further improving the pertinence of the model.
[0130] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the second model to obtain the second paragraph label, the method further includes:
[0131] Filter entity words in template phrases.
[0132] In one or more embodiments, a method for processing a template phrase is introduced. The extensibility of article recall can also be improved by deleting entity words in the template phrase. Assuming that the template phrase includes Changi Airport Timetable, since "Changi" is a location entity and has too many restrictions, the participle "Changi" can be removed to finally obtain "Airport Timetable".
[0133] Secondly, in the embodiment of the present application, a method for processing template phrases is provided. By the above method, entity words in the template phrases are deleted, and the extensibility of the second model is further improved.
[0134] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the second model to obtain the second paragraph label, the method further includes:
[0135] The article paragraphs and matching phrases with relevance scores greater than a second preset score are taken as positive samples;
[0136] Randomly extract template phrases and corresponding articles from the template phrase library as negative samples;
[0137] The second model is trained based on the positive and negative samples.
[0138] In one or more embodiments, a method for training the second model is introduced. It is also possible to select part of the second training data as training samples for the second model, and screen data pairs whose correlation between article paragraphs and matching phrases meets the conditions as positive samples. The embodiment of the present application can filter through the correlation BERT model, and select data pairs of article paragraphs and matching phrases whose correlation scores are greater than the second preset score, wherein the correlation score can also be expressed in the form of a gear, for example, selecting data with a correlation gear greater than or equal to 4 gears as positive sample training data, and the negative sample training data can be a template phrase and a corresponding article randomly extracted from a template phrase library as a negative sample, or the most difficult negative sample can be sampled to participate in the training of the second model.
[0139] Secondly, in the embodiment of the present application, a method for training the second model is provided. Through the above method, positive samples and negative samples are used at the same time, so that these template phrases that are not selected as targets can also participate in the model training process. Accordingly, during the model training process, the trained model can learn the key information contained in such background samples, thereby improving the model performance of the trained model, so that the model can more comprehensively and accurately identify various input data.
[0140] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, inputting the paragraph of the first article into the second model to obtain the second paragraph label includes:
[0141] Inputting the paragraph of the first article into the second model to output a plurality of candidate tags whose cosine similarity between the paragraph of the first article and the template phrase meets the requirement;
[0142] Determine the relevance scores of the plurality of candidate tags and the paragraph of the first article, the second paragraph tag being a candidate tag whose relevance score satisfies the first condition.
[0143] In one or more embodiments, a method for obtaining a second paragraph label is introduced. Candidate labels can be selected by the cosine similarity between the paragraph of the first article and the template phrase, the second model receives the paragraph of the first article, and outputs multiple candidate labels that meet the cosine similarity requirements. The multiple candidate labels can be all template phrases that meet the cosine similarity requirements, or some template phrases that are ranked top. Then, by calculating the relevance score between each candidate label and the paragraph of the first article, the multiple candidate labels are sorted for relevance, and the candidate label whose relevance score meets the first condition can be selected as the second paragraph label. The first condition can be the highest relevance score or meeting a score threshold, etc.
[0144] For example, the twin model sentence-bert can be used as the backbone, such as Figure 4 As shown in the figure, after pre-training, the twin model is used to train the second training data, with paragraph features encoded on the left and template phrase features encoded on the right. The paragraph features are pre-trained by the Bert model to generate a paragraph embedding vector with semantics, which is then converted into a fixed-degree paragraph sentence embedding through a pooling layer, where u is the vector representation of the paragraph sentence embedding, and the template phrase features are pre-trained by the Bert model to generate a template embedding vector with semantics, which is then converted into a fixed-degree template sentence embedding through a pooling layer, where v is the vector representation of the template sentence embedding. The cosine similarity of the two vector representations is then calculated, and finally the 10 patterns with the highest cosine similarity are selected from the pattern library as candidate tags. The correlation Bert model is then used to calculate the correlation score of each candidate tag with the paragraph of the first article, and the correlation score is re-ranked, and finally the top 3 candidate tags with a score threshold greater than or equal to 0.2 are selected as the final second paragraph tag.
[0145] Secondly, in the embodiment of the present application, a method for obtaining a second paragraph label is provided. In the above method, the candidate labels are first screened by cosine similarity, and then the second paragraph label is selected based on the relevance score between the candidate labels and the paragraph of the first article, and the accuracy of the second paragraph label selection is improved through double screening.
[0146] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes:
[0147] generating a recommendation list according to a third query phrase;
[0148] Obtaining a first query phrase;
[0149] When the first query phrase is the same as the third query phrase, a step of inputting the paragraph of the first article into the first model to output a first paragraph tag is triggered.
[0150] In one or more embodiments, a method for triggering article recall is introduced. When sufficient results are recalled during the first query issuance of the query, it indicates that there are many recommended results that meet the object requirements. In this case, it is not necessary to increase the recall through paragraph tags. If the object is not satisfied with the recommended list generated for the third query phrase, a second query can be issued, that is, it is determined to input the first query phrase. When the first query phrase is the same as the third query phrase, the terminal device can determine that it is necessary to trigger the article recall scheme of the embodiment of the present application, that is, trigger step 201. Among them, the first query issuance of the query uses the main path for recall, and the second query issuance can add n-way recalls. n depends on the query issuance granularity, and generally n <= 3, and then combined with the recommended list generated by the main path for display.
[0151] Secondly, in the embodiment of the present application, a method for triggering article recall is provided. Through the above method, the article recall scheme of the embodiment of the present application is triggered only when the query phrase is issued for the second time, saving computing resources.
[0152] Optionally, based on the corresponding embodiments above Figure 2 In another optional embodiment provided by the embodiment of the present application, before determining the matching degrees of the first query phrase to be queried with the first paragraph tag and the second paragraph tag respectively, the method further includes:
[0153] Adjust the rough ranking score of the first article according to the first query phrase, the first paragraph tag, and the second paragraph tag;
[0154] Add a mark to the first article whose rough ranking score meets the second condition;
[0155] Merge the first article with the mark added into the recommended list.
[0156] In one or more embodiments, a method for optimizing rough ranking is introduced. First, determine the hit degree scores in each field of rough ranking for the first article corresponding to the first query phrase and the first paragraph tag and the second paragraph tag respectively, then adjust the name degree score based on the hit relationship between the first paragraph tag and the second paragraph tag and the first query phrase, then rank the first article based on the hit degree score, and mark the first article whose rough ranking score meets the second condition. The second condition can be that the score is greater than the score threshold, or the first few articles in the front row of the ranking. Then, the first article with the mark added can be merged into the recommended list generated by the normal ranking.
[0157] Exemplarily, the rough sorting is usually divided into three stages, namely L1, L2, and L3. L1 mainly focuses on the hit degree scoring of the query and each domain of the doc, such as the title domain, the text domain, the account domain, and the related query domain. The maximum hit score of each domain is taken as the L1 score. In the embodiment of the present application, a new paragraph tag domain (the first paragraph tag and the second paragraph tag) is added. If the keywords in the query hit the tag domain, the L1 score will be multiplied by a coefficient greater than 1 to increase the score. L2 mainly considers the word weight of the query in the doc to score. Generally, the higher the word weight, the higher the score. In this method, if the words in the article appear in the paragraph tag domain at the same time, the word weight of the word is defaulted to the highest level in the embodiment of the present application to improve the L2 score. L3 is not changed. Note that in order to ensure that it does not affect the main recall calculation, this method is only applied in the newly added n-way delivery. Specifically, the embodiment of the present application calculates the L1 and L2 scores of the first article in the newly added n-path recall, sorts the first article from high to low according to the L1 and L2 scores, and iii. selects at most 10 docs that meet the following second condition from high to low, and marks them with flags to indicate that they are recalled by paragraph tags. Among them, the first article may be one that does not appear in the top 250 articles in the final sorting of the main path recall, or there is a word in the query that only hits the label field in the first article, and does not hit the other four fields (title, text, account, related queries), or the required words in the query all hit the label field in the first article. Then the at most 10 first articles carrying the mark can be merged with the articles recalled by the main path.
[0158] Secondly, in the embodiment of the present application, a method for optimizing rough ranking is provided. In the above method, the first query phrase and the first paragraph tag and the second paragraph tag are firstly roughly ranked and scored respectively, the first article that meets the second condition is marked and merged into the recommendation list, and the article recalled by the embodiment of the present application is added without affecting the information of the recommended articles in the main path. The object can obtain the recall information of the main path and other n paths at the same time, which improves the query experience of the object.
[0159] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, adjusting the ranking of the first article in the recommendation list according to the relevance score between the first article and the first query phrase includes:
[0160] The first article in the recommendation list that is marked is moved forward to be inserted before the fourth article, and the relevance score of the fourth article to the first query phrase is less than the relevance score of the first article that is marked to the first query phrase.
[0161] In one or more embodiments, a method for optimizing fine ranking is introduced. On the basis of the first marked article added to the recommendation list in the rough ranking, the articles in the recommendation list are sorted according to the usual sorting method of fine ranking. After re-sorting, the first marked article in the recommendation list is screened again, and according to the relevance score between the article and the first query phrase, the first article is moved forward to the fourth article. The fourth article can be any one before the first article and with a lesser relevance score than the first article, or can be an article with a lesser relevance score than the first article and the highest ranking.
[0162] Exemplarily, the refined ranking is usually represented by L4, which includes relevance calculation and article sorting. The embodiment of the present application can adjust the relevance score obtained by L4. The relevance score is a score that measures the degree of association between the query and the doc content. The deep model and the tree model finally obtain a score. The embodiment of the present application can perform post-processing and weighting on the model output score, specifically: calculate the maximum literal matching rate (expressed as max_match_ratio) and the average literal matching rate (the average of top3, expressed as mean_match_ratio) of the first query phrase and the first paragraph tag and the second paragraph tag of each marked first article respectively. The final literal matching rate match_ratio = 0.7*max_match_ratio+0.3*mean_match_ratio, and then the article relevance score and gear are weighted according to the literal matching rate. Assuming that the relevance gear is expressed as level (ranging from 1 to 5 gears), the relevance score is expressed as score (ranging from 0 to 1.0), and the specific adjustment strategy is as follows:
[0163] 1). If match_ratio is greater than 0.8, score = min(0.95, score + 0.2), level = min(5, level + 1).
[0164] 2). If match_ratio is greater than 0.6, score = min(0.95, score + 0.1), level = min(5, level + 1).
[0165] 3). If match_ratio is greater than 0.3, score = min(0.95, score+0.05).
[0166] Among them, for the first article with a flag mark, if the relevance level is less than or equal to 2, it can be directly filtered; for a strongly time-sensitive query, if there is a first article with a flag mark, since the newly added n channels do not pay attention to timeliness, it can be directly filtered. For example, assume that the first query phrase being queried now is "new case information", and the relevant docs in the first article may be from 2020. Then, according to the distribution of the relevance levels of the entire queue, the docs with flag marks are retained or filtered. For example, assume that most of the relevance levels in the entire queue are 4 and 5, then the docs with flag marks and a relevance level of 3 and below can be directly filtered. Assume that most of the relevance levels in the entire queue are 3, then the docs with flag marks and a relevance level of 3 and above can be retained. Then, use the system sorting function, and this system sorting function can be a way of sorting based on the global information of the article. The embodiment of the present application can determine the sorted queue, and judge the satisfaction degree of the relevance of the top 4 articles. If the number of articles with a relevance level greater than or equal to 4 in the top 4 does not exceed 2, the following strong insertion strategy will be performed:
[0167] 1. Start traversing from the fourth article in the queue from front to back, and judge whether each doc has a flag mark until the first doc with a flag mark is encountered and taken out.
[0168] 2. Compare the relevance of the doc with the flag mark with the third doc in the queue. If the former is greater than or equal to the latter, insert the doc with the flag mark into the third position in the queue, and move the other docs back one position in sequence.
[0169] Exemplarily, please refer to Figure 5 the schematic diagram of refined ranking optimization shown. The left figure is the result after system sorting. Taking the first 10 articles as an example, they are articles 1 - 10 respectively. Assume that article 5 is the first article with a mark, and the relevance of article 5 to the first query phrase is greater than the relevance of article 3 to the first query phrase. Then, as shown in the right figure, article 5 can be inserted before article 3, and the others are moved back in sequence accordingly.
[0170] Exemplarily, the search architecture of the embodiment of the present application can refer to Figure 6As shown in the figure, the search architecture includes an indexing layer 601, a recall layer 602, a ranking layer 603, and a display layer 604. Among them, the indexing layer 601 is used to incorporate the first paragraph label and the second paragraph label into the construction of the index chain to facilitate querying and indexing the corresponding first article; the recall layer is used to recall articles outside the preset ranking range, that is, based on the relationship between the first query phrase and the first paragraph label and the second paragraph label, adjust the rough ranking score of the first article corresponding to the first paragraph label and the second paragraph label, and then select the first article based on the rough ranking score for marking and recalling; the ranking layer 603 is used to rank the recalled first article and the main road recalled articles according to the global information, and then rank based on the relevance score to move the marked first article forward; the display layer 604 is used to display the refined ranking optimized recommendation list to the user for viewing.
[0171] Secondly, in the embodiments of the present application, a method for optimizing refined ranking is provided. Through the above method, the first article is moved forward by comparing the relevance scores, making the first article easier to be noticed by the object and improving the reliability of the refined ranking of the article.
[0172] The following will describe in detail the article recall setting in the present application. Please refer to Figure 7 , Figure 7 which is a schematic diagram of an embodiment of the device in the embodiments of the present application. The article recall device 70 includes:
[0173] An input unit 701, configured to input the paragraphs of the first article into the first model to output the first paragraph label. The first article is an article outside the preset ranking range in the recommendation list. The first model is trained and generated using the first training data, and the first training data includes the first data pair of the article paragraph and the first label; input the paragraphs of the first article into the second model to obtain the second paragraph label. The second model is trained and generated using the second training data, and the second training data includes the second data pair of the article paragraph and the matching phrase. The matching phrase includes the intersection of the first label and the template phrase in the template phrase library. The template phrase is a word that appears at least the first preset number of times in the phrases with a preset periodic traffic less than the threshold;
[0174] A determination unit 702, configured to determine the matching degrees of the first query phrase to be queried with the first paragraph label and the second paragraph label respectively;
[0175] An adjustment unit 703, configured to increase the relevance score of the first article with a matching degree greater than the preset threshold to the first query phrase; adjust the ranking of the first article in the recommendation list according to the relevance score of the first article to the first query phrase.
[0176] In an embodiment of the present application, an article recall device is provided. Through the above device, a model based on paragraph and corresponding label training determines paragraph labels for article paragraphs, and then adjusts the relevance score between the article and the query phrase based on the matching degree between the paragraph label and the query phrase, thereby improving the ranking of the article in the recommendation list, achieving the effect of recalling related articles, and based on the fact that the template phrase in the training data of the second model is a low-frequency phrase, the content extensibility of the recalled article can be improved.
[0177] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes an acquisition unit 704, and the acquisition unit 704 is specifically used for:
[0178] Acquire a fourth data pair, the fourth data pair being a data pair in which the number of clicks on the target object exceeds the second preset number and the relevance score is greater than the first preset score, and the fourth data pair includes the third article and the third query phrase;
[0179] Get public datasets, which include public articles and phrases.
[0180] In an embodiment of the present application, an article recall device is provided. Through the above device, data pairs with a click count exceeding a second preset number and a relevance score greater than a first preset score and a public data set are involved in the training of the first model, so that the first model can combine the information of the fourth data pair and the public data set to improve the accuracy of the first model.
[0181] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes a labeling unit 705, and the labeling unit 705 is specifically used for:
[0182] Label the article paragraphs to obtain the first label.
[0183] In an embodiment of the present application, a device for recalling an article is provided. Through the device, by marking the article paragraphs, the specification of the first data pair meets the model requirements, thereby improving the accuracy of article recall.
[0184] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes a filtering unit 706, and the filtering unit 706 is specifically used to
[0185] Data pairs whose matching degree between the article paragraph and the first label in the first training data is outside a first preset range are filtered.
[0186] In an embodiment of the present application, an article recall device is provided. Through the above device, data pairs whose matching degree between the article paragraph and the first tag is outside a first preset range are filtered, the richness of the content of the model is improved, and the expansibility of article recall is guaranteed.
[0187] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes an acquisition unit 704, and the acquisition unit 704 is specifically used for:
[0188] Obtain unpopular phrases with a preset period flow rate less than a threshold;
[0189] Clustering the unpopular phrases to select high-frequency words that appear at least a first preset number of times;
[0190] Phrases whose matching degree with the corresponding second article in the high-frequency word segmentation is outside a second preset range are filtered to obtain template phrases.
[0191] In an embodiment of the present application, an article recall device is provided. Through the above device, unpopular phrases are screened by thresholds, high-frequency word segments are obtained by clustering, and phrases whose matching degree with the second article is outside the second preset range are filtered to generate template phrases, and unpopular high-frequency phrases are used to participate in the training of the second model to improve the scalability of the model, and the unpopular high-frequency phrases further filter phrases that do not meet the conditions, further improving the pertinence of the model.
[0192] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes a filtering unit 706, and the filtering unit 706 is specifically used for:
[0193] Filter entity words in template phrases.
[0194] In an embodiment of the present application, a device for recalling articles is provided. By means of the device, entity words in the template phrase are deleted, thereby further improving the scalability of the second model.
[0195] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes a training unit 707, and the training unit 707 is specifically used for:
[0196] The article paragraphs and matching phrases with relevance scores greater than a second preset score are taken as positive samples;
[0197] Randomly extract template phrases and corresponding articles from the template phrase library as negative samples;
[0198] The second model is trained based on the positive and negative samples.
[0199] In the embodiment of the present application, an article recall device is provided. Through the above device, positive samples and negative samples are used at the same time, so that these template phrases that are not selected as targets can also participate in the model training process. Accordingly, during the model training process, the trained model can learn the key information contained in such background samples, thereby improving the model performance of the trained model, so that the model can more comprehensively and accurately identify various input data.
[0200] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the input unit 701 is further used for:
[0201] Inputting the paragraph of the first article into the second model to output a plurality of candidate tags whose cosine similarity between the paragraph of the first article and the template phrase meets the requirement;
[0202] The determining unit 702 is further configured to:
[0203] Determine the relevance scores of the plurality of candidate tags and the paragraph of the first article, the second paragraph tag being a candidate tag whose relevance score satisfies the first condition.
[0204] In an embodiment of the present application, an article recall device is provided. Through the above device, candidate tags are first screened by cosine similarity, and then a second paragraph tag is selected based on the relevance score between the candidate tag and the paragraph of the first article, thereby improving the accuracy of the second paragraph tag selection through double screening.
[0205] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the device 70 further includes a trigger unit 708, and the trigger unit 708 is specifically used for:
[0206] generating a recommendation list according to a third query phrase;
[0207] Obtaining a first query phrase;
[0208] When the first query phrase is the same as the third query phrase, a step of inputting the paragraph of the first article into the first model to output a first paragraph tag is triggered.
[0209] In an embodiment of the present application, an article recall device is provided. Through the above device, the article recall solution of the embodiment of the present application is triggered only when the query phrase is issued for the second time, thereby saving computing resources.
[0210] Optionally, in the above Figure 7On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the adjustment unit 703 is further used for:
[0211] Adjusting the rough ranking score of the first article according to the first query phrase, the first paragraph tag, and the second paragraph tag;
[0212] Add a mark to the first article whose rough ranking score meets the second condition;
[0213] Merge the first article that adds the tag into the recommendation list.
[0214] In the embodiment of the present application, an article recall device is provided. Through the above device, the first query phrase and the first paragraph tag and the second paragraph tag are firstly roughly ranked and scored respectively, the first article that meets the second condition is marked and merged into the recommendation list, and the article recalled by the embodiment of the present application is added without affecting the information of the recommended articles in the main path. The object can obtain the recall information of the main path and other n paths at the same time, which improves the query experience of the object.
[0215] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in another embodiment of the article recall device 70 provided in the embodiment of the present application, the adjustment unit 703 is specifically used for:
[0216] The first article in the recommendation list that is marked is moved forward to be inserted before the fourth article, and the relevance score of the fourth article to the first query phrase is less than the relevance score of the first article that is marked to the first query phrase.
[0217] In an embodiment of the present application, an article recall device is provided. Through the above device, the first article is moved forward by comparing the relevance scores, so that the first article is more likely to be noticed by the target, thereby improving the reliability of the article sorting.
[0218] Figure 8It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device. Further, the central processor 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device 300.
[0219] The computer device 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0220] The steps performed by the terminal device in the above embodiments may be based on the Figure 8 computer device structure shown.
[0221] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.
[0222] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.
[0223] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0224] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0225] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0226] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0227] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0228] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for recalling articles. It is characterized in that include: Inputting a paragraph of a first article into a first model to output a first paragraph label, wherein the first article is an article outside a preset sorting range of a recommendation list, and the first model is trained and generated using first training data, wherein the first training data includes a first data pair of an article paragraph and a first label; Inputting the paragraph of the first article into a second model to obtain a second paragraph label, wherein the second model is trained and generated using second training data, wherein the second training data includes a second data pair of the article paragraph and a matching phrase, wherein the matching phrase includes an intersection of the first label and a template phrase in a template phrase library, wherein the template phrase is a word that appears at least a first preset number of times repeatedly in a phrase whose flow rate during a preset period is less than a threshold value; Determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively; increasing the relevance score of the first article and the first query phrase whose matching degree is greater than a preset threshold; The ranking of the first article in the recommendation list is adjusted according to a relevance score between the first article and the first query phrase.
2. The method according to claim 1, It is characterized in that The first training data further includes a fourth data pair and a public data set. Before inputting the paragraph of the first article into the first model to output the first paragraph label, the method further includes: Acquire the fourth data pair, the fourth data pair being a data pair whose target object click times exceed the second preset times and whose relevance score is greater than the first preset score, and the fourth data pair including the third article and the third query phrase; The public dataset is obtained, where the public dataset includes public articles and phrases.
3. The method according to claim 1 or 2, It is characterized in that Before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes: The article paragraph is labeled to obtain the first label.
4. The method according to claim 1, It is characterized in that Before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes: Data pairs in the first training data whose matching degree between the article paragraph and the first label is outside a first preset range are filtered.
5. The method according to claim 1, It is characterized in that Before inputting the paragraph of the first article into the second model to obtain a second paragraph label, the method further includes: Obtain unpopular phrases with a preset period flow rate less than a threshold; Clustering the unpopular phrases to select high-frequency participles that appear at least the first preset number of times; Phrases in the high-frequency word segmentation whose matching degree with the corresponding second article is outside a second preset range are filtered to obtain the template phrases.
6. The method according to claim 1, It is characterized in that Before inputting the paragraph of the first article into the second model to obtain a second paragraph label, the method further includes: Filter entity words in the template phrase.
7. The method according to claim 1, It is characterized in that Before inputting the paragraph of the first article into the second model to obtain a second paragraph label, the method further includes: Taking the article paragraph and the matching phrase whose relevance score is greater than a second preset score as positive samples; Randomly extracting template phrases and corresponding articles from the template phrase library as negative samples; The second model is trained according to the positive samples and the negative samples.
8. The method according to claim 1, It is characterized in that Inputting the paragraph of the first article into the second model to obtain a second paragraph label includes: Inputting the paragraph of the first article into the second model to output a plurality of candidate tags whose cosine similarity between the paragraph of the first article and the template phrase meets the requirement; Determine the relevance scores of the multiple candidate tags and the paragraphs of the first article, and the second paragraph tag is the candidate tag whose relevance score meets the first condition.
9. The method according to claim 1, It is characterized in that Before inputting the paragraph of the first article into the first model to output the first paragraph tag, the method further includes: generating the recommendation list according to a third query phrase; Obtaining the first query phrase; When the first query phrase is the same as the third query phrase, the step of inputting the paragraph of the first article into the first model to output a first paragraph tag is triggered.
10. The method according to claim 1, It is characterized in that Before determining the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively, the method further includes: Adjusting the rough ranking score of the first article according to the first query phrase, the first paragraph tag, and the second paragraph tag; Adding a mark to the first article whose rough ranking score satisfies the second condition; The marked first article is merged into the recommendation list.
11. The method according to claim 10, It is characterized in that The step of adjusting the ranking of the first article in the recommendation list according to the relevance score between the first article and the first query phrase includes: The marked first article in the recommendation list is moved forward to be inserted before a fourth article, and a relevance score of the fourth article to the first query phrase is less than a relevance score of the marked first article to the first query phrase.
12. An article recall device, It is characterized in that include: an input unit, configured to input a paragraph of a first article into a first model to output a first paragraph label, wherein the first article is an article outside a preset sorting range of a recommendation list, and the first model is trained and generated using first training data, wherein the first training data includes a first data pair of an article paragraph and a first label; input the paragraph of the first article into a second model to obtain a second paragraph label, wherein the second model is trained and generated using second training data, wherein the second training data includes a second data pair of the article paragraph and a matching phrase, wherein the matching phrase includes an intersection of the first label and a template phrase in a template phrase library, wherein the template phrase is a word that appears at least a first preset number of times in a phrase whose preset period flow is less than a threshold; A determination unit, configured to determine the matching degree between the first query phrase to be queried and the first paragraph tag and the second paragraph tag respectively; The adjustment unit is used to increase the relevance score between the first article and the first query phrase whose matching degree is greater than a preset threshold; and adjust the ranking of the first article in the recommendation list according to the relevance score between the first article and the first query phrase.
13. A computer device, It is characterized in that include: memories, transceivers, processors, and bus systems; Wherein, the memory is used to store programs; The processor is used to execute the program in the memory, including executing the method according to any one of claims 1 to 11; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
14. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 11.
15. A computer program product, It is characterized in that When the computer program product is executed on a computer, the computer performs the method according to any one of claims 1 to 11.