Vocabulary recommendation methods, devices, terminals, and readable storage media
By mining candidate words related to the target words from the vocabulary database and recommending new words based on relevance, the problem of insufficient relevance of advertising words in existing technologies is solved, and more extensive and accurate advertising word recommendations are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-07
- Publication Date
- 2026-04-03
AI Technical Summary
In existing internet advertising, the determination of advertising keywords is limited, resulting in poor relevance of similar keywords to the client's business, which affects the effectiveness of ad exposure.
By acquiring the target account's keywords, candidate keywords related to the target keywords are identified from the keyword database. Based on the relevance of the candidate keywords to the target recommended content, new keywords are recommended to improve the advertising effect.
It expands the vocabulary coverage, improves the accuracy and efficiency of recommended words, and ensures the relevance of recommended words to the target content.
Smart Images

Figure CN114595377B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a vocabulary recommendation method, apparatus, terminal, and readable storage medium. Background Technology
[0002] Internet advertising refers to commercial advertising that promotes goods or services directly or indirectly through internet media such as websites, web pages, and internet applications, using text, images, audio, video, or other forms. It is an emerging form of advertising media. With the development of cloud technology, its application is gradually appearing in the internet advertising industry. For example, big data is used to mine users' frequently used search terms and utilize these terms in advertising promotion to increase ad exposure.
[0003] Clients can purchase ad keywords related to their business for ad placement. When a user searches for these keywords, relevant ad content will be displayed. For example, if a client purchases the keyword "elementary school English tutoring" from a supplier, then when a user searches on the supplier's search website and the search terms include "elementary school English tutoring," the search results will include the client's relevant business ad content. In this technology, the ad keyword is generated by the client through creative analysis of their own business, and the supplier then recommends similar keywords to the client.
[0004] However, advertising copy determined in this way has limitations, and the resulting similar keywords may be poorly relevant to the client's business. Summary of the Invention
[0005] This application provides a vocabulary recommendation method, apparatus, terminal, and readable storage medium, which can improve the accuracy of vocabulary selection when recommending targeted content to target accounts. The technical solution is as follows:
[0006] On the one hand, a vocabulary recommendation method is provided, the method comprising:
[0007] Obtain the target account's keywords for ad placement. The keywords are the words used by the target account when ad placement of target recommended content. The target recommended content is used as the search results for the keywords for content recommendation.
[0008] Based on the targeted vocabulary, at least one candidate vocabulary that is related to the targeted vocabulary is determined from the vocabulary database;
[0009] Based on the relevance between the at least one candidate word and the target recommended content, recommended words are determined from the at least one candidate word;
[0010] The recommended keywords are sent to the target account. The recommended keywords are words used when delivering the target recommended content to the target account.
[0011] On the other hand, a vocabulary recommendation device is provided, the device comprising:
[0012] The acquisition module is used to acquire the target account's keywords for ad placement. The keywords are the words selected by the target account when ad placement of target recommended content. The target recommended content is used as the search results for the keywords for content recommendation.
[0013] The determining module is used to determine at least one candidate word that is related to the targeted words from the word library based on the targeted words;
[0014] The determining module is further configured to determine recommended words from the at least one candidate words based on the relevance between the at least one candidate word and the target recommended content;
[0015] The sending module is used to send the recommended words to the target account. The recommended words are words selected for delivering the target recommended content to the target account.
[0016] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement any of the vocabulary recommendation methods described in the embodiments of this application.
[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the program code being loaded and executed by a processor to implement any of the vocabulary recommendation methods described in the embodiments of this application.
[0018] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the vocabulary recommendation methods described in the above embodiments.
[0019] The technical solution provided in this application includes at least the following beneficial effects:
[0020] Based on the target account's target keywords, candidate keywords with relevance to the target keywords are identified from the keyword database. Then, recommended keywords are determined based on the relevance between the candidate keywords and the target recommended content, and the recommended keywords are sent to the target account. This enables the mining of new keywords for target recommended content based on existing target keywords, making the recommended keywords more comprehensive without losing their relevance to the target recommended content, thus improving the efficiency and accuracy of pushing recommended keywords to the target account. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;
[0023] Figure 2 This is a flowchart of a vocabulary recommendation method provided in an exemplary embodiment of this application;
[0024] Figure 3 This is a flowchart of a vocabulary recommendation method provided in another exemplary embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the training relevance model features provided in an exemplary embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the structured model branch provided in an exemplary embodiment of this application;
[0027] Figure 6 This is a schematic diagram of structural branching information provided in an exemplary embodiment of this application;
[0028] Figure 7 This is a schematic diagram of a machine translation verification model provided in an exemplary embodiment of this application;
[0029] Figure 8 This is a schematic diagram of a pavilion system for a vocabulary recommendation method provided in an exemplary embodiment of this application;
[0030] Figure 9 This is a schematic diagram of a vocabulary database provided in an exemplary embodiment of this application;
[0031] Figure 10 This is a schematic diagram of a recommended vocabulary generation module provided in an exemplary embodiment of this application;
[0032] Figure 11 This is a schematic diagram of a platform system module provided in an exemplary embodiment of this application;
[0033] Figure 12 This is a structural block diagram of a vocabulary recommendation apparatus provided in an exemplary embodiment of this application;
[0034] Figure 13 This is a structural block diagram of a vocabulary recommendation apparatus provided in another exemplary embodiment of this application;
[0035] Figure 14 This is a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0037] First, a brief introduction to the terms used in the embodiments of this application:
[0038] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.
[0039] Big data refers to data sets that cannot be captured, managed, and processed within a certain timeframe using conventional software tools. It represents massive, rapidly growing, and diverse information assets that require new processing models to achieve stronger decision-making, insightful discovery, and process optimization capabilities. With the advent of the cloud era, big data has attracted increasing attention. Big data requires specialized technologies to effectively process large amounts of data within a tolerable timeframe. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems. In this application embodiment, big data is used to achieve vocabulary mining and expansion of a vocabulary database.
[0040] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0041] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0042] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs. For example, recommended vocabulary for a target account can utilize machine translation technology to generate relevant candidate words based on the target account's target vocabulary.
[0043] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. For example, when verifying candidate words using a preset verification method, an artificial neural network can be combined with machine translation to calculate the relevance between candidate words and the input words. In one example, this artificial neural network is a recurrent neural network (RNN).
[0044] Based on the above definitions, the application scenarios of the embodiments of this application will be described.
[0045] After a user searches for keywords using a search engine, the webpage displays multiple links to webpages related to that keyword, each with corresponding recommended content. Internet advertising, as an emerging advertising medium, is diverse in form and can directly or indirectly promote goods or services. Combining search engines with internet advertising provides a highly efficient content promotion method. In one example, if a user searches for "primary school English tutoring," the webpage displays multiple recommended content items based on that keyword. Advertisers can bid on the keyword "primary school English tutoring," and the highest bidder can target specific recommended content for that keyword. For example, the target recommended content might be "XX Education Primary School English Training Courses." When a user searches for this keyword, the webpage will display "XX Education Primary School English Training Courses" and its corresponding link. Subsequently, the user may click on the link, thus promoting the targeted recommended content and increasing the advertiser's business exposure.
[0046] After advertisers participate in bidding for or purchase keywords, they can target specific content for those keywords. These keywords are then used for keyword bidding. Generally, advertisers conduct creative analysis based on their business scope to determine the keywords they want to bid on, and then participate in the keyword bidding process. For example, if an advertiser's business is piano tutoring, creative analysis might suggest suitable keywords such as "piano tutor," "how to learn piano from scratch," and "piano beginners." However, for such keywords, there are many competitors bidding, resulting in higher bids. In this embodiment, candidate keywords related to the target account's bidding keywords are obtained from a keyword database. Recommended keywords are determined based on the relevance of these candidate keywords to the target content. These recommended keywords are then pushed to the target account as a reference for keyword bidding. This keyword recommendation method provides a wider coverage of recommended keywords without sacrificing relevance to the target content. Furthermore, this method can also uncover keywords with lower bids that can still bring traffic to the target content.
[0047] The above description uses the application of the vocabulary recommendation method in Internet advertising as an example. The vocabulary recommendation method in this application can be applied to other scenarios that require the generation of recommended vocabulary, and is not limited here.
[0048] Secondly, the implementation environment of the embodiments of this application will be described in conjunction with the above application scenarios and definitions. For illustrative purposes, please refer to... Figure 1 The real environment includes terminal 101, server 102 and communication network 103.
[0049] Terminal 101 can be an electronic device such as a mobile phone, tablet computer, e-book reader, multimedia playback device, wearable device, laptop computer, desktop computer, or biometric all-in-one machine. Illustratively, terminal 101 has a target application installed, which enables advertisers to purchase or bid on keywords, generate recommended keywords, and perform other similar functions. Illustratively, this target application can be traditional application software, cloud application software, a mini-program within a host application, or a web platform.
[0050] Server 102 provides keyword purchasing, keyword bidding, and keyword recommendation services to terminal 101. Server 102 stores a keyword database. For example, the keywords in the database are obtained by mining big data collected from search engines. Each keyword in the database also includes keyword information such as keyword category, part of speech, length, purchase price, bidding price, and usage rate. In one example, server 102 receives a keyword selection signal from terminal 101. Based on the keyword identifier carried in the selection signal, it reads the corresponding keyword and its corresponding keyword information from the memory. Based on this keyword information, it returns a corresponding purchase or bidding signal to terminal 101, and terminal 101 displays the corresponding interface based on the signal. In another example, server 102 receives an account identifier from terminal 101. Based on this account identifier, it determines the corresponding purchased keywords and generates recommended keywords based on these keywords. The recommended keywords are then pushed to terminal 101, and terminal 101 displays the recommended keywords. Optionally, server 102 can be a physical server or a cloud server. Server 102 can be a single server, a server cluster consisting of several servers, or a cloud computing service center.
[0051] Server 102 can establish a communication connection with terminal 101 through communication network 103. This network can be a wireless network or a wired network.
[0052] Please refer to Figure 2 The document illustrates a flowchart of a vocabulary recommendation method according to an embodiment of this application. The method may include the following steps:
[0053] Step 201: Obtain the target keywords for the target account.
[0054] In this embodiment, the keyword recommendation method can be applied to content recommendation software or platforms. Taking its application in an advertising platform as an example, the target account is the account used by the first user who needs to place ads; the first user is the advertiser. The first user can use the target account to log in to the aforementioned advertising platform and place ads related to their own business through the platform. Illustratively, the target account can obtain keywords for placement on the advertising platform, for example, through direct purchase or auction. These keywords are the words selected by the target account when placing target recommended content. The target recommended content is used as the search result for the keywords for content recommendation, and there is a corresponding relationship between the keywords and the target account.
[0055] Target accounts can obtain targeted keywords by redeeming virtual resources from the platform. In one example, the first user of the target account selects target keywords related to their business from the keyword library. This is illustrative; the keywords in the keyword library correspond to the quantity of virtual resources on the platform. When the target account redeems the target keywords using the corresponding quantity of virtual resources, those keywords become the target account's targeted keywords, thus establishing a connection with the target account. This virtual resource can be virtual coins or virtual items.
[0056] Target accounts can acquire keywords through competition with other accounts. In one example, because multiple accounts have similar business content, and their respective first users select the same target keyword, the target account can compete with other accounts to acquire the target keyword. When the target account wins the bid, it pays the corresponding bid price, and that target keyword becomes the target account's advertising keyword, establishing a direct relationship with the target account. Alternatively, when multiple accounts bid for the same target keyword, the target recommended content for each participating account is sorted and displayed according to the bid price paid. For example, if account A (recommended content A, bid price 500 yuan), account B (recommended content B, bid price 400 yuan), and account C (recommended content C, bid price 1000 yuan) simultaneously bid for the target keyword, then when searching for the target keyword using a search engine, the displayed content will be in the order of recommended content C, recommended content A, and recommended content B.
[0057] In this embodiment of the application, the vocabulary to be delivered can correspond to a single word, such as "English", or to a vocabulary composed of multiple words, such as "English tutoring", or to a short sentence, such as "What to do if my child is not good at learning English", and there is no limitation here.
[0058] In this embodiment, after obtaining the target keywords, the target account can select its target recommended content and establish a correspondence with the target keywords. The advertising platform will then recommend content to the target account based on the target keywords. Optionally, one target recommended content can correspond to multiple target keywords.
[0059] Once the advertising platform obtains the target account's existing keywords, it can recommend other target keywords to the target account as a reference for keyword selection.
[0060] Step 202: Based on the target vocabulary, identify at least one candidate vocabulary that is related to the target vocabulary from the vocabulary database.
[0061] In this embodiment, the advertising platform's server stores a vocabulary database. Illustratively, the vocabulary in this database is obtained through word mining of search data accumulated by search engines, and includes search terms with a certain search frequency from the search engines; the database also includes terms entered by the target account. Illustratively, there may be relationships between words in the vocabulary database, including at least one of semantic, structural, grammatical, and indexed relationships. For example, "body odor" and "bromhidrosis" have a semantic relationship; "Wuxi cuisine" and "Huishan cuisine" have a structural relationship; "learning piano fingering" and "learning piano fingering" have a grammatical relationship; "parking lot" and "underground parking lot" are adjacent in the inverted index in the vocabulary database, therefore "parking lot" and "underground parking lot" have an indexed relationship.
[0062] Methods for identifying at least one candidate keyword based on the target account's target keywords include at least one of the following methods:
[0063] First, determine the first candidate keywords based on the behavioral characteristics of the targeted keywords. These behavioral characteristics indicate the features corresponding to the targeted keywords when they are searched.
[0064] Behavioral features include co-occurrence features, which are the probabilities of words appearing together in a given context. For example, if the target account's target recommendation content is English tutoring courses from AA Education Institution, and the target account's existing keyword is "poor English learning," based on the co-occurrence features of this keyword, in the context corresponding to "poor English learning," the word with a high probability of co-occurrence is likely "poor English grades," thus "poor English grades" is determined as the first candidate keyword. Behavioral features also include click features, which are the other links clicked by a second user when searching for the keyword, corresponding to the keyword owned by the first user. Here, the second user is a user using a search engine to search for keywords. For example, if the target account's existing keyword is "poor English learning," and after a second user searches for "poor English learning," the webpage displays multiple content recommendation links including the current target account's target recommendation content "English tutoring courses from AA Education Institution," and other accounts' recommendation content "BB Education Institution one-on-one English tutoring," and these other accounts also have other keyword recommendations for "English tutoring," then "English tutoring" is selected as the second candidate keyword. Behavioral characteristics also include historical characteristics, which are the first candidate words that the target account has previously selected in the vocabulary library but has not yet used for advertising.
[0065] Second, determine the second candidate words based on the grammatical features of the target words.
[0066] This syntactic feature is determined by machine translation (MT) technology, that is, machine translation is performed on the input vocabulary to determine the second candidate vocabulary. For example, if the input vocabulary of the target account is "Learn piano fingering", machine translation can obtain multiple results such as "Learn piano fingering", "ピアノの運指を学ぶ", "Learn about piano fingering", etc. At least one of the above machine translation results is determined as the second candidate vocabulary.
[0067] Third, determine the third candidate vocabulary according to the structural features of the input vocabulary.
[0068] This structural feature is used to indicate the composition structure of the vocabulary. Schematically, the input vocabulary is divided into at least two branches according to the structure, and new vocabulary is generated according to the structural information corresponding to each branch as the third candidate vocabulary. In one example, the input vocabulary is structurally divided into three branches: <entity, intent, region>. For example, if the input vocabulary of the target account is "How much is the housing price in Huishan District", its entity corresponds to the housing price, the intent corresponds to the price, and the region corresponds to Huishan District. The third candidate vocabulary that can be generated according to the structural information corresponding to each of its branches includes "How much is the housing price in Wuxi", "What are the real estate projects in Huishan District", etc.
[0069] Fourth, perform a Boolean search on the vocabulary library to obtain the fourth candidate vocabulary.
[0070] This Boolean search refers to using Boolean logic operators to connect the input vocabulary and the vocabulary in the vocabulary library, and then the computer performs corresponding logical operations to determine the fourth candidate vocabulary. In one example, an inverted search is performed on the vocabulary in the vocabulary library to determine the fourth candidate vocabulary that has an index association relationship with the input vocabulary.
[0071] Schematically, according to the above first candidate vocabulary, second candidate vocabulary, third candidate vocabulary, and fourth candidate vocabulary, determine at least one candidate vocabulary. The above first candidate vocabulary, second candidate vocabulary, third candidate vocabulary, and fourth candidate vocabulary are all determined as candidate vocabulary for relevance determination. There may be the same candidate vocabulary among the first candidate vocabulary, second candidate vocabulary, third candidate vocabulary, and fourth candidate vocabulary. At least one candidate vocabulary for relevance determination is determined according to the repetition rate corresponding to the above candidate vocabulary.
[0072] Step 203, determine the recommended vocabulary from at least one candidate vocabulary based on the relevance between at least one candidate vocabulary and the target recommended content.
[0073] In the embodiments of the present application, after determining at least one candidate vocabulary, it is necessary to filter the above at least one candidate vocabulary.
[0074] Optionally, business filtering is performed on at least one of the above candidate words. For illustration, the target account also corresponds to negative words, which are words already negative by the first user in the word library; the target account also corresponds to the target recommended content's target geographic region. Business filtering is performed on at least one of the above candidate words based on the aforementioned negative words and the target recommended content's target geographic region.
[0075] Optionally, the candidate words are filtered based on the relevance of at least one candidate word to the target recommended content. Illustratively, relevance data between at least one candidate word and the target recommended content is determined based on a preset verification method. This preset verification method determines the relevance between the candidate search results obtained when using at least one candidate word as a search request and the target recommended content. Recommended words are then determined from the at least one candidate word based on the relevance data. The relevance data includes at least one of semantic relevance, structural relevance, and grammatical relevance. Semantic relevance represents the degree of semantic association between words. In one example, semantic relevance is obtained from a relevance model trained on a corpus, which is obtained by crawling content from various web pages on the Internet, i.e., a web crawler. Structural relevance represents the degree of structural association between words. In one example, structural relevance is obtained from a structured model that divides the words in the vocabulary into a preset number of structural branches according to their compositional structure and determines the structural relevance between words based on the branch information corresponding to the same structural branches of different words. Syntactic relevance is used to represent the degree of grammatical association between words. In one example, syntactic relevance is obtained by a machine translation verification model, which is trained using machine translation techniques based on recurrent neural networks.
[0076] Step 204: Send recommended keywords to the target account.
[0077] In this embodiment of the application, the recommended words are words selected for use when delivering target recommended content to the target account. That is, the recommended words are pushed to the target account, and the first user of the target account uses them as a reference for bidding on the recommended words.
[0078] In this embodiment of the application, when the number of candidate words is at least two, the candidate words are sorted according to the relevance data to obtain a candidate word sequence; based on the candidate word sequence, a preset number of recommended words are obtained.
[0079] Before sending recommended keywords to the target account, the server can also perform a legal review of the recommended keywords. In one example, this legal review includes AKA (Authentication and Key Agreement) review and competitor filtering. The AKA review checks the authentication information of the target account and whether the recommended keywords contain prohibited words, while the competitor filtering filters out keywords from the recommended keywords that are from competing products of the target account's target recommended content.
[0080] In summary, the vocabulary recommendation method provided in this application, based on the target account's target vocabulary, determines candidate words related to the target vocabulary from the vocabulary library, then determines recommended words based on the relevance of the candidate words to the target recommended content, and sends the recommended words to the target account. This realizes the mining of new words for use when delivering target recommended content based on existing target vocabulary, making the recommended vocabulary coverage wider, without losing relevance to the target recommended content, thus improving the efficiency and accuracy of pushing recommended words to the target account.
[0081] In the process of obtaining candidate words in this application embodiment, candidate words can be determined based on multiple word features of the target words. Furthermore, during the filtering process of candidate words, the relevance data between the candidate words and the target recommended content can be calculated based on various relevance factors. Please refer to... Figure 3 The diagram illustrates a flowchart of a vocabulary recommendation method according to an embodiment of this application. The method is applied in a server and may include the following steps:
[0082] Step 301: Obtain the target keywords for the target account.
[0083] In this embodiment, the server obtains the account identifier corresponding to the target account currently logged in by the terminal, and retrieves the target keywords corresponding to the account identifier from the memory.
[0084] Step 3021: Determine the first candidate word based on the behavioral characteristics of the delivered words.
[0085] This behavioral feature indicates the characteristics corresponding to the targeted keywords when they are searched. Illustratively, the behavioral feature includes co-occurrence features, i.e., the probability of two words co-occurring in a certain context; it also includes click features, i.e., the targeted keywords owned by the first user corresponding to other links clicked by the second user when searching for the targeted keywords; and it also includes historical features, i.e., keywords that the target account has previously selected in the keyword database but has not yet been targeted, as candidate keywords. After determining the behavioral feature of the targeted keyword, the server searches the keyword database for other keywords with the above behavioral features as first candidate keywords. Optionally, the above behavioral feature includes at least one of co-occurrence features, click features, and historical features.
[0086] Step 3022: Determine the second candidate words based on the grammatical features of the submitted words.
[0087] The server performs machine translation on the above-mentioned target words to obtain at least one translated word, which is then identified as a second candidate word. Optionally, the translated word can be a word in the same language as the target words, or it can be a word in a different language.
[0088] Step 3023: Determine the third candidate word based on the structural characteristics of the target words.
[0089] This structural feature is used to indicate the compositional structure of the vocabulary. The server divides the delivered vocabulary into at least two branches according to their structure, and generates new vocabulary based on the structural information corresponding to each branch, which serves as the third candidate vocabulary.
[0090] Step 3024: Perform a Boolean search on the vocabulary database to obtain the fourth candidate vocabulary.
[0091] The server performs an inverted search on the vocabulary database to determine the fourth candidate vocabulary that has an indexed relationship with the target vocabulary.
[0092] Step 303: Determine at least one candidate word based on the first candidate word, the second candidate word, the third candidate word, and the fourth candidate word.
[0093] Optionally, the first, second, third, and fourth candidate words mentioned above can all be selected as candidate words for relevance determination.
[0094] Optionally, there may be the same candidate words among the first, second, third and fourth candidate words. At least one candidate word is determined for relevance determination based on the repetition rate of the above candidate words.
[0095] Optionally, the candidate words determined based on the aforementioned behavioral features, grammatical features, structural features, and Boolean retrieval correspond to quality levels. In one example, candidate words generated based on grammatical features correspond to the first level, candidate words generated based on structural features correspond to the second level, candidate words generated based on Boolean retrieval correspond to the third level, and candidate words generated based on behavioral features correspond to the fourth level. The first level corresponds to the highest quality, and the fourth level corresponds to the lowest quality. Illustratively, the target account can specify the quality level requirements for the recommended words. The server then filters the first, second, third, and fourth candidate words according to these quality level requirements to obtain at least one candidate word.
[0096] Step 3041: Perform a relevance check on at least one candidate word to determine the semantic relevance between at least one candidate word and the target word.
[0097] In this embodiment, at least one candidate word includes a target candidate word. The server extracts a first feature of the target candidate word and a second feature of the target word; the first and second features are input into a relevance model to obtain semantic relevance, wherein the relevance model is trained on corpus data.
[0098] In one example, please refer to Figure 4The diagram illustrates the feature types used to train the relevance model 410, including MT features 420, text attribute class features 430, extended text features 440, DNN (Deep Neural Networks) features 450, and other features 460. The text attribute class features 430 are jointly determined by IDF (Inverse Document Frequency) 431, word vectors 432, part-of-speech tags 433, proper nouns 434, length 435, number of terms 436, and number of non-terms 437. The proper noun feature indicates whether the target word is a proper noun, and the number of terms feature indicates the number of terms in the target word; a term is the basic unit for logical analysis of words by a computer. The extended text features 440 are jointly determined by the BM25 algorithm 441, summary extended terms 442, PLSA (Probabilistic Latent Semantic Analysis) 443, and substitution points 444. DNN feature 450 is determined jointly by single-word vectors 451, phrase2vec (Phrase Embedding Based on Parsing) 452, and word2vec (Word to Vector) 453. Other features 460 include industry features 461 and outliers 462, where outliers are used to represent the degree of difference between the target candidate word and other candidate words.
[0099] As an illustration, the relevance model can be trained using RF (Random Forest) or GBDT (Gradient Boosting Decison Tree). In one example, the relevance model underwent multiple rounds of iterative training, and the results are shown in Table 1. To ensure the quality of the recommended words, the relevance model callback value was set between 0.25 and 0.3, prioritizing quality. The online test set was obtained by sampling and labeling words from the vocabulary database after the recommended words were determined, while the offline test set was obtained by randomly sampling words from the vocabulary database.
[0100] Table 1:
[0101]
[0102] Step 3042: Perform structured verification on the at least one candidate word to determine the structural relevance between the at least one candidate word and the word being delivered.
[0103] In this embodiment, the server inputs target candidate words into a structured model and outputs first branch information corresponding to each structure branch. The structured model includes at least two structure branches. The server inputs target words into the structured model and outputs second branch information corresponding to each structure branch. The server compares the first branch information and the second branch information to determine the structural relevance of the target candidate words.
[0104] This structured model is used to extract structured information from target candidate words. In one example, there are six structural branches: "region, gender, body part, industry, entity, and intent". Please refer to [reference needed]. Figure 5 When the input term in the structured model is "How much does breast augmentation surgery cost in Xierqi?", the following branches are possible: Region 510 corresponds to Beijing; optionally, region 510 can be further divided into primary region 511 and secondary region 512, with primary region 511 corresponding to Beijing and secondary region 512 corresponding to Haidian District; Gender 520 corresponds to female; Body part 530 corresponds to breast; Industry 540 corresponds to medical; optionally, industry 540 can be further divided into primary industry 541 and secondary industry 542, with primary industry 541 corresponding to medical and secondary industry 542 corresponding to cosmetic surgery; Entity 550 corresponds to breast augmentation surgery; Intent 560 corresponds to price.
[0105] The server inputs the target candidate words into the structured model and extracts the corresponding structured information, i.e., the first branch information. Simultaneously, it inputs the target words into the structured model and extracts the corresponding structured information, i.e., the second branch information. The server compares the first and second branch information to determine the structural relevance of the target candidate words. For illustration, this structural relevance is determined based on the degree of matching between the same branches in the first and second branch information. Please refer to [reference needed]. Figure 6 The target word 610 corresponds to "how to treat body odor in a seven-year-old girl in Wuhan" and the target candidate word 620 corresponds to "how to treat body odor in a seven-year-old boy in Wuhan". The target word 610 and the target candidate word 620 are respectively input into the structured model for structured information extraction to obtain the first branch information 611 and the second branch information 621. The first branch information 611 and the second branch information 621 are matched and judged 602 to determine whether to filter the target candidate word 603.
[0106] Step 3043: Perform machine translation verification on at least one candidate word to determine the grammatical relevance between at least one candidate word and the input word.
[0107] In this embodiment of the application, the server inputs at least one candidate word into the machine translation verification model and outputs recombined candidate words; the grammatical relevance of the target candidate word is determined based on the degree of difference between the target candidate word and the recombined candidate word.
[0108] In one example, please refer to Figure 7 The diagram illustrates a machine translation verification model. Candidate word A corresponds to "barbecue franchise" 701, and candidate word B corresponds to "barbecue skewer agency franchise" 702. Candidate words A and B are processed by encoder 703 and semantic extraction 704, respectively, followed by semantic connection 705. After passing through encoder 703 again, each corresponding term is input into a softmax logistic regression model 706 for encoding, ultimately resulting in a recombined candidate word "barbecue skewer franchise". The grammatical relevance of the target candidate word is determined by comparing the degree of difference between the recombined candidate word and the target candidate word. In other words, during the verification process, all candidate words are used as input, and the output words are segmented to obtain multiple terms. The recombined candidate words are determined based on the probability corresponding to the dictionary size generated at the time series point.
[0109] Step 305: Determine the relevance data of at least one candidate word based on semantic relevance, structural relevance, and grammatical relevance.
[0110] Semantic relevance, structural relevance, and grammatical relevance each have different weights, and the relevance data of candidate words is calculated based on these weights. In one example, the weight for semantic relevance is 0.5, the weight for structural relevance is 0.3, and the weight for grammatical relevance is 0.2. The relevance data for each candidate word is calculated based on its semantic, structural, and grammatical relevance.
[0111] Step 306: Based on relevance data, determine recommended words from at least one candidate word.
[0112] In this embodiment, when the number of candidate words is at least two, the server sorts the candidate words according to relevance data to obtain a candidate word sequence; based on the candidate word sequence, a preset number of recommended words are obtained. Indicatively, this preset number can be determined by the target account or pre-set by the server.
[0113] Step 307: Send recommended keywords to the target account.
[0114] In this embodiment of the application, the recommended words are words selected for use when delivering target recommended content to the target account. That is, the recommended words are pushed to the target account, and the first user of the target account uses them as a reference for bidding on the recommended words.
[0115] In summary, the vocabulary recommendation method provided in this application determines candidate words from a vocabulary database based on the behavioral, grammatical, and structural features of the target account's target vocabulary. Then, it determines recommended words based on the semantic, structural, and grammatical relevance of the candidate words to the target recommended content and sends the recommended words to the target account. This method achieves the goal of mining new words for target recommended content based on existing target vocabulary, resulting in a wider coverage of recommended words without sacrificing relevance to the target recommended content, thus improving the efficiency and accuracy of pushing recommended words to the target account.
[0116] Please refer to Figure 8 This illustrates that the vocabulary recommendation method in this application is applied to an internet promotion platform system. The platform system includes a vocabulary selection module 810, a recommended vocabulary generation module 820, and a vocabulary delivery module 830. The data processing and transmission steps between the modules include:
[0117] Step 811: The vocabulary selection module 810 receives the target account's selection signal for the target vocabulary in the vocabulary database.
[0118] Step 812: In response to determining that the target account is qualified to obtain the target keywords, the keyword selection module 810 identifies the target keywords as the target account's keywords for advertising.
[0119] In one example, please refer to Figure 9 The keywords 901 in the vocabulary database 900 are divided into multiple advertising units 902 based on their vocabulary characteristics. For example, the vocabulary in the vocabulary database 900 is classified according to the domain characteristics corresponding to the vocabulary, and is divided into advertising units such as "medical", "education", "entertainment", and "building materials".
[0120] Step 821: The recommended vocabulary generation module 820 obtains the target account's target vocabulary from the target vocabulary selection module 810.
[0121] Step 822: The recommended vocabulary generation module 820 determines at least one candidate vocabulary that is related to the target vocabulary from the vocabulary database.
[0122] Step 823: The recommended vocabulary generation module 820 determines recommended vocabulary from at least one candidate vocabulary based on the relevance of at least one candidate vocabulary to the target recommended content.
[0123] Step 824: The recommended vocabulary generation module 820 sends recommended vocabulary to the terminal corresponding to the target account and the vocabulary delivery module 830.
[0124] The recommended vocabulary generation module 820 is further configured to determine the relevance data between at least one candidate vocabulary and the target recommended content based on a preset verification method. The preset verification method is used to determine the degree of relevance between the candidate search results obtained and the target recommended content when at least one candidate vocabulary is used as a search request. Based on the relevance data, the recommended vocabulary is determined from at least one candidate vocabulary.
[0125] The relevance data includes semantic relevance. The recommended vocabulary generation module 820 is also used to perform relevance verification on at least one candidate vocabulary to determine the semantic relevance between at least one candidate vocabulary and the delivered vocabulary; to perform structural verification on at least one candidate vocabulary to determine the structural relevance between at least one candidate vocabulary and the delivered vocabulary; and to perform machine translation verification on at least one candidate vocabulary to determine the grammatical relevance between at least one candidate vocabulary and the delivered vocabulary.
[0126] At least one candidate word includes a target candidate word. The recommended word generation module 820 is also used to extract a first feature of the target candidate word and a second feature of the proposed word. The first feature and the second feature are input into the relevance model to obtain the semantic relevance. The relevance model is obtained by training on the corpus.
[0127] The recommended vocabulary generation module 820 is also used to input target candidate vocabulary into a structured model and output the first branch information corresponding to each structure branch. The structured model includes at least two structure branches. It also inputs the target vocabulary into the structured model and outputs the second branch information corresponding to each structure branch. By comparing the first branch information and the second branch information, the structural relevance of the target candidate vocabulary is determined.
[0128] The recommended vocabulary generation module 820 is also used to input at least one candidate vocabulary into the machine translation verification model and output recombined candidate vocabulary; and to determine the grammatical relevance of the target candidate vocabulary based on the degree of difference between the target candidate vocabulary and the recombined candidate vocabulary.
[0129] The recommended vocabulary generation module 820 is also used to sort the candidate words according to the relevance data to obtain a candidate word sequence when the number of candidate words is at least two; and to obtain a preset number of recommended words based on the candidate word sequence.
[0130] The recommended vocabulary generation module 820 is also used to determine the first candidate vocabulary based on the behavioral features of the submitted vocabulary, the behavioral features being used to indicate the features corresponding to the submitted vocabulary when it is searched; determine the second candidate vocabulary based on the grammatical features of the submitted vocabulary; determine the third candidate vocabulary based on the structural features of the submitted vocabulary, the structural features being used to indicate the composition structure of the vocabulary; perform a Boolean search on the vocabulary database to obtain the fourth candidate vocabulary; and determine at least one candidate vocabulary based on the first candidate vocabulary, the second candidate vocabulary, the third candidate vocabulary, and the fourth candidate vocabulary.
[0131] In one example, please refer to Figure 10 The recommended vocabulary generation module 1000 is further divided into a candidate vocabulary generation unit 1010 and a candidate vocabulary verification unit 1020. The candidate vocabulary generation unit 1010 includes an MT generation subunit 1011, a behavioral feature mining subunit 1012, a vocabulary database inverted index retrieval subunit 1013, and a structured association subunit 1014. The candidate vocabulary verification unit 1020 includes an MT verification subunit 1021, a relevance verification subunit 1022, a structured verification subunit 1023, and an adoption rate model subunit 1024. To illustrate, the adoption rate model subunit 1024 filters candidate vocabulary based on the historical adoption rate of candidate vocabulary.
[0132] Step 831: The vocabulary delivery module 830 receives the target account's selection operation for at least one target recommended vocabulary from the recommended vocabulary.
[0133] Step 832: In response to determining that the target account is qualified to obtain the target recommended keywords, the keyword delivery module 830 determines the target recommended keywords as the delivery keywords for the target account.
[0134] Step 833: The vocabulary delivery module 830 receives the target account's selection signal for the delivered vocabulary and target recommended content.
[0135] Step 834, the vocabulary delivery module 830 establishes the correspondence between the delivered vocabulary and the target recommended content.
[0136] Step 835: The keyword delivery module 830 delivers the target recommended content based on the keywords to be delivered.
[0137] The keyword delivery module 830 is also used to display the recommendation bid in response to receiving a selection operation of a target recommended keyword from at least one recommended keyword from the target account.
[0138] In one example, please refer to Figure 11The diagram illustrates the corresponding modules of the aforementioned platform system 1100. The recommended vocabulary generation module includes a candidate word generation module 1110, a business filtering module 1120, a relevance module 1130, a structured validation module 1140, and a legal review module 1150. The candidate word generation module 1110 further includes an MT generation unit 1111, a Boolean retrieval generation unit 1112, a behavioral data mining unit 1113, and a structured generation unit 1114. The business filtering module 1120 includes negative word filtering 1121 and geographic targeting filtering 1122. The relevance module 1130 includes a traditional relevance model unit 1131 and an MT feature unit 1132. The structured validation module 1140 includes N structural branches 1141. The legal review module 1150 also includes AKA review 1151 and competitor filtering 1152. The platform system also includes a recommendation bidding module 1161, a matching module 1162, a landing page selection module 1163, and a data selection module 1164. The recommendation bidding module 1161 is used to bid on recommended keywords selected by the target account. The matching module 1162 is used by the target account to set the correspondence between the keywords and the target recommended content. The landing page selection module 1163 is used by the target account to set the page for the target recommended content. The data selection module 1164 is used to detect the traffic of each keyword for the target account and generate corresponding log reports based on the results, providing the target account with a reference for using keywords.
[0139] In summary, the vocabulary recommendation method provided in this application is applied to an advertising platform. Based on the target account's target vocabulary, it determines candidate words that are related to the target vocabulary from the vocabulary library, then determines recommended words based on the relevance of the candidate words to the target recommended content, and sends the recommended words to the target account. This realizes the mining of new words for use when delivering target recommended content based on existing target vocabulary, making the recommended vocabulary coverage wider without losing its relevance to the target recommended content, and improving the efficiency and accuracy of pushing recommended words to the target account.
[0140] Please refer to Figure 12 This is a structural block diagram of a vocabulary recommendation apparatus provided in an exemplary embodiment of this application. The apparatus includes:
[0141] The acquisition module 1210 is used to acquire the target account's keywords, which are the keywords selected by the target account when delivering target recommended content. The target recommended content is used as the search results for the keywords to recommend content.
[0142] The determining module 1220 is used to determine at least one candidate word that is related to the targeted words from the word library based on the targeted words;
[0143] The determining module 1220 is further configured to determine recommended words from the at least one candidate words based on the relevance between the at least one candidate word and the target recommended content;
[0144] The sending module 1230 is used to send the recommended words to the target account, wherein the recommended words are words selected for delivering the target recommended content to the target account.
[0145] In an optional embodiment, the determining module 1220 is further configured to determine the relevance data between the at least one candidate word and the target recommended content based on a preset verification method, wherein the preset verification method is configured to determine the degree of relevance between the candidate search results obtained when the at least one candidate word is used as a search request and the target recommended content.
[0146] The determining module 1220 is further configured to determine the recommended words from the at least one candidate words based on the relevance data.
[0147] In an optional embodiment, the relevance data includes semantic relevance;
[0148] The determining module 1220 is further configured to perform a relevance check on the at least one candidate word and determine the semantic relevance between the at least one candidate word and the delivered word.
[0149] In an optional embodiment, please refer to Figure 13 The at least one candidate word includes the target candidate word;
[0150] The determining module 1220 further includes an extraction unit 1221, used to extract the first feature of the target candidate word and the second feature of the delivered word;
[0151] The determining unit 1222 is used to input the first feature and the second feature into the relevance model to obtain the semantic relevance, wherein the relevance model is obtained by training on the corpus in the corpus.
[0152] In an optional embodiment, the relevance data includes structural relevance;
[0153] The determining module 1220 is further configured to perform structured verification on the at least one candidate word and determine the structural relevance between the at least one candidate word and the delivered word.
[0154] In an optional embodiment, the at least one candidate word includes a target candidate word;
[0155] The determining module 1220 further includes an output unit 1223, which is used to input the target candidate words into the structured model and output the first branch information corresponding to each structure branch. The structured model includes at least two structure branches.
[0156] The output unit 1223 is also used to input the target vocabulary into the structured model and output the second branch information corresponding to the structure branch;
[0157] The determining unit 1222 is further configured to compare the first branch information and the second branch information to determine the structural relevance of the target candidate words.
[0158] In an optional embodiment, the relevance data includes syntactic relevance;
[0159] The determining module 1220 is further configured to perform machine translation verification on the at least one candidate word and determine the grammatical relevance between the at least one candidate word and the delivered word.
[0160] In an optional embodiment, the at least one candidate word includes a target candidate word;
[0161] The output unit 1223 is also used to input the at least one candidate word into the machine translation verification model and output the recombined candidate word;
[0162] The determining unit 1222 is further configured to determine the grammatical relevance of the target candidate words based on the degree of difference between the target candidate words and the recombined candidate words.
[0163] In an optional embodiment, the determining module 1220 further includes a sorting unit 1224, configured to sort the candidate words according to the relevance data to obtain a candidate word sequence when the number of candidate words is at least two.
[0164] The determining unit 1222 is further configured to obtain a preset number of recommended words based on the candidate word sequence.
[0165] In an optional embodiment, the determining module 1220 is further configured to determine a second candidate word based on the grammatical features of the delivered word;
[0166] The determining module 1220 is further configured to determine a third candidate word based on the structural features of the delivered word, wherein the structural features are used to indicate the compositional structure of the word;
[0167] The determining module 1220 is also used to perform a Boolean search on the vocabulary database to obtain a fourth candidate word;
[0168] The determining module 1220 is further configured to determine the at least one candidate word based on the first candidate word, the second candidate word, the third candidate word, and the fourth candidate word.
[0169] In summary, the vocabulary recommendation device provided in this application embodiment determines candidate words that are related to the target words from the vocabulary library based on the target account's target vocabulary, then determines recommended words based on the relevance between the candidate words and the target recommended content, and sends the recommended words to the target account. This realizes the mining of new words for use when delivering target recommended content based on existing target vocabulary, making the recommended vocabulary coverage wider without losing relevance to the target recommended content.
[0170] It should be noted that the vocabulary recommendation device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the vocabulary recommendation device and the vocabulary recommendation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0171] Figure 14 A schematic diagram of the structure of a server provided in an exemplary embodiment of this application is shown. Specifically:
[0172] Server 1400 includes a Central Processing Unit (CPU) 1401, a system memory 1404 including Random Access Memory (RAM) 1402 and Read Only Memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the CPU 1401. Server 1400 also includes a mass storage device 1406 for storing an operating system 1413, application programs 1414, and other program modules 1415.
[0173] Mass storage device 1406 is connected to central processing unit 1401 via a mass storage controller (not shown) connected to system bus 1405. Mass storage device 1406 and its associated computer-readable media provide non-volatile storage for server 1400. That is, mass storage device 1406 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drives.
[0174] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1404 and mass storage device 1406 described above can be collectively referred to as memory.
[0175] According to various embodiments of this application, server 1400 can also be connected to a remote computer on a network, such as the Internet. That is, server 1400 can be connected to network 1412 via network interface unit 1411 connected to system bus 1405, or it can also use network interface unit 1411 to connect to other types of networks or remote computer systems (not shown).
[0176] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.
[0177] Embodiments of this application also provide a computer device including a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the vocabulary recommendation method provided in the above-described method embodiments. Optionally, the computer device may be a terminal or a server.
[0178] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the vocabulary recommendation method provided in the above-described method embodiments.
[0179] Embodiments of this application also provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the vocabulary recommendation methods described in the above embodiments.
[0180] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0181] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0182] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A vocabulary recommendation method, characterized in that, The method includes: Obtain the target account's keywords for ad placement. The keywords are the words used by the target account when ad placement of target recommended content. The target recommended content is used as the search results for the keywords for content recommendation. The first candidate word is determined based on the behavioral characteristics of the target word. The behavioral characteristics are used to indicate the features corresponding to the target word when it is searched. The behavioral characteristics include at least one of co-occurrence features, click features, and historical features. The co-occurrence features are used to indicate the probability of words appearing together in a certain context. The click features are used to indicate the target words corresponding to other links clicked when searching for the target word. The historical features are used to indicate words that the target account has selected in the word library but has not been targeted. The submitted words are machine translated to determine the second candidate words; The target vocabulary is divided into at least two branches according to its structure, and a third candidate vocabulary is generated based on the structural information corresponding to each branch. A Boolean search was performed on the vocabulary database to obtain the fourth candidate word; Based on the repetition rate among the first candidate words, the second candidate words, the third candidate words, and the fourth candidate words, at least one candidate word is determined; and / or, based on the quality levels corresponding to the first candidate words, the second candidate words, the third candidate words, and the fourth candidate words, and the quality level requirements specified by the target account, the at least one candidate word is selected from the first candidate words, the second candidate words, the third candidate words, and the fourth candidate words. Based on the relevance between the at least one candidate word and the target recommended content, recommended words are determined from the at least one candidate word; The recommended keywords are sent to the target account. The recommended keywords are words used when delivering the target recommended content to the target account.
2. The method according to claim 1, characterized in that, The step of determining recommended words from the at least one candidate word based on the relevance between the at least one candidate word and the target recommended content includes: Based on a preset verification method, the relevance data between the at least one candidate word and the target recommended content is determined. The preset verification method is used to determine the degree of relevance between the candidate search results obtained when the at least one candidate word is used as a search request and the target recommended content. Based on the relevance data, the recommended words are determined from the at least one candidate word.
3. The method according to claim 2, characterized in that, The relevance data includes semantic relevance; The determination of the relevance data between the at least one candidate word and the target recommended content based on a preset verification method includes: A relevance check is performed on the at least one candidate word to determine the semantic relevance between the at least one candidate word and the delivered word.
4. The method according to claim 3, characterized in that, The at least one candidate word includes the target candidate word; The step of performing a relevance check on the at least one candidate word to determine the semantic relevance between the at least one candidate word and the delivered word includes: Extract the first feature of the target candidate words and the second feature of the delivered words; The first feature and the second feature are input into the relevance model to obtain the semantic relevance, wherein the relevance model is obtained by training on the corpus.
5. The method according to claim 2, characterized in that, The correlation data includes structural correlation. The determination of the relevance data between the at least one candidate word and the target recommended content based on a preset verification method includes: Perform structured verification on the at least one candidate word to determine the structural relevance between the at least one candidate word and the target word.
6. The method according to claim 5, characterized in that, The at least one candidate word includes the target candidate word; The step of performing structured verification on the at least one candidate word to determine the structural relevance between the at least one candidate word and the target word includes: The target candidate words are input into the structured model, and the first branch information corresponding to each structure branch is output. The structured model includes at least two structure branches. The targeted vocabulary is input into the structured model, and the second branch information corresponding to the structure branch is output; By comparing the information from the first branch and the information from the second branch, the structural relevance of the target candidate words is determined.
7. The method according to claim 2, characterized in that, The relevance data includes syntactic relevance; The determination of the relevance data between the at least one candidate word and the target recommended content based on a preset verification method includes: Machine translation verification is performed on the at least one candidate word to determine the grammatical relevance between the at least one candidate word and the delivered word.
8. The method according to claim 7, characterized in that, The at least one candidate word includes the target candidate word; The step of performing machine translation verification on the at least one candidate word to determine the grammatical relevance between the at least one candidate word and the delivered word includes: The at least one candidate word is input into the machine translation verification model, and the recombined candidate words are output. The grammatical relevance of the target candidate words is determined based on the degree of difference between the target candidate words and the recombined candidate words.
9. The method according to any one of claims 2 to 8, characterized in that, The step of determining the recommended words from the at least one candidate words based on the relevance data includes: When the number of candidate words is at least two, the candidate words are sorted according to the relevance data to obtain a candidate word sequence; Based on the candidate word sequence, a preset number of recommended words are obtained.
10. A vocabulary recommendation device, characterized in that, The device includes: The acquisition module is used to acquire the target account's keywords for ad placement. The keywords are the words selected by the target account when ad placement of target recommended content. The target recommended content is used as the search results for the keywords for content recommendation. The determination module is used to determine a first candidate word based on the behavioral characteristics of the target word. The behavioral characteristics are used to indicate the features corresponding to the target word when it is searched. The behavioral characteristics include at least one of co-occurrence features, click features, and historical features. The co-occurrence features are used to indicate the probability of words appearing together in a certain context. The click features are used to indicate the target words corresponding to other links clicked when searching for the target word. The historical features are used to indicate words that the target account has selected in the word library but has not been targeted. The determining module is further configured to perform machine translation on the delivered words to determine a second candidate word; The determining module is further configured to divide the target words into at least two branches according to their structure, and generate a third candidate word based on the structural information corresponding to each branch; The determining module is also used to perform a Boolean search on the vocabulary database to obtain a fourth candidate word; The determining module is further configured to determine at least one candidate word based on the repetition rate among the first candidate word, the second candidate word, the third candidate word, and the fourth candidate word; and / or to filter the at least one candidate word from the first candidate word, the second candidate word, the third candidate word, and the fourth candidate word based on the quality levels corresponding to the first candidate word, the second candidate word, the third candidate word, and the fourth candidate word, and the quality level requirements specified by the target account. The determining module is further configured to determine recommended words from the at least one candidate words based on the relevance between the at least one candidate word and the target recommended content; The sending module is used to send the recommended words to the target account. The recommended words are words selected for delivering the target recommended content to the target account.
11. The apparatus according to claim 10, characterized in that, The determining module is further configured to determine the relevance data between the at least one candidate word and the target recommended content based on a preset verification method. The preset verification method is configured to determine the degree of relevance between the candidate search results obtained when the at least one candidate word is used as a search request and the target recommended content. The determining module is further configured to determine the recommended words from the at least one candidate words based on the relevance data.
12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by the processor to implement the vocabulary recommendation method as described in any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the vocabulary recommendation method as described in any one of claims 1 to 9.
14. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the vocabulary recommendation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Keyword recommending method and device
CN104636334A
Text similarity calculation method, device and electronic device
CN109472008A