Keyword recommendation method, apparatus, device, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING QIHOOD TECHNOLOGY CO LTD
- Filing Date
- 2021-03-05
- Publication Date
- 2026-06-09
AI Technical Summary
In vertical search scenarios, non-standard user queries or semantic confusion can lead to the inability to retrieve multimedia information, resulting in inefficient demand and revenue from multimedia information and low monetization of media traffic.
By extracting search keywords from user search instructions, a candidate set of multimedia information keywords is obtained. A set of user interest words is determined using a preset fine-grained identification strategy and recall strategy. Keyword recommendation and recall are then performed, and multimedia information recall is carried out using a multimedia information keyword rewriting service.
It improves the display costs for media outlets and the efficiency of multimedia information retrieval, and solves the retrieval problems caused by non-standard or semantically confusing search queries.
Smart Images

Figure CN115017344B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of search technology, and in particular to a keyword recommendation method, apparatus, device, and storage medium. Background Technology
[0002] Searching for multimedia information requires retrieving multimedia information based on user search keywords. However, in some vertical search scenarios, the lack of user search queries or unclear query expressions can lead to a situation where a request to the multimedia information engine fails to retrieve multimedia information. This results in inefficient multimedia information demand and revenue for multimedia information owners, media resource providers, and multimedia information audiences, and low commercialization of media traffic.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide a keyword recommendation method, apparatus, device, and storage medium, aiming to solve the technical problem that existing retrieval queries are not standardized or semantically confusing, resulting in the inability to recall multimedia information.
[0005] To achieve the above objectives, the present invention provides a keyword recommendation method, which includes the following steps:
[0006] Upon receiving a search instruction from a user, the system extracts search keywords from the search instruction.
[0007] When the search keywords do not meet the preset conditions, obtain the candidate set of multimedia information keywords corresponding to the user;
[0008] The set of user interest words is determined based on a preset fine-grained recognition strategy and the multimedia information keyword candidate set;
[0009] According to the preset recall strategy, the multimedia information keyword candidate set and the user interest word set are recalled respectively to obtain multiple keyword update candidate sets;
[0010] The candidate set of updated keywords is coarsely sorted, and keyword recommendations are made based on the sorting results.
[0011] Optionally, the step of obtaining the candidate set of multimedia information keywords corresponding to the user includes:
[0012] Obtain basic training data for keyword rewriting of multimedia information from users across the entire network;
[0013] Obtain the user's permanent IP address, and determine the location characteristics of the permanent IP address based on a preset IP location technology;
[0014] The location of residence is used as the geographical attribute corresponding to the IP permanent address;
[0015] Based on the geographic attributes and the basic training data, a candidate set of multimedia information keywords corresponding to the user is generated.
[0016] Optionally, the step of generating a candidate set of multimedia information keywords for the user based on the geographic attributes and the basic training data includes:
[0017] Generate an initial set of multimedia information keyword candidates for the user based on the geographic attributes and the basic training data;
[0018] Read the initial multimedia information keyword candidate set and aggregate the initial multimedia information keyword candidate set according to a preset time granularity to obtain an aggregated multimedia information keyword candidate set;
[0019] According to the preset horizontal segmentation rules, the aggregated multimedia information keyword candidate set is segmented into a horizontal multimedia information keyword candidate set according to the timestamp.
[0020] The column dimensions of the horizontal multimedia information keyword candidate set are numbered according to the preset vertical dimension columns to obtain the vertical multimedia information keyword candidate set;
[0021] The vertical multimedia information keyword candidate set is compressed according to a preset bitmap algorithm, and the compressed vertical multimedia information keyword candidate set is used as the multimedia information keyword candidate set corresponding to the user.
[0022] Optionally, the step of obtaining the basic training data for rewriting multimedia information keywords from all users on the network includes:
[0023] Obtain the user's historical search network data;
[0024] The historical retrieval network data is subjected to anti-fraud traffic processing to obtain anti-fraud retrieval data;
[0025] The historical search term set is determined based on the aforementioned anti-fraud search data;
[0026] The historical search term set is rewritten to obtain basic training data for rewriting multimedia information keywords from all users across the internet.
[0027] Optionally, the step of determining the historical search term set based on the anti-fraud retrieval data includes:
[0028] Feature extraction is performed on the anti-fraud retrieval data to obtain retrieval feature information;
[0029] A sentence representation vector is generated based on the retrieval feature information, and multi-scale context aggregation is performed on the sentence representation vector to obtain a set of historical search terms.
[0030] Optionally, before the step of obtaining the user's historical search network data, the method further includes:
[0031] Retrieve user search logs and user click logs for each user;
[0032] Based on a pre-defined data warehouse technology, the user's search logs and click logs are extracted, cleaned, transformed, loaded, and processed into a data warehouse to obtain the user's historical search network data.
[0033] Optionally, the step of determining the set of user interest words based on a preset fine-grained recognition strategy and the multimedia information keyword candidate set includes:
[0034] Based on the multimedia information keyword candidate set, determine user click behavior information and user context behavior information within a preset time period;
[0035] A user search session is constructed based on the user click behavior information and the user context behavior information;
[0036] Fine-grained feature identification is performed on the user search session to obtain fine-grained clustering features of several keywords;
[0037] The fine-grained clustering features are linearly weighted to obtain a set of user interest words.
[0038] Furthermore, to achieve the above objectives, the present invention also proposes a keyword recommendation device, the keyword recommendation device comprising:
[0039] The extraction module is used to extract search keywords from the search instructions input by the user when the user inputs a search instruction.
[0040] The acquisition module is used to acquire a candidate set of multimedia information keywords corresponding to the user when the search keywords do not meet the preset conditions;
[0041] The determination module is used to determine the set of user interest words based on a preset fine-grained recognition strategy and the candidate set of multimedia information keywords;
[0042] The recall module is used to recall the multimedia information keyword candidate set and the user interest word set according to a preset recall strategy, so as to obtain multiple keyword update candidate sets.
[0043] The recommendation module is used to perform a coarse sorting of the keyword update candidate set and recommend keywords based on the sorting results.
[0044] Furthermore, to achieve the above objectives, the present invention also proposes a keyword recommendation device, which includes a memory, a processor, and a keyword recommendation program stored in the memory and executable on the processor. The keyword recommendation program is configured to implement the steps of the keyword recommendation method described above.
[0045] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a keyword recommendation program, which, when executed by a processor, implements the steps of the keyword recommendation method as described above.
[0046] This invention extracts search keywords from a user's input search command; when the search keywords do not meet preset conditions, it obtains a candidate set of multimedia information keywords corresponding to the user; it determines a set of user interest words based on a preset fine-grained identification strategy and the candidate set of multimedia information keywords; it recalls the candidate set of multimedia information keywords and the set of user interest words according to a preset recall strategy to obtain multiple updated keyword candidate sets; it performs a coarse ranking of the updated keyword candidate sets and recommends keywords based on the ranking results. In this invention, when the user's input search terms do not meet preset conditions (i.e., are non-standard or semantically ambiguous and cannot be identified), multimedia information keywords are recommended by integrating the user's comprehensive search behavior, i.e., the candidate set of multimedia information keywords corresponding to the user. Simultaneously, using a multimedia information keyword rewriting service for multimedia information recall also increases the display cost for media outlets, thereby solving the technical problem of existing searches failing to recall multimedia information due to non-standard or semantically ambiguous queries. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of the structure of a keyword recommendation device for the hardware operating environment involved in the embodiments of the present invention;
[0048] Figure 2 This is a flowchart illustrating the first embodiment of the keyword recommendation method of the present invention;
[0049] Figure 3 This is a schematic diagram illustrating a recommendation result display in an embodiment of the present invention;
[0050] Figure 4 This is a flowchart illustrating the second embodiment of the keyword recommendation method of the present invention;
[0051] Figure 5 This is a flowchart illustrating the third embodiment of the keyword recommendation method of the present invention;
[0052] Figure 6This is a structural block diagram of the first embodiment of the keyword recommendation device of the present invention.
[0053] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0054] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0055] Reference Figure 1 , Figure 1 This is a schematic diagram of the device structure for recommending keywords related to the hardware operating environment involved in the embodiments of the present invention.
[0056] like Figure 1 As shown, the recommended device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, and optionally, it may also include a standard wired interface or a wireless interface. In this invention, the wired interface of the user interface 1003 may be a USB interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0057] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the keyword recommendation device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0058] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a keyword recommendation program.
[0059] exist Figure 1In the keyword recommendation device shown, the network interface 1004 is mainly used to connect to the backend server and communicate with the backend server; the user interface 1003 is mainly used to connect to the user device; the keyword recommendation device calls the keyword recommendation program stored in the memory 1005 through the processor 1001 and executes the keyword recommendation method provided in this embodiment of the invention.
[0060] Based on the above hardware structure, an embodiment of the keyword recommendation method of the present invention is proposed.
[0061] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the keyword recommendation method of the present invention, which presents the first embodiment of the keyword recommendation method of the present invention.
[0062] In the first embodiment, the keyword recommendation method includes the following steps:
[0063] Step S10: Upon receiving a search instruction input by the user, extract search keywords from the search instruction.
[0064] It should be noted that the execution entity in this embodiment is the keyword recommendation device. The keyword recommendation device can be a detection terminal on a computer, a detection terminal on a mobile phone, a detection terminal on an IoT device, etc., or other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment, a detection terminal on a computer is used as an example. Receiving the search instruction input by the user can be the user entering a statement that needs to be searched for security detection in the search box on the computer, or the user clicking in the search box on the computer but not entering search content. Receiving the search instruction input by the user can also be done by other means. This embodiment does not limit this.
[0065] It should be understood that when a user inputs a search instruction, search keywords are extracted from the search instruction. This embodiment does not limit the specific method of extracting search keywords from the search instruction. This embodiment describes one method: extracting search fields from the user input search instruction, extracting numbers, analyzing professional terms from the search fields to obtain professional terminology keywords; analyzing regular sentences from the search fields to obtain regular sentence keywords; and using the extracted numbers and / or professional terminology keywords and / or regular sentence keywords as search keywords.
[0066] Specifically, for example, if a user issues a search command with the search field "60 square meters or more near No. 25 Middle School, less than 1 million yuan", upon receiving the user's search command, the system extracts the search field "60 square meters or more near No. 25 Middle School, less than 1 million yuan" from the command; it then searches for keywords such as "unit price", "total price", "rent", "yuan", "ten thousand yuan", "area", and "square meters" to extract price-related figures, which in this example are 1 million yuan and 60 square meters; and it analyzes professional terms from the search field to obtain professional terminology keywords, which in this example are... The keyword for the place name is "No. 25 Middle School". Further analysis shows that the No. 25 Middle School is located near Jiangxi Road, Nanjing Road, and Minjiang Road in Shinan District. Analysis of common phrases in the search fields yields the keywords "above" and "below". After logically combining the extracted numerical and / or professional terminology keywords and / or common phrase keywords, the following search elements are obtained: Shinan District, Jiangxi Road, Nanjing Road, Minjiang Road, houses with an area of 60 square meters or more and a total price of less than 1 million yuan, regardless of floor, age, or whether it is a high-rise or multi-story building. These search elements can be used as search keywords.
[0067] Step S20: When the search keywords do not meet the preset conditions, obtain the candidate set of multimedia information keywords corresponding to the user.
[0068] It is easy to understand that receiving a user's input search instruction can be either the user entering a statement for a security check search in the search box on the computer, or the user clicking in the search box on the computer but not entering any search content. The system then determines whether the search keywords meet preset conditions. These preset conditions can include the standardization of the search keywords and clarity of meaning. If the search keywords do not meet the preset conditions, it is determined that the search keywords are not standard, are semantically ambiguous and cannot be identified, or the user has not entered any search terms. By integrating the user's behavior in the comprehensive search, i.e., the candidate set of multimedia information keywords corresponding to the user, multimedia information keywords are recommended.
[0069] It should be noted that the candidate set of multimedia information keywords corresponding to the user can be obtained in the following way: obtain the basic training data of rewriting multimedia information keywords of all users on the network; obtain the user's permanent IP address, and determine the permanent location characteristics of the permanent IP address according to the preset IP positioning technology; use the permanent location characteristics as the geographical attributes corresponding to the permanent IP address; generate the candidate set of multimedia information keywords corresponding to the user according to the geographical attributes and the basic training data.
[0070] Specifically, the process of obtaining basic training data for rewriting multimedia information keywords across the entire network can be as follows: Based on 360 Search user retrieval logs, user click logs, 360 Navigation retrieval logs, map search user retrieval and click logs, user multimedia information click logs, and commercial multimedia information query-CPM data, a series of ETL processes are used to model the above-mentioned multi-source data to form the basic training data for rewriting multimedia information keywords across the entire network. The ETL process includes extracting, transforming, and loading data from the source end to the destination end. Modeling methods can include traffic anti-fraud and context-aligned aggregation, etc., which are not limited in this embodiment.
[0071] Specifically, the system obtains the user's permanent IP address, and based on the distribution of the user's permanent IP address, it mines the permanent location characteristics of the IP address using high-precision IP positioning technology. As part of the multimedia information, the LBS attribute, i.e., the geographical attribute, is considered. Considering that the bid and click-through rate (CTR) of the multimedia information change dynamically over time, the candidate set of keywords for the multimedia information can be aggregated and segmented at multiple time granularities and assigned different weights.
[0072] Step S30: Determine the set of user interest words based on the preset fine-grained recognition strategy and the multimedia information keyword candidate set.
[0073] It should be understood that user click behavior information and user context behavior information within a preset time period are determined based on the candidate set of multimedia information keywords; a user search session is constructed based on the user click behavior information and the user context behavior information; fine-grained feature recognition is performed on the user search session to obtain fine-grained clustering features of several keywords; and the fine-grained clustering features are linearly weighted to obtain a set of user interest words.
[0074] Specifically, the preset time can be set according to actual needs. This embodiment takes 30 minutes as an example. The user's click behavior information and context behavior information within 30 minutes of a user's search can be concatenated into a user search session. Fine-grained feature identification is performed in each user search session. This fine-grained feature identification mainly calculates the word segmentation term, term-idf, term-weight, word frequency, query timeliness, and category of the query / title / cpc query, etc., to obtain fine-grained clustering features of several keyword queries / title / cpc queries. The fine-grained clustering features are linearly weighted to obtain the user interest term set, i.e., the user interest word set.
[0075] Step S40: Recall the multimedia information keyword candidate set and the user interest word set according to the preset recall strategy to obtain multiple keyword update candidate sets.
[0076] It should be noted that multiple recall configuration information sets and multiple inverted index sets are obtained according to the preset recall strategy; based on the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set and the user interest word set are searched and cyclically read respectively to obtain multiple keyword recall sets; the multiple keyword recall sets are used as multiple keyword update candidate sets.
[0077] Specifically, the preset recall strategy can include a variety of recall methods, which are not limited in this embodiment. This embodiment will describe them in the following ways: inverted index recall of multimedia information queries based on interest terms, inverted index recall of multimedia information queries based on semantically similar interest terms, LBS-based regional targeted recall based on the user's frequent location, generalized recall of semantically similar matching CPC queries, and recall of time-sensitive multimedia information queries, etc.
[0078] It is easy to understand that by recalling the multimedia information keyword candidate set and the user interest word set according to the preset recall strategy, multiple keyword update candidate sets can be obtained, which can expand the keyword candidate set and prevent the triggering of multimedia information word frequency control due to the limited keyword candidate set, thus recalling empty multimedia information.
[0079] Step S50: Perform a coarse sorting on the keyword update candidate set, and recommend keywords based on the sorting results.
[0080] It should be noted that in this embodiment, keyword recommendations can be made based on the coarse ranking results, or further fine-tuned from the coarse ranking results to obtain a refined ranking result, and then recommendations can be made based on the refined ranking result. This embodiment does not impose any limitations on this. Figure 3 , Figure 3 This is a schematic diagram showing one of the recommendation results in this embodiment.
[0081] It is easy to understand that the recommendation based on the coarse ranking result can be as follows: obtain the weight coefficient and timeliness coefficient corresponding to the keyword update candidate set; determine the coarse ranking score of the keyword update candidate set according to the weight coefficient and the timeliness coefficient through a preset coarse ranking strategy; normalize the coarse ranking score to the probability space according to a preset function to obtain the ranking result, and make keyword recommendations based on the ranking result.
[0082] Specifically, the N keyword update candidate sets (N≈50) obtained from multi-channel recall are initially sorted based on normalized CPM, weight parameters, decay coefficient, term frequency, etc., since the keyword update candidate set is the union of different recall strategy channels. Then, softmax is used to map each score to the probability space. Here, CPM refers to Cost Per Mille, which is a basic form of measuring the effect of multimedia information (whether it is traditional media or online media). CPM is the cost required to display multimedia information to one thousand people.
[0083] It is easy to understand that for a candidate set S = {S_{n}} containing N results, i Each candidate result receives a coarse-ranked score S in the overall ranking. i It is obtained from the following function:
[0084] i is a positive integer and i∈[1, N=50]
[0085] Where, α i It is the weight coefficient for each CPC query, α i = Term frequency * IDF * Click weight * Source weight; β i It is the timeliness coefficient for each CPC query, β i =0.98 △T △T represents the number of days since the last click of the CPC query during training. For the candidate set S, softmax is used to divide {S} into groups. i Normalized to the probability space, i.e., {S1} * S i * S N *};in,
[0086]
[0087] It should be understood that the method for further fine-tuning the coarse ranking results to obtain fine ranking results and making recommendations based on the fine ranking results can be as follows: coarsely ranking the updated candidate set of keywords to obtain coarse ranking results; determining the click-through rate of each keyword in the coarse ranking results based on the click-through rate model; determining the output value of each keyword in the coarse ranking results based on the keyword cost distribution bidding model; determining the display cost value of each keyword in the coarse ranking results based on the clicks and the output value using a preset display cost algorithm; generating fine ranking results based on the display cost values; and making keyword recommendations based on the fine ranking results.
[0088] Specifically, fine-tuning is used to further refine the coarse-ranked results, prioritizing the exposure of multimedia information keywords with high CPMs. Training samples are constructed based on historical click characteristics, material characteristics, and historical bid distribution of multimedia information keywords, combined with historical user-side characteristics. Negative examples from offline samples are randomly sampled, and overall data augmentation is performed to train the model and deploy the PCTR service. Due to the relatively random distribution of online traffic, offline calls do not perform prediction calibration on the PCTR. The fine-ranking part, in addition to estimating the click-through rate (CTR) of candidate CPC queries based on the CTR model, also integrates 360 Search's real-time query consumption prediction service, which estimates the bid based on the recent multimedia information consumption distribution of a given query. These two metrics, based on CPM = CTR * Bid * PVR, are used to finally rank the candidate set from the coarse-ranked results, yielding the fine-ranked results. Here, PVR represents the multimedia information recall rate for a given CPC query, and the top 10 results from the fine-ranked set are used for recommendation.
[0089] Furthermore, the CTR model used in fine sorting can be encapsulated into an independent online service that can be called. The purpose is to decouple the various modules of the system, make the logic clearer, and improve the iteration efficiency. The CTR model can be trained based on Google's open-source TensorFlow framework. This framework itself also provides a framework for deploying trained models online for online inference (the service request is the model's input feature, and the return is the CTR prediction value), namely tensorflow-serving.
[0090] It should be noted that this embodiment also includes a sensitive word filtering section, which is used to filter prohibited words from user search keywords and candidate multimedia information keywords. A blacklist of prohibited words is built on the server side. Each user-input search command and the offline calculations in steps S10 to S40 above are precisely matched based on this blacklist. If a match is found, it is directly filtered. For example, if the prohibited word is "**", then the entity word "A**B" will be filtered. The sensitive word filtering section is integrated throughout the construction of the word segmentation term and CPC query forward / inverted index. If the candidate set of multimedia information keywords and the inverted key of multimedia information match the sensitive word list, they are filtered out.
[0091] It is easy to understand that steps S10 to S40 in this embodiment are the offline part, which uses Redis for storage (the key is the user_id, and the value is the feature set and recommendation results). Two Redis instances are used online for automatic data updates, alternating between them. The online service determines which Redis instance to connect to based on its status parameters (the Redis instance with an online status and a larger update timestamp is used online). This dual-database mechanism also provides a certain degree of guarantee for service stability. Specifically, the entire data preprocessing and modeling process, index construction, and multi-path recall calculation, from user data acquisition to the final data, is processed offline. The results are in key-value pair format, and the key-value results are updated according to the offline update format after each training cycle, taking effect online.
[0092] This embodiment extracts search keywords from a user's search command upon receipt; when the search keywords do not meet preset conditions, it obtains a candidate set of multimedia information keywords corresponding to the user; it determines a set of user interest words based on a preset fine-grained identification strategy and the candidate set of multimedia information keywords; it recalls the candidate set of multimedia information keywords and the set of user interest words according to a preset recall strategy to obtain multiple updated keyword candidate sets; it performs a coarse ranking of the updated keyword candidate sets and recommends keywords based on the ranking results. In this embodiment, when the user's input search terms do not meet preset conditions (i.e., are non-standard or semantically ambiguous and cannot be identified), multimedia information keywords are recommended by fusing the user's comprehensive search behavior, i.e., the candidate set of multimedia information keywords corresponding to the user. Simultaneously, using a multimedia information keyword rewriting service for multimedia information recall increases the display cost for media outlets, thereby solving the technical problem of existing searches failing to recall multimedia information due to non-standard or semantically ambiguous queries.
[0093] Reference Figure 4 , Figure 4 This is a flowchart illustrating the second embodiment of the keyword recommendation method of the present invention. Based on the first embodiment described above, a second embodiment of the keyword recommendation method of the present invention is proposed.
[0094] In the second embodiment, step S20 includes:
[0095] Step S201: When the search keywords do not meet the preset conditions, obtain the basic training data for rewriting multimedia information keywords of all users on the network.
[0096] It is easy to understand that receiving a user's input search instruction can be either the user entering a statement for a security check search in the search box on the computer, or the user clicking in the search box on the computer but not entering any search content. The system then determines whether the search keywords meet preset conditions. These preset conditions can include the standardization of the search keywords and clarity of meaning. If the search keywords do not meet the preset conditions, it is determined that the search keywords are not standard, are semantically ambiguous and cannot be identified, or the user has not entered any search terms. By integrating the user's behavior in the comprehensive search, i.e., the candidate set of multimedia information keywords corresponding to the user, multimedia information keywords are recommended.
[0097] It should be understood that the process involves: acquiring the user's historical search network data; performing anti-fraud traffic processing on the historical search network data to obtain anti-fraud search data; determining a historical search term set based on the anti-fraud search data; and rewriting the historical search term set to obtain basic training data for rewriting multimedia information keywords for all users across the network. Specifically, feature extraction is performed on the anti-fraud search data to obtain search feature information; sentence representation vectors are generated based on the search feature information; and multi-scale context aggregation is performed on the sentence representation vectors to obtain the historical search term set.
[0098] It should be noted that, in order to improve the processability of the data, network data processing is required before obtaining the user's historical search network data: obtaining the user's search logs and user click logs; and extracting, cleaning, transforming, loading, and performing data warehouse processing on the user's search logs and user click logs based on a preset data warehouse technology to obtain the user's historical search network data.
[0099] Specifically, the process of obtaining basic training data for rewriting multimedia information keywords across the entire network can be as follows: Based on 360 Search user retrieval logs, user click logs, 360 Navigation retrieval logs, map search user retrieval and click logs, user multimedia information click logs, and commercial multimedia information query-CPM data, a series of ETL processes are used to model the above-mentioned multi-source data to form the basic training data for rewriting multimedia information keywords across the entire network. The ETL process includes extracting, transforming, and loading data from the source end to the destination end. Modeling methods can include traffic anti-fraud and context-aligned aggregation, etc., which are not limited in this embodiment.
[0100] Step S202: Obtain the user's permanent IP address and determine the location characteristics of the permanent IP address based on a preset IP location technology.
[0101] Specifically, the system obtains the user's permanent IP address, and based on the distribution of the user's permanent IP address, mines the location characteristics of the permanent IP address using high-precision IP positioning technology, and uses this as part of the multimedia information's LBS attribute, i.e., geographical attribute.
[0102] Step S203: Use the permanent residence feature as the geographical attribute corresponding to the IP permanent residence address.
[0103] It should be noted that by mining the permanent location characteristics of IP addresses using high-precision IP positioning technology, and considering that the bid and click-through rate (CTR) of multimedia information change dynamically over time, the candidate set of keywords for multimedia information can be aggregated and segmented at multiple time granularities and assigned different weights.
[0104] Step S204: Generate a candidate set of multimedia information keywords corresponding to the user based on the geographic attributes and the basic training data.
[0105] It should be noted that, based on the geographical attributes and the basic training data, an initial multimedia information keyword candidate set corresponding to the user is generated; the initial multimedia information keyword candidate set is read and aggregated according to a preset time granularity to obtain an aggregated multimedia information keyword candidate set; the aggregated multimedia information keyword candidate set is divided into horizontal multimedia information keyword candidate sets according to a preset horizontal segmentation rule based on timestamps; the column dimensions of the horizontal multimedia information keyword candidate sets are numbered according to a preset vertical dimension column to obtain a vertical multimedia information keyword candidate set; the vertical multimedia information keyword candidate set is compressed according to a preset bitmap algorithm, and the compressed vertical multimedia information keyword candidate set is used as the multimedia information keyword candidate set corresponding to the user.
[0106] This embodiment obtains basic training data for rewriting multimedia information keywords from users across the entire network; obtains users' permanent IP addresses, and determines the location characteristics of the permanent IP addresses based on preset IP positioning technology; uses the location characteristics as the geographical attributes corresponding to the permanent IP addresses; and generates a candidate set of multimedia information keywords for the user based on the geographical attributes and the basic training data. In this embodiment, when the search terms entered by the user do not meet preset conditions (i.e., are non-standard or semantically ambiguous and cannot be identified), multimedia information keywords are recommended by fusing user behavior from comprehensive search, i.e., the candidate set of multimedia information keywords for the user. Simultaneously, using the multimedia information keyword rewriting service for multimedia information retrieval also increases the display cost for media outlets, thereby solving the technical problem of existing retrieval queries being non-standard or semantically ambiguous, leading to the inability to retrieve multimedia information.
[0107] Reference Figure 5 , Figure 5 This is a flowchart illustrating a third embodiment of the keyword recommendation method of the present invention. Based on the first embodiment described above, a third embodiment of the keyword recommendation method of the present invention is proposed. This embodiment is described based on the first embodiment.
[0108] In the third embodiment, step S40 includes:
[0109] Step S401: Obtain multiple recall configuration information sets and multiple inverted index sets according to the preset recall strategy.
[0110] It should be noted that the preset recall strategy can include various recall methods, and this embodiment does not limit them. This embodiment describes the following methods: inverted index recall of multimedia information queries based on interest terms, inverted index recall of multimedia information queries based on semantically similar interest terms, LBS-based location-oriented recall based on the user's frequent location, generalized recall of semantically similar matching CPC queries, and recall of time-sensitive multimedia information queries, etc. The preset recall strategy is used to recall the multimedia information keyword candidate set and the user's interest word set respectively to obtain multiple updated keyword candidate sets. This can expand the keyword candidate set and prevent the triggering of multimedia information word frequency control due to a limited keyword candidate set, thus recalling empty multimedia information.
[0111] Specifically, the process of constructing an inverted index set may include: obtaining preset pay-per-click multimedia information keywords, and segmenting the preset pay-per-click multimedia information keywords to obtain segmented words; constructing a forward word list and an inverted word list based on the segmented words and the preset pay-per-click multimedia information keywords; and constructing multiple inverted index sets based on the forward word list and the inverted word list.
[0112] Step S402: Based on the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set and the user interest word set are retrieved and read in a loop to obtain multiple keyword recall sets.
[0113] It is easy to understand that one of the recall methods is as follows: First weight values and second weight values are obtained respectively for the multimedia information keyword candidate set and the user interest word set; the multimedia information keyword candidate set is searched and iteratively read according to the first weight value, the multiple recall configuration information sets, and the multiple inverted index sets to obtain multiple first keyword recall sets; the user interest word set is searched and iteratively read according to the second weight value, the multiple recall configuration information sets, and the multiple inverted index sets to obtain multiple second keyword recall sets; correspondingly, the step of using the multiple keyword recall sets as multiple keyword update candidate sets includes: generating multiple keyword update candidate sets according to the multiple first keyword recall sets and the multiple second keyword recall sets.
[0114] It should be understood that the semantic similarity matching CPC query generalization recall method can be as follows: determine the semantic matching model according to a preset recall strategy; generate keywords for the multimedia information keyword candidate set and the user interest word set according to the semantic matching model, respectively, to obtain multiple keyword update candidate sets. The semantic matching model can be the semantic matching model ESIM, and this embodiment does not limit it to this.
[0115] S403: Use the multiple keyword recall sets as multiple keyword update candidate sets.
[0116] It should be noted that, according to the preset recall strategy, the multimedia information keyword candidate set and the user interest word set are recalled respectively to obtain multiple keyword update candidate sets, which can expand the keyword candidate set and prevent the triggering of multimedia information word frequency control due to the limited keyword candidate set, thus recalling empty multimedia information.
[0117] In this embodiment, multiple recall configuration information sets and multiple inverted index sets are obtained according to a preset recall strategy. Based on these sets, the multimedia information keyword candidate set and the user interest word set are retrieved and iteratively read to obtain multiple keyword recall sets. These keyword recall sets are then used as multiple keyword update candidate sets. This embodiment expands the keyword candidate set by retrieving the multimedia information keyword candidate set and the user interest word set according to the preset recall strategy, preventing the triggering of multimedia information word frequency control and the recall of empty multimedia information due to a limited keyword candidate set. When the user-input search terms do not meet preset conditions (i.e., are non-standard or semantically ambiguous and cannot be identified), multimedia information keywords are recommended by integrating the user's comprehensive search behavior, i.e., the user's corresponding multimedia information keyword candidate set. Furthermore, using a multimedia information keyword rewriting service for multimedia information recall increases the media's display cost, thereby solving the technical problem of existing retrieval queries being non-standard or semantically ambiguous, leading to the inability to recall multimedia information.
[0118] In addition, refer to Figure 6 This invention also proposes a keyword recommendation device, which includes:
[0119] The extraction module 10 is used to extract search keywords from the search instruction input by the user when the user inputs a search instruction.
[0120] The acquisition module 20 is used to acquire a candidate set of multimedia information keywords corresponding to the user when the search keywords do not meet the preset conditions.
[0121] The determination module 30 is used to determine the set of user interest words based on a preset fine-grained recognition strategy and the multimedia information keyword candidate set.
[0122] The recall module 40 is used to recall the multimedia information keyword candidate set and the user interest word set according to a preset recall strategy, so as to obtain multiple keyword update candidate sets.
[0123] The recommendation module 50 is used to perform a coarse sorting of the keyword update candidate set and recommend keywords based on the sorting results.
[0124] This embodiment employs an extraction module 10 to extract search keywords from a user-input search command; an acquisition module 20 to acquire a candidate set of multimedia information keywords corresponding to the user when the search keywords do not meet preset conditions; a determination module 30 to determine a set of user interest words based on a preset fine-grained identification strategy and the candidate set of multimedia information keywords; a recall module 40 to recall the candidate set of multimedia information keywords and the set of user interest words respectively according to a preset recall strategy to obtain multiple keyword update candidate sets; and a recommendation module 50 to perform coarse sorting on the keyword update candidate sets and recommend keywords based on the sorting results. In this embodiment, when the user-input search terms do not meet preset conditions (i.e., are non-standard or semantically ambiguous and cannot be identified), multimedia information keywords are recommended by integrating the user's comprehensive search behavior, i.e., the candidate set of multimedia information keywords corresponding to the user. Simultaneously, using a multimedia information keyword rewriting service for multimedia information recall increases the display cost for media outlets, thereby solving the technical problem of existing searches failing to recall multimedia information due to non-standard or semantically ambiguous queries.
[0125] In one embodiment, the acquisition module 20 is further configured to acquire basic training data for rewriting multimedia information keywords of all users across the network;
[0126] Obtain the user's permanent IP address, and determine the location characteristics of the permanent IP address based on a preset IP location technology;
[0127] The location of residence is used as the geographical attribute corresponding to the IP permanent address;
[0128] Based on the geographic attributes and the basic training data, a candidate set of multimedia information keywords corresponding to the user is generated.
[0129] In one embodiment, the acquisition module 20 is further configured to generate an initial set of multimedia information keyword candidates for the user based on the geographic attributes and the basic training data;
[0130] Read the initial multimedia information keyword candidate set and aggregate the initial multimedia information keyword candidate set according to a preset time granularity to obtain an aggregated multimedia information keyword candidate set;
[0131] According to the preset horizontal segmentation rules, the aggregated multimedia information keyword candidate set is segmented into a horizontal multimedia information keyword candidate set according to the timestamp.
[0132] The column dimensions of the horizontal multimedia information keyword candidate set are numbered according to the preset vertical dimension columns to obtain the vertical multimedia information keyword candidate set;
[0133] The vertical multimedia information keyword candidate set is compressed according to a preset bitmap algorithm, and the compressed vertical multimedia information keyword candidate set is used as the multimedia information keyword candidate set corresponding to the user.
[0134] In one embodiment, the acquisition module 20 is further configured to acquire the user's historical search network data;
[0135] The historical retrieval network data is subjected to anti-fraud traffic processing to obtain anti-fraud retrieval data;
[0136] The historical search term set is determined based on the aforementioned anti-fraud search data;
[0137] The historical search term set is rewritten to obtain basic training data for rewriting multimedia information keywords from all users across the internet.
[0138] In one embodiment, the acquisition module 20 is further configured to extract features from the anti-fraud retrieval data to obtain retrieval feature information;
[0139] A sentence representation vector is generated based on the retrieval feature information, and multi-scale context aggregation is performed on the sentence representation vector to obtain a set of historical search terms.
[0140] In one embodiment, the acquisition module 20 is further configured to acquire user search logs and user click logs corresponding to the user;
[0141] Based on a pre-defined data warehouse technology, the user's search logs and click logs are extracted, cleaned, transformed, loaded, and processed into a data warehouse to obtain the user's historical search network data.
[0142] In one embodiment, the determining module 30 is further configured to determine user click behavior information and user context behavior information within a preset time period based on the multimedia information keyword candidate set;
[0143] A user search session is constructed based on the user click behavior information and the user context behavior information;
[0144] Fine-grained feature identification is performed on the user search session to obtain fine-grained clustering features of several keywords;
[0145] The fine-grained clustering features are linearly weighted to obtain a set of user interest words.
[0146] In one embodiment, the recall module 40 is further configured to obtain multiple recall configuration information sets and multiple inverted index sets according to a preset recall strategy;
[0147] Based on the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set and the user interest word set are retrieved and cyclically read to obtain multiple keyword recall sets;
[0148] The multiple keyword recall sets are used as multiple keyword update candidate sets.
[0149] In one embodiment, the recall module 40 is further configured to obtain a first weight value and a second weight value corresponding to the multimedia information keyword candidate set and the user interest word set, respectively;
[0150] Based on the first weight value, the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set is retrieved and read cyclically to obtain multiple first keyword recall sets;
[0151] Based on the second weight value, the multiple recall configuration information sets, and the multiple inverted index sets, the user interest word set is retrieved and read cyclically to obtain multiple second keyword recall sets;
[0152] Accordingly, the step of using the multiple keyword recall sets as multiple keyword update candidate sets includes:
[0153] Multiple keyword update candidate sets are generated based on the multiple first keyword recall sets and the multiple second keyword recall sets.
[0154] In one embodiment, the recall module 40 is further configured to obtain preset click-based pay-multimedia information keywords and perform word segmentation on the preset click-based pay-multimedia information keywords to obtain segmented words;
[0155] Construct a forward word list and an inverted word list based on the segmented words and the preset click-to-pay multimedia information keywords;
[0156] Multiple inverted index sets are constructed based on the forward and inverted word lists.
[0157] In one embodiment, the recall module 40 is further configured to determine a semantic matching model according to a preset recall strategy;
[0158] Based on the semantic matching model, keywords are generated for the multimedia information keyword candidate set and the user interest word set, respectively, to obtain multiple keyword update candidate sets.
[0159] In one embodiment, the recommendation module 50 is further configured to obtain the weight coefficient and timeliness coefficient corresponding to the keyword update candidate set;
[0160] The coarse ranking score of the keyword update candidate set is determined based on the weight coefficient and the timeliness coefficient using a preset coarse ranking strategy.
[0161] The coarse ranking score is normalized to the probability space according to a preset function to obtain the ranking result, and keyword recommendations are made based on the ranking result.
[0162] In one embodiment, the recommendation module 50 is further configured to perform a coarse sorting on the keyword update candidate set to obtain a coarse sorting result;
[0163] The click-through rate of each item in the coarse ranking results is determined based on the click pass rate model;
[0164] The bid value of each item in the coarse ranking result is determined based on the keyword consumption distribution bidding model;
[0165] The display cost values of the coarse ranking results are determined based on the clicks and the output value using a preset display cost algorithm.
[0166] A refined ranking result is generated based on the displayed cost value, and keyword recommendations are made based on the refined ranking result.
[0167] Other embodiments or specific implementations of the keyword recommendation device described in this invention can be found in the above-described embodiments of the keyword recommendation method, and will not be repeated here.
[0168] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0169] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0170] In addition, for technical details not described in detail in this embodiment, please refer to the keyword recommendation method provided in any embodiment of the present invention, which will not be repeated here.
[0171] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0172] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, and third, etc., does not indicate any order and can be interpreted as identifiers.
[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory image (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0174] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A keyword recommendation method, characterized in that, The keyword recommendation method includes the following steps: Upon receiving a search instruction from a user, the system extracts search keywords from the search instruction. When the search keywords do not meet the preset conditions, obtain the candidate set of multimedia information keywords corresponding to the user; The set of user interest words is determined based on a preset fine-grained recognition strategy and the multimedia information keyword candidate set; According to the preset recall strategy, the multimedia information keyword candidate set and the user interest word set are recalled respectively to obtain multiple keyword update candidate sets; The keyword update candidate set is coarsely sorted, and keyword recommendations are made based on the sorting results; The step of coarsely sorting the keyword update candidate set and recommending keywords based on the sorting results includes: Obtain the weight coefficients and timeliness coefficients corresponding to the keyword update candidate set; The coarse ranking score of the keyword update candidate set is determined based on the weight coefficient and the timeliness coefficient using a preset coarse ranking strategy. The coarse ranking score is normalized to the probability space according to a preset function to obtain the ranking result, and keyword recommendations are made based on the ranking result.
2. The keyword recommendation method as described in claim 1, characterized in that, The step of obtaining the candidate set of multimedia information keywords corresponding to the user includes: Obtain basic training data for keyword rewriting of multimedia information from users across the entire network; Obtain the user's permanent IP address, and determine the location characteristics of the permanent IP address based on a preset IP location technology; The location of residence is used as the geographical attribute corresponding to the IP permanent address; Based on the geographic attributes and the basic training data, a candidate set of multimedia information keywords corresponding to the user is generated.
3. The keyword recommendation method as described in claim 2, characterized in that, The step of generating a candidate set of multimedia information keywords for the user based on the geographic attributes and the basic training data includes: Generate an initial set of multimedia information keyword candidates for the user based on the geographic attributes and the basic training data; Read the initial multimedia information keyword candidate set and aggregate the initial multimedia information keyword candidate set according to a preset time granularity to obtain an aggregated multimedia information keyword candidate set; According to the preset horizontal segmentation rules, the aggregated multimedia information keyword candidate set is segmented into a horizontal multimedia information keyword candidate set according to the timestamp. The column dimensions of the horizontal multimedia information keyword candidate set are numbered according to the preset vertical dimension columns to obtain the vertical multimedia information keyword candidate set; The vertical multimedia information keyword candidate set is compressed according to a preset bitmap algorithm, and the compressed vertical multimedia information keyword candidate set is used as the multimedia information keyword candidate set corresponding to the user.
4. The keyword recommendation method as described in claim 2, characterized in that, The steps for obtaining the basic training data for rewriting multimedia information keywords from all users on the network include: Obtain the user's historical search network data; The historical retrieval network data is subjected to anti-fraud traffic processing to obtain anti-fraud retrieval data; The historical search term set is determined based on the aforementioned anti-fraud search data; The historical search term set is rewritten to obtain basic training data for rewriting multimedia information keywords from all users across the internet.
5. The keyword recommendation method as described in claim 4, characterized in that, The step of determining the set of historical search terms based on the anti-fraud retrieval data includes: Feature extraction is performed on the anti-fraud retrieval data to obtain retrieval feature information; A sentence representation vector is generated based on the retrieval feature information, and multi-scale context aggregation is performed on the sentence representation vector to obtain a set of historical search terms.
6. The keyword recommendation method as described in claim 4, prior to the step of obtaining the user's historical search network data, further includes: Retrieve user search logs and user click logs for each user; Based on a pre-defined data warehouse technology, the user's search logs and click logs are extracted, cleaned, transformed, loaded, and processed into a data warehouse to obtain the user's historical search network data.
7. The keyword recommendation method as described in claim 1, characterized in that, The step of determining the set of user interest words based on a preset fine-grained recognition strategy and the multimedia information keyword candidate set includes: Based on the multimedia information keyword candidate set, determine user click behavior information and user context behavior information within a preset time period; A user search session is constructed based on the user click behavior information and the user context behavior information; Fine-grained feature identification is performed on the user search session to obtain fine-grained clustering features of several keywords; The fine-grained clustering features are linearly weighted to obtain a set of user interest words.
8. The keyword recommendation method as described in claim 1, characterized in that, The step of recalling the multimedia information keyword candidate set and the user interest word set according to the preset recall strategy to obtain multiple keyword update candidate sets includes: Based on the preset recall strategy, obtain multiple recall configuration information sets and multiple inverted index sets; Based on the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set and the user interest word set are retrieved and cyclically read to obtain multiple keyword recall sets; The multiple keyword recall sets are used as multiple keyword update candidate sets.
9. The keyword recommendation method as described in claim 8, characterized in that, The step of retrieving and iteratively reading the multimedia information keyword candidate set and the user interest word set according to the multiple recall configuration information sets and the multiple inverted index sets to obtain multiple keyword recall sets includes: Obtain the first weight value and the second weight value corresponding to the multimedia information keyword candidate set and the user interest word set, respectively; Based on the first weight value, the multiple recall configuration information sets and the multiple inverted index sets, the multimedia information keyword candidate set is retrieved and read cyclically to obtain multiple first keyword recall sets; Based on the second weight value, the multiple recall configuration information sets, and the multiple inverted index sets, the user interest word set is retrieved and read cyclically to obtain multiple second keyword recall sets; Accordingly, the step of using the multiple keyword recall sets as multiple keyword update candidate sets includes: Multiple keyword update candidate sets are generated based on the multiple first keyword recall sets and the multiple second keyword recall sets.
10. The keyword recommendation method as described in claim 8, characterized in that, Before the step of obtaining multiple recall configuration information sets and multiple inverted index sets according to a preset recall strategy, the method further includes: Obtain preset keywords for pay-per-click multimedia information, and segment the preset keywords for pay-per-click multimedia information to obtain segmented words; Construct a forward word list and an inverted word list based on the segmented words and the preset click-to-pay multimedia information keywords; Multiple inverted index sets are constructed based on the forward and inverted word lists.
11. The keyword recommendation method as described in claim 1, characterized in that, The step of recalling the multimedia information keyword candidate set and the user interest word set according to the preset recall strategy to obtain multiple keyword update candidate sets includes: The semantic matching model is determined based on the preset recall strategy; Based on the semantic matching model, keywords are generated for the multimedia information keyword candidate set and the user interest word set, respectively, to obtain multiple keyword update candidate sets.
12. The keyword recommendation method as described in claim 1, characterized in that, The step of coarsely ranking the keyword update candidate set and recommending keywords based on the ranking results includes: The keyword update candidate set is coarsely sorted to obtain the coarse sorting result; The click-through rate of each item in the coarse ranking results is determined based on the click pass rate model; The bid value of each item in the coarse ranking result is determined based on the keyword consumption distribution bidding model; The display cost values of the coarse ranking results are determined based on the clicks and the output value using a preset display cost algorithm. A refined ranking result is generated based on the displayed cost value, and keyword recommendations are made based on the refined ranking result.
13. A keyword recommendation device, characterized in that, The keyword recommendation device includes: The extraction module is used to extract search keywords from the search instructions input by the user when the user inputs a search instruction. The acquisition module is used to acquire a candidate set of multimedia information keywords corresponding to the user when the search keywords do not meet the preset conditions; The determination module is used to determine the set of user interest words based on a preset fine-grained recognition strategy and the candidate set of multimedia information keywords; The recall module is used to recall the multimedia information keyword candidate set and the user interest word set according to a preset recall strategy, so as to obtain multiple keyword update candidate sets. The recommendation module is used to perform a coarse sorting of the keyword update candidate set and recommend keywords based on the sorting results; The recommendation module is further configured to obtain the weight coefficient and timeliness coefficient corresponding to the keyword update candidate set; determine the coarse ranking score of the keyword update candidate set according to the weight coefficient and the timeliness coefficient through a preset coarse ranking strategy; normalize the coarse ranking score to the probability space according to a preset function to obtain the ranking result, and recommend keywords according to the ranking result.
14. The keyword recommendation device as described in claim 13, characterized in that, The acquisition module is also used to acquire basic training data for rewriting multimedia information keywords of all users on the network; Obtain the user's permanent IP address, and determine the location characteristics of the permanent IP address based on a preset IP location technology; The location of residence is used as the geographical attribute corresponding to the IP permanent address; Based on the geographic attributes and the basic training data, a candidate set of multimedia information keywords corresponding to the user is generated.
15. The keyword recommendation device as described in claim 14, characterized in that, The acquisition module is further configured to generate an initial set of multimedia information keyword candidates for the user based on the geographic attributes and the basic training data; Read the initial multimedia information keyword candidate set and aggregate the initial multimedia information keyword candidate set according to a preset time granularity to obtain an aggregated multimedia information keyword candidate set; According to the preset horizontal segmentation rules, the aggregated multimedia information keyword candidate set is segmented into a horizontal multimedia information keyword candidate set according to the timestamp. The column dimensions of the horizontal multimedia information keyword candidate set are numbered according to the preset vertical dimension columns to obtain the vertical multimedia information keyword candidate set; The vertical multimedia information keyword candidate set is compressed according to a preset bitmap algorithm, and the compressed vertical multimedia information keyword candidate set is used as the multimedia information keyword candidate set corresponding to the user.
16. The keyword recommendation device as described in claim 14, characterized in that, The acquisition module is also used to acquire the user's historical search network data; The historical retrieval network data is subjected to anti-fraud traffic processing to obtain anti-fraud retrieval data; The historical search term set is determined based on the aforementioned anti-fraud search data; The historical search term set is rewritten to obtain basic training data for rewriting multimedia information keywords from all users across the internet.
17. The keyword recommendation device as described in claim 16, characterized in that, The acquisition module is also used to extract features from the anti-fraud retrieval data to obtain retrieval feature information; A sentence representation vector is generated based on the retrieval feature information, and multi-scale context aggregation is performed on the sentence representation vector to obtain a set of historical search terms.
18. A keyword recommendation device, characterized in that, The keyword recommendation device includes: a memory, a processor, and a keyword recommendation program stored in the memory and executable on the processor, wherein the keyword recommendation program is configured to implement the keyword recommendation method according to any one of claims 1 to 12.
19. A storage medium, characterized in that, The storage medium stores a keyword recommendation program, which, when executed by a processor, implements the steps of the keyword recommendation method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Recall method and device based on user portrait label, equipment and storage medium
CN112231555A
Search data identification method and device, electronic equipment and computer storage medium
CN112307183A