Content Recommendation Method, Apparatus, Server, and Storage Medium
A content recommendation method using content clustering to generate class labels and exclude similar content based on user preferences improves recommendation accuracy by overcoming manual categorization inefficiencies.
Patent Information
- Application Number
- CN202111016634.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In the prior art, the content recommendation system still recommends the same type of content after the account gives negative feedback, resulting in a low recommendation accuracy.
By obtaining the feature information of the account to be recommended, clustering analysis is performed, the correspondence between the content and category labels is generated, and the content associated with the content label whose similarity to the account's feature is less than the preset similarity, and the target content is recommended.
It improves the accuracy of content recommendation, avoids the limitations of artificial classification system, and increases the ability to generalize similar content.
Smart Images

Figure CN113868547B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a content recommendation method, apparatus, server, and storage medium. Background Art
[0002] With the development of Internet technologies, there is a rich variety of content on the Internet, such as pictures, videos, advertisements, articles, etc. A recommendation system will automatically recommend content to an account.
[0003] In related technologies, when making content recommendations, when an account gives negative feedback on a certain piece of content, the recommendation system may recommend content belonging to another category to the account based on manually divided content categories. However, in the case of a large amount of content, the efficiency of manually dividing content categories is low, resulting in the recommendation system still recommending the same type of content to the account when the account gives negative feedback on a certain piece of content, thereby leading to a low accuracy rate of content recommendation. Summary of the Invention
[0004] The present disclosure provides a content recommendation method, apparatus, server, and storage medium to at least solve the problem of low accuracy rate of content recommendation in related technologies. The technical solution of the present disclosure is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a content recommendation method, including:
[0006] Obtaining to-be-recommended content corresponding to a to-be-recommended account;
[0007] Querying a pre-generated correspondence between content and category labels to obtain a category label corresponding to the to-be-recommended content; the correspondence is obtained by clustering feature information of the content;
[0008] Obtaining a sample label of sample content corresponding to the to-be-recommended account, and deleting from the to-be-recommended content the content associated with the category label and any one of the sample labels to obtain target content; the sample content is used to represent content whose feature similarity with the account features of the to-be-recommended account is less than a preset similarity, and the sample label is obtained through the correspondence;
[0009] Recommending the target content to the to-be-recommended account.
[0010] In an exemplary embodiment, before obtaining the to-be-recommended content corresponding to the to-be-recommended account, it further includes:
[0011] Obtaining feature information of content;
[0012] Cluster the content according to the feature information of the content to obtain a content set; the content set is matched with a category label, and the category label of the content in the content set is the same as the category label matched by the content set.
[0013] Determine the correspondence between the content and the category label according to the category label of the content in the content set, as the pre-generated correspondence between the content and the category label.
[0014] In an exemplary embodiment, the clustering the content according to the feature information of the content to obtain a content set includes:
[0015] Input the feature information of the content into a preset clustering model, and through the preset clustering model, combine multiple contents with a similarity greater than a preset threshold between the feature information to obtain the content set.
[0016] In an exemplary embodiment, the obtaining the feature information of the content includes:
[0017] Obtain the text information of the content;
[0018] Preprocess the text information of the content to obtain preprocessed text information;
[0019] Extract the feature information of the preprocessed text information through a feature extraction model as the feature information of the content.
[0020] In an exemplary embodiment, the obtaining the text information of the content includes:
[0021] Obtain the content identifier and description information of the content;
[0022] Combine the content identifier and description information of the content to obtain the text information of the content.
[0023] In an exemplary embodiment, the preprocessing the text information of the content to obtain preprocessed text information includes:
[0024] Delete the invalid information in the content identifier and the invalid information in the description information to obtain a target content identifier and a target description information;
[0025] Perform word segmentation on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information;
[0026] Connect the word segmentation of the target content identifier and the word segmentation of the target description information through a delimiter to obtain the preprocessed text information.
[0027] In an exemplary embodiment, querying the correspondence between the pre-generated content and the category labels to obtain the category labels corresponding to the content to be recommended includes:
[0028] Obtaining an identifier of the content to be recommended for the content to be recommended;
[0029] Querying the correspondence between the pre-generated content and the category labels to obtain the category labels corresponding to the content with the same content identifier as the content to be recommended, as the category labels of the content to be recommended.
[0030] According to a second aspect of the embodiments of the present disclosure, there is provided a content recommendation device, including:
[0031] An obtaining unit, configured to obtain the content to be recommended corresponding to the account to be recommended;
[0032] A querying unit, configured to query the correspondence between the pre-generated content and the category labels to obtain the category labels corresponding to the content to be recommended; the correspondence is obtained by clustering the feature information of the content;
[0033] A deleting unit, configured to obtain the sample labels of the sample content corresponding to the account to be recommended, and delete the content associated with the category labels and any of the sample labels from the content to be recommended to obtain the target content; the sample content is used to represent the content whose feature similarity with the account features of the account to be recommended is less than a preset similarity, and the sample labels are obtained through the correspondence;
[0034] A recommending unit, configured to recommend the target content to the account to be recommended.
[0035] In an exemplary embodiment, the content recommendation device further includes a determining unit, configured to obtain the feature information of the content; cluster the content according to the feature information of the content to obtain a content set; the content set is matched with category labels, and the category labels of the content in the content set are the same as the category labels matched by the content set; determine the correspondence between the content and the category labels according to the category labels of the content in the content set, as the correspondence between the pre-generated content and the category labels.
[0036] In an exemplary embodiment, the determining unit is further configured to input the feature information of the content into a preset clustering model, and combine, through the preset clustering model, multiple contents with a similarity greater than a preset threshold between the feature information to obtain the content set.
[0037] In an exemplary embodiment, the determining unit is further configured to execute obtaining text information of the content; preprocessing the text information of the content to obtain preprocessed text information; and extracting feature information of the preprocessed text information through a feature extraction model as the feature information of the content.
[0038] In an exemplary embodiment, the determining unit is further configured to execute obtaining a content identifier and description information of the content; and combining the content identifier and description information of the content to obtain the text information of the content.
[0039] In an exemplary embodiment, the determining unit is further configured to execute deleting invalid information in the content identifier and invalid information in the description information to obtain a target content identifier and target description information; performing word segmentation processing on the target content identifier and the target description information respectively to obtain word segmentation of the target content identifier and word segmentation of the target description information; and connecting the word segmentation of the target content identifier and the word segmentation of the target description information through a separator to obtain the preprocessed text information.
[0040] In an exemplary embodiment, the querying unit is further configured to execute obtaining a to-be-recommended content identifier of the to-be-recommended content; and querying a pre-generated correspondence between content and category tags to obtain a category tag corresponding to the content with the same content identifier as the to-be-recommended content as the category tag of the to-be-recommended content.
[0041] According to a third aspect of the embodiments of the present disclosure, there is provided a server, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the content recommendation method as described in any one of the embodiments of the first aspect.
[0042] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, including: when instructions in the storage medium are executed by a processor of a server, enabling the server to execute the content recommendation method as described in any one of the embodiments of the first aspect.
[0043] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of the device reads and executes the computer program from the readable storage medium, enabling the device to execute the content recommendation method as described in any one of the embodiments of the first aspect.
[0044] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0045] By obtaining the content to be recommended corresponding to the account to be recommended; then querying the corresponding relationship between the pre-generated content and the category labels to obtain the category labels corresponding to the content to be recommended; the corresponding relationship is obtained by clustering the feature information of the content; then obtaining the sample labels of the sample content corresponding to the account to be recommended, and deleting the content associated with the category labels and any sample labels from the content to be recommended to obtain the target content; the sample content is used to characterize the content whose feature similarity with the account features of the account to be recommended is less than the preset similarity, and the sample labels are obtained through the corresponding relationship; finally, recommending the target content to the account to be recommended; in this way, clustering the content using the feature information of the content to obtain the corresponding relationship between the content and the category labels is beneficial to avoiding the limitations of the manually divided classification system, thereby increasing the generalization ability of similar content, and further improving the determination accuracy of the category labels, so that when recommending content to the account to be recommended, it is possible to accurately not recommend the content associated with the sample labels of the category labels and the sample content with negative feedback from the account to be recommended, thereby improving the recommendation accuracy of the content.
[0046] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0048] Figure 1 is an application environment diagram of a content recommendation method shown according to an exemplary embodiment.
[0049] Figure 2 is a flowchart of a content recommendation method shown according to an exemplary embodiment.
[0050] Figure 3 is a flowchart of the construction steps of the corresponding relationship between content and category labels shown according to an exemplary embodiment.
[0051] Figure 4 is a flowchart of another content recommendation method shown according to an exemplary embodiment.
[0052] Figure 5 is a flowchart of an advertisement recommendation method shown according to an exemplary embodiment.
[0053] Figure 6 is a block diagram of a content recommendation device shown according to an exemplary embodiment.
[0054] Figure 7 is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners
[0055] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0056] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0057] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0058] The content recommendation method provided by the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 110 interacts with the server 120 through the network. Specifically, referring to Figure 1 , the server 120 obtains the content to be recommended corresponding to the account to be recommended logged in the terminal 110; queries the pre-generated corresponding relationship between the content and the category label to obtain the category label corresponding to the content to be recommended; the corresponding relationship is obtained by clustering the feature information of the content; obtains the sample label of the sample content corresponding to the account to be recommended, and deletes the content associated with the category label and any sample label from the content to be recommended to obtain the target content; the sample content is used to represent the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, and the sample label is obtained through the corresponding relationship; recommends the target content to the terminal 110 corresponding to the account to be recommended, and the terminal 110 displays the target content pushed by the server 120 through the terminal interface. Among them, the terminal 110 may be, but is not limited to, various personal computers, smart phones, tablet computers, notebook computers, or smart wearable devices, etc., and the server 120 may be implemented by an independent server or a server cluster composed of multiple servers.
[0059] Figure 2 is a flowchart of a content recommendation method shown according to an exemplary embodiment. As Figure 2 shown, the content recommendation method is used for the server shown in Figure 1 and includes the following steps:
[0060] In step S210, obtain the content to be recommended corresponding to the account to be recommended.
[0061] Among them, an account refers to a registered account of an application program in a terminal, such as a registered account of a short video application program, a registered account of a video browsing program, etc. The account to be recommended refers to an authorized account that needs to be processed and analyzed, specifically referring to the object of content recommendation; in an actual scenario, the account to be recommended can refer to the object of advertisement recommendation, the object of news and information recommendation, or the object of promotional video recommendation, etc.
[0062] Among them, the content to be recommended refers to the content that needs to be recommended to the account to be recommended, which can be a video, a live broadcast, a picture, an advertisement, an article, news and information, etc. For example, the content to be recommended is the advertisement "CCC - A popular and exciting mini-game, interesting and fun".
[0063] Specifically, the server obtains authorized accounts on the network as the accounts to be recommended; from the content database storing content, obtains the content matching the account to be recommended, and uses the content matching the account to be recommended as the content to be recommended corresponding to the account to be recommended.
[0064] For example, the server obtains the content to be recommended corresponding to the account to be recommended from the content database storing content according to the collaborative filtering recommendation algorithm.
[0065] In step S220, query the corresponding relationship between the pre-generated content and the category label, and obtain the category label corresponding to the content to be recommended; the corresponding relationship is obtained by clustering the feature information of the content.
[0066] Among them, the pre-generated corresponding relationship between the content and the category label is used to represent that there is a one-to-one corresponding relationship between the content and the category label; it should be noted that different from the category labels divided manually, the category labels involved in the present disclosure are obtained by clustering the feature information of the content, and they are only used as identifiers and do not necessarily have real semantics; each category label can capture deeper similarities and has a better effect when recommending content according to the account preferences.
[0067] Among them, the feature information of the content is used to represent the key features extracted from the text information of the content (such as the name, description information, etc.).
[0068] Specifically, the server obtains the pre-generated corresponding relationship between the content and the category label from the local database, and queries the pre-generated corresponding relationship between the content and the category label according to the content to be recommended, and obtains the category label corresponding to the content identical to the content to be recommended as the category label corresponding to the content to be recommended.
[0069] For example, assume that in the correspondence between pre-generated content and category labels, the category label corresponding to content a is A, and the content to be recommended is the same as content a, indicating that the category label of the content to be recommended is also A.
[0070] Further, before querying the correspondence between pre-generated content and category labels to obtain the category label corresponding to the content to be recommended, the server first obtains the feature information of the content in the content database, and performs clustering processing on the content in the content database according to the feature information of the content to obtain multiple content sets, and each content set is matched with a category label; according to the category label matched by each content set, the correspondence between the content and the category label is determined as the correspondence between the pre-generated content and the category label.
[0071] It should be noted that in the process of obtaining the feature information of the content in the content database, the server can obtain all the content in the content database or part of the content in the content database, which can be specifically determined according to the actual situation; for example, when the number of content in the content database is relatively small, all the content in the content database can be obtained; when the number of content in the content database is relatively large, popular content in the content database can be obtained.
[0072] In step S230, obtain the sample label of the sample content corresponding to the account to be recommended, and delete the content associated with the category label and any sample label from the content to be recommended to obtain the target content; the sample content is used to represent the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, and the sample label is obtained through the correspondence.
[0073] Among them, the account feature refers to the attribute feature of the account, such as gender, region, interest, etc.; the sample content refers to the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, specifically referring to the content negatively feedback by the account to be recommended, such as the content not liked by the account to be recommended, the content marked as disliked by the account to be recommended, etc. It should be noted that the preset similarity can be 0.7, 0.6, etc., and the specific value can be adjusted according to the actual situation. In addition, the sample label of the sample content is also obtained by querying the correspondence between the pre-generated content and the category label.
[0074] Among them, the content associated with the category label and any sample label refers to the content whose label similarity between the category label and any sample label is greater than the first preset threshold; the first preset threshold can be 0.7, 0.6, etc., and the specific value can be adjusted according to the actual situation; when the label similarity between the category label and any sample label is 1, the content associated with the category label and any sample label specifically refers to the content where the category label and any sample label are the same.
[0075] Specifically, the server determines the content of the negative feedback of the account to be recommended based on the operation behavior log of the account to be recommended, and uses it as the sample content whose feature similarity with the account features of the account to be recommended is less than the preset similarity; queries the pre-generated correspondence between the content and the category label according to the sample content, and obtains the category label corresponding to the content identical to the sample content as the sample label of the sample content; deletes the content associated with the category label and any sample label from the content to be recommended, and obtains the remaining content as the target content. In this way, when recommending content to the account to be recommended, it is possible to accurately not recommend the content associated with the sample label of the sample content with negative feedback from the account to be recommended, thereby improving the accuracy of content recommendation.
[0076] For example, assume that the sample labels of the sample content are A and B, and the category labels of the content to be recommended are A, C, and D. Then the server deletes the content with the category label A from the content to be recommended to obtain the target content.
[0077] In step S240, the target content is recommended to the account to be recommended.
[0078] Specifically, the server pushes the target content to the terminal corresponding to the account to be recommended at a preset recommendation frequency, such as recommending 10 pieces of content per minute. The terminal displays the recommended target content through the terminal interface, meeting the interest needs of the account to be recommended, thereby realizing accurate content recommendation and further improving the accuracy of content recommendation.
[0079] In the above content recommendation method, by obtaining the content to be recommended corresponding to the account to be recommended; then querying the pre-generated correspondence between the content and the category label to obtain the category label corresponding to the content to be recommended; the correspondence is obtained by clustering the feature information of the content; then obtaining the sample label of the sample content corresponding to the account to be recommended, and deleting the content associated with the category label and any sample label from the content to be recommended to obtain the target content; the sample content is used to represent the content whose feature similarity with the account features of the account to be recommended is less than the preset similarity, and the sample label is obtained through the correspondence; finally, the target content is recommended to the account to be recommended; in this way, clustering the content using the feature information of the content to obtain the correspondence between the content and the category label is beneficial to avoiding the limitations of the manually divided classification system, thereby increasing the generalization ability for similar content, and further improving the accuracy of determining the category label, so that when recommending content to the account to be recommended, it is possible to accurately not recommend the content associated with the sample label of the sample content with negative feedback from the account to be recommended, thereby improving the accuracy of content recommendation.
[0080] In an exemplary embodiment, as Figure 3As shown, before obtaining the content to be recommended corresponding to the account to be recommended in step S210, it further includes a step of constructing the correspondence between the content and the category label, which can be specifically implemented through the following steps:
[0081] In step S310, obtain the feature information of the content.
[0082] Among them, the feature information of the content can be represented by a feature vector, such as a feature vector (A1, A2, A3 ······ An) with a preset length of N dimensions; for example, the feature information of the content is the feature vector (1, 3, 2, 4, 2 ······).
[0083] Specifically, the server performs feature extraction processing on the text information of the content (such as the content name, description information, etc.) through a feature extraction model to obtain the feature vector of the content as the feature information of the content.
[0084] For example, the server extracts the features of the content through feature extraction methods such as Word2Vec to obtain a feature vector with a preset length of N dimensions as the feature information of the content. Of course, the server can also input the text information of the content into a deep neural network model and output the feature vector of the content through the deep neural network model as the feature information of the content.
[0085] In step S320, cluster the content according to the feature information of the content to obtain a content set; the content set is matched with a category label, and the category label of the content in the content set is the same as the category label matched by the content set.
[0086] Among them, the number of content sets obtained by clustering generally refers to two or more. Each content set is matched with a category label, and the category label of each content in each content set is the same as the category label matched by the corresponding content set.
[0087] Specifically, the server clusters the content based on the feature information of the content according to a preset clustering instruction to obtain at least two content sets and labels each content set with a category label.
[0088] For example, the server clusters the content based on the feature information of the content according to the kmeans clustering algorithm to obtain multiple content sets, namely content set A, content set B, and content set C, and labels content set A with category label 1, labels content set B with category label 2, and labels content set C with category label 3. In this way, the category label of each content in content set A is category label 1, the category label of each content in content set B is category label 2, and the category label of each content in content set C is category label 3.
[0089] In step S330, according to the category labels of the contents in the content set, the corresponding relationship between the contents and the category labels is determined as the pre-generated corresponding relationship between the contents and the category labels.
[0090] Specifically, the server constructs the corresponding relationship between the contents and the category labels according to the category labels of each content in each content set, and uses this corresponding relationship as the pre-generated corresponding relationship between the contents and the category labels.
[0091] For example, assume that the category labels corresponding to contents A1, A2, and A3 are all category label 1, the category labels corresponding to contents B1, B2, and B3 are all category label 2, and the category labels corresponding to contents C1, C2, and C3 are all category label 3. Then the corresponding relationship between the contents and the category labels is (A1 - category label 1, A2 - category label 1, A3 - category label 1, B1 - category label 2, B2 - category label 2, B3 - category label 2, C1 - category label 3, C2 - category label 3, C3 - category label 3).
[0092] The technical solution provided by the embodiments of the present disclosure clusters the contents according to the feature information of the contents, obtains the category labels of multiple content sets, and obtains the corresponding relationship between the contents and the category labels according to the category labels of the multiple content sets, which is beneficial to avoiding the limitations of the manually divided classification system, thereby increasing the generalization ability of similar contents, and further improving the determination accuracy of the category labels, making the contents recommended based on the category labels more accurate subsequently.
[0093] In an exemplary embodiment, in the above step S320, clustering the contents according to the feature information of the contents to obtain a content set specifically includes: inputting the feature information of the contents into a preset clustering model, and through the preset clustering model, combining multiple contents with a similarity greater than a preset threshold between the feature information to obtain a content set.
[0094] Among them, the preset clustering model is a model used to cluster the contents, such as the kmeans clustering model. The similarity between the feature information is used to measure whether the contents are similar. In addition, it should be noted that the preset threshold can be adjusted according to the actual situation and is not specifically limited here.
[0095] Specifically, the server obtains the feature information of the content and inputs the feature information of the content into a preset clustering model; through the preset clustering model, it determines multiple contents whose similarity between feature information is greater than a preset threshold, and clusters together multiple contents whose similarity between feature information is greater than the preset threshold (such as 0.7) respectively, to obtain multiple content sets, such as content set A (content A1, content A2, content A3), content set B (content B1, content B2, content B3), content set C (content C1, content C2, content C3).
[0096] The technical solution provided by the embodiments of the present disclosure clusters the content based on the feature information of the content through a preset clustering model to obtain the category labels of the content. Throughout the process, it is not restricted by the artificially divided classification system, making the category labels of the subsequent content to be recommended more accurate.
[0097] In an exemplary embodiment, in the above step S310, obtaining the feature information of the content specifically includes: obtaining the text information of the content; preprocessing the text information of the content to obtain the preprocessed text information; through a feature extraction model, extracting the feature information of the preprocessed text information as the feature information of the content.
[0098] Among them, the text information of the content refers to natural language texts such as content description information and content names for clustering. Of course, it can also refer to other natural language texts with similarity between contents. Preprocessing refers to deleting invalid characters in the text information, performing word segmentation on the text information, and performing connection processing on the word segmentation, etc.
[0099] Among them, the feature extraction model refers to a model used to extract the feature information of the text information of the content, such as the Word2Vec model, deep neural network model, etc.
[0100] Specifically, the server obtains the text information of the content according to a preset text information acquisition instruction; preprocesses the text information of the content according to a preset preprocessing instruction to obtain the preprocessed text information; inputs the preprocessed text information into the feature extraction model, and through the feature extraction model, performs feature extraction processing on the preprocessed text information to obtain the feature information of the preprocessed text information, and uses the feature information of the preprocessed text information as the feature information of the content.
[0101] The technical solution provided by the embodiments of the present disclosure first preprocesses the text information of the content, and then extracts the feature information of the preprocessed text information as the feature information of the content, which is beneficial to improving the determination accuracy of the feature information of the content, and further improves the accuracy of the category labels obtained by clustering based on the feature information of the content.
[0102] In an exemplary embodiment, obtaining the text information of the content specifically includes: obtaining the content identifier and the description information of the content; combining the content identifier and the description information of the content to obtain the text information of the content.
[0103] Among them, the content identifier refers to the unique identifier information of the content, such as the content name; the description information refers to the brief introduction of the content.
[0104] For example, the server obtains the advertisement name "CCC" of advertisement A and the description information "A popular and exciting mini-game, interesting and fun", and combines the advertisement name "CCC" and the description information "A popular and exciting mini-game, interesting and fun" of advertisement A to obtain the text information of the content: name "CCC" - description information "A popular and exciting mini-game, interesting and fun".
[0105] The technical solution provided by the embodiments of the present disclosure is beneficial to subsequent clustering of the content based on the information of multiple dimensions of the content by obtaining the information of multiple dimensions of the content, so that the category labels obtained by clustering are closer to the user's cognition, thereby improving the determination accuracy of the category labels of the content to be recommended.
[0106] In an exemplary embodiment, preprocessing the text information of the content to obtain the preprocessed text information includes: deleting the invalid information in the content identifier and the invalid information in the description information to obtain the target content identifier and the target description information; performing word segmentation processing on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information; connecting the word segmentation of the target content identifier and the word segmentation of the target description information through a separator to obtain the preprocessed text information.
[0107] Among them, the invalid information refers to special characters, numbers, line breaks, etc. in the text information; the separator refers to a symbol for connecting the word segmentation of the target content identifier and the word segmentation of the target description information, such as $, %, #, etc.
[0108] Specifically, the server identifies the content identifier and the description information according to the invalid information recognition instruction to obtain the invalid information in the content identifier and the invalid information in the description information, and deletes the invalid information in the content identifier and the invalid information in the description information to obtain the target content identifier and the target description information; according to the Jieba word segmentation instruction, performs word segmentation processing on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information; deletes the stop words in the word segmentation of the target content identifier and the stop words in the word segmentation of the target description information to obtain the target word segmentation of the target content identifier and the target word segmentation of the target description information; obtains a separator, and uses the separator to connect the target word segmentation of the target content identifier and the target word segmentation of the target description information together as the preprocessed text information.
[0109] For example, for the text information of advertisement A: Name "CCC" - Description "A popular and exciting mini-game, interesting and fun", after removing special characters, it becomes:
[0110] Name: "CCC" - Description: "A popular and exciting mini-game interesting and fun";
[0111] After word segmentation, it becomes:
[0112] Name: "C C C" - Description: "A popular and exciting mini-game interesting and fun";
[0113] After removing stop words, it becomes:
[0114] Name: "C C C" - Description: "A popular and exciting mini-game interesting fun";
[0115] After connecting with a delimiter (such as $), it becomes:
[0116] "C C C$A popular and exciting mini-game interesting fun".
[0117] The technical solution provided by the embodiments of the present disclosure preprocesses the text information of the content, making the feature information of the content obtained from the subsequent text information of the content more accurate, thereby improving the determination accuracy of the feature information of the content, and further improving the accuracy of the category labels obtained by clustering based on the feature information of the content.
[0118] In an exemplary embodiment, in the above step S220, querying the pre-generated correspondence between the content and the category label to obtain the category label corresponding to the content to be recommended includes: obtaining the content identifier to be recommended of the content to be recommended; querying the pre-generated correspondence between the content and the category label to obtain the category label corresponding to the content with the same content identifier as the content to be recommended, as the category label of the content to be recommended.
[0119] Among them, the content identifier to be recommended refers to the identification information of the content to be recommended, such as content ID, content name, etc.
[0120] Specifically, the server obtains the content identifier to be recommended of the content to be recommended through a content identifier acquisition instruction; according to the content identifier to be recommended of the content to be recommended, queries the pre-generated correspondence between the content and the category label to obtain the content with the same content identifier as the content to be recommended, and uses the category label corresponding to this content as the category label of the content to be recommended.
[0121] For example, in the pre-generated correspondence between the content and the category label, the category label corresponding to the content with content ID 005 is A, and the content ID of the content to be recommended is also 005, indicating that the category label of the content to be recommended is also A.
[0122] The technical solution provided by the embodiments of the present disclosure obtains the category label corresponding to the content to be recommended by querying the pre-generated correspondence between the content and the category label, which is beneficial for accurately not recommending the content with the same sample label as the category label of the sample content with negative feedback from the account to be recommended when recommending content to the account to be recommended, thereby improving the recommendation accuracy of the content.
[0123] Figure 4 is a flowchart of another content recommendation method shown according to an exemplary embodiment. As Figure 4 shown, the content recommendation method is used for Figure 1 the server shown in, and includes the following steps:
[0124] In step S410, obtain the content identifier and description information of the content; combine the content identifier and description information of the content to obtain the text information of the content.
[0125] In step S420, delete the invalid information in the content identifier and the invalid information in the description information to obtain the target content identifier and the target description information.
[0126] In step S430, perform word segmentation processing on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information; connect the word segmentation of the target content identifier and the word segmentation of the target description information through a delimiter to obtain the preprocessed text information.
[0127] In step S440, extract the feature information of the preprocessed text information through a feature extraction model as the feature information of the content.
[0128] In step S450, input the feature information of the content into a preset clustering model, and through the preset clustering model, combine multiple contents with a similarity greater than a preset threshold between the feature information to obtain a content set.
[0129] Among them, each content set is matched with a category label, and the category label of each content in each content set is the same as the category label matched by the corresponding content set.
[0130] In step S460, determine the correspondence between the content and the category label according to the category label of the content in the content set as the pre-generated correspondence between the content and the category label.
[0131] In step S470, obtain the content to be recommended corresponding to the account to be recommended; obtain the content identifier to be recommended of the content to be recommended; query the pre-generated correspondence between the content and the category label, and obtain the category label corresponding to the content with the same content identifier as the content to be recommended as the category label of the content to be recommended.
[0132] In step S480, the sample label of the sample content corresponding to the account to be recommended is obtained, and the category label and the content associated with any sample label are deleted from the content to be recommended to obtain the target content; the sample content is used to represent the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, and the sample label is obtained through the corresponding relationship.
[0133] In step S490, the target content is recommended to the account to be recommended.
[0134] In the above content recommendation method, the characteristic information of the content is used to cluster the content to obtain the correspondence between the content and the category label, which is conducive to avoiding the limitations of the artificial classification system, thereby increasing the generalization ability of similar content, and further improving the accuracy of determining the category label, so that when recommending content to the account to be recommended, the content whose category label is associated with the sample label of the sample content of the negative feedback of the account to be recommended can be accurately not recommended, thereby improving the recommendation accuracy of the content.
[0135] In an exemplary embodiment, if Figure 5 As shown, an advertisement recommendation method is provided, which mainly clusters advertisements according to text information such as descriptions and names of advertisements to obtain clustering results; the clustering results are queried according to the negative feedback advertisements of users to obtain a label set of negative feedback advertisements, and advertisements with the same label are filtered according to the label set of negative feedback advertisements to obtain target advertisements; the specific contents are as follows:
[0136] (1) Build a clustering environment, including the following: operating system, etc.
[0137] (2) Prepare natural language texts such as descriptions and names for clustering. Each piece of data is a combination of name and description, for example, name: "CCC"-description: "A popular mini-game that is interesting and fun."
[0138] (3) Perform data preprocessing; first, remove special characters, numbers, and line breaks from the text; then use the Jieba word segmentation to segment the text into Chinese words; then, remove stop words from the text; finally, use a separator to connect different texts together as raw data for feature extraction. For example, for name: "CCC" - description: "Hot and popular mini-game, interesting and fun",
[0139] After removing the special characters it becomes:
[0140] Name: "CCC" - Description: "Hot and popular games are fun and interesting"
[0141] After word segmentation, it becomes:
[0142] Name: "CCC" - Description: "A popular and interesting mini-game";
[0143] After removing stop words, it becomes:
[0144] Name: "CCC" - Description: "A popular and interesting mini-game";
[0145] After connecting with a delimiter (such as $), it becomes:
[0146] "CCC$A popular and interesting mini-game".
[0147] (4) Extract features; use feature extraction methods such as Word2Vec provided by the software package to extract features, and obtain a vector with a preset length of N dimensions as the features of this text.
[0148] (5) Feature clustering; use clustering methods such as kmeans to cluster the samples, and sample and evaluate the distribution of samples in each category, and readjust the parameters.
[0149] (6) Label the samples; according to the clustering results, label different samples. This label is only used as an identifier and does not have to have real semantics, such as "Category 1", "Category 2", "Category 3", etc.
[0150] (7) Obtain the negative feedback samples of the user, query the clustering results, and obtain the label set of these samples; when recommending advertisements to the user again, filter the advertisements with the same label according to the label set of the negative feedback.
[0151] It should be noted that the existing methods generally rely on a manually edited label system, and each label has a clear name and meaning. The results based on clustering completely depend on a certain learned metric. Although it does not have the high interpretability of manually defined labels, each category can capture deeper similarities and has better effects when recommending advertisements according to user preferences; in addition, when recommending positive advertising materials, users are often given a label set for them to make preference selections, and advertisements are recommended according to the labels selected by the users. However, when users express negative feedback, they only express their dislike for a certain advertising material, but do not express the advertising categories they are not interested in, resulting in the inability to accurately recommend advertisements to users according to advertising categories.
[0152] The above advertisement recommendation method can achieve the following technical effects: (1) Using the clustering method for classification avoids the limitations of the manually divided classification system and increases the generalization ability of the system for similar advertising materials; (2) By filtering similar advertisements, the negative feedback rate significantly decreases after the user gives negative feedback.
[0153] It should be understood that although Figures 2 - 4The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 - 4 At least some of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0154] Figure 6 is a block diagram of a content recommendation device shown according to an exemplary embodiment. Referring to Figure 6 , the content recommendation device includes an acquisition unit 610, a query unit 620, a deletion unit 630, and a recommendation unit 640.
[0155] The acquisition unit 610 is configured to acquire the content to be recommended corresponding to the account to be recommended.
[0156] The query unit 620 is configured to query the correspondence between the pre-generated content and the category label, and obtain the category label corresponding to the content to be recommended; the correspondence is obtained by clustering the feature information of the content.
[0157] The deletion unit 630 is configured to acquire the sample label of the sample content corresponding to the account to be recommended, delete the content associated with the category label and any sample label from the content to be recommended, and obtain the target content; the sample content is used to characterize the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, and the sample label is obtained through the correspondence.
[0158] The recommendation unit 640 is configured to recommend the target content to the account to be recommended.
[0159] In an exemplary embodiment, the content recommendation device further includes a determination unit, which is configured to acquire the feature information of the content; cluster the content according to the feature information of the content to obtain a content set; the content set is matched with a category label, and the category label of the content in the content set is the same as the category label matched by the content set; determine the correspondence between the content and the category label according to the category label of the content in the content set, and use it as the pre-generated correspondence between the content and the category label.
[0160] In an exemplary embodiment, the determination unit is further configured to input the feature information of the content into a preset clustering model, and through the preset clustering model, combine multiple contents with a similarity greater than a preset threshold between the feature information to obtain a content set.
[0161] In an exemplary embodiment, the determining unit is further configured to execute: obtaining text information of the content; preprocessing the text information of the content to obtain preprocessed text information; and extracting feature information of the preprocessed text information through a feature extraction model as the feature information of the content.
[0162] In an exemplary embodiment, the determining unit is further configured to execute: obtaining the content identifier and description information of the content; and combining the content identifier and description information of the content to obtain the text information of the content.
[0163] In an exemplary embodiment, the determining unit is further configured to execute: deleting invalid information in the content identifier and invalid information in the description information to obtain a target content identifier and target description information; performing word segmentation processing on the target content identifier and the target description information respectively to obtain word segmentation of the target content identifier and word segmentation of the target description information; and connecting the word segmentation of the target content identifier and the word segmentation of the target description information through a delimiter to obtain preprocessed text information.
[0164] In an exemplary embodiment, the querying unit 620 is further configured to execute: obtaining the identifier of the content to be recommended of the content to be recommended; and querying the pre-generated corresponding relationship between the content and the category label to obtain the category label corresponding to the content with the same content identifier as the content to be recommended as the category label of the content to be recommended.
[0165] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0166] Figure 7 is a block diagram of a device 700 for performing the above content recommendation method shown according to an exemplary embodiment. For example, the device 700 may be a server. Referring to Figure 7 , the device 700 includes a processing component 720, which further includes one or more processors, and memory resources represented by a memory 722 for storing instructions executable by the processing component 720, such as application programs. The application programs stored in the memory 722 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 720 is configured to execute instructions to perform the above content recommendation method.
[0167] Device 700 may also include a power supply component 724 configured to perform power management of device 700, a wired or wireless network interface 726 configured to connect device 700 to a network, and an input / output (I / O) interface 728. Device 700 may operate based on an operating system stored in memory 722, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0168] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as memory 722 including instructions, which can be executed by a processor of device 700 to complete the above-described content recommendation method. The computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0169] In an exemplary embodiment, there is also provided a computer program product. The program product includes a computer program stored in a readable storage medium. At least one processor of the device reads and executes the computer program from the readable storage medium, so that the device executes the content recommendation method described in any one of the embodiments of the present disclosure.
[0170] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0171] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A content recommendation method, characterized in that, Including: Obtain the content to be recommended corresponding to the account to be recommended; Query the corresponding relationship between the pre-generated content and the category label to obtain the category label corresponding to the content to be recommended; The corresponding relationship is obtained by clustering the feature information of the content; Obtain the sample label of the sample content corresponding to the account to be recommended, and delete the content associated with the category label and any of the sample labels from the content to be recommended to obtain the target content; The sample content is used to represent the content whose feature similarity with the account feature of the account to be recommended is less than the preset similarity, and the sample label is obtained through the corresponding relationship; Recommend the target content to the account to be recommended; Before obtaining the content to be recommended corresponding to the account to be recommended, it further includes: Obtain the feature information of the content; Cluster the content according to the feature information of the content to obtain a content set; the content set is matched with a category label, and the category label of the content in the content set is the same as the category label matched by the content set; Determine the corresponding relationship between the content and the category label according to the category label of the content in the content set as the corresponding relationship between the pre-generated content and the category label.
2. The content recommendation method according to claim 1, characterized in that, The clustering the content according to the feature information of the content to obtain a content set includes: Input the feature information of the content into a preset clustering model, and through the preset clustering model, combine multiple contents with a similarity greater than a preset threshold between the feature information to obtain the content set.
3. The content recommendation method according to claim 1, characterized in that The obtaining the feature information of the content includes: Obtain the text information of the content; Preprocess the text information of the content to obtain the preprocessed text information; Extract the feature information of the preprocessed text information through a feature extraction model as the feature information of the content.
4. The content recommendation method according to claim 3, wherein The obtaining the text information of the content includes: Obtain the content identifier and description information of the content; Combine the content identifier and description information of the content to obtain the text information of the content.
5. The content recommendation method according to claim 4, characterized in that The preprocessing the text information of the content to obtain the preprocessed text information includes: Delete the invalid information in the content identifier and the invalid information in the description information to obtain the target content identifier and the target description information; Perform word segmentation processing on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information; Connect the word segmentation of the target content identifier and the word segmentation of the target description information through a separator to obtain the preprocessed text information.
6. The content recommendation method according to claim 1, characterized in that The querying the corresponding relationship between the pre-generated content and the category label to obtain the category label corresponding to the content to be recommended includes: Obtain the content identifier to be recommended of the content to be recommended; Query the corresponding relationship between the pre-generated content and the category label to obtain the category label corresponding to the content with the same content identifier as the content to be recommended as the category label of the content to be recommended.
7. A content recommendation device, characterized in that, Including: An obtaining unit configured to execute obtaining the content to be recommended corresponding to the account to be recommended; A query unit, configured to execute querying the corresponding relationship between pre-generated content and category labels, and obtain the category labels corresponding to the content to be recommended; the corresponding relationship is obtained by clustering the feature information of the content. A deletion unit, configured to execute obtaining the sample labels of the sample content corresponding to the account to be recommended, and deleting from the content to be recommended the content associated with the category labels and any of the sample labels, to obtain target content. The sample content is used to represent content whose feature similarity with the account features of the account to be recommended is less than a preset similarity, and the sample labels are obtained through the corresponding relationship. A recommendation unit, configured to execute recommending the target content to the account to be recommended. The content recommendation device further includes a determination unit, configured to execute obtaining the feature information of the content; clustering the content according to the feature information of the content to obtain a content set; the content set is matched with category labels, and the category labels of the content in the content set are the same as the category labels matched by the content set; determining the corresponding relationship between the content and the category labels according to the category labels of the content in the content set, as the pre-generated corresponding relationship between the content and the category labels.
8. The content recommendation device according to claim 7, characterized in that, The determination unit is further configured to execute inputting the feature information of the content into a preset clustering model, and through the preset clustering model, combining multiple contents with a similarity greater than a preset threshold between the feature information respectively, to obtain the content set.
9. The content recommendation device according to claim 7, wherein The determination unit is further configured to execute obtaining the text information of the content; preprocessing the text information of the content to obtain preprocessed text information; extracting the feature information of the preprocessed text information through a feature extraction model, as the feature information of the content.
10. The content recommendation device according to claim 9, characterized in that, The determination unit is further configured to execute obtaining the content identifier and description information of the content; combining the content identifier and description information of the content to obtain the text information of the content.
11. The content recommendation device according to claim 10, characterized in that, The determination unit is further configured to execute deleting the invalid information in the content identifier and the invalid information in the description information to obtain a target content identifier and a target description information; performing word segmentation processing on the target content identifier and the target description information respectively to obtain the word segmentation of the target content identifier and the word segmentation of the target description information; connecting the word segmentation of the target content identifier and the word segmentation of the target description information through a delimiter to obtain the preprocessed text information.
12. The content recommendation device according to claim 7, wherein The query unit is further configured to execute obtaining the content identifier to be recommended of the content to be recommended; querying the pre-generated corresponding relationship between the content and the category labels, and obtaining the category labels corresponding to the content with the same content identifier as the content to be recommended, as the category labels of the content to be recommended.
13. A server, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the content recommendation method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of the server, the server is enabled to execute the content recommendation method according to any one of claims 1 to 6.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the content recommendation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Content recommendation method and device, computer readable storage medium and computer equipment
CN110263242A
Systems and methods for analyzing user activity
US20180336582A1