Information Recommendation Method, Device, and Computer Readable Storage Medium

By constructing multi-dimensional user portraits and real-time updates, combining pre-trained models and display sorting algorithms, the problem of incomplete user portraits in the existing technology is solved, and the real-time and diversity of recommended results are improved.

CN114398523BActive Publication Date: 2025-07-01CHENGDU TUBALONG INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210012274.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-07
Publication Date
2025-07-01
Estimated Expiration
2042-01-07

AI Technical Summary

Technical Problem

The existing technology lacks comprehensiveness and poor real-time performance when building user portraits, resulting in low quality of recommendation results, inability to track user interests in time, and the recommendation results are too single.

Method used

By constructing multi-dimensional user portraits, including positive preference portraits, negative masked portraits and historical behavior portraits, the user portraits are updated in real time, and the recommended results are generated using pre-trained models and display sorting algorithms.

Benefits of technology

It improves the compliance of the recommended results with user expectations, enhances the real-time and diversity of the recommended results, and solves the problem of incomplete interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398523B_ABST
    Figure CN114398523B_ABST
Patent Text Reader

Abstract

The present invention discloses an information recommendation method, device and computer-readable storage medium. Among them, the method includes: obtaining in real time a set of information to be recommended within a preset period and a multi-dimensional user portrait of the current user; filtering the set of information to be recommended according to the historical behavior portrait and the negative shielding portrait to obtain a first sequence of information to be recommended; generating a feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended according to the positive preference portrait; inputting the feature sequence into a pre-trained model to obtain an estimated recommendation value corresponding to each piece of information to be recommended in the first sequence of information to be recommended; sorting the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated recommendation value to obtain a second sequence of information to be recommended; and displaying the information to be recommended in the second sequence of information to be recommended in sequence through a sorting display algorithm to obtain an information recommendation result. The quality, content diversity and user interest compliance of the information recommendation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimedia information dissemination, and particularly to a method and device for information recommendation and a computer-readable storage medium. Background Art

[0002] With the development of the Internet, the release of information has become more convenient, and the number of information has shown an explosive growth. Traditional search engines rely on users to actively input content. When users do not know a new hot topic or do not need to accurately obtain certain aspects of information, but hope to obtain the information they are interested in in a timely manner from a large amount of information, the search engine seems powerless. As an important supplement to the search engine, the recommendation system can better meet the diverse needs of users for a large amount of information. The recommendation system infers the user's interest preferences based on the user's attributes and behaviors, and finds the content that the user may be interested in or has potential interest in a large amount of information, greatly improving the user experience in the era of information overload.

[0003] The existing recommendation technologies mainly construct user portraits, abstract the recorded user behaviors into interest vectors, and recommend the information with higher matching to users according to the similarity between the interest vectors and the information vectors to be recommended. However, the existing technologies have many disadvantages. First, the quality of the recommendation calculation results is not high due to the lack of a comprehensive user portrait. Second, the real-time performance of constructing the portrait is poor, resulting in the inability to track the user's interests in a timely manner for the recommended results. Third, the recommended results are highly similar, resulting in too single content for display.

[0004] In view of the above problems existing in the prior art, there is currently no effective solution. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method and device for information recommendation and a computer-readable storage medium. By constructing a multi-dimensional user portrait, the problem of how to comprehensively construct a user portrait is solved; the user portrait is updated in a timely manner, and the user portrait and the information set to be recommended are extracted in real time, so as to solve the problem of poor real-time performance in the prior art; a pre-trained model and a display sorting algorithm are used to generate recommended results, so as to solve the problems of incomplete interest points, lagging relevance of recommended results, and too single content in the prior art.

[0006] To achieve the above object, the present invention provides an information recommendation method, including: S1. Obtain a set of information to be recommended within a preset period and a multi-dimensional user profile of the current user in real time; wherein, the multi-dimensional user profile includes a positive preference profile, a negative shielding profile, and a historical behavior profile; S2. Filter the set of information to be recommended according to the historical behavior profile and the negative shielding profile to obtain a first sequence of information to be recommended; S3. Generate a feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended one by one according to the positive preference profile; S4. Input the feature sequence into a pre-trained model to obtain an estimated recommendation value corresponding to each piece of information to be recommended in the first sequence of information to be recommended; S5. Sort the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated recommendation value to obtain a second sequence of information to be recommended; S6. Display the information to be recommended in the second sequence of information to be recommended in order through a sorting display algorithm to obtain an information recommendation result.

[0007] Further optionally, generating a feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended one by one according to the positive preference profile includes: S301. Select each field value of the positive preference profile or the feature value of the information to be recommended in the first sequence of information to be recommended as a first-order feature sequence; S302. Establish a second-order feature sequence according to the matching relationship between each field value of the positive preference profile and the feature value of the information to be recommended in the first sequence of information to be recommended; S303. Concatenate the first-order feature sequence and the second-order feature sequence to obtain a feature sequence corresponding to each piece of information to be recommended.

[0008] Further optionally, presenting the to-be-recommended information in the second to-be-recommended information sequence in order through a sorting and presentation algorithm to obtain an information recommendation result includes: S601. Store the to-be-recommended information ranked first in the second to-be-recommended information sequence at the first position of a third to-be-recommended information sequence; wherein, the third to-be-recommended information sequence is preset and is empty in the initial state; S602. Determine whether the second to-be-recommended information sequence has been traversed, or whether the number of to-be-recommended information in the third to-be-recommended information sequence has reached a preset number; if so, use the third to-be-recommended information sequence at the current moment as the information recommendation result; if not, use the next to-be-recommended information in the second to-be-recommended information sequence as the current to-be-recommended information, and execute step S603: S603. Sequentially determine whether the number of available single categories corresponding to the category of the current to-be-recommended information is greater than or equal to a preset single-category number threshold, whether the number of available single channels corresponding to the channel of the current to-be-recommended information is greater than or equal to a preset single-channel number threshold, and whether the similarity between the keyword of the current to-be-recommended information and the keyword of the to-be-recommended information at the end of the third to-be-recommended information sequence is greater than a preset similarity threshold. If any of the judgment results is yes, jump to step S602; if all the judgment results are no, increment the number of available single categories corresponding to the category of the current to-be-recommended information by 1, increment the number of available single channels corresponding to the channel of the current to-be-recommended information by 1, and store the current to-be-recommended information at the end of the third to-be-recommended information sequence, then jump to step S602.

[0009] Further optionally, after obtaining the information recommendation result, it further includes: S7. Update the multi-dimensional user portrait of the current user using a timed decay update method; and / or, S8. Update the multi-dimensional user portrait of the current user using streaming asynchronous computing based on user behavior data and the information recommendation result.

[0010] Further optionally, filtering the to-be-recommended information set according to the historical behavior portrait and the negative shielding portrait includes: S201. Identify the first information ID corresponding to each to-be-recommended information in the to-be-recommended information set, and the second information ID corresponding to the recommended information in the historical behavior portrait. When the first information ID is the same as the second information ID, delete the to-be-recommended information corresponding to the first information ID from the to-be-recommended information set; S202. Identify the shielding value corresponding to each feature in the negative shielding portrait, and delete the features corresponding to the shielding values less than a preset shielding value threshold to obtain an effective negative shielding portrait; wherein, the shielding value decays over time; S203. In the to-be-recommended information set, delete the to-be-recommended information that conforms to the features of the effective negative shielding portrait.

[0011] On the other hand, the present invention also provides an information recommendation device, including: a data acquisition module, configured to acquire in real time a set of information to be recommended within a preset period and a multi-dimensional user profile of the current user; wherein, the multi-dimensional user profile includes a positive preference profile, a negative shielding profile and a historical behavior profile; a first information sequence to be recommended generation module, configured to filter the set of information to be recommended according to the historical behavior profile and the negative shielding profile to obtain a first information sequence to be recommended; a feature sequence generation module, configured to generate, one by one according to the positive preference profile, a feature sequence corresponding to each piece of information to be recommended in the first information sequence to be recommended; an estimated recommendation value generation module, configured to input the feature sequence into a pre-trained model to obtain an estimated recommendation value corresponding to each piece of information to be recommended in the first information sequence to be recommended; a second information sequence to be recommended generation module, configured to sort the information to be recommended in the first information sequence to be recommended from high to low according to the corresponding estimated recommendation value to obtain a second information sequence to be recommended; an information recommendation result generation module, configured to display the information to be recommended in the second information sequence to be recommended in order through a sorting display algorithm to obtain an information recommendation result.

[0012] Further optionally, the feature sequence generation module includes: a first-order feature sequence generation sub-module, configured to select each field value of the positive preference profile or the feature value of the information to be recommended in the first information sequence to be recommended as the first-order feature sequence; a second-order feature sequence generation sub-module, configured to establish a second-order feature sequence according to the matching relationship between each field value of the positive preference profile and the feature value of the information to be recommended in the first information sequence to be recommended; a feature splicing sub-module, configured to splice the first-order feature sequence and the second-order feature sequence to obtain a feature sequence corresponding to each piece of information to be recommended.

[0013] Further optionally, the information recommendation result generation module includes: a first information storage sub-module for storing the to-be-recommended information ranked first in the second to-be-recommended information sequence at the first position of the third to-be-recommended information sequence; wherein, the third to-be-recommended information sequence is preset and is empty in the initial state; a traversal sub-module for determining whether the second to-be-recommended information sequence has been traversed or whether the number of to-be-recommended information in the third to-be-recommended information sequence reaches a preset number; if so, taking the third to-be-recommended information sequence at the current moment as the information recommendation result; if not, taking the next to-be-recommended information in the second to-be-recommended information sequence as the current to-be-recommended information and performing the following steps: sequentially determining whether the number of available single classifications corresponding to the classification of the current to-be-recommended information is greater than or equal to a preset single classification number threshold, whether the number of available single channels corresponding to the channel of the current to-be-recommended information is greater than or equal to a preset single channel number threshold, and whether the similarity between the keyword of the current to-be-recommended information and the keyword of the last to-be-recommended information in the third to-be-recommended information sequence is greater than a preset similarity threshold; if any of the judgment results is yes, jumping to the above judgment step; if all the judgment results are no, adding 1 to the number of available single classifications corresponding to the classification of the current to-be-recommended information, adding 1 to the number of available single channels corresponding to the channel of the current to-be-recommended information, and storing the current to-be-recommended information at the end of the third to-be-recommended information sequence, then jumping to the above judgment step.

[0014] Further optionally, the device further includes: a first update module for updating the multi-dimensional user portrait of the current user by using a timed attenuation update method; a second update module for updating the multi-dimensional user portrait of the current user by using streaming asynchronous calculation according to user behavior data and information recommendation results.

[0015] Further optionally, the first to-be-recommended information sequence generation module includes: a first filtering sub-module for identifying the first information ID corresponding to each to-be-recommended information in the to-be-recommended information set and the second information ID corresponding to the recommended information in the historical behavior portrait, and when the first information ID is the same as the second information ID, deleting the to-be-recommended information corresponding to the first information ID from the to-be-recommended information set; a shielding value identification sub-module for identifying the shielding value corresponding to each feature in the negative shielding portrait, and deleting the features corresponding to the shielding values less than a preset shielding value threshold to obtain an effective negative shielding portrait; wherein, the shielding value decays with time; a second filtering sub-module for deleting the to-be-recommended information in the to-be-recommended information set that conforms to the features of the effective negative shielding portrait.

[0016] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above information recommendation method is implemented.

[0017] The above technical solution has the following beneficial effects: By constructing the multi-dimensional positive and negative user portraits, the problem of incomplete points of interest in the prior art is solved, and the compliance of the generated recommendation results with user expectations is improved; By time decay and real-time streaming computing, the problems that user interests change over time, user interests are not updated in a timely manner, and the computational complexity of updating user portraits in a specified time period is large are solved, and the real-time performance of obtaining recommendation results is improved; Through user portraits, information pools, and pre-trained models, the quality of recommendation results is improved; Through the display sorting algorithm, the problem that the content of recommendation results is too concentrated is solved, and the diversity of recommendation results is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a flowchart of an information recommendation method provided by an embodiment of the present invention;

[0020] Figure 2 It is a flowchart of a feature sequence generation method provided by an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of an information recommendation result provided by an embodiment of the present invention;

[0022] Figure 4 It is a flowchart of a sorting display algorithm provided by an embodiment of the present invention;

[0023] Figure 5 It is a flowchart of a multi-dimensional user portrait update method provided by an embodiment of the present invention;

[0024] Figure 6 It is a data flow diagram of the recommendation process and portrait update provided by an embodiment of the present invention;

[0025] Figure 7 It is a flowchart of a first to-be-recommended information sequence generation method provided by an embodiment of the present invention;

[0026] Figure 8 It is a structural schematic diagram of an information recommendation device provided by an embodiment of the present invention;

[0027] Figure 9 It is a structural schematic diagram of a feature sequence generation module provided by an embodiment of the present invention;

[0028] Figure 10It is a schematic structural diagram of an information recommendation result generation module provided by an embodiment of the present invention;

[0029] Figure 11 It is a schematic structural diagram of a first update module and a second update module provided by an embodiment of the present invention;

[0030] Figure 12 It is a schematic structural diagram of a first information sequence to be recommended generation module provided by an embodiment of the present invention.

[0031] Reference numerals: 100 - data acquisition module; 200 - first information sequence to be recommended generation module; 2001 - first filtering sub-module; 2002 - shielding value identification sub-module; 2003 - second filtering sub-module; 300 - feature sequence generation module; 3001 - first-order feature sequence generation sub-module; 3002 - second-order feature sequence generation sub-module; 3003 - feature splicing sub-module; 400 - estimated recommendation value generation module; 500 - second information sequence to be recommended generation module; 600 - information recommendation result generation module; 6001 - first information storage sub-module; 6002 - traversal sub-module; 700 - first update module; 800 - second update module Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] To solve the problems of incomplete construction of points of interest, lag in relevance of recommendation results, and single content in the prior art, an embodiment of the present invention provides an information recommendation method. Figure 1 It is a flowchart of the information recommendation method provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0034] S1. Real-time obtain the information set to be recommended within a preset period and the multi-dimensional user portrait of the current user; wherein, the multi-dimensional user portrait includes a positive preference portrait, a negative shielding portrait, and a historical behavior portrait;

[0035] Construct an information pool to be recommended, that is, an information set to be recommended, by using the information collected within a specified time range. The specified time range is a preset period, which can be a fixed period set manually or a time range dynamically adjusted according to the size of the information to be recommended uploaded.

[0036] The information set to be recommended is continuously updated over time. When receiving an information recommendation request, based on the current moment, obtain the information set to be recommended within a specified time range to ensure the timeliness and data volume of the information set to be recommended.

[0037] The multi-dimensional user profile is continuously updated over time. When receiving an information recommendation request, obtain the multi-dimensional user profile at the current moment. The recommendation results obtained by analyzing with the multi-dimensional user profile at this moment are more accurate, improving the relevance and compliance of the recommendation results.

[0038] The categories of the multi-dimensional user profile include: positive preference profile, negative shielding profile, and historical behavior profile.

[0039] As an optional implementation method, the positive preference profile includes: interest preference profile, classification preference profile, channel preference profile, and collection preference profile.

[0040] As an optional implementation method, the negative shielding profile includes: blocked keyword profile, blocked channel profile, and disliked interest profile.

[0041] As an optional implementation method, the historical behavior profile is a record of the information that has been shown to the user within a specified time.

[0042] To construct the user's multi-dimensional profile, based on the user behavior data or recommendation result data, combined with the information features, calculate the key-value pairs of keywords. When calculating, the value of the information feature is defaulted to 1.0. For example, the calculation result of a certain field user profile is {"football": 1.0, "basketball": 3.0}. The specific calculation methods for each field of the multi-dimensional user profile are shown in Table 1:

[0043] User Portrait Category User Portrait Field Calculation Method Positive Preference Portrait Interest Preference Portrait Accumulated User Behavior Weight Multiplier Information Keywords Positive Preference Portrait Classification Preference Portrait Accumulated Information Categories Corresponding to User Reading Behavior Positive Preference Portrait Channel Preference Portrait Accumulated Information Channels Corresponding to User Reading Behavior Positive Preference Portrait Favorite Preference Portrait Accumulated Information Keywords Corresponding to User Favorite Behavior Negative Blocking Portrait Blocking Keyword Portrait Accumulated Information Keywords Corresponding to User Blocking Behavior Negative Blocking Portrait Blocking Channel Portrait Accumulated Information Channels Corresponding to User Blocking Behavior Negative Blocking Portrait Disliked Interest Portrait Accumulated Information Keywords Corresponding to User Dislike Behavior Historical Behavior Portrait Historical Behavior Portrait Concatenate the Information IDs That Have Been Recommended to the User

[0044] Table 1

[0045] Among them, the user behavior data includes the user reading a certain piece of information, collecting a certain piece of information, blocking the keywords of a certain piece of information, and disliking a certain piece of information.

[0046] The user behavior weight multiple table can be manually specified, and the user behavior weights are specifically reading, collecting, blocking, and disliking.

[0047] As an optional implementation method, the weight value table is shown in Table 2:

[0048]

[0049]

[0050] Table 2

[0051] Cumulative calculation refers to adding the value generated by a new behavior to the existing value of a specific field in the user profile. For example, if the user's existing interest preference profile is {"Football": 1.0, "Basketball": 3.0}, and the user newly generates a favorite behavior with the information keywords "Football, David Beckham", the cumulative result is {"Football": 6.0, "Basketball": 3.0, "David Beckham": 5.0}.

[0052] Information features refer to information ID, information keywords, information classification, and information channels. Among them, the information ID is an identifier that can uniquely identify an information; the information keywords are one or a group of words that can represent the theme content of the information. The method of extracting information keywords can be specified by manual editing or generated by a computer program. For example, after segmenting the information content text, the top K words with higher weights are found using algorithms such as TF-IDF as the keywords of the information; the information classification refers to the category of the information, such as sports, entertainment, health preservation, etc. The generation of information classification can be specified by manual editing or by using a computer program. For example, a classification model is trained using classification algorithms such as Naive Bayes, and the information classification is generated using this model; the information channel refers to the source of the information, such as Xinhua News Agency, People's Daily, etc.

[0053] Building a multi-dimensional user profile enriches the feature space for user analysis, improves the relevance and richness of information recommendation results, and avoids the problem of single information recommendation results.

[0054] S2. Filter the set of information to be recommended according to the historical behavior profile and the negative shielding profile to obtain the first sequence of information to be recommended;

[0055] After obtaining the set of information to be recommended, check each piece of information to be recommended in the set of information to be recommended one by one, and delete the information to be recommended that conforms to the content of the historical behavior profile and the content of the negative shielding profile to obtain the first sequence of information to be recommended.

[0056] These deleted information to be recommended may be the information that has been recommended to the user within a specified time or the information that the user has subjectively shielded. The deletion of this information can reduce the duplication of information recommended to the user within a specified time and improve the compliance of the recommended information with the user.

[0057] S3. Generate the feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended one by one according to the positive preference profile;

[0058] Use the features of the positive preference profile and the features of each piece of information to be recommended in the first sequence of information to be recommended to generate the feature sequences of all the information to be recommended in the first sequence of information to be recommended.

[0059] S4. Input the feature sequence into the pre-trained model to obtain the estimated recommendation value corresponding to each piece of information to be recommended in the first sequence of information to be recommended;

[0060] A pre-trained model refers to a probability prediction model generated by a classification algorithm using user portrait data, historical information data, and records of whether a user has read a certain piece of information, which are composed of the user's historical behavior data. Using this model, the probability value of a user reading a certain piece of information can be obtained. As an alternative implementation, the classification algorithm is Naive Bayes, Logistic Regression, Decision Tree, etc.

[0061] Input the feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended into the pre-trained model, and the estimated recommendation value corresponding to each piece of information to be recommended can be obtained. Among them, the estimated recommendation value represents the degree of compliance of this piece of information with the user's needs. The higher the estimated recommendation value of the information to be recommended, the more it meets the user's needs. On the contrary, the lower the estimated recommendation value of the information to be recommended, the less it meets the user's needs.

[0062] S5. Sort the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated recommendation value to obtain the second sequence of information to be recommended;

[0063] Sort the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated value to obtain the sorted second sequence of information to be recommended. Among the second sequence of information to be recommended, the information to be recommended at the head is more in line with the user's expectations than the information to be recommended at the tail.

[0064] S6. Display the information to be recommended in the second sequence of information to be recommended in order through a sorting display algorithm to obtain the information recommendation result.

[0065] Display the information to be recommended in the second sequence of information to be recommended with a sorting display algorithm to obtain the final information recommendation result.

[0066] As an alternative implementation Figure 2 is the flowchart of the feature sequence generation method provided by the embodiments of the present invention. As Figure 2 shown, according to the positive preference portrait, generate the feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended one by one, including:

[0067] S301. Select each field value of the positive preference portrait or the feature value of the information to be recommended in the first sequence of information to be recommended as the first-order feature sequence;

[0068] The first-order feature sequence refers to the sequence composed of each field value of the direct user's positive preference portrait or the feature value of the information to be recommended.

[0069] As an alternative implementation, the generation process of the first-order feature sequence is specifically as follows: Use the interest preference profile to generate key-value pairs above a specified threshold (for example, the generation result is: "F11_football": 6.0, "F11_basketball": 3.0, "F11_David Beckham": 5.0, where F11 refers to the feature index name), use the classification preference profile to generate key-value pairs above a specified threshold, use the channel preference profile to generate key-value pairs above a specified threshold, use the collection preference profile to generate key-value pairs above a specified threshold, use the information keywords and assign a value of 1.0 to each keyword, use the information classification and assign a value of 1.0 to this classification, and use the information channel and assign a value of 1.0 to this channel.

[0070] S302. Establish a second-order feature sequence according to the matching relationship between the field values of the positive preference profile and the feature values of the to-be-recommended information in the first to-be-recommended information sequence;

[0071] The second-order feature sequence refers to a feature sequence formed by using the matching relationship between the field values of the user's positive preference profile and the feature values of the to-be-recommended information.

[0072] As an alternative implementation, the generation process of the second-order feature sequence is as follows: Whether there is an overlap between the keywords above the specified threshold in the user interest preference profile and the information keywords (for example, the generation result is: "F21": 1, where F21 refers to the feature index name, and its value 1 indicates that there is an overlap), whether there is an overlap between the classifications above the specified threshold in the user classification preference profile and the information classification, whether there is an overlap between the channels above the specified threshold in the user channel preference profile and the information channel, and whether there is an overlap between the keywords above the specified threshold in the user collection preference profile and the information keywords.

[0073] S303. Concatenate the first-order feature sequence and the second-order feature sequence to obtain the feature sequence corresponding to each to-be-recommended information.

[0074] The first-order feature sequence and the second-order feature sequence are concatenated end to end to form a feature sequence. Since the feature sequence includes the first-order feature sequence and the second-order feature sequence, and both the first-order feature sequence and the second-order feature sequence are constructed by different methods, the feature space of the concatenated feature sequence is rich, which improves the accuracy of the estimated recommendation value obtained by inputting it into the pre-trained model.

[0075] As an alternative implementation, Figure 3 is the flowchart of the sorting and display method provided by the embodiment of the present invention. As Figure 3 shown, the to-be-recommended information in the second to-be-recommended information sequence is displayed in sequence through the sorting and display algorithm to obtain the information recommendation result, including:

[0076] S601. Store the to-be-recommended information ranked first in the second to-be-recommended information sequence at the first position of the third to-be-recommended information sequence; where the third to-be-recommended information sequence is preset in advance and is empty in the initial state.

[0077] S602. Determine whether the second to-be-recommended information sequence has been traversed completely, or whether the number of to-be-recommended information in the third to-be-recommended information sequence has reached the preset quantity; if so, use the third to-be-recommended information sequence at the current moment as the information recommendation result; if not, use the next to-be-recommended information in the second to-be-recommended information sequence as the current to-be-recommended information, and execute step 603:

[0078] S603. Sequentially determine whether the number of available single classifications corresponding to the classification of the current to-be-recommended information is greater than or equal to the preset single classification quantity threshold, whether the number of available single channels corresponding to the channel of the current to-be-recommended information is greater than or equal to the preset single channel quantity threshold, and whether the similarity between the keywords of the current to-be-recommended information and the keywords of the last to-be-recommended information in the third to-be-recommended information sequence is greater than the preset similarity threshold. If any of the judgment results is yes, jump to step S602; if all the judgment results are no, increment the number of available single classifications corresponding to the classification of the current to-be-recommended information by 1, increment the number of available single channels corresponding to the channel of the current to-be-recommended information by 1, and store the current to-be-recommended information at the end of the third to-be-recommended information sequence, then jump to step S602.

[0079] For the sorting and display algorithm, various thresholds need to be preset for the information sequence generated by a single recommendation, specifically including: preset quantity, preset similarity threshold, preset single channel quantity threshold, and preset single classification quantity threshold.

[0080] Among them, the preset quantity is the number of recommended results set for the recommendation.

[0081] The preset similarity threshold is the maximum value of the similarity between keywords. If the similarity between the keywords of two pieces of information is greater than the preset similarity threshold, the two pieces of information are considered highly similar.

[0082] The preset single channel quantity threshold is the maximum value of the same channels from which the to-be-recommended information in the second to-be-recommended information sequence comes. If the number of available single channels corresponding to the channel of a certain piece of information is greater than the preset single channel quantity threshold, it is considered that there are enough pieces of information corresponding to this channel to be recommended to the user, and this piece of information does not need to be recommended anymore.

[0083] The preset single classification quantity threshold is the maximum value of the same categories from which the to-be-recommended information in the second to-be-recommended information sequence comes. If the number of available single classifications corresponding to the category of a certain piece of information is greater than the preset single classification quantity threshold, it is considered that there are enough pieces of information corresponding to this category to be recommended to the user, and this piece of information does not need to be recommended anymore.

[0084] As an alternative implementation, the similarity calculation method is to divide the number of overlapping keywords of two pieces of information by the total number of de-duplicated keywords of the two pieces of information. For example, if the keywords of information 1 are "sports, Ronaldo" and the keywords of information 2 are "sports, basketball", then the similarity calculation process is 1÷3≈0.33.

[0085] Figure 4 This is the flowchart of the sorting and display algorithm provided by the embodiments of the present invention. As Figure 4 shown, the specific process of the sorting and display algorithm is as follows:

[0086] 1. Obtain the first piece of information to be recommended in the second sequence of information to be recommended (sequence of information to be recommended), and store it at the head of the third sequence of information to be recommended (sequence of available information).

[0087] 2. Determine whether the second sequence of information to be recommended has been traversed or the number of pieces of information in the third sequence of information to be recommended reaches the preset number. If any of the judgment results is yes, return the sequence of available information and the algorithm ends; if not, obtain the next piece of information in sequence.

[0088] 3. Sequentially judge the following conditions of this piece of information: the number of available single classifications corresponding to the classification of this piece of information ≥ the preset single classification number threshold, the number of available single channels corresponding to the channel of this piece of information ≥ the preset single channel number threshold, and whether the similarity between the keywords of this piece of information and the keywords of the last piece of information in the sequence of available information exceeds the preset similarity threshold. If any judgment is yes, jump to step 2, otherwise continue.

[0089] 4. Increase the number of available single classifications corresponding to the classification of this piece of information by 1, increase the number of available single channels corresponding to the channel of this piece of information by 1, and store this piece of information at the end of the third sequence of information to be recommended, then jump to step 2.

[0090] Restricting the number of classifications and source channels of each recommended information sequence, and restricting the similarity of adjacent information in the generated information sequence reduces the reading fatigue caused by the similarity of content between recommended results, improves the diversity of recommended results, and solves the problem that the content of a single recommended result is too single.

[0091] As an alternative implementation, Figure 5 This is the flowchart of the multi-dimensional user portrait update method provided by the embodiments of the present invention. As Figure 5 shown, after obtaining the information recommendation result, it further includes:

[0092] S7. Update the multi-dimensional user portrait of the current user by using a timed decay update method;

[0093] And / or, S8. Update the multi-dimensional user portrait of the current user by using streaming asynchronous calculation according to the user behavior data and the information recommendation result.

[0094] When updating the multi-dimensional user portrait, only the timed attenuation update method can be used to update the multi-dimensional user portrait. For example, the user portrait is attenuated and updated once at a fixed time every day; or only the streaming asynchronous calculation can be used to implement the update of the multi-dimensional user portrait. For example, after the user generates an action, the streaming asynchronous calculation is used to update the multi-dimensional user portrait; or both the timed attenuation update method and the streaming asynchronous calculation update method can be used. That is, on the basis of updating the multi-dimensional user portrait by the timed attenuation method every day, when the user generates the latest action, the streaming asynchronous calculation is also used to update the multi-dimensional user portrait in real time.

[0095] The timed attenuation update method means that except for the historical behavior portrait, each field value in the user portrait of all users in the library is attenuated to α times the original value every day. In this embodiment, α = 0.926. At the same time, the key-value pairs with field values less than β are deleted. In this embodiment, β = 0.1.

[0096] The following beneficial effects can be achieved through the above timed attenuation update method:

[0097] First, when the user is no longer interested in the information corresponding to a certain type of keyword, that is, when no new user actions are generated for a certain keyword, after 30 days, the weight of the keyword in the user portrait can be attenuated from 1.0 to less than 0.1. By adopting a reasonable threshold in subsequent calculations, effective information can be obtained (for example, those greater than the threshold of 0.1 represent the long-term interests of the user within 30 days, and those greater than the threshold of 0.58 represent the short-term interests of the user within 7 days).

[0098] Second, when the user portrait is updated every day, this method only needs to traverse and update the entire library once, instead of recalculating all the data within a specified date range (such as the most recent 30 days), which greatly reduces the amount of calculation. In addition, during the timed attenuation update process, for the historical behavior portrait, the records of the information shown to the user outside the specified time period need to be cleared. Clearing the records outside the specified time period is because the user's reading memory decays over time, and most users will forget the display records a long time ago. By setting a reasonable time period range, the storage space requirement of the historical behavior portrait can be reduced, and at the same time, the user portrait data communication and calculation time in the single recommendation calculation process can be reduced.

[0099] The user behavior streaming calculation and recommendation result streaming calculation method means that using streaming asynchronous calculation, the user behavior data and recommendation result data to be calculated are written into the calculation queue, and the user portrait update program reads the content of the queue and updates the corresponding user portrait fields according to the user portrait construction method.

[0100] Figure 6 It is the data flow diagram of the recommendation process and portrait update provided by the embodiment of the present invention, as Figure 6As shown by the dashed line in the figure, in this embodiment, the open-source kafka and spark data processing platforms are used as examples to illustrate the implementation process of the streaming computing method. When a specified behavior occurs for a client user, data is uploaded to the backend server. On the server side, kafka is used to collect the data. Additionally, when a recommendation result is generated, kafka is also used to collect the data. That is, kafka collects two types of data, namely user behavior data and recommendation result data. The spark streaming component in spark is used to subscribe to the data in kafka. When new data is generated, spark streaming can obtain the data in real time, and the user portrait construction method in the present invention is used to update the user portrait.

[0101] The following beneficial effects can be achieved through streaming computing:

[0102] First, newly generated behaviors can be quickly applied to the user portrait, making the data in the user portrait more timely and accurate.

[0103] Second, in the recommendation calculation process and the user portrait update process, data is synchronized through a queue, and the asynchronous calculation between the two reduces the influence between processes and improves the stability of the system.

[0104] As an alternative implementation manner, Figure 7 is a flowchart of the first method for generating a sequence of recommended information to be recommended provided by an embodiment of the present invention. As Figure 7 shown, filtering the set of information to be recommended according to the historical behavior portrait and the negative shielding portrait includes:

[0105] S201. Identify the first information ID corresponding to each piece of information to be recommended in the set of information to be recommended, and the second information ID corresponding to the recommended information in the historical behavior portrait. When the first information ID is the same as the second information ID, delete the information to be recommended corresponding to the first information ID from the set of information to be recommended.

[0106] S202. Identify the shielding value corresponding to each feature in the negative shielding portrait, and delete the features corresponding to the shielding values less than the preset shielding value threshold to obtain an effective negative shielding portrait. Among them, the shielding value decays with time.

[0107] S203. In the set of information to be recommended, delete the information to be recommended that conforms to the features of the effective negative shielding portrait.

[0108] Since the historical behavior portrait already covers the historical information that has been recommended to the user, traverse the information to be recommended in the set of information to be recommended. If there is information with the same information ID as the information in the historical behavior portrait, then delete this piece of information from the set of information to be recommended.

[0109] An effective negative shielding image refers to, due to the time decay property of user interest preferences, using the key-value pairs greater than the set threshold in each shielding image by setting the shielding value threshold as the effective negative shielding image. For example, if a user's keyword shielding image is {"Football": 0.2, "Pele": 0.8}, and the shielding threshold is taken as 0.5, then the effective keyword shielding image is "Pele". Check the features of the information in the information pool one by one, and remove the information that meets the content of the effective shielding image from the set of information to be recommended.

[0110] The above method can preliminarily filter invalid data, reduce the system burden, and improve the data processing efficiency.

[0111] The present invention also provides an information recommendation device. Figure 8 It is a schematic structural diagram of the information recommendation device provided by an embodiment of the present invention, as Figure 8 shown. The device includes:

[0112] A data acquisition module 100, configured to acquire in real time a set of information to be recommended within a preset period and a multi-dimensional user profile of the current user; wherein, the multi-dimensional user profile includes a positive preference profile, a negative shielding profile, and a historical behavior profile.

[0113] Construct an information pool to be recommended, that is, a set of information to be recommended, using the information collected within a specified time range. Wherein, the specified time range is a preset period, and this period can be a fixed period set manually or a time range dynamically adjusted according to the size of the information to be recommended uploaded.

[0114] The set of information to be recommended is updated continuously over time. When receiving an information recommendation request, based on the current moment, acquire the set of information to be recommended within the specified time range to ensure the timeliness and data volume of the set of information to be recommended.

[0115] The multi-dimensional user profile is updated continuously over time. When receiving an information recommendation request, acquire the multi-dimensional user profile at the current moment. The recommendation result obtained by analyzing using the multi-dimensional user profile at this moment is more accurate, improving the relevance and compliance of the recommendation result.

[0116] The categories of the multi-dimensional user profile include: a positive preference profile, a negative shielding profile, and a historical behavior profile.

[0117] As an optional implementation manner, the positive preference profile includes: an interest preference profile, a classification preference profile, a channel preference profile, and a collection preference profile.

[0118] As an optional implementation manner, the negative shielding profile includes: a keyword shielding profile, a channel shielding profile, and a disliked interest profile.

[0119] As an alternative implementation, the historical behavior portrait is a record of the information that has been shown to the user within a specified time.

[0120] To construct a multi-dimensional user portrait, key-value pairs of keywords are calculated based on user behavior data or recommendation result data, combined with information features. When calculating, the value of the information feature is defaulted to 1.0. For example, the calculation result of a certain field user portrait is {"football": 1.0, "basketball": 3.0}. The specific calculation methods for each field of the multi-dimensional user portrait are shown in Table 1.

[0121] Among them, user behavior data includes the keywords of a user reading a piece of information, collecting a piece of information, blocking a piece of information, and not liking a piece of information.

[0122] The user behavior weight multiple table can be specified manually. The user behavior weights are specifically reading, collecting, blocking, and not liking.

[0123] As an alternative implementation, the weight value table is shown in Table 2.

[0124] Accumulative calculation means adding the value generated by a new behavior to the existing value of a specific field in the user portrait. For example, if the user's existing interest preference portrait is {"football": 1.0, "basketball": 3.0}, and the user newly generates a collection behavior with the information keywords "football, Beckham", the accumulative result is {"football": 6.0, "basketball": 3.0, "Beckham": 5.0}.

[0125] Information features refer to information ID, information keywords, information classification, and information channels. Among them, the information ID is an identifier that can uniquely identify an information; the information keyword is one or a group of words that can represent the theme content of the information. The information keyword extraction method can be specified manually or generated using a computer program. For example, after segmenting the information content text, the top K words with higher weights are found as the keywords of the information using algorithms such as TF-IDF; the information classification refers to the category of the information, such as sports, entertainment, health preservation, etc. The generation of the information classification can be specified manually or using a computer program. For example, a classification model is trained using classification algorithms such as Naive Bayes, and the model is used to generate the information classification; the information channel refers to the source of the information, such as Xinhua News Agency, People's Daily, etc.

[0126] Establishing a multi-dimensional user portrait enriches the feature space for user analysis, improves the relevance and richness of information recommendation results, and avoids the problem of single information recommendation results.

[0127] The first to-be-recommended information sequence generation module 200 is used to filter the to-be-recommended information set according to the historical behavior portrait and the negative shielding portrait to obtain the first to-be-recommended information sequence;

[0128] After obtaining the set of information to be recommended, check each piece of information to be recommended in the set of information to be recommended one by one, and delete the information to be recommended that conforms to the content of the historical behavior portrait and the content of the negative shielding portrait, so as to obtain the first sequence of information to be recommended.

[0129] These deleted pieces of information to be recommended may be the information that has been recommended to the user within a specified time, or may be the information that the user subjectively shields. The deletion of these pieces of information can reduce the duplication of information recommended to the user within a specified time and improve the compliance of the recommended information with respect to the user.

[0130] The feature sequence generation module 300 is used to generate the feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended one by one according to the positive preference portrait;

[0131] Using the features of the positive preference portrait and the features of each piece of information to be recommended in the first sequence of information to be recommended, generate the feature sequences of all pieces of information to be recommended in the first sequence of information to be recommended.

[0132] The estimated recommendation value generation module 400 is used to input the feature sequence into the pre-trained model to obtain the estimated recommendation value corresponding to each piece of information to be recommended in the first sequence of information to be recommended;

[0133] The pre-trained model refers to a probability estimation model generated by a classification algorithm using the user portrait data, historical information data, and the record of whether the user has read a certain piece of information composed of the user's historical behavior data. Using this model, the probability value of the user reading a certain piece of information can be obtained. As an alternative implementation, the classification algorithm is Naive Bayes, Logistic Regression, Decision Tree, etc.

[0134] Input the feature sequence corresponding to each piece of information to be recommended in the first sequence of information to be recommended into the pre-trained model, and the estimated recommendation value corresponding to each piece of information to be recommended can be obtained. Among them, the estimated recommendation value represents the degree of compliance of this piece of information with the user's needs. The higher the estimated recommendation value of the information to be recommended, the more it conforms to the user's needs. On the contrary, the lower the estimated recommendation value of the information to be recommended, the less it meets the user's needs.

[0135] The second sequence of information to be recommended generation module 500 is used to sort the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated recommendation value to obtain the second sequence of information to be recommended;

[0136] Sort the information to be recommended in the first sequence of information to be recommended from high to low according to the corresponding estimated value to obtain the sorted second sequence of information to be recommended. Among the second sequence of information to be recommended, the information to be recommended at the head is more in line with the user's expectations than the information to be recommended at the tail.

[0137] An information recommendation result generation module 600 is configured to display the to-be-recommended information in the second to-be-recommended information sequence in order through a sorting and display algorithm to obtain an information recommendation result.

[0138] Display the to-be-recommended information in the second to-be-recommended information sequence through a sorting and display algorithm to obtain the final information recommendation result.

[0139] As an optional implementation manner, Figure 9 is a schematic structural diagram of a feature sequence generation module provided by an embodiment of the present invention. As shown in Figure 9 shown, the feature sequence generation module 300 includes:

[0140] A first-order feature sequence generation sub-module 3001 is configured to select each field value of the positive preference profile or the feature value of the to-be-recommended information in the first to-be-recommended information sequence as the first-order feature sequence;

[0141] The first-order feature sequence refers to a sequence composed of each field value of the direct user's positive preference profile or the feature value of the to-be-recommended information.

[0142] As an optional implementation manner, the specific process of generating the first-order feature sequence is as follows: Use the interest preference profile to generate key-value pairs above a specified threshold (for example, the generation result is: "F11_football": 6.0, "F11_basketball": 3.0, "F11_David Beckham": 5.0, where F11 refers to the feature index name), use the classification preference profile to generate key-value pairs above the specified threshold, use the channel preference profile to generate key-value pairs above the specified threshold, use the collection preference profile to generate key-value pairs above the specified threshold, use the information keywords and assign a value of 1.0 to each keyword, use the information classification and assign a value of 1.0 to this classification, and use the information channel and assign a value of 1.0 to this channel.

[0143] A second-order feature sequence generation sub-module 3002 is configured to establish a second-order feature sequence according to the matching relationship between each field value of the positive preference profile and the feature value of the to-be-recommended information in the first to-be-recommended information sequence;

[0144] The second-order feature sequence refers to a feature sequence formed by using the matching relationship between each field value of the user's positive preference profile and the feature value of the to-be-recommended information.

[0145] As an alternative implementation, the second-order feature sequence generation process is as follows: whether there is an overlap between the keywords above the specified threshold in the user interest preference profile and the information keywords (for example, the generation result is: "F21": 1, where F21 refers to the feature index name, and its value 1 indicates an overlap between the two), whether there is an overlap between the categories above the specified threshold in the user classification preference profile and the information category, whether there is an overlap between the channels above the specified threshold in the user channel preference profile and the information channel, and whether there is an overlap between the keywords above the specified threshold in the user favorite preference profile and the information keywords.

[0146] The feature splicing sub-module 3003 is used to splice the first-order feature sequence and the second-order feature sequence to obtain the feature sequence corresponding to each piece of information to be recommended.

[0147] The first-order feature sequence and the second-order feature sequence are spliced end to end to form a feature sequence. Since the feature sequence includes the first-order feature sequence and the second-order feature sequence, and the first-order feature sequence and the second-order feature sequence are both constructed by different methods, the feature space of the spliced feature sequence is rich, which improves the accuracy of the estimated recommendation value obtained by inputting it into the pre-trained model.

[0148] As an alternative implementation, Figure 10 is a schematic structural diagram of the information recommendation result generation module provided by the embodiments of the present invention. As Figure 10 shown, the information recommendation result generation module 600 includes:

[0149] The first-information storage sub-module 6001 is used to store the piece of information to be recommended ranked first in the second piece of information to be recommended sequence at the first position of the third piece of information to be recommended sequence; wherein, the third piece of information to be recommended sequence is preset and is empty in the initial state;

[0150] The traversal sub-module 6002 is used to determine whether the traversal of the second sequence of recommended information is completed, or whether the number of pieces of recommended information in the third sequence of recommended information reaches a preset quantity; if so, the third sequence of recommended information at the current moment is used as the information recommendation result; if not, the next piece of recommended information in the second sequence of recommended information is used as the current piece of recommended information, and the following steps are executed: successively determine whether the number of available single classifications corresponding to the classification of the current piece of recommended information is greater than or equal to the preset single classification quantity threshold, whether the number of available single channels corresponding to the channel of the current piece of recommended information is greater than or equal to the preset single channel quantity threshold, and whether the similarity between the keywords of the current piece of recommended information and the keywords of the last piece of recommended information in the third sequence of recommended information is greater than the preset similarity threshold. If any of the judgment results is yes, jump to the above judgment step; if all the judgment results are no, increase the number of available single classifications corresponding to the classification of the current piece of recommended information by 1, increase the number of available single channels corresponding to the channel of the current piece of recommended information by 1, and store the current piece of recommended information at the end of the third sequence of recommended information, and then jump to the above judgment step.

[0151] The sorting and display algorithm needs to preset various thresholds for the information sequence generated by a single recommendation, specifically including: preset quantity, preset similarity threshold, preset single channel quantity threshold, and preset single classification quantity threshold.

[0152] Among them, the preset quantity is the number of recommended result pieces set for the recommendation;

[0153] The preset similarity threshold is the maximum value of the similarity between keywords. If the similarity between the keywords of two pieces of information is greater than the preset similarity threshold, it is considered that the two pieces of information are highly similar;

[0154] The preset single channel quantity threshold is the maximum value of the same channels from which the pieces of recommended information in the second sequence of recommended information come. If the number of available single channels corresponding to the channel of a certain piece of information is greater than the preset single channel quantity threshold, it is considered that there are enough pieces of information corresponding to this channel to be recommended to the user, and this piece of information does not need to be recommended anymore.

[0155] The preset single classification quantity threshold is the maximum value of the same categories from which the pieces of recommended information in the second sequence of recommended information come. If the number of available single classifications corresponding to the category of a certain piece of information is greater than the preset single classification quantity threshold, it is considered that there are enough pieces of information corresponding to this category to be recommended to the user, and this piece of information does not need to be recommended anymore.

[0156] As an optional implementation manner, the similarity calculation method is to divide the number of overlapping keywords of two pieces of information by the total number of non-repeated keywords of the two pieces of information. For example, if the keywords of information 1 are "sports, Ronaldo" and the keywords of information 2 are "sports, basketball", then the similarity calculation process is 1÷3≈0.33.

[0157] Such asFigure 4 As shown in the figure, the specific process of the sorting and display algorithm is as follows:

[0158] 1. Obtain the first piece of information to be recommended in the second information sequence to be recommended (information sequence to be recommended), and store it at the first position of the third information sequence to be recommended (recommendable information sequence).

[0159] 2. Determine whether the second information sequence to be recommended has been traversed completely or the number of information in the third information sequence to be recommended reaches the preset quantity. If any of the judgment results is yes, return the recommendable sequence and the algorithm ends; if not, obtain the next piece of information in sequence.

[0160] 3. Judge the following conditions of this piece of information in turn: the number of available single classifications corresponding to the classification of this piece of information ≥ the preset single classification quantity threshold, the number of available single channels corresponding to the channel of this piece of information ≥ the preset single channel quantity threshold, and whether the keyword similarity between the keyword of this piece of information and the keyword of the last piece of information in the available information sequence exceeds the preset similarity threshold. If any judgment is yes, jump to step 2; otherwise, continue.

[0161] 4. Add 1 to the number of available single classifications corresponding to the classification of this piece of information, add 1 to the number of available single channels corresponding to the channel of this piece of information, and store this piece of information at the end of the third information sequence to be recommended, then jump to step 2.

[0162] Restricting the number of classifications and source channels of each recommended information sequence, and restricting the similarity of adjacent information in the generated information sequence reduces the reading fatigue caused by the similarity of content between recommended results, improves the diversity of recommended results, and solves the problem that the information content of a single recommended result is too single.

[0163] As an optional implementation manner, Figure 11 is a schematic structural diagram of the first update module and the second update module provided by an embodiment of the present invention. As Figure 11 shown, the device further includes:

[0164] The first update module 700 is used to update the multi-dimensional user portrait of the current user by using a timed attenuation update method;

[0165] The second update module 800 is used to update the multi-dimensional user portrait of the current user by using streaming asynchronous calculation according to user behavior data and information recommendation results.

[0166] When updating the multi-dimensional user portrait, only the timed decay update method can be used to update the multi-dimensional user portrait. For example, the user portrait is decayed and updated once at a fixed time every day; or only the streaming asynchronous calculation can be used to implement the update of the multi-dimensional user portrait. For example, when the user generates an action, the streaming asynchronous calculation is used to update the multi-dimensional user portrait; or both the timed decay update method and the streaming asynchronous calculation update method can be used. That is, on the basis of updating the multi-dimensional user portrait by the timed decay method every day, when the user generates the latest action, the streaming asynchronous calculation is also used to update the multi-dimensional user portrait in real time.

[0167] The timed decay update method means that except for the historical behavior portrait, the field values in the user portraits of all users in the database are decayed to α times the original value every day. In this embodiment, α = 0.926. At the same time, the key-value pairs with field values less than β are deleted. In this embodiment, β = 0.1.

[0168] The following beneficial effects can be achieved through the above timed decay update method:

[0169] First, when the user is no longer interested in the information corresponding to a certain type of keyword, that is, when no new user actions are generated for a certain keyword, after 30 days of the user portrait, the weight of this keyword can be decayed from 1.0 to less than 0.1. By adopting a reasonable threshold in subsequent calculations, effective information can be obtained (for example, those greater than the threshold of 0.1 represent the long-term interests of the user within 30 days, and those greater than the threshold of 0.58 represent the short-term interests of the user within 7 days);

[0170] Second, when the user portrait is updated every day, this method only needs to traverse and update the entire database once, instead of recalculating all the data within a specified date range (such as the most recent 30 days) once, which greatly reduces the calculation amount. In addition, during the timed decay update process, for the historical behavior portrait, the records of the information that has been shown to the user outside the specified time period need to be cleared. Clearing the records outside the specified time period is because the user's reading memory decays over time, and most users will forget the display records a long time ago. By setting a reasonable time period range, the storage space requirement of the historical behavior portrait can be reduced, and at the same time, the user portrait data communication and calculation time in the single recommendation calculation process can be reduced.

[0171] The user behavior streaming calculation and recommendation result streaming calculation method means that using the streaming asynchronous calculation, the user behavior data and recommendation result data to be calculated are written into the queue to be calculated, and the user portrait update program reads the content of this queue and updates the corresponding user portrait fields according to the user portrait construction method.

[0172] Such as Figure 6As shown by the dashed line in the figure, in this embodiment, the open-source Kafka and Spark data processing platforms are used as examples to illustrate the implementation process of the streaming computing method. When a specified behavior occurs for a client user, data is uploaded to the backend server. On the server side, Kafka is used to collect the data. Additionally, when a recommendation result is generated, Kafka is also used to collect the data. That is, Kafka collects two types of data, namely user behavior data and recommendation result data. The Spark Streaming component in Spark is used to subscribe to the data in Kafka. When new data is generated, Spark Streaming can obtain the data in real time and use the user profile construction method in the present invention to update the user profile.

[0173] The following beneficial effects can be achieved through streaming computing:

[0174] First, newly generated behaviors can be quickly applied to the user profile, making the data in the user profile more timely and accurate.

[0175] Second, in the recommendation calculation process and the user profile update process, data is synchronized through a queue, and the asynchronous calculation between the two reduces the impact between processes and improves the stability of the system.

[0176] As an alternative implementation, the first to-be-recommended information sequence generation module 200 includes:

[0177] A first filtering sub-module 2001, which is used to identify the first information ID corresponding to each to-be-recommended information in the to-be-recommended information set and the second information ID corresponding to the recommended information in the historical behavior profile. When the first information ID is the same as the second information ID, the to-be-recommended information corresponding to the first information ID is deleted from the to-be-recommended information set;

[0178] A shielding value identification sub-module 2002, which is used to identify the shielding value corresponding to each feature in the negative shielding profile, delete the features corresponding to the shielding values less than the preset shielding value threshold, and obtain an effective negative shielding profile; wherein, the shielding value decays with time;

[0179] A second filtering sub-module 2003, which is used to delete the to-be-recommended information in the to-be-recommended information set that conforms to the features of the effective negative shielding profile.

[0180] Since the historical behavior profile already covers the historical information that has been recommended to the user, when traversing the to-be-recommended information in the to-be-recommended information set, if there is information with the same information ID as the information in the historical behavior profile, then this piece of information is deleted from the to-be-recommended information set.

[0181] An effective negative shielding image refers to an image that, due to the time-decaying nature of user interest preferences, uses key-value pairs greater than a set threshold in each shielding image as an effective negative shielding image by setting a shielding value threshold. For example, if a user's keyword shielding image is {"football": 0.2, "Pelé": 0.8} and the shielding threshold is set to 0.5, the effective keyword shielding image is "Pelé". Check the characteristics of the information in the information pool one by one, and remove the information that meets the content of the effective shielding image from the set of information to be recommended.

[0182] The above method can initially filter out invalid data, reduce the system burden, and improve data processing efficiency.

[0183] An embodiment of the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it implements the above-mentioned information recommendation method.

[0184] The above software is stored in the above storage medium, and the storage medium includes but is not limited to: optical discs, floppy disks, hard disks, rewritable memories, etc.

[0185] The above technical solution has the following beneficial effects: By constructing multiple dimensions of positive and negative user portraits, it solves the problem of incomplete interest points in the prior art and improves the compliance of the generated recommendation results with user expectations; Through time decay and real-time streaming computing, it solves the problems of users' interests changing over time, untimely updates of users' interests, and large computational amounts for updating user portraits within a specified time period, and improves the real-time nature of obtaining recommendation results; Through user portraits, information pools, and pre-trained models, the quality of recommendation results is improved; Through the display sorting algorithm, it solves the problem of over-concentration of the content of recommendation results and improves the diversity of recommendation results.

[0186] The recommendation method of the embodiment of the present invention can be applied to products for information flow recommendation. When a user needs to display an information flow, the method of the present invention can generate a specified number of information recommendation results. When the user performs an action on the recommendation results (such as reading, collecting, shielding), the user portrait can be updated in real time through streaming computing. When the user finishes browsing the current display results and needs more displays, for example, guiding the user to scroll down on the client side, at this time, calling the method of the present invention can generate the next batch of specified number of information recommendation results, and repeating such operations can realize the reading product of the information flow.

[0187] The above specific implementation manners of the invention further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above content is only the specific implementation manners of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An information recommendation method, characterized in that, Including: S1. Obtain the set of to-be-recommended information within a preset period and the multi-dimensional user profile of the current user in real time; wherein, the multi-dimensional user profile includes a positive preference profile, a negative shielding profile, and a historical behavior profile; S2. Filter the set of to-be-recommended information according to the historical behavior profile and the negative shielding profile to obtain a first to-be-recommended information sequence; S3. Generate a feature sequence corresponding to each piece of to-be-recommended information in the first to-be-recommended information sequence one by one according to the positive preference profile; S4. Input the feature sequence into a pre-trained model to obtain an estimated recommendation value corresponding to each piece of to-be-recommended information in the first to-be-recommended information sequence; S5. Sort the to-be-recommended information in the first to-be-recommended sequence from high to low according to the corresponding estimated recommendation value to obtain a second to-be-recommended information sequence; S6. Display the to-be-recommended information in the second to-be-recommended information sequence in order through a sorting display algorithm to obtain an information recommendation result; The step of displaying the to-be-recommended information in the second to-be-recommended information sequence in order through a sorting display algorithm to obtain an information recommendation result includes: S601. Store the to-be-recommended information ranked first in the second to-be-recommended information sequence at the first position of a third to-be-recommended information sequence; wherein, the third to-be-recommended information sequence is preset and is empty in the initial state; S602. Judge whether the second to-be-recommended information sequence has been traversed, or whether the number of to-be-recommended information in the third to-be-recommended information sequence has reached a preset number; If so, use the third to-be-recommended information sequence at the current moment as the information recommendation result; If not, use the next to-be-recommended information in the second to-be-recommended information sequence as the current to-be-recommended information, and execute step S603: S603. Judge in turn whether the number of available single classifications corresponding to the classification of the current to-be-recommended information is greater than or equal to a preset single classification number threshold, whether the number of available single channels corresponding to the channel of the current to-be-recommended information is greater than or equal to a preset single channel number threshold, and whether the similarity between the keyword of the current to-be-recommended information and the keyword of the last to-be-recommended information in the third to-be-recommended information sequence is greater than a preset similarity threshold. If any judgment result is yes, jump to step S602; if all judgment results are no, add 1 to the number of available single classifications corresponding to the classification of the current to-be-recommended information, add 1 to the number of available single channels corresponding to the channel of the current to-be-recommended information, and store the current to-be-recommended information at the end of the third to-be-recommended information sequence, then jump to step S602.

2. The information recommendation method according to claim 1, wherein Generating a feature sequence corresponding to each piece of to-be-recommended information in the first to-be-recommended information sequence one by one according to the positive preference profile includes: S301. Select the field values of the positive preference profile or the feature values of the to-be-recommended information in the first to-be-recommended information sequence as the first-order feature sequence; S302. Establish a second-order feature sequence according to the matching relationship between the field values of the positive preference profile and the feature values of the to-be-recommended information in the first to-be-recommended information sequence; S303. Concatenate the first-order feature sequence and the second-order feature sequence to obtain the feature sequence corresponding to each piece of information to be recommended.

3. The information recommendation method according to claim 1, wherein After obtaining the information recommendation result, it further includes: S7. Update the multi-dimensional user profile of the current user by using a timed decay update method; And / or, S8. Update the multi-dimensional user profile of the current user by using streaming asynchronous computing according to user behavior data and information recommendation results.

4. The information recommendation method according to claim 1, wherein The filtering of the information set to be recommended according to the historical behavior profile and the negative shielding profile includes: S201. Identify the first information ID corresponding to each piece of information to be recommended in the information set to be recommended, and the second information ID corresponding to the recommended information in the historical behavior profile. When the first information ID is the same as the second information ID, delete the information to be recommended corresponding to the first information ID from the information set to be recommended; S202. Identify the shielding value corresponding to each feature in the negative shielding profile, and delete the features corresponding to the shielding values less than the preset shielding value threshold to obtain an effective negative shielding profile; wherein, the shielding value decays with time; S203. In the information set to be recommended, delete the information to be recommended that conforms to the features of the effective negative shielding profile.

5. An information recommendation device, characterized in that, It includes: A data acquisition module, configured to acquire in real time the information set to be recommended within a preset period and the multi-dimensional user profile of the current user; wherein, the multi-dimensional user profile includes a positive preference profile, a negative shielding profile, and a historical behavior profile; A first information sequence generation module to be recommended, configured to filter the information set to be recommended according to the historical behavior profile and the negative shielding profile to obtain a first information sequence to be recommended; A feature sequence generation module, configured to generate the feature sequence corresponding to each piece of information to be recommended in the first information sequence to be recommended one by one according to the positive preference profile; An estimated recommendation value generation module, configured to input the feature sequence into a pre-trained model to obtain the estimated recommendation value corresponding to each piece of information to be recommended in the first information sequence to be recommended; A second information sequence generation module to be recommended, configured to sort the information to be recommended in the first information sequence to be recommended from high to low according to the corresponding estimated recommendation value to obtain a second information sequence to be recommended; An information recommendation result generation module, configured to display the information to be recommended in the second information sequence to be recommended through a sorting display algorithm in order to obtain an information recommendation result; The information recommendation result generation module includes: A first information storage sub-module to be recommended, configured to store the information to be recommended ranked first in the second information sequence to be recommended at the first position of a third information sequence to be recommended; wherein, the third information sequence to be recommended is preset and is empty in the initial state; A traversal sub-module is used to determine whether the traversal of the second sequence of to-be-recommended information is completed, or whether the number of to-be-recommended information in the third sequence of to-be-recommended information reaches a preset number; if so, the third sequence of to-be-recommended information at the current moment is used as the information recommendation result; if not, the next to-be-recommended information in the second sequence of to-be-recommended information is used as the current to-be-recommended information, and the following steps are executed: successively determine whether the number of available single categories corresponding to the category of the current to-be-recommended information is greater than or equal to a preset single-category number threshold, whether the number of available single channels corresponding to the channel of the current to-be-recommended information is greater than or equal to a preset single-channel number threshold, and whether the similarity between the keyword of the current to-be-recommended information and the keyword of the last to-be-recommended information in the third sequence of to-be-recommended information is greater than a preset similarity threshold. If any of the judgment results is yes, jump to the above judgment step; if all judgment results are no, increase the number of available single categories corresponding to the category of the current to-be-recommended information by 1, increase the number of available single channels corresponding to the channel of the current to-be-recommended information by 1, and store the current to-be-recommended information at the end of the third sequence of to-be-recommended information, and then jump to the above judgment step.

6. The information recommendation device according to claim 5, characterized in that, The feature sequence generation module includes: A first-order feature sequence generation sub-module, which is used to select each field value of the positive preference portrait or the feature value of the to-be-recommended information in the first sequence of to-be-recommended information as the first-order feature sequence; A second-order feature sequence generation sub-module, which is used to establish a second-order feature sequence according to the matching relationship between each field value of the positive preference portrait and the feature value of the to-be-recommended information in the first sequence of to-be-recommended information; A feature splicing sub-module, which is used to splice the first-order feature sequence and the second-order feature sequence to obtain the feature sequence corresponding to each to-be-recommended information.

7. The information recommendation system according to claim 5, wherein It also includes: A first update module, which is used to update the multi-dimensional user portrait of the current user by using a timed decay update method; A second update module, which is used to update the multi-dimensional user portrait of the current user by using streaming asynchronous calculation according to user behavior data and information recommendation results.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the information recommendation method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Internet intelligence crawling and recommendation system for field of investment attraction

    CN106960063A

  • Growth incentive book recommendation method and recommendation system

    CN110990670A