Information recommendation method and related equipment
By analyzing the relevance and semantic features of users' historical browsing data, fine-grained primary and secondary interest features are determined, solving the problem of inaccurate information recommendation in existing technologies and achieving more accurate information recommendation results.
Patent Information
- Application Number
- CN202511052334.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to accurately identify user interests, leading to inaccurate information recommendations.
By analyzing the relevance and semantic features of users' historical browsing data, at least two types of user interest features are determined, including fine-grained primary interest features and fine-grained secondary interest features, which respectively represent the degree of user interest in different data items, and information recommendation is made based on these features.
It enables more refined identification of user interests, improves the accuracy and personalization of information recommendations, and meets the actual needs of users.
Smart Images

Figure CN120974004A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information recommendation, and in particular to an information recommendation method and device, equipment, a computer readable storage medium and a computer program product. BACKGROUND
[0002] Information recommendation refers to automatically pushing information content that a user may be interested in to the user according to the user's interests and / or historical behaviors and other factors. Information recommendation technology has been widely applied in various fields such as e-commerce platforms, social media, news websites, and government websites, greatly improving user experience and information acquisition efficiency.
[0003] At present, how to accurately identify user interests and then implement accurate information recommendation based on user interests is a technical problem that needs to be continuously solved in this technical field. SUMMARY
[0004] Embodiments of the present application provide an information recommendation method to solve the problem of how to accurately identify user interests and then implement accurate information recommendation in the prior art.
[0005] Embodiments of the present application also provide an information recommendation device, equipment, a computer readable storage medium and a computer program product.
[0006] Embodiments of the present application adopt the following technical solutions:
[0007] An information recommendation method includes: determining at least two first interest features of a user according to the relevance of different data entries historically browsed by the user in a same browsing sequence of the user and the semantic features of the different data entries; the at least two first interest features respectively represent the interest degrees of the user for the data entries with different relevance in the browsing sequence; and selecting a data entry from data entries for recommendation to recommend to the user according to the at least two first interest features.
[0008] An information recommendation device includes: an interest determination unit configured to determine at least two first interest features of a user according to the relevance of different data entries historically browsed by the user in a same browsing sequence of the user and the semantic features of the different data entries; the at least two first interest features respectively represent the interest degrees of the user for the data entries with different relevance in the browsing sequence; and an information recommendation unit configured to select a data entry from data entries for recommendation to recommend to the user according to the at least two first interest features.
[0009] A computing device includes: a memory and a processor, wherein the memory is used to store a computer program; and the processor is coupled to the memory and used to execute the computer program stored in the memory to perform the method described above.
[0010] A computer-readable storage medium storing a computer program that, when executed by a computer, enables the implementation of the above-described method.
[0011] A computer program product storing instructions that, when executed by a computer, cause the computer to perform the method described above.
[0012] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:
[0013] Because this solution can identify at least two primary interest features of a user, and these features respectively characterize the user's degree of interest in data items with varying degrees of relevance within the browsing sequence, this solution achieves the separation of "the user's degree of interest in the same browsing sequence," effectively refining the identification of user interests. Therefore, compared to existing technologies that recommend information based on a single user interest level, this solution, based on at least two primary interest features, can achieve more refined information recommendations, making the recommended information more accurate and better suited to the user's actual needs. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0015] Figure 1 A flowchart illustrating the specific implementation of an information recommendation method provided in this application embodiment;
[0016] Figure 2a This is a schematic diagram illustrating the application process of the coarse-fine ranking and recall recommendation system framework in this application embodiment in a real-world application scenario.
[0017] Figure 2b A schematic diagram illustrating the specific implementation process of the information recommendation method provided in this application embodiment applied to a government website;
[0018] Figure 2c This is a schematic diagram of each sub-step included in step 21;
[0019] Figure 2d A schematic diagram illustrating a specific example of the implementation process of step 21;
[0020] Figure 2e This is a schematic diagram of each sub-step included in step 24;
[0021] Figure 3 A schematic diagram of the specific structure of an information recommendation device provided in this application embodiment;
[0022] Figure 4 A schematic diagram of the specific structure of the computing device provided in the embodiments of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] As will be known to those skilled in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0025] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0026] Example 1
[0027] Embodiment 1 of this application provides an information recommendation method to solve the problem of how to accurately identify user interests and then make accurate information recommendations in the prior art.
[0028] The subject executing this method can be any computing device capable of implementing the method, such as a server, mobile phone, personal computer, smart wearable device, smart robot, etc.
[0029] Different steps of this method can be implemented by the same execution entity or by different execution entities. This application does not limit which execution entity is used to implement the method.
[0030] Furthermore, the embodiments of this application do not limit the execution order of different steps. When using the method provided in the embodiments of this application, the execution order of different steps can be adjusted according to actual needs.
[0031] For ease of description, the following uses an application (APP) with information recommendation function as the execution subject of this method as an example to provide a detailed description of the method provided in this application embodiment.
[0032] like Figure 1 The diagram shown is a flowchart illustrating the specific implementation of an information recommendation method provided in this application, including the following steps:
[0033] Step 11: The APP determines at least two primary interest features of the user based on the relevance of different data items viewed in the user's history within the same browsing sequence, as well as the semantic features of different data items;
[0034] The term "user" here refers to any user who browses data entries using the app. For example, it could be user A.
[0035] Data entries refer to various information units that users browse within an app, such as, but not limited to, articles, products, images, videos, etc. In a specific example, taking a government affairs app as an example, data entries could include, but are not limited to, policies and regulations, announcements and notices, public service guides, service procedures, and public information. A government affairs app is a mobile application that provides services to the government or government agencies, allowing users to browse or handle various government affairs.
[0036] A browsing sequence refers to a collection of data items viewed by a user within an app, arranged sequentially according to the user's browsing time. For example, if user A browses article 1, product 2, image 3, and video 4 sequentially between 9:00 AM and 10:00 AM, these data items would form a browsing sequence. In other words, the "browsing sequence" described in this embodiment is a time series.
[0037] Generally, the browsing sequence can correspond to a historical time period. For example, "9 o'clock to 10 o'clock" in the example above is a historical time period corresponding to the browsing sequence.
[0038] The relevance of different data items within a user's browsing sequence refers to the degree of mutual association between different data items within that same browsing sequence. This relevance can be determined based on one or more factors, such as the content, topic, category, tags, user behavior (e.g., clicks, dwell time, sharing), and the time interval between data items. For example, if a user has currently browsed two articles about healthy eating consecutively, then these two articles have a high degree of relevance within the user's current browsing sequence.
[0039] In one alternative implementation, the autocorrelation coefficients of data entries in the browsing sequence can be calculated. The calculated autocorrelation coefficients of different data entries can characterize the degree of correlation between different data entries in the same browsing sequence of the user.
[0040] In a specific example, the autocorrelation coefficient can be calculated as follows:
[0041] Assume the browsing sequence is S u ,D1,...,D s Let S represent the topic of the data items viewed by the user—for example, the title of an article about healthy eating. u Noted as: S u ={D1,...,D s For browsing sequence S u Each topic in the dataset can have its autocorrelation coefficient calculated. For example, with D... s For example, in one alternative implementation, the autocorrelation coefficient η can be calculated using the following formula [1]. S :
[0042] η S =σ(E S ·E u ′) [1]
[0043] The explanation of formula [1] is as follows:
[0044] It is the sigmoid function;
[0045] E S D s The embedding vector is used to characterize D. s The semantics of;
[0046] Among them, E u ={E1,E2,...,E S} represents D1,...,D sThe set of embedding vectors formed by their respective embedding vectors; S represents the browsing sequence S u The total number of topics included.
[0047] By E u As can be seen from the calculation method of ′, it is equivalent to calculating E. u Except for E S The mean of the external embedding vectors—the purpose is to construct the browsing sequence S. u Vector representation of E; S ·E u ′ is the value used to calculate the mean and E. S The dot product—its result is the product of the cosine of the angle between two vectors and the product of the lengths of the two vectors. The significance of the dot product lies in its ability to calculate the projection between two vectors; σ(E S ·E u ′) is the value of the sigmoid function used to calculate the result of the dot product. The sigmoid function can map any real number to the range (0,1), and is therefore particularly suitable for probability estimation.
[0048] In summary, it can be seen that D s The autocorrelation coefficient η S It can be used to represent D s With browsing sequence S u The degree of correlation with other topics. Furthermore, if the autocorrelation coefficient is calculated according to formula [1], the larger the coefficient, the higher the degree of correlation; conversely, the lower the coefficient, the lower the correlation.
[0049] In one alternative implementation, for topic D s Regarding the embedding vector, topic D can be obtained in the following way. s Embedding vector:
[0050] First, each topic is converted into its corresponding text description. Then, using a pre-trained word embedding model, such as Word2Vec, BERT, or GPT, each word in the text description is converted into a word embedding vector. Next, these word embedding vectors are weighted and averaged or weighted and summed using an attention mechanism to obtain the embedding vector for the entire topic. Generally, embedding vectors can not only capture the latent semantic features of words, but also reflect the semantic relationships and contextual information between words.
[0051] The semantic features of the data entries mentioned in step 11 refer to the vector representation obtained after processing the data entries. This vector can capture the latent semantic features of the data entries. Specifically, when the data entries are in text format, they can be encoded by inputting them into a pre-trained semantic model to obtain the corresponding semantic feature vectors. These semantic feature vectors not only reflect the literal meaning of the text information but also contain deeper semantic features such as contextual information and semantic relationships.
[0052] In an alternative implementation, the semantic features of a data entry can also be characterized by an embedding vector.
[0053] In this embodiment, at least two primary interest features of the user can be determined based on "the degree of relevance of different data entries browsed in the user's historical browsing sequence" and "the semantic features of different data entries". The at least two primary interest features of the user respectively characterize the degree of interest the user has in data entries with different degrees of relevance in the browsing sequence.
[0054] In one alternative implementation, at least two primary interest features may be, for example, fine-grained primary interest features and fine-grained secondary interest features.
[0055] Among them, fine-grained primary interest features characterize the degree of user interest in data items in the primary interest browsing sequence, which is split from the same browsing sequence; fine-grained secondary interest features characterize the degree of user interest in data items in the secondary interest browsing sequence, which is split from the same browsing sequence.
[0056] The relevance of data entries in the primary interest browsing sequence to the same browsing sequence is greater than the relevance of data entries in the secondary interest browsing sequence to the same browsing sequence.
[0057] In one alternative implementation, the user's fine-grained primary interest features and fine-grained secondary interest features can be determined by performing the following sub-steps:
[0058] Sub-step 1: Determine the autocorrelation coefficients of different data entries in the browsing sequence;
[0059] For details on how to determine the autocorrelation coefficient of a data item, please refer to the previous text; it will not be repeated here.
[0060] Sub-step 2: Based on the autocorrelation coefficient, split the browsing sequence into a primary interest browsing sequence and a secondary interest browsing sequence;
[0061] Following the previous example, assume that by executing sub-step 1, the browsing sequence S is calculated. u ={D1,...,D sThe autocorrelation coefficient of each data entry in the browsing sequence is denoted as η. i The value of i ranges from [1, S]. Based on this, the specific implementation of sub-step 2 may include:
[0062] According to η i In ascending order, for the corresponding data entries D1,...,D s Sort;
[0063] For the sorted data entries D1,...,D s Select the top-ranked (η) groups according to a preset ratio (e.g., 70%). i The system selects a number of data items (of relatively large size) that are ranked high and whose number in the browsing sequence equals a preset proportion, thus forming the user's main interest sequence. Where u is an identifier used to represent the user;
[0064] Construct the user's secondary interest sequence based on the remaining data items after selection (e.g., the remaining 30%).
[0065] As mentioned before, for example, D s The autocorrelation coefficient η S It can be used to represent D s With browsing sequence S u The degree of correlation with other topics. Furthermore, if the autocorrelation coefficient is calculated according to formula [1], the larger the coefficient, the higher the degree of correlation; conversely, the lower the coefficient. Therefore, it can be inferred that: relative to η S The sequence of secondary interests formed by the lower-ranked data items In terms of main interest sequence The data entries in the sequence are relatively more correlated with other data entries in the browsing sequence. In other words, compared to the secondary interest sequence... In terms of main interest sequence The data items in the sequence are those that users are more interested in and are likely to continue browsing; while those relative to the main interest sequence... In terms of secondary interest sequences The data entries in the middle are those that users are less interested in (but still somewhat interested in), and are likely data entries that they browsed casually.
[0066] Sub-step 3: Determine the user's fine-grained main interest features based on the browsing frequency of data items in the browsing sequence and the semantic features of data items in the main interest browsing sequence;
[0067] The browsing frequency of a data entry refers to the number of times a data entry is viewed within a specified time period. This specified time period can be a recent unit of time, such as within the last hour or the last 24 hours; alternatively, it can be a historical time period corresponding to a browsing sequence. In an optional implementation, the browsing frequency of a data entry can be obtained by the app from its corresponding server. In a specific example, the app can send a "frequency retrieval request containing the unique identifier of the data entry" to the server, requesting the server to provide the app with the number of times the data entry with that unique identifier has been viewed within the specified time period.
[0068] In one alternative implementation, the number of times a data entry is viewed may refer to the total number of times a data entry is viewed by any user.
[0069] Continuing with the previous example, assume that the main interest sequence has been obtained by performing sub-step 2. In a specific example, sub-step 3 can be implemented as follows:
[0070] Obtain the main interest sequence The embedding vectors of each data entry constitute the main interest data embedding vector set. This is essentially an implementation of obtaining the semantic features of data entries in the main interest browsing sequence described in sub-step 3;
[0071] Get the frequency f of each data item in the browsing sequence. i The value of i ranges from [1, S];
[0072] According to f i Calculate the average frequency of browsing data entries in the browsing sequence.
[0073] according to and The fine-grained main interest embedding vector representation of the user is calculated according to the following formula [2] .
[0074]
[0075] In formula [2], E j express The j-th embedding vector, where j takes values in the range [1, p], and p is... The maximum number of embedded vectors contained therein; Indicates calculation The total number of embedded vectors contained therein.
[0076] As can be seen from formula [2], since Based on the main interest data embedding vector set Calculated, and The corresponding main interest sequence The data items in the data are those that users are more interested in and are likely to continue browsing. It can characterize the degree of interest in the main interest browsing sequence. Based on This function can As a fine-grained primary interest feature of users.
[0077] Sub-step 4: Determine the user's fine-grained secondary interest features based on the browsing frequency of data entries in the secondary interest browsing sequence and the semantic features of the data entries in the secondary interest browsing sequence.
[0078] Continuing with the previous example, assume that the sub-interest sequence has been obtained by performing sub-step 2. In a specific example, sub-step 4 can be implemented as follows:
[0079] Obtaining secondary interest sequences The embedding vectors of each data entry constitute the set of embedding vectors for secondary interest data. This is essentially an implementation of obtaining the semantic features of data entries in the secondary interest browsing sequence described in sub-step 4;
[0080] according to and The fine-grained sub-interest embedding vector representation of the user is calculated according to the following formula [3].
[0081]
[0082] In formula [2], E k express The k-th embedding vector, where k takes values in the range [1, q], and q is... The maximum number of embedded vectors contained therein; Indicates calculation The total number of embedded vectors contained therein.
[0083] As can be seen from formula [3], since Based on secondary interest data embedding vector set Calculated, and The corresponding secondary interest sequence The data items in the data are those that users are more interested in and are likely to continue browsing. This can characterize the degree of interest in a secondary interest browsing sequence. Based on This function can As a fine-grained secondary interest feature of users.
[0084] because and Different, correspondingly, calculated and They are also different, as they represent the different levels of interest a user has in different data items within the browsing sequence.
[0085] Step 12: Based on the user's at least two primary interest features, select data entries from the recommended data entries and recommend them to the user.
[0086] In one alternative implementation, data entries whose semantic features meet a preset similarity requirement (e.g., similarity less than a preset similarity threshold) can be selected from the recommended data entries based on the at least two first interest features and recommended to the user.
[0087] In an optional implementation, when the at least two first interest features include the "fine-grained primary interest features" and "fine-grained secondary interest features" as described above, step 12 may specifically include:
[0088] From the recommended data entries, select the first data entry that meets the preset first similarity requirement with the fine-grained main interest features;
[0089] From the other recommended data entries besides the first data entry, select a second data entry that meets the preset second similarity requirement with the fine-grained secondary interest features;
[0090] The first and second data entries are sorted and then recommended to the user.
[0091] For example, the first and second data items can be sorted by placing the first data item before the second data item, and then a list of recommended information composed of the sorted data items can be generated. The list of recommended information can be presented to the user all at once; or, data items can be retrieved from the list in order of their sorting position and presented to the user in multiple installments.
[0092] In this embodiment, by distinguishing user interest features into "fine-grained primary interest features" and "fine-grained secondary interest features," the primary and secondary interests of users can be separated. This allows for a more precise understanding of user interests, improving the accuracy and personalization of recommendations. For example, in product recommendations on an e-commerce platform, if a user's primary interest is purchasing sports equipment and their secondary interest is focusing on healthy eating, by separating these two types of interest features, the system can prioritize recommending sports equipment that matches the user's primary interest, while also appropriately recommending healthy foods based on the secondary interest. This detailed interest segmentation not only enhances the user experience but also strengthens the practicality and intelligence of the recommendation system.
[0093] Similarly, if user interests are categorized into more layers—such as three or more—it also achieves the separation of different levels of user interest. Similar to the effects described above, this allows for a more precise understanding of user interests, improving the accuracy and personalization of recommendations, thereby enhancing the user experience and increasing the practicality and intelligence of the recommendation system.
[0094] In this embodiment of the application, the "data entries for recommendation" refers to data entries to be recommended to target users for browsing.
[0095] In one optional implementation, these data entries can be stored on the server corresponding to the APP, and the APP can retrieve these data entries from the server when it needs to make recommendations to the user; in another optional implementation, these data entries may also be stored in the storage space that the APP can occupy on the user's terminal, and the APP can retrieve these data entries from the storage space when it needs to make recommendations to the user.
[0096] In one optional implementation, the recommended data items can be pre-determined by the app's operators based on the app's business type and intended for recommendation to users. In another optional implementation, the recommended data items can be data items that the app automatically crawls from the internet based on the user's browsing history, search history, and other information, which may be of interest to the user.
[0097] In one alternative implementation, the recommended data items can be selected from a pool of data items. For example, assuming the server may contain several data items that the user has not yet viewed, the app can retrieve the recommended data items by performing the following operations:
[0098] First, the user's second interest characteristics can be determined based on the semantic features of the data entries the user has browsed in the past;
[0099] Then, based on the user's second interest feature, data entries that meet the preset third similarity requirement with the second interest feature are selected from the data entries that the user has not yet browsed, and these are used as data entries for recommendation.
[0100] The user's second interest feature can characterize the degree of interest the user has in the historically viewed data items. In one optional implementation, the semantic features of the user's historically viewed data items can be the embedding vectors of each historically viewed data item; while the user's second interest feature can be the mean vector of the embedding vectors of each data item—that is, the vector composed of the average value of each element of the embedding vector of each data item.
[0101] Based on this mean vector, the distance between the embedding vector of each selected data entry and the mean vector can be calculated—this distance measures the similarity between vectors and is therefore also called a similarity value. Then, the selected data entries can be sorted in descending order of similarity value, and, for example, the data entries that rank at the top and account for 50% of the total number can be selected as recommended data entries. Alternatively, data entries with similarity values greater than a preset similarity threshold can be selected as recommended data entries.
[0102] In this embodiment, the preset third similarity requirement can be adjusted according to actual needs to adapt to the recommendation needs of different users or different scenarios.
[0103] In this embodiment, by selecting data entries that meet a preset third similarity requirement with the second interest feature from the data entries that the user has not yet browsed, it is possible to first roughly filter out data entries that are relatively matched with the user's current interests, avoiding the problem of excessive data processing volume caused by using all data entries that the user has not yet browsed as data entries for recommendation; at the same time, this method can also ensure that the data entries for recommendation are as close as possible to the user's interests, making it easier for the user to find information of interest later, and improving the user's satisfaction and trust in the recommendation system.
[0104] The method provided in this application can determine at least two primary interest features of the user. These features represent the user's degree of interest in data items with varying degrees of relevance within the browsing sequence. Therefore, this solution separates the "user's degree of interest in the same browsing sequence," effectively refining the identification of user interests. Based on this, compared to existing technologies that recommend information based on a single user interest level, this solution, using at least two primary interest features, can achieve more refined information recommendation, making the recommended information more accurate and better suited to the user's actual needs.
[0105] Example 2
[0106] Example 2 mainly introduces a specific implementation method for applying the information recommendation method provided in Example 1 of this application to a real-world scenario. This specific implementation method is also used to solve the problem of how to accurately identify user interests and then make accurate information recommendations in the prior art.
[0107] Specifically, in Example 2, the information recommendation method provided in Example 1 is applied to a government website. A "coarse-fine ranking and recall recommendation system framework" is encapsulated in the website server. Through two-stage screening, the framework can match users' potential data browsing interests from a massive database. At the same time, it uses sequence debiasing to remove the influence of potential conditional variables on the system's capture of users' true interests. By separating users' primary and secondary interests, it can mine the real correlation between the data themselves, complete the data recommendation task with a higher hit rate, reduce the time users spend searching for government data, and improve data utilization efficiency.
[0108] The “coarse-fine ranking and recall recommendation system framework” includes two major modules: “coarse ranking model” and “fine ranking and recall model”. The fine ranking and recall model further includes a “sequence debiasing module” to complete fine-grained user interest mining.
[0109] In practical application scenarios, the application process of the "coarse-fine ranking and recall recommendation system framework" is illustrated in the diagram below. Figure 2a As shown: First, the "coarse ranking model" filters out potential data items related to the user's browsing sequence from massive amounts of government data based on text relevance, generating a coarse ranking candidate data item list; then, the "fine ranking recall model" uses the "sequence debiasing module" to process the user's interest sequence, and models the user's true data browsing intent by separating the primary and secondary interest sequences, and then performs fine ranking matching on the candidate data item list to generate the final relevant data recommendation results for the user.
[0110] The following details the specific steps for implementing information recommendation methods on government websites. For example... Figure 2b As shown, these steps include:
[0111] Step 21: The coarse ranking model segments the data entries in the government website database using the word segmentation module, obtains the word vectors of each segmented word in the title of each data entry based on the word vector embedding model, and calculates the embedding vector of the title of the data entry based on the word vectors;
[0112] Specifically, in one alternative implementation, step 21 may include, for example: Figure 2c The following sub-steps are shown:
[0113] Sub-step 211: The coarse-grained model in the government website server can obtain all data entries (specifically, titles, referred to as data entry titles below) available for visitors to browse on the data recommendation website, forming a complete set of data entry titles.
[0114] in, Represents the total number of headers for all data entries.
[0115] For a specific example, please refer to Figure 2d As shown, the data entry title set can contain various data entry titles, such as "Annual Afforestation Area Information of Dongying District", "Information on Drought Relief Agricultural Disaster Information Release in Linyi County", and "Sensor Data of High-Standard Farmland in Rencheng District", etc.
[0116] Sub-step 212: For each data entry title D n (The value of n ranges from [1, N]), the coarse-ranking model splits it into D using the word segmentation module in the coarse-ranking model. n =[w1,w2,...,w L ];
[0117] Where w l This represents the l-th word in the title of this data entry, where l ranges from [1, L]; L is the total length of the word list after word segmentation of the title of this data entry.
[0118] For a specific example, please refer to Figure 2d As shown, taking the data entry title "Rencheng District High-Standard Farmland Sensor Data" as an example, the word segmentation module can split it into the words "Rencheng District", "High-Standard", "Farmland", "Sensor", and "Data".
[0119] Sub-step 213: The coarse-ranking model uses a pre-trained word embedding model to generate a list of phrases D for each data entry title. n (The value of n is in the range [1, N]) Mapped to word vector groups
[0120] v l =Embedding(w l The value of l ranges from [1, L].
[0121] Where d is the dimension of the low-dimensional embedding vector, and Embedding(·) is the mapping function of the pre-trained vector embedding model.
[0122] For a specific example, please refer to Figure 2d As shown, the word embedding model can map various word segmentation results (specifically, a list of word groups) input into the model to corresponding word vector groups.
[0123] Sub-step 214: The coarse-ranking model, based on word vector groups, uses an average aggregation strategy to obtain the title embedding vector of the data entry titles. The value of n ranges from [1, N]. Therefore, the set of embedding vector representations of all data entry titles in the embedding space can be obtained as E = {E1, E2, ..., E...}. N}
[0124] The average aggregation strategy involves averaging the vectors of each word group to obtain the embedding vector for the data entry title. This embedding vector can better preserve the semantic features of the data entry title.
[0125] For a specific example, please refer to Figure 2d As shown, the word vector set obtained by the word vector embedding model can be input into the title embedding vector representation module. This module then uses an average aggregation strategy on the word vector set to obtain the set of embedding vector representations of the data entry title in the embedding space. Figure 2d The set of title embedding vector representations shown.
[0126] Step 22: The data recommendation website's server obtains a "set of browsed data entries" consisting of the titles of the data entries viewed by the user. The coarse-ranking model uses this set to calculate the embedding vector representation set E = {E1, E2, ..., E...} of all data entry titles. N In the process, the embedding vector of the title of each browsed data item in the set is determined, and the coarse-grained interest embedding representation of the user is obtained by calculating the mean vector based on the embedding vector of each browsed data item title.
[0127] The purpose of step 22 is to determine the user's browsing interests based on the titles of the data items the user has already viewed.
[0128] Specifically, taking user u as an example, assuming the server obtains S the "set of browsing data item titles" for that user. u ={D1,...,D S} where the subscript S represents the length of the user's browsing sequence. Therefore, the coarse-ranking model, based on S, can determine the embedding vector of S—specifically, the set of embedding vectors E. u ={E1,E2,...,E S Furthermore, by calculating the mean vector, a coarse embedding representation of user interests can be obtained. The obtained user interest coarse embedding representation It can represent the browsing interests of user u.
[0129] Step 23: The coarse-ranking model uses the embedded vector representation set E = {E1, E2, ..., E...} of all data entry titles obtained from the calculation. N In this process, the embedding vectors of the titles of unviewed data items are obtained, excluding the embedding vectors of the titles of each viewed data item. Based on the embedding vectors of the titles of unviewed data items and the user's coarse-grained interest embedding representation, the similarity score between the titles of unviewed data items and the user's interests is calculated. Based on the similarity score, the titles of unviewed data items are sorted, and the list of sorted titles of unviewed data items is used as the "coarse-ranked data recommendation list".
[0130] Specifically, for example, the following formula [4] can be used to calculate the similarity score δ between the title of the unviewed data item and the user's interest. x :
[0131]
[0132] The explanation of formula [4] is as follows:
[0133] E x Let E represent the embedding vector of the title of an unbrowsed data entry in set E. x ∈EE u The range of x is [1, N / S].
[0134] After calculating the similarity score δ x Then, the coarse-ranking model can sort the titles of the corresponding unviewed data items in descending order of similarity scores; subsequently, the top N1 unviewed data item titles can be selected to construct a coarse-ranked data recommendation list.
[0135] Step 24: The fine-grained ranking and recall model obtains the user's browsing sequence and uses the sequence debiasing module to generate the user's fine-grained main interest embedding vector and fine-grained secondary interest embedding vector;
[0136] In one alternative implementation, step 24 may include, for example: Figure 3 The sub-steps shown:
[0137] Sub-step 241: The fine-grained ranking and recall model targets the set of data item titles browsed by user u. Each data entry title D t (The value of t is in the range of [1, S]), and the autocorrelation coefficient η can be calculated according to a formula [1] similar to that in Example 1. t ;
[0138] Sub-step 242: Based on the calculated η t (The value of t ranges from [1, S]) The sequence debiasing module can process data in descending order of t value. The corresponding data entry titles are sorted sequentially.
[0139] Sub-step 243: The sequence debiasing module selects the first 70% of the data entry titles after sorting to form the main interest sequence of user u. The remaining 30% of the data entry titles are used to form a sequence of user u's secondary interests.
[0140] Sub-step 244: Based on Sequence debiasing module obtains the main interest sequence The embedding vectors of the titles of each data entry constitute the main interest data embedding vector set. based on Sequence debiasing module obtains secondary interest sequences The embedding vectors of the titles of each data entry constitute the set of embedding vectors for secondary interest data.
[0141] Sub-step 245: The sequence debiasing module obtains the set of browsing data item titles. Each data entry title D t (The value of t ranges from [1, S]) The frequency of browsing f on the government website s ;
[0142] Sub-step 246: The sequence debiasing module, based on f i Calculate the average frequency of browsing data entries in the browsing sequence.
[0143] Sub-step 247: Finally, the sequence debiasing module is based on and According to formulas [2] and [3] described in Example 1, the fine-grained main interest embedding vector representation of the user is calculated respectively. and fine-grained secondary interest embedding vector representation
[0144] The fine-grained primary interest embedding vector representation and the fine-grained secondary interest embedding vector representation of a user can be collectively referred to as the user's "fine-grained dual interest embedding vector representation".
[0145] By executing step 24, the user's interest representation is further refined using the sequence debiasing module, and embedding vectors for primary and secondary interests are generated respectively. This can accurately distinguish the user's degree of interest in different data items, thereby effectively separating primary and secondary interests.
[0146] Step 25: The fine-ranking recall model calculates the fine-ranking similarity score based on the data item titles and the user's fine-grained dual-interest embedding vector representation in the coarse-ranked data recommendation list; based on the fine-ranking similarity score, it generates the final data item recommendation list and recommends it to the user.
[0147] Specifically, the fine-ranking recall model can obtain the coarse-ranked data recommendation list generated by the coarse-ranking model. And based on get The set of embedding vectors E, which consists of the embedding vectors of the data entry titles in the dataset. c ;
[0148] Based on E c And fine-grained main interest embedding vector representation of users Calculate according to formula [5] Interest point scores between the embedding vectors of each data entry title and the user's fine-grained main interest embedding vector.
[0149]
[0150] in, For the sigmoid function, E a ∈E c The value of a ranges from [1, N1]. It should be noted that in this embodiment, the nonlinear transformation by the sigmoid function can make the calculation of the similarity score more accurate and enable the calculated similarity score to better reflect the user's preference for different data item titles.
[0151] Interest point scores are calculated. Then, according to From high to low order, proceed sequentially. Sort the data entry titles in the data;
[0152] from From the sorted data item titles, select the top N2 data item titles to generate a main interest recommendation list containing these N2 data item titles.
[0153] According to formula [6], calculate The set of data item titles that were not selected Interest point scores between the embedding vectors of each data entry title and the user's fine-grained sub-interest embedding vectors.
[0154]
[0155] in, For the sigmoid function, E b ∈ set The value of b is in the range of [1, N1-N2];
[0156] Based on interest point scores From high to low order, proceed sequentially. Sort the data entry titles in the data;
[0157] from From the sorted data item titles, select the top N2 data item titles to generate a secondary interest recommendation list containing these N2 data item titles.
[0158] Depend on and As can be seen from the method of determination, the main interest recommendation list It includes the titles of data items that users are most interested in, and a secondary interest recommendation list. This includes the titles of data entries that the user might be interested in but with a slightly lower level of interest.
[0159] In this embodiment of the application, it can be and Combine them into a single data recommendation list to obtain the final data item recommendation list. Will Recommendations are then made to users. This step considers not only the user's direct interests but also their potential interests, thus ensuring the accuracy and diversity of the recommendations.
[0160] In implementing the final data item recommendation list to users. Furthermore, in Example 2, step 26 can also be performed.
[0161] Step 26: Determine if the user has viewed the new data entry title. If yes, return to step 22; otherwise, the recommendation service can be terminated.
[0162] Applying the information recommendation method provided in this application to government websites, by adopting a data recommendation paradigm based on a coarse-fine ranking and recall framework, can reduce inference computation costs and improve the response speed of recommendation services when dealing with massive amounts of government data. Simultaneously, it captures users' browsing interests to achieve accurate data recommendations. By removing bias from the user's data item title browsing sequence in the fine-ranking and recall model, separating the user's primary and secondary interests, it increases the diversity of recommended data items while ensuring high relevance of the recommendation results, mines potential user data browsing intentions, models user interests in a fine-grained manner, and alleviates the information cocoon problem in data recommendation.
[0163] Example 3
[0164] To address the problem of how to accurately identify user interests and thus make accurate information recommendations in the existing technology, based on the same inventive concept as the above embodiments of this application, Embodiment 3 of this application provides an information recommendation device.
[0165] A schematic diagram of the specific structure of the device is shown below. Figure 4 As shown, it includes the following functional units:
[0166] The first interest determination unit 31 is used to determine at least two first interest features of the user based on the relevance of different data items browsed in the same browsing sequence of the user and the semantic features of the different data items; the at least two first interest features respectively characterize the degree of interest of the user in different data items in the browsing sequence within the historical time period corresponding to the browsing sequence.
[0167] The information recommendation unit 32 is used to select data items from the data items to recommend to the user based on the at least two first interest features.
[0168] In one optional implementation, the first interest determination unit 31 may be specifically used to: determine the user's fine-grained primary interest features and fine-grained secondary interest features based on the relevance and the semantic features.
[0169] Among them, fine-grained primary interest features characterize the degree of user interest in data items within a primary interest browsing sequence segmented from the same browsing sequence; fine-grained secondary interest features characterize the degree of user interest in data items within a secondary interest browsing sequence segmented from the same browsing sequence. The relevance of data items in the primary interest browsing sequence within the same browsing sequence is greater than the relevance of data items in the secondary interest browsing sequence within the same browsing sequence.
[0170] In one optional implementation, the first interest determination unit 31 may specifically be used for:
[0171] Determine the autocorrelation coefficients of the different data entries in the browsing sequence;
[0172] Based on the autocorrelation coefficient, the browsing sequence is divided into a primary interest browsing sequence and a secondary interest browsing sequence;
[0173] Based on the browsing frequency of data entries in the browsing sequence and the semantic features of data entries in the main interest browsing sequence, the user's fine-grained main interest features are determined.
[0174] Based on the browsing frequency and the semantic features of the data entries in the secondary interest browsing sequence, the user's fine-grained secondary interest features are determined.
[0175] In one alternative implementation, the information recommendation unit 32 may specifically be used for:
[0176] From the recommended data entries, select a first data entry that meets a preset first similarity requirement with the fine-grained main interest feature;
[0177] From the other recommended data entries besides the first data entry, select a second data entry that meets a preset second similarity requirement with the fine-grained secondary interest feature;
[0178] The first data entry and the second data entry are sorted and then recommended to the user.
[0179] In one optional implementation, the apparatus provided in this application embodiment may further include:
[0180] The second interest determination unit is used to determine the user's second interest features based on the semantic features of the data entries the user has browsed in the past; the user's second interest features characterize the degree of interest the user has in the data entries they have browsed in the past.
[0181] The data entry selection unit is used to select, based on the second interest feature, data entries that meet a preset third similarity requirement with the second interest feature from the data entries that the user has not yet browsed, as the data entries to be recommended.
[0182] The apparatus provided in this application can determine at least two primary interest features of the user. These features represent the user's degree of interest in data items with varying degrees of relevance within a browsing sequence. Therefore, this solution separates the "user's degree of interest in the same browsing sequence," effectively refining the identification of user interests. Based on this, compared to existing technologies that recommend information based on a single user interest level, this solution, using at least two primary interest features, can achieve more refined information recommendation, making the recommended information more accurate and better suited to the user's actual needs.
[0183] Example 4
[0184] Based on the same inventive concept as the foregoing embodiments of this application, Embodiment 4 of this application provides a computing device to solve the problem of how to accurately identify user interests and then make accurate information recommendations in the prior art.
[0185] like Figure 4As shown, the computing device includes a memory 41 and a processor 42. The memory 41 can be configured to store various other data to support operation on the electronic device. Examples of such data include instructions for any application or method used to operate on the electronic device. The memory 41 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0186] The processor 42, coupled to the memory 41, is used to execute the computer program stored in the memory 41 for performing an information recommendation method as described in Embodiment 1 of this application.
[0187] When the processor 42 executes the computer program in the memory 41, in addition to the functions described above, it can also perform other functions, as detailed in the descriptions of the preceding embodiments.
[0188] Furthermore, such as Figure 4 As shown, the computing device also includes other components such as a display 44, a communication component 43, a power supply component 45, and an audio component 46. Figure 4 The diagram only shows some components and does not mean that the computing device includes only these components. The components shown.
[0189] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a computer, can implement the methods provided in the above embodiments.
[0190] Accordingly, this application also provides a computer program product, which stores instructions that, when executed by a computer, cause the computer to implement the methods provided in the above embodiments.
[0191] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An information recommendation method, characterized in that, include: Based on the relevance of different data entries viewed in the user's historical browsing sequence within the same browsing sequence, and the semantic features of the different data entries, at least two primary interest features of the user are determined; The at least two first interest features respectively characterize the user's degree of interest in the data items with different degrees of relevance in the browsing sequence; Based on the at least two primary interest features, data entries are selected from the recommended data entries and recommended to the user.
2. The method as described in claim 1, characterized in that, Based on the relevance of different data entries viewed in the user's historical browsing sequence within the same browsing sequence, and the semantic features of the different data entries, at least two primary interest features of the user are determined, including: Based on the relevance and semantic features, the user's fine-grained primary interest features and fine-grained secondary interest features are determined; The fine-grained primary interest feature represents the degree of interest a user has in data items in a primary interest browsing sequence extracted from the same browsing sequence; the fine-grained secondary interest feature represents the degree of interest a user has in data items in a secondary interest browsing sequence extracted from the same browsing sequence. The relevance of data entries in the primary interest browsing sequence to the same browsing sequence is greater than the relevance of data entries in the secondary interest browsing sequence to the same browsing sequence.
3. The method as described in claim 2, characterized in that, Based on the relevance and semantic features, the user's fine-grained primary interest features and fine-grained secondary interest features are determined, including: Determine the autocorrelation coefficients of the different data entries in the browsing sequence; Based on the autocorrelation coefficient, the browsing sequence is divided into a primary interest browsing sequence and a secondary interest browsing sequence; Based on the browsing frequency of data entries in the browsing sequence and the semantic features of data entries in the main interest browsing sequence, the user's fine-grained main interest features are determined. Based on the browsing frequency and the semantic features of the data entries in the secondary interest browsing sequence, the user's fine-grained secondary interest features are determined.
4. The method as described in claim 2 or 3, characterized in that, Selecting data entries from the list of recommended data entries to recommend to the user based on the at least two first interest features includes: From the recommended data entries, select a first data entry that meets a preset first similarity requirement with the fine-grained main interest feature; From the other recommended data entries besides the first data entry, select a second data entry that meets a preset second similarity requirement with the fine-grained secondary interest feature; The first data entry and the second data entry are sorted and then recommended to the user.
5. The method as described in claim 1, characterized in that, The method further includes: Based on the semantic features of the data entries browsed in the user's history, the user's second interest feature is determined; the user's second interest feature represents the degree of interest the user has in the data entries browsed in history. Based on the second interest feature, data entries that meet a preset third similarity requirement with the second interest feature are selected from the data entries that the user has not yet browsed, and these are used as the recommended data entries.
6. An information recommendation device, characterized in that, include: An interest determination unit is used to determine at least two first interest features of the user based on the degree of relevance of different data entries browsed in the same browsing sequence of the user in the user's history, and the semantic features of the different data entries; The at least two first interest features respectively characterize the user's degree of interest in the data items with different degrees of relevance in the browsing sequence; An information recommendation unit is configured to select data items from the data items to recommend to the user based on the at least two first interest features.
7. The apparatus as claimed in claim 6, characterized in that, The interest determination unit is specifically used for: Based on the relevance and semantic features, the user's fine-grained primary interest features and fine-grained secondary interest features are determined; The degree of interest represented by the fine-grained primary interest feature is different from the degree of interest represented by the fine-grained secondary interest feature.
8. A computing device, characterized in that, include: Memory and processor, among which, The memory is used to store computer programs; The processor, coupled to the memory, is configured to execute the computer program stored in the memory for performing the method according to any one of claims 1 to 5.
9. A computer-readable storage medium storing a computer program, which, when executed by a computer, enables the implementation of the method described in any one of claims 1 to 5.
10. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 5.