An Information Recommendation Method and System Based on VSM and AMMK-means

The user interest model is constructed through VSM and AMMK-means, and the clustering center is determined using the maximum and minimum distance clustering algorithm, which solves the problem of deviation caused by artificial classification standards in the collaborative filtering algorithm and the poor selection of K-means initial centers, achieving higher recommendation accuracy and efficiency.

CN114625952BActive Publication Date: 2025-07-18CHINESE PEOPLES LIBERATION ARMY UNIT 93216
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011432407.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-10
Publication Date
2025-07-18
Estimated Expiration
2040-12-10

AI Technical Summary

Technical Problem

In the existing information recommendation system, the recommendation algorithm based on collaborative filtering relies on manual classification standards, resulting in a deviation in recommendation effect. The K-means algorithm needs to set the number of clusters in advance when clustering and the initial center selection is poor, which affects the accuracy of recommendation.

Method used

The information recommendation method based on VSM and AMMK-means is adopted to determine the clustering center through the maximum and minimum distance clustering algorithm, build a user interest model, and calculate the similarity between candidate information and user portraits, and recommend the information with the highest similarity.

Benefits of technology

It improves the accuracy of information recommendation, avoids deviations from manual classification standards, optimizes the initial cluster center selection, reduces the amount of calculation, and enhances the accuracy of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625952B_ABST
    Figure CN114625952B_ABST
Patent Text Reader

Abstract

The present invention discloses an information recommendation method and system based on VSM and AMMK-means, including: obtaining the item portraits of each candidate information; substituting the item portraits of each candidate information into a pre-constructed interest model to obtain the similarity between each candidate information and the user portrait; recommending the candidate information with the highest similarity to the user; the interest model is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed. Since the interest model in the present invention is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed, it is equivalent to being customized based on the item portraits that the user is interested in, avoiding the deviation between the information categories recommended to the user according to the classification criteria set by the editors and the information categories that the user is actually interested in. Compared with the traditional collaborative filtering algorithm, the accuracy of the recommendation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval, and particularly to an information recommendation method and system based on VSM and AMMK-means. Background Art

[0002] Situation assessment refers to understanding the relationships between battlefield objects based on target assessment, analyzing, reasoning, and judging multi-source information, and is used to support decision-making at the command level. Due to the characteristics of large data volume and diverse data types of battlefield information, each seat (i.e., user) has difficulty in information selection. This requires an auxiliary decision-making system to recommend information of interest to different seats according to the browsing records of each seat and the characteristics of the battlefield information itself.

[0003] Generally, recommendation algorithms based on collaborative filtering make recommendations according to user browsing records or feedback records. In addition to according to user browsing records or feedback records, some studies consider adding information categories when generating recommendation information. However, the information categories are obtained by manual classification by information editors, and the classification criteria can only represent the opinions of the editors, resulting in deviation of the recommendation effect. Summary of the Invention

[0004] In order to solve the above-mentioned deficiencies in the prior art, the present invention provides an information recommendation method and system based on VSM and AMMK-means.

[0005] In a first aspect, an information recommendation method based on VSM and AMMK-means is provided, including:

[0006] Obtaining the item portraits of each candidate information;

[0007] Substituting the item portraits of the candidate information into a pre-constructed interest model to obtain the similarity between the candidate information and the user portrait;

[0008] Recommending the candidate information with the highest similarity to the user;

[0009] The interest model is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed.

[0010] Preferably, the construction of the interest model includes:

[0011] Obtaining the item portraits of the information that the user has browsed, and using VSM to characterize the item portraits of the information that the user has browsed;

[0012] Clustering the item portraits of the information that the user has browsed by AMMK-means, and using the clustering result as the information categories that the user is interested in;

[0013] Calculate the weights of each information category based on the number of information items browsed by the user in each information category and the total number of information items the user has browsed respectively;

[0014] Generate a user profile based on the information categories the user is interested in and the weights of the information categories;

[0015] Calculate the similarity between the item profile of the candidate information and the item profile in the user profile.

[0016] Further, the clustering of the item profiles of the information browsed by the user through AMMK-means, and taking the clustering result as the information categories the user is interested in, includes:

[0017] Generate a data set based on the item profiles of the information browsed by the user;

[0018] Determine the cluster centers and the number of cluster centers for the samples in the data set using the maximum-minimum distance clustering algorithm;

[0019] Take the number of cluster centers as the K value in the K-means algorithm, and take all the obtained cluster centers as the initial cluster centers in the K-means clustering algorithm;

[0020] Based on the distances between each sample in the data set and each initial cluster center, obtain the clustering result when the set constraint conditions are met;

[0021] Take the clustering result as the information categories the user is interested in.

[0022] Further, the determination of the cluster centers and the number of cluster centers for the samples in the data set using the maximum-minimum distance clustering algorithm includes:

[0023] Calculate the average value of the sample attributes, calculate the distances between each sample and the average value, and take the sample corresponding to the minimum distance as the first cluster center C1;

[0024] Select the sample farthest from C1 as the second cluster center C2;

[0025] Calculate the distances D i1 and D i2 from all the remaining samples to C1 and C2. If D l = max{min(D i1 , D i2 ), i = 1, 2,...n}, and D l > θD 12 , where θ is a given value and D 12 is the distance between C1 and C2, then take x l as the third cluster center C3;

[0026] If C3 exists, calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,... n. If D j > θD 12 then establish the fourth clustering center;

[0027] And so on, until the maximum and minimum distance is not greater than θD 12 End the calculation of finding the clustering center, and obtain the clustering center and the number of clustering centers.

[0028] Preferably, the expression of the interest model is as follows:

[0029] V seat = (w1 * T1, w2 * T2,..., w m * T m ) T

[0030] In the formula: V seat represents the user portrait; w m represents the weight of the m-th information category; T m represents the feature vector of the m-th information category.

[0031] Preferably, the similarity is calculated as follows:

[0032]

[0033] In the formula: seat is the item portrait in the user portrait, w i is the weight of the information category to which the candidate information d i belongs, T i T is the feature vector of the information category to which the candidate information d i belongs, is the feature vector of d i .

[0034] In a second aspect, an information recommendation system based on VSM and AMMK-means is provided, including:

[0035] An acquisition module, configured to acquire the item portraits of each candidate information;

[0036] A similarity calculation module, configured to substitute the item portraits of the candidate information into a pre-constructed interest model to obtain the similarity between the candidate information and the user portrait;

[0037] A recommendation module, configured to recommend the candidate information with the highest similarity to the user;

[0038] The interest model is constructed based on the VSM, AMMK-means, and the item portraits of the information that the user has browsed.

[0039] Preferably, the system further includes a construction module for the interest model; the construction module for the interest model includes:

[0040] A first construction unit, configured to obtain the item portraits of the information that the user has browsed, and use VSM to represent the item portraits of the information that the user has browsed;

[0041] An information category construction unit, configured to cluster the item portraits of the information that the user has browsed by using AMMK-means, and use the clustering result as the information categories that the user is interested in;

[0042] A weight calculation unit, configured to calculate the weights of each information category respectively according to the number of information browsed by the user included in each information category and the total number of information that the user has browsed;

[0043] A user portrait construction unit, configured to generate a user portrait based on the information categories that the user is interested in and the weights of the information categories;

[0044] A calculation unit, configured to calculate the similarity between the item portrait of the candidate information and the item portrait in the user portrait.

[0045] In a third aspect, a storage device is provided, in which multiple program codes are stored, and the program codes are adapted to be loaded and run by a processor to execute the information recommendation method based on VSM and AMMK-means described in any one of the above technical solutions.

[0046] In a fourth aspect, a control device is provided, including a processor and a storage device, the storage device is adapted to store multiple program codes, and the program codes are adapted to be loaded and run by the processor to execute the information recommendation method based on VSM and AMMK-means described in any one of the above technical solutions.

[0047] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:

[0048] In this embodiment, first, the item portraits of each candidate information are obtained; then, the item portraits of each candidate information are substituted into a pre-constructed interest model to obtain the similarity between each candidate information and the user portrait; finally, the candidate information with the highest similarity is recommended to the user. Since the interest model in this embodiment is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed, it is equivalent to being customized based on the item portraits that the user is interested in, avoiding the deviation between the information categories recommended to the user according to the classification criteria set by the editors and the information categories that the user is actually interested in. Compared with the traditional collaborative filtering algorithm, the accuracy of the recommendation is improved. Brief Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of the main steps of an information recommendation method based on VSM and AMMK-means according to an embodiment of the present invention;

[0050] Figure 2 It is a main step flowchart of the improved algorithm AMMK-means in the embodiment of the present invention;

[0051] Figure 3 It is a main structural block diagram of an information recommendation method based on VSM and AMMK-means according to an embodiment of the present invention. Detailed Embodiment

[0052] To better understand the present invention, the content of the present invention will be further described below in conjunction with the accompanying drawings of the specification and examples.

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] To solve the problems existing in the existing information recommendation models, this embodiment provides an information recommendation method based on VSM and AMMK-means, where AMMK-means refers to an improved K-means clustering algorithm based on the maximum minimum distance clustering algorithm. This method constructs an interest model representing the information content according to the important influence of the information content and information category on the user interest and classifies the quantified information at the same time. In this process, the VSM is first used to represent the text information of the information, and then the AMMK-means algorithm is used to cluster the information to construct the user's interest model. Finally, the similarity between the user's interest model and the candidate information is calculated, and the information that the user is interested in is recommended to the user. While inheriting the interpretability advantage of the existing content recommendation algorithm, this embodiment uses the improved clustering algorithm to avoid the problems existing in the manual classification. This embodiment proves through experiments that compared with the traditional collaborative filtering algorithm, the method proposed in this embodiment has high prediction accuracy.

[0055] In the embodiment of the present invention, refer to the attached Figure 1 , Figure 1 is a flowchart of an information recommendation method based on VSM and AMMK-means. As Figure 1 shown, the information recommendation method based on VSM and AMMK-means in the embodiment of the present invention mainly includes the following steps:

[0056] S1. Obtain the item portraits of each candidate information;

[0057] S2. Substitute the item portraits of the candidate information into the pre-constructed interest model to obtain the similarity between each candidate information and the user portrait;

[0058] S3. Recommend the candidate information with the highest similarity to the user;

[0059] The interest model is constructed based on VSM, AMMK-means and the item portraits of the information that the user has browsed.

[0060] In this embodiment, the so-called information fusion, also known as data fusion and multi-sensor fusion, can be defined as an information processing process that uses computer technology to automatically analyze and synthesize several sensor observation information obtained in time series under certain criteria in order to complete task decision-making and evaluation.

[0061] In this embodiment, the information keyword refers to the most representative word in the information, which can represent the uniqueness and distinctiveness of the information. Usually, the information keyword is extracted through a text processing algorithm.

[0062] In this embodiment, the information feature vector referred to is that since the content included in the information belongs to the text type, a multi-dimensional vector d = {w1, w2,..., w i ,...w n} is often used to represent the information content, and the result of vectorizing the information text is called the information feature vector.

[0063] In one implementation, the construction of the interest model in S2 includes:

[0064] Obtain the item portrait of the information that the user has browsed, and use VSM to represent the item portrait of the information that the user has browsed;

[0065] Cluster the item portraits of the information that the user has browsed through AMMK-means, and use the clustering result as the information category that the user is interested in;

[0066] Calculate the weights of each information category according to the number of information browsed by the user in each information category and the total number of information that the user has browsed;

[0067] Generate a user portrait based on the information categories that the user is interested in and the weights of the information categories;

[0068] Calculate the similarity between the item portrait of the candidate information and the item portrait in the user portrait.

[0069] In this embodiment, VSM can be used to represent the item portrait of the information that the user has browsed, including:

[0070] In the vector space model VSM, each document is represented by a feature vector to represent the multi-dimensional information in the document. The feature vector is the item portrait. Constructing the item portrait of the information can not only reflect the high-dimensionality of the information but also facilitate clustering as clustering information to construct the user interest model;

[0071] This embodiment uses the vector space model to represent the information feature vector. Given the information set X = {X1, X2,..., X i ,...X n}, the vectorization representation of the information is:

[0072]

[0073] where w ij represents the weight of keyword j in information i.

[0074] In the VSM construction process, the dimension m of the keyword set needs to be determined first. Keywords are used to characterize the features of documents. When the number of keywords increases, as m increases, the time complexity increases. On the premise of ensuring the representation effect, in order to reduce the time cost, in this embodiment, the first 5 keywords in each piece of information are extracted to represent the piece of information (generally, taking 3 and 5 has the best effect). Then, the TF-IDF algorithm is used to obtain the dimension m of the keyword set in the information set, and the TF-IDF algorithm is used to calculate the weight w ij 。

[0075] The process of calculating the weight w using the TF-IDF algorithm ij includes: The calculation of the TF-IDF algorithm can be divided into two parts: term frequency (TF) and inverse document frequency (IDF). The product of these two parts jointly determines the weight of the document terms.

[0076] The calculation formula of TF is as follows:

[0077]

[0078] where count(i, j) represents the frequency of keyword i in information document j, and size(j) represents the total number of information j.

[0079] The IDF calculation formula is:

[0080]

[0081] N represents the total number of the information set, and n(i) represents the number of information in which keyword i appears.

[0082] The weight is calculated from TF and IDF as:

[0083] w ij = TF(i, j) * IDF(i)

[0084] After processing the weight using the normalization method, w ij :

[0085]

[0086] In one embodiment, the inventors found that the traditional recommendation method based on VSM + Kmeans has the following disadvantages:

[0087] 1) Usually, the recommendation algorithm based on collaborative filtering makes recommendations according to the user's browsing records or feedback records. In addition, some methods consider adding information category factors, but the information categories are manually classified by information editors, and the classification criteria can only represent the opinions of the editors, resulting in a deviation in the recommendation effect.

[0088] 2) There are limitations in using the K-means algorithm for clustering: First, the algorithm needs to preset the number of clusters in clustering, but it is very difficult to give an accurate number of clusters in actual applications. For different data sets, there is no basis for choosing the reference of the number of clusters, and a large number of training experiments are required. Second, the initial cluster centers of the algorithm are obtained randomly. If the initial center positions are not selected appropriately, it is very likely to increase the computational amount and fail to obtain the global optimal solution.

[0089] Therefore, the maximum-minimum distance clustering algorithm was first used in the field of pattern recognition. By exploring the Euclidean distance between clusters and using the sample points that are as far apart as possible as the initial centers for clustering, it can effectively avoid the situation where the clustering result is poor due to the too-close selection of the initial centers. And after completing the selection of the initial cluster centers, the number of clusters that are expected to be generated is naturally obtained, making up for the deficiency of the unknown number of classes in K-means clustering.

[0090] Specifically, clustering the item portraits of the information that the user has browsed through AMM K-means, and using the clustering result as the information categories that the user is interested in, including:

[0091] Generating a data set based on the item portraits of the information that the user has browsed through;

[0092] Determining the cluster centers and the number of cluster centers for the samples in the data set using the maximum-minimum distance clustering algorithm;

[0093] Taking the number of cluster centers as the K value in the K-means algorithm, and taking all the obtained cluster centers as the initial cluster centers in the K-means clustering algorithm;

[0094] Based on the distances between each sample in the data set and each initial cluster center, obtaining the clustering result when the set constraint conditions are met;

[0095] Taking the clustering result as the information categories that the user is interested in.

[0096] Specifically, the step of determining the cluster centers and the number of cluster centers for the samples in the data set using the maximum-minimum distance clustering algorithm includes:

[0097] Calculating the average value of the sample attributes, calculating the distances between each sample and the average value, and taking the sample corresponding to the minimum distance as the first cluster center C1;

[0098] Selecting the sample that is farthest from C1 as the second cluster center C2;

[0099] Calculating the distances D i1 and D i2 , if D l= max{min(D i1 , D i2 ), i = 1, 2,...n}, and D l > θD 12 , θ is a given value, D 12 is the distance between C1 and C2, then take x l as the third cluster center C3;

[0100] If C3 exists, calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,...n, if D j > θD 12 then establish the fourth cluster center;

[0101] And so on, until the maximum minimum distance is not greater than θD 12 to end the calculation of finding the cluster centers, and obtain the cluster centers and the number of cluster centers.

[0102] This clustering algorithm does not need to set the number of clusters, nor does it need to pre-estimate the number of clusters using a large number of experiments, and the computational complexity is smaller; it not only optimizes the selection of the initial cluster centers, but also avoids the deficiency of the local optimal solution caused by the random way of obtaining the initial cluster centers in the previous algorithms.

[0103] In one embodiment, the improved algorithm AMMK-means refers to Appendix Figure 2 , Figure 2 which is the main step flowchart of the improved algorithm AMMK-means in the embodiments of the present invention. The specific steps of the improved algorithm AMMK-means are as follows:

[0104] For the given data set X = {X1, X2,..., X n}:

[0105] Step 1: Calculate the average value of the sample attributes, calculate the distance between each sample and the average value, and take the sample corresponding to the minimum distance as the first cluster center C1;

[0106] Step 2: Given θ, 0 < θ < 1;

[0107] Step 3: Select the sample point corresponding to the farthest D i1 from C1 as the second cluster center C2;

[0108]

[0109] Step 4: Calculate the distances D i1 and D i2, if D l = max{min(D i1 , D i2 ), i = 1, 2,... n}, and D l > θD 12 , D 12 is the distance between C1 and C2, take x l as the third cluster center C3;

[0110] Step 5. If C3 exists, calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,... n. If D j > θD 12 then continue to search for and establish cluster centers, and so on until the distance is not greater than θD 12 ;

[0111] Step 6. Take the number of cluster centers obtained in Steps 3, 4, and 5 as the K value in the K-means algorithm, and take all the obtained cluster centers as the initial cluster centers in the K-means clustering algorithm;

[0112] Step 7. Calculate the distance between each sample in the sample set and the cluster center, and assign the remaining samples to the centroid cluster with the closest distance according to the nearest distance principle, and update the centroid of each cluster;

[0113] Step 8. Repeat Step 7 until the sum of squared errors criterion is satisfied to minimize the objective function to complete the clustering and obtain the clustering result.

[0114] The objective function in Step 8 above is shown as follows:

[0115]

[0116] In the formula, p represents the data object, C i represents the centroid, and J c represents the sum of squared errors of all objects in the dataset.

[0117] In the method proposed in this embodiment, the AMMK-means algorithm does not require presetting the number of clusters in clustering, does not require a large number of experiments in advance to estimate the number of clusters, and has a smaller computational amount; and the initial cluster centers of this algorithm are selected appropriately, avoiding the deficiencies of increased computational amount and failure to obtain the global optimal solution caused by obtaining the initial cluster centers in a random manner.

[0118] In the existing process of clustering based on the Chamelon algorithm, on the one hand, when constructing the K-nearest neighbor graph (Gk graph), it is necessary to calculate the similarity between pairwise data points and take the top K values in the order of decreasing similarity. The K value of the K-nearest neighbor graph and the threshold of the similarity function need to be given artificially, and a large amount of prior knowledge is required to give these parameters, which is quite difficult. The clustering algorithm provided in this embodiment avoids artificially giving the K value of the K-nearest neighbor graph and the threshold of the similarity function when constructing the G K graph, avoids the large amount of prior knowledge required to give these parameters, and reduces the difficulty. On the other hand, when dividing the G K graph into unconnected subgraphs and using them as the initial clusters of clustering, it is necessary to divide the G K graph into two approximately equal subgraphs according to the minimum cut principle, and then use the divided subgraphs as the initial clusters, and continuously repeat the above process until the division standard is reached to complete the division process. However, the division technique used when dividing the G K graph increases the complexity of the algorithm, and it is difficult to use the minimum bisection selection. The clustering method provided in this embodiment reduces the algorithm complexity and avoids the difficulty of using the minimum bisection selection.

[0119] In one embodiment, the user constructs a user interest model through the information that has been browsed. The first-layer nodes of this model are users, the second-layer nodes are information categories, and the third-layer nodes are the battlefield information browsed by the user.

[0120] If the user has browsed m different pieces of battlefield information, the user interest model can be expressed as:

[0121] seat={(T1, w1, n1),..., (T m , w m , n m )}.

[0122] Where T i represents the feature vector of the i-th information category, w i represents the weight of the i-th information category, and n i represents the number of pieces of information browsed by the user included in the i-th information category.

[0123] The feature vector of a certain information category is obtained by weighted averaging of all the feature vectors of the information browsed included in this category.

[0124] The calculation formula for the feature vector T i of the i-th information category is:

[0125]

[0126] Where, E jDenote the set of information browsed by users in information category i, e j Denote the information feature vector, I j Denote the user interest degree for the j-th information in this category. The information browsed by users indicates the users' interest in this information. Therefore, set I j to 1, and the formula can be simplified as:

[0127]

[0128] Furthermore, the value of w i is calculated according to the ratio of the information browsed by users in the i-th information category to the total browsed information. The calculation formula is:

[0129]

[0130] When calculating, the user interest model is expressed as:

[0131] V seat =(w1*T1, w2*T2,..., w m *T m ) T

[0132] Finally, use the cosine similarity to calculate the similarity between the candidate battlefield information d i and the user. The calculation formula:

[0133] where w i *T i T is the feature vector of the battlefield information category to which the candidate news d i belongs, is the feature vector of d i .

[0134] This embodiment can accurately express the users' interests, improve the recommendation effect, and avoid the defect that when using the collaborative filtering-based recommendation algorithm to make recommendations based on users' browsing records or feedback records, since the information categories are manually classified by information editors, the classification criteria can only represent the opinions of the editors.

[0135] In a specific implementation, first obtain the feature attributes of the information; then analyze the information browsed by the users to generate a user profile and calculate the feature similarity between the user profile and the candidate information. Finally, recommend the information with a high similarity to the users according to the similarity.

[0136] The above content-based recommendation method generally includes three steps: item profile, user profile, and recommendation generation.

[0137] An item portrait represents an item using feature information. The attributes describing the item include structured data and unstructured data. The unstructured data needs to be converted into structured data before it can be used in the model.

[0138] Currently, the commonly used method for item portrait is the Vector Space Model (VSM) based on TF-IDF weights.

[0139] VSM converts text documents into spatial vectors, and TF-IDF is used to calculate the keyword weights of each document. Since there are phenomena such as synonyms and polysemy among the words in the document, this will reduce the robustness and accuracy of the recommendation model. To enhance the generalization ability of the model for problems such as polysemy and synonyms, semantic analysis and knowledge graphs are applied to the recommendation system.

[0140] A user portrait constructs a user interest model based on the features that the user has browsed or evaluated in the past. Such models mainly include two parts: text classification and constructing a hierarchical user interest model. That is, first, it is necessary to cluster the item portraits of the information browsed by the user to obtain information categories and the corresponding features of each information category (i.e., item portraits); second, calculate the weights of the information categories; finally, count the number of information browsed by the user included in the information category. Traditional text classification models include the nearest neighbor algorithm, Rocchio algorithm, decision tree method, linear classification method, Bayesian classifier, etc. The construction process of the user interest hierarchical model has a hierarchical model with a three-layer structure of "user-category-item" or a hierarchical model with a three-layer structure of "user-interest-item". Recommendation is to recommend a set of items with the highest relevance to the user by comparing the feature similarity between the user portrait and the candidate items. Commonly used similarity calculation methods include the Pearson correlation coefficient and cosine similarity.

[0141] The item portrait referred to in this embodiment is a series of labels for each item. One of the functions of the item portrait is that it can be used as the item features in the recommendation model. On the other hand, in the recommendation system, the item portrait is the basis of the user portrait: item portrait + user behavior = user portrait.

[0142] It should be noted that although the above steps are described in a specific order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of the present invention, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these changes are all within the protection scope of the present invention.

[0143] Based on the same inventive concept, as Figure 3 shown, the present invention also provides an information recommendation system based on VSM and AMMK-means, including:

[0144] An acquisition module, configured to acquire item portraits of each candidate information;

[0145] A similarity calculation module, configured to substitute the item portraits of the candidate information into a pre-constructed interest model to obtain the similarity between the candidate information and the user portrait;

[0146] A recommendation module, configured to recommend the candidate information with the highest similarity to the user;

[0147] The interest model is constructed based on VSM, AMMK-means, and item portraits of information that the user has browsed.

[0148] In the embodiment, the system further includes a construction module for the interest model; the construction module for the interest model includes:

[0149] A first construction unit, configured to acquire item portraits of information that the user has browsed, and use VSM to represent the item portraits of information that the user has browsed;

[0150] An information category construction unit, configured to cluster the item portraits of information that the user has browsed through AMMK-means, and use the clustering result as the information categories that the user is interested in;

[0151] A weight calculation unit, configured to calculate the weights of each information category according to the number of information browsed by the user in each information category and the total number of information browsed by the user;

[0152] A user portrait construction unit, configured to generate a user portrait based on the information categories that the user is interested in and the weights of the information categories;

[0153] A calculation unit, configured to calculate the similarity between the item portrait of the candidate information and the item portrait in the user portrait.

[0154] The maximum minimum distance method adopted in this embodiment is based on the Euclidean distance, and selects objects that are as far apart as possible as the clustering centers, avoiding the situation where the clustering centers may be too close when the initial values are selected by the K-means method. It not only intelligently determines the number of initial clustering centers, but also improves the efficiency of dividing the initial data set.

[0155] In the embodiment, the information category construction unit is specifically configured to:

[0156] Generate a data set based on the item portraits of information that the user has browsed;

[0157] Use the maximum minimum distance clustering algorithm to determine the clustering centers and the number of clustering centers for the samples in the data set;

[0158] Take the number of the clustering centers as the value of K in the K-means algorithm, and take all the obtained clustering centers as the initial clustering centers in the K-means clustering algorithm;

[0159] Based on the distances between each sample in the dataset and each initial clustering center, obtain the clustering result when the set constraint conditions are met;

[0160] Take the clustering result as the information categories that the user is interested in.

[0161] In the embodiment, determining the clustering centers and the number of clustering centers for the samples in the dataset by using the maximum-minimum distance clustering algorithm includes:

[0162] Calculate the average value of the sample attributes, calculate the distances between each sample and the average value, and take the sample corresponding to the minimum distance as the first clustering center C1;

[0163] Select the sample with the farthest distance from C1 as the second clustering center C2;

[0164] Calculate the distances D i1 and D i2 of the remaining all samples to C1 and C2, if D l = max{min(D i1 , D i2 ), i = 1, 2,...n}, and D l > θD 12 , θ is a given value, D 12 is the distance between C1 and C2, then take x l as the third clustering center C3;

[0165] If C3 exists, then calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,...n, if D j > θD 12 then establish the fourth clustering center;

[0166] And so on, until the maximum-minimum distance is not greater than θD 12 to end the calculation of finding the clustering centers, and obtain the clustering centers and the number of clustering centers.

[0167] In the embodiment, the expression of the interest model is shown as follows:

[0168] V seat = (w1*T1, w2*T2,..., w m *T m ) T

[0169] Where: V seat represents the user profile; w m represents the weight of the m-th information category; T m represents the feature vector of the m-th information category.

[0170] In the embodiment, the similarity is calculated according to the following formula:

[0171]

[0172] Where: seat is the item profile in the user profile, w i is the weight of the information category to which the candidate information d i belongs, T i T is the weight of the information category to which the candidate information d i belongs, T i *T i T is the weight of the information category to which the candidate information d i belongs, is the feature vector of d i and is the feature vector of d.

[0173] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0174] Further, the present invention also provides a storage device. In an embodiment of the storage device according to the present invention, the storage device may be configured to store a program for executing the information recommendation method based on VSM and AMMK-means in the above method embodiment. This program can be loaded and run by a processor to implement the above information recommendation method based on VSM and AMMK-means. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The storage device may be a storage device formed by various electronic devices. Optionally, the storage in the embodiments of the present invention is a non-transitory computer-readable storage medium.

[0175] Further, the present invention also provides a control device. In an embodiment of the control device according to the present invention, the control device includes a processor and a storage device. The storage device may be configured to store a program for executing the information recommendation method based on VSM and AMMK-means in the above method embodiment. The processor may be configured to execute the program in the storage device, and this program includes but is not limited to the program for executing the information recommendation method based on VSM and AMMK-means in the above method embodiment. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The control device may be a control device formed by various electronic devices.

[0176] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0177] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0178] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in one or more of the blocks.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in one or more of the blocks.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still modifications or equivalent replacements can be made to the specific embodiments of the present invention, and any modifications or equivalent replacements made without departing from the spirit and scope of the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. An information recommendation method based on VSM and AMMK - means, characterized in that Including: Obtain the item portraits of each candidate information; Substitute the item portraits of the candidate information into a pre-constructed interest model to obtain the similarity between the candidate information and the user portrait; Recommend the candidate information with the highest similarity to the user; The interest model is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed; The construction of the interest model includes: Obtain the item portraits of the information that the user has browsed, and use VSM to represent the item portraits of the information that the user has browsed; Cluster the item portraits of the information that the user has browsed through AMMK-means, and use the clustering result as the information categories that the user is interested in; Calculate the weights of each information category according to the number of information browsed by the user in each information category and the total number of information that the user has browsed; Generate a user portrait based on the information categories that the user is interested in and the weights of the information categories; Calculate the similarity between the item portrait of the candidate information and the item portrait in the user portrait; The clustering of the item portraits of the information that the user has browsed through AMMK-means, and using the clustering result as the information categories that the user is interested in, includes: Generate a data set based on the item portraits of the information that the user has browsed; Use the maximum-minimum distance clustering algorithm to determine the clustering centers and the number of clustering centers for the samples in the data set; Use the number of clustering centers as the K value in the K-means algorithm, and use all the obtained clustering centers as the initial clustering centers in the K-means clustering algorithm; Based on the distances between each sample in the data set and each initial clustering center, obtain the clustering result when the set constraint conditions are met; Use the clustering result as the information categories that the user is interested in; The use of the maximum-minimum distance clustering algorithm to determine the clustering centers and the number of clustering centers for the samples in the data set includes: Calculate the average value of the sample attributes, calculate the distances between each sample and the average value, and use the sample corresponding to the minimum distance as the first clustering center C1; Select the sample with the farthest distance from C1 as the second clustering center C2; Calculate the distances D from all the remaining samples to C1 and C2 i1 and D i2 , if D l = max{min(D i1 , D i2 ), i = 1, 2,... n}, and D l > θD 12 , θ is a given value, D 12 is the distance between C1 and C2, then take x l as the third cluster center C3; If C3 exists, then calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,... n, if D j > θD 12 Then establish the fourth cluster center; By analogy, until the maximum and minimum distance is not greater than θD 12 End the calculation of finding the cluster centers, and obtain the cluster centers and the number of cluster centers; The expression of the interest model is shown as follows: Where: V seat represents the user profile; w m represents the weight of the m-th information category; T m represents the feature vector of the m-th information category; The similarity is calculated as follows: where: seat is the item portrait in the user portrait, w i is the weight of the information category to which the candidate information d i belongs, and T i T is the feature vector of the information category to which the candidate information d i belongs, and is the feature vector of d i .

2. An information recommendation system based on VSM and AMMK-means, characterized in that, Including: An acquisition module for obtaining the item portraits of each candidate information; A similarity calculation module for substituting the item portraits of the candidate information into a pre-constructed interest model to obtain the similarity between the candidate information and the user portrait; A recommendation module for recommending the candidate information with the highest similarity to the user; The interest model is constructed based on VSM, AMMK-means, and the item portraits of the information that the user has browsed; The system further includes a construction module for the interest model; The construction module for the interest model includes: A first construction unit for obtaining the item portraits of the information that the user has browsed, and using VSM to represent the item portraits of the information that the user has browsed; An information category construction unit, which is used to cluster the item portraits of the information that the user has browsed through AMMK-means, and use the clustering result as the information categories that the user is interested in; A weight calculation unit, which is used to calculate the weights of each information category according to the number of information browsed by the user in each information category and the total number of information that the user has browsed; A user portrait construction unit, which is used to generate a user portrait based on the information categories that the user is interested in and the weights of the information categories; A calculation unit, which is used to calculate the similarity between the item portrait of the candidate information and the item portrait in the user portrait; The information category construction unit specifically includes: Generating a data set based on the item portraits of the information that the user has browsed; Using the maximum-minimum distance clustering algorithm to determine the clustering centers and the number of clustering centers for the samples in the data set; Taking the number of clustering centers as the value of K in the K-means algorithm, and taking all the obtained clustering centers as the initial clustering centers in the K-means clustering algorithm; Based on the distances between each sample in the data set and each initial clustering center, obtaining a clustering result when the set constraint conditions are met; Taking the clustering result as the information categories that the user is interested in; The step of using the maximum-minimum distance clustering algorithm to determine the clustering centers and the number of clustering centers for the samples in the data set in the information category construction unit specifically includes: Calculating the average value of sample attributes, calculating the distances between each sample and the average value, and taking the sample corresponding to the minimum distance as the first clustering center C1; Selecting the sample with the farthest distance from C1 as the second clustering center C2; Calculate the distances D from all the remaining samples to C1 and C2 i1 and D i2 , if D l = max{min(D i1 , D i2 ), i = 1, 2,... n}, and D l > θD 12 , θ is a given value, D 12 is the distance between C1 and C2, then take x l as the third clustering center C3; If C3 exists, then calculate D j = max{min(D i1 , D i2 , D i3 ), i = 1, 2,... n, if D j > θD 12 Then establish the fourth cluster center; And so on until the maximum and minimum distances are not greater than θD 12 End the calculation of finding the cluster centers, and obtain the cluster centers and the number of cluster centers; The expression of the interest model is shown as follows: V seat =(w1*T1, w2*T2, …, w m *T m ) T Where: V seat represents the user profile; w m represents the weight of the m-th information category; T m represents the feature vector of the m-th information category; The similarity is calculated as follows: where: seat is the item portrait in the user portrait, w i is the weight of the information category to which the candidate information d i belongs, T i T is the feature vector of the information category to which the candidate information d i belongs, is the feature vector of d i .

3. A storage device that stores multiple program codes, characterized in that, The program code is suitable to be loaded and run by a processor to execute the information recommendation method based on VSM and AMMK-means described in claim 1.

4. A control device, comprising a processor and a storage device, the storage device being adapted to store a plurality of program codes, characterized in that, The program code is suitable to be loaded and run by the processor to execute the information recommendation method based on VSM and AMMK-means described in claim 1.

Citation Information

Patent Citations

  • College library user portrait model construction method based on multi-view binary k-means

    CN110532306A

  • Air target clustering method based on K-means clustering

    CN110781963A