User portrait recommendation method and system based on support vector machine
Through the hypergraph structure dataset and support vector machine model, the problem that traditional graph structures are difficult to model the multi-dimensional coupling relationship between users, items and scenes is solved, and high-precision user-item interaction probability prediction and personalized recommendations are achieved.
Patent Information
- Application Number
- CN202510714483.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Traditional graph structures find it difficult to effectively model the multi-dimensional coupling relationship between users, items, and scenes, resulting in insufficient feature expression and loss of semantic information. Existing recommendation methods have significant deficiencies in processing multimodal and high-order correlation information.
Using a hypergraph structured dataset, we generate embedding vectors by defining nodes and hyperedges of users, items, and scenes. We then use a support vector machine model combined with positive and negative sample training of log records to predict the probability of user-item interaction and generate a personalized recommendation list.
It improves the prediction accuracy of user-item interaction probability, enhances the generalization ability of the model and the accuracy of recommendations, improves the degree of personalization, and enhances user experience.
Smart Images

Figure CN120653834A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a user portrait recommendation method and system based on a support vector machine. Background Art
[0002] With the rapid development of Internet technology and the continuous expansion of user scale, personalized recommendation systems have become an important means to enhance user experience and platform commercial value. Early recommendation systems were mainly based on collaborative filtering methods and completed recommendations by analyzing the similarities between users or items. However, such methods have shown obvious limitations when faced with sparse data, cold start problems, and the inability to effectively model complex multi-dimensional interaction relationships. In recent years, the combination of graph structures and deep learning technologies has brought new breakthroughs to recommendation systems, especially the use of graph neural networks (GNNs) to model user-item interaction relationships, which has improved recommendation accuracy and generalization capabilities. In addition, the development of embedding learning technology has enabled high-dimensional discrete features to be effectively mapped to low-dimensional continuous vector spaces, thereby supporting more efficient feature fusion and model training.
[0003] Although existing recommendation methods have been optimized in multiple dimensions, they still have significant deficiencies in processing multimodal and high-order correlation information. Especially in complex scenarios, traditional graph structures find it difficult to effectively model the multidimensional coupling relationship between users, items, and scenes, resulting in insufficient feature expression and loss of semantic information. Existing technologies usually adopt multi-graph fusion or heterogeneous graph modeling strategies, constructing multiple sub-graphs to respectively characterize binary relationships such as user-item and item-scene, and then perform feature integration through simple splicing or attention mechanisms. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a user portrait recommendation method based on support vector machine to solve the problem that traditional graph structures are difficult to effectively model the multi-dimensional coupling relationship between users, objects and scenes.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a user portrait recommendation method based on a support vector machine, comprising: obtaining user portrait recommendation data and preprocessing the data, wherein the user portrait recommendation data includes a user unique identifier, an item unique identifier, an operation type, a timestamp, an item category, an item price, an item description text, and a scene label;
[0008] Generate a hypergraph structured dataset using user profile recommendation data preprocessed by hypergraph association;
[0009] Convert hypergraph structured datasets into embedding vectors of users, items, and scenes;
[0010] The embedded vectors of users, items, and scenes are concatenated to generate an enhanced feature vector. The positive and negative samples in the log records are then used to train a support vector machine model to predict the probability of user-item interaction and generate a personalized recommendation list based on the user profile.
[0011] As a preferred solution of the user portrait recommendation method based on support vector machine of the present invention, wherein: the user portrait recommendation data after hypergraph association preprocessing is used to generate a hypergraph structure data set, specifically as follows:
[0012] Define user nodes, item nodes and scene nodes;
[0013] Connect the combination of users and multiple items under the same scene label as a hyperedge;
[0014] Determine the hyperedge weight based on the frequency of operation types and the importance of scene labels, and filter out low-frequency hyperedges;
[0015] Count the frequency of pre-processed operation types and generate user node feature vectors;
[0016] Process the item category field through one-hot encoding to generate a category feature vector;
[0017] The pre-trained word embedding model is used to process the item description text to generate word vectors, which are then concatenated with the category feature vector to form the item node feature vector.
[0018] Convert the scene label field into a scene node category feature vector through one-hot encoding;
[0019] Integrate user nodes, item nodes, scene nodes, hyperedges, hyperedge weights, user node feature vectors, item node feature vectors, and scene node category feature vectors to form a hypergraph structured dataset.
[0020] As a preferred solution of the user portrait recommendation method based on support vector machine described in the present invention, the conversion of the hypergraph structured data set into embedding vectors of users, items and scenes refers to processing the hypergraph structured data set through an input layer, two message passing layers and an output layer to generate embedding vectors of users, items and scenes, and capture the interactive relationship between user behavior patterns, item attributes and scene associations.
[0021] As a preferred solution of the user portrait recommendation method based on support vector machine of the present invention, wherein: the hypergraph structure data set is processed through the input layer, two message passing layers and output layer, specifically as follows:
[0022] The input layer inputs the user node feature vector, item node feature vector, scene node category feature vector and hyperedge weight in the hypergraph structure dataset;
[0023] The first message passing layer aggregates the user node feature vector, item node feature vector, and scene node category feature vector within the same hyperedge, and generates the first-layer intermediate feature vector through nonlinear transformation;
[0024] The second message passing layer receives the intermediate feature vectors of the first layer, aggregates the neighbor hyperedge features of the same user node in the hypergraph, and generates the second layer intermediate feature vectors through nonlinear transformation;
[0025] The output layer linearly transforms and compresses the intermediate feature vectors of the second layer to generate user embedding vectors, item embedding vectors, and scene embedding vectors.
[0026] As a preferred solution of the user portrait recommendation method based on support vector machine described in the present invention, the hyperedge feature is a weighted combination of the user node feature vector, item node feature vector and scene node category feature vector within the hyperedge.
[0027] As a preferred solution of the user portrait recommendation method based on support vector machine of the present invention, wherein: the training support vector machine model is as follows:
[0028] Extract positive and negative samples from log records and combine them with the enhanced feature vectors generated by concatenation to form a training set;
[0029] A kernel function is used to map the enhanced feature vectors in the training set into a high-dimensional space, and the user-item interaction pattern is learned based on positive and negative samples;
[0030] The classification hyperplane is optimized using the sequential minimal optimization algorithm.
[0031] As a preferred solution of the user portrait recommendation method based on support vector machine of the present invention, wherein: the log record is the user unique identifier, item unique identifier, operation type and timestamp in the user portrait recommendation data;
[0032] The positive and negative samples are determined based on the operation type;
[0033] The classification hyperplane is a dividing line that separates positive and negative samples in a high-dimensional space.
[0034] In a second aspect, the present invention provides a user portrait recommendation system based on support vector machine, comprising:
[0035] An acquisition module is used to acquire and preprocess user portrait recommendation data, wherein the user portrait recommendation data includes a user unique identifier, an item unique identifier, an operation type, a timestamp, an item category, an item price, an item description text, and a scene label;
[0036] The association module is used to generate a hypergraph structured dataset by using the user profile recommendation data preprocessed by the hypergraph association;
[0037] The conversion module is used to convert the hypergraph structure dataset into embedding vectors of users, items and scenes;
[0038] The prediction module is used to concatenate the embedding vectors of users, items, and scenes to generate an enhanced feature vector, and train a support vector machine model based on the positive and negative samples in the log records to predict the probability of user-item interaction and generate a personalized recommendation list based on the user profile.
[0039] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the user portrait recommendation method based on support vector machine as described in the first aspect of the present invention is implemented.
[0040] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the user portrait recommendation method based on support vector machine as described in the first aspect of the present invention.
[0041] The present invention achieves the following beneficial effects: By concatenating user, item, and scenario embedding vectors to generate an enhanced feature vector, which is then fed into a support vector machine model for training, the method achieves high-precision prediction of user-item interaction probabilities. By mapping features to a high-dimensional space using a radial basis function kernel, the separability of positive and negative samples is enhanced, improving the generalization capability of the support vector machine model. Platt scaling is used to output interaction probabilities, and scenario weights are introduced for dynamic adjustment, making recommendation results more tailored to actual scenario requirements. Compared to traditional models, the proposed method is more robust in processing high-dimensional sparse data, improving recommendation accuracy and personalization, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1Flowchart of the user profile recommendation method based on support vector machine.
[0044] Figure 2 Schematic diagram of the user portrait recommendation system based on support vector machine.
[0045] Figure 3 Construct a flowchart for a hypergraph.
[0046] Figure 4 Generate a neural network diagram for embeddings. DETAILED DESCRIPTION
[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0049] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0050] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a user portrait recommendation method based on a support vector machine, comprising the following steps:
[0051] S1. Obtain user portrait recommendation data and preprocess it.
[0052] Furthermore, a data collection program is run on the user interaction platform to query the user interaction platform's log records, item attribute database, and application scenario labeling rules to collect user portrait recommendation data, as follows:
[0053] Query the log records of the user interaction platform to extract the user unique identifier, item unique identifier, operation type (click, purchase, and favorite) and timestamp;
[0054] Query the item attribute database to extract item category, item price and item description text;
[0055] Apply scenario labeling rules to generate scenario labels; the scenario labeling rules include a timestamp rule: determining Monday to Friday as workdays and Saturday to Sunday as weekend shopping based on the timestamp; a promotion calendar rule: determining promotion activity dates based on the promotion calendar; and a user location rule: inferring a home scenario based on the user's location being the home address.
[0056] Combine the user unique identifier, item unique identifier, operation type, timestamp, item category, item price, item description text, and scene label to form user portrait recommendation data; the user unique identifier reflects the user identity, the item unique identifier reflects the recommended object, the operation type reflects the interaction behavior, the timestamp reflects the interaction time, the item category reflects the recommendation level, the item price reflects the economic characteristics, the item description text reflects the content characteristics, and the scene label reflects the context environment;
[0057] The collected user profile recommendation data is transmitted to the central control center via the Internet protocol;
[0058] After receiving the user portrait recommendation data, the central control center standardizes the item price, item description text, and operation type frequency to ensure that the dimensions of the three indicators are consistent; among them, the item price and operation type frequency are standardized through Z-Score (Z-score normalization), and the item description text is first processed through a pre-trained word embedding model and then standardized through Z-Score; the word embedding model is an existing model, such as the Word2Vec model, which is a model well known to technicians in this field and is widely used in the field of natural language processing; the operation type frequency is determined based on the sum of the user's clicks, purchases, and favorites.
[0059] It should be noted that the pre-trained Word2Vec model is used to process the item description text as follows:
[0060] Enter the item description text (e.g. "Smartphone, 8GB RAM, black");
[0061] Use a word segmentation tool (such as Jieba or NLTK) to split the item description text into words (such as "smartphone", "8GB", "RAM", "black");
[0062] For each word, query the Word2Vec model's vocabulary to obtain a 300-dimensional word vector. If the word is not in the Word2Vec model's vocabulary (e.g., "8GB"), ignore the word that is not in the vocabulary. Output the queried 300-dimensional word vector set.
[0063] The average of all the 300-dimensional word vectors found is taken to generate a 300-dimensional text vector; if no 300-dimensional word vector is found (that is, all words are not in the vocabulary of the Word2Vec model), then a 300-dimensional 0 text vector is generated (that is, [0, 0, ..., 0] to ensure that subsequent steps such as Z-Score normalization and hypergraph construction do not report errors due to the lack of 300-dimensional text vectors).
[0064] S2. Generate a hypergraph structured dataset by using the user portrait recommendation data preprocessed by hypergraph association.
[0065] Based on the standardized user profile recommendation data, a hypergraph is constructed to associate the user profile recommendation data, as follows:
[0066] Define user nodes based on the user's unique identifier;
[0067] Define item nodes based on the item’s unique identifier;
[0068] Define scene nodes based on scene tags;
[0069] A hyperedge connects a combination of a user's unique identifier and multiple item unique identifiers under the same scenario label. For example, if user 001 clicks on item 101 and buys item 102 in the weekend shopping scenario, the generated hyperedge includes user 001, item 101, item 102, and the scenario label "weekend shopping";
[0070] The hyperedge weight is calculated by multiplying the frequency of the operation type and the importance of the scene label, and is expressed as:
[0071] W e =F o ×I s ;
[0072] Among them, W e is the hyperedge weight, which indicates the strength of the hyperedge and reflects the importance of the user's interaction with the item in a specific scenario. The larger the value, the higher the influence of the hyperedge in the hypergraph structure dataset. e is the index of the hyperedge, and F o is the frequency of the operation type, o is the index of the operation type, I s The importance of the scene tag is based on the preset weight of the recommendation target, ranging from 0.5 to 2.0. It is stored in the configuration file of the central control center and reflects the relative importance of different scene tags (such as promotions and weekend shopping) to the recommendation. For example, if the importance of promotions is 1.5 and that of weekend shopping is 1.0, then promotions will be prioritized. s is the index of the scene tag.
[0073] Filter hyperedge weights and remove hyperedges with operation type frequencies less than 2 to reduce computational costs;
[0074] Extract click, purchase, and favorite frequencies from the standardized operation type frequencies, and directly combine click, purchase, and favorite frequencies to generate a three-dimensional feature vector for the user node;
[0075] The item category field in the user profile recommendation data is converted into a category feature vector of the item node through one-hot encoding. The conversion is specifically as follows: a category set is extracted from the item attribute database, for example, 5 categories, including electronic products, clothing, books, food, and furniture; a 5-dimensional category feature vector is constructed for each item, with each dimension corresponding to a category; if the item is an electronic product, the electronic product dimension is set to 1, and the clothing, book, food, and furniture dimensions are set to 0; if the item is clothing, the clothing dimension is set to 1, and the electronic products, books, food, and furniture dimensions are set to 0; among them, one-hot encoding is a conventional technique widely used in machine learning, data processing and other fields;
[0076] The standardized item description text is processed through a pre-trained word embedding model (Word2Vec model) to generate word vectors for item nodes;
[0077] Directly concatenate the category feature vector and word vector of the item node to generate a node feature vector of the item node, for example, 305-dimensional. The category feature vector of the item node is a 5-dimensional vector generated based on the one-hot encoding of the item category, indicating the category affiliation. The node feature vector of the item node is a 305-dimensional comprehensive vector obtained by concatenating the category feature vector and the 300-dimensional word vector, indicating the complete attributes of the item.
[0078] The scene label field in the user portrait recommendation data is converted into a category feature vector of the scene node through one-hot encoding. The conversion is as follows: a scene label set is extracted from the scene label rule, for example, 4 categories, including weekdays, weekend shopping, promotions, and family scenes; a 4-dimensional category feature vector is constructed for each scene label, with each dimension corresponding to a scene; if it is a weekday, the weekday dimension is set to 1, and the weekend shopping, promotions, and family scene dimensions are set to 0; if it is weekend shopping, the weekend shopping dimension is set to 1, and the weekday, promotions, and family scene dimensions are set to 0;
[0079] Based on the three-dimensional feature vector of user nodes, the node feature vector of item nodes, the category feature vector of scene nodes, user nodes, item nodes, scene nodes, hyperedges and hyperedge weights, a hypergraph is constructed to generate a hypergraph structure dataset.
[0080] It should be noted that this step effectively integrates the multi-dimensional information of users, items, and scene labels by constructing a hypergraph structured dataset. The design of hyperedges can not only represent the complex interactive relationships between users and multiple items in specific scenarios, but also highlight the importance of high-value interactions by combining the frequency of operation types and the importance of scenarios to calculate weights; removing low-frequency interactions reduces computational costs while maintaining core related information. In addition, by applying one-hot encoding to item categories and combining pre-trained word embedding models to process item description text to generate comprehensive node feature vectors, item attributes are more comprehensively expressed; scene labels are also converted into feature vectors through one-hot encoding, enhancing the understanding of the context; finally, the completed hypergraph construction provides a high-quality data foundation for user portrait recommendations, which helps to improve the accuracy and relevance of personalized recommendations.
[0081] S3. Convert the hypergraph structure dataset into embedding vectors of users, items, and scenes.
[0082] Furthermore, the central control center receives the hypergraph structured data set and constructs a hypergraph neural network model based on the hypergraph structured data set;
[0083] The hypergraph neural network model is based on a multi-layer neural network model, which includes an input layer, two message passing layers, and an output layer to generate a comprehensive set of embedding vectors.
[0084] The input layer receives the node feature vectors (three-dimensional feature vectors of user nodes, node feature vectors of item nodes, and category feature vectors of scene nodes) and hyperedge weights of the hypergraph structure dataset;
[0085] The first message passing layer aggregates the node feature vectors of user nodes, item nodes, and scene nodes within the same hyperedge in the hypergraph. For example, in the hyperedge (user 001, item 101, item 102, and weekend shopping), the node feature vectors of item 101, item 102, and the scene node "weekend shopping" are weighted and aggregated to user 001, generating a first-layer 64-dimensional intermediate feature vector. The hyperedge feature is a weighted combination of the user node feature vector, item node feature vector, and scene node category feature vector within the hyperedge.
[0086] The second message passing layer receives the 64-dimensional intermediate feature vector of the first layer, aggregates the neighbor hyperedge features of the same user node in the hypergraph, enhances the context information, and generates the second layer 64-dimensional intermediate feature vector;
[0087] Each message passing layer uses the ReLU activation function for nonlinear transformation to ensure the expressiveness of features;
[0088] The output layer linearly transforms and compresses the 64-dimensional intermediate feature vectors of the second layer to generate 32-dimensional user embedding vectors for user nodes (based on log records of clicks, purchases, and collections, encoding user behavior preferences, such as Aidian electronic products), 32-dimensional item embedding vectors for item nodes (based on the co-occurrence relationship between the category, price, description text and log records in the item attribute database, encoding item attributes and interaction relationships, such as mobile phones are often selected together with headphones), and 32-dimensional scene embedding vectors for scene nodes (based on log records of scene labels, encoding scene associations, such as weekend shopping), which are used to recommend matching users and items.
[0089] The log records in the user profile recommendation data are used as training samples for the hypergraph neural network model. The <user, item> corresponding to the operation type is extracted from the log records as a positive sample with a label of 1. The <user, item> that is not in the log record is randomly sampled from the user unique identifier and the item unique identifier as a negative sample with a label of 0.
[0090] The training samples are divided into 90% training set and 10% validation set. The training set is used to optimize the parameters of the hypergraph neural network model, and the validation set is used to evaluate the generalization ability of the hypergraph neural network model to prevent overfitting. The details are as follows:
[0091] The training set is input into the hypergraph neural network model for training. First, the hypergraph neural network model receives the <user, item> of the training set samples through the input layer and processes them through the message passing layer and output layer. The 32-dimensional user embedding vector and the 32-dimensional item embedding vector are multiplied bit by bit and summed to generate a user-item interaction score. Then, the user-item interaction score is mapped to an interaction probability using a sigmoid function. Finally, the cross-entropy loss function is used to compare the interaction probability with the true interaction label of the training set samples (positive samples are 1, negative samples are 0) to generate a loss value and optimize the hypergraph neural network model parameters.
[0092] The hypergraph neural network model parameters are optimized through mini-batch gradient descent, processing 128 training samples per batch for a maximum of 20 epochs. When the cross-entropy loss of the validation set stops decreasing for several consecutive epochs, training is stopped early. The optimal hypergraph neural network model parameters are selected so that the 32-dimensional user embedding vector encodes user behavior patterns (based on the clicks, purchases, and favorites in the log records, reflecting user interaction preferences, such as I love electronic products), and the 32-dimensional item embedding vector encodes item attributes and interaction relationships (based on the item category, price, description text in the item attribute database and the co-occurrence relationship between items in the log records, such as mobile phones and headphones are often selected together).
[0093] For new users without log records, operation type frequencies are generated based on external data (such as registration information). Cosine similarity is calculated based on the operation type frequencies of existing users in the log records to determine the similarity between the new user and existing users. The 32-dimensional user embedding vectors of the existing users are weighted averaged to generate a 32-dimensional user embedding vector for the new user. Chordal similarity is a commonly used standard calculation method that measures similarity by calculating the cosine value of the angle between two vectors. The closer the value is to 1, the more similar the two are. This is a technique well known in the art.
[0094] Based on the trained hypergraph neural network model, the central control center generates a comprehensive embedding vector set, which includes 32-dimensional user embedding vectors of all user nodes, 32-dimensional embedding vectors of item nodes, and 32-dimensional embedding vectors of scene nodes.
[0095] It should be noted that this step significantly improves the understanding and expression of user behavior patterns, item attributes, and scene-related interactions by converting the hypergraph structured dataset into embedding vectors of users, items, and scenes. This conversion enables each entity (user, item, and scene) to be accurately represented in the form of a low-dimensional vector, which not only captures individual characteristics but also reveals the complex interactions between each entity. In addition, the frequency of operation types is combined with the importance of scenes to strengthen the importance of key interactions, and the richness of feature expression is ensured through nonlinear transformations. For the processing of new users, cosine similarity is used to calculate the similarity with existing users, and then a reliable embedding vector is generated, which effectively solves the cold start problem. Overall, this step improves the depth of understanding of user preferences, optimizes the effect of personalized recommendations, makes the recommendation results more in line with the user's real needs and interests, enhances the user experience, and improves the quality and accuracy of recommendations.
[0096] S4. Concatenate the embedding vectors of users, items, and scenes to generate an enhanced feature vector. Combine the positive and negative samples in the log records to train a support vector machine model to predict the probability of user-item interaction and generate a personalized recommendation list for user profiles.
[0097] Furthermore, the 32-dimensional user embedding vector and the 32-dimensional item embedding vector are directly combined into a 64-dimensional user-item feature vector;
[0098] The 32-dimensional scene embedding vector is appended to the 64-dimensional user-item feature vector as the context feature of the 64-dimensional user-item feature vector to form a 96-dimensional enhanced feature vector to capture the impact of the interaction scene;
[0099] Positive and negative samples are extracted from log records and directly combined with 96-dimensional enhanced feature vectors to generate the training set of the support vector machine model.
[0100] Train the support vector machine model classifier using the training set as follows:
[0101] The kernel function of the radial basis function is used to map the 96-dimensional enhanced feature vector to a high-dimensional space, making it easier to separate positive and negative samples;
[0102] The user-item similarity is calculated by the radial basis function inner product (e.g., the similarity between the feature vector of user 001 and item 101 and the positive sample), providing a distance metric for the support vector machine model classifier to distinguish between interactions (positive samples, label 1) and no interactions (negative samples, label 0). The similarity expression is:
[0103] K(x,x′)=exp(-γ||xx′|| 2 );
[0104] Among them, K(x,x′) is the user-item similarity, that is, the kernel function value between two 96-dimensional enhanced feature vectors, x is the 96-dimensional enhanced feature vector of the current user-item, such as the feature vector of user 001 and item 101, x′ is the 96-dimensional enhanced feature vector of another user-item, usually a positive or negative sample in the training set, exp is an exponential function (based on the natural logarithm base e), which converts the input value to a value between 0 and 1, γ is the parameter of the radial basis function, that is, the kernel parameter, which is a positive number (γ>0) that controls the decay rate of the similarity, ||xx′|| 2 is the square of the Euclidean distance between the two 96-dimensional enhanced feature vectors x, x′, indicating the degree of difference between the two 96-dimensional enhanced feature vectors;
[0105] It should be noted that this formula is based on the radial basis function, relies on Mercer's theorem, and is widely used in interactive prediction in the recommendation field.
[0106] The sequential minimum optimization algorithm is used to iteratively solve the support vector and weight of the support vector machine model classifier, and optimize the hyperplane of the support vector machine model classifier (the dividing line that separates positive and negative samples in high-dimensional space), as follows:
[0107] Randomly select some training samples from the training set as initial support vectors;
[0108] Set the initial weight to zero, the penalty parameter (which determines the tolerance of the support vector machine model classifier to classification errors and is usually initially set to 1), and the radial basis function kernel parameter (which is initially set to 1 / 96 and determines the strictness of similarity judgment based on the 96-dimensional enhanced feature vector) as preliminary parameters of the support vector machine model classifier;
[0109] Check whether each training sample in the training set meets the classification conditions, that is, whether the training sample meets the optimization requirements of the hyperplane classification, such as being misclassified, close to the hyperplane, or not reaching the maximum interval;
[0110] Prioritize training samples that seriously violate the classification conditions (such as positive or negative samples that are misclassified, or training samples close to the hyperplane). Determine candidate support vectors by calculating the similarity between the training samples that seriously violate the classification conditions and the current support vector samples (training samples close to the hyperplane) (high similarity means close to the hyperplane).
[0111] Select training sample point pairs from the candidate support vectors (two training samples, which can be two positive samples, two negative samples, or one positive sample and one negative sample). Adjust the weights of the support vector machine model classifier to make the hyperplane classification boundary of the support vector machine model classifier more accurate. Optimize the position of the hyperplane to ensure that positive and negative samples are separated as much as possible in high-dimensional space. Learn user-item interaction patterns (using the 96-dimensional enhanced feature vectors and positive and negative labels of positive and negative samples to identify the interaction between user preferences and item attributes, such as users who like to click on electronic products are more likely to click on mobile phones).
[0112] Positive and negative samples were randomly extracted from log records and directly combined with the 96-dimensional augmentation vector to generate a validation set for the SVM model, independent of the SVM training set. The SVM classifier's classification accuracy was evaluated. Grid search was used to adjust the penalty parameter and the kernel parameter of the radial basis function to select the optimal parameters and ensure generalization. Grid search is a standard parameter optimization technique in machine learning and recommendation, and is widely used in parameter adjustment for SVM models.
[0113] Output the optimized support vector machine model classifier to ensure that positive and negative samples are separated as much as possible in high-dimensional space and learn user-item interaction patterns.
[0114] For new users without log records or existing users without interaction with user-item combinations, a 96-dimensional enhanced feature vector is generated;
[0115] The support vector machine model classifier uses the hyperplane generated by training and the kernel function of the radial basis function to calculate the distance between the 96-dimensional enhanced feature vector and the hyperplane (positive values indicate a tendency towards interaction, and negative values indicate a tendency towards no interaction) and outputs the classification score;
[0116] The classification scores are converted into interaction probabilities (0 to 1, e.g., the probability of user 001 for item 101 is 0.9) through Platt scaling. The interaction probabilities are then weighted and adjusted based on the 32-dimensional scene embedding vector (e.g., weekend shopping increases the probability of promotional items, with a 10% weight increase) to generate the final interaction probability. Platt scaling is an existing technique in machine learning and recommendation, and is widely used in probability conversion for support vector machine models.
[0117] According to the final interaction probability, all items of each user are sorted in descending order to generate a personalized recommendation list for the user portrait (for example, recommending mobile phones and headphones to user 001 who loves electronic products, and intercepting the top 10 high-probability items).
[0118] It should be noted that this step generates a 96-dimensional enhanced feature vector by concatenating a 32-dimensional user embedding vector, a 32-dimensional item embedding vector, and a 32-dimensional scene embedding vector. This vector comprehensively captures the influence of user behavior patterns, item attributes, and scene associations. The 96-dimensional enhanced feature vector is then combined with positive and negative samples to train a support vector machine model, and a radial basis function kernel is employed to enhance the discriminability of features in high-dimensional space, ensuring accurate modeling of user-item interaction patterns. The optimal hyperplane is solved using a sequential minimum optimization algorithm, and Platt scaling is used to convert classification scores into interaction probabilities. These probabilities are then adjusted based on scene importance, ultimately generating a personalized recommendation list based on interaction probability sorting. This step significantly improves the relevance of recommendation results and user experience, making recommendations more tailored to users' actual interests and current contextual needs.
[0119] This embodiment also provides a user portrait recommendation system based on a support vector machine, including:
[0120] The acquisition module is used to obtain and preprocess user portrait recommendation data. The user portrait recommendation data includes user unique identifier, item unique identifier, operation type, timestamp, item category, item price, item description text and scene label;
[0121] The association module is used to generate a hypergraph structured dataset by using the user profile recommendation data preprocessed by the hypergraph association;
[0122] The conversion module is used to convert the hypergraph structure dataset into embedding vectors of users, items and scenes;
[0123] The prediction module is used to concatenate the embedding vectors of users, items, and scenes to generate an enhanced feature vector, and train a support vector machine model based on the positive and negative samples in the log records to predict the probability of user-item interaction and generate a personalized recommendation list based on the user profile.
[0124] This embodiment also provides a computer device, which is suitable for the user portrait recommendation method based on a support vector machine, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the user portrait recommendation method based on a support vector machine proposed in the above embodiment.
[0125] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0126] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the user portrait recommendation method based on a support vector machine as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0127] In summary, this method achieves high-precision prediction of user-item interaction probabilities by concatenating user, item, and scene embedding vectors to generate enhanced feature vectors, which are then fed into a support vector machine model for training. Using a radial basis function kernel to map features to a high-dimensional space enhances the separability of positive and negative samples and improves the generalization capability of the support vector machine model. Platt scaling is used to output interaction probabilities, and scene weights are introduced for dynamic adjustment, making recommendation results more tailored to actual scenario requirements. Compared to traditional models, this method is more robust in processing high-dimensional sparse data, improving recommendation accuracy and personalization, and enhancing the user experience.
[0128] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A user profile recommendation method based on support vector machine, characterized by: include, Obtain and preprocess user portrait recommendation data, wherein the user portrait recommendation data includes a user unique identifier, an item unique identifier, an operation type, a timestamp, an item category, an item price, an item description text, and a scene label; Generate a hypergraph structured dataset using user profile recommendation data preprocessed by hypergraph association; Convert hypergraph structured datasets into embedding vectors of users, items, and scenes; The embedded vectors of users, items, and scenes are concatenated to generate an enhanced feature vector. The positive and negative samples in the log records are then used to train a support vector machine model to predict the probability of user-item interaction and generate a personalized recommendation list based on the user profile.
2. The user profile recommendation method based on support vector machine according to claim 1, characterized in that: The user portrait recommendation data after hypergraph association preprocessing is used to generate a hypergraph structure dataset, as follows: Define user nodes, item nodes and scene nodes; Connect the combination of users and multiple items under the same scene label as a hyperedge; Determine the hyperedge weight based on the frequency of operation types and the importance of scene labels, and filter out low-frequency hyperedges; Count the frequency of pre-processed operation types and generate user node feature vectors; Process the item category field through one-hot encoding to generate a category feature vector; The pre-trained word embedding model is used to process the item description text to generate word vectors, which are then concatenated with the category feature vector to form the item node feature vector. Convert the scene label field into a scene node category feature vector through one-hot encoding; Integrate user nodes, item nodes, scene nodes, hyperedges, hyperedge weights, user node feature vectors, item node feature vectors, and scene node category feature vectors to form a hypergraph structured dataset.
3. The user profile recommendation method based on support vector machine according to claim 1, characterized in that: The conversion of a hypergraph structured dataset into embedding vectors of users, items, and scenes refers to processing the hypergraph structured dataset through an input layer, two message passing layers, and an output layer to generate embedding vectors of users, items, and scenes, thereby capturing the interactive relationships among user behavior patterns, item attributes, and scene associations.
4. The user profile recommendation method based on support vector machine according to claim 3, characterized in that: The hypergraph structure dataset is processed through the input layer, two message passing layers and output layer, as follows: The input layer inputs the user node feature vector, item node feature vector, scene node category feature vector and hyperedge weight in the hypergraph structure dataset; The first message passing layer aggregates the user node feature vector, item node feature vector, and scene node category feature vector within the same hyperedge, and generates the first-layer intermediate feature vector through nonlinear transformation; The second message passing layer receives the intermediate feature vectors of the first layer, aggregates the neighbor hyperedge features of the same user node in the hypergraph, and generates the second layer intermediate feature vectors through nonlinear transformation; The output layer linearly transforms and compresses the intermediate feature vectors of the second layer to generate user embedding vectors, item embedding vectors, and scene embedding vectors.
5. The user profile recommendation method based on support vector machine according to claim 4, characterized in that: The hyperedge feature is a weighted combination of the user node feature vector, the item node feature vector, and the scene node category feature vector within the hyperedge.
6. The user profile recommendation method based on support vector machine according to claim 1, characterized in that: The training support vector machine model is as follows: Extract positive and negative samples from log records and combine them with the enhanced feature vectors generated by concatenation to form a training set; A kernel function is used to map the enhanced feature vectors in the training set into a high-dimensional space, and the user-item interaction pattern is learned based on positive and negative samples; The classification hyperplane is optimized using the sequential minimal optimization algorithm.
7. The user profile recommendation method based on support vector machine according to claim 6, characterized in that: The log record is the user unique identifier, item unique identifier, operation type and timestamp in the user portrait recommendation data; The positive and negative samples are determined based on the operation type; The classification hyperplane is a dividing line that separates positive and negative samples in a high-dimensional space.
8. A user portrait recommendation system based on a support vector machine, based on the user portrait recommendation method based on a support vector machine according to any one of claims 1 to 7, characterized in that: include, An acquisition module is used to acquire and preprocess user portrait recommendation data, wherein the user portrait recommendation data includes a user unique identifier, an item unique identifier, an operation type, a timestamp, an item category, an item price, an item description text, and a scene label; The association module is used to generate a hypergraph structured dataset by using the user profile recommendation data preprocessed by the hypergraph association; The conversion module is used to convert the hypergraph structure dataset into embedding vectors of users, items and scenes; The prediction module is used to concatenate the embedding vectors of users, items, and scenes to generate an enhanced feature vector, and train a support vector machine model based on the positive and negative samples in the log records to predict the probability of user-item interaction and generate a personalized recommendation list based on the user profile.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the user portrait recommendation method based on support vector machine described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the user portrait recommendation method based on support vector machine according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Order pushing method, device and system
CN105894359A
Information recommendation method and device, server and readable storage medium
CN111274472A
Sequence recommendation method and device based on hypergraph neural network
CN115082147A
Information recommendation method and device based on portrait system, electronic equipment and medium
CN117372094A
Advertisement marketing recommendation method based on deep reinforcement learning
CN119151616A