An E-commerce Personalized Recommendation Service Method Based on Genetic Fuzzy Clustering
By genetically fuzzy clustering of products and user groups on e-commerce platforms according to color richness, a personalized recommendation list is generated, and dynamically adjusted based on user feedback, the problem of poor user experience in the existing technology is solved, and more accurate and satisfactory product recommendations are achieved.
Patent Information
- Application Number
- CN202411190700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-28
AI Technical Summary
Existing intelligent recommendation solutions for e-commerce products rarely consider users' psychological hints, which leads to poor experience when choosing products. Especially for users who prefer monochromatic or low color richness, excessive colors may cause visual fatigue and confusion.
By converting product images into RGB color space, using the genetic fuzzy clustering algorithm to divide the user group into multiple categories according to color richness, establish a preference prediction model, generate a personalized recommendation list, and dynamically adjust the recommendation strategy through user feedback to improve the accuracy of recommendations and user satisfaction.
It enhances targeted recommendations from the user group, improves the accuracy and user satisfaction of recommendations, ensures the timeliness and personalization of the recommendation list, and improves the user's experience in selecting products.
Smart Images

Figure CN119090588B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to an e-commerce personalized recommendation service method based on genetic fuzzy clustering. Background Art
[0002] In e-commerce platforms, the number of colors of products has an important impact on users' purchase intention; products with a rich color quantity are more likely to attract users who prefer multi-color products, but for users who prefer single-color or low-color-rich products, too many colors may cause visual fatigue and confusion, reducing their purchase intention.
[0003] Existing intelligent recommendation solutions for e-commerce products rarely consider the psychological implications of colors for users, resulting in a poor experience for users in the process of selecting required products. Summary of the Invention
[0004] Based on this, it is necessary to provide an e-commerce personalized recommendation service method based on genetic fuzzy clustering for the above technical problems.
[0005] An e-commerce personalized recommendation service method based on genetic fuzzy clustering provided by this application includes:
[0006] Convert the product images corresponding to each product obtained and loaded into the RGB color space, extract the colors of the product images, and divide the corresponding products into single-color, low-color-richness, medium-color-richness, and high-color-richness according to the number of colors of the product images from low to high in terms of color richness;
[0007] Based on the basic information data of each user group obtained and the behavioral data of the corresponding user group, analyze the preference relationship between the user group and the color richness of the product, and establish a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group; the user group includes users of different genders, occupations, and ages; the behavioral data includes the browsing records, purchase records, and collection records of the user group for products;
[0008] According to the behavioral data of the user group and the preference feature vector, use the fuzzy clustering algorithm to analyze the differences and similarities of different user groups in terms of color richness, and use the genetic algorithm to optimize the fuzzy clustering algorithm;
[0009] Based on the analysis results of the optimized fuzzy clustering algorithm, divide the user group into multiple categories of users according to color richness, and obtain the color preference characteristics of each category of users; generate a recommended list of products based on color richness according to the color preference characteristics and the products corresponding to the corresponding color richness;
[0010] Collect and dynamically adjust the recommendation list based on the feedback information of users in each category to optimize the color preference features.
[0011] The above-mentioned e-commerce personalized recommendation service method based on genetic fuzzy clustering can first divide commodities into multiple categories of color richness by converting commodities into the RGB color space, then obtain the behavior data of each user group to get the preference relationship between user behavior and color richness, establish a prediction model according to the preference relationship, convert the user preferences into feature vectors to obtain structured data, and then further divide the user groups into categories through the fuzzy clustering algorithm optimized by the genetic algorithm, enhance the pertinence of the user groups, and associate with the categories of color richness. Finally, generate a recommendation list according to the association results, recommend commodities to the user groups specifically, and dynamically adjust the recommendation list in real time to maintain the timeliness of the recommendation list and improve the accuracy and user satisfaction of the recommendation, effectively solving the problem that the existing intelligent recommendation solutions for e-commerce products rarely consider the psychological hint of color to users, resulting in a poor experience for users in the process of selecting required commodities. Brief Description of the Drawings
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic flowchart of an e-commerce personalized recommendation service method based on genetic fuzzy clustering in an embodiment. Detailed Embodiments
[0014] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0015] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0016] In this application, unless otherwise clearly specified and limited, the terms "initial", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0017] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element, or connected to the other element through an intermediate element. In addition, the "connection" in the following embodiments should be understood as "electrical connection", "communication connection", etc. if there is transmission of electrical signals or data between the connected objects.
[0018] Users of different occupations, genders, and ages have different preferences for the number of colors. Therefore, how to recommend suitable color quantity products to users based on their personal characteristics has become a technical problem that e-commerce platforms need to solve. On the basis of collecting basic information and behavioral data of users, establishing a correlation model between gender, occupation, age and product color quantity preferences is the key to realizing personalized color quantity product recommendations. However, in the modeling process, how to select a suitable algorithm and optimize the performance of the algorithm to achieve efficient classification of user color quantity preference characteristics is a technical problem that needs to be solved urgently. At the same time, after the model is established, how to use historical data and feedback to verify and adjust the model to ensure the accuracy of the recommendation is also an issue that needs to be explored in depth. In addition, in the actual recommendation process, how to design a reasonable color quantity grading scheme and provide corresponding product recommendations based on the grading is also a technical issue worth thinking about.
[0019] In an exemplary embodiment, Figure 1 As shown, a method for e-commerce personalized recommendation service based on genetic fuzzy clustering is provided, including the following steps 202 to 210. Among them:
[0020] Step 202, converting the acquired and loaded product images corresponding to the products into RGB color space, extracting the colors of the product images using a fuzzy clustering algorithm, and dividing the corresponding products into monochrome, low color richness, medium color richness, and high color richness according to the number of colors of the product images from low to high in terms of color richness.
[0021] Specifically, first load the product image and convert it to the RGB color space. According to business requirements and data analysis, set the threshold for classifying color richness. Select the threshold based on product categories and color preference factors to ensure the rationality and practicality of the classification results.
[0022] In an exemplary embodiment, the steps of using a fuzzy clustering algorithm to extract the colors of the product image and classifying the corresponding products into monochromatic, low color richness, medium color richness, and high color richness according to the number of colors in the product image from low to high color richness include:
[0023] Process the converted RGB color space using a fuzzy clustering algorithm; the fuzzy clustering algorithm is the k-means fuzzy clustering algorithm; allocate the data points in the RGB color space to K cluster centers in an iterative manner and adjust the distance from each data point to the cluster center to be minimized;
[0024] Cluster the pixel points of the product image into several representative colors to obtain the color distribution of the product image with respect to the representative colors;
[0025] Judge that the color richness of products with the number of colors in the color distribution below the first preset threshold is monochromatic, the color richness of products with the number of colors in the color distribution exceeding the first preset threshold and below the second preset threshold is low color richness, the color richness of products with the number of colors in the color distribution exceeding the second preset threshold and below the third preset threshold is medium color richness, and the color richness of products with the number of colors in the color distribution exceeding the third preset threshold is high color richness; among them, the first preset threshold, the second preset threshold, and the third preset threshold are all set according to the pre-stored business requirements and business data.
[0026] In the specific implementation process, an image processing library such as OpenCV can be used to load the product image, and further use the cvtColor function in OpenCV to convert the image from the BGR color space (or other types of color spaces) to the RGB color space. Then, according to business requirements and data analysis, the threshold for classifying color richness can be set. For example, the first preset threshold can be set to 5, the second preset threshold can be set to 10, and the third preset threshold can be set to 20. More than 20 is defined as high color richness. Among them, the first preset threshold, the second preset threshold, and the third preset threshold can be adjusted according to the actual situation.
[0027] Further, during the color clustering process, the K-means algorithm can be used to cluster the pixel points in the RGB color space by setting the number of clustering centers K to 10, the number of iterations to 100, and the initial center point selection method to random selection. The clustering result can be obtained by calculating the number of pixel points of each clustering center to obtain several representative color numbers in the image.
[0028] Specifically, the number of representative colors can be set to the number of main colors of the seven-color light.
[0029] Further, to evaluate the classification effect, 100 product images can be randomly selected and manually judged by multiple professional designers, and then compared with the algorithm classification results. By calculating indicators such as accuracy and recall rate, if the accuracy rate reaches more than 90% and the recall rate reaches more than 95%, it indicates that the classification method for color richness is feasible.
[0030] Step 204, based on the obtained basic information data of each user group and the behavior data of the corresponding user group, use the association rule mining algorithm to analyze and obtain the preference relationship between the user group and the color richness of the product, and establish a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group; the user group includes users of different genders, occupations and ages; the behavior data includes the browsing records, purchase records and collection records of the user group for the product.
[0031] Specifically, the basic information data collected includes users of the same gender, age and occupation who form a user group, and these data are cleaned and preprocessed, and converted into a structured feature vector to obtain user features. Obtain the behavior data of the user on the e-commerce platform, including browsing records, purchase records, and collection records. Perform feature selection on the user features and behavior data, and select the features that have an impact on color preference prediction as the preferred user features to improve the effect of the model. For each user in the user group, count the number and duration of their interaction behaviors on products with different color richness as the preference measure for the corresponding color richness of the user. Normalize the measurement values of these preference measures to obtain the preference distribution of the user for different color richness, and then use the association rule mining algorithm, such as the Apriori algorithm or the FP-growth algorithm, to perform correlation analysis on the user feature vector and the color preference data, and discover the association rules between different feature combinations and color preferences. According to the mined association rules, construct a user color richness preference prediction model to obtain the preference feature vector of the user group.
[0032] In an exemplary embodiment, the steps of using the association rule mining algorithm to analyze and obtain the preference relationship between the user group and the color richness of the product, and establishing a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group include:
[0033] Convert the basic information data into a structured feature vector to obtain user features;
[0034] Perform feature selection on the user features according to the behavior data of user groups to obtain preference user features associated with the behavior data;
[0035] Statistically analyze the interaction behaviors of different user groups on products with different color richness levels respectively, as the preference measurement of user groups for color richness; Interaction behaviors include the number and duration of browsing records, purchase records, and collection records generated for products;
[0036] Normalize each preference measurement to obtain the preference distribution of different user groups for color richness;
[0037] Use the association rule mining algorithm to perform a correlation analysis on the preference user features and preference relationships to obtain the association rules between different feature combinations and color preferences; Different feature combinations include color combinations of color richness; The association rule mining algorithm includes the Apriori algorithm or the FP-growth algorithm;
[0038] Construct a preference prediction model based on the association rules between different feature combinations and color preferences to obtain the preference feature vector of the user group, and use the collaborative filtering algorithm and matrix factorization algorithm to optimize the preference prediction model;
[0039] Use the cross-validation method to evaluate the preference prediction model, calculate the accuracy rate, recall rate, and F1 metric of the preference prediction model, and adjust the parameters of the preference prediction model and perform feature selection of user features according to the evaluation results.
[0040] In an exemplary embodiment, the steps of optimizing the preference prediction model using the collaborative filtering algorithm and matrix factorization algorithm include:
[0041] Use the collaborative filtering algorithm to construct an interaction matrix between user groups and products, and use the matrix factorization algorithm to calculate the interaction matrix to obtain a similarity matrix; Find the similar users of the target user in the set user group according to the similarity matrix, and calculate the preference threshold of the similar users for the products corresponding to each color richness level as the preference prediction value of the target user; Train the preference prediction model according to the basic information data of the target user and the preference prediction value of the target user.
[0042] In an exemplary embodiment, the steps of using the association rule mining algorithm to analyze and obtain the preference relationship between the user group and the color richness of products, and establishing a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group further include:
[0043] When the target user is a new user, based on the basic information data of the new user, find similar users through a preference prediction model, calculate the average value of the preference prediction values of the similar users, and use it as the initial preference prediction value of the new user;
[0044] When the new user generates interaction behaviors, update the preference prediction value according to the behavior data of the new user; introduce the color richness of the commodity, and generate an initial recommendation list for the new user through a recommendation method based on the content of the commodity to promote the interaction behaviors of the new user.
[0045] Specifically, use collaborative filtering to construct a user-item interaction matrix, calculate the similarity matrix between users or between items, find the K most similar users to the target user according to the similarity matrix, and calculate the average preference value of these users for each commodity with high color richness as the preference prediction value of the target user.
[0046] Furthermore, for a new user, use its basic information data to find the user group most similar to its feature combination according to the association rules, and use the average preference of this group as the initial preference prediction of the new user. After the new user generates certain interaction behaviors, gradually add its behavior data to update the preference prediction result.
[0047] Furthermore, introduce the content feature of the color quantity of the commodity, generate an initial recommendation list for the new user through a content-based recommendation method to promote the interaction between the new user and the system. Evaluate and verify the effect of the prediction model, use the cross-validation method, calculate the accuracy rate, recall rate, and F1 value index (F1 score) of the model, and adjust the model parameters and feature selection according to the evaluation results to improve the prediction performance. Apply the optimized prediction model to the actual commodity recommendation and personalized display scenarios, and recommend suitable commodities for users according to their color preferences to enhance the shopping experience and satisfaction of users.
[0048] In the specific implementation process, basic information of users such as gender, age, occupation, etc. can be collected through a user registration form, and the pandas library in Python can be used to clean and preprocess the data, converting it into a structured feature vector to obtain user features. At the same time, the behavioral data of users on the platform such as browsing records, purchase records, favorite records, etc. can be obtained from the database, and feature selection algorithms such as chi-square test, mutual information, etc. can be used to select the top-N features that have a greater impact on color preference prediction as the preferred user features. For each user, the number of interaction times and duration on products with different color richness (such as single color, low color richness, medium color richness, and high color richness) can be counted, and these measurement values can be mapped to an interval using Min-Max normalization to obtain the preference distribution of the user for different color richness. Next, the Apriori algorithm can be used to mine association rules for the user feature vector and color preference data, setting the minimum support to 0.01, the minimum confidence to 0.5, and the maximum item set length to 3, to obtain the association rules between different feature combinations and color preferences. Based on the association rules, a user color richness preference prediction model can be constructed to obtain the preference feature vector of the user group.
[0049] It can be understood that since the association rules are obtained through correlation analysis of the preferred user features and preference relationships, and the preference prediction model is established based on the association rules, the output of the preference prediction model can be a preference feature vector with structured features for quantitative analysis.
[0050] Furthermore, the UserCF method in the collaborative filtering algorithm can be adopted. First, a user-product (user group and product) interaction matrix is constructed, and the matrix elements represent the preference measurement of the user for the product. Then, the cosine similarity is used to calculate the similarity matrix between users. For the target user, find the top-K (K = 20) users who are most similar to it, and calculate the average preference value of these users for products with different color richness as the preference prediction value of the target user.
[0051] In one embodiment, for new users, according to their basic information, the top-M (M = 5) user groups that are most similar to their feature combination can be matched using association rules, and the average preference of these groups for products with different color richness is used as the initial preference prediction of the new user. After the new user generates interaction behavior, their behavior data is gradually integrated to update the preference prediction result. The content features such as the color and style of the product can also be introduced, and a content-based recommendation algorithm can be used to generate a personalized initial recommendation list for the new user.
[0052] Furthermore, 5-fold cross-validation can also be used to evaluate the model, resulting in an average accuracy of 85%, an average recall rate of 78%, and an average F1 value of 0.82. The optimized model is applied to the actual product recommendation scenario, and the top-N products with the highest matching degree (N = 10) are recommended to users according to their color preferences, improving the user's shopping experience and satisfaction.
[0053] Step 206: According to the behavior data and preference feature vectors of the user group, use the fuzzy clustering algorithm to analyze the differences and similarities in color richness among different user groups, and use the genetic algorithm to optimize the fuzzy clustering algorithm. By simulating the natural selection and genetic evolution mechanisms, continuously iterate and optimize the clustering center and membership function of the fuzzy clustering algorithm.
[0054] In an exemplary embodiment, the steps of using the fuzzy clustering algorithm to analyze the differences and similarities in color richness among different user groups according to the behavior data and preference feature vectors of the user group include:
[0055] Perform data cleaning and preprocessing operations on the behavior data; the preprocessing operations include extracting the product ID, behavior type, and behavior time of the product and converting them into a structured format; the behavior type includes the actions corresponding to the behavior data, and the actions include browsing, purchasing, or collecting;
[0056] Obtain the color richness of the corresponding product according to the product ID; cluster the pixel points in the product image corresponding to the product ID into several representative colors through the K-means fuzzy clustering algorithm, calculate the first color richness of the representative colors according to the number of representative colors, and extract color-related keywords from the attribute metadata of the product. According to the number and weight of the keywords, obtain the second color richness of the representative colors; perform weighted fusion on the first color richness and the second color richness to obtain the final color richness of the product;
[0057] Sort the behaviors of the users in the user group in chronological order, and use the sliding window method with a certain time span as the window to count the number of actions of the users in the user group on products with different color richness in each window, and obtain the color preference distribution of the users in the user group in different time periods;
[0058] In the case where the user in the user group is a new user, use the average value of the preference thresholds of each user in the user group corresponding to the new user, or the average value of the preference thresholds of the similar users of the new user found by the preference prediction model as the preference threshold of the new user;
[0059] Represent the color preference distribution of users in the user group as a preference feature vector of length N, and convert the user features into a numerical user feature vector using the One-hot encoding or Word2Vec method; where: each element in the preference feature vector represents the degree of preference of the user for the corresponding color.
[0060] Use the fuzzy C-means algorithm to perform clustering analysis on the user preference feature vector, and divide users with similar preference feature vectors into the same user category.
[0061] Optimize the fuzzy C-means algorithm by adjusting its parameters until the clustering performance index reaches the expectation; the parameters of the fuzzy C-means algorithm include the number of clusters and the membership threshold; the clustering performance index includes the silhouette coefficient and the Davies-Bouldin index.
[0062] According to the results of the clustering analysis, analyze the color preference characteristics of different user groups, and explore the differences and similarities in color preferences among different user groups.
[0063] Specifically, by analyzing the user's historical browsing, purchase, and collection behavior data, combined with the preference feature vector of the user for products with different color richness, use fuzzy clustering to analyze the differences and similarities in color preferences among different groups.
[0064] Furthermore, collect the user's historical browsing and purchase behavior data on the platform, clean and preprocess the data, extract keyword fields such as user ID, product ID, behavior type, and behavior time, and convert the data into a structured format.
[0065] Furthermore, according to the product ID, obtain the color richness information of each product; by performing K-means clustering algorithm analysis on the product image, cluster the pixel points of the image into representative main colors, and calculate the color richness score according to the number and distribution of colors; extract color-related keywords from the product attribute metadata, count the number and weight of the keywords, and obtain the color richness score. Weight and fuse the color richness scores obtained by the two methods to obtain the final product color richness. Sort the user's behavior data in chronological order, and use the sliding window method with a certain time span as the window to count the number of times the user browses, purchases, or collects products with different color richness within each window, and obtain the color preference distribution of the user at different time periods (behavior time).
[0066] Furthermore, for users in the user group of new users, they are classified into similar user groups according to their basic attributes. The average preference distribution of the group is used as their initial preference, or the collaborative filtering algorithm is adopted to estimate and fill in the missing preference data according to the preference distributions of similar users. The color preference distribution of a user is represented as a vector of length N, and each element represents the user's preference degree for the corresponding color richness. Then, other user characteristics of the user, including age, gender, and occupation, are transformed into numerical user feature vectors using the One-hot encoding or Word2Vec method. The color preference vector and the attribute feature vector are concatenated or weighted and fused to construct the final user color preference feature vector.
[0067] The fuzzy C-means clustering algorithm in fuzzy clustering is used to perform clustering analysis on the preference feature vector, and users with similar preferences are classified into the same group. The fuzzy C-means clustering algorithm minimizes the objective function, adjusts the distance between the cluster center and the sample points it contains to be the smallest, and at the same time allows a sample point to belong to multiple clusters, which is suitable for processing data with fuzziness and uncertainty such as user preferences. By adjusting the parameters of the fuzzy C-means clustering algorithm, including the number of clusters and the membership threshold, the clustering effect is continuously optimized until the expected clustering performance indicators, including the silhouette coefficient and the Davies-Bouldin index, are achieved. According to the clustering results, the color preference characteristics of different user groups are analyzed to explore the differences and similarities in color preferences among each group.
[0068] In the specific implementation process, various behavioral data of users on the platform, such as browsing, clicking, collecting, and purchasing, can be collected through data logging, and big data processing frameworks such as Hadoop are used to clean and preprocess the data. Key fields such as user ID (basic information data), product ID, behavior type, and timestamp are extracted, and the data is stored in a NoSQL database such as HBase for convenient subsequent querying and analysis. For the color richness information of products, computer vision libraries such as OpenCV can be used to cluster the pixel points of product images through the K-means clustering algorithm. The number of clusters K is set to 16, and the number of iterations is set to 50 to obtain 16 main colors. Then, the standard deviation of the color quantity and distribution is calculated, and after normalization, the color richness score is obtained. At the same time, word segmentation tools such as jieba can be used to segment the product attribute metadata, extract color-related keywords, and calculate the TF-IDF value as the keyword weight. After normalization, another color richness score is obtained. The two scores are weighted and fused according to a ratio of 7:3 to obtain the final product color richness.
[0069] Specifically, for the analysis of users' color preferences, a 30-day sliding window can be adopted to count the interaction times and durations of each user with products of different color richness within the window, and use Min-Max normalization to convert them into preference score vectors. For new users or users with an activity level lower than 10 times per month, the UserCF collaborative filtering algorithm can be used to find the top 50 most similar users based on attributes such as age and gender, and use their average preference vectors for filling. Finally, the user attribute features are converted into 0-1 vectors using one-hot encoding and concatenated with the color preference vector to obtain the final user feature vector with a dimension of 128. During the clustering analysis, the fuzzy C-means clustering algorithm is adopted, with the number of clusters C set to 20, the fuzzification coefficient m set to 2, and the maximum number of iterations set to 100. The cluster centers and membership matrix are obtained by minimizing the objective function. The clustering effect is evaluated through the silhouette coefficient and the DBI index, and the parameters are continuously adjusted until the silhouette coefficient is greater than 0.7 and the DBI index is less than 1.5. The clustering results show that there are significant differences in color preferences among different user groups. For example, male users prefer products with low saturation, while female users prefer products with high saturation; young users prefer bright and colorful products, while middle-aged and elderly users prefer calm and simple products. Based on these insights, personalized products and marketing content can be recommended for different groups to improve the transaction conversion rate and user stickiness of the platform.
[0070] In an exemplary embodiment, a genetic algorithm is used to optimize the fuzzy clustering algorithm. The steps of continuously iteratively optimizing the cluster centers and membership functions of the fuzzy clustering algorithm by simulating the natural selection and genetic evolution mechanisms include:
[0071] Encode the initial parameters of the fuzzy clustering algorithm as genotypes in the genetic algorithm; the initial parameters of the fuzzy clustering algorithm include the cluster centers and the coefficients of the membership functions; the cluster centers use real number encoding, and each gene position represents the coordinate value of a cluster center in the corresponding dimension; the membership functions use floating-point encoding, and each gene position identifies a parameter of the membership function; the length of the gene depends on the dimension of the cluster center and the number of parameters of the membership function;
[0072] Randomly generate an initial population, with the number of individuals in the population being a preset number, and each individual represents a set of possible combinations of clustering parameters;
[0073] Decode the parameters of the clustering algorithm according to the genotype of the individual, and obtain the clustering result based on the training samples and the parameters of the clustering algorithm;
[0074] The clustering performance of each individual is evaluated using a fitness function, and the sum of squared errors of clustering is introduced to measure the clustering error; the fitness function includes the silhouette coefficient, Davies-Bouldin index, Calinski-Harabasz index, and Dunn index; the clustering performance includes evaluating the clustering of each individual, including accuracy, compactness, and separation metrics;
[0075] The fitness function is designed as a weighted combination of metrics, and the weights are adjusted according to actual needs. The larger the value of the fitness function, the better the clustering parameters corresponding to the individual;
[0076] Selection, crossover, and mutation genetic operations are performed on the individuals; the clustering centers are optimized through arithmetic crossover with real number coding and Gaussian mutation to ensure that the newly generated clustering centers are still within the data space; uniform crossover and uniform compilation are used to optimize the membership function to ensure that the generated membership coefficients are within [0, 1]; the selection operation adopts the elitist retention strategy to prevent excellent individuals from being lost during the evolution process;
[0077] The newly generated offspring individuals are added to the population, replacing some individuals with insufficient fitness to maintain the diversity and evolution ability of the population; the clustering parameters are continuously iteratively optimized until the preset number of iterations or fitness threshold is reached to obtain the optimal clustering parameters.
[0078] Specifically, the initial parameters of the fuzzy clustering algorithm, including the clustering centers and the coefficients of the membership function, can be encoded as the genotype in the genetic algorithm. The clustering centers use real number coding, and each gene position represents the coordinate value of a clustering center in a certain dimension. The membership function uses floating-point coding, and each gene position represents a parameter of the membership function. The length of the gene depends on the dimension of the clustering center and the number of parameters of the membership function. A initial population is randomly generated, and the number of individuals in the population is generally set to 50 - 100. Each individual represents a set of possible clustering parameter combinations. According to the genotype of the individual, it is decoded into the parameters of the clustering algorithm, and the fuzzy clustering algorithm is run on the training samples to obtain the clustering results.
[0079] A fitness function is used to evaluate the clustering performance of each individual. Combining the accuracy, compactness, and separation metrics of the clustering results, including the silhouette coefficient, Davies-Bouldin index, Calinski-Harabasz index, and Dunn index, the sum of squared errors of clustering is introduced to measure the clustering error. The fitness function is designed as a weighted combination of metrics, and the weights are adjusted according to actual needs. The larger the value of the fitness function, the better the clustering parameters corresponding to the individual. Selection, crossover, and mutation genetic operations are performed on the individuals in the population.
[0080] Furthermore, for the optimization of the cluster centers, arithmetic crossover with real - number encoding and Gaussian mutation are adopted to ensure that the newly generated cluster centers still lie within the data space. For the optimization of the membership functions, uniform crossover and uniform mutation are adopted to ensure that the generated membership coefficients are within the range of [0, 1]. The selection operation adopts the elitist retention strategy to prevent excellent individuals from being lost during the evolution process. The newly generated offspring individuals are added to the population to replace some individuals with insufficient fitness, maintaining the diversity and evolutionary ability of the population. The clustering parameters are continuously iteratively optimized until the preset number of iterations or fitness threshold is reached, obtaining the optimal combination of clustering parameters. According to the optimal clustering parameters, fuzzy clustering is performed on the user data to obtain the membership matrix of the users.
[0081] Furthermore, for each user, find the cluster with the largest membership degree and assign it to the corresponding user group. For each group, analyze its characteristics in terms of user attributes and behavior preferences, and construct a targeted personalized recommendation model. When generating the recommendation list, use the membership degree of the user as an important personalized factor to adjust the recommendation weights of different items, making the recommendation results more in line with the user's true preferences, and applying it to subsequent user group division and personalized recommendation tasks.
[0082] In the specific implementation process, the cluster centers can be represented as a 10 - dimensional real - number vector, and each dimension corresponds to an attribute in the space of user features, such as age, gender, etc. The membership function can be selected as a Gaussian function, and 2 floating - point parameters are used to represent the mean and variance. Suppose the users are to be divided into 8 groups, then the genotype length of each individual is 82. Randomly generate 100 individuals as the initial population, and each individual represents a set of possible clustering parameters. For each individual, decode it into the parameters of the cluster centers and the membership function, and then run the fuzzy C - means clustering algorithm on the training data of 50,000 users, and iterate 50 times to obtain the membership matrix. Use the weighted average of the silhouette coefficient, CH index, and Dunn index as the fitness function, and the weights can be 0.4, 0.3, and 0.3 respectively. The higher the fitness, the better the clustering effect.
[0083] Furthermore, in genetic operations, ranking selection is adopted, and parents are selected from the top 20% of individuals in terms of fitness with a probability of 0.8 for crossover. Arithmetic crossover is used for the cluster centers, and uniform crossover is used for the membership functions, with a crossover probability of 0.7. In the mutation operation, Gaussian noise is added to each dimension of the cluster centers with a probability of 0.1, and each parameter of the membership function is randomly perturbed within a range with a probability of 0.05. The population is iterated 200 times, and the clustering parameters corresponding to the optimal individual obtained are taken as the final result. These parameters are applied to the full 10 million user data for fuzzy clustering to obtain the membership degrees of each user to 8 groups. Users are assigned to the group with the highest membership degree, and user portraits, purchase behaviors, hobbies, etc. of different groups are analyzed, and significant differences are found in aspects such as age, gender, consumption ability, and category preference among different groups. Accordingly, personalized recommendation models are constructed for each group, comprehensively considering the historical behaviors and membership degrees of users to generate recommendation lists.
[0084] Step 208: Based on the analysis results of the optimized fuzzy clustering algorithm, divide the user groups into multiple categories of users according to the color richness, and obtain the color preference characteristics of users in each category; adopt the item-based collaborative filtering algorithm, and generate a recommendation list of items based on the color richness according to the color preference characteristics and the items corresponding to the corresponding color richness.
[0085] In an exemplary embodiment, the steps of dividing the user groups into multiple categories of users according to the color richness based on the analysis results of the optimized fuzzy clustering algorithm and obtaining the color preference characteristics of users in each category include:
[0086] Apply the optimal clustering parameters to the preference feature vectors of users in the user group, and use the fuzzy C-means clustering algorithm to obtain the membership degrees of each user to different clusters;
[0087] According to the optimal clustering parameters, perform fuzzy clustering on the basic information data of users in the user group to obtain the membership matrix of users, and divide each user into the cluster with the highest membership degree according to the membership matrix to form different color richness preference categories, and the number of categories is obtained from the number of clusters in the optimal clustering parameters;
[0088] Statistically analyze the average preference degrees of each color richness preference category at different color richness levels to obtain the preference distribution characteristics at each color richness level, and statistically analyze the distribution of users in each color preference category in terms of age, gender, education level, and income demographic attributes. Through chi-square test and t-test methods, find the differences in demographic attributes among different color richness preference categories to form demographic characteristic rules;
[0089] Mine association rules and frequent patterns for the behavioral data of users in each category, obtain the key features of users in the corresponding category in terms of category preference, price preference, and activity preference, and form a user portrait with color richness preference; store the color richness preference category and the corresponding user feature rules in the user classification table as the basis for subsequent personalized recommendations and marketing; when the user in the user group is a new user, match the most similar color richness preference category according to his color richness preference characteristics.
[0090] Specifically, obtain the fuzzy clustering parameters optimized by the genetic algorithm, including the cluster centers and membership functions, apply the fuzzy clustering parameters to the color preference feature vector data of users, run the fuzzy C-means clustering algorithm, and obtain the membership degrees of each user to different clusters. According to the membership degree matrix, divide each user into the cluster with the largest membership degree to form different color richness preference categories, and the number of categories is determined by the number of clusters in the optimal clustering parameters.
[0091] Furthermore, for each color richness preference category, statistically analyze the average preference degrees of the users it contains at different color richness levels to obtain the preference distribution characteristics of this category at each color richness level. For each color richness preference category, further analyze its user characteristics. Statistically analyze the distribution of its users in terms of age, gender, region, and income demographic attributes, and use the chi-square test and t-test methods to find the significant differences between different categories in these attributes to form demographic feature rules.
[0092] Furthermore, perform association rule mining and frequent pattern mining on the behavioral data of users' browsing, collecting, and purchasing in each category to obtain the key features of users in this category in terms of category preference, price preference, and activity preference, and form a complete user portrait with color richness preference. Store the color richness preference category and the corresponding user feature rules in the user classification table as an important basis for subsequent personalized recommendations and marketing. When a new user enters the system, match the most similar preference category according to his color richness preference characteristics to quickly achieve user classification. According to the color richness preference category to which the user belongs, select the products that match the preference characteristics of this category from the product library as the candidate recommendation set.
[0093] Furthermore, by combining the similarity between the color richness feature of the product and the category preference distribution, as well as the product category and price metadata, the matching degree is calculated to describe the degree of conformity with the user characteristics of this category. For users with different color richness preference categories, the number of purchases and the number of purchasing users of the products at each color richness level are counted, and the purchase probability (number of purchasing users / total number of users in the category) and purchase frequency (number of purchases / number of purchasing users) are calculated to obtain the association rules between color richness and purchase behavior, which are used to guide the personalized recommendation strategy of the product. When generating the recommendation list, for each candidate product, the cosine similarity between its color richness feature and the user's category preference distribution is calculated as the color richness matching degree weight. At the same time, combined with other product features, including category and price, and the matching degree with user characteristics, the comprehensive matching degree score is calculated.
[0094] Finally, the candidate products are sorted according to the comprehensive matching degree score, and the products with high matching degree are ranked in the front. Set the threshold of the color richness weight. For products exceeding the threshold, further improve their rankings to highlight the influence of color richness preference, make the recommendation results more in line with the user's color preference, and improve the conversion rate and user satisfaction of the recommendation.
[0095] In the specific implementation process, the optimized fuzzy clustering parameters, such as 10 clustering centers and the corresponding Gaussian membership functions, can be used to cluster the color preference feature vectors of 1 million users. By calculating the membership degree of each user to each clustering center, it is assigned to the clustering with the largest membership degree to form 10 different color richness preference categories.
[0096] Furthermore, for each category, the average preference score of its users at 5 color richness levels is counted, and a preference distribution radar chart is drawn. At the same time, an independence test is performed on the user attribute distribution of each category. For example, it is found that the proportion of users aged 18 - 25 in the high color richness preference category is significantly higher than other categories (p < 0.01). The Apriori algorithm is also used for association rule mining to obtain the purchase patterns of users in each category in different price ranges, categories, etc. For example, users with high color richness preference are more inclined to purchase T-shirts and coats priced at 100 - 200 yuan (support > 5%, confidence > 60%). According to the user classification results, 1000 products are selected from 100,000 products as the initial recommendation set, and the 500 products with the highest matching degree are selected by calculating the JS divergence between the product color characteristics and the user preference distribution. After comprehensively considering the matching degrees of other characteristics such as category and price, the top 50 products are selected to generate a personalized recommendation list. Finally, the differences in purchase behavior of different clustering users are also analyzed. For example, the average monthly purchase frequency of users with high color richness preference is 1.5 times, which is 2 times that of users with low color richness, and the unit price per pen is also 30% higher, which provides an important reference for subsequent marketing strategies and inventory management.
[0097] In an exemplary embodiment, the steps of generating a recommended list of products based on color richness by using a product-based collaborative filtering algorithm according to color preference features and corresponding products with color richness include:
[0098] Dividing the products in the product library into corresponding product subsets according to color richness, and selecting candidate recommended products from the product subsets corresponding to the color richness according to the color richness preference category to which each user belongs;
[0099] Setting corresponding user attribute values according to the user's color richness preference category; the user attribute values are positively correlated with the number of colors in the color richness; for users with a high color richness preference, selecting products with a color richness attribute value greater than or equal to 4 as candidates; for users with a medium color richness preference, selecting products with an attribute value between 2 and 4; for users with a low color richness preference, selecting products with an attribute value less than 2; and combining the user's historical behavior data, setting weights positively correlated with the occurrence frequency of the user's browsing, purchase or collection records; calculating and sorting the similarity of the candidate recommended products by using a product-based collaborative filtering algorithm;
[0100] Judging the similarity between products by color richness and other similarity metrics, and obtaining the color richness difference degree according to the difference in color richness between products; other similarity metrics include whether there are browsing, purchase or collection records of multiple users with the same user attribute values between products; and other similarity metrics are positively correlated with the occurrence frequency of the browsing, purchase or collection records of users with the same user attribute values; the difference in color richness between products is the difference in the number of colors; mapping the color richness difference degree to the interval of 0 to 1 through a mapping function to obtain the color richness similarity; the mapping function includes a Gaussian kernel function; weighting and summing the color richness similarity and other similarity metrics to obtain the comprehensive similarity;
[0101] Using the K-means clustering algorithm to cluster the candidate recommended products according to color richness, dividing the products with similar color richness into the same cluster to form product subsets with different color richness; sorting the candidate recommended products and dividing them into different color richness subsets to form a hierarchical recommended list; the recommended list is used to preferentially recommend products that match the user's color richness preference category; when generating the final recommended list of products with color richness, combining factors such as real-time inventory, promotional activities and new product launches of products to dynamically adjust the recommended list to ensure the diversity and timeliness of the recommended products and continuously optimize the recommended results.
[0102] Specifically, obtain the basic information and color preference characteristics of each user, and determine the color quantity preference category to which the user belongs according to the fuzzy clustering result generated in the previous step. Extract the metadata information of each product in the product library, and focus on obtaining the color richness attribute of the product, that is, the number of main colors contained in the product image obtained through the color clustering algorithm. Divide the products in the product library according to the color richness attribute to obtain product subsets with different color richness levels, including products with low color richness, products with medium color richness, and products with high color richness.
[0103] Specifically, for each user, select candidate recommended products from the product subsets of the corresponding color richness level according to the color quantity preference category to which the user belongs. Set different screening thresholds according to the color richness preference category of the user. For users with a high color richness preference, select products with a color richness attribute value greater than or equal to 4 as candidates; for users with a medium color richness preference, select products with an attribute value between 2 and 4; for users with a low color richness preference, select products with an attribute value less than 2. At the same time, combine the user's historical behavior data, and give higher weights to products that appear frequently in their browsing, collection, and purchase records to increase their selection probability.
[0104] Furthermore, for the set of candidate recommended products, use the item-based collaborative filtering algorithm to calculate similarity and sort. When calculating the similarity between every two products, in addition to the factors of co-purchase and co-browsing, the similarity of the color richness attribute should also be introduced. For two products A and B, extract their color richness attribute values colorA and colorB respectively, calculate the absolute value of the difference between the two abs(colorA - colorB) to obtain the color richness difference degree. Map the difference degree to the interval from 0 to 1 to obtain the color richness similarity. The mapping function can use the Gaussian kernel function, in the form of exp(-abs(colorA - colorB)^2 / 2σ^2), where σ is a parameter that controls the similarity decay rate.
[0105] Furthermore, sum the color richness similarity and other similarity metrics with weights to obtain the comprehensive similarity. For the set of candidate recommended products, use the K-means clustering algorithm to cluster according to the color richness attribute values of the products. The purpose of clustering is to divide products with similar attribute values into the same cluster to form product subsets with different color richness levels. The number of clusters K can be set according to the scale of the candidate products and the color richness distribution, generally taking 3 to 5.
[0106] Specifically, when clustering, other important attributes of the products can also be incorporated into the clustering attributes to improve the clustering effect. Based on the product similarity matrix and the clustering results, the candidate recommended products are sorted, and the products are divided into subsets with different color richness levels to form a hierarchical recommendation list. Priority is given to recommending products in the color richness level with the highest matching degree to the user's preferences. When generating the final color richness product recommendation list, factors such as real-time inventory, promotional activities, and new product launches of the products are combined to dynamically adjust the recommendation list to ensure the diversity and timeliness of the recommended products and continuously optimize the recommendation results.
[0107] Exemplarily, through user registration information and behavior logs, basic attributes such as the age and gender of each user and color preference characteristics can be obtained. According to the user's behaviors such as browsing, clicking, collecting, and purchasing on products with different color richness levels, the corresponding preference scores are statistically calculated, and then the users are divided into three color richness preference categories: high, medium, and low through fuzzy clustering.
[0108] For 1 million products, the RGB histograms of their images are extracted. Through peak detection and threshold filtering, the number of main colors contained in each product is obtained, and it is normalized to the range of 0 - 1 as the color richness attribute value. According to this attribute value, the K-means clustering algorithm is used to cluster the products into 5 clusters, which respectively represent five color richness levels: very low, low, medium, high, and very high. For each user, according to their preference category, products in the corresponding color richness level are selected as candidates, and at the same time, a weight of 1.5 - 2 times is given to the products in the user's historical behavior.
[0109] Specifically, when calculating the product similarity, the color richness attribute values of the products are extracted, the absolute values of the differences are calculated, and then they are mapped to the range of 0 - 1 through a Gaussian kernel function with σ = 0.1 to obtain the color richness similarity. Then, combined with similarities such as co-purchase and co-browsing, they are weighted and summed according to a ratio of 2:2:1 to obtain a comprehensive similarity matrix. The candidate products are further clustered, and products with similar color richness are clustered into 3 - 5 clusters. Within each cluster, they are sorted according to the similarity to form a hierarchical recommendation list. Finally, the recommendation list also needs to be dynamically adjusted, such as filtering out out-of-stock products and giving a ranking boost to products during the promotion period, etc., to ensure the real-time and effectiveness of the recommendation.
[0110] Step 210, collect and, according to the feedback information of users in each category on the recommendation list, dynamically adjust the recommendation list to optimize the color preference characteristics.
[0111] In an exemplary embodiment, the steps of collecting and, according to the feedback information of users in each category on the recommendation list, dynamically adjusting the recommendation list to optimize the color preference characteristics include:
[0112] Collect feedback information from users on historical recommended products; the feedback information includes behavioral data and explicit feedback; the explicit feedback includes the scoring and commenting situations of the recommended products by users according to the set rules; through feature extraction of the feedback information, obtain the preference performance of users for products with different color richness; the preference performance is used to indicate the click-through rate and purchase rate;
[0113] Compare the feedback information with the preference portrait of the corresponding user's color richness preference category; the preference portrait comparison is used to indicate the calculation of the satisfaction of the recommended products for the user preference category; the satisfaction includes click-through rate and purchase rate indicators;
[0114] In response to the click-through rate being lower than the preset click-through rate or the purchase rate being lower than the preset purchase rate, by constructing the heat distribution of the recommended products, analyze whether there is a situation of insufficient recommendation of long-tail products or over-recommendation of head products, and obtain the analysis result based on the heat distribution;
[0115] According to the analysis result based on the heat distribution, dynamically adjust the recommendation list, including: if the satisfaction is higher than the first preset satisfaction, increase the recommendation proportion of the corresponding recommended product and give priority to the recommendation, and increase the weight of the corresponding recommended product in the process of similarity calculation; if the satisfaction is lower than the second preset satisfaction, reduce the recommendation proportion of the corresponding recommended product and reduce the weight of the corresponding recommended product in the process of similarity calculation.
[0116] Specifically, collect feedback information from users on historical recommended products, including users' click, browse, favorite, and purchase behavioral data, as well as explicit feedback on the scoring and commenting of the recommended products by users. Preprocess and extract features from the user feedback data to obtain the preference performance of users for products with different color richness, including the click-through rate and purchase rate of products with high color richness. Compare the user feedback features with the user's original color richness preference portrait to calculate the acceptance degree and satisfaction of the user for the recommended products. By analyzing the behavioral feedback of users on the recommended products, calculate the overall click-through rate and purchase rate indicators of the recommended products and compare them with the average level of the platform. If the indicators of the recommended products are lower than the average level, it indicates that there are deficiencies in the recommendation strategy.
[0117] Furthermore, by constructing the popularity distribution of recommended products, analyze whether there is a phenomenon of insufficient recommendation of long-tail products or over-recommendation of head products, and identify problems in the recommendation strategy. According to user feedback, optimize and adjust each link in the recommendation strategy. When selecting candidate products, calculate the recommendation satisfaction of each color richness level according to the user's feedback on products with different color richness levels. For the levels with high satisfaction, increase their proportion in the candidate products; for the levels with low satisfaction, correspondingly reduce their proportion. When calculating similarity, increase the weight of product features with positive user feedback; for product features with negative feedback, reduce the weight. The adjustment range of the weight can be determined according to the intensity and persistence of user feedback. When generating the recommendation list, give priority to recommending products with positive user feedback, and downgrade or filter products with negative feedback. At the same time, when clustering candidate products, in addition to the color richness attribute, introduce new clustering features in combination with user feedback, including user satisfaction and recommendation success rate. Use the K-means algorithm for clustering, and determine the optimal number of clusters by calculating the silhouette coefficient. According to user feedback, dynamically adjust the clustering center, gather products with positive user feedback together to form a product cluster with high satisfaction; remove or sink products with negative user feedback to optimize the clustering results and make them more in line with the true preferences of users. Regularly update the user's color richness preference profile, take the long-term and stable behavior feedback of users as the main basis for preferences, and reduce the influence of accidental and sudden behaviors. Continuously optimize the user's color richness classification to accurately reflect the true preferences of users. Conduct an AB test to evaluate the optimization effect of the recommendation strategy. Randomly divide users into an experimental group and a control group. The experimental group adopts the optimized recommendation strategy, and the control group adopts the original recommendation strategy.
[0118] Furthermore, compare the click-through rate, purchase rate, and satisfaction indicators of the two groups of users within a certain period of time, and judge whether the optimized recommendation strategy has brought a statistically significant improvement through a significance test. Use the AB test to quantify the effect of recommendation optimization and verify the effectiveness of the optimization measures. Establish an evaluation and feedback mechanism for the recommendation strategy, continuously track the user's feedback on the recommendation results, promptly discover problems and deficiencies in the recommendation, and quickly make adjustments and optimizations to form a closed-loop improvement of the recommendation strategy, and continuously improve the accuracy of the recommendation and user satisfaction.
[0119] In the specific implementation process, various feedback behavior data of users on the recommended products can be collected through data logging and logs, such as clicks, collections, adding to cart, purchases, etc., as well as explicit feedback such as ratings and comments. For a recommendation system with 10 million users and 1 million products, TB-level user feedback data can be collected every day. Through big data processing platforms such as Spark, the feedback data is cleaned, transformed, and feature extracted to obtain the preference performance indicators of users on products with different color richness levels, such as click-through rate and conversion rate. These indicators are correlated with the original color preference portraits of users, and measures such as Pearson correlation coefficient are used to measure the satisfaction degree of users with the recommended products.
[0120] Furthermore, if it is found that the overall click-through rate of the recommended products is more than 10% lower than the platform average, or the top 1% of the popular products account for more than 50% of the recommended exposure, then the recommendation strategy optimization process is triggered. Dynamically adjust the proportion of different color richness levels in the candidate product pool, increase the sampling volume of the level with the highest satisfaction by 20%, and decrease the sampling volume of the level with the lowest satisfaction by 20%. In the similarity calculation of ItemCF, multiply the weights of users' click and purchase behaviors by 1.5 and 2.0 respectively, and multiply the weights of users' browsing and collection behaviors by 0.8 and 1.2. For the newly added clustering features, such as satisfaction and success rate, assign weights of 0.3 and 0.5 respectively, and divide the candidate products into 5 - 8 clusters through KModes clustering. For the product clusters with a satisfaction greater than 0.7 and a recommendation success rate higher than 80%, increase their ranking by 20% when generating the recommendation list; for the product clusters with a satisfaction lower than 0.3 or a recommendation success rate lower than 30%, lower their ranking by 50%. When updating the user preference portrait, assign a weight of 1.5 times to the color characteristics of the products with click and collection behaviors occurring continuously for 3 days, and assign a weight of 2 times to the color characteristics of the products with purchase behaviors occurring in the last 7 days, to increase the weight of the stable preference characteristics.
[0121] Furthermore, through monthly regular updates, the compliance rate of the color richness user division with the true preference can be increased from 85% to over 93%. Conduct an AB test of the recommendation strategy once a week, randomly select 5% of the users as the experimental group, use the optimized strategy to generate personalized recommendations for them, observe whether indicators such as the CTR and conversion rate of the experimental group exceed those of the control group by more than 20%, and conduct paired t-tests and Mann-Whitney U-tests to verify the significance of the improvement in the experimental group. Through continuous strategy optimization and effect evaluation, increase the average CGR from 8% to 12%, and increase the NDCG from 0.32 to 0.41, effectively improving the recommendation effect and user satisfaction.
[0122] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0123] In an exemplary embodiment, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a low-altitude monitoring method for an unmanned aerial vehicle as provided above in the present application.
[0124] In an exemplary embodiment, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a low-altitude monitoring method for an unmanned aerial vehicle as provided above in the present application.
[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0127] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An e-commerce personalized recommendation service method based on genetic fuzzy clustering, characterized in that, The method includes: Converting the commodity images corresponding to each acquired and loaded commodity into the RGB color space, extracting the colors of the commodity images, and dividing the corresponding commodities into monochromatic, low color richness, medium color richness, and high color richness according to the number of colors of the commodity images from low to high in terms of color richness; Based on the basic information data of each user group obtained and the behavioral data of the corresponding user group, analyzing the preference relationship between the user group and the color richness of the commodity, and establishing a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group; the user group includes users of different genders, occupations, and ages; the behavioral data includes the browsing records, purchase records, and collection records of the user group for the commodity; the steps of analyzing the preference relationship between the user group and the color richness of the commodity and establishing a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group include: converting the basic information data into a structured feature vector to obtain user features; performing feature selection on the user features according to the behavioral data of the user group to obtain preference user features associated with the behavioral data; respectively counting the interaction behaviors of different user groups on the commodities with each color richness as the preference measure of the user group for color richness; the interaction behaviors include the number and duration of generating browsing records, purchase records, and collection records for the commodity; normalizing each preference measure to obtain the preference distribution of different user groups for color richness; performing a correlation analysis on the preference user features and the preference relationship to obtain the association rules between different feature combinations and the color preference; the different feature combinations include color combinations of color richness; constructing a preference prediction model based on the association rules between the different feature combinations and the color preference to obtain the preference feature vector of the user group, and optimizing the preference prediction model using collaborative filtering algorithms and matrix factorization algorithms; evaluating the accuracy rate, recall rate, and F1 metric of the preference prediction model, and adjusting the parameters of the preference prediction model and performing feature selection of user features according to the evaluation results; the steps of optimizing the preference prediction model using collaborative filtering algorithms and matrix factorization algorithms include: constructing an interaction matrix between the user group and the commodity, and using a matrix factorization algorithm to calculate the interaction matrix to find similar users of the target user among the set user groups, and calculating the preference threshold of the similar users for the commodities corresponding to each color richness as the preference prediction value of the target user; training the preference prediction model according to the basic information data of the target user and the preference prediction value of the target user; According to the behavioral data of the user group and the preference feature vector, using the fuzzy clustering algorithm to analyze the differences and similarities of different user groups in terms of color richness, and optimizing the fuzzy clustering algorithm using the genetic algorithm; Based on the analysis results of the optimized fuzzy clustering algorithm, the user group is divided into multiple categories of users according to the color richness, and the color preference characteristics of the users in each category are obtained; according to the color preference characteristics and the corresponding products corresponding to the color richness, a recommendation list of products based on the color richness is generated; Collect and, according to the feedback information of the users in each category on the recommendation list, dynamically adjust the recommendation list to optimize the color preference characteristics.
2. The method according to claim 1, wherein The steps of extracting the colors of the product images and classifying the corresponding products into monochromatic, low color richness, medium color richness, and high color richness according to the number of colors of the product images from low to high color richness include: Use the fuzzy clustering algorithm to process the converted RGB color space; through an iterative method, allocate the data points in the RGB color space to K cluster centers and adjust the distance from each data point to the cluster center to be the minimum; Cluster the pixel points of the product image into several representative colors to obtain the color distribution of the product image with respect to the representative colors; Judge that the color richness of the product with the number of colors in the color distribution below the first preset threshold is monochromatic, the color richness of the product with the number of colors in the color distribution exceeding the first preset threshold and below the second preset threshold is low color richness, the color richness of the product with the number of colors in the color distribution exceeding the second preset threshold and below the third preset threshold is medium color richness, and the color richness of the product with the number of colors in the color distribution exceeding the third preset threshold is high color richness; wherein, the first preset threshold, the second preset threshold, and the third preset threshold are all set according to the pre-stored business requirements and business data.
3. The method according to claim 1, wherein The steps of analyzing the preference relationship between the user group and the color richness of the product and establishing a preference prediction model according to the preference relationship to obtain the preference feature vector of the user group further include: In the case that the target user is a new user, according to the basic information data of the new user, find the similar users through the preference prediction model, calculate the average value of the preference prediction values of the similar users, and use it as the initial preference prediction value of the new user; In the case that the new user generates an interaction behavior, update the preference prediction value according to the behavior data of the new user; introduce the color richness of the product, and generate an initial recommendation list for the new user through a recommendation method based on the content of the product to promote the interaction behavior of the new user.
4. The method according to claim 3, characterized in that The steps of analyzing the differences and similarities of different user groups in color richness by using the fuzzy clustering algorithm according to the behavior data of the user group and the preference feature vector include: Perform data cleaning and preprocessing operations on the behavior data; the preprocessing operations include extracting the product ID, behavior type, and behavior time of the product and converting them into a structured format; the behavior type includes the actions corresponding to the behavior data, and the actions include browsing, purchasing, or collecting; Obtain the color richness of the corresponding product according to the product ID; cluster the pixel points in the product image corresponding to the product ID into several representative colors through the fuzzy clustering algorithm, calculate the first color richness of the representative colors according to the number of the representative colors, and extract color-related keywords from the attribute metadata of the product, and obtain the second color richness of the representative colors according to the number and weight of the keywords; perform weighted fusion on the first color richness and the second color richness to obtain the final color richness of the product; Sort the behaviors of the users in the user group in chronological order, and use the sliding window method with a certain time span as the window to count the number of actions of the users in the user group on products with different color richness in each window, so as to obtain the color preference distribution of the users in the user group in different time periods; In the case that the user in the user group is a new user, use the average value of the preference thresholds of each user in the user group corresponding to the new user, or the average value of the preference thresholds of the similar users of the new user found by using the preference prediction model as the preference threshold of the new user; Represent the color preference distribution of the users in the user group as a preference feature vector of length N, and convert it into a numerical user feature vector; where: each element in the preference feature vector represents the preference degree of the user for the corresponding color; Use the fuzzy C-means algorithm to perform clustering analysis on the preference feature vector of the user, and divide the users with similar preference feature vectors into the same user category; Optimize the fuzzy C-means algorithm by adjusting the parameters of the fuzzy C-means algorithm until the clustering performance index reaches the expectation; the parameters of the fuzzy C-means algorithm include the number of clusters and the membership threshold; the clustering performance index includes the silhouette coefficient and the Davies-Bouldin index; According to the results of the clustering analysis, analyze the color preference characteristics of different user groups, and explore the differences and similarities in color preferences among different user groups.
5. The method according to claim 1, characterized in that, The steps of optimizing the fuzzy clustering algorithm by using the genetic algorithm include: Encode the initial parameters of the fuzzy clustering algorithm as the genotype in the genetic algorithm; the initial parameters of the fuzzy clustering algorithm include the cluster centers and the coefficients of the membership function; the cluster centers are encoded by real numbers, and each gene bit represents the coordinate value of a cluster center in the corresponding dimension; the membership function is encoded by floating-point numbers, and each gene bit identifies a parameter of the membership function; the length of the gene depends on the dimension of the cluster center and the number of parameters of the membership function; Randomly generate an initial population, the number of individuals in the population is a preset number, and each individual represents a set of possible combinations of clustering parameters; Decode the parameters of the clustering algorithm according to the genotype of the individual, and obtain the clustering result based on the training samples and the parameters of the clustering algorithm; The clustering performance of each of the individuals is evaluated using a fitness function, and the sum of squared errors of clustering is introduced to measure the clustering error; the fitness function includes the silhouette coefficient, Davies-Bouldin index, Calinski-Harabasz index, and Dunn index; the clustering performance includes indicators such as accuracy, compactness, and separation for evaluating the clustering of each of the individuals; The fitness function is designed as a weighted combination of the indicators, and the weights are adjusted according to actual needs. The larger the value of the fitness function, the better the clustering parameters corresponding to the individual; Selection operations, crossover operations, and mutation genetic operations are performed on the individuals; the cluster centers are optimized to ensure that the newly generated cluster centers are still within the data space; uniform crossover and uniform compilation are used to optimize the membership function to ensure that the generated membership coefficients are within [0, 1]; the selection operation adopts an elitist retention strategy to prevent excellent individuals from being lost during the evolution process; The newly generated offspring individuals are added to the population, replacing some individuals with insufficient fitness to maintain the diversity and evolution ability of the population; the clustering parameters are continuously iteratively optimized until a preset number of iterations or fitness threshold is reached to obtain the optimal clustering parameters.
6. The method according to claim 5, wherein The steps of dividing the user group into multiple categories of users according to the color richness based on the analysis result of the optimized fuzzy clustering algorithm and obtaining the color preference characteristics of the users in each category include: Applying the optimal clustering parameters to the preference feature vectors of the users in the user group and using the fuzzy C-means clustering algorithm to obtain the membership degrees of each user to different clusters; According to the optimal clustering parameters, fuzzy clustering is performed on the basic information data of the users in the user group to obtain the membership matrix of the users, and each user is divided into the cluster with the largest membership degree according to the membership matrix to form different color richness preference categories, and the number of categories is obtained from the number of clusters in the optimal clustering parameters; The average preference degrees of each color richness preference category on different color richness levels are statistically analyzed to obtain the preference distribution characteristics on each color richness level, and the distribution of the users in each color preference category in terms of age, gender, education level, and income demographic attributes is statistically analyzed. Through the chi-square test and t-test methods, the differences in the demographic attributes among different color richness preference categories are found to form demographic characteristic rules; Association rule mining and frequent pattern mining are performed on the behavior data of the users in each category to obtain the key characteristics of the users in each category in terms of category preference, price preference, and activity preference, forming a user portrait of color richness preference; the color richness preference categories and the corresponding user characteristic rules are stored in the user classification table as the basis for subsequent personalized recommendation and marketing; in the case where the user in the user group is a new user, the most similar color richness preference category is matched according to his color richness preference characteristics.
7. The method according to claim 5, wherein The steps of generating a recommendation list of products based on color richness according to the color preference characteristics and the products corresponding to the corresponding color richness include: Commodities in the commodity library are divided into corresponding commodity subsets according to the color richness, and candidate recommended commodities are selected from the commodity subsets with corresponding color richness according to the color richness preference category to which each user belongs; According to the color richness preference category stated by the user, set the corresponding user attribute value; the user attribute value is positively correlated with the number of colors in the color richness; for users with a high color richness preference, select commodities with a color richness attribute value greater than or equal to 4 as candidates; for users with a medium color richness preference, select commodities with an attribute value between 2 and 4; for users with a low color richness preference, select commodities with an attribute value less than 2; and in combination with the user's historical behavior data, set a weight positively correlated with the occurrence frequency of the user's browsing, purchase, or collection records; use a collaborative filtering algorithm based on the commodities to calculate similarity and sort the candidate recommended commodities; Judge the similarity between the commodities through color richness and other similarity metrics, and obtain the color richness difference degree according to the difference in color richness between the commodities; the other similarity metrics include whether there are browsing, purchase, or collection records of multiple users with the same user attribute value between the commodities; and the other similarity metrics are positively correlated with the occurrence frequency of the browsing, purchase, or collection records of users with the same user attribute value; the difference in color richness between the commodities is the difference in the number of colors; map the color richness difference degree to the interval of 0 to 1 through a mapping function to obtain the color richness similarity; the mapping function includes a Gaussian kernel function; weight and sum the color richness similarity and other similarity metrics to obtain the comprehensive similarity; Use the K-means clustering algorithm to cluster the candidate recommended commodities according to the color richness, divide the commodities with similar color richness into the same cluster to form commodity subsets with different color richness; sort the candidate recommended commodities and divide them into subsets with different color richness to form a hierarchical recommendation list; the recommendation list is used to preferentially recommend commodities that match the user's color richness preference category.
8. The method according to claim 7, characterized in that, The step of collecting and dynamically adjusting the recommendation list according to the feedback information of each category of users on the recommendation list to optimize the color preference characteristics includes: Collect the feedback information of users on historical recommended commodities; the feedback information includes behavior data and explicit feedback; the explicit feedback includes the scoring and commenting situations of users on the recommended commodities according to the set rules; through feature extraction of the feedback information, obtain the preference performance of users for commodities with different color richness; the preference performance is used to indicate the click-through rate and purchase rate. Compare the feedback information with the preference portrait of the user's color richness preference category; the preference portrait comparison is used to indicate the calculation of the satisfaction of the user with the recommended commodities in the preference category; the satisfaction includes the click-through rate and purchase rate indicators; In response to the click-through rate being lower than a preset click-through rate or the purchase rate being lower than a preset purchase rate, by constructing the popularity distribution of recommended products, analyze whether there is a shortage of recommended long-tail products or over-recommendation of head products, and obtain an analysis result based on the popularity distribution; According to the analysis result based on the popularity distribution, dynamically adjust the recommendation list, including: if the satisfaction is higher than the first preset satisfaction, increase the recommendation proportion of the corresponding recommended product and give priority to the recommendation, and increase the weight of the corresponding recommended product during the similarity calculation process; if the satisfaction is lower than the second preset satisfaction, reduce the recommendation proportion of the corresponding recommended product and reduce the weight of the corresponding recommended product during the similarity calculation process.
Citation Information
Patent Citations
Matching method for users and commodities and commodity matching recommendation method based on color
CN106354768A
Comprehensive decision-making method and system for preference learning based on graph neural network, and medium
CN117540247A