User classification method based on prepolymerization storage table

By comprehensively collecting, fine processing and efficient storage of user data, combined with advanced clustering and association rule mining algorithms, the problem of poor user classification effect in the existing technology is solved, and accurate user classification and personalized services are achieved.

CN120336604AActive Publication Date: 2025-07-18北京蜂创科技有限公司

Patent Information

Application Number
CN202510426529.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing technology has problems with data legality in user classification, low data quality, inaccurate feature extraction, unreasonable storage structure design, and low efficiency of analysis algorithms, resulting in poor user classification results and unable to meet the needs of enterprises for precise marketing and personalized services.

Method used

Data is collected comprehensively through network crawlers, log analysis and third-party data interfaces, strictly follow laws and regulations, adopt the 3σ principle and multiple missing value processing methods to pre-process data, deeply extract features, build pre-aggregated storage tables, and design an efficient index structure. Combined with K-Means algorithm and association rule mining, clustering and association rule mining, and adjust classification results based on business experience.

Benefits of technology

Ensure that data is legal and compliant, improve data quality, reduce redundancy, improve data processing efficiency, accurately divide customer groups, provide enterprises with the basis for precise marketing and personalized services, optimize product recommendation strategies, and improve operational results and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336604A_ABST
    Figure CN120336604A_ABST
Patent Text Reader

Abstract

The invention discloses a user classification method based on a preaggregation storage table, and relates to the technical field of data processing, and the method comprises the steps: S1, data collection, S2, data preprocessing: carrying out all-directional cleaning on customer data collected in the step S1, employing a 3 sigma principle based on a statistical method to recognize and remove noise data and abnormal values, and carrying out data classification; s3, feature extraction: deep feature extraction is carried out on the customer data preprocessed in the step S2, S4, a prepolymerization storage table is constructed, and S5, data analysis is carried out. Through a data collection stage, a web crawler technology, a log analysis tool and a third-party data interface are comprehensively utilized, customer behaviors and basic information data are comprehensively collected, laws and regulations and website protocols are strictly followed, the data are ensured to be legal and compliant, and the user experience is improved. Meanwhile, during data preprocessing, a 3 sigma principle, a plurality of missing value processing methods and a data smoothing and normalization technology are adopted, noise is effectively removed, missing values are filled, data quality is improved, and a classification result can truly reflect customer characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a user classification method based on a pre-aggregated storage table. Background Art

[0002] With the rapid development of the Internet and information technology, a huge amount of user data has been accumulated in various industries. Enterprises expect to use this data to deeply understand user behavior, consumption habits, etc., so as to achieve precise marketing, personalized services, and product optimization, and enhance their own competitiveness; in the field of data processing, the explosive growth of data volume has promoted the wide application of big data analysis technology. At the same time, technologies such as web crawlers, log analysis, and third-party data interfaces have become important means for obtaining data, and algorithms such as clustering analysis and association rule mining have also been continuously developed, providing technical support for user classification and behavior analysis. In terms of storage, distributed storage technology and various index structures are used to efficiently manage data. In this context, how to integrate these technologies to build a complete and efficient user classification method has become a research hotspot.

[0003] There are many deficiencies in the existing technology for user classification. In the data collection stage, some data collection methods may violate laws, regulations, or website agreements, resulting in doubts about the legality of data acquisition. At the same time, the obtained data may be incomplete and inaccurate; in the data preprocessing link, the processing of noise data and outliers is not fine enough, and the missing value processing method is single, making it difficult to adapt to complex data scenarios and affecting data quality; when extracting features, there is a lack of effective feature selection algorithms, and the extracted features have a high redundancy and cannot accurately reflect the key information of users; in terms of data storage and analysis, the storage structure design is unreasonable, and the index technology is inefficient, resulting in slow data query and processing speeds; the selection and application of data analysis algorithms are not flexible enough, the clustering results are inaccurate, and the depth of association rule mining is insufficient, making it difficult to deeply understand user behavior. Eventually, the user classification effect is not good and cannot meet the needs of enterprises for precise marketing and personalized services. There is a need to design a user classification method based on a pre-aggregated storage table to solve the above-mentioned problems. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a user classification method based on a pre-aggregated storage table to solve the problems that appear in the above technical solutions.

[0005] To achieve the above object, the present invention is realized through the following technical solutions: A user classification method based on a pre-aggregated storage table includes the following classification methods: S1. Data collection: First, use diverse technical means and channels, as well as based on historical data and browsing situations, to comprehensively collect customer behavior data and basic information data. And through web crawler technology, according to the set rules and strategies, accurately capture the browsing records of customers from specified web pages, e-commerce platforms, and social media network resources, and detailedly record information such as the page URLs visited by customers, access times, and browsing sequences. Then use log analysis tools to deeply analyze website server logs, APP operation logs, etc., to obtain the operation trajectories of customers within the application, including search records, click behaviors, stay times, etc. Secondly, by docking with third-party data interfaces, obtain authoritative customer basic information data, and obtain information such as the age, gender, and region of customers through the demographic data platform, and obtain the occupation information of customers through industry data service providers. Finally, the collected customer behavior data includes but is not limited to customers' browsing records, purchase records, search records, and stay times, and at the same time collect customers' basic information data; S2. Data preprocessing: Perform a full range of cleaning on the customer data collected in step S1. Adopt the 3σ principle based on statistical methods to identify and remove noise data and outliers. The 3σ principle is based on the assumption of the normal distribution of data. If a data point is more than 3 times the standard deviation away from the mean, it is determined as an outlier. For missing values, different methods are used for processing according to the data characteristics. For numerical data, the mean filling method is used when the data distribution is relatively uniform, and the regression prediction filling method is used when there is a linear relationship in the data. For complex data scenarios, the multiple imputation method is used to generate multiple reasonable filling values. At the same time, use data smoothing technology to eliminate the minor fluctuations in the data and regularize the data. Adopt normalization technology to improve the data quality and perform zero-mean and unit-variance processing on the data; S3. Feature extraction: Deeply extract features from the preprocessed customer data in step S2. Deeply extract features from the preprocessed customer behavior data. Through statistical analysis of the login times and operation frequency data of customers over a period of time, obtain the activity and usage time of customers. By calculating the number of purchases of customers within a certain period, obtain the purchase frequency. By accumulating the purchase amount of each customer, obtain the purchase amount. By analyzing the categories of goods purchased by customers, determine the preferred categories. And use feature selection algorithms to screen out representative and discriminative features from the extracted features; The feature selection algorithms include the information gain method and the chi-square test method. The information gain method calculates the information gain of each feature for the classification target and selects features with high information gain. The chi-square test method is based on the independence assumption between the feature and the classification target and measures the importance of the feature; S4. Construct a pre-aggregated storage table: Based on the customer data preprocessed in step S3, construct a pre-aggregated storage table according to the extracted features, and perform aggregated storage and classification aggregation on the data according to the time dimension and attribute dimension. At the same time, aggregate and store the corresponding data according to the behavior type dimension, and use distributed storage technology to store the pre-aggregated storage table. Design an efficient data index structure for different feature dimensions. For range queries on the time dimension, use the B-tree index method; for range queries on the attribute and behavior dimensions, use the hash index method. S5. Data analysis: Use big data analysis technology to deeply analyze customer behavior data based on the pre-aggregated storage table constructed in step S4. The specific implementation steps are as follows: T1. Cluster analysis: Use a clustering algorithm to cluster customers into different customer groups. Use the K-Means algorithm and combine it with the elbow method to calculate the clustering error under different K values; use the silhouette coefficient method to calculate the silhouette coefficient of each sample. T2. Association rule mining: Use the Apriori algorithm and the FP-Growth algorithm to mine the association rules between customer behaviors within different customer groups for the customer data obtained from the cluster analysis in step T1. The Apriori algorithm is based on the prior principle, and by generating candidate frequent item sets and validating them in the dataset, gradually mines out the frequent item sets and association rules; the FP-Growth algorithm constructs an FP tree to compress the data storage space and improve the mining efficiency, directly mines the frequent item sets from the FP tree, and avoids the process of generating a large number of candidate sets in the Apriori algorithm. T3. Customer group analysis: Use the customer behavior results obtained from the association rule mining in step T2 to classify the customers into customer groups, determine the customer group category to which each customer belongs, and at the same time, combine business experience and domain knowledge to manually review and adjust the classification results.

[0006] Further, in the cluster analysis step of T1, the specific implementation methods are as follows: When using the K-Means algorithm, randomly initialize the cluster centers multiple times, and combine the elbow method or the silhouette coefficient method to determine the optimal K value, and perform K-Means clustering calculation to obtain different clustering results. The elbow method plots the relationship curve between the clustering error and the K value, observes the change trend of the curve. When the K value is small, as K increases, the clustering error drops sharply. When K increases to a certain extent, the decline trend of the clustering error slows down. When the curve shape is similar to an elbow, select the K value corresponding to the elbow as the optimal value. The silhouette coefficient method calculates the silhouette coefficient for each sample. The calculation formula for the silhouette coefficient is: si = (bi - ai) / max(ai, bi), where ai represents the average distance between sample i and other samples in the same cluster; bi represents the average distance between sample i and the nearest sample in other clusters; the value range of the silhouette coefficient is [-1, 1]. The closer the value is to 1, the better the clustering effect of the sample. By traversing different K values, the K value when the silhouette coefficient is the largest is selected as the optimal K value.

[0007] Further, the customer group analysis steps of T3 include customer group feature analysis steps: First, comprehensively analyze the characteristics of each customer group, divide the distribution ratio of the customer group according to different ages, genders, regions and occupations, and understand the basic composition of the customer group; Then, in terms of consumption behavior characteristics, analyze the average purchase amount, purchase frequency and preferred purchase categories of the customer group to judge the consumption ability and consumption habits of the customer group; Secondly, by analyzing the product categories and search keywords browsed and purchased by the customer group, excavate the interest and hobby characteristics of the customer group; Finally, use data visualization tools to display the customer group portrait in an intuitive chart form.

[0008] Further, the customer group analysis steps of T3 also include customer group dynamic update steps: First, collect new customer behavior data in real time, and import the new data into the system through message queue and scheduled task technical means; Then update the pre-aggregated storage table, adopt an incremental update strategy, and only update the changed data; Secondly, re-perform big data analysis and customer group classification, and update the customer group portrait through customer group feature analysis; At the same time, set an update threshold, and when the new data volume reaches 3% of the total data volume of the pre-aggregated storage table, trigger the update process.

[0009] Further, in the data collection step of S1, collect customer behavior data and basic information data through web crawler technology, log analysis tools and third-party data interface channels. When using web crawler technology, relevant laws, regulations and the robots protocol of the website need to be followed to ensure legal and compliant data acquisition; According to different website structures and data formats, write corresponding crawler programs and adopt the Scrapy framework of Python; In terms of log analysis tools, adopt the combined technology of Elasticsearch, Logstash and Kibana. Logstash is responsible for collecting and preprocessing log data, Elasticsearch is used to store and retrieve log data, and Kibana provides a visualization interface for log analysis; When docking with a third-party data interface, sign a data use agreement to ensure the security and compliance of the data, and call the interface according to the interface document specification to obtain data and enrich the customer information dimension.

[0010] Further, in the data preprocessing step of S2, the regression prediction filling method uses the linear relationship between other variables in the data and the variable where the missing value is located to establish a regression model for prediction and filling; the multiple imputation method generates multiple complete data sets by randomly generating filling values multiple times, analyzes each data set, and comprehensively analyzes the results to improve the accuracy and reliability of data processing.

[0011] Further, in the step of constructing the pre-aggregated storage table of S4, each node of the B-tree stores multiple key-value pairs and pointers to child nodes, determines the storage location and query path of the data by comparing the key values, and quickly locates the data in the specified time interval; the hash index quickly finds the unique customer identifier, maps the unique customer identifier to a hash value of a fixed length through a hash function, and directly locates the storage location of the data according to the hash value.

[0012] Further, in the S5 data analysis step, the elastic computing resources provided by the cloud computing platform are used to dynamically adjust the computing resource configuration of the analysis task, improve the analysis efficiency and reduce the cost; when performing big data analysis tasks, the computing resources are dynamically adjusted according to the load situation of the task, and when the analysis task is completed or the load decreases, the redundant computing resources are released in time.

[0013] In summary, the present invention provides a user classification method based on a pre-aggregated storage table, which has the following beneficial effects: 1. In the data collection stage of the present invention, by comprehensively using web crawler technology, log analysis tools and third-party data interfaces, customer behavior and basic information data are comprehensively collected, and laws, regulations and website agreements are strictly followed to ensure the legality and compliance of the data. At the same time, in data preprocessing, the 3σ principle, various missing value processing methods, and data smoothing and normalization techniques are adopted to effectively remove noise, fill in missing values, improve data quality, provide an accurate and complete data basis for subsequent feature extraction and analysis, ensure the accuracy of user classification, enable the classification results to truly reflect customer characteristics, and provide a reliable basis for enterprise decision-making.

[0014] 2. The present invention deeply extracts key features such as activity and purchase frequency from the preprocessed data, uses the information gain method and the chi-square test method to screen valuable features, reduces data redundancy, highlights the key information for classification, and at the same time stores the data in a multi-dimensional aggregated manner when constructing the pre-aggregated storage table, combined with distributed storage technology to ensure the security and reliability of the data; designs B-trees and hash indexes for different feature dimensions, can quickly locate and query data, greatly improves data processing efficiency, facilitates enterprises to quickly obtain the required information, and enhances the overall efficiency of data management and analysis, providing an efficient solution for large-scale data processing.

[0015] 3. The present invention accurately divides customer groups through cluster analysis and by using the K-Means algorithm in combination with the elbow method and the silhouette coefficient method, providing a basis for precision marketing and personalized services. Association rule mining uses the Apriori algorithm and the FP-Growth algorithm to discover potential connections in customer behavior, optimize product recommendation strategies, and improve sales conversion rates. Through customer group analysis, not only are the characteristics of customer groups comprehensively analyzed, but also customer group portraits are displayed through data visualization. Combining business experience to review and adjust classification results enables enterprises to deeply understand customers, formulate targeted operation strategies, and improve operation effects and customer satisfaction.

[0016] 4. The present invention ensures that the classification results closely follow changes in customer behavior by setting up a mechanism for dynamically updating customer groups, collecting new data in real time, adopting an incremental update strategy to update the storage table, and re-analyzing and updating customer group portraits when the amount of new data reaches a threshold. At the same time, during the data analysis process, with the help of elastic computing resources on the cloud computing platform, the configuration is dynamically adjusted according to the task load, and resources are released in a timely manner when the task is completed or the load decreases, avoiding waste, reducing computing costs, enhancing the adaptability and economy of the system, enabling enterprises to efficiently utilize resources, and maintaining an advantage in market competition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic structural diagram of the process architecture of a user classification method based on a pre-aggregated storage table according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment: Please refer to Figure 1 As shown, the present invention provides a technical solution: a user classification method based on a pre-aggregated storage table, including the following classification methods: S1. Data collection: First, use diverse technical means and channels, as well as based on historical data and browsing situations, to comprehensively collect customer behavior data and basic information data. And through web crawler technology, according to the set rules and strategies, accurately capture the browsing records of customers from specified web pages, e-commerce platforms, and social media network resources, and detailedly record information such as the page URLs visited by customers, access times, and browsing sequences. Then use log analysis tools to deeply analyze website server logs, APP operation logs, etc., to obtain the operation trajectories of customers within the application, including search records, click behaviors, stay times, etc. Secondly, by docking with third-party data interfaces, obtain authoritative customer basic information data, and obtain information such as the age, gender, and region of customers through the demographic data platform, and obtain the occupation information of customers through industry data service providers. Finally, the collected customer behavior data includes but is not limited to customer browsing records, purchase records, search records, and stay times. At the same time, collect customer basic information data. Collecting data through multiple channels can comprehensively obtain customer information, providing rich basis for accurate classification. At the same time, accurately capturing through web crawler technology and deeply analyzing logs can capture customer behavior details, and third-party data docking supplements basic information, improving the accuracy and comprehensiveness of classification; S2. Data preprocessing: Thoroughly clean the customer data collected in step S1. Adopt the 3σ principle based on statistical methods to identify and remove noise data and outliers. The 3σ principle is based on the assumption of the normal distribution of data. If a data point is more than 3 times the standard deviation away from the mean, it is determined as an outlier. For missing values, different methods are used for processing according to the data characteristics. For numerical data, the mean filling method is used when the data distribution is relatively uniform, and the regression prediction filling method is used when there is a linear relationship in the data. For complex data scenarios, multiple imputation methods are used to generate multiple reasonable imputed values. At the same time, use data smoothing technology to eliminate minor fluctuations in the data and regularly process the data. Adopt normalization technology to improve data quality, and perform zero-mean and unit-variance processing on the data. Through data cleaning and processing, data quality can be improved, noise and outliers can be removed to avoid interfering with the classification results. At the same time, by adopting appropriate missing value processing methods for different data situations, data integrity is ensured, and through normalization processing, the data is more comparable, which is conducive to subsequent analysis; S3. Feature Extraction: Deeply extract features from the preprocessed customer data in step S2. Deeply extract features from the preprocessed customer behavior data. By statistically analyzing the login times and operation frequencies of customers over a period of time, obtain the activity and usage time of customers; by calculating the purchase times of customers within a certain period, obtain the purchase frequency; by accumulating the purchase amount of each customer, obtain the purchase amount; by analyzing the categories of products purchased by customers, determine the preferred categories; and use feature selection algorithms to screen out representative and discriminative features from the extracted features. By deeply extracting features, customer behavior and consumption characteristics can be accurately depicted. By using feature selection algorithms to screen out key features, data redundancy can be reduced, and information valuable for user classification can be highlighted, making the classification more targeted and effective. Feature selection algorithms include the information gain method and the chi-square test method. The information gain method calculates the information gain of each feature for the classification target and selects features with high information gain; the chi-square test method, based on the independence hypothesis between features and the classification target, measures the importance of features. S4. Construct a pre-aggregated storage table: Based on the preprocessed customer data in step S3, construct a pre-aggregated storage table according to the extracted features, and perform aggregated storage and classification aggregation on the data according to the time dimension and attribute dimension. At the same time, aggregate and store the corresponding data according to the behavior type dimension. Use distributed storage technology to store the pre-aggregated storage table. Design an efficient data index structure for different feature dimensions. Use the B-tree index method for range queries on the time dimension; use the hash index method for range queries on the attribute and behavior dimensions. By constructing a pre-aggregated storage table, data management and analysis are facilitated. Aggregated storage according to multiple dimensions improves data query and processing efficiency. Distributed storage ensures data security and reliability. Different index methods meet various query requirements and improve system performance. S5. Data Analysis: Use big data analysis technology to deeply analyze customer behavior data based on the pre-aggregated storage table constructed in step S4. The specific implementation steps are as follows: T1. Cluster Analysis: Use a clustering algorithm to cluster customers, divide customers into different customer groups. Use the K-Means algorithm combined with the elbow method to calculate the clustering error under different K values; use the silhouette coefficient method to calculate the silhouette coefficient of each sample. Cluster analysis can segment customers. The K-Means algorithm combined with the elbow method and the silhouette coefficient method can determine the optimal number of clusters, accurately divide customer groups, provide a basis for precise marketing and personalized services, and improve marketing effectiveness and user satisfaction. T2. Association Rule Mining: Use the Apriori algorithm and the FP-Growth algorithm to mine the association rules between customer behaviors within different customer groups from the customer data obtained by clustering analysis in step T1. The Apriori algorithm is based on the prior principle. By generating candidate frequent item sets and validating them in the dataset, it gradually mines out frequent item sets and association rules. The FP-Growth algorithm constructs an FP tree to compress the data storage space and improve the mining efficiency. It directly mines frequent item sets from the FP tree, avoiding the large number of candidate set generation processes in the Apriori algorithm. Mining association rules can discover potential connections between customer behaviors, understand the combination of customer purchase habits and preferences. Based on this, enterprises can optimize product recommendations and marketing strategies, improving sales conversion rates and customer loyalty. T3. Customer Group Analysis: Based on the customer behavior results obtained from association rule mining in step T2, classify customers into different customer groups to determine the customer group category to which each customer belongs. At the same time, combine business experience and domain knowledge to conduct manual review and adjustment of the classification results. Customer group classification clarifies customer attribution. By combining business experience and domain knowledge for review and adjustment, it ensures that the classification results meet the actual business needs, enabling enterprises to formulate precise operation strategies for different customer groups and improve operation effects.

[0020] In the clustering analysis step of T1, the following specific implementation methods are included: When using the K-Means algorithm, by randomly initializing the clustering centers multiple times and combining the elbow method or the silhouette coefficient method to determine the optimal K value, perform K-Means clustering calculations to obtain different clustering results. The elbow method plots the relationship curve between the clustering error and the K value, and observes the change trend of the curve. When the K value is small, as K increases, the clustering error drops sharply. When K increases to a certain extent, the decline trend of the clustering error slows down. When the curve shape is similar to an elbow, select the K value corresponding to the elbow as the optimal value. The silhouette coefficient method calculates the silhouette coefficient for each sample. The calculation formula for the silhouette coefficient is: si = (bi - ai) / max(ai, bi), where ai represents the average distance between sample i and other samples in the same cluster; bi represents the average distance between sample i and the nearest sample in other clusters; the silhouette coefficient value ranges from [-1, 1]. The closer the value is to 1, the better the clustering effect of the sample. Traverse different K values and select the K value with the largest silhouette coefficient as the optimal K value. The method of randomly initializing the clustering centers multiple times and selecting the optimal K value can avoid the clustering results falling into local optima, improve the accuracy and stability of clustering, more accurately divide customer groups, and help enterprises carry out targeted marketing activities more effectively.

[0021] The customer group analysis steps of T3 include the customer group characteristics analysis step: First, comprehensively analyze the characteristics of each customer group, divide the customer groups according to different ages, genders, regions, and occupations in terms of distribution ratios, and understand the basic composition of the customer groups; Then, in terms of consumption behavior characteristics, analyze the average purchase amount, purchase frequency, and preferred purchase categories of the customer groups to judge the consumption ability and consumption habits of the customer groups; Secondly, by analyzing the product categories and search keywords browsed and purchased by the customer groups, explore the interest and hobby characteristics of the customer groups; Finally, use data visualization tools to display the customer group portraits in an intuitive chart form. Comprehensive customer group characteristics analysis can help enterprises deeply understand the characteristics of different customer groups, and visual display makes the analysis results more intuitive and understandable, facilitating enterprises to formulate product strategies and marketing plans that meet the needs of customer groups and enhance market competitiveness.

[0022] The customer group analysis steps of T3 also include the customer group dynamic update step: First, collect new customer behavior data in real time, and import the new data into the system through technical means such as message queues and scheduled tasks; Then, update the pre-aggregated storage table, adopting an incremental update strategy to only update the changed data; Secondly, re-perform big data analysis, customer group classification, and customer group characteristics analysis to update the customer group portraits; At the same time, set an update threshold. When the amount of new data reaches 3% of the total amount of data in the pre-aggregated storage table, trigger the update process. Customer group dynamic update ensures that the classification results closely follow changes in customer behavior, timely reflect new customer needs and new trends, incremental update saves resources, and setting the update threshold balances update costs and timeliness, enabling enterprises to continuously provide accurate services.

[0023] In the data collection step of S1, collect customer behavior data and basic information data through web crawler technology, log analysis tools, and third-party data interface channels. When using web crawler technology, relevant laws, regulations, and the robots protocol of the website need to be followed to ensure legal and compliant acquisition of data; According to different website structures and data formats, write corresponding crawler programs, adopting the Scrapy framework of Python; In terms of log analysis tools, adopt the combined technology of Elasticsearch, Logstash, and Kibana. Logstash is responsible for collecting and preprocessing log data, Elasticsearch is used for storing and retrieving log data, and Kibana provides a visualization interface for log analysis; When docking with third-party data interfaces, sign a data usage agreement to ensure the security and compliance of the data, and call the interfaces according to the interface document specifications to obtain data, enriching the dimension of customer information. A legal and compliant data collection method guarantees the operational safety of enterprises. The application of different technical tools efficiently obtains multi-source data, enriches the dimension of customer information, and lays a solid foundation for building an accurate user classification model.

[0024] In the data preprocessing step of S2, the regression prediction filling method uses the linear relationship between other variables in the data and the variable where the missing value is located to establish a regression model for prediction filling; the multiple imputation method generates multiple complete data sets by randomly generating filling values multiple times, analyzes each data set, and synthesizes the analysis results to improve the accuracy and reliability of data processing. By using the regression prediction and multiple imputation methods, the data missing problem is effectively handled, the data integrity and accuracy are improved, high-quality data is provided for subsequent feature extraction, model construction and analysis, and the reliability of the user classification result is guaranteed.

[0025] In the step of constructing the pre-aggregated storage table of S4, each node of the B-tree stores multiple key-value pairs and pointers to child nodes. By comparing the key values, the storage location and query path of the data are determined, and the data in the specified time interval can be quickly located; the hash index enables fast lookup of the customer unique identifier. The customer unique identifier is mapped to a hash value of a fixed length through a hash function, and the storage location of the data is directly located according to the hash value. Through the design of the B-tree and hash index, the data query efficiency is greatly improved, the data can be quickly located, time and resources are saved, it is convenient for enterprises to quickly obtain the required customer information, the data processing and analysis speed are improved, and the business response ability is optimized.

[0026] In the data analysis step of S5, the elastic computing resources provided by the cloud computing platform are used to dynamically adjust the computing resource configuration of the analysis task, improve the analysis efficiency and reduce the cost; when performing big data analysis tasks, the computing resources are dynamically adjusted according to the load of the task. When the analysis task is completed or the load decreases, the redundant computing resources are released in time. By using the elastic computing resources of the cloud computing, the resources can be flexibly allocated according to the task requirements, the analysis efficiency is improved, the computing cost is reduced, resource waste is avoided, the adaptability and economy of the system are enhanced, and the resource utilization efficiency of the enterprise is improved.

[0027] The above is only a preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A user classification method based on a pre-aggregated storage table, characterized in that: Including the following classification methods: S1. Data collection: First, use diverse technical means and channels, as well as based on historical data and browsing situations, to comprehensively collect customer behavior data and basic information data. And through web crawler technology, according to the set rules and strategies, accurately capture the browsing records of customers from designated web pages, e-commerce platforms, and social media network resources, and detail the information such as the page URL visited by customers, access time, and browsing order. Then use log analysis tools to deeply analyze website server logs, APP operation logs, etc., to obtain the operation trajectories of customers within the application, including search records, click behaviors, stay times, etc. Secondly, by docking with third-party data interfaces, obtain authoritative customer basic information data, and obtain information such as the age, gender, and region of customers through the demographic data platform, and obtain the occupation information of customers through industry data service providers. Finally, the collected customer behavior data includes, but is not limited to, customer browsing records, purchase records, search records, and stay times, and at the same time collect customer basic information data; S2. Data preprocessing: Conduct a comprehensive cleaning of the customer data collected in step S1. Adopt the 3σ principle based on statistical methods to identify and remove noise data and outliers. The 3σ principle is based on the assumption of the normal distribution of data. If a data point is more than 3 times the standard deviation away from the mean, it is determined as an outlier. For missing values, different methods are used according to the data characteristics. For numerical data, if the data distribution is relatively uniform, the mean filling method is used. If there is a linear relationship in the data, the regression prediction filling method is used. For complex data scenarios, the multiple imputation method is used to generate multiple reasonable filling values. At the same time, use data smoothing technology to eliminate the small fluctuations in the data and regularize the data; Adopt normalization technology to improve data quality, and process the data with zero mean and unit variance; S3. Feature extraction: Deeply extract features from the preprocessed customer data in step S2. Deeply extract features from the preprocessed customer behavior data. Through statistical analysis of the login times and operation frequency data of customers over a period of time, obtain the activity and usage time of customers. By calculating the number of purchases of customers within a certain period, obtain the purchase frequency; By accumulating the purchase amount of each customer, obtain the purchase amount. By analyzing the categories to which the purchased goods of customers belong, determine the preferred categories; and use feature selection algorithms to screen out representative and discriminative features from the extracted features; The feature selection algorithms include the information gain method and the chi-square test method. The information gain method calculates the information gain of each feature for the classification target and selects the features with high information gain. The chi-square test method, based on the independence assumption between the feature and the classification target, measures the importance of the feature; S4. Construct the pre-aggregated storage table: Based on the customer data preprocessed in step S3, construct a pre-aggregated storage table according to the extracted features, and perform aggregated storage and classification aggregation on the data according to the time dimension and attribute dimension. At the same time, aggregate and store the corresponding data according to the behavior type dimension, and use distributed storage technology to store the pre-aggregated storage table. Design an efficient data index structure for different feature dimensions, and use the B-tree index method for range queries in the time dimension; Use the hash index method for range queries in the attribute and behavior dimensions; S5. Data analysis: Use big data analysis technology to deeply analyze the customer behavior data based on the pre-aggregated storage table constructed in step S4. The specific implementation steps are as follows: T1. Cluster analysis: Use a clustering algorithm to cluster customers, divide customers into different customer groups, use the K-Means algorithm combined with the elbow method to calculate the clustering error under different K values; use the silhouette coefficient method to calculate the silhouette coefficient of each sample; T2. Association rule mining: Use the Apriori algorithm and the FP-Growth algorithm to mine the association rules between customer behaviors within different customer groups from the customer data obtained by the cluster analysis in step T1; The Apriori algorithm is based on the prior principle, and gradually mines the frequent item sets and association rules by generating candidate frequent item sets and validating them in the dataset; The FP-Growth algorithm constructs an FP tree to compress the data storage space and improve the mining efficiency, directly mines the frequent item sets from the FP tree, and avoids the large number of candidate set generation processes in the Apriori algorithm; T3. Customer group analysis: Use the customer behavior results obtained by the association rule mining in step T2 to classify the customers into customer groups, determine the customer group category to which each customer belongs, and at the same time, combine business experience and domain knowledge to manually review and adjust the classification results.

2. The user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: In the clustering analysis step of T1, the specific implementation methods are as follows: When using the K-Means algorithm, determine the optimal K value by randomly initializing the clustering center multiple times and combining the elbow method or the silhouette coefficient method, and perform K-Means clustering calculation to obtain different clustering results; The elbow method draws a curve of the clustering error versus the K value, observes the change trend of the curve. When the K value is small, as K increases, the clustering error drops sharply. When K increases to a certain extent, the drop trend of the clustering error slows down. When the curve shape is similar to an elbow, select the K value corresponding to the elbow as the optimal value; The silhouette coefficient method calculates the silhouette coefficient for each sample. The calculation formula of the silhouette coefficient is: si = (bi - ai) / max(ai, bi), where ai represents the average distance between sample i and other samples in the same cluster; bi represents the average distance between sample i and the nearest sample in other clusters; The value range of the silhouette coefficient is [-1, 1]. The closer the value is to 1, the better the clustering effect of the sample. Traverse different K values and select the K value with the largest silhouette coefficient as the optimal K value.

3. A user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: The customer group analysis steps of T3 include the customer group characteristic analysis step: First, comprehensively analyze the characteristics of each customer group, divide the customer group according to different ages, genders, regions and occupations in terms of distribution ratios, and understand the basic composition of the customer group; Then, in terms of consumer behavior characteristics, analyze the average purchase amount, purchase frequency and preferred purchase categories of the customer group to judge the consumption ability and consumption habits of the customer group; Secondly, by analyzing the product categories and search keywords browsed and purchased by the customer group, excavate the interest and hobby characteristics of the customer group; Finally, use data visualization tools to display the customer group portrait in an intuitive chart form.

4. A user classification method based on a pre-aggregated storage table according to claim 3, characterized in that: The customer group analysis steps of T3 also include the customer group dynamic update step: First, collect new customer behavior data in real time, and import the new data into the system through technical means such as message queues and scheduled tasks; Then, update the pre-aggregated storage table, adopt an incremental update strategy, and only update the changed data; Secondly, re-perform big data analysis and customer group classification, and update the customer group portrait through customer group characteristic analysis; At the same time, set an update threshold, and when the amount of new data reaches 3% of the total amount of data in the pre-aggregated storage table, trigger the update process.

5. A user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: In the data collection step of S1, collect customer behavior data and basic information data through web crawler technology, log analysis tools and third-party data interface channels. When using web crawler technology, relevant laws, regulations and the robots protocol of the website need to be followed to ensure the legal and compliant acquisition of data; According to different website structures and data formats, write corresponding crawler programs and adopt the Scrapy framework of Python; In terms of log analysis tools, adopt the combined technology of Elasticsearch, Logstash and Kibana. Logstash is responsible for collecting and preprocessing log data, Elasticsearch is used to store and retrieve log data, and Kibana provides a visualization interface for log analysis; When docking with third-party data interfaces, sign a data use agreement to ensure the security and compliance of the data, and call the interface according to the interface document specifications to obtain data and enrich the dimension of customer information.

6. A user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: In the data preprocessing step of S2, the regression prediction filling method uses the linear relationship between other variables in the data and the variable where the missing value is located to establish a regression model for prediction and filling. The multiple imputation method generates multiple complete data sets after randomly generating imputation values multiple times, analyzes each data set, and synthesizes the analysis results to improve the accuracy and reliability of data processing.

7. A user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: In the step of constructing the pre-aggregated storage table of S4, each node of the B-tree stores multiple key-value pairs and pointers to child nodes. By comparing the key values, determine the storage location and query path of the data, and quickly locate the data in the specified time interval; The hash index enables fast lookup of the customer unique identifier. The customer unique identifier is mapped to a fixed-length hash value through a hash function, and the storage location of the data is directly located according to the hash value.

8. A user classification method based on a pre-aggregated storage table according to claim 1, characterized in that: In the S5 data analysis step, elastic computing resources provided by the cloud computing platform are utilized to dynamically adjust the computing resource configuration of the analysis task, improving the analysis efficiency and reducing costs; when performing big data analysis tasks, the computing resources are dynamically adjusted according to the load situation of the task, and when the analysis task is completed or the load decreases, the redundant computing resources are released in a timely manner.

Citation Information

Patent Citations

  • Data processing method, device and system, computer equipment and storage medium

    CN110008257A

  • Data pre-aggregation method and system, calculation device and storage medium

    CN111090670A

  • Data processing method and device

    CN115098029A

  • Client information management system based on data analysis

    CN118628147A

  • Network user behavior analysis and result presenting system and method thereof

    US20200160390A1

Cited By

  • Entertainment culture content transmission mode recognition method based on improved FP-GROWTH algorithm

    CN121188730A