A method and system for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features
By introducing technical means of big data processing and multi-dimensional features in user portrait construction, the shortcomings of traditional methods in processing massive data and describing users are solved, and more efficient and accurate user portrait construction and personalized services are achieved.
Patent Information
- Application Number
- CN202410937208.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-12
AI Technical Summary
Traditional personalized user portrait construction methods are difficult to efficiently process massive data, cannot describe users in a comprehensive and accurate manner, and cannot integrate multiple data sources and multi-dimensional characteristics.
Using a method based on big data processing and multi-dimensional features, through unified storage of data lakes and Hive data acquisition, combined with Spark's distributed LightGBM model training, static and dynamic user portrait models are built, and personalized user portrait applications are carried out through the scoring computing system.
It realizes efficient processing and analysis of massive user data, can describe users' needs and behavioral characteristics more comprehensively and accurately, provide more personalized services, and improve data processing efficiency and user service perception.
Smart Images

Figure CN118780837B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of new information technology, and specifically, relates to a method and system for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features. Background Art
[0002] With the rapid development of information technology and the popularization of the Internet, personalized services have become an important means to improve user experience and market competition. Personalized user portraits are the key to achieving personalized services. They can form a comprehensive understanding and description of users by analyzing and organizing data such as personal characteristics, interest preferences, and behavioral habits of users.
[0003] However, there are some problems with the traditional method of building personalized user portraits. First, with the rapid development of the Internet, the data generated by users has exploded, and how to efficiently process this big data has become a challenge. Second, personalized user portraits require the integration of multiple data sources and multi-dimensional features, but current methods often only focus on certain specific data or features, making it difficult to fully and accurately describe users.
[0004] Therefore, it is necessary to provide a method and system for building a personalized user portrait application based on big data processing and multi-dimensional features to solve the above problems. The method and system can efficiently process and analyze massive amounts of user data through effective big data processing technology to extract key features and behavior patterns of users. At the same time, it can also integrate multiple data sources and multi-dimensional features to fully describe the personalized needs and preferences of users.
[0005] The existing technology has some shortcomings, such as:
[0006] 1. Data processing speed is slow and big data cannot be processed efficiently.
[0007] 2. Focusing only on certain data or features makes it difficult to fully and accurately describe users.
[0008] 3. It is impossible to integrate multiple data sources and multi-dimensional features, and it is impossible to fully describe users' personalized needs and preferences. Summary of the invention
[0009] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a method and system for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features. The method and system can efficiently process and analyze massive amounts of user data through effective big data processing technology, and extract key features and behavior patterns of users. At the same time, it can also integrate multiple data sources and multi-dimensional features to fully describe the personalized needs and preferences of users.
[0010] To solve the above problems, the technical solution provided by the present invention is:
[0011] A method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features includes the following steps:
[0012] S1. Based on business needs, data from different sources are imported into the data lake for unified storage and management. Then, data is collected from the data lake based on Hive, and ETL-related data processing is completed to form a large wide table.
[0013] S2. Based on business and modeling requirements, determine and process the features in the large wide table in step S1, and use a three-level hierarchical architecture to classify the dimensions, i.e., feature objects, primary classification, and secondary classification, to obtain a data wide table;
[0014] S3. For the data wide table obtained in step S2, Spark-based distributed LightGBM model training is used to find important feature factors and target labels:
[0015] S4. According to the important characteristic factors and target labels obtained in step S3, target labels and judgment probabilities are obtained for users based on static portrait models and dynamic portrait models, and re-screened and labeled through thresholds to further refine the user group;
[0016] S5. Combined with the labels, a scoring calculation system is constructed based on the acquired dynamic and static important characteristic factors and the multi-dimensional classification;
[0017] S6. According to the scoring calculation system determined in step S5, dynamic and static portrait scores are performed on the users tagged with the target tags, and the scores are combined to complete the personalized user portrait application.
[0018] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0019] In step S1, data from different sources are imported into the data lake based on business needs, where the data from different sources are related platform data and text data;
[0020] In step S1, data collection is performed from the data lake based on Hive, where the data collection involves traffic statistics, subscription relationships, billing, terminals, locations, and channel services.
[0021] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0022] The processing in step S2 includes basic processing, classification processing, dimensionality reduction and merging processing, and content conversion processing; among which the basic processing involves type conversion, unit format conversion consistency, and function statistics; among which the classification processing involves labeling a large amount of data in combination with business classification; among which the dimensionality reduction and merging processing involves dimensionality reduction processing of high-dimensional features to meet the subsequent model availability standards; among which the content conversion processing involves the conversion between numerical content and character content, and selecting a more suitable content representation form to express it.
[0023] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0024] Step S3 also includes constructing a static user portrait model and a dynamic user portrait model, wherein the static user portrait model is used to describe user characteristics; and the dynamic user portrait model is constructed to observe recent changes in users.
[0025] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0026] The method of constructing a static user portrait model in step S3 is as follows:
[0027] Take this month's data from the data wide table as a static user portrait modeling dataset and use it for subsequent Spark-based distributed LightGBM model training;
[0028] The method of constructing a dynamic user portrait model in step S3 is as follows:
[0029] The changes in data between this month and last month in the data wide table are taken as the dynamic user portrait modeling data set and used for subsequent Spark-based distributed LightGBM model training.
[0030] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0031] In step S4, the operations of obtaining target labels and judging probabilities for users based on the static portrait model and the dynamic portrait model are as follows:
[0032] Output labels and probabilities for users on the entire network based on the optimal static user model, then further filter static label users based on the set threshold, and then output labels and probabilities for users on the entire network based on the optimal dynamic user model.
[0033] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0034] In step S5, for the character features of the dynamic and static important feature factors, a character feature score calculation rule is constructed, and the proportion of each category in each feature in the target label and the non-target label is calculated respectively, and the score is calculated using 3 times the standard deviation;
[0035] In step S5, for the numerical features of the dynamic and static important feature factors, a numerical feature scoring calculation rule is constructed to determine the relevance of each feature with the target label, and the data is mapped using maximum and minimum normalization to calculate the score.
[0036] The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features,
[0037] The operation of scoring the dynamic and static portraits of the user with the target tag in step S6 is as follows:
[0038] Perform static portrait scoring on the statically labeled users after labeling to find out the user's shortcomings and classify them into groups;
[0039] Conduct dynamic portrait scoring on the dynamic tag users who have been screened after labeling to find the timing and strategy for intervention.
[0040] A personalized traffic user portrait application construction system based on big data processing and multi-dimensional features.
[0041] The system comprises a data collection and processing module, a multi-dimensional feature construction module, a dynamic and static user portrait construction module and a portrait application module connected to each other;
[0042] The data collection and processing module is connected to the multi-dimensional feature construction module and the dynamic and static user portrait construction module respectively, and the multi-dimensional feature construction module is connected to the portrait application module through the dynamic and static user portrait construction module;
[0043] Data collection and processing module: complete the collection and preliminary processing of big data, and complete the data processing work before user portrait modeling;
[0044] Multi-dimensional feature construction module: Combines business and modeling requirements to complete feature judgment and processing, and multi-dimensional hierarchical processing;
[0045] Dynamic and static user portrait construction module: With the support of the data collection and processing module, complete the construction of dynamic and static user portrait models and output important feature factors and target labels;
[0046] Portrait application module: Filter and circle users with target tags, build a user rating calculation system based on dynamic and static user portrait models, and complete personalized user portrait application.
[0047] Beneficial effects:
[0048] The present invention provides a method and system for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features to solve the problems existing in traditional methods when constructing personalized traffic user portraits. The method and system make full use of big data analysis technology, integrate different types of data and multi-dimensional features, and achieve a comprehensive understanding and description of users, thereby providing users with more personalized services. The method and system of the present invention can more comprehensively and accurately describe the needs and behavioral characteristics of users, thereby providing users with more personalized services. By constructing a personalized traffic user portrait, user needs can be better met and the quality of personalized services can be improved. The method and system of the present invention make full use of big data analysis technology to efficiently process and analyze massive traffic user data. Compared with traditional methods, the present invention can complete the construction of user portraits more quickly and accurately, and improve data processing efficiency. Through technical means based on big data processing and multi-dimensional features, the method and system of the present invention can provide enterprises with accurate traffic user portraits and personalized services, thereby improving user service perception, enhancing the competitiveness of enterprises, and creating more commercial value. Therefore, the present invention can provide a method and system that can make full use of big data analysis technology, integrate different types of data and multi-dimensional features, and construct personalized traffic user portraits, so as to improve the quality of personalized services, improve data processing efficiency, and create commercial value.
[0049] This method and system not only solves the speed and accuracy problems of traditional methods in processing massive data, but also achieves accurate mining and comprehensive description of user needs by introducing multi-dimensional feature data and dynamic and static user portrait scoring systems, providing fine support for personalized recommendations and services. Specifically, as follows:
[0050] (1) Introducing more dimensional feature data to form a three-level hierarchical architecture:
[0051] Introducing more dimensional feature data is a key step in building an efficient user portrait. Through the three-level hierarchical architecture, the user portrait can be expressed more clearly and intuitively, and more dimensional feature data can be reasonably combined and utilized.
[0052] The details include:
[0053] Demographic information: includes name, gender, phone number, email address, home address, etc. This information can help companies contact customers and promote their products and services to them.
[0054] Credit information: used to describe the user's income potential and income situation, and ability to pay. The customer's occupation, income, assets, liabilities, education, credit score, etc. are all credit information.
[0055] Consumption characteristics: used to describe customers’ main consumption habits and preferences, and to find high-frequency and high-value customers. It helps companies recommend relevant financial products and services based on customers’ consumption characteristics, and the conversion rate will be very high.
[0056] (2) Build two types of user portraits, dynamic and static, and combine them with a user portrait scoring system:
[0057] The purpose of building dynamic and static user portraits is to better explore user needs, understand users more accurately, and improve user service perception, thereby providing more sophisticated support for personalized recommendations and services. Specifically, it includes the following:
[0058] Static user portrait: focus on the user's static attribute characteristics, such as age, gender, region, occupation, education level, etc.
[0059] Dynamic user portrait: focus on the user's dynamic behavior characteristics, such as browsing history, search history, purchase history, etc.
[0060] (3) Improve the user portrait scoring system:
[0061] The user portrait scoring system is designed to quantify the characteristics of user portraits. By scoring user portraits, we can better understand user needs and preferences and improve user service perception. The specifics include the following:
[0062] Scoring calculation rules: score the characteristics of the user portrait, such as the user's age, gender, occupation, education level, etc.
[0063] Scoring weight: Determine the weight of each feature based on business needs. For example, the weight of age may be higher than that of gender.
[0064] Rating results: Based on the rating calculation rules and rating weights, the user's total rating is calculated for personalized recommendations and services.
[0065] (4) Expand application scenarios:
[0066] Personalized recommendations: Recommend relevant financial products and services based on user profiles and rating results to improve user satisfaction and conversion rates.
[0067] Precision marketing: Carry out precision marketing based on user portraits and scoring results to improve marketing effectiveness and efficiency.
[0068] Customer relationship management: Manage customer relationships and improve customer satisfaction and loyalty based on user profiles and scoring results.
[0069] This technical effect introduces more dimensional feature data to form a three-level hierarchical architecture, constructs two types of dynamic and static user portraits, and combines a user portrait scoring system. It can better explore user needs, understand users more accurately, and improve user service perception, thereby providing more refined support for personalized recommendations and services. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A flowchart of the method for building a personalized traffic user profiling application based on big data processing and multi-dimensional features;
[0071] Figure 2 This is a schematic diagram of a personalized traffic user portrait application system based on big data processing and multi-dimensional features. DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] Combination Figure 1 From the perspective of the present invention, the personalized traffic user portrait application construction method based on big data processing and multi-dimensional features includes the following steps:
[0074] S1. Based on business needs, data from different sources are imported into the data lake for unified storage and management. After data is collected from the data lake based on Hive, data processing such as ETL is completed to form a large wide table.
[0075] The specific steps of step S1 include:
[0076] S11. According to business needs, relevant platform data and text data from different sources are put into the data lake for unified storage and management; in the modern data processing environment, the data lake is a very effective data storage and management solution. The data lake can store structured and unstructured data from multiple sources, including platform data, text data, log data, etc. By storing these data in the data lake, subsequent data processing and analysis can be facilitated.
[0077] Diversified data sources: Data lakes can receive data from different platforms, such as customer relationship management systems (CRM), enterprise resource planning systems (ERP), social media platforms, sensor data, etc.
[0078] Unified storage management: Data lakes use distributed storage technology to efficiently manage and store massive amounts of data, ensuring high data availability and reliability.
[0079] S12. Collect data from the data lake based on Hive (involving traffic statistics, ordering relationships, billing, terminals, locations, and channel services); Hive is a data warehouse tool based on Hadoop that can map structured data files to a database table and provide SQL-like query functions. Through Hive, you can easily collect the required data from the data lake.
[0080] The scope of data collection is wide: the collected data includes traffic statistics, subscription relationships, billing information, terminal information, location data, channel services, etc. These data can fully reflect the user's behavior and characteristics.
[0081] Efficient data query: Hive provides efficient data query capabilities and can quickly extract required information from massive data.
[0082] S13. After collection, a series of data processing operations such as ETL are performed; ETL (extraction, transformation, loading) is a key step in data processing. Through ETL operations, the original data can be converted into a format suitable for analysis and use.
[0083] Data extraction: Extract the required data from the data lake.
[0084] Data conversion: Clean the extracted data, convert the format, aggregate the data, and ensure the quality and consistency of the data. For example, remove duplicate data, process missing values, and perform unit conversion.
[0085] Data loading: Load the processed data into the target data storage to form a large wide table.
[0086] S14. Aggregate the processed data into a large wide table. A large wide table is a table containing a large number of fields that can fully reflect the various characteristics and behavior data of users. Aggregating the processed data into a large wide table can provide rich data support for subsequent data analysis and modeling.
[0087] Data integration: Integrate data from different sources and different types into a large wide table to form a unified data view.
[0088] Data richness: Large wide tables contain multi-dimensional feature data of users and can fully reflect user behaviors and characteristics.
[0089] S2. After judging and partially processing the features in the large wide table in combination with the business, these features are dimensionally classified using a three-level hierarchical architecture, namely: feature object, primary classification, and secondary classification.
[0090] The specific steps of step S2 include:
[0091] S21. Combine business and modeling requirements to determine whether the features in the large wide table are attributes to be processed, and then process them. The processing mainly involves basic processing, classification processing, dimensionality reduction and merging processing, and content conversion processing:
[0092] In practical application,
[0093] In the process of big data processing and analysis, feature engineering is a crucial part. The purpose of feature engineering is to generate more meaningful features that are more suitable for model use by processing and transforming raw data. The following are the specific steps of feature processing:
[0094] (1) Basic processing
[0095] Basic processing is the first step in feature engineering, which mainly includes the following:
[0096] Type conversion: Convert data types uniformly, such as converting string data to numeric data, to facilitate subsequent calculations and processing.
[0097] Unit format conversion: Ensure that the units and formats of all data are consistent, such as converting all time data into a unified time format and converting all currency data into a unified currency unit.
[0098] Function statistics: Perform basic statistical analysis on the data, such as calculating the mean, median, standard deviation, etc., to provide a reference for subsequent feature processing.
[0099] (2) Classification processing
[0100] Classification processing is the process of classifying and labeling data according to business needs:
[0101] Business classification and labeling: Classify and label a large amount of data according to business needs. For example, classify user behavior data according to different behavior types, and classify user consumption data according to different consumption categories.
[0102] (3) Dimensionality reduction and merging processing
[0103] Dimensionality reduction and merging processing is to reduce the dimension of high-dimensional features to reduce the dimension of data and improve the computational efficiency and performance of the model:
[0104] Dimensionality reduction processing: Dimensionality reduction algorithms such as principal component analysis (PCA) and linear discriminant analysis (LDA) are used to convert high-dimensional features into low-dimensional features to ensure the availability of data in the model.
[0105] Feature merging: Merge highly correlated features to reduce the number of features and simplify the complexity of the model.
[0106] (4) Content conversion processing
[0107] Content conversion processing is to convert numeric content into character content and vice versa, and select a more suitable content representation to express the characteristics:
[0108] Numerical content conversion: For example, convert "total user overflow traffic" into "whether user traffic overflows", and convert "user usage traffic" of various apps into "whether users generate high-traffic app traffic (top 20)".
[0109] Character content conversion: For example, convert the "permanent base station coverage type" into a numeric type of 1 to 9 to facilitate model processing.
[0110] S22. Combined with the business, the processed features are classified into multiple dimensions using a three-level hierarchical architecture. The feature itself is the feature object at the lowest level. On this basis, a preliminary secondary classification is summarized and all features are labeled; (Note: The secondary classification includes package attributes, usable quantity, business handling, SMS reminders, outbound calls, complaints, rate changes, payments, consultations, complaints, channel interactions, Wing Payment, roaming, historical data, resident base station features, time period cumulative values, hardware, basic features and other dimensions).
[0111] The specific application is as follows:
[0112] After feature processing is completed, the features need to be classified in order to better organize and manage data. The three-level hierarchical architecture includes feature objects, primary classification, and secondary classification.
[0113] Trait Objects
[0114] Feature objects are the features themselves, and are at the lowest level. Each feature object represents a specific data feature, such as the user's age, gender, and spending amount.
[0115] First level classification
[0116] The first-level classification is a preliminary classification of feature objects, which mainly includes the following six dimensions:
[0117] Behavior: User behavior data, such as browsing history, click history, purchase history, etc.
[0118] Package: User’s package information, such as the type of package ordered, package cost, etc.
[0119] Content: information about the content accessed by users, such as web pages visited, videos watched, etc.
[0120] Location: User's geographic location information, such as the city where they live, the base station where they live, etc.
[0121] Terminal: information about the terminal device used by the user, such as mobile phone model, operating system, etc.
[0122] User basic information: basic information of the user, such as age, gender, occupation, etc.
[0123] Secondary classification
[0124] On the basis of the primary classification, the features are further subdivided to form the secondary classification. The secondary classification includes but is not limited to the following dimensions:
[0125] Package attributes: Specific attributes of the package, such as data package, voice package, etc.
[0126] Available capacity: The amount of resources available to the user, such as remaining traffic, remaining voice minutes, etc.
[0127] Business processing: information about business processed by users, such as activated value-added services, ordered business packages, etc.
[0128] SMS reminder: SMS reminder information received by users, such as bill reminders, traffic reminders, etc.
[0129] Outbound calls: user’s outbound call records, such as the number of calls made, call duration, etc.
[0130] Complaints: User complaint records, such as the number of complaints, complaint types, etc.
[0131] Rate changes: information on changes in user rates, such as rate adjustments, package changes, etc.
[0132] Payment: User’s payment record, such as payment amount, payment time, etc.
[0133] Consultation: User’s consultation records, such as the questions asked, number of consultations, etc.
[0134] Channel interaction: records of interactions between users and channels, such as online customer service interactions, offline store interactions, etc.
[0135] YiPay: records of users using YiPay, such as payment amount, number of payments, etc.
[0136] Roaming: The user's roaming records, such as roaming countries, roaming duration, etc.
[0137] Historical data: historical data records of users, such as historical consumption records, historical behavior records, etc.
[0138] Permanent base station characteristics: characteristics of the user's permanent base station, such as base station coverage type, base station signal strength, etc.
[0139] Cumulative value during a time period: The cumulative value of a user during a specific time period, such as monthly traffic usage, quarterly consumption amount, etc.
[0140] Hardware: Information about the hardware devices used by users, such as device model, device brand, etc.
[0141] Basic characteristics: basic characteristic information of users, such as age, gender, occupation, etc.
[0142] S23. On the basis of the secondary classification, the top-level classification is further summarized, and all features are also labeled. (Note: The primary classification is divided into six dimensions: behavior, package, content, location, terminal, and basic user information).
[0143] After completing the secondary classification, it is necessary to further summarize the top-level primary classification and label all features. The primary classification includes the following six dimensions:
[0144] Behavior: User behavior data.
[0145] Package: User’s package information.
[0146] Content: Information about the content accessed by the user.
[0147] Location: User's geographic location information.
[0148] Terminal: information about the terminal device used by the user.
[0149] User basic information: basic information of the user.
[0150] Through the classification of the three-level hierarchical architecture, feature data can be better organized and managed to provide support for subsequent modeling and analysis. This classification method can not only improve the efficiency of data processing, but also enhance the interpretability and usability of data.
[0151] Step S2 judges and processes the features in the large wide table in combination with business needs, and uses a three-level hierarchical architecture for dimension classification, ensuring the high quality and high availability of feature data. More meaningful feature data is generated through basic processing, classification processing, dimensionality reduction and merging processing, and content conversion processing. Using a three-level hierarchical architecture for classification not only improves the efficiency of data processing, but also enhances the interpretability and availability of data, providing a solid foundation for subsequent modeling and analysis.
[0152] S3. Based on the data wide table after completing big data processing and multi-dimensional feature fusion, Pyspark in Spark is used to train the distributed machine learning LightGBM model to find important feature factors and target labels: build a static user portrait model to describe user characteristics; build a dynamic user portrait model to observe recent changes in users.
[0153] The specific steps of step S3 include:
[0154] S31. Take the data of this month from the large wide table as the static user portrait modeling data set. Because the distribution of traffic users is uneven, there are ultra-high traffic users (5G cards for testing and cards used as hotspots) at the top and a large number of low-zero users at the bottom. Therefore, the "box curve method" is used to take the upper quartile as the high value of high-traffic users. Based on this data, 700,000 users across the province are taken as the low value of high-traffic users. Based on this range, the target label 1 is circled, and the remaining users are label 0;
[0155] S32, after performing a series of big data processing operations on the above data set using Pyspark in Spark, the data set is divided into a training set and a test set in a ratio of 8:2, and static important feature factors and weights and the optimal model are output;
[0156] Among them, a series of big data processing operations include: missing filling processing, outlier processing, dimensionality reduction processing based on Pearson correlation coefficient, discrete feature one-hot conversion, data standardization processing, etc.
[0157] S33. For the large wide table, the changes in data between this month and last month are taken as the dynamic user portrait modeling data set, where the numerical data are taken as the change difference and change range, and the character data are taken as whether it has changed; therefore, the data presents a normal distribution, and "change difference > median" and "change range > median" are selected as traffic growth users, and they are circled as target label 1, and the remaining users are label 0;
[0158] S34. After performing a series of big data processing operations on the above data set using Pyspark in Spark, the data set is divided into a training set and a test set in a ratio of 8:2, and the dynamic important feature factors and weights and the optimal model are output.
[0159] The above practical application can be implemented in the following ways:
[0160] Step S3: Based on the data wide table after completing big data processing and multi-dimensional feature fusion, use Pys park in Spark to train the distributed machine learning LightGBM model.
[0161] The specific steps of step S3 include:
[0162] S31: Take this month's data from the large wide table as a static user portrait modeling dataset
[0163] When building a static user portrait model, you first need to extract this month's data from the large wide table. Since the distribution of traffic users is uneven, there are ultra-high traffic users at the top (such as tested 5G cards and hotspot cards) and a large number of low-zero users at the bottom, so the "box line method" is needed for data processing.
[0164] -Dataset selection: Use the "box plot method" to take the upper quartile as the high value of high-traffic users, and use this data to take 700,000 users across the province as the low value of high-traffic users. Based on this range, the target label 1 is circled, and the remaining users are label 0.
[0165] S32: After a series of big data processing operations are performed on the above data set using Pyspark in Spark, it is divided into a training set and a test set in a ratio of 8:2.
[0166] After the data set selection is completed, a series of big data processing operations need to be performed on the data to ensure the quality and consistency of the data. The specific operations include:
[0167] - Missing value filling: Use methods such as mean, median or mode to fill missing values to ensure data integrity.
[0168] -Outlier processing: Identify and process outliers in the data to avoid negative impact on model training.
[0169] -Dimensionality reduction based on Pearson correlation coefficient: calculate the correlation between features, remove redundant features, and reduce data dimensions.
[0170] -Discrete feature one-hot conversion: Convert discrete features into one-hot encoding for easier model processing.
[0171] -Data standardization: Scale the data to a uniform range to improve the training effect of the model.
[0172] After completing the above processing, the data is divided into a training set and a test set in a ratio of 8:2 for model training. Through training, the static important feature factors and their weights are output, and the optimal model is obtained.
[0173] S33: Take the changes in data between this month and last month from the large wide table as the dynamic user portrait modeling data set
[0174] When building a dynamic user portrait model, it is necessary to consider changes in user behavior. The specific steps are as follows:
[0175] -Dataset selection: The change between this month and last month's data is used as the dynamic user portrait modeling data set. For numerical data, the change difference and change range are used respectively, and for character data, whether the data has changed is used.
[0176] - Target label determination: Since the data presents a normal distribution, we select "change difference > median" and "change amplitude > median" as traffic growth users, circle them as target label 1, and the remaining users are label 0.
[0177] S34: After performing a series of big data processing operations on the above data set using Pyspark in Spark, the data set is divided into a training set and a test set in a ratio of 8:2.
[0178] Similarly, a series of big data processing operations are performed on the dynamic user portrait modeling dataset, including:
[0179] - Missing value filling: use methods such as mean, median or mode to fill missing values.
[0180] -Outlier handling: Identify and handle outliers in the data.
[0181] -Dimensionality reduction based on Pearson correlation coefficient: remove redundant features and reduce data dimension.
[0182] -Discrete feature one-hot conversion: Convert discrete features into one-hot encoding.
[0183] -Data normalization: scaling the data to a uniform range.
[0184] After processing, the data is divided into a training set and a test set in a ratio of 8:2 for model training. Through training, dynamic important feature factors and their weights are output, and the optimal model is obtained.
[0185] Use the LightGBM model in Pyspark for training. The specific implementation steps are as follows:
[0186] Environment configuration: First, configure the Spark environment and introduce necessary dependency packages.
[0187]
[0188] Data preprocessing: Preprocess the data, including missing value filling, outlier processing, one-hot encoding and standardization.
[0189] ```Python
[0190] from pyspark.ml.feature import StringIndexer,OneHotEncoder,VectorAssembler
[0191] from pyspark.ml.feature import StandardScaler
[0192] Sample code
[0193] indexer = StringIndexer(inputCol="category", outputCol="categoryIndex")
[0194] encoder = OneHotEncoder(inputCol="categoryIndex", outputCol="categoryVec")
[0195] assembler = VectorAssembler(inputCols=["feature1", "feature2", "categoryVec"], outputCol="features")
[0196] scaler = StandardScaler(inputCol="features", outputCol="scaledFeatures")
[0197] Model training: Use LightGBMClassifier to train the model.
[0198] ```python
[0199] classifier = LightGBMClassifier(featuresCol="scaledFeatures", labelCol="label", learningRate = 0.3, numIterations = 150, numLeaves = 100)
[0200] pipeline = Pipeline(stages=[indexer, encoder, assembler, scaler, classifier])
[0201] model = pipeline.fit(trainingData)
[0202] Model evaluation: Evaluate the model and output important feature factors and their weights.
[0203] ```Python
[0204] from pyspark.ml.evaluation import BinaryClassificationEvaluator
[0205] predictions=model.transform(testData)
[0206] evaluator=BinaryClassificationEvaluator(labelCol="label",rawPredictionCol="prediction")
[0207] accuracy=evaluator.evaluate(predictions)
[0208] print(f"Test Accuracy:{accuracy}")
[0209] Through the above steps, based on the data wide table after completing big data processing and multi-dimensional feature fusion, the distributed machine learning LightGBM model is trained using Pyspark in Spark, which can effectively build static and dynamic user portrait models. The static model is used to describe user characteristics, and the dynamic model is used to observe recent changes in users. Through model training and evaluation, important feature factors and their weights can be output to provide support for subsequent personalized recommendations and services.
[0210] S4. Obtain target labels and judgment probabilities for users based on static portrait models and dynamic portrait models, re-screen and label through thresholds to further refine the user group.
[0211] The specific steps of step S4 include:
[0212] S41, output labels and probabilities for all network users based on the optimal static user model; after completing the training of the static user portrait model, the next step is to apply the model to all network users and output each user's label and its corresponding probability. The specific implementation of this step is as follows:
[0213] Model application: Use the trained static user portrait model to predict the data of users across the entire network. The model will output the probability of each user belonging to the target label (such as high-value users, potential churn users, etc.) based on the user's characteristics.
[0214] Label output: According to the prediction results of the model, each user is labeled accordingly. For example, if the probability that a user is predicted to be a high-value user is 0.85, the user will be labeled as a "high-value user".
[0215] S42, further filter static label users according to the set threshold: after obtaining the labels and probabilities of all users, it is necessary to further filter out static label users that meet the conditions according to the set threshold. The specific implementation of this step is as follows:
[0216] Threshold setting: Set a reasonable probability threshold based on business needs and model evaluation results. For example, if the threshold is set to 0.7, only users with a predicted probability greater than 0.7 will be screened out.
[0217] User screening: All users are screened according to the set threshold. Only users with predicted probabilities greater than the threshold are retained, and other users are filtered out.
[0218] S43, output labels and probabilities for all network users based on the optimal dynamic user model; after completing the label output and screening of static user portraits, the next step is to apply the dynamic user portrait model to all network users and output each user's dynamic label and its corresponding probability. The specific implementation of this step is as follows:
[0219] Model application: Use the trained dynamic user portrait model to predict the data of users across the entire network. The model will output the probability of each user belonging to the target label (such as recent traffic growth users, recent active users, etc.) based on the user's recent behavior changes.
[0220] Label output: According to the prediction results of the model, each user is labeled with a corresponding dynamic label. For example, if the probability that a user is predicted to be a user with recent traffic growth is 0.75, the user will be labeled as a "user with recent traffic growth".
[0221] S44, further screening dynamic tag users according to the set threshold. After obtaining the dynamic tags and probabilities of all users, it is necessary to further screen out dynamic tag users that meet the conditions according to the set threshold. The specific implementation of this step is as follows:
[0222] Threshold setting: Set a reasonable probability threshold based on business needs and model evaluation results. For example, if the threshold is set to 0.6, only users with a predicted probability greater than 0.6 will be screened out.
[0223] User screening: All users are screened according to the set threshold. Only users with predicted probabilities greater than the threshold are retained, and other users are filtered out.
[0224] The following is a sample code for model prediction and threshold screening using Pyspark:
[0225] Python
[0226] from pyspark.sql import SparkSession
[0227] from pyspark.ml.classification import LightGBMClassifier
[0228] from pyspark.ml.feature import VectorAssembler
[0229] from pyspark.ml import Pipeline
[0230] Initializing SparkSession
[0231] spark=SparkSession.builder.appName("User Profiling").getOrCreate()
[0232] Loading data
[0233] data=spark.read.csv("user_data.csv", header=True, inferSchema=True)
[0234] Feature Engineering
[0235] assembler=VectorAssembler(inputCols=["feature1","feature2","feature3"],outputCol="features")
[0236] data=assembler.transform(data)
[0237] Load the trained static user portrait model
[0238] static_model=LightGBMClassifier.load("static_model_path")
[0239] Predict static labels and probabilities
[0240] static_predictions=static_model.transform(data)
[0241] static_predictions=static_predictions.withColumnRenamed("prediction","static_label")
[0242] static_predictions=static_predictions.withColumnRenamed("probability","static_probability")
[0243] Filter users with static tags based on thresholds
[0244] static_threshold=0.7
[0245] static_filtered=static_predictions.filter(static_predictions.static_probability[1]>static_threshold)
[0246] Load the trained dynamic user portrait model
[0247] dynamic_model=LightGBMClassifier.load("dynamic_model_path")
[0248] Predict dynamic labels and probabilities
[0249] dynamic_predictions=dynamic_model.transform(data)
[0250] dynamic_predictions=dynamic_predictions.withColumnRenamed("prediction","dynami c_label")
[0251] dynamic_predictions=dynamic_predictions.withColumnRenamed("probability","dynam ic_probability")
[0252] Filter dynamic tag users based on thresholds
[0253] dynamic_threshold=0.6
[0254] dynamic_filtered=dynamic_predictions.filter(dynamic_predictions.dynamic_probability[1]>dynamic_threshold)
[0255] Output the filtered users
[0256] static_filtered.show()
[0257] dynamic_filtered.show()
[0258] Through the above steps, based on static and dynamic user portrait models, we can effectively predict and filter labels for users across the entire network. Static models are used to describe the long-term characteristics of users, and dynamic models are used to observe recent changes in users. By setting reasonable thresholds, users can be further filtered and labeled, thereby achieving precision in user groups and providing support for personalized recommendations and services.
[0259] S5. Construct a scoring calculation system based on the acquired dynamic and static important characteristic factors and the multi-dimensional classification.
[0260] The specific steps of step S5 include:
[0261] S51. Construct character feature scoring calculation rules for character features. The specific steps are as follows:
[0262] (1) Calculate the proportion of each category in each feature for target label 1 and non-target label 0 respectively;
[0263] (2) If the target label 1 accounts for a larger proportion of the category than the label 0, the excess is calculated;
[0264] (3) If the target label 1 accounts for a smaller proportion of the class than label 0, then the magnitude of the lower ratio is calculated;
[0265] (4) For the above two amplitudes (both positive values), based on the “3σ criterion”, if it is higher than 3 times the standard deviation, the score is 0, and for others, the score is the ratio of the amplitude to 3 times the standard deviation.
[0266] The formula for calculating 3 times the standard deviation is:
[0267]
[0268] S52. For numerical features, construct numerical feature scoring calculation rules, the specific steps are as follows:
[0269] (1) Determine the correlation between each feature and the target label, which can be divided into three cases: positive correlation, negative correlation, and low correlation.
[0270] (2) For the case of positive correlation, since the data in the real world may contain negative numbers, we first convert the sample data through y=e x The positive correlation is retained and mapped to the range (0, +∞), and then the data is mapped to the range (0, 1] through maximum and minimum normalization as the score.
[0271] The principle of maximum and minimum normalization is as follows:
[0272]
[0273] (3) For the case of negative correlation, since the data in the real world may contain negative numbers, we first convert the sample data through y=e -x Negative correlations are retained and mapped to the range (0, +∞), and then the data is mapped to the range (0, 1] through maximum and minimum normalization as the score.
[0274] (4) For low correlation, the score is recorded as 0.
[0275] S53: Obtain the detailed score of each feature of the scoring user. The obtaining rules are as follows:
[0276] User feature detailed score = feature score + feature weight
[0277] S54. Based on the new dimension of the secondary classification, the detailed scores of the scoring users under each new dimension are obtained. The obtaining rule is: based on the secondary classification, the detailed scores of the features of the user under the same category are summed up. At the same time, the feature weights of the features under the same category can also be summed up to obtain the feature dimension proportion under the new classification dimension.
[0278] S55. Based on the new dimension of the primary classification, the actual score of the scoring user under each new dimension is obtained. The obtaining rule is: based on the primary classification, the detailed scores of the user's secondary classification dimension features under the same primary classification are summed up. At the same time, the feature weights of the features under the same classification can also be summed up to obtain the feature dimension proportion under the new classification dimension.
[0279] S56. Summarize the actual scores of the six dimensions of the first-level classification to obtain the total score of the user.
[0280] S6. According to the determined scoring calculation system, the users with the target label are scored with static portraits to find the user's shortcomings and classify them into groups; the users with the target label are scored with dynamic portraits to find the time and strategy for intervention. The dynamic and static situations are combined to complete the personalized user portrait application.
[0281] The specific steps of step S6 include:
[0282] S61. According to the constructed scoring calculation system, the static label users screened after labeling are scored for static portraits, user shortcomings are found, and grouping and classification are performed; after completing the training and application of the static user portrait model, it is necessary to score the static label users screened after labeling in order to better understand user characteristics, find user shortcomings, and group and classify. The specific steps are as follows:
[0283] Rating calculation: Each static tag user is scored according to the constructed rating calculation system. The rating system usually includes character feature scoring and numerical feature scoring.
[0284] Character feature scoring: Calculate the proportion of each feature category in the target label 1 and the non-target label 0, and score according to the degree of excess or deficiency.
[0285] Numerical feature scoring: Determine the correlation between each feature and the target label, perform positive or negative correlation mapping and normalization, and obtain the final score.
[0286] Find user weaknesses: Identify user weaknesses in various feature dimensions through scoring results. For example, a user's low score in the "package usage" dimension may indicate that the user is not satisfied with the existing package.
[0287] Grouping and classification: Users are grouped and classified according to the scoring results. For example, users with higher scores are classified as high-value users, and users with lower scores are classified as potential churn users.
[0288] S62. According to the constructed scoring calculation system, the dynamic label users screened after labeling are scored for dynamic portraits to find the timing and strategy for intervention; after completing the training and application of the dynamic user portrait model, it is necessary to score the dynamic label users screened after labeling in order to observe the recent changes of users and find the timing and strategy for intervention. The specific steps are as follows:
[0289] Rating calculation: Each dynamic tag user is scored according to the constructed rating calculation system. The rating system usually includes character feature scoring and numerical feature scoring.
[0290] Character feature scoring: Calculate the proportion of each feature category in the target label 1 and the non-target label 0, and score according to the degree of excess or deficiency.
[0291] Numerical feature scoring: Determine the correlation between each feature and the target label, perform positive or negative correlation mapping and normalization, and obtain the final score.
[0292] Find the right time to intervene: Use the scoring results to identify recent changes in user behavior. For example, if a user's score on the "data usage" dimension increases significantly, it may indicate that the user's demand for data traffic has increased recently.
[0293] Formulate intervention strategies: formulate corresponding intervention strategies according to the dynamic changes of users. For example, push traffic package discount information to users with increased traffic demand.
[0294] S63, combining dynamic and static situations to complete personalized user portrait application; after completing the static and dynamic user portrait scoring, the two need to be combined to form a comprehensive personalized user portrait application. The specific steps are as follows:
[0295] Comprehensive score: The static score and the dynamic score are combined to form a comprehensive score for the user. For example, a weighted average method can be used to weight the static score and the dynamic score according to a certain weight to obtain a comprehensive score.
[0296] Improve user portrait: Based on the comprehensive score, further improve the user portrait. For example, the user's feature description can be further refined based on the comprehensive score.
[0297] Personalized recommendations and services: Based on a comprehensive user profile, personalized recommendations and services are provided. For example, high-end product recommendations are pushed to high-value users, and preferential activities are provided to potential churn users.
[0298] Combination Figure 2 From the perspective of building a personalized traffic user portrait application system based on big data processing and multi-dimensional features,
[0299] The system comprises a data collection and processing module, a multi-dimensional feature construction module, a dynamic and static user portrait construction module and a portrait application module connected to each other;
[0300] The data collection and processing module is connected to the multi-dimensional feature construction module and the dynamic and static user portrait construction module respectively, and the multi-dimensional feature construction module is connected to the portrait application module through the dynamic and static user portrait construction module;
[0301] Data collection and processing module: complete the collection and preliminary processing of big data, and complete the data processing work before user portrait modeling;
[0302] Multi-dimensional feature construction module: Combines business and modeling requirements to complete feature judgment and processing, and multi-dimensional hierarchical processing;
[0303] Dynamic and static user portrait construction module: With the support of the data collection and processing module, complete the construction of dynamic and static user portrait models and output important feature factors and target labels;
[0304] Portrait application module: Filter and circle users with target tags, build a user rating calculation system based on dynamic and static user portrait models, and complete personalized user portrait application.
[0305] The personalized traffic user portrait application construction system based on big data processing and multi-dimensional features is a complex and multi-level system that aims to build and apply user portraits through comprehensive data collection and processing to achieve precision marketing and personalized services. The following is a detailed summary of the system and an introduction to actual cases:
[0306] System Configuration
[0307] Data acquisition and processing module
[0308] This module is responsible for collecting user data from various data sources, including static data (such as basic user attributes) and dynamic data (such as user behavior data). After data collection is completed, data cleaning and preprocessing are required to ensure data accuracy and consistency.
[0309] Multi-dimensional feature building blocks
[0310] In this module, user data is feature extracted and processed in combination with business needs and modeling requirements. Features can be divided into multiple dimensions, such as user's social attributes, consumption habits, behavioral preferences, etc. By layering and grading these features, different levels of users can be better understood.
[0311] Dynamic and static user portrait construction module
[0312] This module builds a user portrait model with the support of the data collection and processing module. Static user portraits mainly describe the basic attributes of users, such as age, gender, occupation, etc. Dynamic user portraits focus on changes in user behavior, such as recent purchases, browsing history, etc. By combining this information, a comprehensive user portrait is generated, and important feature factors and target labels are output.
[0313] Portrait application module
[0314] In this module, based on the user portrait model, the target user group is screened and identified. Combined with dynamic and static user portraits, a user rating calculation system is built to carry out personalized recommendations, precision marketing and other applications.
[0315] Actual Cases
[0316] Case: Personalized recommendation system for e-commerce platforms
[0317] background
[0318] An e-commerce platform hopes to improve user experience and sales through a user portrait system. The platform has a large amount of user data, including user registration information, browsing history, purchase history, evaluation, etc.
[0319] Implementation steps
[0320] Data collection and processing
[0321] -Collect user registration information (static data) and user browsing and purchase records (dynamic data).
[0322] -Perform data cleaning to remove duplicate and invalid data and ensure data accuracy.
[0323] Multi-dimensional feature construction
[0324] - Extract basic attributes of users (such as age, gender, geographic location, etc.).
[0325] -Extract user behavior characteristics (such as browsing frequency, purchase frequency, preferred product categories, etc.).
[0326] - Layer features, such as dividing users into high-frequency purchasing users, low-frequency browsing users, etc.
[0327] Dynamic and static user portrait construction
[0328] -Build static user portraits to describe the basic attributes of users.
[0329] -Build dynamic user portraits to describe changes in user behavior.
[0330] - Combine static and dynamic information to generate comprehensive user portraits and output target labels, such as "high-value users", "potential churn users", etc.
[0331] Portrait Application
[0332] -Based on user portraits, filter out high-value users and recommend high-profit products to them.
[0333] -Push personalized discount information to potential lost users to win them back.
[0334] - Make personalized recommendations through the user rating system to improve user purchase conversion rate.
[0335] Through the user portrait system, e-commerce platforms can understand user needs more accurately, make personalized recommendations and precision marketing, and significantly improve user experience and sales. For example, the purchase conversion rate of high-value users increased by 20%, and the recovery rate of potential lost users increased by 15%. Based on big data processing and multi-dimensional features, the personalized traffic user portrait application construction system can comprehensively and accurately describe users through data collection, feature construction, portrait modeling and application, support personalized recommendations and precision marketing, and improve business results. The e-commerce platform in the actual case has achieved a dual improvement in user experience and sales through this system, verifying the effectiveness and practicality of the system.
[0336] In general, the method and system constructed by the present invention mainly address the deficiencies of the prior art:
[0337] Traditional personalized user portrait construction methods have slow processing speeds when faced with massive amounts of data, cannot accurately understand user needs, and have difficulty providing personalized services. The present invention solves these problems through technical means based on big data processing and multi-dimensional features, and has the following advantages:
[0338] Efficient data processing: Using ETL (extraction, transformation, loading) technology, massive data can be efficiently processed and managed to ensure the accuracy and completeness of the data.
[0339] Accurately describe user needs: Through the integration of multi-dimensional feature data, it is possible to describe user needs and behavioral characteristics more comprehensively and accurately.
[0340] Provide personalized services: Based on dynamic and static user portraits and scoring systems, we can better explore user needs and provide personalized recommendations and services.
[0341] Optimize user experience: By dynamically adjusting user portraits, timely respond to changes in user needs and improve user service perception.
[0342] Compared with existing technologies, the main advantages are:
[0343] Efficiency and stability in processing massive amounts of big data:
[0344] The use of big data processing technology can quickly process and analyze massive user data to ensure the efficiency and stability of the system.
[0345] Introduction of multi-dimensional feature data:
[0346] By introducing multi-dimensional feature data such as behavior, packages, content, location, terminal, and basic user information, combined with dynamic and static user portrait construction methods, users can be described more comprehensively and accurately than traditional methods.
[0347] Dynamic and static user portrait scoring system:
[0348] By building a dynamic and static user portrait scoring system, we can quantify user characteristics, identify user needs, and provide personalized user portrait services and applications.
[0349] Complete system architecture design:
[0350] The system architecture includes modules such as data collection and processing, multi-dimensional feature construction, dynamic and static user portrait construction, and portrait application, which can efficiently and accurately build personalized user portraits and apply them in actual marketing to enhance user service perception.
[0351] In addition, possible application scenarios in the future:
[0352] Applications in traffic and similar scenarios:
[0353] In various traffic scenarios, we can achieve data-driven development, capture users' growing service demands, more accurately explore user needs, and help companies develop more refined marketing strategies.
[0354] Optimize product design:
[0355] Help companies build more accurate user portraits, thereby optimizing the design of various products and increasing the user base.
[0356] Customer Relationship Management:
[0357] Provide enterprises with in-depth customer insights, optimize customer relationships and improve customer service perception and satisfaction.
[0358] In the current era of information explosion, accurately grasping user needs and behaviors is the key for enterprises to gain competitive advantages. This patent provides an efficient method and system that can help enterprises build personalized user portraits, better understand and tap user needs, achieve more refined marketing, and improve user satisfaction and purchase conversion rates. In summary, the present invention significantly improves data processing efficiency and the accuracy of user portraits by introducing big data processing technology and multi-dimensional feature data, combined with dynamic and static user portraits and scoring systems, and provides strong support for personalized services and marketing.
[0359] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, characterized by: The steps include: S1. Based on business needs, data from different sources are imported into the data lake for unified storage and management. Then, data is collected from the data lake based on Hive, and ETL-related data processing is completed to form a large wide table. S2. Based on business and modeling requirements, determine and process the features in the large wide table in step S1, and use a three-level hierarchical architecture to classify the dimensions, i.e., feature objects, primary classification, and secondary classification, to obtain a data wide table; S3. For the data wide table obtained in step S2, Spark-based distributed LightGBM model training is used to find important feature factors and target labels: S4. According to the important characteristic factors and target labels obtained in step S3, target labels and judgment probabilities are obtained for users based on static portrait models and dynamic portrait models, and re-screened and labeled through thresholds to further refine the user group; S5. Combined with the labels, a scoring calculation system is constructed based on the acquired dynamic and static important characteristic factors and the multi-dimensional classification; S6. According to the scoring calculation system determined in step S5, the dynamic and static portrait scores are performed on the users marked with the target tags, and the scores are combined to complete the personalized user portrait application; The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, in step S1, the operation of importing data from different sources into the data lake based on business needs, wherein the data from different sources are related platform data and text data; In step S1, data is collected from the data lake based on Hive, where the data collected involves traffic statistics, subscription relationships, billing, terminals, locations, and channel services; The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, the processing in step S2 includes basic processing, classification processing, dimensionality reduction and merging processing, and content conversion processing; wherein the basic processing involves type conversion, unit format conversion consistency, and function statistics; wherein the classification processing involves labeling a large amount of data in combination with business classification; wherein the dimensionality reduction and merging processing involves dimensionality reduction processing of high-dimensional features to meet the standards for the availability of subsequent models; wherein the content conversion processing involves the conversion between numerical content and character content, and the selection of a more suitable content representation form to express; The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, step S3 also includes constructing a static user portrait model and a dynamic user portrait model, wherein the static user portrait model is used to describe user characteristics; Among them, a dynamic user portrait model is built to observe recent changes in users; In the method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, the method for constructing a static user portrait model in step S3 is as follows: Take this month's data from the data wide table as a static user portrait modeling dataset and use it for subsequent Spark-based distributed LightGBM model training; The method of constructing a dynamic user portrait model in step S3 is as follows: The changes between the data of this month and last month in the data wide table are taken as the dynamic user portrait modeling data set and used for the subsequent distributed LightGBM model training based on Spark; The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, in step S4, the operations of obtaining target labels and judging probabilities for users based on static portrait models and dynamic portrait models are as follows: output labels and probabilities for users of the entire network based on the optimal static user model, then further filter static label users according to a set threshold, and then output labels and probabilities for users of the entire network based on the optimal dynamic user model; The method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, in step S5, constructs a character feature score calculation rule for the character features of the dynamic and static important feature factors, calculates the proportion of each category in each feature in the target label and the non-target label, and calculates the score using 3 times the standard deviation; In step S5, for the numerical features of the important dynamic and static feature factors, a numerical feature scoring calculation rule is constructed to determine the relevance of each feature with the target label, and the data is mapped using maximum and minimum normalization to calculate the score; In the method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features, the operation of scoring the dynamic and static portraits of the target-tagged user in step S6 is as follows: Perform static portrait scoring on the statically labeled users after labeling to find out the user's shortcomings and classify them into groups; Conduct dynamic portrait scoring on the dynamic tag users who have been screened after labeling to find the timing and strategy for intervention.
2. A system obtained by the method for constructing a personalized traffic user portrait application based on big data processing and multi-dimensional features as claimed in claim 1, characterized in that: The system comprises a data collection and processing module, a multi-dimensional feature construction module, a dynamic and static user portrait construction module and a portrait application module connected to each other; The data collection and processing module is connected to the multi-dimensional feature construction module and the dynamic and static user portrait construction module respectively, and the multi-dimensional feature construction module is connected to the portrait application module through the dynamic and static user portrait construction module; Data collection and processing module: complete the collection and preliminary processing of big data, and complete the data processing work before user portrait modeling; Multi-dimensional feature construction module: Combines business and modeling requirements to complete feature judgment and processing, and multi-dimensional hierarchical processing; Dynamic and static user portrait construction module: With the support of the data collection and processing module, complete the construction of dynamic and static user portrait models and output important feature factors and target labels; Portrait application module: Filter and circle users with target tags, combine dynamic and static user portrait models to build a user rating calculation system, and complete personalized user portrait application.
Citation Information
Patent Citations
Accurate marketing method and system for establishing user portrait based on big data
CN109978630A
User portrait generation method based on characteristic variable scoring, device, vehicle, and storage medium
WO2024067387A1