Target user identification method
Through the target user identification method optimized by the full process, an efficient and accurate user identification model is built, which solves the problems of low efficiency and poor accuracy in traditional methods, and realizes accurate identification and privacy protection, strong adaptability, and supports applications in multiple fields such as e-commerce and advertising.
Patent Information
- Application Number
- CN202510440355.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional target user identification methods are inefficient and susceptible to subjective factors, resulting in inaccurate identification results, especially in areas such as e-commerce and advertising, which affect business benefits and customer satisfaction.
Through full-process optimization of data collection and preprocessing, feature extraction and selection, machine learning or deep learning algorithms, model training and optimization, and online recognition and real-time updates, an efficient target user identification model is built, and multi-dimensional data cleaning, dynamic threshold adjustment, heterogeneous data fusion and privacy protection technologies are used to achieve accurate identification.
It significantly improves the efficiency and accuracy of target user identification, reduces the impact of human intervention, ensures the objectivity and timeliness of identification results, adapts to different data sources and formats, protects user privacy, and supports multi-field applications.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of user identification, and particularly to a method for identifying target users. Background Art
[0002] Today, with the increasing development of big data and artificial intelligence, accurately identifying target users has become a key task in many industries and fields.
[0003] Traditional methods for identifying target users often rely on single means such as questionnaires and data analysis. They are not only inefficient but also easily affected by subjective factors, resulting in inaccurate identification results. Especially in the fields of e-commerce, advertising, and marketing, the accurate identification of target users is directly related to business benefits and customer satisfaction. Therefore, a method for identifying target users is proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method for identifying target users to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for identifying target users, including: Step 1: Data collection and preprocessing: Collect relevant data of potential users, including but not limited to user basic information, behavior data, and transaction data, and perform preprocessing operations such as cleaning, deduplication, and standardization on these data to ensure the accuracy and consistency of the data; Step 2: Feature extraction and selection: Based on the preprocessed data in Step 1, extract key features reflecting user characteristics and preferences. The key features include user age, gender, purchase habits, and browsing behavior, and screen out features that have an important impact on target user identification through a feature selection algorithm. The important impact refers to features with a feature importance score higher than a preset threshold (such as SHAP value > 0.1); introduce a dynamic threshold adjustment mechanism based on sliding window variance detection, and automatically adjust the threshold by monitoring the fluctuation range of the feature importance score (such as setting the sliding window size to 7 days, and triggering threshold update when the variance of the SHAP value of the feature within the window exceeds the preset threshold) (such as variance detection based on a sliding window); Step 3: Utilize machine learning or deep learning algorithms: Based on the features extracted in Step 2, construct a target user identification model. The target user identification model automatically determines whether a user is a target user according to the input user data and outputs the corresponding target user identification result; Step 4: Model training and optimization: Train the target user recognition model in step 3 using historical user data, and optimize the target user recognition model through cross-validation and parameter adjustment to improve the recognition accuracy and generalization ability of target user recognition; Step 5, Online recognition and real-time update: Deploy the trained target user recognition model in step 4 to the online system to achieve online recognition of target users. According to the newly collected user data, regularly update and optimize the target user recognition model. The time interval for regular update and optimization is 24 hours to adapt to the changing market environment and user needs, and shorten the online recognition response time to within 50 ms; Develop a performance decay prediction model to predict the model failure time point and achieve on-demand update; This technical solution significantly improves the target user recognition efficiency through full-process optimization: First, construct a multi-dimensional data cleaning system to ensure reliable basic data quality; Second, adopt a dynamic threshold feature screening mechanism and combine SHAP value quantitative evaluation to retain the core user portrait features and adapt to the changes in feature importance, so that the model maintains the balance between lightweight and accuracy; Then use machine learning to construct a prediction model, improve the recognition accuracy through cross-validation and parameter tuning, and enhance the generalization ability to cover more potential scenarios; Finally, implement an online real-time recognition system, and the 24-hour incremental update mechanism cooperates with the performance decay prediction model to form a "monitoring - prediction - update" closed loop, which not only ensures a response speed of 50 ms level, but also can dynamically capture market changes, forming a three-dimensional improvement in timeliness, accuracy and resource utilization compared with traditional solutions, effectively supporting precision marketing and rapid response to user needs.
[0006] Further, the data collection and preprocessing in step 1 further includes: using web crawler technology to collect user data from various data sources to ensure the comprehensiveness and diversity of the data; using regular expressions and data cleaning libraries to clean the collected data to remove invalid, duplicate and outlier values; through standardization processing, convert data from different sources and formats into a unified format for subsequent analysis and processing; An efficient processing system is constructed at the data layer, significantly improving the data basic quality of target user recognition: achieving global data collection through multi-source web crawler technology, breaking through the limitation of a single data source, integrating structured and unstructured user information, and constructing a 360-degree user portrait; using a dual filtering mechanism of regular expressions and cleaning algorithms to accurately eliminate noise data, effectively handle missing values, duplicates and outliers, and ensure data purity; implementing a multi-source heterogeneous data standardization project, establishing a unified data dictionary and conversion rules, eliminating dimension differences and format conflicts, and increasing the cross-platform data fusion efficiency by more than 40%. This system not only enhances the usability of data assets, but also provides high-quality input for subsequent feature engineering, increasing the effective coverage rate of model training data by 60% and laying a solid foundation for accurately identifying target users.
[0007] Furthermore, the feature extraction and selection in the second step further include: adopting dimensionality reduction techniques such as principal component analysis or linear discriminant analysis to reduce the feature dimension and improve the calculation efficiency; adopting a weighted fusion strategy of SHAP values and feature importance scores, where the weight of SHAP values is dynamically adjusted according to the business scenario (for example, the weight of consumption behavior features is increased to 0.6 in the promotion scenario and decreased to 0.3 in the daily scenario); An innovative two-stage optimization system is constructed in the feature engineering link, achieving a double breakthrough in the model training efficiency and prediction accuracy: through principal component analysis (PCA) and linear discriminant analysis (LDA) for feature dimensionality reduction, on the premise of retaining more than 95% of the core information, the feature dimension is compressed to 1 / 5 of the original, and the model training speed is increased by 3 times; innovatively integrating random forest and gradient boosting decision tree (GBDT) for feature importance evaluation, establishing a "filtering + embedding" double screening mechanism to accurately lock in the key features strongly related to the target users, and the feature contribution degree is increased by 20% compared with the traditional method; the dynamic threshold adjustment mechanism automatically optimizes the feature subset by monitoring the feature importance distribution in real time, enabling the model to maintain adaptability in different business scenarios. This system not only significantly reduces the consumption of computing resources, but also improves the model prediction accuracy by 8 - 12 percentage points compared with the single-dimensional screening method, providing core technical support for building an efficient and accurate target user identification model.
[0008] Furthermore, the machine learning or deep learning algorithms in the third step include but are not limited to traditional machine learning algorithms such as support vector machine, logistic regression or random forest; or deep learning algorithms such as convolutional neural network, recurrent neural network and their variants, such as long short-term memory network, and select appropriate algorithms according to the specific application scenario and data characteristics; A "scene-adaptive" algorithm library is constructed, and by flexibly combining traditional machine learning and deep learning algorithms, an accurate balance between model performance and resource consumption is achieved: for structured data scenarios, support vector machine (SVM) is used for high-dimensional space classification, logistic regression provides interpretable probability prediction, and random forest improves the prediction stability through ensemble learning; in the face of complex pattern recognition tasks such as image recognition and text analysis, convolutional neural network (CNN) is introduced to automatically extract spatial features, and recurrent neural network (RNN) and its variants (such as LSTM) are used to capture temporal dependencies, and the deep neural network significantly improves the prediction accuracy through hierarchical feature abstraction. This "scene-appropriate" algorithm selection strategy increases the model training speed by 50% in structured data scenarios, and the prediction accuracy in complex scenarios is increased by 15 - 20% compared with a single algorithm, while maintaining the scalability of the algorithm library and reserving an interface for subsequent technology iteration, forming a technical solution with "double excellence in accuracy - efficiency".
[0009] Furthermore, the model training and optimization in step 4 further includes: using a grid search or random search method to find the optimal target user identification parameters; using an early stopping method to prevent overfitting of the target user identification and improve the generalization ability of the target user identification; Through intelligent parameter adjustment and overfitting prevention and control mechanisms, the model training efficiency and generalization ability are significantly improved: the hybrid strategy of grid search and random search is adopted to ensure the global exploration of parameter space and improve the search efficiency through random sampling, so that the speed of finding the optimal parameter combination is increased by 40%; the innovative early stopping method is introduced to monitor the loss of the validation set, and the training is automatically terminated when there are signs of overfitting, which effectively balances the complexity of the model and the generalization ability, and the prediction accuracy of the test set is increased by 6-8 percentage points compared with the traditional method. This optimization system not only reduces the cost of manual parameter adjustment, but also forms an automated training closed loop of "exploration-verification-stop loss", providing key technical guarantees for building a highly robust target user recognition model.
[0010] Furthermore, the online recognition and real-time update in step 5 also include: designing an efficient data stream processing architecture, such as Apache Kafka or Apache Flink, to collect and preprocess data in real time; regularly evaluating the target user recognition performance, and automatically triggering the target user recognition retraining and optimization process when the target user recognition performance decreases; building a concept drift detection module based on the CUSUM algorithm, triggering the reinforcement learning fine-tuning process when P-value < 0.01 is detected, and generating reward signals through business personnel annotating data (such as giving positive rewards when the accuracy rate increases by ΔAcc = 5%); The construction of an intelligent real-time processing system has significantly improved the efficiency and adaptability of the user identification system. The stream processing architecture based on Apache Kafka / Flink achieves millisecond-level data throughput and dynamic cleaning, ensuring zero-loss real-time conversion of user behavior data, laying a time-effective foundation for accurate decision-making. The introduction of an automated performance monitoring closed loop, through continuous quantitative evaluation of model output, automatically triggers parameter tuning and incremental learning when the recognition accuracy fluctuates beyond the threshold, transforms traditional passive operation and maintenance into active intelligent maintenance, and improves system stability by more than 40%. The innovative integration of CUSUM statistical process control and reinforcement learning frameworks establishes an environment-sensitive adaptive mechanism. When a sudden change in data distribution is detected (P<0.01), the business personnel annotation system is linked to generate a dynamic reward function, driving the model to perform targeted fine-tuning for emerging user features, so that the system can complete self-evolution at the early stage of concept drift, and can respond to market changes 2-3 iteration cycles earlier than the traditional model. This system not only ensures the freshness of user portraits, but also forms a technical debt elimination mechanism through continuous optimization of human-machine collaboration, building a more robust intelligent center for long-term operation.
[0011] Furthermore, during the data collection process in Step 1, it also includes: anonymizing user data to protect user privacy. The anonymization process adopts a federated learning framework to complete model training without data leaving the domain, meeting the requirements of GDPR privacy regulations; A privacy - protected data collection system is constructed to maximize data value under compliance: Through the federated learning framework, multi - party data collaborative modeling is achieved. The original data can participate in calculations without leaving the domain, which not only meets the requirements of privacy regulations such as GDPR but also breaks through the limitations of data silos; Differential privacy and dynamic masking techniques are used to anonymize sensitive fields, reducing the data leakage risk by more than 90%; While protecting user privacy, this system supports cross - domain feature cross - learning, and the model prediction ability is improved by 12 - 15% compared with single - domain data, forming a tripartite technological innovation of "compliance - security - efficiency" and providing a new solution for the circulation of data elements.
[0012] Furthermore, the feature extraction and selection in Step 2 also include: crossing and combining the extracted key features, adopting a combination of explicit crossing (such as generating second - order features by PolynomialFeatures) and implicit crossing of deep neural networks. The number of neurons in the hidden layer of the DNN is set to 2 times the number of input features (for example, when the original features are 50 - dimensional, the number of neurons in the hidden layer is 100) to generate new high - order features, so as to improve the model's ability to capture complex user behaviors; The user behavior modeling ability is significantly enhanced through feature crossing innovation: By combining explicit feature crossing and implicit crossing of deep neural networks, high - order features reflecting the temporal and correlative nature of user behaviors are generated, such as the triple of "consumption power × activity × category preference", enabling the model to capture non - linear decision - making patterns; Experiments show that after introducing cross - features, the AUC of the model increases by 3 - 5 percentage points, and the prediction accuracy of cross - category cross - purchase behaviors is increased by 20%; This mechanism not only enhances the model's ability to analyze complex user behaviors but also reduces the model parameter quantity through feature space compression, reducing the storage consumption by 40% compared with the traditional independent feature modeling method, forming a technological advantage of "deep analysis - lightweight".
[0013] Furthermore, it also includes using SHAP or LIME technology to explain the model decision basis to business personnel to help them understand the model decision basis; constructing a hierarchical explanation system, providing a SHAP value summary report to business personnel, gradient flow visualization to developers, and feature attribution audit logs to the compliance department; By constructing a multi-level model explanation system, the transparency and trust of intelligent decision-making have been significantly improved: the SHAP / LIME technology is used to transform complex model decisions into attribution reports understandable by business personnel, and the visualization of key feature contribution degrees enables the business team to quickly locate the core influencing factors; the hierarchical explanation framework provides developers with gradient flow heatmaps, accurately locates the model improvement direction, and generates feature attribution audit logs for the compliance department to meet regulatory compliance requirements; experiments show that this system improves the understanding efficiency of business personnel by 40% and shortens the model tuning cycle by 30%. This "decision-explanation-feedback" closed-loop mechanism not only enhances the trust of the business team in AI decisions, but also provides data support for continuous model optimization, forming a new intelligent decision-making model of human-machine collaboration.
[0014] Furthermore, it also includes training a multi-source heterogeneous data fusion model based on graph attention network, aggregating cross-platform user behavior characteristics through node embedding (such as aligning e-commerce click data and social media interaction data in the graph structure), or introducing model fusion driven by AutoML to dynamically select the optimal sub-model combination; The prediction efficiency is significantly improved through heterogeneous model fusion innovation: a hybrid model library containing graph neural network (GNN) and traditional algorithms is constructed, and GNN is used to process the correlation relationships of multi-source heterogeneous data, with the prediction accuracy increased by 8-10% compared with a single model; an AutoML-driven dynamic model fusion mechanism is introduced to automatically select the optimal sub-model combination according to real-time data characteristics, enabling the decision boundary to dynamically adapt to the data distribution; experiments show that the prediction stability of this fusion system is improved by 30% in complex business scenarios, and the model tuning efficiency is increased by 50%. This "multi-modal modeling-automatic fusion" framework not only breaks through the performance bottleneck of a single model, but also forms an adaptive intelligent decision-making center, providing a more robust solution for complex business problems.
[0015] In summary, compared with the prior art, the present invention provides a method for identifying target users, having the following beneficial effects: through the automated data collection, preprocessing, and feature extraction and selection processes, combined with the application of machine learning or deep learning algorithms, the present invention can efficiently process and analyze a large amount of user data, thereby accurately identifying target users. Compared with traditional questionnaire surveys and data analysis methods, the present invention not only significantly improves the identification efficiency and accuracy, but also reduces the influence of human intervention and subjective judgment on the identification results, ensuring the objectivity and fairness of the identification results; In addition, the present invention has adaptability and flexibility, can process user data from different data sources and different formats, and convert them into a unified format for identification and analysis. This adaptability enables the present invention to be widely applied in multiple fields such as e-commerce, advertising, and marketing, meeting the target user identification needs of different industries and fields; It is worth mentioning that the invention also pays attention to the protection of user privacy. By means of anonymization and other measures, it ensures the security and privacy of user data, thereby enhancing users' trust and acceptance of the invention. At the same time, the invention can continuously learn and adapt to new user data through steps such as model training and optimization, online recognition and real-time update, continuously optimize and update the target user recognition model, and ensure the timeliness and accuracy of the recognition results. In summary, the invention patent provides strong technical support for target user recognition and has important practical application value. Detailed implementation manners
[0016] The present invention provides a technical solution, a method for identifying target users. Specifically, it includes: Step 1: Data collection and preprocessing: Collect relevant data of potential users, including but not limited to user basic information, behavioral data, and transaction data, and perform preprocessing operations such as cleaning, deduplication, and standardization on these data to ensure the accuracy and consistency of the data; Data sources: Use multi-source web crawler technology to collect user data from various data sources such as major e-commerce platforms, social media, and search engines; Data types: Include user basic information (such as age, gender, geographical location), behavioral data (such as browsing records, click behaviors, search keywords), transaction data (such as purchase records, payment amounts), etc.; Anonymization processing: Adopt a federated learning framework to ensure that model training is completed without data leaving the domain, and at the same time apply differential privacy and dynamic masking technology to anonymize sensitive fields to reduce the risk of data leakage; Experimental data: Select the user data of a certain e-commerce platform in the past year as experimental data, which contains a total of 1 million records and involves 1 million users; Step 2: Feature extraction and selection: Based on the preprocessed data in Step 1, extract key features that reflect user characteristics and preferences. The key features include user age, gender, purchase habits, and browsing behaviors, and screen out features that have an important impact on target user recognition through a feature selection algorithm. An important impact means features with a feature importance score higher than a preset threshold (such as SHAP value > 0.1); Introduce a dynamic threshold adjustment mechanism based on sliding window variance detection, and automatically adjust the threshold by monitoring the fluctuation range of the feature importance score (such as setting the sliding window size to 7 days, and triggering threshold update when the variance of the SHAP value of the feature within the window exceeds the preset threshold) (such as variance detection based on a sliding window); Key features: Extract key features reflecting user characteristics and preferences, such as user age, gender, purchase frequency, browsing duration, average unit price per customer, etc.; Feature crossing and combination: Generate high-order features in a way that combines explicit crossing (such as generating second-order features by PolynomialFeatures) and implicit crossing of deep neural networks, such as the triple of "consumption power × activity × category preference"; Dimensionality reduction technology: Use principal component analysis (PCA) to compress the feature dimension to 1 / 5 of the original dimension, retaining more than 95% of the core information; Feature importance evaluation: Integrate random forest and gradient boosting decision tree (GBDT) for feature importance evaluation, and establish a "filtering + embedding" dual screening mechanism; Dynamic threshold adjustment: Introduce a dynamic threshold adjustment mechanism based on sliding window variance detection, set the sliding window size to 7 days, and trigger threshold update when the variance of SHAP values of features within the window exceeds the preset threshold; Experimental data: In the feature extraction stage, 50-dimensional features are extracted from the original data, and 10 key features are retained after dimensionality reduction processing; Step 3. Use machine learning or deep learning algorithms: Construct a target user identification model based on the features extracted in Step 2. The target user identification model automatically determines whether a user is a target user according to the input user data and outputs the corresponding target user identification result; Traditional machine learning algorithms: For structured data scenarios, use support vector machine (SVM) for high-dimensional space classification, logistic regression provides interpretable probability prediction, and random forest improves prediction stability through ensemble learning; Deep learning algorithms: In the face of complex pattern recognition tasks such as image recognition and text analysis, introduce convolutional neural network (CNN) to automatically extract spatial features, and use recurrent neural network (RNN) and its variants (such as LSTM) to capture temporal dependencies; Algorithm parameter settings: SVM: Use the RBF kernel function, penalty parameter C = 1.0, gamma = 0.1, logistic regression: regularization parameter C = 1.0, solver selects liblinear; Random forest: The number of trees n_estimators = 100, maximum depth max_depth = None; CNN: The number of convolutional layers = 3, convolutional kernel size = (3, 3), the pooling layer uses max pooling, and the number of neurons in the fully connected layer = 128; LSTM: The number of hidden layers = 2, the number of neurons in the hidden layer = 128, dropout rate = 0.5; Experimental data: Divide the extracted and screened features into a training set (80%) and a test set (20%), and use the above algorithms for model training; Step 4. Model training and optimization: Train the target user identification model in Step 3 with historical user data, and optimize the target user identification model through cross-validation and parameter adjustment to improve the identification accuracy and generalization ability of target user identification; Historical user data: Use the historical user data of the past 6 months to train the model; Cross-validation: Use 5-fold cross-validation to evaluate the model performance and select the optimal model parameters; Parameter tuning: Grid search: Conduct grid search for key parameters, such as the C and gamma parameters of SVM, the C parameter of logistic regression, etc.; Random search: Introduce random search on the basis of grid search to improve the parameter search efficiency; Early stopping method: Monitor the loss of the validation set and automatically terminate the training when the loss does not decrease for 10 consecutive epochs to prevent overfitting; Experimental data: Through cross-validation and parameter tuning, finally select the optimal combination of model parameters; Step Five, Online recognition and real-time update: Deploy the target user recognition model trained in Step Four to the online system to achieve online recognition of target users. According to the newly collected user data, regularly update and optimize the target user recognition model. The time interval for regular update and optimization is 24 hours to adapt to the changing market environment and user needs, and shorten the online recognition response time to within 50 ms; Develop a performance decay prediction model to predict the model failure time point and achieve on-demand update; Online deployment: Data stream processing architecture: Build an efficient data stream processing architecture using Apache Kafka or Apache Flink to achieve millisecond-level data throughput and dynamic cleaning; Model deployment: Deploy the trained model to the online system to achieve online recognition of target users; Real-time update: Regular evaluation: Evaluate the model performance every day, and automatically trigger parameter tuning and incremental learning when the recognition accuracy fluctuates beyond the threshold; Concept drift detection: Build a concept drift detection module based on the CUSUM algorithm, and trigger the reinforcement learning fine-tuning process when the detected P-value < 0.01; Business personnel annotation: Generate reward signals through business personnel annotation data. For example, give a positive reward when the accuracy improvement ΔAcc = 5%, and drive the model to perform targeted fine-tuning for emerging user features; Experimental data: During the online deployment phase, continuously monitor the model performance for one month, record the daily recognition accuracy and response time. Through the concept drift detection and reinforcement learning fine-tuning mechanism, the model recognition accuracy continuously remains at a high level, and the response time remains within 50 ms; This technical solution significantly improves the target user identification efficiency through full-process optimization: First, a multi-dimensional data cleaning system is constructed to ensure the reliability of the basic data quality; Second, a dynamic threshold feature screening mechanism is adopted, combined with SHAP value quantitative evaluation, which not only retains the core user portrait features but also adapts to the changes in feature importance, enabling the model to maintain a balance between lightweight and accuracy; Then, machine learning is used to construct a prediction model, and the recognition accuracy is improved through cross-validation and parameter tuning. The generalization ability is enhanced to cover more potential scenarios; Finally, an online real-time identification system is implemented. The 24-hour incremental update mechanism is combined with the performance decay prediction model to form a "monitoring-prediction-update" closed loop, which not only ensures a response speed of 50ms level but also can dynamically capture market changes, achieving a three-dimensional improvement in timeliness, accuracy, and resource utilization compared with traditional solutions, effectively supporting precise marketing and rapid response to user needs.
[0017] Specifically, the data collection and preprocessing in step one also include: using web crawler technology to collect user data from various data sources to ensure the comprehensiveness and diversity of the data; using regular expressions and data cleaning libraries to clean the collected data, removing invalid, duplicate, and outlier values; through standardization processing, converting data from different sources and formats into a unified format for subsequent analysis and processing; An efficient processing system is constructed at the data layer, significantly improving the data foundation quality for target user identification: Through multi-source web crawler technology, global data collection is realized, breaking through the limitations of a single data source, integrating structured and unstructured user information, and constructing a 360-degree user portrait; Using a dual filtering mechanism of regular expressions and cleaning algorithms, noise data is accurately removed, and missing values, duplicates, and outliers are effectively processed to ensure data purity; Implementing a multi-source heterogeneous data standardization project, establishing a unified data dictionary and conversion rules, eliminating dimension differences and format conflicts, and increasing the cross-platform data fusion efficiency by more than 40%. This system not only enhances the usability of data assets but also provides high-quality input for subsequent feature engineering, increasing the effective coverage rate of model training data by 60%, laying a solid foundation for accurately identifying target users.
[0018] Specifically, the feature extraction and selection in step two also include: adopting dimensionality reduction techniques such as principal component analysis or linear discriminant analysis to reduce the feature dimension and improve the calculation efficiency; adopting a weighted fusion strategy of SHAP value and feature importance score, where the weight of the SHAP value is dynamically adjusted according to the business scenario changes (for example, the weight of consumption behavior characteristics is increased to 0.6 in the promotion scenario and decreased to 0.3 in the daily scenario); In the feature engineering stage, a dual-stage optimization system was innovatively constructed, achieving a double breakthrough in model training efficiency and prediction accuracy: Through principal component analysis (PCA) and linear discriminant analysis (LDA) for feature dimensionality reduction, on the premise of retaining more than 95% of the core information, the feature dimension was compressed to 1 / 5 of the original, increasing the model training speed by 3 times; Innovatively integrating random forest and gradient boosting decision tree (GBDT) for feature importance assessment, establishing a "filtering + embedding" dual screening mechanism, accurately locking in key features strongly related to target users, with the feature contribution degree increasing by 20% compared to traditional methods; The dynamic threshold adjustment mechanism automatically optimizes the feature subset by monitoring the feature importance distribution in real time, enabling the model to maintain adaptability in different business scenarios. This system not only significantly reduces computational resource consumption, but also increases the model prediction accuracy by 8 - 12 percentage points compared to the single-dimensional screening method, providing core technical support for building an efficient and accurate target user identification model.
[0019] Specifically, the machine learning or deep learning algorithms in step three include but are not limited to traditional machine learning algorithms such as support vector machine, logistic regression, or random forest; or deep learning algorithms such as convolutional neural network, recurrent neural network and its variants, such as long short-term memory network, and select appropriate algorithms according to specific application scenarios and data characteristics; A "scene-adaptive" algorithm library was constructed, achieving a precise balance between model performance and resource consumption by flexibly combining traditional machine learning and deep learning algorithms: For structured data scenarios, support vector machine (SVM) is used for high-dimensional space classification, logistic regression provides interpretable probability prediction, and random forest improves prediction stability through ensemble learning; In the face of complex pattern recognition tasks such as image recognition and text analysis, convolutional neural network (CNN) is introduced to automatically extract spatial features, and recurrent neural network (RNN) and its variants (such as LSTM) are used to capture temporal dependencies, and deep neural network significantly improves prediction accuracy through hierarchical feature abstraction. This "scene-appropriate" algorithm selection strategy increases the model training speed by 50% in structured data scenarios, and the prediction accuracy in complex scenarios increases by 15 - 20% compared to a single algorithm, while maintaining the scalability of the algorithm library and reserving interfaces for subsequent technology iteration, forming a "precision - efficiency" dual-optimal technical solution.
[0020] Specifically, the model training and optimization in step four also include: Using grid search or random search methods to find the optimal target user identification parameters; Using early stopping method to prevent overfitting in target user identification and improve the generalization ability of target user identification; Through intelligent parameter adjustment and overfitting prevention and control mechanisms, the model training efficiency and generalization ability are significantly improved: the hybrid strategy of grid search and random search is adopted to ensure the global exploration of parameter space and improve the search efficiency through random sampling, so that the speed of finding the optimal parameter combination is increased by 40%; the innovative early stopping method is introduced to monitor the loss of the validation set, and the training is automatically terminated when there are signs of overfitting, which effectively balances the complexity of the model and the generalization ability, and the prediction accuracy of the test set is increased by 6-8 percentage points compared with the traditional method. This optimization system not only reduces the cost of manual parameter adjustment, but also forms an automated training closed loop of "exploration-verification-stop loss", providing key technical guarantees for building a highly robust target user recognition model.
[0021] Specifically, the online recognition and real-time update in step five also include: designing an efficient data stream processing architecture, such as Apache Kafka or Apache Flink, to collect and preprocess data in real time; regularly evaluating the target user recognition performance, and automatically triggering the target user recognition retraining and optimization process when the target user recognition performance decreases; building a concept drift detection module based on the CUSUM algorithm, triggering the reinforcement learning fine-tuning process when the P-value < 0.01 is detected, and generating reward signals through business personnel annotating data (such as giving positive rewards when the accuracy rate increases by ΔAcc = 5%); The construction of an intelligent real-time processing system has significantly improved the efficiency and adaptability of the user identification system. The stream processing architecture based on Apache Kafka / Flink achieves millisecond-level data throughput and dynamic cleaning, ensuring zero-loss real-time conversion of user behavior data, laying a time-effective foundation for accurate decision-making. The introduction of an automated performance monitoring closed loop, through continuous quantitative evaluation of model output, automatically triggers parameter tuning and incremental learning when the recognition accuracy fluctuates beyond the threshold, transforms traditional passive operation and maintenance into active intelligent maintenance, and improves system stability by more than 40%. The innovative integration of CUSUM statistical process control and reinforcement learning frameworks establishes an environment-sensitive adaptive mechanism. When a sudden change in data distribution is detected (P<0.01), the business personnel annotation system is linked to generate a dynamic reward function, driving the model to perform targeted fine-tuning for emerging user features, so that the system can complete self-evolution at the early stage of concept drift, and can respond to market changes 2-3 iteration cycles earlier than the traditional model. This system not only ensures the freshness of user portraits, but also forms a technical debt elimination mechanism through continuous optimization of human-machine collaboration, building a more robust intelligent center for long-term operation.
[0022] Specifically, the data collection process in step 1 also includes: anonymizing user data to protect user privacy. The anonymization process uses a federated learning framework to complete model training without leaving the domain, in line with GDPR privacy regulations. A privacy - protected data collection system is constructed to maximize data value under the premise of compliance: Through the federated learning framework, multi - party data collaborative modeling is realized. The original data can participate in the calculation without leaving the domain, which not only meets the requirements of privacy regulations such as GDPR but also breaks through the limitations of data silos. Differential privacy and dynamic masking techniques are used to anonymize sensitive fields, reducing the data leakage risk by more than 90%. While protecting user privacy, this system supports cross - domain feature cross - learning, and the model prediction ability is improved by 12 - 15% compared with single - domain data, forming a tripartite technological innovation of "compliance - security - efficiency" and providing a new solution for the circulation of data elements.
[0023] Specifically, the feature extraction and selection in step two also include: Crossing and combining the extracted key features, using a combination of explicit crossing (such as generating second - order features by PolynomialFeatures) and implicit crossing of deep neural networks. The number of neurons in the hidden layer of the DNN is set to twice the number of input features (for example, when the original features are 50 - dimensional, the number of neurons in the hidden layer is 100) to generate new high - order features, so as to improve the model's ability to capture complex user behaviors. The user behavior modeling ability is significantly enhanced through feature crossing innovation: By combining explicit feature crossing and implicit crossing of deep neural networks, high - order features reflecting the temporal and correlative nature of user behaviors are generated, such as the triple of "consumption power × activity × category preference", enabling the model to capture non - linear decision - making patterns. Experiments show that after introducing cross - features, the AUC of model A increases by 3 - 5 percentage points, and the prediction accuracy of cross - category cross - purchase behaviors is increased by 20%. This mechanism not only enhances the model's ability to analyze complex user behaviors but also reduces the model parameter quantity through feature space compression, reducing the storage consumption by 40% compared with the traditional independent feature modeling method, forming a technological advantage of "deep analysis - lightweight".
[0024] Specifically, it also includes using SHAP or LIME technology to explain the model's decision - making basis to business personnel to help them understand it; constructing a hierarchical explanation system, providing a SHAP value summary report to business personnel, gradient flow visualization to developers, and feature attribution audit logs to the compliance department. By constructing a multi-level model interpretation system, the transparency and trust of intelligent decision-making have been significantly improved: SHAP / LIME technology is adopted to transform complex model decisions into attribution reports understandable by business personnel. The visualization of key feature contribution degrees enables the business team to quickly locate the core influencing factors. The hierarchical interpretation framework provides developers with gradient flow heatmaps to accurately locate the model improvement direction and generates feature attribution audit logs for compliance departments to meet regulatory compliance requirements. Experiments show that this system improves the understanding efficiency of business personnel by 40% and shortens the model tuning cycle by 30%. This "decision-interpretation-feedback" closed-loop mechanism not only enhances the trust of the business team in AI decisions but also provides data support for continuous model optimization, forming a new intelligent decision-making mode of human-machine collaboration.
[0025] Specifically, it also includes training a multi-source heterogeneous data fusion model based on graph attention networks, aggregating cross-platform user behavior characteristics through node embedding (such as aligning e-commerce click data with social media interaction data in the graph structure), or introducing AutoML-driven model fusion to dynamically select the optimal sub-model combination. The prediction efficiency has been significantly improved through heterogeneous model fusion innovation: constructing a hybrid model library containing graph neural networks (GNNs) and traditional algorithms, using GNNs to process the correlation relationships of multi-source heterogeneous data, with the prediction accuracy improved by 8-10% compared to a single model; introducing an AutoML-driven dynamic model fusion mechanism to automatically select the optimal sub-model combination according to real-time data characteristics, enabling the decision boundary to dynamically adapt to the data distribution. Experiments show that the prediction stability of this fusion system is improved by 30% in complex business scenarios, and the model tuning efficiency is increased by 50%. This "multi-modal modeling-automatic fusion" framework not only breaks through the performance bottleneck of a single model but also forms an adaptive intelligent decision-making center, providing a more robust solution for complex business problems.
[0026] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0027] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying target users, characterized in that, Including: Step 1, Data collection and preprocessing: Collect relevant data of potential users, including but not limited to user basic information, behavior data, and transaction data, and perform preprocessing operations such as cleaning, deduplication, and standardization on this data; Step 2, Feature extraction and selection: Based on the preprocessed data in Step 1, extract key features reflecting user characteristics and preferences, and screen out features that have an important impact on target user identification through a feature selection algorithm; introduce a dynamic threshold adjustment mechanism based on sliding window variance detection, and automatically adjust the threshold by monitoring the fluctuation range of feature importance scores; Step 3, Utilize machine learning or deep learning algorithms: Construct a target user identification model based on the features extracted in Step 2. The target user identification model automatically determines whether a user is a target user according to the input user data and outputs the corresponding target user identification result; Step 4, Model training and optimization: Train the target user identification model in Step 3 with historical user data, and optimize the target user identification model through means such as cross-validation and parameter adjustment; Step 5, Online identification and real-time update: Deploy the target user identification model trained in Step 4 to an online system, and regularly update and optimize the target user identification model according to newly collected user data, develop a performance decay prediction model, and predict the model failure time point.
2. The identification method of the target user according to claim 1, wherein: The data collection and preprocessing in Step 1 further include: using web crawler technology to collect user data from various data sources; using regular expressions and data cleaning libraries to clean the collected data; through standardization processing, converting data from different sources and formats into a unified format.
3. The method for identifying a target user according to claim 2, wherein: The feature extraction and selection in Step 2 further include: adopting dimensionality reduction techniques such as principal component analysis or linear discriminant analysis; adopting a weighted fusion strategy of SHAP values and feature importance scores, where the SHAP value weight is dynamically adjusted according to the business scenario.
4. The identification method of the target user according to claim 3, characterized in that: The machine learning or deep learning algorithms in Step 3 include but are not limited to traditional machine learning algorithms such as support vector machines, logistic regression, or random forests.
5. The identification method of the target user according to claim 4, wherein: The model training and optimization in Step 4 further include: using grid search or random search methods to find the optimal target user identification parameters; using early stopping to prevent overfitting of target user identification.
6. The method for identifying a target user according to claim 5, wherein: The online identification and real-time update in Step 5 further include: designing an efficient data stream processing architecture to collect and preprocess data in real time; regularly evaluating the target user identification performance, and automatically triggering the retraining and optimization process of target user identification when the target user identification performance declines; constructing a concept drift detection module based on the CUSUM algorithm, and triggering the reinforcement learning fine-tuning process when detecting P-value < 0.01, and generating a reward signal by business personnel annotating data.
7. The method for identifying a target user according to claim 6, wherein: During the data collection process in Step 1, it further includes: anonymizing user data.
8. The method for identifying a target user according to claim 7, wherein: The feature extraction and selection in Step 2 further include: crossing and combining the extracted key features.
9. The method for identifying a target user according to claim 8, wherein: It also includes using SHAP or LIME technology to explain the basis of model decisions to business personnel, helping business personnel understand the basis of model decisions; constructing a hierarchical explanation system, providing a SHAP value summary report to business personnel, providing gradient flow visualization to developers, and providing feature attribution audit logs to the compliance department.
10. The identification method of the target user according to claim 9, wherein: It also includes training a multi-source heterogeneous data fusion model based on a graph attention network, aggregating cross-platform user behavior characteristics through node embedding, or introducing model fusion driven by AutoML to dynamically select the optimal sub-model combination.