User behavior analysis method based on multi-source information

By collecting, cleaning and fusing multi-source data and combining deep learning and clustering algorithms, the limitations of traditional user behavior analysis methods are overcome, multi-dimensional description and real-time analysis of user behavior are achieved, and the accuracy of analysis and predictive ability are improved.

CN120670968APending Publication Date: 2025-09-19BEIJING QICHUANG TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510747148.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional user behavior analysis methods are based on a single data source and cannot fully reflect the complexity and diversity of user behavior. They also lack effective multi-source information integration and mining methods, making it difficult to discover the deep connections and patterns behind the data.

Method used

A multi-step approach is adopted: by collecting multi-source data, cleaning and preprocessing, fusing multi-source data, using deep learning and clustering algorithms for feature extraction and pattern mining, building a predictive model, and updating it in real time to adapt to changes in user behavior.

Benefits of technology

It realizes the multi-dimensional description of user behavior, overcomes the limitations of a single data source, improves the accuracy of analysis and predictive capabilities, can process newly generated user behavior data in real time, promptly discover changing trends, and dynamically adapt to market and user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670968A_ABST
    Figure CN120670968A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of user behavior analysis methods, in particular to a user behavior analysis method based on multi-source information, which comprises the following steps: step 1, collecting multi-source data related to user behaviors, including user basic information data, equipment data, behavior log data, geographic position data and social data, respectively storing the collected data in different data sources; 2, cleaning and preprocessing the collected multi-source data, removing noise data and abnormal values, and processing missing values by adopting a filling method; according to the method, the user behavior can be comprehensively described from multiple dimensions by fusing the multi-source information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user behavior analysis methods, and in particular to a user behavior analysis method based on multi-source information. Background Art

[0002] In today's digital age, user behavior on various platforms generates massive amounts of data, rich with information about user characteristics and behavioral patterns. Traditional user behavior analysis methods often rely on a single data source, such as website logs or app operation records, to describe user behavior through simple statistical analysis (such as number of clicks and dwell time). For example, a user's interest in a product on an e-commerce website can be judged solely based on the time they browse and the number of clicks they make. However, this approach has significant limitations.

[0003] A single data source cannot fully reflect the complexity and diversity of user behavior. User behavior is affected by many factors, including but not limited to the user's own attributes (age, gender, occupation, etc.), the characteristics of the device used (screen size, operating system, network environment, etc.), the spatiotemporal environment (geographic location, time, weather, etc.), and social relationships. At the same time, traditional methods lack effective integration and mining methods when processing high-dimensional and heterogeneous data, making it difficult to discover the deep connections and patterns hidden behind the data. With the development of technologies such as big data and artificial intelligence, there is an urgent need for a method that can integrate multi-source information and deeply analyze user behavior to meet the needs of enterprises in various aspects such as precision marketing, personalized recommendations, and user experience optimization.

[0004] However, in the existing technology, there is a problem of difficulty in data integration. Multi-source information has different data structures and characteristics. How to efficiently and accurately integrate data from different data sources to avoid information redundancy and conflict is the primary problem to be solved. Summary of the Invention

[0005] In order to solve the technical problem that data integration is difficult, the present invention provides a user behavior analysis method based on multi-source information.

[0006] The technical solution adopted by the present invention is: a user behavior analysis method based on multi-source information, comprising the following steps:

[0007] Step 1: Collect multi-source data related to user behavior, including basic user information data, device data, behavior log data, geographic location data, and social data, and store the collected data in different data sources;

[0008] Step 2: Clean and preprocess the collected multi-source data to remove noise data and outliers. For missing values, fill them in using the filling method.

[0009] Step 3: Fuse the cleaned and preprocessed multi-source data to build a unified user behavior dataset;

[0010] Step 4: Use convolutional neural networks and recurrent neural networks in deep learning to extract features from the fused user behavior data;

[0011] Step 5: Use a method based on mutual information and principal component analysis to perform feature selection and dimensionality reduction;

[0012] Step 6: Use the improved DBSCAN clustering algorithm to cluster the user behavior features after dimensionality reduction and mine user behavior patterns;

[0013] Step 7: Use support vector machine to classify the clustered behavior patterns and build a prediction model;

[0014] Step 8: Use the accuracy, recall, and F1 value indicators to evaluate the constructed behavior analysis model;

[0015] Step 9: Real-time update and dynamic analysis;

[0016] Step 10: Presentation and application of results.

[0017] The further setting is that the noise data processing method is as follows:

[0018] For numerical data, a statistical method is used, adopting the 3σ principle. If a data point x satisfies |x-μ|>3σ, then the data point is considered an outlier, where μ is the mean of the data and σ is the standard deviation of the data.

[0019] The calculation formula is as follows:

[0020]

[0021] Where n is the number of data points, x i represents the i-th data point;

[0022] The standard deviation formula is as follows:

[0023]

[0024] The missing value filling method is as follows:

[0025] For categorical data, the mode is used for filling; for numerical data, the mean or median is used for filling.

[0026] A further setting is to use a feature vector concatenation-based method to integrate data related to the same user from different data sources into one vector;

[0027] Assume that the multi-source data feature vector of user u is V u :

[0028] V u =[v u1 ,v u2 ,…,v un ];

[0029] in,

[0030] v u i represents the i-th feature value of user u;

[0031] n is the total number of features.

[0032] Further settings are as follows: The specific method of using convolutional neural networks and recurrent neural networks in deep learning to extract features from the fused user behavior data is as follows:

[0033] The convolutional neural network feature extraction method is as follows:

[0034] The user behavior data matrix X is used as the input of the convolutional neural network feature. The local features and abstract features of the data are extracted through the convolution layer, pooling layer and fully connected layer. In the convolution layer, the convolution kernel K is used to perform a convolution operation on the input data. The formula for the convolution operation is:

[0035]

[0036] Among them, y ij Represents the value of the output feature map at position (i, j) after convolution, x i+m,j+n Indicates the value of the input data at position (i+m,j+n), k m n represents the value of the convolution kernel K at the (m,n) position, M and N are the number of rows and columns of the convolution kernel respectively, and b is the bias term;

[0037] The pooling layer uses the maximum pooling method to downsample the feature map output by the convolution layer. The formula is y ij :

[0038] y ij =max m,n∈R x i×s+m,j×s+n ;

[0039] Among them, R is the pooling area, s is the pooling step size;

[0040] The recurrent neural network feature extraction method is as follows:

[0041] For a hidden unit in a recurrent neural network, the update formula is h t :

[0042] h t =σ(Wxh x t +W hh h t-1 +b h );

[0043] Among them, h t represents the hidden state at time t, x t represents the input at time t, W xh and W hh is the weight matrix, b h is the bias vector, σ is the activation function;

[0044] The convolutional neural network features and the features extracted by the recurrent neural network are combined to obtain the comprehensive feature vector F of the user behavior. u .

[0045] Further setting is that in step 5, the feature selection method based on mutual information is as follows: calculate the mutual information between each feature and the user behavior target variable, and the calculation formula of mutual information is MI(X;Y):

[0046]

[0047] Where X and Y represent two random variables, p(x,y) represents the joint probability distribution of X and Y, and p(x) and p(y) represent the marginal probability distributions of X and Y respectively.

[0048] Select features whose mutual information is greater than the threshold τ to form the candidate feature set C;

[0049] The dimensionality reduction method based on PCA is as follows:

[0050] Perform PCA dimensionality reduction on the candidate feature set C to map high-dimensional features to low-dimensional space;

[0051] Let the original feature matrix be X C , its covariance matrix is

[0052] in, For X C The mean matrix of the covariance matrix Σ is decomposed into eigenvalues ​​to obtain the eigenvalue λ i and the corresponding eigenvector e i , select the eigenvectors corresponding to the first k largest eigenvalues ​​to form the projection matrix P = [e1, e2, ..., e k ], the feature matrix after dimensionality reduction Y = X C P.

[0053] Further setting is that in step 6, a distance measurement method based on cosine similarity is used, and the cosine similarity calculation formula of two feature vectors x and y is:

[0054]

[0055] In the improved DBSCAN algorithm, the core object is defined as follows: if the number of data points contained in the neighborhood with a radius of ∈ centered on a data point p is not less than the minimum number of samples MinPts, then p is a core object. By continuously expanding the neighborhood of the core object, a cluster is formed, and users with similar behavior patterns are divided into the same cluster.

[0056] Further, in step 7, the support vector machine is used to classify the behavior patterns obtained by clustering, and the specific method of constructing the prediction model is as follows:

[0057] For the linearly separable case, the goal of SVM is to find a hyperplane w T x+b=0, which maximizes the distance between the two types of data points and the hyperplane;

[0058] The interval is calculated as follows:

[0059]

[0060] By solving the optimization problem min w , The constraint condition is y i (w T x i +b)≥1(i=1,2,…,n), where x i is the training sample, y i is the category label of the sample, and the optimal hyperplane parameters w and b are obtained;

[0061] For the case of nonlinear separability, the radial basis function K(x i ,x j );

[0062] K(x i ,x j )=exp(-γ||x i -x j || 2 );

[0063] Map the data to a high-dimensional space, perform linear classification in the high-dimensional space, and use the trained SVM model to classify and predict new user behavior data to determine which behavior pattern the user belongs to.

[0064] The further setting is, in step 8,

[0065] Accuracy

[0066] Recall

[0067]

[0068] Among them, TP represents true positive examples, TN represents true negative examples, FP represents false positive examples, and FN represents false negative examples.

[0069] Further settings are as follows: in step 9, a real-time data processing mechanism is established. When new user behavior data is generated, the data is collected, cleaned, and fused for preprocessing, and the new data is input into the trained model for real-time analysis. At the same time, the model is retrained regularly.

[0070] It is further configured that, in step 10, the results of the user behavior analysis are presented to the user in the form of intuitive charts and reports.

[0071] The beneficial effects of the present invention are: compared with the existing technology, the present invention can comprehensively describe user behavior from multiple dimensions by fusing multi-source information, overcome the limitations of a single data source, and more accurately portray user portraits. It can efficiently and accurately integrate data from different data sources to avoid information redundancy and conflict. By using advanced feature extraction and pattern mining algorithms, it can deeply explore the potential laws and patterns of user behavior, improve the accuracy and predictive ability of behavior analysis, and provide a more reliable basis for corporate decision-making. Through technical means such as dimensionality reduction, it can effectively reduce data dimensions, reduce the amount of calculation, improve analysis efficiency, and realize rapid processing and analysis of massive user behavior data. It can process newly generated user behavior data in real time, promptly discover changing trends in user behavior, and enable analysis results to dynamically adapt to changes in market and user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0073] In the description of the present invention, it should be noted that the terms "front", "up", "down", "left", "right", "vertical", "horizontal", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limiting the present invention.

[0074] refer to Figure 1 In order to solve the problems existing in the background technology, this application proposes the following technical solution: a user behavior analysis method based on multi-source information, comprising the following steps:

[0075] Step 1: Collect multi-source data related to user behavior, including basic user information data (such as age A, gender G, occupation O), device data (such as screen size S, operating system OS, network type NT), behavior log data (such as click timestamp T click , browse page P), geographic location data (latitude and longitude L l at、L l ng) and social data (number of friends F n um, social interaction frequency F int ). The collected data is stored in different data sources, such as relational databases, NoSQL databases, log files, etc.

[0076] The above technical solution is explained as follows: By extensively collecting multi-source data such as basic user information, devices, behavior logs, geographic location, and social media, it comprehensively covers all factors that influence user behavior. It can characterize user characteristics from multiple dimensions, no longer limited to a single perspective. For example, combining basic information such as age, gender, and occupation, as well as device usage and geographic location, can more accurately understand user needs and preferences, providing a rich and comprehensive data foundation for subsequent in-depth analysis. It helps to discover potential connections between data of different dimensions, unearth more valuable user behavior patterns, and overcome the one-sidedness of traditional single data source analysis.

[0077] Step 2: Clean and preprocess the collected multi-source data to remove noise data and outliers. For missing values, fill them in using the filling method.

[0078] In step 2, the noise data processing method is as follows:

[0079] For numerical data, such as age and screen size, a statistical method is used, adopting the 3σ principle (assuming the data follows a normal distribution). If a data point x satisfies |x-μ|>3σ, then the data point is considered an outlier, where μ is the mean of the data and σ is the standard deviation of the data.

[0080] The calculation formula is as follows:

[0081]

[0082] Where n is the number of data points, x i represents the i-th data point;

[0083] The standard deviation formula is as follows:

[0084]

[0085] The missing value filling method is as follows:

[0086] For categorical data, such as gender, occupation, etc., mode filling is used; for numerical data, mean or median filling is used. Taking mean filling as an example, if a numerical feature X has missing values, the mean of the feature is used. Fill in, where m is the number of non-missing values ​​of the feature, x j is the jth non-missing data point.

[0087] The above technical solution is explained as follows: Using statistical methods to remove noise and outliers effectively purifies the data environment, preventing erroneous data from interfering with analytical results. A rational filling strategy for missing values ​​ensures data integrity. Using the 3σ principle for numerical data as an example, outliers are accurately identified and removed, ensuring that the data more closely reflects the true distribution. Mode filling for categorical data and mean-median filling for numerical data maintain the statistical properties of the data. These operations improve data quality, making subsequent analysis and mining based on this data more reliable and accurate, and enhancing the stability and generalization capabilities of the model.

[0088] Step 3: Fuse the cleaned and preprocessed multi-source data to build a unified user behavior dataset;

[0089] In step 3, a feature vector concatenation method is used to integrate data related to the same user from different data sources into one vector;

[0090] Assume that the multi-source data feature vector of user u is V u :

[0091] V u =[v u1 ,v u2 ,…,v un ];

[0092] in,

[0093] v u i represents the i-th feature value of user u;

[0094] n is the total number of features;

[0095] For example,

[0096] The technical solution is explained as follows: A multi-source data fusion method based on feature vector concatenation integrates user-related data scattered across different data sources into a unified vector. This process breaks down data silos, allowing different types of data to complement each other and form a complete user behavior dataset. This comprehensive dataset allows businesses to gain a comprehensive user picture, facilitating holistic analysis and insights. For example, in e-commerce scenarios, integrating data such as page views, click timestamps, and basic user information can more accurately determine user purchase intent, providing strong support for personalized recommendations and marketing campaigns, and enhancing user experience and commercial value.

[0097] Step 4: Use convolutional neural networks and recurrent neural networks in deep learning to extract features from the fused user behavior data;

[0098] In step 4, the specific method for extracting features from the fused user behavior data using convolutional neural networks and recurrent neural networks in deep learning is as follows:

[0099] The convolutional neural network feature extraction method is as follows:

[0100] The user behavior data matrix X is used as the input of the convolutional neural network feature. The local features and abstract features of the data are extracted through the convolution layer, pooling layer and fully connected layer. In the convolution layer, the convolution kernel K is used to perform a convolution operation on the input data. The formula for the convolution operation is:

[0101]

[0102] Among them, y ij Represents the value of the output feature map at position (i, j) after convolution, x i+m,j+n Indicates the value of the input data at position (i+m,j+n), k m n represents the value of the convolution kernel K at the (m,n) position, M and N are the number of rows and columns of the convolution kernel respectively, and b is the bias term;

[0103] The pooling layer uses the maximum pooling method to downsample the feature map output by the convolution layer. The formula is y ij :

[0104] y ij =max m,n∈R x i×s+m,j×s+n ;

[0105] Among them, R is the pooling area, s is the pooling step size;

[0106] The recurrent neural network feature extraction method is as follows: Considering that user behavior data has time series characteristics, a recurrent neural network is used to process the data to capture the time dependency in the data.

[0107] For a hidden unit in a recurrent neural network, the update formula is h t :

[0108] h t =σ(W xh x t +W hh h t-1 +b h );

[0109] Among them, h t represents the hidden state at time t, x t represents the input at time t, W xh and W hh is the weight matrix, b h is the bias vector, σ is the activation function (such as Sigmoid function or RelU function);

[0110] The convolutional neural network features and the features extracted by the recurrent neural network are combined to obtain the comprehensive feature vector F of the user behavior. u .

[0111] The above technical solution is explained as follows: Feature extraction leverages the strengths of convolutional neural networks and recurrent neural networks. CNNs excel at capturing local features in data, extracting key patterns and structures in user behavior data through convolution and pooling operations. RNNs effectively process time series characteristics, exploring the temporal evolution and dependencies of user behavior. The resulting combined feature vector comprehensively reflects user behavior characteristics, providing more representative features for subsequent behavioral pattern analysis. This improves the model's ability to understand and analyze complex user behaviors, making the analysis more aligned with actual behavior patterns.

[0112] Step 5: Use a method based on mutual information and principal component analysis to perform feature selection and dimensionality reduction;

[0113] In step 5, the feature selection method based on mutual information is as follows: calculate the mutual information between each feature and the user behavior target variable (such as purchase behavior, user retention, etc.). The calculation formula of mutual information is MI(X; Y):

[0114]

[0115] Where X and Y represent two random variables (feature and target variable), p(x,y) represents the joint probability distribution of X and Y, and p(x) and p(y) represent the marginal probability distributions of X and Y, respectively.

[0116] Select features whose mutual information is greater than the threshold τ to form the candidate feature set C;

[0117] The dimensionality reduction method based on PCA is as follows:

[0118] Perform PCA dimensionality reduction on the candidate feature set C to map high-dimensional features to low-dimensional space;

[0119] Let the original feature matrix be X C , its covariance matrix is

[0120] in, For X C The mean matrix of the covariance matrix Σ is decomposed into eigenvalues ​​to obtain the eigenvalue λ i and the corresponding eigenvector e i , select the eigenvectors corresponding to the first k largest eigenvalues ​​to form the projection matrix P = [e1, e2, ..., e k ], the feature matrix after dimensionality reduction Y = X C P.

[0121] The above technical solution is explained as follows: Based on mutual information and principal component analysis, mutual information is first used to screen out features with strong correlations with the target variable of user behavior to form a candidate feature set, removing a large number of irrelevant or redundant features to reduce data interference. PCA is then used to map the high-dimensional candidate feature set to a low-dimensional space, reducing the data dimension while retaining the key information. This not only improves computational efficiency and reduces storage and computing resource consumption, but also makes model training and analysis more efficient, and clarifies the relationships between features, thereby improving model performance and analytical accuracy.

[0122] Step 6:

[0123] Use the improved DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm to cluster the user behavior features after dimensionality reduction and explore user behavior patterns;

[0124] In step 6, a distance measurement method based on cosine similarity is used. The cosine similarity calculation formula of two feature vectors x and y is:

[0125]

[0126] In the improved DBSCAN algorithm, the core object is defined as follows: if the number of data points contained in the neighborhood with a radius of ∈ centered on a data point p is not less than the minimum number of samples MinPts, then p is a core object. By continuously expanding the neighborhood of the core object, a cluster is formed, and users with similar behavior patterns are divided into the same cluster.

[0127] The above technical solution is explained as follows: The improved DBSCAN clustering algorithm combined with the cosine similarity metric effectively adapts to high-dimensional user behavior feature data. Using cosine similarity to measure the similarity between data points can better reflect the directional relationship between feature vectors, avoiding the limitations of traditional Euclidean distance in high-dimensional space. By defining core objects and expanding the neighborhood to form clusters, users with similar behavior patterns can be accurately grouped together. This allows companies to clearly distinguish the behavior patterns of different user groups, such as the differences between active users and potential churners, providing a basis for developing targeted operational strategies and achieving refined management and precision marketing.

[0128] Step 7: Use support vector machine to classify the clustered behavior patterns and build a prediction model;

[0129] In step 7, the support vector machine is used to classify the behavior patterns obtained by clustering and the specific method of building a prediction model is as follows:

[0130] For the linearly separable case, the goal of SVM is to find a hyperplane w T x+b=0, which maximizes the distance between the two types of data points and the hyperplane;

[0131] The interval is calculated as follows:

[0132]

[0133] By solving the optimization problem min w , The constraint condition is y i (w T x i +b)≥1(i=1,2,…,n), where x i is the training sample, y i is the category label of the sample (value is +1 or -1), and the optimal hyperplane parameters w and b are obtained;

[0134] For the case of nonlinear separability, the radial basis function K(x i ,x j );

[0135] K(x i ,x j )=exp(-γ||x i -x j || 2 );

[0136] Map the data to a high-dimensional space, perform linear classification in the high-dimensional space, and use the trained SVM model to classify and predict new user behavior data to determine which behavior pattern the user belongs to.

[0137] The technical solution described above is explained as follows: Support vector machines are used to classify and predict behavioral patterns. For linearly separable data, accurate classification is achieved by finding the optimal hyperplane that maximizes the separation between the two data types. For nonlinearly separable data, linear classification is achieved by mapping the data to a high-dimensional space using a kernel function. The trained model can quickly classify and predict new user behavior data, determining the corresponding behavioral pattern. In practical applications, this can help companies predict user purchasing decisions, product usage preferences, and other behaviors in advance, providing precise guidance for product recommendations, advertising placement, and other services, improving marketing effectiveness and user satisfaction, and enhancing the company's market competitiveness.

[0138] Step 8: Use the accuracy, recall, and F1 value indicators to evaluate the constructed behavior analysis model;

[0139] In step 8,

[0140] Accuracy

[0141] Recall

[0142]

[0143] TP represents a true positive, TN represents a true negative, FP represents a false positive, and FN represents a false negative. Based on the evaluation results, the model parameters are adjusted and optimized, such as the convolutional neural network features, the recurrent neural network structure, and the SVM kernel function parameters, to improve model performance.

[0144] The above technical solution is explained as follows: Using metrics such as accuracy, recall, and F1-score to evaluate models comprehensively measures model performance from different perspectives. Accuracy reflects the model's overall ability to correctly classify data, recall focuses on its ability to identify positive examples, and F1-score provides a comprehensive balance between the two. Adjusting model parameters based on the evaluation results, such as optimizing the neural network structure and SVM kernel function parameters, can improve model performance in a targeted manner. Continuous optimization allows the model to better adapt to evolving user behavior data, maintain the accuracy and effectiveness of analysis, provide businesses with more reliable decision-making, and ensure the model maximizes its value in real-world applications.

[0145] Step 9: Real-time update and dynamic analysis;

[0146] In step 9, a real-time data processing mechanism is established. When new user behavior data is generated, the data is collected, cleaned, and pre-processed for integration. The new data is then fed into the trained model for real-time analysis. Simultaneously, the model is regularly retrained, and model parameters are updated based on new data to adapt to changing user behavior patterns. Using a sliding window approach, the most recent data is selected as the training set to ensure that the model can capture dynamic changes in user behavior in a timely manner.

[0147] The technical solution is explained as follows: A real-time data processing mechanism is established to promptly collect, clean, and integrate newly generated user behavior data, feeding it into model analysis in real time. Simultaneously, the model is regularly retrained, updating parameters based on new data and using a sliding window to select recent data for training. This enables the model to track changes in user behavior in real time, capturing new trends and patterns. This allows businesses to adjust marketing strategies and optimize product features in real time, quickly responding to changes in the market and user needs, and consistently staying in sync with user behavior, improving user experience and enhancing the timeliness and flexibility of their business operations.

[0148] Step 10: Results presentation and application;

[0149] In step 10, the results of the user behavior analysis are presented to users in the form of intuitive charts (such as line charts, bar charts, and heat maps) and reports. Based on these analysis results, companies can conduct targeted marketing campaigns, such as delivering personalized advertising and product recommendations to users with different behavior patterns; optimizing user experience, improving product design and functionality; and providing user churn warnings and proactively taking measures to retain potential users, thereby achieving data-driven decision-making and business growth.

[0150] The technical solution is explained as follows: User behavior analysis results are presented in intuitive charts and reports, making them easy for businesses to understand and interpret. Based on these results, businesses can deliver personalized advertising and product recommendations to users with different behavioral patterns, improving marketing conversion rates. They can also optimize product design and functionality based on user behavior feedback to enhance the user experience. Furthermore, by analyzing signs of user churn, they can provide early warnings and implement retention measures. This data-driven decision-making approach enables more scientific and precise business operations, achieving business growth, enhancing user stickiness and loyalty, and promoting sustainable development.

[0151] To sum up, in the present invention, by fusing multi-source information, it is possible to comprehensively describe user behavior from multiple dimensions, overcome the limitations of a single data source, and more accurately portray user portraits. By utilizing advanced feature extraction and pattern mining algorithms, it is possible to deeply explore the potential laws and patterns of user behavior, improve the accuracy and predictive ability of behavior analysis, and provide a more reliable basis for corporate decision-making. Through technical means such as dimensionality reduction, it is possible to effectively reduce data dimensions, reduce the amount of calculation, improve analysis efficiency, and achieve rapid processing and analysis of massive user behavior data. It is possible to process newly generated user behavior data in real time, promptly discover changing trends in user behavior, and enable analysis results to dynamically adapt to changes in market and user needs.

[0152] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A user behavior analysis method based on multi-source information, characterized by: The following steps are involved: Step 1: Collect multi-source data related to user behavior, including basic user information data, device data, behavior log data, geographic location data, and social data, and store the collected multi-source data in different data sources; Step 2: Clean and preprocess the collected multi-source data to remove noise data and outliers. For missing values, fill them in using the filling method. Step 3: Fuse the cleaned and preprocessed multi-source data to build a unified user behavior dataset; Step 4: Use convolutional neural networks and recurrent neural networks in deep learning to extract features from the fused user behavior data; Step 5: Use a method based on mutual information and principal component analysis to perform feature selection and dimensionality reduction; Step 6: Use the improved DBSCAN clustering algorithm to cluster the user behavior features after dimensionality reduction and mine user behavior patterns; Step 7: Use support vector machines to classify the user behavior patterns obtained through clustering and build a behavior analysis model; Step 8: Use the precision, recall, and F1 value indicators to evaluate the constructed behavior analysis model; Step 9: Real-time update and dynamic analysis; Step 10: Presentation and application of results.

2. The user behavior analysis method based on multi-source information according to claim 1, characterized in that: In step 2, The noise data processing method is as follows: For numerical data, a statistical method is used, adopting the 3σ principle. If a data point x satisfies |x-μ|>3σ, then the data point is considered an outlier, where μ is the mean of the data and σ is the standard deviation of the data. The calculation formula is as follows: Where n is the number of data points, x i represents the i-th data point; The standard deviation formula is as follows: The missing value filling method is as follows: For categorical data, the mode is used for filling; for numerical data, the mean or median is used for filling.

3. The user behavior analysis method based on multi-source information according to claim 2, characterized in that: In step 3, a feature vector concatenation method is used to integrate data related to the same user from different data sources into one vector; Assume that the multi-source data feature vector of user u is V u , V u =[v u1 ,v u2 ,…,v un ]; Among them, v u i represents the i-th feature value of user u; n is the total number of features.

4. The user behavior analysis method based on multi-source information according to claim 3 is characterized in that: In step 4, the specific method for extracting features from the fused user behavior data using convolutional neural networks and recurrent neural networks in deep learning is as follows: The convolutional neural network feature extraction method is as follows: The user behavior data matrix X is used as the input of the convolutional neural network feature. The local features and abstract features of the data are extracted through the convolution layer, pooling layer and fully connected layer. In the convolution layer, the convolution kernel K is used to perform a convolution operation on the input data. The formula for the convolution operation is: Among them, y ij Represents the value of the output feature map at position (i, j) after convolution, x i+m,j+n Indicates the value of the input data at position (i+m,j+n), k m n represents the value of the convolution kernel K at the (m,n) position, M and N are the number of rows and columns of the convolution kernel respectively, and b is the bias term; The pooling layer uses the maximum pooling method to downsample the feature map output by the convolution layer. The formula is y ij : and ij =max m,n∈R x i×s+m,j×s+n ; Among them, R is the pooling area, s is the pooling step size; The recurrent neural network feature extraction method is as follows: For a hidden unit in a recurrent neural network, the update formula is h t : h t =σ(W xh x t +W hh h t-1 +b h ); Among them, h t represents the hidden state at time t, x t represents the input at time t, W xh and W hh is the weight matrix, b h is the bias vector, σ is the activation function; The convolutional neural network features and the features extracted by the recurrent neural network are combined to obtain the comprehensive feature vector F of the user behavior. u .

5. The user behavior analysis method based on multi-source information according to claim 4 is characterized in that: In step 5, The feature selection method based on mutual information is as follows: calculate the mutual information between each feature and the user behavior target variable. The calculation formula of mutual information is MI(X; Y): Where X and Y represent two random variables, p(x,y) represents the joint probability distribution of X and Y, and p(x) and p(y) represent the marginal probability distributions of X and Y respectively. Select features whose mutual information is greater than the threshold τ to form the candidate feature set C; The dimensionality reduction method based on PCA is as follows: Perform PCA dimensionality reduction on the candidate feature set C to map high-dimensional features to low-dimensional space; Let the original feature matrix be X C , its covariance matrix is in, For X C The mean matrix of the covariance matrix Σ is decomposed into eigenvalues ​​to obtain the eigenvalue λ i and the corresponding eigenvector e i , select the eigenvectors corresponding to the first k largest eigenvalues ​​to form the projection matrix P = [e1, e2, ..., e k ], the feature matrix after dimensionality reduction Y = X C P.

6. The user behavior analysis method based on multi-source information according to claim 5, characterized in that: In step 6, Using the distance measurement method based on cosine similarity, the cosine similarity calculation formula of two feature vectors x and y is: In the improved DBSCAN algorithm, the core object is defined as follows: if the number of data points contained in the neighborhood with a radius of ∈ centered on a data point p is not less than the minimum number of samples MinPts, then p is a core object. By continuously expanding the neighborhood of the core object, a cluster is formed, and users with similar behavior patterns are divided into the same cluster.

7. The user behavior analysis method based on multi-source information according to claim 6, characterized in that: In step 7, the support vector machine is used to classify the behavior patterns obtained by clustering and the specific method of building a prediction model is as follows: For the linearly separable case, the goal of SVM is to find a hyperplane w T x+b=0, which maximizes the distance between the two types of data points and the hyperplane; The interval is calculated as follows: By solving the optimization problem The constraint condition is y i (w T x i +b)≥1(i=1,2,…,n), where x i is the training sample, y i is the category label of the sample, and the optimal hyperplane parameters w and b are obtained; For the case of nonlinear separability, the radial basis function K(x i ,x j ); K(x i ,x j )=exp(-γ||x i -x j || 2 ); Map the data to a high-dimensional space, perform linear classification in the high-dimensional space, and use the trained SVM model to classify and predict new user behavior data to determine which behavior pattern the user belongs to.

8. The user behavior analysis method based on multi-source information according to claim 7, characterized in that: In step 8, Accuracy Recall Among them, TP represents true positive examples, TN represents true negative examples, FP represents false positive examples, and FN represents false negative examples.

9. The user behavior analysis method based on multi-source information according to claim 8, characterized in that: In step 9, a real-time data processing mechanism is established. When new user behavior data is generated, the data is collected, cleaned, and fused for pre-processing. The new data is then input into the trained model for real-time analysis. At the same time, the model is retrained regularly.

10. The user behavior analysis method based on multi-source information according to claim 9, characterized in that: In step 10, the results of the user behavior analysis are presented to the user in the form of charts and reports.