User behavior analysis method and related equipment

By obtaining, cleaning and desensitizing the operational behavior data of cloud mobile phone users, a user-application heterogeneous graph network is built, and the graph neural network is used to predict the user's life cycle value, which solves the accuracy and privacy protection problems of cloud mobile phone user behavior analysis, and improves the service quality and personalized recommendations.

CN120337138APending Publication Date: 2025-07-18启朔(深圳)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510422287.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

How to gain an in-depth understanding of the behavior of cloud mobile phone users to improve service quality, the existing technology lacks effective analysis methods.

Method used

By obtaining the operational behavior data of cloud mobile phone users, cleaning and desensitizing, building a user-application heterogeneous graph network, and using graph neural network to predict the user's life cycle value.

Benefits of technology

It improves the accuracy and security of user behavior analysis, provides data support for personalized recommendation strategies, and ensures the security and compliance of user privacy information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337138A_ABST
    Figure CN120337138A_ABST
Patent Text Reader

Abstract

The invention discloses a user behavior analysis method and related equipment, and relates to the technical field of cloud mobile phone application, the method comprises the following steps: obtaining operation behavior data of a cloud mobile phone user, the operation behavior data comprising application use duration, application starting frequency and page browsing behavior; cleaning and desensitizing the operation behavior data based on a preset data processing rule to obtain standardized behavior data; based on the standardized behavior data, a graph neural network model is adopted to construct a user-application heterogeneous graph network, and the heterogeneous graph network comprises user nodes, application nodes and edge weights representing user and application interaction strength; and predicting the life cycle value of the user according to the user-application heterogeneous graph network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cloud mobile phone applications, and in particular, to a user behavior analysis method and related devices. Background Art

[0002] With the rapid development of the mobile Internet, cloud mobile phones, as a new type of mobile terminal solution, have gradually attracted market attention and user favor. By migrating computing and storage resources to the cloud, cloud mobile phones provide users with a more lightweight, efficient and convenient usage experience. However, in the operation of cloud mobile phones, how to deeply understand user behavior has become a key issue in improving the service quality of cloud mobile phones. Therefore, there is an urgent need for a user behavior analysis method to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further described in detail in the Detailed Implementation section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a user behavior analysis method, which includes:

[0005] Obtain the operation behavior data of cloud mobile phone users, where the operation behavior data includes application usage duration, application launch frequency, and page browsing behavior;

[0006] Clean and desensitize the operation behavior data based on preset data processing rules to obtain standardized behavior data;

[0007] Based on the standardized behavior data, construct a user-application heterogeneous graph network using a graph neural network model, where the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between users and applications;

[0008] Predict the user life cycle value according to the user-application heterogeneous graph network.

[0009] In some embodiments, obtaining the operation behavior data of cloud mobile phone users includes:

[0010] Capture the user interface operation event stream through a preset buried point interface, where the operation event stream includes the original touch coordinate sequence, gesture type, and operation timestamp;

[0011] Perform regional blurring processing on the original touch coordinate sequence, map it to the block coding set of the preset screen grid division, and generate a blurred coordinate sequence;

[0012] The gesture types are pattern normalized, and non-standard gestures are mapped to matching codes in a preset standard gesture coding library to generate a standardized gesture coding set.

[0013] In some implementations, the original touch coordinate sequence is subjected to regional fuzzification processing, and is mapped to a block code set divided by a preset screen grid to generate a fuzzy coordinate sequence, including:

[0014] Divide the user operation screen into an N×M grid block matrix, where N and M are preset positive integers, and each grid block corresponds to a unique code;

[0015] Based on the grid block matrix, a space filling curve algorithm is used to traverse each coordinate point in the original touch coordinate sequence, and each coordinate point is mapped to a block code of the grid block to which it belongs, wherein the space filling curve algorithm includes a Z-order curve or a Hilbert curve;

[0016] A fuzzy coordinate sequence is generated according to the mapping results of all coordinate points, wherein the fuzzy coordinate sequence is a sequence composed of block codes corresponding to each coordinate point.

[0017] In some embodiments, the operation behavior data is cleaned and desensitized based on preset data processing rules to obtain standardized behavior data, including:

[0018] Detect abnormal operation sequences in the operation behavior data. If the duration of a single operation exceeds a first preset threshold or the interval between adjacent operations is less than a second preset threshold, mark it as abnormal data and remove it to generate valid operation data;

[0019] Perform irreversible desensitization on the user identification information in the valid operation data, and use the salted hash algorithm combined with a preset number of iterations to generate a desensitized user identification;

[0020] Perform regular expression matching on URL data in page browsing behavior, retain access records that comply with preset domain name whitelist rules, and generate compliant access records;

[0021] Integrate effective operation data, desensitized user identification and compliance access records to generate standardized behavioral data.

[0022] In some implementations, based on standardized behavior data, a graph neural network model is used to construct a user-application heterogeneous graph network, including:

[0023] According to the application usage time and application startup frequency in the standardized behavior data, the function usage depth weight factor of the user and each application node is calculated, where the weight factor is obtained by normalizing the product of the average daily usage time plus 1, logarithmically transformed and the startup frequency;

[0024] Construct weighted edges between user nodes and application nodes based on the depth weight factor of function usage to generate an initial heterogeneous graph network;

[0025] Perform multi-order neighborhood aggregation on the initial heterogeneous graph network through a graph attention network to extract user node embedding vectors and application node embedding vectors, forming a user-application heterogeneous graph network.

[0026] In some embodiments, according to the user-application heterogeneous graph network, predict the user lifetime value, including:

[0027] Extract user node embedding vectors, application node embedding vectors, and edge weight features from the user-application heterogeneous graph network to generate a user behavior topology feature set;

[0028] Fuse the user behavior topology feature set with the user historical behavior time series data, where the time series data includes the application usage duration volatility and the payment behavior interval period within a continuous preset number of days;

[0029] Input the fused features into a time series enhanced prediction model to output the probability distribution of the user lifetime value within a future preset time window.

[0030] In some embodiments, it further includes:

[0031] Determine the user value level according to the probability distribution of the user lifetime value;

[0032] Construct a recommendation action space, where the action space includes a set of recommendation channels, recommendation content types, and recommendation trigger times, and the trigger times are dynamically adjusted according to the user active period distribution;

[0033] Adopt the proximal policy optimization algorithm to dynamically adjust the recommendation policy parameters according to the user real-time behavior feedback, and generate a personalized recommendation policy matching the user value level.

[0034] In a second aspect, the present application proposes a user behavior analysis device, which includes:

[0035] A data acquisition unit for acquiring the operation behavior data of cloud mobile phone users, where the operation behavior data includes application usage duration, application startup frequency, and page browsing behavior;

[0036] A data processing unit for cleaning and desensitizing the operation behavior data based on preset data processing rules to obtain standardized behavior data;

[0037] A model construction unit for constructing a user-application heterogeneous graph network based on the standardized behavior data by using a graph neural network model, where the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between users and applications;

[0038] A behavior prediction unit for predicting the user lifecycle value based on the user-application heterogeneous graph network.

[0039] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, the steps of the user behavior analysis method according to any one of the first aspects are implemented.

[0040] In a fourth aspect, the present application provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, the user behavior analysis method according to any one of the first aspects is implemented.

[0041] In summary, the present application obtains multi-dimensional operation behavior data and performs standardized cleaning and desensitization processing, significantly improving data quality and privacy security; constructs a user-application heterogeneous graph network based on a graph neural network, which can effectively capture complex topological relationships in user behavior, and combines time-series data fusion and multi-hop neighborhood aggregation technologies to achieve accurate prediction of user lifecycle value. This method not only enhances the depth and accuracy of user behavior analysis, but also provides reliable data support for the generation of subsequent personalized recommendation strategies. At the same time, through technologies such as block coding obfuscation and salted hash desensitization, it ensures the security and compliance of user privacy information throughout the collection and processing process. Description of the Drawings

[0042] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit this specification. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0043] Figure 1 It is a schematic flowchart of a user behavior analysis method provided by an embodiment of the present application;

[0044] Figure 2 It is a schematic structural diagram of a user behavior analysis device provided by an embodiment of the present application;

[0045] Figure 3 It is a structural diagram of an electronic device for user behavior analysis provided by an embodiment of the present application. Detailed Embodiments

[0046] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments.

[0047] Please refer to Figure 1 , which is a schematic flowchart of a user behavior analysis method provided by an embodiment of this application, and specifically may include:

[0048] S110. Obtain the operation behavior data of cloud mobile phone users. Among them, the operation behavior data includes the application usage duration, the application startup frequency, and the page browsing behavior;

[0049] Exemplarily, obtaining the operation behavior data of cloud mobile phone users is a basic step in user behavior analysis, aiming to construct a complete portrait of user behavior through multi-dimensional data collection. As a cloud virtualized terminal, the operation behavior data of cloud mobile phones needs to be extracted from the interaction between users and applications, covering core dimensions such as application usage duration, application startup frequency, and page browsing behavior. These data are captured in real time through lightweight data embedding technology, ensuring that the collection process is non-intrusive to the user experience, and at the same time, following the principle of privacy protection to perform preliminary obfuscation processing on sensitive information, providing a high-fidelity and low-risk original data source for subsequent analysis.

[0050] The application usage duration reflects the degree of functional dependence of users on specific applications. The application startup frequency reveals the usage habits and active cycles of users, and the page browsing behavior depicts the interest distribution and interaction paths of users within the application. The combination of the three can three-dimensionally describe user behavior characteristics from the three levels of time, frequency, and content, providing key inputs for constructing a user-application interaction model. Such data not only supports the mining of short-term behavior patterns, but also can form a baseline for user life cycle analysis through long-term accumulation, laying a data foundation for subsequent prediction and recommendation services.

[0051] S120. Clean and desensitize the operation behavior data based on preset data processing rules to obtain standardized behavior data;

[0052] Exemplarily, the data cleaning and desensitization process is based on a preset rule system to systematically manage the noise and sensitive information in the original operation behavior data. Through a multi-dimensional anomaly detection mechanism, invalid or abnormal behavior records (such as ultra-short operation time, high-frequency abnormal triggers, etc.) are identified and eliminated. At the same time, pattern matching technology is used to filter non-compliant data entries to ensure the validity and consistency of the input data, providing a high-quality data foundation for subsequent analysis.

[0053] In the desensitization process, irreversible encryption algorithms are used to transform sensitive fields such as user identifiers, and privacy perturbation techniques are combined to add controllable noise to group behavior characteristics, while eliminating individual identifiability and retaining the statistical characteristics of the data. This stage strictly follows the principle of minimizing data collection, and limits the data processing scope through a dynamic authorization mechanism to ensure that user privacy information is always within the safe boundary during the standardization process.

[0054] S130. Based on the standardized behavior data, a user-application heterogeneous graph network is constructed using a graph neural network model. Among them, the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between users and applications.

[0055] Exemplarily, based on the standardized user behavior data, a two-way interaction relationship network between users and applications is constructed through graph structure modeling technology. User nodes and application nodes are the core entities in the network, and edge weights dynamically represent the functional dependence intensity of users on specific applications through quantization indicators (such as usage frequency, interaction duration, etc.), forming a multi-dimensional heterogeneous relationship graph. The network topology can intuitively reflect the complex association patterns between user groups and the application ecosystem.

[0056] The construction of the heterogeneous graph network focuses on deeply mining the implicit association features in the behavior data, and transforming discrete user operation trajectories into computable graph embedding representations. Through the dynamic adjustment mechanism of edge weights, the differential interaction behaviors between individual users and different applications are accurately characterized, providing a structured data foundation for subsequent in-depth analysis based on graph neural networks to support the requirements of high-level tasks such as user value prediction.

[0057] S140. Predict the user lifecycle value according to the user-application heterogeneous graph network.

[0058] Exemplarily, the prediction of user lifecycle value is based on the fusion analysis of the topological features of the user-application heterogeneous graph network and the behavior time series data. By extracting the embedding vectors of user nodes, the functional attributes of application nodes, and the interaction intensity represented by edge weights in the heterogeneous graph network, combined with the time series features such as the active period and payment mode in the user's historical behavior, a multi-dimensional prediction index system is constructed to dynamically evaluate the potential value contribution degree of users within a preset time window.

[0059] This prediction process utilizes the modeling ability of graph neural networks for heterogeneous relationships to capture the long-term dependencies and cross-application association characteristics in user behavior patterns. By integrating the dynamic change laws of time-series data, it quantifies the decay or growth trend of user value and outputs the probability distribution of life-cycle value in stages, providing a core basis for accurately identifying high-value user groups and formulating differentiated operation strategies.

[0060] In summary, the user behavior analysis method of the embodiments of this application improves the accuracy and security of user behavior analysis through multi-dimensional operation behavior data collection, hierarchical cleaning and desensitization processing, and graph neural network modeling. Specifically, the technical solution realizes capturing high-fidelity behavior data based on regional fuzzification and pattern normalization technologies while protecting user privacy; ensuring data quality and compliance through multi-stage processing of anomaly detection, invalid data filtering, and encryption desensitization; and using a user-application heterogeneous graph network to deeply mine user behavior patterns and interaction characteristics, providing a reliable basis for accurately predicting user life-cycle value, thereby supporting the generation of dynamic and personalized recommendation services.

[0061] In some instances, obtaining the operation behavior data of cloud mobile phone users includes:

[0062] Capturing the user interface operation event stream through a preset buried point interface, where the operation event stream includes the original touch coordinate sequence, gesture type, and operation timestamp;

[0063] Performing regional fuzzification processing on the original touch coordinate sequence, mapping it to the block coding set of the preset screen grid division, and generating a fuzzy coordinate sequence;

[0064] Performing pattern normalization processing on the gesture type, mapping non-standard gestures to the matching codes in the preset standard gesture coding library, and generating a standardized gesture coding set.

[0065] Exemplarily, capturing the user interface operation event stream through a preset buried point interface, embedding a lightweight SDK into the underlying layer of the cloud mobile phone system, and dynamically intercepting user interaction behaviors in a non-intrusive manner. The buried point interface captures the original touch coordinate sequence, gesture type, and operation timestamp in real time through the Hook system API, where the touch coordinate sequence records the continuous trajectory of the user's touch operation, the gesture type covers interaction modes such as sliding, long pressing, and double clicking, and the operation timestamp is accurate to the millisecond level, ensuring the temporal integrity and operation scenario coverage of the behavior data. The data collection process follows the principle of minimization, only capturing the application-level interface information authorized by the user to avoid additional load on the system performance.

[0066] Perform regional blurring on the original touch coordinate sequence, divide the screen into an N×M grid block matrix, and each block corresponds to a unique code (such as using two-dimensional coordinate indexing or hash value identification). Through a space-filling curve algorithm (including Z-order curve or Hilbert curve), encode and transform the original coordinates, map the continuous coordinate points to the central area of the grid block they belong to, and generate a blurred coordinate sequence composed of block codes. This processing process completely eliminates precise location information. For example, map the coordinate (120, 240) to the "B3" block code, and at the same time retain the topological characteristics of the user operation path through the path continuity of the space-filling curve, providing a desensitized data basis for subsequent behavior pattern analysis.

[0067] Perform pattern normalization on the gesture types, and classify and match the original gesture data based on a preset standard gesture coding library (including rules such as a sliding angle tolerance of ±15° and a long-press time threshold of ≥500 ms). For non-standard gestures, calculate the similarity with the standard template through the dynamic time warping (DTW) algorithm and map it to the closest code (such as mapping a non-standard "L"-shaped slide to a "horizontal slide - code H01"). The generated standardized gesture coding set represents the user operation pattern with a discrete coding sequence, eliminating device differences and operation habit noise, ensuring the comparability and consistency of gesture behavior data among different users, and providing a standardized input for constructing the user behavior feature vector.

[0068] In some instances, perform regional blurring on the original touch coordinate sequence, map it to the block code set of a preset screen grid division, and generate a blurred coordinate sequence, including:

[0069] Divide the user operation screen into an N×M grid block matrix, where N and M are preset positive integers, and each grid block corresponds to a unique code;

[0070] Based on the grid block matrix, use a space-filling curve algorithm to traverse each coordinate point in the original touch coordinate sequence, and map each coordinate point to the block code of the grid block it belongs to, where the space-filling curve algorithm includes a Z-order curve or a Hilbert curve;

[0071] Generate a blurred coordinate sequence according to the mapping results of all coordinate points, where the blurred coordinate sequence is a sequence composed of the block codes corresponding to each coordinate point.

[0072] Exemplarily, the user operation screen is divided into an N×M grid block matrix according to preset rules, where N and M are independently configured positive integer parameters, and each grid block generates a unique code through a two-dimensional coordinate index or a hash algorithm (for example, the block code format is "row number-column number" or a hexadecimal string). The density of the grid division is determined by the privacy protection requirements and the accuracy of the behavior analysis. For example, in an 8×8 grid, each block covers 1 / 64 of the screen area, ensuring that the macroscopic distribution characteristics of the user's operation path can be retained after the coordinates are blurred, while eliminating precise location sensitive information. The grid boundary coordinates are dynamically calculated through the screen resolution to adapt to cloud phone terminal display devices of different sizes.

[0073] Based on the grid block matrix, the space filling curve algorithm is used to traverse and map the original touch coordinate sequence. Specifically, for each original coordinate point (x, y), the row and column indexes of the grid block to which it belongs are calculated, and the row and column indexes are converted into a globally unique one-dimensional sequence code based on the spatial continuity characteristics of the space filling curve (such as the Morton code of the Z-order curve or the continuous path code of the Hilbert curve). For example, in the Hilbert curve algorithm, the code of block (2,3) in the third-order curve is 0x23, which also reflects the neighboring relationship of the blocks in the two-dimensional space. This process converts the continuous coordinate sequence into a discrete block code sequence, while eliminating the details of the coordinate values, retaining the spatial correlation characteristics of the user operation trajectory through the path continuity of the space filling curve.

[0074] A fuzzy coordinate sequence is generated based on the mapping results of all coordinate points. The sequence consists of block codes with ordered timestamps. For example, the original coordinate sequence [(120,240),(130,250)] is mapped to ["B3","B4"] in an 8×8 grid. This sequence fully represents the spatiotemporal path of the user's operation, but only retains the block-level location information, and the original coordinates cannot be restored in reverse. By combining the preset encoding rules with the space filling algorithm, a balance is achieved between user privacy protection and the usability of behavior analysis, providing an input data foundation that meets privacy compliance requirements for the subsequent construction of a user-application heterogeneous graph network.

[0075] In some instances, the operation behavior data is cleaned and desensitized based on preset data processing rules to obtain standardized behavior data, including:

[0076] Detect abnormal operation sequences in the operation behavior data. If the duration of a single operation exceeds a first preset threshold or the interval between adjacent operations is less than a second preset threshold, mark it as abnormal data and remove it to generate valid operation data;

[0077] Perform irreversible desensitization on the user identification information in the valid operation data, and use the salted hash algorithm combined with a preset number of iterations to generate a desensitized user identification;

[0078] Perform regular expression matching on the URL data in the page browsing behavior, retain the access records that conform to the preset domain name whitelist rules, and generate compliant access records;

[0079] Integrate the valid operation data, desensitized user identifiers, and compliant access records to generate standardized behavior data.

[0080] Exemplarily, perform anomaly detection on the original operation behavior data based on a preset dual-threshold mechanism: when the single operation duration exceeds the first preset threshold (such as a long unresponsive operation) or the time interval between adjacent operations is less than the second preset threshold (such as high-frequency anomaly triggering), calculate the anomaly score through the Isolation Forest algorithm. If it exceeds the set threshold (such as 0.65), it is marked as abnormal data and excluded. This process uses a sliding window technique to scan the operation sequence in real time to ensure the exclusion of invalid records caused by device lags, malicious scripts, etc., and generate highly confident valid operation data, providing reliable input for subsequent analysis. It should be noted that the first preset threshold and the second preset threshold are determined based on the actual situation.

[0081] Desensitize the user identifiers (such as IMEI, account ID) in the valid operation data using the salted hash algorithm. The specific process is as follows: generate a random salt value and concatenate it with the original identifier, and perform hash operations for a preset number of iterations (such as 10,000 iterations) through a hash function that conforms to the GM / T0003-2012 standard (such as SM3) to generate an irreversible desensitized user identifier. This process effectively prevents rainbow table attacks through the uniqueness of the salt value and the computational complexity of the iteration, ensuring that the user identity information cannot be reversely restored and meeting the requirements of privacy protection regulations.

[0082] Perform regular expression matching on the URL data in the page browsing behavior. The preset domain name whitelist rule is ^https?: / / ([a-z0-9-]+.)+[a-z]{2,6}$, and only retain the access records that conform to this rule. For example, filter out URLs that contain direct IP connections, illegal characters, or unauthorized domain names (such as http: / / 192.168.1.1), and only retain compliant records such as https: / / www.example.com. This process efficiently eliminates potential risk data through the fast pattern matching ability of the regular expression engine and generates a standardized set of compliant access records.

[0083] Integrate the effective operation data (such as the sequence of application usage duration after cleaning), the desensitized user identifier (such as the hashed anonymous ID), and the compliant access record (such as the browsing log under the whitelist domain name) according to the preset field mapping rules. Generate the standardized behavior data stored in a structured manner through data format unification (such as converting the timestamp to the ISO 8601 standard format), field alignment (such as associating the user ID with the behavior record), and elimination of redundant information. This data is output in the form of key-value pairs or columnar storage to ensure the consistency of the data input format for subsequent graph neural network modeling and life cycle value prediction, and to meet the compatibility requirements of cross-module data interaction.

[0084] In some instances, based on the standardized behavior data, a user-application heterogeneous graph network is constructed using a graph neural network model, including:

[0085] According to the application usage duration and application startup frequency in the standardized behavior data, calculate the functional usage depth weight factor between the user and each application node. Among them, the weight factor is obtained by normalizing the product of the natural logarithm transformation after adding 1 to the single-day average usage duration and the startup frequency.

[0086] Based on the functional usage depth weight factor, construct the weighted edges between the user node and the application node to generate the initial heterogeneous graph network.

[0087] Through the graph attention network, perform multi-order neighborhood aggregation on the initial heterogeneous graph network to extract the user node embedding vector and the application node embedding vector, forming the user-application heterogeneous graph network.

[0088] Exemplarily, based on the user application usage duration and startup frequency metrics in the standardized behavior data, a composite quantization model is used to calculate the functional usage depth weight factor between the user and the application node. Specifically, the weight factor calculates the initial value through the product of the natural logarithm transformation after adding 1 to the single-day average usage duration (unit: minute) and the single-day startup frequency (unit: times / day), and is mapped to the [0, 1] interval through min-max normalization. This calculation method reflects both the user usage intensity and frequency characteristics, and eliminates the impact of data dimension differences on the graph structure modeling.

[0089] Taking the user and the application as heterogeneous nodes, construct bidirectional weighted edges based on the weight factor. The edge weight between the user node and the application node directly uses the functional usage depth weight factor, which characterizes the degree of functional dependence of the user on a specific application. For example, the edge weight between the user node U1 and the application node A1 is 0.75, and the edge weight with A2 is 0.62, forming a star topology centered on the user. The initial network retains the quantitative characteristics of the original interaction relationship, and at the same time realizes the preliminary structured expression of the user behavior pattern through the weighted edges, providing the basic graph data for subsequent high-order feature extraction.

[0090] The initial heterogeneous graph network is enhanced in features by using the Graph Attention Network (GAT), and the multi-hop neighbor information is aggregated through the multi-head attention mechanism. Specifically, for each user node, the attention coefficients within its k-hop neighborhood (such as directly associated application nodes and indirectly associated other user nodes) are calculated to dynamically adjust the contribution weights of different neighbor nodes to the current node. For example, the first-order neighbors of user node U1 are application nodes A1 and A2, and the second-order neighbors are other user nodes using A1. The feature vectors of these nodes are weighted and aggregated through the attention coefficients to generate user node embeddings that fuse local and global topological information. This process iteratively performs multi-order propagation to gradually capture the high-order association patterns in user behavior.

[0091] After multi-order neighborhood aggregation, the user nodes and application nodes respectively generate low-dimensional dense embedding vectors (such as 128-dimensional), and these vectors encode the structural characteristics and semantic relationships of the nodes in the heterogeneous graph. For example, the user node embedding vector can represent the comprehensive behavior preferences of the user, and the application node embedding vector reflects its functional attributes and the distribution characteristics of the user group. The finally formed user-application heterogeneous graph network optimizes the node embeddings through end-to-end training, making the nodes with similar topologies closer in the vector space. This network, as the core data structure for behavior analysis, provides an input representation containing multi-dimensional features such as spatio-temporal associations and functional dependencies for subsequent prediction of user lifecycle value.

[0092] In some instances, according to the heterogeneous graph network, the user lifecycle value is predicted, including:

[0093] Extract the user node embedding vector, application node embedding vector, and edge weight features from the user-application heterogeneous graph network to generate a user behavior topological feature set;

[0094] Fuse the user behavior topological feature set with the user historical behavior time series data, and the time series data includes the volatility of application usage duration and the interval period of payment behavior within a continuous preset number of days;

[0095] Input the fused features into the time series enhanced prediction model to output the probability distribution of the user lifecycle value within a future preset time window.

[0096] Exemplarily, multi-dimensional features are extracted from the user-application heterogeneous graph network, including the user node embedding vector (representing user behavior preferences), the application node embedding vector (reflecting application functional attributes), and the edge weight features (quantifying the interaction intensity between the user and the application). The embedding vectors are generated through multi-order neighborhood aggregation of the graph neural network. For example, the user node embedding vector contains the weighted functional usage depth information of its associated applications, and the edge weight features directly adopt the functional usage depth weight factor between the user and the application. Through feature splicing and dimensionality reduction processing, a user behavior topological feature set is generated, and this feature set encodes the structural position and interaction pattern of the user in the graph network in vector form.

[0097] Fuse the user behavior topological feature set with the user historical behavior time series data. The time series data includes the application usage duration volatility within a continuous preset number of days (calculated as the ratio of the standard deviation to the mean of the daily usage duration) and the payment behavior interval period (such as the time difference between two payment behaviors). The fusion process adopts a feature concatenation and attention weighting mechanism. For example, the topological features and the time series volatility are matched according to a time window through time series alignment technology, and the contribution ratio of the two types of features is dynamically adjusted using attention weights to form a fused multi-modal feature vector.

[0098] Input the fused features into a time series enhanced prediction model (such as the Temporal Fusion Transformer architecture), and capture long-term and short-term behavior dependencies through the multi-head self-attention mechanism. The model first performs hierarchical processing on the time series data to extract periodic and trend components; then performs cross-modal interaction with the topological features, and uses the cross-attention mechanism to identify key association patterns. Finally, output the probability distribution of the user lifecycle value within a future preset time window (such as 30 days, 90 days, 365 days). The distribution represents the confidence level of the user belonging to the high-value, medium-value, or low-value interval in the form of classification probabilities. For example, the probability of the high-value interval is 0.7, the medium-value is 0.2, and the low-value is 0.1.

[0099] The probability distribution of the user lifecycle value is output in the form of a multi-dimensional vector, including the expected value, confidence interval, and risk index within a preset time interval. For example, the 30-day prediction focuses on short-term activity and payment potential, the 90-day prediction evaluates the medium-term retention trend, and the 365-day prediction reveals long-term loyalty. The result is provided to the operation system through a visualization dashboard or an API interface to support dynamic resource allocation (such as preferentially pushing value-added services to high-value users), risk user intervention (such as triggering a retention strategy for users with a high probability of churn), and cross-cycle strategy optimization, ultimately improving the efficiency of user lifecycle management on the cloud mobile phone platform.

[0100] In some instances, it also includes:

[0101] Determine the user value level according to the probability distribution of the user lifecycle value;

[0102] Construct a recommendation action space, where the action space includes a set of recommendation channels, recommendation content types, and recommendation trigger times, and the trigger times are dynamically adjusted according to the distribution of user active periods;

[0103] Adopt the proximal policy optimization algorithm to dynamically adjust the recommendation policy parameters according to the real-time behavior feedback of the user, and generate a personalized recommendation policy that matches the user value level.

[0104] Exemplarily, when determining the user value level based on the probability distribution of the user life cycle value, the probability distribution is mapped to discrete value level labels through a preset threshold. For example, users with a high-value probability of ≥70% in the next 30 days are classified as "high-value users", those with a probability between 30%-70% are "medium-value users", and those with a probability <30% are "low-value users". The classification process combines auxiliary features such as the user's historical payment records and active days, and uses a clustering algorithm (such as K-means) to stratify the user group, ensuring that the level classification reflects both the confidence of the prediction result and the requirements of the actual business scenario.

[0105] The recommended action space consists of three dimensions: the set of recommended channels, the content type, and the triggering timing. The parameters of each dimension are preset according to business requirements. For example, the set of channels includes in-app pop-ups, lock screen ads, and system notification bars; the content types cover discount offers, feature recommendations, and membership upgrades; the triggering timing dynamically sets a time window (such as the active period ±1 hour) based on the distribution of user active periods (such as obtaining the active peak period through clustering analysis of historical operation timestamps). Multidimensional action combinations (such as the number of channels × the number of content types × the number of timings) are generated through the Cartesian product to form an action space covering all feasible strategies, ensuring the completeness and flexibility of strategy generation.

[0106] The Proximal Policy Optimization (PPO) algorithm is adopted to maximize the long-term value gain. The optimal policy combination is iteratively selected in the action space. The policy network adopts a deep neural network architecture, the hidden layer dimension is set to [256, 128], the discount factor γ = 0.95, and the action selection probability is output through the Softmax function. User feedback data (such as click-through rate, conversion delay duration, and negative feedback marks) is collected in real time, the policy gradient is calculated through the advantage function, and the network parameters are updated using the Adam optimizer (learning rate 0.001). The full-scale update of the policy model is triggered every 24 hours, and the hyperparameters (such as the exploration rate) are dynamically adjusted through the Bayesian optimization algorithm. If the revenue gain of K consecutive iterations is less than the preset convergence threshold (0.05), the incremental learning mechanism is started to optimize the generalization ability of the model.

[0107] By constructing a multi-dimensional matching mechanism for user value levels and recommendation action spaces, precise strategy adaptation is achieved. Specifically, for high-value users, high-discount offers and exclusive channel (such as the system notification bar) pushes are triggered preferentially, while for low-value users, low-cost awakening strategies (such as lock screen ads) are adopted; during the strategy execution process, conversion rates, resource consumption costs (such as push bandwidth occupancy rates), and user satisfaction indicators are monitored in real time. When the click-through conversion rate of a certain channel drops by more than 10%, its action selection probability is automatically reduced, and the action weights are dynamically adjusted. The finally generated personalized recommendation strategy is encapsulated in JSON format, including channel identifiers, content parameters, and trigger timestamps, and is distributed to the terminal for execution in real time through the cloud phone message middleware. This strategy integrates user value levels, action space options, and real-time feedback data to form dynamic decision rules. For example, exclusive discounts are pushed to high-value users during active periods, and the discount strength and push frequency are optimized in real time based on click feedback. After the strategy is executed, user behavior data (such as conversion rates, stay durations) flows back to the analysis system, triggering a new round of life cycle value prediction and strategy optimization, forming a closed-loop link of "prediction → recommendation → feedback → iteration" to continuously improve the recommendation accuracy and user experience.

[0108] In some instances, it also includes:

[0109] When it is detected that the user logs in to other associated devices, the multi-terminal behavior data is aggregated based on differential privacy technology. Specifically, Laplace noise is added to the local behavior data of each terminal (such as application usage records, operation frequencies) before data aggregation. The noise magnitude is inversely proportional to the user group size (such as the larger the group size, the smaller the noise amplitude), ensuring the statistical validity of the aggregation result (such as the cross-device usage hot zone distribution), while meeting the ε-differential privacy requirements (such as ε = 0.5). This process breaks the direct correlation between individual data and the aggregation result through noise injection, preventing the leakage of specific user behavior details through data reverse inference, and achieving a balance between privacy protection and data availability.

[0110] In the federated learning framework, each terminal locally trains a user life cycle value prediction model and only uploads the model parameter gradients to the cloud aggregation server. The parameter transmission uses a secure multi-party computing protocol to ensure that the gradient information remains in ciphertext state during transmission and aggregation. The cloud aggregation server generates global model parameters through weighted averaging (such as allocating weights according to device data volume) and distributes them to each terminal to update the local model. For example, the mobile phone side and the PC side respectively calculate the locally encrypted gradient values, which are decrypted by the cloud and then aggregated to generate new parameters, and then encrypted and sent back to the terminal. This mechanism achieves the dual goals of improving model performance and zero data exposure.

[0111] The minimum permissions for cross-device data sharing are defined by the scope parameter of the OAuth 2.0 protocol. The scope parameter is set to a preset collection of required fields (e.g., scope=basic_profile:read app_usage:read), which restricts third-party devices to only accessing basic profile data and application usage statistics fields and prohibits the acquisition of sensitive information (such as precise geographical location, device identifier). When generating the authorization token, the legitimacy of the requesting party is checked through a permission verification engine to ensure that data transmission follows the principle of minimization. For example, the associated PC can only obtain the user's active period and application preference tags and cannot access the original operation sequence or the user identifier before desensitization.

[0112] Multi-end user profile fusion is completed in a trusted execution environment (TEE), and Intel SGX technology is used to create an encrypted enclave to isolate the calculation process. The specific process is as follows: The encrypted behavior data of each device is input into the enclave, and operations such as feature alignment and weight fusion are decrypted and executed in the hardware-level encrypted memory. For example, the model parameters aggregated through federated learning are decrypted within the enclave and jointly modeled with local profile data to generate a unified cross-device user profile. After the calculation is completed, the output result is re-encrypted and returned to the external system, and the enclave memory state is immediately destroyed. This mechanism ensures that sensitive data cannot be snooped by the operating system or other processes throughout the entire processing flow through hardware-level isolation and memory encryption, meeting the security requirements of regulations such as GDPR for data processing environments.

[0113] Please refer to Figure 2 , which is a schematic structural diagram of a user behavior analysis device provided by an embodiment of the present application, including:

[0114] A data acquisition unit 21, configured to acquire the operation behavior data of a cloud phone user, where the operation behavior data includes application usage duration, application startup frequency, and page browsing behavior;

[0115] A data processing unit 22, configured to clean and desensitize the operation behavior data based on preset data processing rules to obtain standardized behavior data;

[0116] A model construction unit 23, configured to construct a user-application heterogeneous graph network based on the standardized behavior data by using a graph neural network model, where the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between users and applications;

[0117] A behavior prediction unit 24, configured to predict the user lifecycle value according to the user-application heterogeneous graph network.

[0118] Please refer to Figure 3, Embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for user behavior analysis.

[0119] Since the electronic device introduced in this embodiment is the device adopted for implementing a user behavior analysis device in an embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope to be protected by the present application.

[0120] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0121] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0122] Those skilled in the art should understand that the embodiments of the present application can provide methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.

[0123] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in one or more blocks or multiple blocks.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in one or more blocks or multiple blocks.

[0126] An embodiment of the present application also provides a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a user behavior analysis method in the corresponding embodiment.

[0127] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are fully or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium, an optical medium, or a semiconductor medium, etc.

[0128] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0129] In several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0130] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware and / or software functional units.

[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device to execute all or part of the steps of the methods in each embodiment of this application.

[0133] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.

[0134] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0135] Obviously, those skilled in the art can make various modifications and variations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and variations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification is also intended to include these modifications and variations.

Claims

1. A user behavior analysis method, characterized in that, The method includes: Obtaining the operation behavior data of cloud mobile phone users, where the operation behavior data includes application usage duration, application startup frequency, and page browsing behavior; Performing cleaning and desensitization processing on the operation behavior data based on preset data processing rules to obtain standardized behavior data; Based on the standardized behavior data, constructing a user-application heterogeneous graph network using a graph neural network model, where the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between users and applications; Predicting the user lifecycle value according to the user-application heterogeneous graph network.

2. The method according to claim 1, wherein The obtaining of the operation behavior data of cloud mobile phone users includes: Capturing the user interface operation event stream through a preset buried point interface, where the operation event stream includes the original touch coordinate sequence, gesture type, and operation timestamp; Performing regional blurring processing on the original touch coordinate sequence, mapping it to the block coding set of a preset screen grid division, and generating a blurred coordinate sequence; Performing pattern normalization processing on the gesture type, mapping non-standard gestures to the matching codes in a preset standard gesture coding library, and generating a standardized gesture coding set.

3. The method according to claim 2, wherein The performing of regional blurring processing on the original touch coordinate sequence, mapping it to the block coding set of a preset screen grid division, and generating a blurred coordinate sequence includes: Dividing the user operation screen into a grid block matrix, where each grid block corresponds to a unique code; Based on the grid block matrix, traversing each coordinate point in the original touch coordinate sequence using a space filling curve algorithm, and mapping each coordinate point to the block code of the grid block to which it belongs, where the space filling curve algorithm includes the Z-order curve or the Hilbert curve; Generating the blurred coordinate sequence according to the mapping results of all coordinate points, where the blurred coordinate sequence is a sequence composed of the block codes corresponding to each coordinate point.

4. The method according to claim 1, characterized in that, The performing of cleaning and desensitization processing on the operation behavior data based on preset data processing rules to obtain standardized behavior data includes: Detecting abnormal operation sequences in the operation behavior data. If the single operation duration exceeds a first preset threshold or the adjacent operation interval time is less than a second preset threshold, it is marked as abnormal data and excluded to generate valid operation data; Performing irreversible desensitization processing on the user identification information in the valid operation data, and generating a desensitized user identification using a salted hash algorithm combined with a preset number of iterations; Performing regular expression matching on the URL data in the page browsing behavior, retaining the access records that conform to the preset domain name whitelist rules, and generating compliant access records; Integrating the valid operation data, the desensitized user identification, and the compliant access records to generate the standardized behavior data.

5. The method according to claim 1, characterized in that, The constructing of a user-application heterogeneous graph network using a graph neural network model based on the standardized behavior data includes: Calculate the functional usage depth weight factor between the user and each application node according to the application usage duration and application startup frequency in the standardized behavior data, where the weight factor is obtained by normalizing the product of the logarithm transformation of the daily average usage duration plus 1 and the startup frequency; Based on the functional usage depth weight factor, construct weighted edges between the user node and the application node to generate an initial heterogeneous graph network; Perform multi-order neighborhood aggregation on the initial heterogeneous graph network through a graph attention network to extract the user node embedding vector and the application node embedding vector, forming a user-application heterogeneous graph network.

6. The method according to claim 1, wherein The predicting the user lifecycle value according to the user-application heterogeneous graph network includes: Extract the user node embedding vector, the application node embedding vector and the edge weight feature from the user-application heterogeneous graph network to generate a user behavior topology feature set; Fuse the user behavior topology feature set with the user historical behavior time series data, where the time series data includes the application usage duration volatility and the payment behavior interval period within a continuous preset number of days; Input the fused features into a time series enhanced prediction model to output the probability distribution of the user lifecycle value within a future preset time window.

7. The method according to claim 6, characterized in that It further includes: Determine the user value level according to the probability distribution of the user lifecycle value; Construct a recommendation action space, where the action space includes a set of recommendation channels, a type of recommended content, and a recommendation trigger time, and the trigger time is dynamically adjusted according to the user active period distribution; Adopt the proximal policy optimization algorithm to dynamically adjust the recommendation policy parameters according to the user real-time behavior feedback to generate a personalized recommendation policy matching the user value level.

8. A user behavior analysis device, characterized in that, The device includes: A data acquisition unit for acquiring the operation behavior data of the cloud mobile phone user, where the operation behavior data includes the application usage duration, the application startup frequency, and the page browsing behavior; A data processing unit for cleaning and desensitizing the operation behavior data based on preset data processing rules to obtain standardized behavior data; A model construction unit for constructing a user-application heterogeneous graph network by using a graph neural network model based on the standardized behavior data, where the heterogeneous graph network includes user nodes, application nodes, and edge weights representing the interaction intensity between the user and the application; A behavior prediction unit for predicting the user lifecycle value according to the user-application heterogeneous graph network.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the user behavior analysis method according to any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program, when executed by the processor, implements the user behavior analysis method according to any one of claims 1 to 7.