A method and system for constructing user portraits based on user habit analysis

By conducting fine-grained analysis of historical data on users' mobile data plans and assessing bandwidth instability, we build a user behavior profile, solving the problem of inaccurate network usage status analysis in traditional methods and achieving more accurate package recommendations and improved user experience.

CN120317918BActive Publication Date: 2025-09-12GUANGDONG LEGEND COMM CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510803681.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-12
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The traditional user profile construction method based on user habit analysis is inaccurate in analyzing network usage status, resulting in large errors in constructing user usage habit profiles and low accuracy in target package recommendations.

Method used

By obtaining the historical usage data of users' mobile phone data packages, we conduct fine-grained analysis of data usage, identify the data usage status between different apps, analyze sudden network occupancy and assess bandwidth instability, build a user behavior profile, and recommend target package types.

Benefits of technology

It improves the accuracy of analysis of user network usage status, reduces the error in constructing user usage habit portraits, improves the accuracy of target package recommendations, and enhances user experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317918B_ABST
    Figure CN120317918B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of user portrait construction, and in particular to a method and system for constructing a user portrait based on user habit analysis. The method comprises the following steps: first, by obtaining the historical usage data of the user's mobile phone traffic package, a fine-grained analysis is performed to obtain the traffic usage status data of different APPs. Then, based on these data, the sudden network occupancy when multiple APPs are used simultaneously is analyzed, and the fluctuation characteristics of bandwidth instability are evaluated to obtain bandwidth instability clustering feature data. Next, a dependency structure analysis is performed on the bandwidth instability clustering features to construct the dependency relationship of the user's unstable behavior, and based on this, a user usage behavior portrait is generated. Finally, the user portrait and the bandwidth instability characteristics are combined to recommend the target package type and provide personalized package suggestions. The present invention makes the user portrait construction technology more perfect by optimizing the user portrait construction technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user portrait construction, and in particular to a user portrait construction method and system based on user habit analysis. Background Art

[0002] With the increasing volume of data and the diversity of user behavior, relying solely on traditional data analysis methods can no longer meet the requirements for accurate user profiles. This is especially true in complex scenarios such as sudden network usage and bandwidth instability. Previous methods often struggle to effectively capture and evaluate user behavior. Therefore, a user profile construction method based on fine-grained analysis has emerged. This method not only focuses on user behavior patterns but also comprehensively analyzes multiple factors (such as network usage and device performance) to provide a more accurate profile of user habits. However, traditional user profile construction methods based on user habit analysis suffer from inaccurate analysis of user network usage, resulting in large errors in user habit profile construction and low accuracy in target package recommendations. Summary of the Invention

[0003] Based on this, it is necessary to provide a user portrait construction method and system based on user habit analysis to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a method for constructing a user profile based on user habit analysis is provided, the method comprising the following steps:

[0005] Step S1: Obtain historical usage data of the user's mobile phone data package; perform fine-grained traffic usage analysis on the historical usage data of the user's mobile phone data package to obtain fine-grained data on traffic usage status between different apps;

[0006] Step S2: Based on the fine-grained data of traffic usage status between different apps, the sudden network occupancy status between multiple apps in synchronous usage is analyzed to obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated for the sudden network occupancy status to obtain bandwidth instability clustering feature data;

[0007] Step S3: Perform dependency structure analysis on the bandwidth instability clustering feature data to obtain the unstable clustering habit behavior dependency structure; construct a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile; recommend a target recommended package type based on the user usage behavior profile and the bandwidth instability clustering feature data to obtain target recommended package type data.

[0008] Preferably, step S1 includes the following steps:

[0009] Step S11: Obtaining historical usage data of the user's mobile phone data package;

[0010] Step S12: Clean the historical usage data of the user's mobile phone data package to obtain the cleaned data of the historical usage of the data package;

[0011] Step S13: extracting the traffic usage status between different APPs from the traffic package history usage cleaning data to obtain the traffic usage status between different APPs;

[0012] Step S14: Perform fine-grained traffic usage analysis on the traffic usage status between different APPs to obtain fine-grained data on the traffic usage status between different APPs.

[0013] Preferably, step S2 includes the following steps:

[0014] Step S21: performing traffic usage intensity analysis between different APPs on the fine-grained data of traffic usage status between different APPs to obtain traffic usage intensity data between different APPs;

[0015] Step S22: analyzing the sudden network occupancy status of multiple apps in synchronous use based on the traffic usage intensity data between different apps to obtain the sudden network occupancy status;

[0016] Step S23: performing a network bandwidth fluctuation instability assessment on the sudden network occupancy state to obtain network bandwidth fluctuation instability data;

[0017] Step S24: Based on the network bandwidth fluctuation instability data, weighted sampling and clustering processing is performed on the traffic usage intensity data between different apps to obtain bandwidth instability clustering feature data.

[0018] Preferably, step S23 includes the following steps:

[0019] Step S231: performing sudden network bandwidth fluctuation analysis on the sudden network occupancy state to obtain sudden network bandwidth fluctuation data;

[0020] Step S232: quantifying the multi-frequency fluctuation peak integral of the sudden network bandwidth fluctuation data to obtain the bandwidth multi-frequency fluctuation peak integral;

[0021] Step S233: performing delay jitter regression analysis on the sudden network bandwidth fluctuation data according to the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth delay jitter regression data;

[0022] Step S234: performing congestion equal-amount proportional mapping based on the bandwidth delay jitter regression data and the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth congestion equal-amount proportional mapping data;

[0023] Step S235: perform network bandwidth fluctuation instability assessment based on bandwidth multi-frequency fluctuation peak integral, bandwidth delay jitter regression data, and bandwidth congestion equal-amount proportional mapping data to obtain network bandwidth fluctuation instability data.

[0024] Preferably, step S234 includes the following steps:

[0025] Calculate the time series mean difference of the bandwidth delay jitter regression data to obtain the bandwidth delay jitter time series mean difference;

[0026] The bandwidth congestion factor distribution data is obtained by performing numerical orthogonal projection transformation based on the bandwidth delay jitter time series mean difference and the bandwidth multi-frequency fluctuation peak integral;

[0027] Congestion equal proportion mapping processing is performed according to the bandwidth congestion factor distribution data to obtain bandwidth congestion equal proportion mapping data.

[0028] Preferably, step S24 includes the following steps:

[0029] Step S241: performing amplitude discrete trend fitting processing on the network bandwidth fluctuation instability data to obtain amplitude discrete trend fitting data;

[0030] Step S242: performing weighted reconstruction processing on the traffic usage intensity data of different apps based on the amplitude discrete trend fitting data to obtain traffic usage intensity reconstructed data;

[0031] Step S243: performing clustering iterative optimization on the traffic usage intensity reconstruction data to obtain traffic usage intensity reconstruction cluster data;

[0032] Step S244: reconstructing cluster data according to traffic usage intensity and performing weighted sampling clustering processing to obtain bandwidth instability clustering feature data.

[0033] Preferably, step S3 includes the following steps:

[0034] Step S31: mining habit behavior chains on bandwidth instability clustering feature data to obtain instability clustering habit behavior chain data;

[0035] Step S32: performing dependency structure analysis on the unstable clustering habit behavior chain data to obtain the unstable clustering habit behavior dependency structure;

[0036] Step S33: constructing a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile;

[0037] Step S34: Recommend target package types based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommended package type data.

[0038] Preferably, step S33 includes the following steps:

[0039] Step S331: performing normalization processing on the unstable clustering behavior frequency of the unstable clustering habit behavior dependency structure to obtain normalized unstable clustering behavior frequency data;

[0040] Step S332: performing logic back-inference merging processing on the normalized unstable clustering behavior frequency data to obtain a behavior combination deduction structure;

[0041] Step S333: performing behavior coupling degree division processing on the behavior combination deduction structure to obtain a set of multiple behavior coupling patterns of the user;

[0042] Step S334: construct a user usage behavior profile based on the user's multiple behavior coupling pattern set to obtain the user usage behavior profile.

[0043] Preferably, the present invention further provides a user profile construction system based on user habit analysis, which is used to execute the user profile construction method based on user habit analysis as described above. The user profile construction system based on user habit analysis includes:

[0044] The traffic usage fine-grained analysis module is used to obtain the historical usage data of the user's mobile phone traffic package; the traffic usage fine-grained analysis is performed on the historical usage data of the user's mobile phone traffic package to obtain fine-grained data on the traffic usage status of different apps;

[0045] The bandwidth instability cluster analysis module is used to analyze the sudden network occupancy status of multiple apps in simultaneous use based on the fine-grained data of traffic usage status between different apps, and obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated based on the sudden network occupancy status, and the bandwidth instability clustering feature data is obtained;

[0046] A behavior profile construction module is used to perform dependency structure analysis on bandwidth instability clustering feature data to obtain an unstable clustering habit behavior dependency structure; a user usage behavior profile is constructed based on the unstable clustering habit behavior dependency structure to obtain a user usage behavior profile; and target recommendation package type data is recommended based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommendation package type data.

[0047] The present invention has the beneficial effect of accurately identifying traffic usage across different apps by acquiring and fine-grainedly analyzing historical usage data for a user's mobile data plan. The key to this step is in-depth analysis of traffic usage, which can reveal the user's traffic distribution across different applications. This analysis allows for an understanding of a user's frequently used apps and usage habits, providing a reliable data foundation for subsequent analysis. Fine-grained data not only provides insight into a user's traffic needs but also lays a solid foundation for accurately predicting their future traffic needs and recommending appropriate plans. Based on traffic usage data across different apps, the system analyzes the sudden network occupancy that occurs when multiple apps are used simultaneously. This analysis process can reveal fluctuations in a user's network bandwidth demand when using multiple apps. In particular, in the case of unstable network bandwidth, the system can proactively identify the risk of bandwidth fluctuations, conduct instability assessments, and generate clustering feature data related to bandwidth instability. This system can accurately capture anomalies in network usage and provide valuable reference data for subsequent network optimization and user behavior analysis, improving user experience and reducing the risk of network congestion. Dependency structure analysis of the bandwidth instability clustering feature data can deeply uncover underlying patterns in user behavior and identify the core factors that influence the user's network experience. By constructing an unstable clustering habit behavior dependency structure, we can clearly understand which usage habits will cause bandwidth fluctuation instability and which behavior patterns will aggravate network problems. Based on this structure, we can further construct a user usage behavior portrait, which can not only comprehensively depict the user's traffic needs, but also accurately grasp their network usage preferences. Through this process, it is possible to accurately identify user needs and provide personalized and targeted package recommendations, thereby improving service quality and increasing user satisfaction. Therefore, the present invention is an optimization of a traditional user portrait construction method based on user habit analysis, which solves the problem that a traditional user portrait construction method based on user habit analysis has inaccurate analysis of the user's network usage status, resulting in large errors in the construction of user usage habit portraits and low accuracy in target package recommendations. It improves the accuracy of the analysis of the user's network usage status, reduces the errors in the construction of user usage habit portraits, and improves the accuracy of target package recommendations. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A flowchart of the steps for building a user portrait based on user habit analysis;

[0049] Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG.

[0050] Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG. DETAILED DESCRIPTION

[0051] See also Figures 1 to 3 , a method for constructing a user portrait based on user habit analysis, the method comprising the following steps:

[0052] Step S1: Obtain historical usage data of the user's mobile phone data package; perform fine-grained traffic usage analysis on the historical usage data of the user's mobile phone data package to obtain fine-grained data on traffic usage status between different apps;

[0053] Step S2: Based on the fine-grained data of traffic usage status between different apps, the sudden network occupancy status between multiple apps in synchronous usage is analyzed to obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated for the sudden network occupancy status to obtain bandwidth instability clustering feature data;

[0054] Step S3: Perform dependency structure analysis on the bandwidth instability clustering feature data to obtain the unstable clustering habit behavior dependency structure; construct a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile; recommend a target recommended package type based on the user usage behavior profile and the bandwidth instability clustering feature data to obtain target recommended package type data.

[0055] In the embodiment of the present invention, reference Figure 1 The above is a schematic flow chart of the steps of a method for constructing a user profile based on user habit analysis of the present invention. In this example, the method for constructing a user profile based on user habit analysis includes the following steps:

[0056] Step S1: Obtain historical usage data of the user's mobile phone data package; perform fine-grained traffic usage analysis on the historical usage data of the user's mobile phone data package to obtain fine-grained data on traffic usage status between different apps;

[0057] In an embodiment of the present invention, the user's mobile data package usage data for the past 12 months is first retrieved from the communication operator's database and the data is desensitized. The data desensitization method adopts a multi-level protection mechanism to ensure that user privacy is not leaked. Specifically, information related to the user's identity, such as mobile phone number, IMSI number, IMEI number, precise coordinates in location information, base station number, etc., is first desensitized. Among them, the user identity identifier is irreversibly encrypted using a hash salt algorithm to ensure that the original identity cannot be restored; the location information is fuzzified to reduce the accuracy to the district or county level or the center coordinates of the cluster area; at the same time, records involving communication behavior (such as using timestamps) are processed using time offset or time period classification to avoid accurate trajectory restoration. In addition, all user data must be verified for permissions through a data access control system before processing, and ensure that it is only collected and used under the premise of explicit authorization by the user. The scope of desensitization includes, but is not limited to: user unique identification codes, device identification information, precise geographic location information, IP addresses, MAC addresses, and raw timestamp data strongly related to user behavior. Users are also asked to consent to the collection of their mobile data usage data for research purposes. Research can only be conducted after obtaining user authorization. During the data collection phase, end-to-end encryption protocols (such as TLS 1.3) are used to ensure the security of data during transmission between the collection device and the server. During data transmission, VPNs or dedicated network channels are used to prevent man-in-the-middle attacks and illegal eavesdropping. During data storage, all sensitive data is stored in an encrypted database, encrypted at rest using the AES-256 algorithm, and access control policies are implemented to ensure that only authorized personnel can access it. During data processing, a sandbox environment and desensitization processing are executed in parallel to ensure that data is not exported or exposed before desensitization. At the same time, the system has established behavioral auditing and anomaly detection mechanisms to record and monitor all data operations in real time, promptly identifying and blocking abnormal behavior, thereby building a closed-loop security protection system covering all links of collection, transmission, storage, and processing. This includes the total monthly package traffic, actual traffic used, remaining traffic, package traffic allocation method (such as targeted traffic, general traffic), usage details for each type of traffic, and usage time distribution. The user's app usage details are extracted on an hourly basis, and the uplink and downlink traffic data for each app in each time period is recorded in KB. The user's access network type (4G, 5G, WiFi, etc.), geographic location (latitude and longitude corresponding to the base station number), and signal strength (RSRP value) are also extracted. Data cleaning operations are performed on the above data, including processing packet loss data, data time alignment (uniformly segmented by minute level), and invalid record removal. Afterwards, the data is segmented using a time series segmentation method based on a sliding window. The usage records of each app are analyzed by sliding in a 10-minute time window, recording the traffic value, access network, and activity duration in each time window, and marking whether it is in the background or foreground state.By conducting comprehensive statistics on the differences in upstream and downstream traffic of each APP in the same time window, the activity status ratio, the fluctuation amplitude within the time window span, etc., a fine-grained data table of traffic usage status with "time window-APP" as the basic unit is constructed. The fields include APP identification, start time, end time, foreground and background status, upstream and downstream traffic, access type, average signal strength, maximum rate, etc.

[0058] Step S2: Based on the fine-grained data of traffic usage status between different apps, the sudden network occupancy status between multiple apps in synchronous usage is analyzed to obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated for the sudden network occupancy status to obtain bandwidth instability clustering feature data;

[0059] In this embodiment of the present invention, the "time window-APP" data generated in step S1 is first analyzed, and a sliding pairing method is used to construct the synchronous usage status of multiple apps within the same time window. Specifically, if two or more apps are in the foreground within any 10-minute time window and their traffic usage exceeds 80% of their respective average traffic usage, it is considered a synchronous usage state. For each synchronous usage time window, statistics are collected to determine the number of participating apps, total upstream and downstream traffic, average signal strength, and whether the user has experienced a network switch (e.g., switching from 4G to 5G). Next, based on these synchronous usage states, a time window sudden increase determination algorithm (based on first-order differential rate of change and rising slope determination) is used to identify sudden network occupancy states. If the total traffic surge exceeds twice the historical average and the average rate increases by more than 50% within two consecutive time windows, it is marked as a sudden state. Based on this, the sudden network occupancy state is evaluated for network bandwidth fluctuation and instability. During the evaluation process, the channel change amplitude analysis method is used to record multiple channel indicators such as changes in user signal strength before and after the burst state, average rate changes, RSRP standard deviation, RSRQ index, packet loss rate, etc., and calculate the weighted average bandwidth fluctuation value after multi-dimensional data normalization. Clustering is performed based on the fluctuation period length (i.e., the duration of continuous high fluctuations). The clustering algorithm uses the DBSCAN algorithm with a minimum cluster size of 5 and a neighborhood radius of 0.6 to cluster the bandwidth instability state. The representative characteristics of each cluster are obtained, including the average fluctuation value, fluctuation period, and degree of channel degradation, thereby forming bandwidth instability clustering feature data.

[0060] Step S3: Perform dependency structure analysis on the bandwidth instability clustering feature data to obtain the unstable clustering habit behavior dependency structure; construct a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile; recommend a target recommended package type based on the user usage behavior profile and the bandwidth instability clustering feature data to obtain target recommended package type data.

[0061] In this embodiment of the present invention, the bandwidth instability clustering feature data obtained in step S2 is first subjected to behavioral dependency analysis. The behavioral dependency structure is constructed using a directed graph-based behavior chain extraction method. The instability clustering features in each user's behavior time series are mapped chronologically, recording the causal relationship between each behavior node (e.g., a burst of usage caused by the use of a certain app) and subsequent behavior nodes (e.g., the activation of other apps or a sudden increase in traffic). Directed edges are established by calculating the transition frequency and behavior trigger latency (i.e., the average time from the occurrence of the previous behavior to the initiation of the next behavior) between different behavior nodes. The dependency strength is quantified as edge weights to construct a dependency structure for unstable clustering habitual behaviors. This structure graph is then subjected to frequent path mining (using the Apriori algorithm with a minimum support of 0.4) to extract high-frequency user behavior chains and categorize them into different behavioral patterns, such as the "video app + social app + disconnection and reconnection" pattern and the "download tool + map navigation + background audio" pattern. Next, based on the aforementioned dependency structure and frequent paths, user usage patterns are encoded along multiple dimensions, such as behavior type, traffic characteristics, burst frequency, and switching rate, to form a multidimensional feature vector. Principal component analysis (PCA) is then applied to reduce the dimensionality to a three-dimensional space. Within this space, similar user profiles are distributedly represented to form a user behavior profile. Finally, based on the constructed user behavior profile and bandwidth instability clustering feature data, similar profile vectors are retrieved (using Cosine similarity) to match historical package types with user behavior profiles in the profile database. Historical profile records with a similarity greater than 0.85 with the current user profile are identified and their corresponding recommended package types are extracted. For all matching package types, user satisfaction scores (e.g., low complaint rate, high renewal rate, and average monthly remaining traffic less than 5%) are calculated and prioritized based on the scores. The top three package types are selected and the highest-scoring package type is output as the target recommended package type. The recommended results include fields such as package name, traffic allowance, profile ID, and recommendation confidence value.

[0062] Step S1 includes the following steps:

[0063] Step S11: Obtaining historical usage data of the user's mobile phone data package;

[0064] Step S12: Clean the historical usage data of the user's mobile phone data package to obtain the cleaned data of the historical usage of the data package;

[0065] Step S13: extracting the traffic usage status between different APPs from the traffic package history usage cleaning data to obtain the traffic usage status between different APPs;

[0066] Step S14: Perform fine-grained traffic usage analysis on the traffic usage status between different APPs to obtain fine-grained data on the traffic usage status between different APPs.

[0067] In an embodiment of the present invention, the specific process of obtaining the historical usage data of the user's mobile phone data package is carried out by connecting with the historical billing data interface of the communication operator. On the premise that the data collection period is set to nearly 12 months, a detailed record containing the user's unique identification identifier, package type number, package effectiveness and expiration time, monthly package total traffic, directional traffic details (such as video, social, office), general traffic quota, daily traffic consumption, hourly traffic usage distribution, and APP identification and traffic attribution is obtained. The data is collected in hours, and the data is desensitized. Users are asked whether they agree to collect their mobile phone traffic usage data for research. Research can only be carried out after obtaining user authorization. Daily traffic is constructed into original data sets according to fields such as APP name (APP type is identified by operator mapping IP address port segment), usage time, upstream and downstream traffic, network type (4G, 5G), base station number, user access point location (mapped to latitude and longitude through LAC-CI), and traffic billing type (general / directional). The data field structure is uniformly in CSV format and encoded in UTF-8. The daily data volume for each user is no less than 288 items (calculated based on a 5-minute time window) to ensure high-precision sampling granularity.

[0068] A cleansing process is performed on the acquired historical data on data plan usage. The first step is a field integrity check. All records are identified for null values, duplicate records, and illegal characters. Records with missing identifier fields are deleted, duplicate data with the same timestamp is removed, and records with negative or abnormally high traffic values ​​(e.g., traffic exceeding 2GB in a single 5-minute period) are eliminated. The second step is time standardization. Any differences in timestamp precision (e.g., seconds or minutes) among different records are converted to a 5-minute granularity window in Beijing time format. The UNIX timestamp is divided by 300 and rounded down to the nearest integer to obtain a standard time period number. The third step verifies the accuracy of app identification. By comparing the IP address to the app identification table provided by the operator (including feature information such as IP address range, port number, SNI, and DNS domain name), IP attribution is re-verified to eliminate misidentified and unidentified app records. The fourth step is noise data removal. By calculating the standard deviation of each user's traffic distribution on a daily basis, records with excessively low fluctuation coefficients (e.g., entries with only a small amount of directional traffic, accounting for less than 0.1% of the total traffic over 24 consecutive hours) are identified and eliminated. The resulting cleaned data set contains fields such as user ID, time period number, app name, uplink and downlink traffic (in KB), access network, location identifier, and package traffic type (targeted or universal). The data is sorted in ascending order by user ID and time. The cleaned data set's historical data is then used to extract usage status by app. First, the app records appearing in each time period for each user are merged. The uplink and downlink traffic for the same app in the same time window across multiple records is summed. The app's active foreground status within that time window is then determined. This determination is based on whether the network data packet frequency exceeds once every five seconds, combined with the app's activation status tag recorded in the system log. A traffic usage status table is then constructed for each app. The table header contains the following fields: time window number, uplink and downlink traffic, network type, signal strength (average RSRP), activity status, traffic type (general / targeted), and location information. Then, the APP aggregation operation under the user dimension is performed on the status table to calculate the average daily usage times, peak usage period (the time window number with the most occurrences), and maximum daily traffic usage of each APP. A timeline graph of APP usage is created to provide a basis for subsequent analysis of synchronous usage behavior and bandwidth usage behavior.

[0069] Based on the extracted inter-app traffic usage data, a fine-grained traffic usage analysis is conducted. A sliding time window analysis method is used, with a window width of 10 minutes and a sliding step of 5 minutes, to construct a continuous time window sequence. The upstream and downstream traffic of each app within each time window is differentiated to calculate the rate of change of traffic per unit time. The traffic fluctuation amplitude (the difference between the maximum and minimum values) and the traffic growth slope (the linear fitting regression coefficient) are introduced. Furthermore, the number of times an app is active in the foreground and the duration of active activity within the continuous time window are counted to generate an app activity spectrum. Using a frequency distribution density estimation method, a three-dimensional density distribution function is constructed based on active time periods, traffic values, and network types. The peak traffic usage intervals for each app within a day are determined, and the co-occurrence probability between apps within the same time window is statistically analyzed. For app combinations with a co-occurrence probability exceeding 0.6, their traffic share structure is further analyzed and labeled as potential traffic synchronization clusters, laying the foundation for subsequent synchronization behavior identification and unstable bandwidth pattern analysis. The resulting fine-grained dataset contains fields such as user ID, app name, time window number, upstream and downstream traffic, traffic change rate, foreground and background status, co-occurrence probability, network access type, signal strength mean, fluctuation range, and active duration. This dataset provides behaviorally granular support for further building user profiles.

[0070] Step S2 includes the following steps:

[0071] Step S21: performing traffic usage intensity analysis between different APPs on the fine-grained data of traffic usage status between different APPs to obtain traffic usage intensity data between different APPs;

[0072] Step S22: analyzing the sudden network occupancy status of multiple apps in synchronous use based on the traffic usage intensity data between different apps to obtain the sudden network occupancy status;

[0073] Step S23: performing a network bandwidth fluctuation instability assessment on the sudden network occupancy state to obtain network bandwidth fluctuation instability data;

[0074] Step S24: Based on the network bandwidth fluctuation instability data, weighted sampling and clustering processing is performed on the traffic usage intensity data between different apps to obtain bandwidth instability clustering feature data.

[0075] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:

[0076] Step S21: performing traffic usage intensity analysis between different APPs on the fine-grained data of traffic usage status between different APPs to obtain traffic usage intensity data between different APPs;

[0077] In an embodiment of the present invention, when analyzing the traffic usage intensity of fine-grained data on traffic usage status between different apps, the app usage records for each user within each 5-minute time window are first summarized, and the upstream and downstream traffic and activity status of each app within each time window are counted separately by app name. Next, a traffic intensity vector for each app within 24 hours is constructed. The vector dimension is 288 (one dimension every 5 minutes), and each dimension records the total traffic value (KB) in the corresponding time window, with upstream and downstream models modeled separately. Then, the traffic values ​​are normalized using standard deviation normalization to eliminate the impact of differences in the total usage of different apps on the intensity analysis results. The intensity synergy coefficient of each pair of apps in the same time window is calculated. This coefficient is defined as the ratio of the product of the traffic values ​​of the two apps in the same time window to their overall mean, which is used to measure the intensity of the two apps' simultaneous occupation of bandwidth resources. For each user per day, a traffic intensity synergy matrix is ​​constructed for multiple app pairs. Each matrix has an N×N dimension, where N is the number of active apps for the user on that day, and each matrix element records the average intensity synergy value of the corresponding app pair. In addition, the daily average maximum traffic window position, intensity standard deviation, and all-day active window ratio are calculated for each APP, and the results are summarized into an APP intensity attribute data table to provide parameter input for sudden occupancy status analysis.

[0078] Step S22: analyzing the sudden network occupancy status of multiple apps in synchronous use based on the traffic usage intensity data between different apps to obtain the sudden network occupancy status;

[0079] In this embodiment of the present invention, based on the traffic usage intensity data obtained above between different apps, bursty network occupancy during simultaneous usage of multiple apps is identified. First, a sliding analysis is performed on the traffic intensity coordination matrix constructed in the previous step. Within each 15-minute analysis period, a Fast Fourier Transform (FFT) is performed on the app-pair coordination values ​​to extract high-frequency variations, thereby identifying transient jumps in bandwidth usage. The threshold for determining bursty occupancy is set to a coordination intensity that is 1.5 times higher than the mean and persists for more than two analysis windows (i.e., 30 minutes). Next, app combinations meeting the aforementioned conditions are identified in the app-pair coordination matrix. The time period during which a bursty state occurs is recorded as a bursty simultaneous occupancy window. Within each bursty window, features such as the maximum total uplink and downlink traffic, the number of apps involved, the average active time window, the average network type (e.g., 4G / 5G), and the time period (e.g., whether it is peak nighttime) are extracted to form a bursty network occupancy record. Then, a daily burst occupancy event set is constructed for each user. Each record contains: start time, end time, occupied APP group, maximum instantaneous traffic value, occupancy intensity index, occupancy duration, terminal access type, and area number (base station location identifier), which serves as the basic input for subsequent bandwidth fluctuation and instability assessment.

[0080] Step S23: performing a network bandwidth fluctuation instability assessment on the sudden network occupancy state to obtain network bandwidth fluctuation instability data;

[0081] In this embodiment of the present invention, network bandwidth fluctuation and instability are assessed based on the acquired bursty network occupancy status. First, the raw uplink and downlink traffic data for the time window involved in each burst is arranged in chronological order. A third-order difference operation is performed on each record to extract the fluctuation trend. The local fluctuation energy is calculated using the wavelet transform method to obtain a fluctuation intensity spectrum. Then, based on the start and end time periods of the burst, the mean, standard deviation, and maximum jump amplitude (i.e., the maximum difference value per unit time) of the uplink and downlink fluctuations per unit time are calculated. The fluctuation duration (the proportion of time within the burst period when the fluctuation exceeds a certain threshold) is introduced for feature quantification. Subsequently, a distribution fitting process is performed, using a normal mixture distribution to estimate the burst fluctuation indicators for each user. The fitted mean, variance, and crest factor are extracted as the bandwidth fluctuation parameter features of the user. Next, a mapping table is established from bursts to fluctuation features. Each burst record is appended with fields such as the corresponding fluctuation intensity value, burst duration, network signal mean, access mode identifier, and region to which it belongs. This forms a complete network bandwidth fluctuation and instability dataset, providing input for cluster analysis in the next step.

[0082] Step S24: Based on the network bandwidth fluctuation instability data, weighted sampling and clustering processing is performed on the traffic usage intensity data between different apps to obtain bandwidth instability clustering feature data.

[0083] In this embodiment of the present invention, based on the aforementioned network bandwidth fluctuation and instability data, a weighted sampling and clustering process is performed on traffic usage intensity data across different apps. First, a bandwidth fluctuation feature vector is constructed for each user's daily emergency event record. The vector dimensions include 10 dimensions, including the fluctuation mean, maximum amplitude, duration, number of active apps, average synchronization intensity, and network type index. Subsequently, bandwidth fluctuation intensity is introduced as a weight, and weighted sampling is performed on each feature vector. High-intensity emergency event samples are prioritized. The sample ratio is set to retain 100% of the samples with high fluctuation intensity that account for the top 20% of all events, 60% of the samples with medium fluctuation intensity, and 20% of the samples with low fluctuation intensity that are randomly sampled. Density-based spatial clustering (DBSCAN) processing is performed on all samples, with a distance threshold ε of 0.8 and a minimum number of sample points of 10. Clustering is performed based on the Euclidean distance of the fluctuation feature vectors, resulting in several clusters of unstable bandwidth patterns. For each cluster, statistics are collected on its APP combination characteristics, fluctuation characteristic center value, active time period preference, and network signal environment average value, to generate a bandwidth instability cluster feature dataset. The fields include: cluster number, center fluctuation value, representative APP set, average synchronization index, typical network type, representative area number, etc., which provide the input basis for subsequent behavior dependency structure analysis and user portrait modeling.

[0084] Step S23 includes the following steps:

[0085] Step S231: performing sudden network bandwidth fluctuation analysis on the sudden network occupancy state to obtain sudden network bandwidth fluctuation data;

[0086] Step S232: quantifying the multi-frequency fluctuation peak integral of the sudden network bandwidth fluctuation data to obtain the bandwidth multi-frequency fluctuation peak integral;

[0087] Step S233: performing delay jitter regression analysis on the sudden network bandwidth fluctuation data according to the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth delay jitter regression data;

[0088] Step S234: performing congestion equal-amount proportional mapping based on the bandwidth delay jitter regression data and the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth congestion equal-amount proportional mapping data;

[0089] Step S235: perform network bandwidth fluctuation instability assessment based on bandwidth multi-frequency fluctuation peak integral, bandwidth delay jitter regression data, and bandwidth congestion equal-amount proportional mapping data to obtain network bandwidth fluctuation instability data.

[0090] In this embodiment of the present invention, the process of analyzing sudden network bandwidth fluctuations in response to bursty network occupancy states begins by extracting the network traffic time series corresponding to each burst event. This time series consists of uplink and downlink traffic values ​​at a granularity of 1 second during the event. Wireless signal parameters such as RSRP (reference signal received power), RSRQ (reference signal received quality), and SINR (signal-to-noise ratio) are simultaneously collected from the network to construct a basic dataset for fluctuation analysis. A sliding window difference calculation is performed on each traffic time series, with a window length of 5 seconds and a sliding step of 1 second, to obtain a bandwidth change rate series. Points in the change rate series where the amplitude of consecutive jumps exceeds the baseline (set by the 5% quantile difference value across the entire network) by more than twice are identified as initial fluctuation points. Subsequently, a local extreme point extraction algorithm is used to locate frequent jump intervals, and the variance, range, and root mean square value (RMS) of each interval are extracted. Each event fluctuation interval data entry includes the time of occurrence, duration, maximum jump amplitude, average frequency (number of fluctuations per unit time), and average jump directionality (percentage of increases or decreases). All fields are aggregated to form a dataset of sudden network bandwidth fluctuations. For example, a user recorded four fluctuation intervals during a sudden event at 18:42 on May 12, 2025. The maximum jump amplitude was 740 KB / s, the frequency was 7 fluctuations every 10 seconds, and the average duration was approximately 23 seconds. Based on this fluctuation data, multi-frequency fluctuation peak integral quantification is performed. First, a frequency-domain discrete Fourier transform (DFT) is performed on the bandwidth change rate sequence for each sudden fluctuation interval to extract the main frequency and the first five harmonic frequencies. The corresponding amplitude values ​​of each frequency band in the spectrum are then extracted as energy indicators. After obtaining the frequency-domain features, the energy values ​​of each frequency band are integrated. A spectrum integral curve is constructed with frequency as the horizontal axis and energy as the vertical axis, and the total spectrum integral value is calculated. The integral is calculated by weighting the amplitude values ​​of each frequency band. The weights are set based on the frequency location, with the primary frequency receiving a weight of 1.0 and the secondary frequencies decreasing gradually to 0.2. This reflects the dominant contribution of primary frequency bandwidth jitter to network instability. After integrating the spectral energy for each sudden fluctuation segment, the peak integral of the bandwidth multi-frequency fluctuations is generated and used as a quantitative indicator of the intensity of the multi-frequency instability in that event. For example, for a sudden fluctuation segment, its DFT primary frequency is 0.07 Hz and its amplitude is 2.1 MB / s, resulting in an integral value of 24.6 MB·Hz.

[0091] Based on the above integral values, latency jitter regression analysis is performed on the fluctuation data. First, the network latency data collected synchronously during the period of the sudden fluctuation event is time-aligned. The per-second ping delay records are used as the basic input to construct a delay sequence of the same length as the traffic change rate. Next, the delay sequence is subjected to the same sliding window statistical processing, extracting the maximum delay, minimum delay, average delay, and fluctuation amplitude (the difference between the maximum and minimum) within the window as window-level latency metrics. Then, using the spectrum integral value as the independent variable and the delay fluctuation amplitude as the dependent variable, a least squares linear regression is performed between the two. Metrics such as the regression slope, residual sum of squares, and correlation coefficient are obtained for each sudden event, forming a bandwidth latency jitter regression data table. These steps use an overlapping sliding window technique with a fixed window length of 10 seconds and a step size of 5 seconds to enhance data smoothness and avoid outlier interference. For example, the spectrum integral value of a certain event is 26.3 MB·Hz. The corresponding delay fluctuation jumps from an average of 21 ms to a maximum of 85 ms during the burst segment. The regression slope is 2.35, and the correlation coefficient R² is 0.86.

[0092] Based on the bandwidth delay jitter regression data and the peak integral value of the fluctuation obtained above, a congestion equivalent ratio mapping process is performed. Specifically, the equivalent ratio is defined as the unit delay jitter change caused by the integral of the unit spectrum fluctuation. First, a standard congestion mapping reference set is established. A total of 8,000 typical congestion case samples from the base station level across the entire network are collected. A standard bandwidth congestion curve cluster is constructed, with each curve recording the correspondence between the integral value and the delay slope. Then, for each user incident, the integral value and corresponding delay regression slope are matched to the closest curve template in the standard curve set. Classification is performed using the minimum Euclidean distance matching method, and the matching results are mapped to the congestion equivalent ratio. This ratio is defined as the ratio of the current event to the standard congestion unit value. For example, if a sudden event has an integral value of 20.4 and a regression slope of 1.88, and is matched to congestion level L3 in the standard curve cluster, its equivalent ratio mapping value is 0.73, indicating that the congestion level of the event is 73% of the standard L3 level. The mapped ratio values ​​of all incidents are stored in the bandwidth congestion equivalent ratio mapping data table. Fields include event ID, integral value, regression slope, matching curve number, and equivalent ratio value. Based on the bandwidth multi-frequency fluctuation peak integral, bandwidth delay jitter regression data, and bandwidth congestion equivalent ratio mapping data obtained above, a network bandwidth fluctuation instability assessment mechanism is jointly constructed. First, a multi-evaluation indicator system is defined, including the integral intensity score (normalized by the standard deviation of the integral value distribution), the delay regression significance (weighted by R²), and the congestion equivalent ratio factor (mapped to the score based on its normalized level). These three scores are normalized to the range of 0-1 and then weighted together to produce a comprehensive instability index. The default weights are set as: integral intensity 0.4, regression significance 0.3, and congestion ratio 0.3. Subsequently, a statistical analysis of the instability index of all incidents per user per day is performed, calculating daily maximum, average, and fluctuation coefficient indicators to generate network bandwidth fluctuation instability data. Each instability data record includes: user ID, event ID, timestamp, instability index, instability level (mapped to L1-L4 using 0.3 / 0.6 / 0.8 segments), number of apps involved, primary active time period, access base station number, and network type tag. This ultimately forms a complete network bandwidth fluctuation and instability assessment dataset, providing multi-dimensional dynamic indicator input for subsequent behavioral habit dependency analysis and profiling.

[0093] Step S234 includes the following steps:

[0094] Calculate the time series mean difference of the bandwidth delay jitter regression data to obtain the bandwidth delay jitter time series mean difference;

[0095] The bandwidth congestion factor distribution data is obtained by performing numerical orthogonal projection transformation based on the bandwidth delay jitter time series mean difference and the bandwidth multi-frequency fluctuation peak integral;

[0096] Congestion equal proportion mapping processing is performed according to the bandwidth congestion factor distribution data to obtain bandwidth congestion equal proportion mapping data.

[0097] In an embodiment of the present invention, during the time series mean difference calculation of bandwidth delay jitter regression data, the delay sequence recorded in each burst event is first strictly aligned to ensure that the sampling frequency is uniformly one data point per second. Then, using the start and end times of the bandwidth burst fluctuation as boundaries, equal time series segments are extracted from the original data set, and the time series mean difference calculation is performed on the delay data in each sequence segment. The method used is the fixed window sliding mean difference method, which specifically sets the sliding window length to 10 seconds and the sliding step to 1 second. The average delay is calculated in each window, and then the difference is calculated with the average value of the previous window to generate a mean difference sequence. The mean and standard deviation of all mean difference values ​​are statistically calculated and used as the mean difference indicator of bandwidth delay jitter in the burst event. For each event, the final values ​​obtained are: average time series mean difference, mean difference standard deviation, and maximum jump difference amplitude, which constitute the bandwidth delay jitter time series mean difference data set. For example, an event lasted 80 seconds, resulting in 80 delay samples collected across 71 windows. The mean difference was 12 ms, the standard deviation was 5 ms, and the maximum inter-window hop error was 32 ms. A numerical orthogonal projection transformation was performed based on the bandwidth delay jitter time series mean difference and the obtained bandwidth multi-frequency fluctuation peak integral. First, the bandwidth multi-frequency fluctuation peak integral was normalized to the interval [0, 1] using the maximum and minimum normalization algorithm for all samples. Then, the bandwidth delay jitter time series mean difference was normalized in the same manner. After obtaining these two normalized values, a vector space model was constructed using two variables in a two-dimensional space, representing the X-axis and Y-axis, respectively. The standard orthogonal basis vectors u1 = (1, 0) and u2 = (0, 1) were defined, meaning that the integral value is projected along the X-axis and the mean difference value is projected along the Y-axis. A projection transformation is performed on the vectors of each set of event data. The projection values ​​of the data point in the u1 and u2 directions are calculated using the Euclidean projection formula, thereby obtaining the orthogonal projection distribution points. Aggregate statistics are then performed on all data points, and the coordinate points are divided into five intervals based on the projection value range, corresponding to five categories of congestion potential: low, medium-low, medium, medium-high, and high. This constructs a congestion factor distribution matrix. This matrix is ​​a two-dimensional array structure, with each grid point recording the number of sample points within the corresponding range, the mean position, and the density estimate, forming the bandwidth congestion factor distribution data. Congestion proportional mapping is then performed based on the bandwidth congestion factor distribution data. The specific processing method is to extract and construct a standard congestion factor template table from historical base station-level network congestion events. This template is based on the orthogonal projection results of over 8,000 real-world network congestion events nationwide. Congestion levels L1 to L5 are delineated, and the corresponding projection point distribution center, boundary range, and density weight values ​​are clearly defined for each level. The projection points in the current bandwidth congestion factor distribution data are mapped to the corresponding level area of ​​the standard template according to the nearest neighbor distance criterion. The K-neighbor algorithm (K is set to 5) is used to find the closest historical point group for each point, and the frequency of the level of the point is counted as the mapping result.To further enhance stability, the projected regional density of each event point group is compared with the density in the template to calculate the congestion equivalent scaling factor, which ranges from [0, 1]. This factor represents the ratio of the current event's congestion level to the standard level. For example, if the orthogonal projection coordinates of a sudden event are (0.72, 0.81), the L4 level regional density in the matching historical template is 0.62, and the local density of the current point is 0.49, the mapping scaling factor is 0.79. The resulting congestion equivalent scaling factor for this event is 0.79, corresponding to congestion level L4. The mapping results and equivalent scaling values ​​for all events are aggregated to form a bandwidth congestion equivalent scaling mapping dataset. The dataset includes fields such as event ID, projected coordinates, mapping level, scaling value, corresponding historical matching point group number, and matching accuracy metrics (such as mean Euclidean distance). This data serves as an input to subsequent instability assessments and contributes to comprehensive bandwidth stability assessments.

[0098] Step S24 includes the following steps:

[0099] Step S241: performing amplitude discrete trend fitting processing on the network bandwidth fluctuation instability data to obtain amplitude discrete trend fitting data;

[0100] Step S242: performing weighted reconstruction processing on the traffic usage intensity data of different apps based on the amplitude discrete trend fitting data to obtain traffic usage intensity reconstructed data;

[0101] Step S243: performing clustering iterative optimization on the traffic usage intensity reconstruction data to obtain traffic usage intensity reconstruction cluster data;

[0102] Step S244: reconstructing cluster data according to traffic usage intensity and performing weighted sampling clustering processing to obtain bandwidth instability clustering feature data.

[0103] In this embodiment of the present invention, the amplitude discrete trend fitting process for network bandwidth fluctuation and instability data begins by normalizing the bandwidth fluctuation time series for each sudden event into a vector of length N. N is set to 60, indicating that the bandwidth fluctuations within each event duration are sampled as 60 consecutive points. A sliding polynomial regression-based fitting method is then used to perform trend decomposition on the fluctuation series. Specifically, a sliding window with a window width of 10 is selected, and a third-order polynomial regression is performed on the data within each window to calculate the fitted residual sequence. The standard deviation of the residual sequence within each window is then used as the discrete amplitude indicator for that time point. This process is repeated for the entire series, resulting in a discrete amplitude sequence of length 51. Next, a discrete Fourier transform (DFT) is used to extract the dominant frequency trend from this discrete sequence, recording characteristic parameters such as the amplitude peak change point, dominant frequency oscillation range, and average fluctuation rate. Taking actual collected data as an example, the average standard deviation of the fluctuation series for a user sudden event is 3.2 Mbps, the dominant frequency is 0.15 Hz, the maximum discrete peak is 7.4 Mbps, and the fluctuation duration is 34 seconds. Finally, structured data containing fields such as event identifier, discrete amplitude sequence, fitted dominant frequency, maximum peak location and value, average amplitude, oscillation period, and rate of change is constructed. This data serves as amplitude discrete trend fitting data for subsequent processing. Based on this amplitude discrete trend fitting data, weighted reconstruction of traffic intensity data across different apps is performed. This process first extracts traffic intensity data for apps that overlap with bandwidth fluctuation events from the original intensity dataset. The average rate, peak rate, and fluctuation range for each app during the corresponding time period are then calculated to form a basic indicator triplet (avg, peak, var). The dominant frequency and maximum discrete amplitude in the amplitude discrete trend fitting data are then used as weighting factors, and the data for each app is weighted based on the matching relationship. Specifically, events with higher dominant frequencies and larger discrete amplitudes are assigned higher weights, resulting in amplification or compression of the intensity data for the corresponding apps. The weighting factor is calculated by multiplying the normalized dominant frequency by the normalized maximum discrete amplitude value. The average rate and fluctuation amplitude of the app are then adjusted based on the weighting factor. For example, in an event, the normalized value of the main frequency is 0.83 and the discrete amplitude is 0.75. The weighting factor is 0.6225. If the original app average rate is 0.6 Mbps, the reconstructed rate is adjusted to 0.3735 Mbps. After all app intensity data is processed in this way, traffic usage intensity reconstruction data is generated. The fields include information such as the app ID, the original intensity triplet, the weighting factor, the adjusted triplet, and the event match number.

[0104] Iterative clustering optimization is performed on the traffic usage intensity reconstruction data. This process first performs initial grouping based on the K-means clustering algorithm, setting the number of clusters K to 6. The app reconstruction intensity triplet (avg, peak, var) is used as the three-dimensional feature space. Each app is clustered as a data point, and the initial cluster center is determined based on the sixth quantile of the data distribution. The silhouette coefficient (Silhouette Coefficient) of each clustering round is recorded during the clustering process, and the optimal clustering result is retained. After the initial clustering is completed, local fine-tuning is performed on the data of each cluster center. A high-density cluster point weighted centroid optimization method is used to recalculate the weighted centroid in the local subspace, with the weights determined by the weighting factors of each point. After updating the cluster center, clustering is iterated again until the silhouette coefficient converges or the maximum number of iterations (set to 100) is reached. The clustering process uses Euclidean distance as the distance function. The final output of traffic usage intensity reconstruction cluster data includes fields such as the app ID, cluster number, reconstruction intensity triplet, cluster center coordinates, local density value, and silhouette coefficient.

[0105] Based on the above clustering results, weighted sampling clustering is performed to generate the final bandwidth instability cluster feature data. Specifically, the top 20% of data points with the highest density within each cluster are extracted from the traffic usage intensity reconstructed cluster data as sampling candidates. A weighting coefficient is calculated based on the weighting factor and local density value of each point. Importance sampling is used for sampling, with a sampling ratio of 30% within each cluster. A global clustering optimization based on DBSCAN density clustering is then performed on the sampled points, with a minimum sample number of 5 and a neighborhood radius ε of 0.45. Clustering is performed based on the weighted triple feature space to form the final bandwidth instability cluster feature set. Outliers are excluded from the clustering results, retaining only the central pattern features of high-density areas. The resulting bandwidth instability cluster feature data includes fields such as the instability pattern number, feature center value, representative app distribution ratio, cluster compactness, number of corresponding events, average bandwidth fluctuation level, and the main frequency of event fluctuations. This data structure serves as the input for subsequent behavioral dependency structure analysis and user profiling modeling.

[0106] Step S3 includes the following steps:

[0107] Step S31: mining habit behavior chains on bandwidth instability clustering feature data to obtain instability clustering habit behavior chain data;

[0108] Step S32: performing dependency structure analysis on the unstable clustering habit behavior chain data to obtain the unstable clustering habit behavior dependency structure;

[0109] Step S33: constructing a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile;

[0110] Step S34: Recommend target package types based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommended package type data.

[0111] As an example of the present invention, refer to Figure 3 As shown, in this example, step S3 includes:

[0112] Step S31: mining habit behavior chains on bandwidth instability clustering feature data to obtain instability clustering habit behavior chain data;

[0113] In this embodiment of the present invention, habitual behavior chain mining is performed on bandwidth instability cluster feature data. First, all multiple clustered event sequences associated with the same user ID are extracted from the bandwidth instability cluster feature data. For each user ID, a chronological list of app usage events is constructed, where each event contains fields such as a timestamp, app name, bandwidth usage intensity characteristic value, instability cluster ID, and event duration. Based on this event sequence, a fixed-length sliding window approach is used to segment the event sequence, with a window size of 5 and a step size of 1, to generate a set of sliding subsequences. Frequent pattern mining is then performed on the app combination patterns within each subsequence, using the Apriori algorithm to calculate support, with a minimum support threshold of 0.15 and a minimum confidence level of 0.7. When these thresholds are met, high-frequency app combination paths are extracted, and the temporal position and instability status of each app behavior are labeled. For a specific user, the extracted high-frequency chain is: "short video → game → instant messaging → video conferencing → map navigation," with each step corresponding to metrics such as average bandwidth and the dominant frequency of instability. The final output of the unstable clustering habit behavior chain data includes user ID, chain number, chain APP sequence, cluster number of each APP, chain frequency, average occurrence time period, average instability index, etc.

[0114] Step S32: performing dependency structure analysis on the unstable clustering habit behavior chain data to obtain the unstable clustering habit behavior dependency structure;

[0115] In one embodiment of the present invention, dependency structure analysis is performed on unstable clustered habit behavior chain data. The process first constructs a weighted directed graph based on all chain data. Each node in the graph represents an app usage behavior, and each edge represents a directed temporal dependency relationship from the previous app behavior to the next app behavior. The edge weight is the sum of the frequencies of occurrence of that path across all chains. After the graph is constructed, a structured dependency mining operation is performed. First, using the maximum weight path extraction method, all paths are ranked by path score. Paths whose summed weights exceed a threshold (e.g., 85% of the total path frequency) are selected as candidate dependency chains. The graph is then subgraphed and the Girvan-Newman algorithm is used to segment communities based on edge betweenness, yielding multiple highly cohesive behavioral dependency substructures. Subsequently, the in-degree and out-degree distributions of nodes in each subgraph are analyzed to identify core nodes and key jump paths. For example, in a certain user group, the "short video" node has an out-degree of 4, connecting "games," "social networking," "e-commerce," and "live streaming," making it a high-frequency dependency initiation behavior. Finally, an unstable clustering habit behavior dependency structure is formed, and the output fields include user group ID, dependency path ID, path composition APP sequence, path cumulative frequency, dependency subgraph number, subgraph node set, subgraph edge set, node importance score, etc.

[0116] Step S33: constructing a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile;

[0117] In an embodiment of the present invention, a user usage behavior portrait is constructed based on the unstable clustering habitual behavior dependency structure. In the specific operation, the habitual behavior dependency structure of each user is first vectorized to establish a vector space behavior feature model. The behavior path sequence is used as the basic vector construction unit, and each dimension represents the frequency of occurrence of a standard path pattern. After normalization, a behavior vector is formed. At the same time, the average bandwidth usage intensity, unstable cluster number, and APP usage time density of the user in each APP are encoded into the extended behavior vector dimension to form a complete multi-dimensional behavior portrait feature vector. For different users, the Cosine similarity is used to calculate the similarity score with the standard behavior path vector set, and this is used as the basis for portrait label matching. Finally, each user portrait contains the following fields: user ID, main behavior chain label, behavior vector code, cluster preference vector, high-frequency usage time period, behavior dependency main path number, unstable mode indicator flag, behavior graph structure entropy value, etc.

[0118] Step S34: Recommend target package types based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommended package type data.

[0119] In this embodiment of the present invention, target package types are recommended based on the user behavior profile and bandwidth instability clustering feature data. The recommendation process consists of two stages: matching user behavior features with package templates and ranking recommendations. First, a package template library is established. The templates contain fields such as package number, data traffic limit, time-of-day priority policy, targeted app preferential treatment type, delay sensitivity level, and bandwidth guarantee threshold. Next, the user profile vector is matched with the package templates. Feature comparison is used to cross-check the frequently used app categories, fluctuation characteristics, and high-frequency nodes in the instability dependency path with the package policy fields. For example, if the user profile frequently uses live streaming and short videos, and high-frequency burst cluster nodes appear in the instability path, package templates with bandwidth guarantees and targeted video optimization policies are prioritized. Next, all candidate packages are ranked based on the matching score between the user profile and the package template. The score calculation takes into account the degree of behavioral overlap, the estimated traffic adaptation rate, and the cluster matching weight. Finally, the top N (e.g., N = 3) target recommended package types are output. Recommendation result fields include user ID, recommended package number, matching score, key profile-package matching fields, and estimated bandwidth guarantee coefficient. The entire process does not involve machine learning models, but instead achieves deterministic recommendations by comparing the structured features of user profiles with policy rules.

[0120] Step S33 includes the following steps:

[0121] Step S331: performing normalization processing on the unstable clustering behavior frequency of the unstable clustering habit behavior dependency structure to obtain normalized unstable clustering behavior frequency data;

[0122] Step S332: performing logic back-inference merging processing on the normalized unstable clustering behavior frequency data to obtain a behavior combination deduction structure;

[0123] Step S333: performing behavior coupling degree division processing on the behavior combination deduction structure to obtain a set of multiple behavior coupling patterns of the user;

[0124] Step S334: construct a user usage behavior profile based on the user's multiple behavior coupling pattern set to obtain the user usage behavior profile.

[0125] In the embodiment of the present invention, the unstable clustering habit behavior dependency structure is subjected to unstable clustering behavior frequency normalization processing. First, all behavior nodes and their corresponding frequency counts are extracted from the unstable clustering habit behavior dependency structure corresponding to each user. The frequency count is derived from the actual number of occurrences of a certain behavior path or node by the user within a specific time window (such as the past 30 days). The behavior node "short video" that appears most frequently in the dependency structure of user U is used as a reference benchmark, and its frequency is set to the maximum value. , and then the frequency values ​​of all behavior nodes Perform normalization operations to calculate the normalized frequency of each behavior node . During the implementation process, in order to avoid the loss of information in subsequent analysis due to the low weight of low-frequency behavior nodes, low-frequency smoothing processing is performed on behavior nodes with a frequency lower than the global frequency median, and the normalized value is adjusted to no less than 30% of the global normalized frequency mean. Finally, a normalized unstable cluster behavior frequency data table is formed. The table structure includes the behavior node ID, original frequency, normalized frequency value, dependency chain number, unstable cluster number, etc. The normalized unstable cluster behavior frequency data is logically back-pushed and merged. First, the behavior continuity score is calculated based on the normalized frequency difference between the previous and next nodes in the behavior dependency path. The difference threshold θ is set (such as θ = 0.2). If the normalized frequency difference between adjacent nodes is less than the threshold, it is determined to have strong logical coherence and merged into a behavior combination fragment. In the specific implementation, a path backtracking-based merging strategy is adopted, which gradually back-pushes from the end behavior node to construct all combined behavior path groups. For each merged path segment, we calculated its logical similarity index, including the difference in behavior frequency, the overlap rate of time periods, and the association weight of behavior types. We used a linear weighted combination to calculate the path combination score, with a minimum combination score threshold of 0.7. All path segments with scores exceeding the threshold were retained as behavioral combination deduction structural units. Each unit recorded the path ID, the sequence of constituent behaviors, the logical back-link number, the combination strength score, and the average instability cluster label.

[0126] The behavior combination deduction structure is partitioned by behavioral coupling. A behavioral coupling calculation method is used to analyze three metrics: concurrent usage frequency, temporal overlap, and unstable state co-occurrence rate between behaviors in the combination path. The coupling calculation uses a three-factor normalized weighted approach, with weights set to 0.4 for concurrent usage frequency, 0.3 for temporal overlap, and 0.3 for unstable state co-occurrence rate. In practice, for each behavior combination deduction structure unit, all original instance fragments from the user behavior log are extracted. The number of concurrent occurrences of multiple behaviors within a 10-minute window, the timestamp overlap ratio, and the proportion of instances simultaneously marked as unstable states during burst bandwidth fluctuations are counted. Based on this, a coupling score is calculated for each pair of behaviors. A hierarchical aggregation method is then used to partition coupled behaviors into groups, with the intra-group coupling criterion being greater than 0.6. The resulting output is a set of user multi-behavior coupling patterns, consisting of a coupling pattern ID, a list of included behavior IDs, a coupling matrix, temporal synchronization weights, and an unstable state co-occurrence factor. User behavior profiles are constructed based on the set of multi-behavior coupling patterns, generating a multidimensional feature vector based on the coupling pattern. Specifically, metrics such as the behavior type composition, average behavior initiation time, behavior sequence stability index, mean instability intensity, and average bandwidth usage within each coupling pattern are vectorized and concatenated into a complete profile code. For example, user X's coupling pattern is "video playback + social commenting," with a sequence stability of 0.8, an average concurrent duration of 180 seconds, an average bandwidth usage of 2.5 MB / s, and an instability co-occurrence rate of 0.45. This information is encoded as part of the behavior profile vector. Furthermore, to characterize the overall profile, all coupling patterns are aggregated and analyzed to form the overall behavioral feature center of the user profile. This includes fields such as user ID, primary behavior coupling type (e.g., entertainment-social interaction), average behavior cycle length, multi-behavior synchronization rate, behavior instability sensitivity level, and typical bandwidth pressure index. The resulting user behavior profile is a structured feature set that can be used to support subsequent package recommendations and network resource allocation strategies.

[0127] The present invention also provides a user profile construction system based on user habit analysis, which is used to execute the user profile construction method based on user habit analysis as described above. The user profile construction system based on user habit analysis includes:

[0128] The traffic usage fine-grained analysis module is used to obtain the historical usage data of the user's mobile phone traffic package; the traffic usage fine-grained analysis is performed on the historical usage data of the user's mobile phone traffic package to obtain fine-grained data on the traffic usage status of different apps;

[0129] The bandwidth instability cluster analysis module is used to analyze the sudden network occupancy status of multiple apps in simultaneous use based on the fine-grained data of traffic usage status between different apps, and obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated based on the sudden network occupancy status, and the bandwidth instability clustering feature data is obtained;

[0130] A behavior profile construction module is used to perform dependency structure analysis on bandwidth instability clustering feature data to obtain an unstable clustering habit behavior dependency structure; a user usage behavior profile is constructed based on the unstable clustering habit behavior dependency structure to obtain a user usage behavior profile; and target recommendation package type data is recommended based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommendation package type data.

[0131] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a user portrait based on user habit analysis, characterized in that: The following steps are involved: Step S1: Obtain historical usage data of the user's mobile phone data package; perform fine-grained traffic usage analysis on the historical usage data of the user's mobile phone data package to obtain fine-grained data on traffic usage status between different apps; Step S2 includes: Step S21: performing traffic usage intensity analysis between different APPs on the fine-grained data of traffic usage status between different APPs to obtain traffic usage intensity data between different APPs; Step S22: analyzing the sudden network occupancy status of multiple apps in synchronous use based on the traffic usage intensity data between different apps to obtain the sudden network occupancy status; Step S23: performing a network bandwidth fluctuation instability assessment on the sudden network occupancy state to obtain network bandwidth fluctuation instability data; wherein step S23 includes: Step S231: performing sudden network bandwidth fluctuation analysis on the sudden network occupancy state to obtain sudden network bandwidth fluctuation data; Step S232: quantifying the multi-frequency fluctuation peak integral of the sudden network bandwidth fluctuation data to obtain the bandwidth multi-frequency fluctuation peak integral; Step S233: performing delay jitter regression analysis on the sudden network bandwidth fluctuation data according to the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth delay jitter regression data; Step S234: performing congestion equal-amount proportional mapping based on the bandwidth delay jitter regression data and the bandwidth multi-frequency fluctuation peak integral to obtain bandwidth congestion equal-amount proportional mapping data; wherein step S234 is specifically as follows: Calculate the time series mean difference of the bandwidth delay jitter regression data to obtain the bandwidth delay jitter time series mean difference; The bandwidth congestion factor distribution data is obtained by performing numerical orthogonal projection transformation based on the bandwidth delay jitter time series mean difference and the bandwidth multi-frequency fluctuation peak integral; Performing congestion equal proportion mapping processing according to bandwidth congestion factor distribution data to obtain bandwidth congestion equal proportion mapping data; Step S235: Perform network bandwidth fluctuation instability assessment based on bandwidth multi-frequency fluctuation peak integral, bandwidth delay jitter regression data, and bandwidth congestion equal-amount proportional mapping data to obtain network bandwidth fluctuation instability data; Step S24: performing weighted sampling and clustering processing on the traffic usage intensity data between different apps based on the network bandwidth fluctuation and instability data to obtain bandwidth instability clustering feature data; Step S3: Perform dependency structure analysis on the bandwidth instability clustering feature data to obtain the unstable clustering habit behavior dependency structure; construct a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile; recommend a target recommended package type based on the user usage behavior profile and the bandwidth instability clustering feature data to obtain target recommended package type data.

2. The method for constructing a user profile based on user habit analysis according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Obtaining historical usage data of the user's mobile phone data package; Step S12: Clean the historical usage data of the user's mobile phone data package to obtain the cleaned data of the historical usage of the data package; Step S13: extracting the traffic usage status between different APPs from the traffic package history usage cleaning data to obtain the traffic usage status between different APPs; Step S14: Perform fine-grained traffic usage analysis on the traffic usage status between different APPs to obtain fine-grained data on the traffic usage status between different APPs.

3. The method for constructing a user profile based on user habit analysis according to claim 1, characterized in that: Step S24 includes the following steps: Step S241: performing amplitude discrete trend fitting processing on the network bandwidth fluctuation instability data to obtain amplitude discrete trend fitting data; Step S242: performing weighted reconstruction processing on the traffic usage intensity data of different apps based on the amplitude discrete trend fitting data to obtain traffic usage intensity reconstructed data; Step S243: performing clustering iterative optimization on the traffic usage intensity reconstruction data to obtain traffic usage intensity reconstruction cluster data; Step S244: reconstructing cluster data according to traffic usage intensity and performing weighted sampling clustering processing to obtain bandwidth instability clustering feature data.

4. The method for constructing a user profile based on user habit analysis according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: mining habit behavior chains on bandwidth instability clustering feature data to obtain instability clustering habit behavior chain data; Step S32: performing dependency structure analysis on the unstable clustering habit behavior chain data to obtain the unstable clustering habit behavior dependency structure; Step S33: constructing a user usage behavior profile based on the unstable clustering habit behavior dependency structure to obtain the user usage behavior profile; Step S34: Recommend target package types based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommended package type data.

5. The method for constructing a user profile based on user habit analysis according to claim 4, characterized in that: Step S33 includes the following steps: Step S331: performing normalization processing on the unstable clustering behavior frequency of the unstable clustering habit behavior dependency structure to obtain normalized unstable clustering behavior frequency data; Step S332: performing logic back-inference merging processing on the normalized unstable clustering behavior frequency data to obtain a behavior combination deduction structure; Step S333: performing behavior coupling degree division processing on the behavior combination deduction structure to obtain a set of multiple behavior coupling patterns of the user; Step S334: construct a user usage behavior profile based on the user's multiple behavior coupling pattern set to obtain the user usage behavior profile.

6. A user portrait construction system based on user habit analysis, characterized in that: The method for constructing a user profile based on user habit analysis according to claim 1 is configured to execute the method. The system for constructing a user profile based on user habit analysis comprises: The traffic usage fine-grained analysis module is used to obtain the historical usage data of the user's mobile phone traffic package; the traffic usage fine-grained analysis is performed on the historical usage data of the user's mobile phone traffic package to obtain fine-grained data on the traffic usage status of different apps; The bandwidth instability cluster analysis module is used to analyze the sudden network occupancy status of multiple apps in simultaneous use based on the fine-grained data of traffic usage status between different apps, and obtain the sudden network occupancy status; the network bandwidth fluctuation instability is evaluated based on the sudden network occupancy status, and the bandwidth instability clustering feature data is obtained; A behavior profile construction module is used to perform dependency structure analysis on bandwidth instability clustering feature data to obtain an unstable clustering habit behavior dependency structure; a user usage behavior profile is constructed based on the unstable clustering habit behavior dependency structure to obtain a user usage behavior profile; and target recommendation package type data is recommended based on the user usage behavior profile and bandwidth instability clustering feature data to obtain target recommendation package type data.

Citation Information

Patent Citations

  • SIM package recommendation method and system based on user behavior analysis

    CN118890610A

  • Broadband marketing strategy making method and device and readable storage medium

    CN119398826A

  • Intelligent router traffic management method for home scene

    CN119728521A