An e-sim-based user behavior analysis method

By unifying the time-series alignment and labeling of E-SIM configuration switching logs and user behavior data, and combining factor decomposition and graph matching techniques, the problem of spurious fluctuations in E-SIM user behavior analysis was solved, and accurate user behavior modeling was achieved in complex network environments.

CN121037829BActive Publication Date: 2026-05-22GUANGDONG LEGEND COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG LEGEND COMM CO LTD
Filing Date
2025-09-24
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing user behavior analysis methods fail to effectively distinguish between genuine behavioral changes and spurious fluctuations caused by E-SIM configuration switching when processing E-SIM user behavior data that involves frequent network switching, leading to biased analysis results.

Method used

By collecting E-SIM configuration switching logs, network policy parameters, and user behavior event sequences, time alignment and labeling are performed to construct a configuration switching trigger window. The correlation between network migration and behavioral events is analyzed. Factor decomposition and causal testing are used to evaluate device-side configuration change factors, and graph matching and topology consistency verification are combined to evaluate network-side environmental change factors. Pseudo-migration scores are generated to determine identity perception drift and correct user behavior.

Benefits of technology

It effectively eliminates analytical biases caused by time inconsistencies in multi-source data, accurately locates changes in user behavior, accurately eliminates drift interference, and improves the accuracy and stability of user behavior modeling. It is suitable for complex network environments such as cross-carrier and international roaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037829B_ABST
    Figure CN121037829B_ABST
Patent Text Reader

Abstract

The application discloses a user behavior analysis method based on E-SIM, and particularly relates to the technical field of user behavior analysis; the method comprises the following steps: collecting E-SIM configuration switching logs, network policy parameters and user behavior event sequences, and constructing a unified time sequence dataset; determining a configuration switching trigger window based on the unified time sequence dataset, analyzing the correlation degree between network environment migration and user behavior change, and obtaining a candidate drift segment; respectively performing factor deconstruction on the device side and the network side, and outputting device side and network side pseudo migration scores; comprehensively judging identity recognition drift based on the pseudo migration scores, calculating a behavior correction vector; reconstructing the user behavior sequence by using the behavior correction vector, correcting an abnormal behavior label, and outputting a corrected user behavior analysis result. The accuracy and reliability of user behavior analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of user behavior analysis technology, and more specifically, to a user behavior analysis method based on E-SIM. Background Technology

[0002] With the continuous development of communication technology and mobile terminals, embedded user identification modules (E-SIM) are widely used in international roaming and cross-carrier service scenarios.

[0003] Existing user behavior analysis methods, when processing E-SIM user behavior data with frequent network switching, typically treat behavioral changes as genuine shifts in user behavior, ignoring the spurious fluctuations in user behavior data caused by implicit changes in device status and network environment parameters due to E-SIM configuration switching, leading to biased user behavior analysis results. Summary of the Invention

[0004] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a user behavior analysis method based on E-SIM to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A user behavior analysis method based on E-SIM includes the following steps:

[0007] 1. A user behavior analysis method based on E-SIM, characterized by comprising the following steps:

[0008] S1: Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences, and perform time alignment and labeling to output a unified time-series dataset;

[0009] S2: Construct a configuration switching trigger window based on a unified time series dataset, analyze the correlation between network migration and behavioral events, and output candidate drift segments;

[0010] S3: Deconstruct the candidate drift segments using equipment-side factors, evaluate the impact of equipment-side configuration change factors using factor decomposition and causal tests, and output the equipment-side pseudo-migration score.

[0011] S4: Deconstruct the network-side factors of the candidate drift segments, use graph matching and topology consistency verification to evaluate the impact of network-side environmental change factors, and output the network-side pseudo-migration score.

[0012] S5: Based on the pseudo-transfer scores on the device side and the network side, determine identity perception drift and output a behavior correction vector;

[0013] S6: Reconstruct user behavior sequences based on behavior correction vectors, update user behavior features, correct abnormal behavior labels, and output corrected user behavior analysis results.

[0014] In a preferred embodiment, S1 specifically refers to:

[0015] Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences;

[0016] Synchronize and align E-SIM configuration handover logs, network policy parameters, and user behavior event sequences according to a unified time base.

[0017] The synchronized E-SIM configuration handover logs, network policy parameters, and user behavior event sequences are classified and labeled to generate a unified time-series dataset.

[0018] In a preferred embodiment, S2 specifically refers to:

[0019] Construct a configuration switching trigger window centered on the configuration switching timestamp in the E-SIM configuration switching log and with a duration of a preset time window;

[0020] Extract network policy parameters and user behavior event sequences that fall into the configuration switching trigger window from the unified time series dataset, and combine the extracted results into window data blocks;

[0021] Based on window data blocks, calculate the temporal cross-correlation coefficient between the network strategy parameter change vector and the user behavior event sequence change vector to generate a network migration and behavior event correlation index.

[0022] Threshold determination is performed on the correlation indicators between network migration and behavioral events, and time series segments that meet the threshold conditions are selected to output candidate drift segments.

[0023] In a preferred embodiment, S3 specifically refers to:

[0024] Extract device-side configuration data from candidate drift segments;

[0025] The device-side configuration data is combined with the user behavior change data corresponding to the candidate drift segments to construct a feature matrix;

[0026] Factor decomposition of the feature matrix yields multiple independent influencing factors;

[0027] A causal test was performed on each independent influencing factor and the user behavior change data, and the causal contribution of the independent influencing factors was calculated.

[0028] The pseudo-migration score on the device side is calculated by weighting the independent impact factors based on causal contribution.

[0029] In a preferred embodiment, S4 specifically refers to:

[0030] Extract network-side configuration data from candidate drift segments;

[0031] Construct a network environment topology diagram based on network-side configuration data;

[0032] Graph matching is performed on the network environment topology map. By calculating the similarity of the topology structure before and after the network configuration switch in the candidate drift segment, the topology structure matching index is obtained.

[0033] Topology consistency verification is performed based on the network environment topology diagram. By calculating the degree of change in network-side configuration data, topology consistency indicators are obtained.

[0034] The network-side pseudo-migration score is calculated based on the topology matching index and the topology consistency index.

[0035] In a preferred embodiment, S5 specifically refers to:

[0036] The device-side pseudo-migration scores and network-side pseudo-migration scores are normalized to obtain normalized device-side pseudo-migration scores and normalized network-side pseudo-migration scores.

[0037] An identity perception drift determination matrix is ​​constructed based on normalized device-side pseudo-transfer scores and normalized network-side pseudo-transfer scores.

[0038] Based on the identity perception drift determination matrix, a comprehensive score for identity perception drift is calculated;

[0039] The overall score of identity perception drift is compared with the preset drift judgment threshold, and the drift confidence score is output.

[0040] Calculate the behavior correction vector based on the drift confidence.

[0041] In a preferred embodiment, S6 specifically refers to:

[0042] The behavior correction vector is used to perform vector operations with the user behavior change data corresponding to the candidate drift segment to generate corrected behavior data.

[0043] Replace the corresponding user behavior event sequence in the unified time series dataset with the corrected behavior data to obtain the corrected user behavior sequence;

[0044] Based on the corrected user behavior sequence, behavioral statistical features, temporal features and interaction features are extracted to generate an updated user behavior feature set;

[0045] The abnormal behavior label identifier is reset based on the updated user behavior feature set to form an abnormal label update table;

[0046] The corrected user behavior sequence, updated user behavior feature set, and anomaly label update table are encapsulated into the corrected user behavior analysis results.

[0047] The technical effects and advantages of the user behavior analysis method based on E-SIM of this invention are as follows:

[0048] By unifying and labeling the time sequence of E-SIM configuration switching logs, network policy parameters, and user behavior event sequences, analytical biases caused by time inconsistencies in multi-source data are effectively eliminated. Constructing a configuration switching trigger window and extracting candidate drift segments allows for precise location of key segments in user behavior changes before and after network switching, enabling rapid identification of drift phenomena. Factor decomposition and causal testing quantitatively assess the impact of device-side configuration change factors, generating device-side pseudo-migration scores to distinguish behavioral changes caused by device state reconstruction. Graph matching and topology consistency checks are used to quantitatively evaluate network-side environmental change factors, outputting network-side pseudo-migration scores to effectively reveal the pseudo-impact of network routing, access points, and resolution services changes on behavioral data. Based on pseudo-migration scores, identity perception drift is comprehensively determined, and behavioral correction vectors are calculated to achieve refined correction of misjudgment biases. The user behavior sequence is reconstructed using correction vectors, and feature labels are updated. Corrected user behavior analysis results obtained after accurately eliminating drift interference improve the accuracy and stability of user behavior modeling, making it suitable for intelligent terminal data processing scenarios in complex network environments such as cross-carrier and international roaming. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of a user behavior analysis method based on E-SIM according to the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0051] Example

[0052] Figure 1 This invention presents a user behavior analysis method based on E-SIM, which includes the following steps:

[0053] 1. A user behavior analysis method based on E-SIM, characterized by comprising the following steps:

[0054] S1: Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences, and perform time alignment and labeling to output a unified time-series dataset;

[0055] S2: Construct a configuration switching trigger window based on a unified time series dataset, analyze the correlation between network migration and behavioral events, and output candidate drift segments;

[0056] S3: Deconstruct the candidate drift segments using equipment-side factors, evaluate the impact of equipment-side configuration change factors using factor decomposition and causal tests, and output the equipment-side pseudo-migration score.

[0057] S4: Deconstruct the network-side factors of the candidate drift segments, use graph matching and topology consistency verification to evaluate the impact of network-side environmental change factors, and output the network-side pseudo-migration score.

[0058] S5: Based on the pseudo-transfer scores on the device side and the network side, determine identity perception drift and output a behavior correction vector;

[0059] S6: Reconstruct user behavior sequences based on behavior correction vectors, update user behavior features, correct abnormal behavior labels, and output corrected user behavior analysis results.

[0060] S1: Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences, perform time alignment and labeling, and output a unified time-series dataset, including:

[0061] Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences;

[0062] An embedded subscriber identity module (E-SIM) is an electronic subscriber identity module built into a user terminal device. E-SIMs offer remote configuration and management capabilities, allowing users to remotely select or change different operator network services as needed without physically replacing the subscriber identity module. The embedded user identification module (UIM) configuration switching log is a collection of data recording each change in the identity information of the UIM on the user's terminal device. This includes the timestamp of each configuration switching event, the operator configuration information used before each event, the operator configuration information used after each event, and the activation status data of each UIM identity information configuration. For example, if a user's smartphone needs to switch operator configurations in the UIM from operator A to operator B while roaming internationally, the UIM configuration switching log will record the timestamp of the configuration switch, the access point name (APN) information of the original operator A, the network routing information of operator A, and the domain name resolution server information of operator A. It will also record the APN information, network routing information, and domain name resolution server information of the new operator B, as well as the activation status data (e.g., the configuration activation flag is "activated").

[0063] Network policy parameters are communication configuration parameters used by user equipment when connecting to the network, including access point name information, network routing information, and domain name resolution server information. Among them, access point name information is the access parameter specified by the device when connecting to the mobile operator's network; network routing information indicates the address of the network path used by the user equipment when transmitting network data packets; and domain name resolution server information is the resolution server address used by the user equipment to convert domain names to IP addresses when accessing Internet applications.

[0064] User behavior event sequences are a series of user interaction behavior operation records generated when a user accesses Internet applications through a terminal device. This includes all interaction operation data records that occur when a user accesses a specific online application. For example, when a user opens an e-commerce application on a mobile phone, browses a product page, clicks a promotional advertisement button on the page, and performs an online payment operation, user behavior data is generated respectively. For example, browsing record data of product pages (such as product number, dwell time), advertisement click event data (such as the type of advertisement clicked, click timestamp), and payment operation data (such as payment method, payment amount, and payment occurrence timestamp).

[0065] Synchronize and align E-SIM configuration handover logs, network policy parameters, and user behavior event sequences according to a unified time base.

[0066] By selecting a unified timestamp benchmark, such as Beijing time, and assuming that the configuration switch timestamp of the embedded user identification module configuration switch log is 13:45:20 on May 10, 2025, Beijing time, all recorded data of network policy parameters and user behavior event sequences will be aligned according to the unified time benchmark, so that data from different sources have a unified time axis feature.

[0067] The synchronized and aligned E-SIM configuration handover logs, network policy parameters, and user behavior event sequences are classified and labeled to generate a unified time-series dataset.

[0068] Classification and labeling refers to tagging data according to its inherent attributes. For example, configuration switch log data is labeled as configuration switch type data, network policy parameter data is labeled as network environment configuration type data, and user behavior event sequence data is labeled as user interaction event type data. For instance, operator configuration data in the configuration switch log is labeled as operator configuration change type, network routing information and domain name resolution server information are labeled as network environment parameter type, and browsing, clicking, and payment interaction events are labeled as user behavior event types. After classification and labeling, the E-SIM configuration switch logs, network policy parameters, and user behavior event sequences are integrated to generate a unified time-series dataset, including operator configuration change type data, network environment parameter type data, and user behavior event type data.

[0069] S2: Construct a configuration switching trigger window based on a unified time-series dataset, analyze the correlation between network migration and behavioral events, and output candidate drift segments, including:

[0070] Construct a configuration switching trigger window centered on the configuration switching timestamp in the E-SIM configuration switching log and with a duration of a preset time window;

[0071] The configuration switching trigger window is a specific time range centered on the configuration switching timestamp corresponding to each configuration switching event recorded in the configuration switching log of the embedded user identification module, and extending simultaneously in both directions before and after the configuration switching timestamp for a duration of a preset time window; the configuration switching trigger window is used to determine network policy parameter data and user behavior event sequence data that have a temporal relationship with the configuration switching event of the embedded user identification module. For example, if the timestamp of the configuration switch event recorded in the configuration switch log of the embedded user identification module is 13:45:20 Beijing time on May 10, 2025, and the preset time window is set to 10 minutes, then the constructed configuration switch trigger window time range is from 13:40:20 Beijing time on May 10, 2025 to 13:50:20 Beijing time on May 10, 2025, lasting for a total of 10 minutes. Each configuration switch event corresponds to a separate configuration switch trigger window, so that each configuration switch log event of the embedded user identification module can correspond to a determined window time range, ensuring that the extraction process of network policy parameter data and user behavior event sequence data has a time domain boundary, thereby realizing the effectiveness of data association.

[0072] Extract network policy parameters and user behavior event sequences that fall into the configuration switching trigger window from the unified time series dataset, and combine the extracted results into window data blocks;

[0073] After constructing the configuration switching trigger window, network policy parameter data and user behavior event sequence data falling within the constructed configuration switching trigger window are extracted based on a unified time-series dataset. Using the start and end times of the constructed configuration switching trigger window as a benchmark, all network policy parameter data and user behavior event sequence data whose timestamps fall within the window range are selected from the unified time-series dataset and integrated into a window data block. The window data block is a data set affected by the configuration switching event time of the embedded user identification module, including network routing parameter data, access point name parameter data, domain name resolution server parameter data, and user interaction behavior data within the window time range. For example, all data records of a user browsing a product page, clicking an ad button, and completing a payment operation; for example, within the window time range (i.e., from 13:40:20 to 13:50:20 Beijing time on May 10, 2025), if a user browses a specific e-commerce product page at 13:43:00 Beijing time, clicks a promotional ad button on the e-commerce product page at 13:45:30, and makes an online payment at 13:47:50 Beijing time, then the above user behavior data, together with the network routing parameter data, access point name parameter data, and domain name resolution server parameter data corresponding to the same window range, together form a window data block.

[0074] Based on window data blocks, calculate the temporal cross-correlation coefficient between the network strategy parameter change vector and the user behavior event sequence change vector to generate a network migration and behavior event correlation index.

[0075] Based on the obtained window data blocks, the temporal cross-correlation coefficient between the network policy parameter change vector and the user behavior event sequence change vector is calculated to generate a network migration and behavior event correlation index. The network policy parameter change vector is a data vector used to describe the trend of network environment parameters changing over time within the configuration switchover trigger window. For example, if the network routing parameter changes from the initial value 192.168.10.1 to 172.16.20.1 within the configuration switchover trigger window, then the data element in the network policy parameter change vector corresponding to the network routing parameter is recorded as changing from 192.168.10.1 to 172.16.20.The change information for 1; the user behavior event sequence change vector represents the changes in user interaction behavior data over time within the same configuration switch trigger window. For example, it represents the data vector showing the changes in a user's dwell time on a product page, the number of times an ad button is clicked, or the number of payment operations over time. For instance, if a user viewed the page 3 times in the 5 minutes before the configuration switch event and then viewed it 6 times in the 5 minutes after the event, the user behavior event sequence change vector would contain the trend data from 3 to 6 views. By comparing the values ​​of the network policy parameter change vector and the user behavior event sequence change vector within the time range of the configuration switch trigger window... Calculations were performed, and the correlation between the two vectors was quantitatively evaluated using a standardized time cross-correlation coefficient formula. The calculation process of the time cross-correlation coefficient formula is as follows: The network policy parameter change vector and the user behavior event sequence change vector are defined as two discrete time series, denoted as X and Y, respectively. The elements of the network policy parameter change vector X represent the numerical changes in the network policy parameters relative to their initial state at each time point within the configuration switching trigger window, and the elements of the user behavior event sequence change vector Y represent the number of user behavior events within the configuration switching trigger window. Based on the numerical changes relative to the initial state; calculate the average values ​​of the network policy parameter change vector X and the user behavior event sequence change vector Y within the configuration switching trigger window; calculate the differences between each element in vector X and vector Y and their own average values, generating two new difference vectors, denoted as the difference sequence of vector X and the difference sequence of vector Y; multiply the corresponding elements of the difference sequence of vector X and the difference sequence of vector Y and sum them to obtain the accumulated covariance; calculate the standard deviation of each difference sequence of vector X and the difference sequence of vector Y; sum the squares of each element in the difference sequence of Y, divide by the total number of elements, and then take the square root to obtain the difference sequence of vector Y. The standard deviation is calculated; finally, the calculated cumulative covariance is divided by the product of the standard deviations of the vector X difference sequence and the vector Y difference sequence to obtain the time cross-correlation coefficient between the network policy parameter change vector and the user behavior event sequence change vector. This coefficient is defined as the network migration and behavior event correlation index, used to represent the degree and strength of the time correlation between changes in network policy parameters and changes in user behavior events. The network migration and behavior event correlation index is used to quantify the correlation strength and direction between changes in network policy parameters and changes in user behavior events caused by the configuration switching of the embedded user identification module. The closer to 1, the higher the correlation strength; the closer to 0, the weaker the correlation strength.

[0076] Threshold determination is performed on the correlation indicators between network migration and behavioral events, time series segments that meet the threshold conditions are selected, and candidate drift segments are output.

[0077] Based on the correlation index between network migration and behavioral events, a threshold determination is performed to identify time segments that meet the threshold conditions. The threshold conditions are preset, for example, the threshold condition is 0.8, then all time segments with a correlation index between network migration and behavioral events greater than or equal to 0.8 will be selected. For example, if the network migration and behavioral event correlation index calculated for the configuration switch trigger window is 0.85, then the data segment in the configuration switch trigger window meets the threshold condition and becomes a candidate drift segment.

[0078] S3: Deconstruct the candidate drift segments using equipment-side factors, employ factor decomposition and causal testing to assess the impact of equipment-side configuration change factors, and output equipment-side pseudo-migration scores, including:

[0079] Extract device-side configuration data from candidate drift segments;

[0080] Device-side configuration data refers to the data set formed by changes in the internal configuration or state of the user terminal device when or after a configuration switching event occurs. This includes location cache state data, privacy policy version data, and application process restart state data. Among them, location cache state data refers to the device location information cache data stored in the user terminal device, recording the geographical coordinate data cached by the user terminal device under different network operator environments, such as latitude and longitude, base station location information, and Wi-Fi hotspot location information. For example, when a user's smartphone switches from operator A to operator B, the location cache state data will record the location coordinates cached by operator A as 31.22 degrees north latitude and 121.48 degrees east longitude. After the network environment of operator B is activated, the location cache state data will be updated to 31.20 degrees north latitude and 121.47 degrees east longitude. Privacy policy version data refers to the version number and configuration data of the privacy permission policy settings adopted by the user's terminal device, including the permission version number, permission configuration list, and permission status parameters. For example, if the privacy policy version number of a user's smartphone is "V1.0" under the original operator A network environment, after switching to the operator B network environment, the privacy policy version will be automatically or manually updated to "V1.2". At the same time, the changes in the permission list of the privacy policy will be recorded, such as the location permission changing from allowed to prohibited, and the microphone permission changing from prohibited to allowed. Application process restart status data represents the status information of the application after it is restarted by the terminal device's system or user operation during the configuration switching process. This includes the application process identifier, the timestamp of the application process restart, and the restart reason information. For example, when the user's terminal device switches from operator A to operator B, the e-commerce application is forcibly restarted during the switching process due to the change in network environment. At this time, the application process identifier is recorded as "e-commerce application A", the process restart time is recorded as 13:45:25 on May 10, 2025, and the restart reason is recorded as automatic process restart caused by the change in network environment.

[0081] The device-side configuration data is combined with the user behavior change data corresponding to the candidate drift segments to construct a feature matrix;

[0082] User behavior change data corresponding to device-side configuration data is extracted from candidate drift fragments. This user behavior change data refers to data on changes in user interaction behavior within the user's terminal device before and after a configuration switch event, including user behavior event type, behavior frequency, behavior duration, number of interactions, or intensity. For example, if a user clicks an ad button on an e-commerce application page 3 times in the 5 minutes before the configuration switch, and the number of clicks increases to 9 in the 5 minutes after the switch, then the user behavior change data is recorded as click count trend data. Location cache status data, privacy policy version data, and application process restart status data are then compared with user behavior data. The changed data are combined to form a feature matrix. For example, each row of the feature matrix represents the combined features of device state and user behavior corresponding to a specific configuration switching event, and each column represents the location cache state data features, privacy policy version data features, application process restart state data features, and corresponding user behavior change features. For example, the first row of the feature matrix shows a change of 0.02 degrees (longitude) in location cache state data, a change of privacy policy version data from "V1.0" to "V1.2", and an application process restart marker of 1 (representing a restart event). The corresponding user behavior change feature is an increase of 6 ad clicks.

[0083] Factor decomposition of the feature matrix yields multiple independent influencing factors;

[0084] Factor decomposition decomposes the relationship between device configuration features and user behavior change features in the feature matrix into a small number of independent factors. Each independent factor represents a potential device configuration factor that has a specific impact on user behavior changes. For example, using matrix decomposition, location cache status data change features, privacy policy version data change features, and application process restart features can be independently decomposed to obtain clearly distinguishable independent factors. For example, location cache status change is decomposed into a "displacement distance" factor, privacy policy version data change is decomposed into a "permission change intensity" factor, and application process restart status is decomposed into an "application restart frequency" factor.

[0085] A causal test was performed on each independent influencing factor and the user behavior change data, and the causal contribution of the independent influencing factors was calculated.

[0086] For each independent influencing factor and user behavior change data, causal relationship hypotheses were established and statistical tests were performed to quantify the contribution of each independent influencing factor to the user behavior change data. For any independent influencing factor, such as the "displacement distance" factor generated by changes in location cache status data, a causal relationship hypothesis was established, assuming a causal relationship between the "displacement distance" factor and user behavior change data (e.g., changes in the number of ad clicks). The causal relationship hypothesis was tested using regression analysis in statistical testing: with the independent influencing factor (displacement distance) as the independent variable and the user behavior change data (changes in the number of ad clicks) as the dependent variable, a regression equation was established, expressed as: User behavior change value = Regression coefficient × Independent influencing factor value + Intercept + Error term, where the regression coefficient represents the contribution of the independent influencing factor to the user behavior change data. The impact of user behavior change data is defined by the intercept, which represents the theoretical baseline value of the user behavior change data when the independent impact factor value is zero. The error term represents the random error that the regression model cannot explain. The least squares method is used to solve for the parameters of the regression equation, i.e., to calculate the regression coefficients and the intercept. The least squares method is implemented as follows: based on the independent impact factor values ​​of multiple historical data points in the device-side configuration data and the corresponding user behavior change data, the sum of squares of the errors between the theoretical and actual values ​​for all data points is calculated. The parameter that minimizes the sum of squares of errors is selected as the final value of the regression equation parameters. For example, assuming that the data is calculated from the data points... The regression coefficient of the displacement distance factor is 2.5, and the intercept is 1.0. A significance test is performed on the regression coefficient using the t-test method in statistics. Specifically, the t-statistic is calculated by dividing the regression coefficient by its standard error, where the standard error represents the accuracy of the estimate of the regression coefficient in the sample. For example, if the standard error of the displacement distance factor regression coefficient of 2.5 is 0.5, then the t-statistic is calculated as: t = 2.5 ÷ 0.5 = 5.0. Based on the given degrees of freedom, the probability value P corresponding to the t-statistic is obtained from the standard t-distribution table. Assuming that the table shows... If P < 0.05, it indicates that the regression coefficient is statistically significant, meaning that the displacement distance factor has a significant causal relationship with the user behavior change data. Finally, the causal contribution of the independent influencing factors is calculated based on the significance level of the regression coefficients. Specifically, the variance of user behavior change caused by the independent influencing factors and the total variance of user behavior are calculated separately. Then, the variance of user behavior change caused by the independent influencing factors is divided by the total variance of user behavior to obtain the causal contribution. For example, if the variance of user behavior change caused by the displacement distance factor is 4.0 and the total variance of user behavior is 10.0, then the causal contribution of the displacement distance factor is: 4.0 ÷ 10.0 = 0.4 indicates that the independent impact factor explains 40% of the changes in user behavior. Following the above method, causal tests were performed on the "Permission Change Intensity" factor of the privacy policy version data and the "Application Restart Frequency" factor of the application process restart status data. The causal contribution of permission change intensity and application restart frequency were both 0.3. The causal contribution of each independent impact factor represents the quantitative explanatory power of that factor on the user behavior change data.

[0087] The pseudo-migration score on the device side is calculated by weighting independent impact factors based on causal contribution.

[0088] The device-side pseudo-migration score is a comprehensive score that quantifies and represents the degree of influence of device-side configuration change factors on user behavior changes during configuration switching events. The weighted calculation method is to use the causal contribution of each independent influencing factor as its weight, and sum them together. For example, assuming that within the candidate drift segment, the displacement distance factor is 0.02, the permission change intensity factor is 0.5, and the application restart frequency factor is 1, then the device-side pseudo-migration score is: (0.02×0.4)+(0.5×0.3)+(1×0.3)=0.008+0.15+0.3=0.458; therefore, the device-side pseudo-migration score corresponding to this configuration switching event is clearly 0.458.

[0089] S4: Deconstruct the network-side factors of the candidate drift segments, and use graph matching and topology consistency verification to assess the impact of network-side environmental changes, outputting network-side pseudo-migration scores, including:

[0090] Extract network-side configuration data from candidate drift segments;

[0091] Network-side configuration data refers to the network-related configuration parameters of the mobile communication network environment to which the user terminal device is connected when switching the configuration of the embedded user identification module. Changes in network-related configuration parameters directly affect the data transmission path and network service quality between the user terminal device and the Internet. Network-side configuration data includes network routing parameter data, access point name parameter data, and domain name system resolution address parameter data. Among them, network routing parameter data represents the network node and address information of the data transmission path used by the user terminal device when connecting to the mobile communication network, such as the gateway address of the route, the number of route hops, and the connection method of the network route (such as dynamic host configuration protocol address allocation method or static configuration method). Access point name parameter data represents the network access parameters used by the user terminal device when accessing the mobile communication network, such as access point name information such as "internet.providerA.com" or "roam.providerB.net". Domain name system resolution address parameter data represents the server address used by the user terminal device to convert domain names to Internet Protocol addresses when connecting to the network, such as Google's public domain name resolution server address "8.8.8.8" or Cloudflare's public domain name resolution server address "1.1.1.1". For example, when a user terminal device switches from operator A to operator B at 13:45:20 Beijing time on May 10, 2025, the network routing parameter data changes from the original routing address '100.64.10.1' to the new routing address '100.70.20.1', the access point name parameter data changes from "internet.providerA.com" to the new "roam.providerB.net", and the domain name system resolution address parameter data changes from the original address "8.8.8.8" to the new address "1.1.1.1". All of the above are recorded in the candidate drift segment.

[0092] Construct a network environment topology diagram based on network-side configuration data;

[0093] A network topology diagram is a data model that represents the relationships between network environment configuration parameters in a graph structure. Network nodes represent key network devices or servers, such as routers, gateways, and DNS server nodes. The edges connecting network nodes in the topology diagram represent the paths and connections for actual network data transmission. The topology diagram illustrates changes in the network environment structure. The method for constructing a network topology diagram involves determining the identity and location of each network node based on network routing parameters, access point name parameters, and DNS resolution address parameters, and then marking the connection status between nodes according to the actual network connections. For example, in the network environment of Carrier A, the constructed network topology diagram includes a network routing node identified by the address '100.64.10.1', an access gateway node with the access point name "internet.providerA.com", and a domain name resolution node with the domain name system resolution server address "8.8.8.8". The nodes form the initial topology diagram according to the actual network connection relationship. In the network environment of Carrier B, a new topology diagram is constructed, consisting of a routing address node '100.70.20.1', an access gateway node "roam.providerB.net", and a domain name resolution node "1.1.1.1". The above two network environment topology diagrams represent the network environment structure before and after the configuration switch, respectively.

[0094] Graph matching is performed on the network environment topology map. By calculating the similarity of the topology structure before and after the network configuration switch in the candidate drift segment, the topology structure matching index is obtained.

[0095] Graph matching is a method for comparing and calculating the network topology before and after a configuration switch. Specifically, it uses a topology similarity algorithm to calculate the degree of similarity between two network topologies, quantifying the change in network configuration before and after the switch, and outputting a topology matching index. The implementation of the topology similarity algorithm is as follows: The network topologies before and after the configuration switch are represented as adjacency matrices, where each matrix element indicates whether there is a direct connection between corresponding nodes. The two adjacency matrices are compared element-by-element. For example, elements in the same position in the two adjacency matrices are considered matching elements, and those that are different are considered non-matching elements. The total number of matching elements is counted and divided by the total number of elements in the adjacency matrices to obtain the topology matching index, which ranges from 0 to 1. For example, if an adjacency matrix has 20 elements, and 16 elements are the same in the two adjacency matrices before and after the configuration switch, the topology matching index is calculated as 16 ÷ 20 = 0.8, indicating that the network topology has 80% similarity before and after the configuration switch.

[0096] Topology consistency verification is performed based on the network environment topology diagram. By calculating the degree of change in network-side configuration data, topology consistency indicators are obtained.

[0097] Network-side configuration data includes network routing parameters, access point name parameters, and domain name system resolution address parameters. Each type of data undergoes numerical encoding and change calculation before and after configuration switching.

[0098] To calculate the degree of change in network routing parameter data, let's take the network address before and after the configuration switch (e.g., changing from '100.64.10.1' to '100.70.20.1') as an example. Specifically, the IP address in the network routing parameter data is divided into four numerical segments. The address '100.64.10.1' before the configuration switch is recorded as the initial address, and the address '100.70.20.1' after the configuration switch is recorded as the changed address. The absolute value of the numerical difference between the corresponding numerical segments is calculated, i.e., the difference of the first numerical segment is 0, the difference of the second numerical segment is 10, the difference of the third numerical segment is 10, and the difference of the fourth numerical segment is 0. The differences of each numerical segment are accumulated, i.e., the total change in network routing parameter data is 20. The total change is divided by the theoretical maximum possible range of change in the network address value (e.g., the theoretical maximum possible value is set to 1020) to obtain the degree of change in network routing parameter data, i.e., 182 ÷ 1020 ≈ 0.019, which represents the quantitative degree of change in network routing parameter data.

[0099] To calculate the degree of change in access point name parameter data, let's take "internet.providerA.com" before the configuration switch and "roam.providerB.net" after the configuration switch as examples. Specifically: convert the access point name string into a character encoding sequence, and calculate the ASCII value of each character. For example, the character "a" corresponds to the ASCII value 97. Calculate the absolute value of the difference in ASCII values ​​of corresponding characters in the two strings. If the string lengths are different, pad the shorter parts with null characters (ASCII value 0). For example, suppose the first string is "internet.providerA.com". The first string "iderA.com" has a length of 22 characters, and the second string "roam.providerB.net" has a length of 18 characters. The second string, which is shorter than the first, is padded with 4 empty characters to make its length match the first string's 22 characters. The absolute value of the ASCII value difference at each character's position is calculated, and all absolute values ​​are summed to obtain the total character difference. For example, the total difference is calculated to be 256. The total character difference is divided by the theoretical maximum possible character difference (e.g., 2000) to obtain the degree of change, i.e., 256 ÷ 2000 = 0.128, representing the degree of change in the access point name parameter data.

[0100] To calculate the degree of change in the DNS resolution address parameters, let's take the change from "8.8.8.8" to "1.1.1.1" as an example. Divide the DNS resolution address into four numerical segments, and calculate the difference for each segment as |8 - 1| = 7. The total difference is 7 + 7 + 7 + 7 = 28. Divide this by the theoretical maximum possible difference (e.g., 1020) to get the degree of change, which is 28 ÷ 1020 ≈ 0.027, representing the degree of change in the DNS resolution address parameters.

[0101] After calculating the degree of change of the network node parameter data, standardization is performed on each parameter while maintaining the ratio between the degrees of change. For example, if the degree of change of network routing parameter data is 0.019, the degree of change of access point name parameter data is 0.128, and the degree of change of domain name system resolution address parameter data is 0.027, the sum is 0.019 + 0.128 + 0.027 = 0.174. The standardized values ​​are 0.019 ÷ 0.174 ≈ 0.110, 0.128 ÷ 0.174 ≈ 0.735, and 0.027 ÷ 0.174 ≈ 0.155, respectively.

[0102] The standardized change values ​​are summed to obtain the topology consistency index, which is 0.534 + 0.384 + 0.082 = 1.0. This indicates that the overall change in the network-side configuration data is 1.0. The closer it is to 1, the higher the overall change is, and the closer it is to 0, the lower the overall change is.

[0103] Based on the topology matching index and the topology consistency index, the network-side pseudo-migration score is calculated.

[0104] The network-side pseudo-migration score is used to represent the impact of network-side environmental changes on user behavior. It is calculated by summing the topology matching index and the topology consistency index using a preset weighting. For example, if the preset weight of the topology matching index is 0.6 and the weight of the topology consistency index is 0.4, then the network-side pseudo-migration score is calculated as (topology matching index × 0.6) + (topology consistency index × 0.4). For example, if the topology matching index is 0.8 and the topology consistency index is 0.9, then the network-side pseudo-migration score is (0.8 × 0.6) + (0.9 × 0.4) = 0.48 + 0.36 = 0.84, which represents the degree of impact of network-side configuration changes on user behavior as 0.84.

[0105] S5: Based on the device-side pseudo-transfer score and the network-side pseudo-transfer score, determine identity perception drift and output a behavior correction vector, including:

[0106] The device-side pseudo-migration scores and network-side pseudo-migration scores are normalized to obtain normalized device-side pseudo-migration scores and normalized network-side pseudo-migration scores.

[0107] Normalization processing refers to converting the device-side pseudo-migration scores and network-side pseudo-migration scores to the same data range to eliminate differences in numerical range or units. First, the device-side pseudo-migration score dataset and the network-side pseudo-migration score dataset are collected from multiple historical embedded user identification module configuration switching events of the user terminal device. The maximum and minimum values ​​of the device-side pseudo-migration score dataset and the network-side pseudo-migration score dataset are determined respectively. Then, the range standardization method is applied to the currently obtained device-side pseudo-migration scores and network-side pseudo-migration scores respectively. The formula for calculating the range standardization method is: Normalized score value = (Score value to be processed - Minimum value of historical score set) divided by (Maximum value of historical score set - Minimum value of historical score set). For example, if the device-side pseudo-migration score calculated for the current candidate drift segment is 0.458, and the maximum value of the historical device-side pseudo-migration score dataset is 0.9 and the minimum value is 0.1, then the normalized calculation of the current device-side pseudo-migration score is: (0.458 - 0.1) divided by (0.9 - 0.1) = 0.358 ÷ 0.8 = 0.4475. Meanwhile, assuming the network-side pseudo-migration score calculated for the current candidate drift segment is 0.84, and the maximum value of the historical network-side pseudo-migration score dataset is 1.0 and the minimum value is 0.2, then the normalized calculation of the current network-side pseudo-migration score is: (0.84 - 0.2) divided by (1.0 - 0.2) = 0.64 ÷ 0.8 = 0.8. After normalization, the normalized device-side pseudo-transfer score is 0.4475, and the normalized network-side pseudo-transfer score is 0.8. The normalized device-side pseudo-transfer score and the normalized network-side pseudo-transfer score are used to ensure that pseudo-transfer scores from two different sources can be quantitatively compared and analyzed on the same scale when constructing the identity recognition drift determination matrix.

[0108] An identity perception drift determination matrix is ​​constructed based on normalized device-side pseudo-transfer scores and normalized network-side pseudo-transfer scores.

[0109] The identity recognition drift determination matrix comprehensively describes the relationship between normalized device-side pseudo-migration scores and normalized network-side pseudo-migration scores in matrix form. The rows of the matrix represent pseudo-migration score types from different sources: device-side pseudo-migration and network-side pseudo-migration. The columns represent the normalized scores calculated under different embedded user identification module configuration switching event conditions. Each element records the normalized device-side pseudo-migration score and the normalized network-side pseudo-migration score under the corresponding condition. For example, under the current candidate drift segment condition, a 2×1 matrix is ​​constructed, where the first row element value is recorded as the device-side normalized pseudo-migration score of 0.4475, and the second row element value is recorded as the network-side normalized pseudo-migration score of 0.8. If multiple historical candidate drift segment conditions exist, the determination matrix is ​​further expanded to a 2×N matrix, where N represents the total number of historical events, and each column records the normalized pseudo-migration score explicitly calculated under a single candidate drift segment event condition.

[0110] Based on the identity perception drift determination matrix, a comprehensive score for identity perception drift is calculated;

[0111] The comprehensive score for identity perception drift is a quantitative indicator used to represent the combined impact of device-side and network-side configuration factors on user behavior change data under the current candidate drift segment conditions. The calculation method is to perform a weighted comprehensive calculation on each score value in the identity perception drift judgment matrix, and the weight values ​​are determined in advance through historical data verification. For example, after historical data analysis, the weight value of the normalized device-side pseudo-migration score is determined to be 0.45, and the weight value of the normalized network-side pseudo-migration score is determined to be 0.55. Then the comprehensive score for identity perception drift is: Comprehensive score = (Normalized device-side pseudo-migration score × Device-side weight) + (Normalized network-side pseudo-migration score × Network-side weight), Comprehensive score = (0.4475 × 0.45) + (0.8 × 0.55) = 0.201375 + 0.44 = 0.641375, indicating that the overall comprehensive impact of identity perception drift under the current candidate drift segment conditions is 0.641375.

[0112] The overall score of identity perception drift is compared with the preset drift judgment threshold, and the drift confidence score is output.

[0113] The drift judgment threshold is obtained in advance through analysis of a large amount of historical configuration switching event data. It represents the critical value of the acceptable degree of identity perception drift impact, and is usually preset to 0.5. If the overall score of the current identity perception drift is greater than or equal to the preset drift judgment threshold, it is determined that identity perception drift exists under the current event conditions, and the drift confidence is output as high confidence. If the overall score of the current identity perception drift is less than the preset drift judgment threshold, it is determined that identity perception drift does not exist under the current event conditions, and the drift confidence is output as low confidence. For example, under the current candidate drift segment conditions, the overall score of identity perception drift is 0.641375, which is greater than the preset drift judgment threshold of 0.5, so the drift confidence is output as high confidence.

[0114] Calculate the behavior correction vector based on the drift confidence level;

[0115] The behavior correction vector is used to correct user behavior change data to eliminate interference from device-side or network-side configuration factors caused by the configuration switching event of the embedded user identification module. The method for calculating the behavior correction vector is to use the drift confidence level as the correction strength coefficient to correct and adjust the fluctuating user behavior features in the user behavior change data. For example, if the recorded change in the number of ad clicks in the user behavior change data is 6 times, and the normal change in the number of clicks determined by historical user behavior data is 2 times, then there are 4 abnormal fluctuations caused by the configuration switching of the embedded user identification module. Using the drift confidence level of 0.641375 as the correction strength coefficient, the behavior correction vector is calculated as: abnormal fluctuation amplitude × drift confidence level = 4 times × 0.641375 ≈ 2.5655 times. The calculated behavior correction vector of 2.5655 times indicates that 2.5655 clicks in the current change in the number of ad clicks are errors caused by identity recognition drift, which can be subtracted in the correction process to restore the true trend of user behavior change data.

[0116] S6: Reconstruct user behavior sequences based on behavior correction vectors, update user behavior features, correct abnormal behavior labels, and output corrected user behavior analysis results, including:

[0117] The behavior correction vector is used to perform vector operations with the user behavior change data corresponding to the candidate drift segment to generate corrected behavior data.

[0118] The behavior correction vector is a set of values ​​obtained through drift confidence calculation to correct abnormal changes in user behavior data. User behavior change data is the set of user behavior event change values ​​extracted within the configuration switch trigger window corresponding to the candidate drift segment. The vector operation is implemented as follows: using the change value of each behavior indicator recorded in the user behavior change data as an initial benchmark, subtract the corresponding correction value from the behavior correction vector. The result is the corrected behavior data value. For example, if the change in ad clicks recorded in the user behavior change data within the candidate drift segment is 6 times, and the behavior correction vector calculates it to 2.5655 times, performing vector subtraction, the corrected change in ad clicks is 6 times minus 2.5655 times, resulting in a true change in ad clicks of 3.4345 times in the corrected behavior data. Repeating this operation on all user behavior indicator change data yields all corrected behavior data corresponding to the candidate drift segment. Through vector operation, the abnormal fluctuations caused by configuration switching of the embedded user identification module can be accurately eliminated, thus obtaining corrected behavior data that truly reflects changes in user behavior trends.

[0119] Replace the corresponding user behavior event sequence in the unified time series dataset with the corrected behavior data to obtain the corrected user behavior sequence;

[0120] Based on the structure of the unified time-series dataset, the start and end positions of the user behavior event sequence corresponding to the current candidate drift segment are located within the unified time-series dataset. The original user behavior event data within this position interval is then replaced sequentially with corrected behavior data, using a numerical overwrite update method. For example, if the user behavior event sequence corresponding to the candidate drift segment in the unified time-series dataset is an increase in ad clicks from 2 to 8, then the calculated corrected behavior data (the change in ad clicks is 3.4345) is used to overwrite the original user behavior data, updating the corresponding change in ad clicks in the unified time-series dataset from 2 to 5.4345. Through these operations, the user behavior event sequence data originally recorded in the unified time-series dataset that was affected by abnormal factors is corrected to a true user behavior change sequence, thus obtaining a corrected user behavior sequence that accurately reflects the actual trend changes in user behavior.

[0121] Based on the corrected user behavior sequence, behavioral statistical features, temporal features and interaction features are extracted to generate an updated user behavior feature set;

[0122] Behavioral statistics are the overall statistics of user behavior data, including the number of behavioral events, the frequency of behavioral events, and the duration of behavioral events; for example, statistically correcting the total number and frequency of ad click events in a user behavior sequence. Temporal features are based on the regularity of changes in corrected user behavior sequence data over time, such as calculating the trend of user ad clicks over time or the distribution of user payment operations over time. Interaction features represent the relationship between user interactions with web applications, such as the interaction relationships between different events in e-commerce applications, such as browsing products, clicking ads, and completing payments. These features are extracted from corrected user behavior sequence data: the extracted behavioral statistics feature is that the corrected ad click event occurred 12 times, with an average click frequency of 2 times per hour; the extracted temporal features show a gradual increase in ad clicks over time, indicating a gradual increase in user interest; the extracted interaction features show that the probability of a user clicking an ad within 30 seconds of browsing a product page is 60%. By performing feature extraction using the methods described above, an updated user behavior feature set can be formed. This updated user behavior feature set depicts the user's true behavioral state and trend characteristics, effectively eliminating abnormal data interference caused by the configuration switching of the embedded user identification module, and ensuring the accuracy and authenticity of user behavior analysis.

[0123] The abnormal behavior label identifier is reset based on the updated user behavior feature set to form an abnormal label update table;

[0124] Abnormal behavior tags are used to label data segments or events in user behavior data that may contain anomalies or risks. The original abnormal behavior tags may be mislabeled due to pseudo-migration factors caused by configuration switching of the embedded user identification module. First, based on the updated user behavior feature set, new rules for judging abnormal behavior are established. For example, the rule that an ad click count exceeds three times the average click count is redefined as abnormal. Each data item in the updated user behavior feature set is then re-evaluated based on these new rules, and the abnormal behavior tags are reset. For instance, if the original abnormal behavior tag determined an ad click count of 8 as abnormal, after updating the user behavior feature set, the average ad click count is corrected to 5, so the abnormal judgment threshold is reset to 15 (5 x 3). If a user clicks 12 times, it is re-judged as non-abnormal. All re-judged abnormal tag data is recorded in the abnormal tag update table. The abnormal tag update indicates which user behavior events have been re-verified as abnormal or non-abnormal, effectively improving the accuracy and effectiveness of abnormal behavior identification.

[0125] The corrected user behavior sequence, updated user behavior feature set, and anomaly label update table are encapsulated into the corrected user behavior analysis results;

[0126] The corrected user behavior analysis results are the output after accurately correcting the impact of user behavior drift caused by the configuration switching event of the embedded user identification module. This result is provided to the user behavior analysis platform or management system for implementing precise user behavior analysis, personalized service recommendations, or security risk warnings. Structured data encapsulation is used, such as using Extended Markup Language (Extended Markup Language) or JavaScript object representation. The corrected user behavior sequence is recorded as a data sequence structure, the updated user behavior feature set is recorded as a data feature list, and the anomaly label update table is recorded as an anomaly behavior identifier list, all encapsulated in a single data package. For example, the corrected user behavior sequence represents a data sequence where the number of ad clicks changes from 2 to 5.4345. The updated user behavior feature set records specific feature values ​​such as ad click frequency and interaction probability. The anomaly label update indicates that the data with 12 ad clicks is updated to a non-anomaly label. Through this encapsulation, the corrected user behavior analysis results are formed, significantly improving the reliability, accuracy, and practicality of the user behavior analysis results.

[0127] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0128] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0129] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0130] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0132] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0133] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0134] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0136] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A user behavior analysis method based on E-SIM, characterized in that, Includes the following steps: S1: Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences, perform time alignment and labeling, and output a unified time-series dataset; network policy parameters include access point name information, network routing information, and domain name resolution server information; S2: Construct a configuration handover trigger window based on a unified time series dataset. Analyze the correlation between network migration and behavioral events within the configuration handover trigger window and output candidate drift segments. The configuration handover trigger window refers to a fixed-length time period extending forward and backward from the time when the configuration handover event occurs in the E-SIM configuration handover log. S3: Deconstruct candidate drift segments using device-side factors, and use factor decomposition and causal testing to assess the impact of device-side configuration change factors, outputting a device-side pseudo-migration score; device-side configuration data includes location cache status data, privacy policy version data, and application process restart status data; the device-side pseudo-migration score is a comprehensive score that quantifies and represents the degree of impact of device-side configuration change factors on user behavior change data during configuration switching events; S4: Deconstruct the candidate drift segments using network-side factors, and use graph matching and topology consistency verification to assess the impact of network-side environmental changes, outputting a network-side pseudo-migration score; network-side configuration data includes network routing parameters, access point name parameters, and domain name system resolution address parameters; the network-side pseudo-migration score is obtained by a weighted sum of the topology matching index and the topology consistency index; S5: Based on the pseudo-transfer scores on the device side and the network side, determine identity perception drift and output a behavior correction vector; Identity perception drift refers to the deviation of user behavior data from the actual change in user behavior caused by the combined effects of device-side configuration changes and network-side environmental changes during the E-SIM configuration switch process. S6: Reconstruct user behavior sequences based on behavior correction vectors, update user behavior features, correct abnormal behavior labels, and output corrected user behavior analysis results.

2. The user behavior analysis method based on E-SIM according to claim 1, characterized in that, S1, specifically: Collect E-SIM configuration handover logs, network policy parameters, and user behavior event sequences; Synchronize and align E-SIM configuration handover logs, network policy parameters, and user behavior event sequences according to a unified time base. The synchronized and aligned E-SIM configuration handover logs, network policy parameters, and user behavior event sequences are classified and labeled to generate a unified time-series dataset. Classification labeling refers to labeling data according to its own attribute characteristics.

3. The user behavior analysis method based on E-SIM according to claim 2, characterized in that, S2, specifically: Construct a configuration switching trigger window centered on the configuration switching timestamp in the E-SIM configuration switching log and with a duration of a preset time window; Extract network policy parameters and user behavior event sequences that fall into the configuration switching trigger window from the unified time series dataset, and combine the extracted results into window data blocks; Based on window data blocks, calculate the temporal cross-correlation coefficient between the network strategy parameter change vector and the user behavior event sequence change vector to generate a network migration and behavior event correlation index. Threshold determination is performed on the correlation indicators between network migration and behavioral events, and time series segments that meet the threshold conditions are selected to output candidate drift segments.

4. The user behavior analysis method based on E-SIM according to claim 3, characterized in that, S3, specifically: Extract device-side configuration data from candidate drift segments; The device-side configuration data is combined with the user behavior change data corresponding to the candidate drift segments to construct a feature matrix; Factor decomposition of the feature matrix yields multiple independent influencing factors; A causal test was performed on each independent influencing factor and the user behavior change data, and the causal contribution of the independent influencing factors was calculated. The pseudo-migration score on the device side is calculated by weighting the independent impact factors based on causal contribution.

5. The user behavior analysis method based on E-SIM according to claim 4, characterized in that, S4, specifically: Extract network-side configuration data from candidate drift segments; Construct a network environment topology diagram based on network-side configuration data; Graph matching is performed on the network environment topology map. By calculating the similarity of the topology structure before and after the network configuration switch in the candidate drift segment, the topology structure matching index is obtained. Topology consistency verification is performed based on the network environment topology diagram. By calculating the degree of change in network-side configuration data, topology consistency indicators are obtained. The network-side pseudo-migration score is calculated based on the topology matching index and the topology consistency index.

6. The user behavior analysis method based on E-SIM according to claim 5, characterized in that, S5, specifically: The device-side pseudo-migration scores and network-side pseudo-migration scores are normalized to obtain normalized device-side pseudo-migration scores and normalized network-side pseudo-migration scores. An identity perception drift determination matrix is ​​constructed based on normalized device-side pseudo-transfer scores and normalized network-side pseudo-transfer scores. Based on the identity perception drift determination matrix, a comprehensive score for identity perception drift is calculated; The overall score of identity perception drift is compared with the preset drift judgment threshold, and the drift confidence score is output. Calculate the behavior correction vector based on the drift confidence.

7. The user behavior analysis method based on E-SIM according to claim 6, characterized in that, S6, specifically: The behavior correction vector is used to perform vector operations with the user behavior change data corresponding to the candidate drift segment to generate corrected behavior data. Replace the corresponding user behavior event sequence in the unified time series dataset with the corrected behavior data to obtain the corrected user behavior sequence; Based on the corrected user behavior sequence, behavioral statistical features, temporal features and interaction features are extracted to generate an updated user behavior feature set; The abnormal behavior label identifier is reset based on the updated user behavior feature set to form an abnormal label update table; The corrected user behavior sequence, updated user behavior feature set, and anomaly label update table are encapsulated into the corrected user behavior analysis results.