User property safety risk early warning method, device, medium and product
By constructing a multi-dimensional anomaly feature matrix and performing multi-model fusion analysis, the limitations of single-dimensional detection in elderly telecommunications patterns have been solved, enabling accurate risk warning and dynamic protection for elderly telecommunications patterns.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP JILIN BRANCH
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, telecommunications security attack protection solutions for the elderly population rely on single-dimensional data detection, which cannot adapt to individual differences, resulting in serious false alarms or missed alarms, and it is difficult to identify complex risks involving multiple minor anomalies.
By acquiring users' communication behavior data, location data, and application usage data, an anomaly feature matrix is constructed. This matrix is then combined with statistical generative models, deep learning discriminative models, and knowledge-driven rule systems to conduct multi-dimensional collaborative analysis and identify complex risks.
It enables accurate risk warnings for telecommunications patterns used by the elderly, reduces false alarms and missed alarms, and provides dynamic abnormal feature matrix identification from group protection to individual protection, achieving a technological breakthrough from post-event blocking to pre-event warning.
Smart Images

Figure CN121786758B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, specifically to a method, device, medium, and product for early warning of user property security risks. Background Technology
[0002] In existing technologies, telecommunications operators' security attack protection solutions for the elderly population are mainly based on rule engines, which detect abnormal behavior by analyzing single-dimensional data such as communication signaling data, location data, and application usage data. However, these methods mainly rely on preset thresholds or general machine learning models to identify security attack behaviors, and the group profiles used cannot adapt to the vastly different living habits among elderly individuals, leading to serious false positives or false negatives.
[0003] Furthermore, existing technologies lack the ability to collaboratively analyze multi-dimensional data, making it difficult to effectively identify comprehensive risks composed of multiple minor anomalies. For example, detection methods based on communication signaling data may misjudge the normal silence of elderly people living alone as abnormal behavior, while detection methods based on location data may misjudge the normal activities of active elderly people as abnormal. These single-dimensional detection methods not only increase the risk of false alarms and false negatives but also struggle to address security attack patterns arising from complex telecommunications patterns composed of multiple minor anomalies. Therefore, existing technologies have significant limitations in identifying security attack risks associated with telecommunications patterns of the elderly, necessitating more advanced technological solutions to address these issues. Summary of the Invention
[0004] At least one embodiment of this application provides a user property security risk warning method, device, medium, and product to address the problem that existing security attack protection schemes for telecommunications modes, which employ a single-dimensional detection method, have significant limitations in identifying security attack risks in telecommunications modes involving the elderly.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a method for early warning of user property security risks, including:
[0007] Based on the user's carrier network behavior data, obtain the user's communication behavior data, location data, and application usage data;
[0008] Based on the communication behavior data, the location data, and the application usage data, target features are generated, and an anomaly feature matrix is constructed based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features;
[0009] Based on the aforementioned anomaly feature matrix, a first generation result, a second generation result, and a third generation result containing the user's risk status are generated using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system, respectively.
[0010] Based on the first generation result, the second generation result, and the third generation result, the predicted status result of the user is determined, and the corresponding property security warning behavior is executed based on the predicted status result.
[0011] Optionally, based on the user's carrier network behavior data, the user's communication behavior data is obtained, including:
[0012] Based on the user's carrier network behavior data, obtain the user's call record data;
[0013] The call record data is cleaned to determine the number of calls between the user and the contact, the call duration between the user and the contact, the number of valid call days between the user and the contact, and the maximum number of calls, the maximum call duration, and the maximum number of call days between the user and all contacts.
[0014] The connection strength calculation model between the user and the contact is determined based on the first ratio of the number of calls to the maximum number of calls, the second ratio of the call duration to the maximum call duration, the third ratio of the effective call days to the maximum call days, and the preset weight value corresponding to each ratio.
[0015] Based on the aforementioned connection strength calculation model, data ranked within the first range are selected as the social circle baseline;
[0016] Based on the call record data and the preset seasonal autoregressive integral moving average model, a communication frequency baseline is determined;
[0017] The social circle baseline and the communication frequency baseline are determined as the user's communication behavior data.
[0018] Optionally, based on the user's carrier network behavior data, the user's location data and application usage data are obtained, including:
[0019] Based on the user's operator network behavior data, obtain the user's user equipment location data and External Data Identification Protocol (XDR) signaling data;
[0020] Based on the user equipment location data, location areas are divided according to the user dwell time, and wireless cells whose dwell time ranks in the second range within each location area are extracted to form location data.
[0021] Based on application usage habits, XDR signaling data and preset application categories are used to construct application feature vectors for corresponding categories as application usage data; the preset application categories are used to indicate financial applications and non-financial applications.
[0022] Optionally, based on the communication behavior data, the location data, and the application usage data, target features are generated, and based on the target features, an anomaly feature matrix is constructed, including:
[0023] Based on the user's historical communication frequency data, a time series prediction model is constructed. The communication behavior data is input into the time series prediction model to determine communication dimension anomaly features. The communication dimension anomaly features are used to quantify the degree of deviation between the user's actual communication behavior on the current day and the historical normal pattern.
[0024] The system obtains the user's current actual social group and determines the first call duration between the current actual social group and the contact person; it also parses the communication behavior data to determine the social circle baseline in the communication behavior data and obtains the second call duration between the social circle baseline and the contact person; it determines social dimension anomaly features based on the first call duration and the second call duration; the social dimension anomaly features are used to indicate the degree of social circle connection decay.
[0025] Based on the application usage data, first-time use marker information is obtained, and based on the usage duration deviation corresponding to the current usage duration and the first-time use marker information, the abnormal features of the financial application usage dimension are determined; the abnormal features of the financial application usage dimension are used to indicate the user's financial application usage habits.
[0026] Based on the application usage data and a preset database, abnormal features of the user's device security dimension are obtained; these abnormal features of the device security dimension are used to indicate the user's non-baseline application usage.
[0027] Based on the location data, determine the spatial dimension anomaly characteristics of the user's indicated location change;
[0028] Based on the application usage data, identify the time-dimensional anomalies in user nighttime data activity.
[0029] The abnormal features in the communication dimension, the abnormal features in the social dimension, the abnormal features in the financial application usage dimension, the abnormal features in the device security dimension, the abnormal features in the spatial dimension, and the abnormal features in the time dimension are taken as target features;
[0030] Based on the target features, an anomaly feature matrix with a preset matrix dimension is constructed.
[0031] Optionally, the user risk status is a discrete risk level, including:
[0032] The first security state indicates that the communication behavior data, the location data, and the application usage data all conform to the historical behavior baseline, and all identified abnormal features are below a preset risk threshold.
[0033] The second security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, one or two features are abnormal, and the degree of abnormality does not reach the threshold for triggering a preset intervention standard; the second security state requires continuous monitoring.
[0034] The third security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, at least three features are abnormal, or there are abnormal features in the financial application usage features or location anomalies; the third security state requires immediate intervention.
[0035] The fourth security state indicates that the user exhibits abnormalities across multiple dimensions of the communication behavior data, the location data, and the application usage data; the fourth security state requires emergency intervention.
[0036] Optionally, based on the anomaly feature matrix, a first generation result, a second generation result, and a third generation result containing the user's risk status are generated using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system, respectively, including:
[0037] When the preset statistical generative model uses a pre-trained Hidden Markov Model, the abnormal feature matrix is input into the Hidden Markov Model, and the Hidden Markov Model is used to analyze the deviation of the current behavior pattern from the baseline to generate a first generation result; the Hidden Markov Model is trained and learned using a preset normal behavior baseline; the first generation result is used to determine whether it is in the fourth safe state;
[0038] When the deep learning discriminative model uses a pre-trained long short-term memory network sequence classifier, an anomaly feature matrix with preset time features is input into the long short-term memory network sequence classifier to generate a second generation result; the long short-term memory network sequence classifier includes an attention mechanism to assign weights to the anomaly features at different time steps; the second generation result is used to determine whether it is in a third safe state;
[0039] The abnormal feature matrix is input into the knowledge-driven rule system to generate a third generation result; the third generation result is used to determine whether the system is in a second safe state.
[0040] Optionally, based on the first generation result, the second generation result, and the third generation result, the predicted status result of the user is determined, and a corresponding property security warning action is executed based on the predicted status result, including:
[0041] If the first generated result, the second generated result, and the third generated result are all greater than the confidence threshold, the predicted state result of the user is determined by a multi-voting method based on the first generated result, the second generated result, and the third generated result.
[0042] If the predicted state indicates the first security state, the property security early warning action will not be issued;
[0043] If the predicted state result indicates the second security state, the property security early warning action will be logged.
[0044] If the predicted state indicates the third security state, the property security early warning action will automatically execute an early warning operation.
[0045] When the predicted state indicates the fourth security state, the property security early warning action adopts an active intervention strategy.
[0046] Secondly, embodiments of this application provide a user property security risk early warning device, comprising:
[0047] The first acquisition module is used to acquire the user's communication behavior data, location data, and application usage data based on the user's operator network behavior data;
[0048] The first processing module is used to generate target features based on the communication behavior data, the location data, and the application usage data, and to construct an anomaly feature matrix based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features;
[0049] The second processing module is used to generate a first generation result, a second generation result, and a third generation result containing the user's risk status based on the abnormal feature matrix, using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system.
[0050] The third processing module is used to determine the user's predicted status result based on the first generation result, the second generation result, and the third generation result, and to execute the corresponding property security warning action based on the predicted status result.
[0051] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method as described in any one of the first aspects.
[0052] Fourthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method as described in any one of the first aspects.
[0053] Compared with existing technologies, the user property security risk early warning method, device, medium, and product provided in this application, based on the user's operator network behavior data, acquires the user's communication behavior data, location data, and application usage data. Based on these data, target features are generated, and an abnormal feature matrix is constructed. By integrating the three types of data—communication, location, and application usage—the abnormal feature matrix enables multi-dimensional collaborative analysis, identifying complex risks from multiple minor anomaly combinations and avoiding misjudgments based on single data. By combining statistical generative models, deep learning discriminative models, and a knowledge-driven rule system, the results of these three models are cross-validated to determine the risk status. This leverages machine learning's ability to identify complex patterns while retaining the deterministic judgment of the rule system, achieving accurate early warning and property security protection. Attached Figure Description
[0054] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0055] Figure 1 A flowchart illustrating the user property security risk warning method provided in this application embodiment;
[0056] Figure 2 This is a schematic diagram of a decision tree structure provided in an embodiment of this application;
[0057] Figure 3 A schematic diagram of the user property security risk warning device provided in the embodiments of this application. Detailed Implementation
[0058] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0059] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc.; an indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0060] To enable those skilled in the art to better understand the embodiments of this application, the following description is provided first:
[0061] Telecommunications security attacks refer to harmful acts such as property security attacks, information theft, and malicious control carried out by criminals using telecommunications network technology against users (especially the elderly). The core objective is to illegally obtain economic benefits or sensitive personal information.
[0062] As described in the background section, existing technologies employ a "one-size-fits-all" approach to group profiling, which fails to adapt to the unique behavioral patterns of individual users, especially the elderly, leading to high false alarms or missed alarms. Secondly, relying on single-dimensional data analysis lacks the fusion analysis of multimodal data such as communication, location, and application usage, making it difficult to identify complex cross-dimensional security attack patterns in telecommunications. Thirdly, based on static rules and thresholds, there is a lack of perception of the dynamic evolution of behavioral patterns, triggering intervention only at critical stages of security attacks in telecommunications, failing to achieve early warning. Finally, intervention strategies are singular, mostly involving blocking or reminders, lacking a refined response mechanism based on risk levels. To address at least one of the above problems, this application provides a user property security risk warning method, device, medium, and product that can reduce or avoid the occurrence of the above situations. By constructing a multi-dimensional personal behavior baseline, fusing multimodal data such as communication, location, and application usage, establishing a dynamic abnormal feature matrix, and using a state transition model to identify risk evolution patterns, this application achieves a technological breakthrough from group protection to individual protection, from single-point detection to multi-dimensional fusion analysis, from post-event blocking to in-event intervention and even pre-event warning, ultimately constructing a precise and proactive user property security protection system.
[0063] This application provides a method, apparatus, medium, and product for early warning of user property security risks. The method and apparatus are based on the same concept, and since the principles by which they solve problems are similar, their implementations can be mutually referenced; repeated details will not be repeated.
[0064] Please refer to Figure 1 This application provides a method for early warning of user property security risks, comprising:
[0065] Step 11: Based on the user's carrier network behavior data, obtain the user's communication behavior data, location data, and application usage data.
[0066] In this application, communication behavior data, location data, and application usage data are extracted and processed based on users' operator network behavior data, providing a comprehensive and accurate data source for subsequent risk identification. By integrating operator network behavior data, multiple key dimensions such as communication, location, and applications are covered, laying a solid data foundation for building a multi-dimensional risk assessment model.
[0067] Communication behavior data includes at least the core social circle and the communication frequency baseline. The core social circle is used to accurately locate the people the user mainly contacts on a daily basis (such as family members and close friends), which is crucial for identifying security attacks that attempt to cut off the elderly from their contact with the outside world. The communication frequency baseline is established for each user through a time series model, which enables the system to accurately identify abnormal communication behaviors that do not conform to the user's normal behavior patterns.
[0068] Step 12: Generate target features based on the communication behavior data, the location data, and the application usage data, and construct an anomaly feature matrix based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features.
[0069] It should be noted that the target feature vector includes, but is not limited to, the following six dimensions of anomaly features:
[0070] Anomaly characteristics in communication dimension: Based on the communication behavior data, extract anomaly characteristics of indicators such as call frequency, call duration, data traffic consumption, and the relationship network structure of communication objects, such as a sudden increase or decrease in call frequency, abnormally long call duration, high-frequency communication during non-working hours, and communication objects changing from normal contacts to unknown numbers.
[0071] Social dimension anomaly features: Based on the social application interaction records in the application usage data, extract features such as abnormal importance of social network nodes, sudden changes in interaction patterns, broken or abnormal addition of social relationship chains, etc. For example, a large number of new unfamiliar friends are added in a short period of time, the scale of the social circle expands abnormally, and the interaction content changes from normal communication to sensitive topics.
[0072] Anomaly characteristics in financial application usage dimension: Based on the operation records of financial applications in the application usage data, extract anomaly characteristics of indicators such as transaction frequency, transaction amount, payment channel, and operation time period, such as large transactions outside of working hours, multiple device logins of the same account in a short period of time, and abnormal fluctuations in transaction amount exceeding the user's historical average.
[0073] Device security dimension anomaly characteristics: Based on the device operation records in the application usage data, extract characteristics such as device permission changes, system vulnerability exploitation, malicious software behavior, and abnormal data transmission encryption status, such as unauthorized system permission granting, frequent application crashes or restarts, and abnormal network connection requests.
[0074] Spatial dimension anomaly features: Based on the location data, extract anomaly features of indicators such as the continuity of location trajectory, the distribution of dwell time, the rate of location change, and the degree of deviation from historical activity areas. For example, staying in inactive areas for a long time, obvious abnormal jumps in location trajectory, and high-frequency movement across regions in a short period of time.
[0075] Time-dimensional anomaly features: Based on the timestamps of the communication behavior data, the location data, and the application usage data, features such as the distribution of activity periods, the interval between events, and deviations from periodic patterns are extracted. For example, normal work and rest patterns are broken, multiple abnormal events occur in a short period of time, and the interval between events is shorter than the historical average.
[0076] Through step 12 above, multi-source heterogeneous data is transformed into a structured abnormal feature matrix, providing a data foundation for subsequent abnormal behavior identification and risk warning.
[0077] Step 13: Based on the abnormal feature matrix, using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system, generate a first generation result, a second generation result, and a third generation result containing the user's risk status, respectively.
[0078] Step 13 integrates the outputs of different types of models to obtain a comprehensive judgment on the user's risk status. The input data is an anomaly feature matrix. This matrix is the output of step 12 and contains key risk features and their quantified values extracted from multi-source data such as user behavior data, environmental data, and social data.
[0079] It's important to note that statistical generative models learn patterns and distributions of normal behavior based on historical data. They take an anomaly feature matrix as input and output a probability distribution representing the likelihood that a user's state conforms to a normal pattern. Deep learning discriminative models, through learning from historical risk cases, possess powerful pattern recognition and classification capabilities. They take an anomaly feature matrix as input and directly output a risk level or category. Knowledge-driven rule systems are a series of logical rules built upon domain expert knowledge and industry best practices. For example, if a user repeatedly attempts to make large transfers to unfamiliar accounts within a short period, a high-risk rule is triggered. It takes an anomaly feature matrix as input, determines whether to trigger the rule based on preset rules, and outputs a Boolean value (triggered or not triggered) or the risk type identified by the rule.
[0080] It should also be noted that the first generated result is output by a statistical generative model, representing a probability distribution about the user's state. The second generated result is output by a deep learning discriminative model, representing a clear risk level or category. The third generated result is output by a knowledge-driven rule system, representing the judgment result triggered by the rule.
[0081] Step 14: Determine the user's predicted status result based on the first generation result, the second generation result, and the third generation result, and execute the corresponding property security warning action based on the predicted status result.
[0082] In this embodiment, step 14 integrates the outputs of multiple models, makes a final judgment, and takes corresponding actions. The first generated result, the second generated result, and the third generated result are fused, and based on the fused information, the predicted state result for the user is finally determined. This result clarifies the degree of property security risk currently faced by the user. Based on the determined predicted state result (e.g., high risk), a corresponding early warning strategy is triggered.
[0083] Optional examples of warning actions are as follows: Push Notification: Send warning information to the user, their family, or guardian, informing them of the risk situation. Telephone Warning: For high-risk users, the system may automatically or through customer service initiate a telephone warning to provide risk alerts and guidance. Operation Restriction: Real-time blocking or restriction of extremely high-risk operations (such as large-amount transfers). Manual Intervention: For situations exceeding the system's automatic processing capacity, the system will push the warning information to relevant risk control or customer service personnel for manual intervention.
[0084] The proposed solution constructs a multi-dimensional personal behavior baseline, integrates multi-modal data such as communication, location, and application usage, establishes a dynamic anomaly feature matrix, and adopts a state transition model to identify the risk evolution pattern. This achieves technological breakthroughs from group protection to individual protection, from single-point detection to multi-dimensional fusion analysis, from post-event blocking to in-event intervention and even pre-event warning, ultimately building a precise and proactive user property security protection system.
[0085] It should be noted that the communication behavior data, location data, and application usage data determined in step 11 are respectively the determined communication behavior baseline, location trajectory baseline, and application usage baseline. The combination of the three forms the individual behavior baseline.
[0086] Furthermore, in step 11 above, obtaining the user's communication behavior data based on the user's carrier network behavior data includes:
[0087] Based on the user's carrier network behavior data, obtain the user's call record data;
[0088] The call record data is cleaned to determine the number of calls between the user and the contact, the call duration between the user and the contact, the number of valid call days between the user and the contact, and the maximum number of calls, the maximum call duration, and the maximum number of call days between the user and all contacts.
[0089] The connection strength calculation model between the user and the contact is determined based on the first ratio of the number of calls to the maximum number of calls, the second ratio of the call duration to the maximum call duration, the third ratio of the effective call days to the maximum call days, and the preset weight value corresponding to each ratio.
[0090] Based on the aforementioned connection strength calculation model, data ranked within the first range are selected as the social circle baseline;
[0091] Based on the call record data and the preset seasonal autoregressive integral moving average model, a communication frequency baseline is determined;
[0092] The social circle baseline and the communication frequency baseline are determined as the user's communication behavior data.
[0093] In this embodiment, the core network signaling system of the operator is used to obtain the user's call record data based on the user's operator network behavior data. This call record data is a complete detailed record of the user's calls, and each call record includes at least the following fields: unique call identifier (call_id), user ID (user_id), call timestamp (timestamp), communication type (call_type), call direction (direction), peer number (peer_number), call duration (duration_seconds), location information (location_info), and call result (call_result). Furthermore, the call record data is cleaned to remove invalid, duplicate, or abnormal call records to ensure the accuracy and reliability of subsequent analysis. A set of valid records is selected using preset rules (such as call duration greater than 0, call result being "answered," etc.). .
[0094] Based on the set of valid records A model for calculating the connection strength between user u and contact c is constructed. The output of this model is the connection strength S(u,c), which quantifies the closeness of the social relationship between user u and contact c. The formula for calculating the connection strength S(u,c) is as follows: ,in, Used to represent the set of valid records The number of calls between user u and contact c; Used to represent the set of valid records The call duration between user u and contact c; D(u,c) represents the set of valid records. The number of valid call days between user u and contact c, i.e., the number of days during which the two people have called each other at least once; , and These are used to represent the valid record set. The maximum number of calls, duration, and number of days between user u and all their contacts; α, β, and γ are preset weighting coefficients, α+β+γ=1. Based on the descending order of the connection strength S(u,c) of all contacts of user u output by the connection strength calculation model, the data in the first range are selected as the baseline of the social circle. Here, "ranked in the first range" can be selected as the top K contacts to construct the core social circle, which can be expressed as: .
[0095] Optionally, based on Dunbar's number theory and the stable size of human social networks (approximately 150 people), K=20 is selected here as the size of the core social circle, focusing on key contacts.
[0096] Furthermore, based on the call record data and a preset seasonal autoregressive integral moving average model, a communication frequency baseline is determined, including:
[0097] The seasonal autoregressive integral moving average (SARIMA) model is represented as: SARIMA(p,d,q) Where p represents the non-seasonal autoregressive order (describing the dependency between the current value and the previous p periods); d represents the non-seasonal differencing order (eliminating non-seasonal trends); q represents the non-seasonal moving average order (describing the dependency between the current value and the previous q periods' error term); P represents the seasonal autoregressive order (describing the dependency within the seasonal cycle); D represents the seasonal differencing order (eliminating seasonal trends); Q represents the seasonal moving average order (describing the dependency of the error term within the seasonal cycle); and s represents the seasonal cycle (here set to 7 days, reflecting the periodic changes in call frequency within a week).
[0098] The model equations of the SARIMA model are expressed as follows: ;in, This represents the seasonal autoregressive operator. These are the autoregressive coefficients. This represents the seasonal shift operator, which functions to shift the variable... Shift back 7 cycles (assuming a cycle of 7 days); B represents a non-seasonal autoregressive operator; similarly, B is a non-seasonal shift operator, shifting the variable... Shift back one cycle. Let (1-B) represent the non-seasonal difference operator, where (1-B) is a first-order difference operator. It is a d-order non-seasonal difference operator, which will convert the original sequence Perform d-fold differencing to eliminate non-seasonal trends or periodicity; B is the non-seasonal shift operator. It is a D-order seasonal difference operator, which will convert the original sequence Perform D seasonal differencing to eliminate seasonal trends in the data. It is a seasonal moving average operator. It is the moving average coefficient. It is a seasonal shift operator. This represents the current random error term. The linear relationship between the current random error term and the random error term from 7 periods ago; θ is the moving average coefficient, and θ(B) represents the current random error term. The linear relationship between the random error term and the previous period; It can be a white noise sequence. This represents the differencing and modeling sequence. In the SARIMA model, this sequence is typically a stationary sequence obtained after one or more differencing processes. This equation describes a mixed model that includes non-seasonal and seasonal autoregressives, non-seasonal and seasonal differencing, and non-seasonal and seasonal moving averages.
[0099] First, the SARIMA model analyzes the original sequence. Processing, through Perform d-fold non-seasonal differencing and D-fold seasonal differencing to make it a stationary series. Then, use... and These two autoregressive operators are used to capture patterns in stationary series caused by short-term (non-seasonal) and long-term (seasonal) correlations. Simultaneously, using... and These two moving average operators are used to capture patterns in stationary sequences caused by short-term (non-seasonal) and long-term (seasonal) random fluctuations. Ultimately, the output of the entire model is... It is used to express, and it explains It includes all the predictable components. Simply put, this is a more complex SARIMA model that considers not only the seasonality of the data (with a period of 7) but also the non-seasonal long-term trends.
[0100] It should be noted that in the SARIMA model, It is a "differentiated and modeled sequence", which is a core time series extracted and transformed from historical call record data. This is the result of preprocessing the raw call log data. Its core function is to transform discrete call events into continuous time series for analysis by the SARIMA model. The communication frequency baseline refers to the user's expected communication frequency over a future period (e.g., the average number of calls per day). The SARIMA model learns... Historical patterns are used to predict future communication frequencies, thereby determining the baseline.
[0101] This application uses historical call record data to obtain a sequence. For the sequence Performing seasonal differencing is the first step in eliminating seasonal trends. The formula is: Where s is the seasonal cycle (e.g., 7 days). These are the observations at time point t; At a certain point in time The observed values are , where s is the seasonal period. For example, for weekly data, s can be 7 (representing each week), and for monthly data, s can be 12 (representing each year). This is the result after seasonal differencing. This step eliminates fluctuations in the data with a period of s. Then, regular differencing (non-seasonal differencing) is performed. If the data still exhibits a trend or non-stationarity after seasonal differencing, regular differencing is necessary. The formula is: , Represented as the previous moment Observed values; Represents the previous moment s time units ago. The observed values. This step will eliminate long-term trends in the data.
[0102] Model order identification: The order (p, d, q) and (P, D, Q) of a model need to be determined through data analysis. This application identifies the model order using the autocorrelation function (ACF) and partial autocorrelation function (PACF), and solves for the model parameters using maximum likelihood estimation. ACF describes the degree of linear correlation between current and past values. Non-seasonal ACF is used to determine q (moving average order). Seasonal ACF is used to determine Q (seasonal moving average order). The partial autocorrelation function (PACF) represents the degree of direct correlation between current and past values after removing the influence of intermediate variables. Non-seasonal PACF is used to determine p (autoregression order). Seasonal PACF is used to determine P (seasonal autoregression order). Parameter solving (maximum likelihood estimation): Once the approximate order of the model is determined, maximum likelihood estimation (MLE) can be used to precisely solve for all the model parameters (including autoregression coefficients). Moving average coefficient Θ, and error variance ).
[0103] The goal of maximum likelihood estimation in the log-likelihood function is to find a set of parameters that maximizes the probability of the observed data. This is achieved by maximizing the log-likelihood function.
[0104] Log-likelihood function: Where n is the number of samples, This represents all the parameters to be estimated in the SARIMA model. These are model parameters; The variance of the model residuals represents the degree of data variation. Let be the error term, representing the deviation at the t-th observation, which is the difference between the observed value and its expected value. In the log-likelihood function... It is the constant term of the normal distribution, determined by the sample size n and the residual variance; in the log-likelihood function... This is the sum of squared residuals, representing the contribution of each observation to the likelihood. Smaller errors (biases) result in a larger likelihood function, indicating a better fit between the observations and the model. Simply put, this formula is the core tool for SARIMA model parameter estimation. By maximizing the log-likelihood function, we can find the parameter combination that best fits the model to historical call data.
[0105] Furthermore, once the parameters are estimated, the model can be used to predict future communication frequencies.
[0106] The predicted values are calculated based on a pre-estimated SARIMA model and can predict time series values for the next h steps (e.g., the next 7 or 30 days). Based on the estimated SARIMA model, the predicted values for the next h steps are... Then the 95% confidence interval is: , This represents the predicted value for the future time t+h, which is an estimate of the target variable; 1.96 is the z-score of the 95% confidence interval in a normal distribution, indicating that under a standard normal distribution, 95% of the data falls within ±1.96 standard deviations of the mean. Indicates the error term The standard deviation reflects the uncertainty of the forecast; Var is an abbreviation for variance, representing the degree of variation in the data. This formula provides an interval estimate, representing the range of confidence in future predicted values. Within this interval, there is a 95% probability that the actual value will fall within this range.
[0107] By inputting call log data into a trained SARIMA model, the communication frequency baseline can be directly determined.
[0108] Finally, the social circle baseline and the communication frequency baseline are determined as the user's communication behavior data.
[0109] Furthermore, based on step 11 above, the user's location data and application usage data are obtained based on the user's carrier network behavior data, including:
[0110] Based on the user's operator network behavior data, obtain the user's user equipment location data and External Data Representation (XDR) signaling data;
[0111] Based on the user equipment location data, location areas are divided according to the user dwell time, and wireless cells whose dwell time ranks in the second range within each location area are extracted to form location data.
[0112] Based on application usage habits, XDR signaling data and preset application categories are used to construct application feature vectors for corresponding categories as application usage data; the preset application categories are used to indicate financial applications and non-financial applications.
[0113] In this embodiment of the application, in order to comprehensively evaluate the communication behavior patterns of elderly users and construct a benchmark model of their normal behavior, this application further establishes a location trajectory baseline and an application usage baseline from two dimensions: spatial behavior and application usage behavior, that is, to construct the user's location data and application usage data.
[0114] The data source location data originates from the mobile operator's base station signaling system and contains periodic location update records of user equipment in different wireless cells. The core fields of each record include the user's unique identifier (user_id), timestamp, location type, and specific geographic coordinate information, such as location area code (lac), cell identifier (cell_id), latitude, and longitude.
[0115] Based on the user's operator network behavior data, data preprocessing is performed. This preprocessing cleans the raw location records in the operator network behavior data, filtering out invalid or abnormal location points to ensure data accuracy and integrity. Core location point identification is based on the user's dwell time in different wireless cells, extracting the wireless cells with the second-highest dwell time ranking in each location area to form location data. For example, the user's spatial activity trajectory is divided into three categories of core location points: Residence: The top 3 wireless cells with the highest cumulative dwell time during the day (e.g., 06:00-22:00). Workplace: The top 3 wireless cells with the highest cumulative dwell time at night (e.g., 22:00-06:00). Other points of interest: All wireless cells with the highest dwell time, excluding the residential and workplace cells mentioned above. Location trajectory baseline set generation integrates the latitude and longitude, dwell time, and other information of the above three categories of core location points (a total of 3+3+10=16 wireless cells) to form the user's location trajectory baseline set. This set represents typical patterns of users' daily spatial activities and serves as an important reference benchmark for judging whether their spatial behavior is abnormal.
[0116] The data source primarily utilizes XDR (Extended Data Recording) extracted from user network traffic by mobile operators using Deep Packet Inspection (DPI) technology. This collection method is passive monitoring, ensuring data privacy and compliance without interfering with user usage. Each record includes fields such as user unique identifier (user_id), timestamp, application package name (app_package), application name (app_name), usage duration (usage_duration), data volume (data_volume), and network type (network_type). Application classification and feature extraction are performed to accurately capture user application usage habits, categorizing applications into two main dimensions: Financial applications: including but not limited to applications related to financial transactions such as banking, securities, insurance, and wealth management; Other applications: including non-financial applications such as social communication, lifestyle services, and leisure and entertainment.
[0117] For each user, their usage characteristics across each application are parsed and extracted from the XDR data, including:
[0118] Average daily usage frequency: The average number of times the application is used each day within the statistical period.
[0119] Average daily data consumption: The average amount of data traffic consumed by the application each day within the statistical period.
[0120] Average duration per use: The average duration of using the application per use within the statistical period.
[0121] Dispersion: Measures the stability of user use of the application, defined as Dispersion = Cumulative usage days / Total number of days in the statistics, where the cumulative usage days refer to the number of days in the statistical period during which the application has been used at least once.
[0122] Feature vector construction: A feature vector (AppProfile) is created for each application, with the following format: AppProfile=[Application name, average daily usage frequency, average daily data consumption, average duration per session, usage dispersion]. This feature vector comprehensively describes the user's behavior patterns on the application.
[0123] The application uses baselines to generate application feature vectors for each user, combined with their overall application usage patterns, to generate baselines for both financial and non-financial (i.e., other application) categories. Top M Financial Applications: From all financial applications, the top M applications by user frequency or data consumption are selected, and their feature vectors constitute the baseline for financial applications. Top N Other Applications: From all other application categories, the top N applications by user frequency or data consumption are selected, and their feature vectors constitute the baseline for other application categories. This application utilizes the financial and non-financial baselines to jointly construct the user's application usage baseline, which is used to compare with the user's future real-time behavior data to detect any abnormal behavior deviating from the normal pattern, such as a sudden increase in frequent use of financial apps or abnormal activity at night.
[0124] Optionally, step 12 above includes:
[0125] Based on the user's historical communication frequency data, a time series prediction model is constructed. The communication behavior data is input into the time series prediction model to determine communication dimension anomaly features. The communication dimension anomaly features are used to quantify the degree of deviation between the user's actual communication behavior on the current day and the historical normal pattern.
[0126] The system obtains the user's current actual social group and determines the first call duration between the current actual social group and the contact person; it also parses the communication behavior data to determine the social circle baseline in the communication behavior data and obtains the second call duration between the social circle baseline and the contact person; it determines social dimension anomaly features based on the first call duration and the second call duration; the social dimension anomaly features are used to indicate the degree of social circle connection decay.
[0127] Based on the application usage data, first-time use marker information is obtained, and based on the usage duration deviation corresponding to the current usage duration and the first-time use marker information, the abnormal features of the financial application usage dimension are determined; the abnormal features of the financial application usage dimension are used to indicate the user's financial application usage habits.
[0128] Based on the application usage data and a preset database, abnormal features of the user's device security dimension are obtained; these abnormal features of the device security dimension are used to indicate the user's non-baseline application usage.
[0129] Based on the location data, determine the spatial dimension anomaly characteristics of the user's indicated location change;
[0130] Based on the application usage data, identify the time-dimensional anomalies in user nighttime data activity.
[0131] The abnormal features in the communication dimension, the abnormal features in the social dimension, the abnormal features in the financial application usage dimension, the abnormal features in the device security dimension, the abnormal features in the spatial dimension, and the abnormal features in the time dimension are taken as target features;
[0132] Based on the target features, an anomaly feature matrix with a preset matrix dimension is constructed.
[0133] In this embodiment, six core abnormal feature dimensions are constructed. Each feature reflects different abnormal situations of user behavior and is derived by comparing actual data with baseline or historical normal pattern data.
[0134] (I) Communication Dimension Anomaly Features (F1): This dimension feature is used to quantify the degree of deviation between a user's actual communication behavior on a given day and their historical normal pattern. Based on the user's historical communication frequency data, a time series prediction model is constructed. This time series prediction model can be defined by the SARIMA model prediction residuals of the communication frequency baseline. ,in, This represents the actual observed value, here it is the total actual call duration for the day (a core indicator in communication behavior data); This represents the predicted value output by the SARIMA model; This represents the standard deviation of the forecast, reflecting the uncertainty of the forecast. This represents the minimum standard deviation threshold (to prevent division by zero) used to ensure the stability of the model within a specific range; This represents the call anomaly weighting coefficient, used to adjust the ratio between the actual and predicted values. Normalize it. ,in, This represents the eigenvalues after normalization. This is a constant parameter used to adjust the model's sensitivity to the F1 score, affecting how the model responds to prediction errors. Specifically, it is used to control the growth rate of abnormal scores; exp is an exponential function, a commonly used function in mathematics, which has the characteristic of rapid growth.
[0135] (II) Social Dimension Anomaly Features (F2): Quantifying the degree of weakening of connections within the core social circle, used to represent the degree of weakening of connections within the social circle, the method is as follows: ;in, This indicates the duration of the first call between the current group of people in contact and contact person c. This represents the second call duration between the social circle baseline and contact c (obtained from communication behavior data analysis).
[0136] (III) Abnormal Features in Financial Application Usage Dimension (F3): Detects abnormal usage of financial applications and is used to indicate users' financial application usage habits. ,in, This represents the set of financial applications in the application usage data, and the scope of financial applications selected from the application usage data; First_use(a) represents the first use indicator function, which is 1 when application a is first observed to be launched in the user's entire history, and 0 otherwise; This is a preset constant used to ensure that the denominator is not zero; T(a) represents the current usage duration; This indicates the weight for abnormal usage duration. Furthermore, F3 is standardized: ;in, Represented as risk weights for financial applications, ; This is an empirical upper limit.
[0137] (iv) Device security dimension abnormal features (F4) are used to detect non-baseline application usage and to indicate the user's non-baseline application usage. ,in, This represents a set of suspicious applications in the application usage data, which is a non-baseline set of applications identified by matching application usage data with a preset database. This represents the collection of all financial applications in the data used by the application.
[0138] (v) Spatial Dimension Anomaly Features (F5): Anomaly detection is performed based on location data (i.e., location trajectory baseline) to indicate the user's location changes. ;in, This indicates the number of visits to unfamiliar locations in the location data. This indicates that the location data includes all locations, both familiar and unfamiliar. The Mahalanobis distance measures the overall deviation of the activity trajectory corresponding to the current location data from the historical activity baseline. Relative to the active baseline The method for calculating Mahalanobis distance is as follows: ;in, Σ is the current position, which can be understood as a point to be evaluated; μ is the mean vector of the data distribution, i.e., the activity baseline; Σ is the covariance matrix of the data, which describes the variance of each dimension and the covariance between dimensions. It is the inverse of the covariance matrix, used to "correct" the distribution pattern of the data; Represents the transpose of a vector or matrix. This represents the confidence threshold for Mahalanobis distance. Under the assumption of a normal distribution, approximately 95% of historical data points will have a Mahalanobis distance lower than this value. This represents the distance to the anomaly weight, here . This represents the maximum value of the Mahalanobis distance.
[0139] (vi) Time-dimensional anomaly features (F6) detect abnormal nighttime network activity, which is used to indicate the user's nighttime data activity. ,in, This refers to nighttime data traffic, such as the total data traffic during the current night (e.g., 0:00-6:00). This represents the average historical nighttime data traffic for that user. This represents the standard deviation of the user's historical nighttime data traffic. It represents the standard deviation of non-nighttime data traffic, characterizing the fluctuation of daytime data traffic. This represents the maximum of the standard deviations for nighttime and non-nighttime data, used in the normalization process to ensure the stability of the calculation results. The formula compares the difference between nighttime and non-nighttime data traffic. Further normalization of F6 ensures it lies within the (0,1) interval; this comparison quantifies the relative intensity of nighttime traffic, providing an assessment of the stability of nighttime activity. , This is the standardized value.
[0140] After determining the features from F1 to F6 using the above method, an anomaly feature matrix is constructed using a sliding time window. .
[0141] Where T=7 is the size of the time window, and the matrix dimension is... For the abnormal feature matrix Each feature dimension in the dataset is standardized using z-score: ,in Let be the mean and standard deviation of feature i on the training set. Missing values are handled using time series interpolation.
[0142] Optionally, the user risk status is a discrete risk level, including:
[0143] The first security state indicates that the communication behavior data, the location data, and the application usage data all conform to the historical behavior baseline, and all identified abnormal features are below a preset risk threshold.
[0144] The second security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, one or two features are abnormal, and the degree of abnormality does not reach the threshold for triggering a preset intervention standard; the second security state requires continuous monitoring.
[0145] The third security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, at least three features are abnormal, or there are abnormal features in the financial application usage features or location anomalies; the third security state requires immediate intervention.
[0146] The fourth security state indicates that the user exhibits abnormalities across multiple dimensions of the communication behavior data, the location data, and the application usage data; the fourth security state requires emergency intervention.
[0147] In this embodiment, four discrete risk states are defined. The first security state S0 is a safe state, indicating that the communication behavior data, the location data, and the application usage data, when compared with the corresponding preset historical behavior baseline, all conform to the historical behavior base station, representing that all abnormal features are below the preset risk threshold. The second security state S1 is a state of concern, indicating that the abnormal feature matrix If one or two features exhibit anomalies, these anomalies may be minor, requiring attention but not immediate intervention. The third safety state, S2, is a warning state, representing the anomaly feature matrix. If at least three features are abnormal, or if key features show significant anomalies, such as abnormal usage characteristics of financial applications or abnormal location changes, the third security state requires immediate intervention. The fourth security state, S3, is a high-risk state, indicating severe anomalies across multiple dimensions, accompanied by high-risk behavioral patterns, and requires emergency intervention.
[0148] Furthermore, based on the aforementioned anomaly feature matrix, and utilizing a pre-defined statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system, a first generation result, a second generation result, and a third generation result containing the user's risk status are generated, including:
[0149] When the preset statistical generative model uses a pre-trained Hidden Markov Model, the abnormal feature matrix is input into the Hidden Markov Model, and the Hidden Markov Model is used to analyze the deviation of the current behavior pattern from the baseline to generate a first generation result; the Hidden Markov Model is trained and learned using a preset normal behavior baseline; the first generation result is used to determine whether it is in the fourth safe state;
[0150] When the deep learning discriminative model uses a pre-trained long short-term memory network sequence classifier, an anomaly feature matrix with preset time features is input into the long short-term memory network sequence classifier to generate a second generation result; the long short-term memory network sequence classifier includes an attention mechanism to assign weights to the anomaly features at different time steps; the second generation result is used to determine whether it is in a third safe state;
[0151] The abnormal feature matrix is input into the knowledge-driven rule system to generate a third generation result; the third generation result is used to determine whether the system is in a second safe state.
[0152] It should be noted that, when the preset statistical generative model uses a pre-trained Hidden Markov Model, this application sets state transition path constraints to avoid state jumps, allowing only transitions between adjacent states to ensure the continuity of state evolution. The specific rules are as follows: If ,but ;in, Indicates time Given state i, the probability of transitioning to state j at time t; i represents the current state; j represents the next state; This represents the difference between states, reflecting the distance between them. When When the distance between states is greater than 1, the probability of transition is set to 0. This means that states cannot jump within a single transition, but can only transition between adjacent states.
[0153] In the first approach, the preset statistical generative model is a Hidden Markov Model (HMM) for state recognition. The preset HMM structure is as follows: Hidden states: Observed variables: Statistical summary features of the anomaly feature matrix: ;in, Representing the abnormal feature matrix The mean of each feature dimension; Representing the abnormal feature matrix Standard deviation of each feature dimension; Representing the abnormal feature matrix The maximum value of each feature dimension; Representing the abnormal feature matrix The trend slope of each feature dimension.
[0154] Based on the pre-defined structure of the Hidden Markov Model (HMM), the observation vectors follow a multivariate Gaussian distribution, and the HMM is trained using a forward algorithm. The anomaly feature matrix is input into the trained HMM, and the deviation of the current behavior pattern from the baseline is analyzed using the Hidden Markov Model to generate a first generated result. This first generated result indicates the predicted state at the current time: And the first generated result can be used to detect whether the current state has changed: . Represented as the forward probability generated using the forward algorithm The state i that takes the maximum value. The conditions or processes representing state changes typically involve how to transition from one state to another; the first generated result here can be used to determine whether the state is in the fourth safe state.
[0155] In the second approach, a Long Short-Term Memory (LSTM) sequence classifier is used for state recognition. This application pre-constructs an LSTM-based sequence classification model to output the current transition probability. The LSTM sequence classification model is constructed based on the LSTM hidden state, an anomaly feature matrix incorporating temporal feature samples, an attention mechanism, and state probability prediction.
[0156] When the deep learning discriminative model employs a pre-trained long short-term memory network sequence classifier, the abnormal feature matrix... Adding preset time features, such as flattening the anomaly feature matrix and then adding time features, is represented as follows: , It is represented as an input feature vector, which is an anomaly feature matrix containing time information. This represents the observation data vector at the current moment. This indicates the day of the week. The binary feature indicates whether the current time is a weekend. This application inputs an anomaly feature matrix with preset time features into the Long Short-Term Memory network sequence classifier to generate a second generation result; the second generation result is used to determine whether the system is in a third safe state.
[0157] In the third approach, the knowledge-driven rule system can be selected as a rule-based state discriminator. This system summarizes and generalizes knowledge from security attack experts, analyzes current and recent anomaly characteristics based on a pre-defined, interpretable rule set, and directly outputs a state determination result. The anomaly characteristic matrix is then input into the knowledge-driven rule system to generate a third result; this third result is used to determine whether the system is in a second security state.
[0158] Specifically, refer to Figure 2 As shown, the system represents four security states: S0 as the safe state, S1 as the attention state, S2 as the warning state, and S3 as the high-risk state. Based on the anomaly feature matrix, a pre-defined statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system are used to generate three generation results, each containing the user's risk state: a first generation result, a second generation result, and a third generation result. The anomaly feature matrix includes features F1 to F6 as described above. For example, the S3 (high-risk state) rules include: Rule 1.1 (fund operation mode): (F3>=0.8 and F1>=0.7); or Rule 1.2 (deeply deceived mode): (F2>=0.6 and F5>=0.7 and F6>=0.6). The S2 (warning state) rules include: Rule 2.1 (mid-stage of security attack development): (F1>=0.6) lasting at least 3 days; or Rule 2.2 (suspicious software risk): (F4>=0.5 and F6>=0.6). The S1 (Attention State) rules include: Rule 3.1 (Initial Contact Signal): (F1>=0.4 and F2>=0.3). The S0 (Safe State) rules: When none of the above rules are met, the system is defaulted to the S0 state.
[0159] Furthermore, based on the first generation result, the second generation result, and the third generation result, the predicted status result of the user is determined, and based on the predicted status result, corresponding property security early warning actions are executed, including:
[0160] If the first generated result, the second generated result, and the third generated result are all greater than the confidence threshold, the predicted state result of the user is determined by a multi-voting method based on the first generated result, the second generated result, and the third generated result.
[0161] If the predicted state indicates the first security state, the property security early warning action will not be issued;
[0162] If the predicted state result indicates the second security state, the property security early warning action will be logged.
[0163] If the predicted state indicates the third security state, the property security early warning action will automatically execute an early warning operation.
[0164] When the predicted state indicates the fourth security state, the property security early warning action adopts an active intervention strategy.
[0165] In this embodiment, the first, second, and third generation results output by the HMM state recognizer, LSTM sequence classifier, and rule-based state discriminator are integrated to achieve risk state prediction. Specifically, the model confidence level is: Here, i and j are states in the state space, namely S0, S1, S2, and S3. i represents the state with the highest probability, i.e., the most likely current state according to the model, and j represents the state with the second highest probability, i.e., the second most likely current state according to the model. Confidence levels are determined for the first, second, and third generated results, and the confidence thresholds for these three results are specified. All can be selected as When it is determined to be greater than the corresponding confidence threshold If the first, second, and third generated results are deemed reliable, then the current result is considered low, triggering manual review and temporarily suspending automatic decision-making. Based on the first, second, and third generated results, a multi-voting method is used to determine the user's predicted state. Specifically, this can be achieved through... , This is the first generated result; This is the second generated result; The third generated result is determined by a multi-voting method where the result with the most votes among the three wins. That is, the predicted state result of the user is determined based on the states corresponding to at least two of the three results.
[0166] Furthermore, based on the predicted state results, different measures are taken, including:
[0167] No risk (S0): The system determines that the user's behavior is normal, does not trigger any warnings, and only performs routine data recording.
[0168] Low Risk (S1): The system identifies a minor anomaly, but it has not yet reached the warning criteria. The usual approach is to log the anomaly in the background and begin to monitor the user's subsequent behavior more closely, but without immediately notifying the user or their contacts.
[0169] Medium Risk (S2): The system determines that there is a clear risk. At this time, an automatic warning will be activated, such as sending a reminder message to the preset emergency contact (children).
[0170] High Risk (S3): The system determines that there is an urgent risk. Strong intervention measures will be initiated immediately, such as the operator's customer service proactively contacting the user and temporarily restricting suspicious operations.
[0171] In summary, traditional property security attack protection solutions rely on static rules based on group profiles and single-dimensional threshold detection, which are difficult to adapt to the huge differences in individual behavioral patterns among the elderly. Moreover, they only provide post-incident blocking at the critical stage of property security attacks, resulting in a strong lag in early warning. In contrast, the solution proposed in this application constructs an individualized multi-dimensional behavioral baseline to accurately depict the unique behavioral patterns of each elderly person. It uses a state transition model to dynamically capture the progressive risk evolution process from a safe state to a high-risk state. By fusing weakly correlated signals such as communication, location, and application through a multi-dimensional abnormal feature matrix, the system can identify risks in the early stages of property security attacks (such as the trust-building stage). Based on reinforcement learning algorithms, it achieves adaptive intervention that is accurately matched with the risk level. This fundamentally solves the technical bottlenecks of existing technologies, such as insensitivity to individual differences, delayed early warning, and limited intervention methods. It achieves a technological leap from group protection to individual protection, and from post-incident blocking to pre-incident early warning.
[0172] The various methods of the embodiments of this application have been described above. Apparatus for implementing the above methods will now be provided.
[0173] Please refer to Figure 3 This application embodiment also provides a user property security risk warning device, including:
[0174] The first acquisition module 31 is used to acquire the user's communication behavior data, location data, and application usage data based on the user's operator network behavior data.
[0175] The first processing module 32 is used to generate target features based on the communication behavior data, the location data, and the application usage data, and to construct an anomaly feature matrix based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features;
[0176] The second processing module 33 is used to generate a first generation result, a second generation result, and a third generation result containing the user's risk status based on the abnormal feature matrix and using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system.
[0177] The third processing module 34 is used to determine the user's predicted status result based on the first generation result, the second generation result and the third generation result, and to execute the corresponding property security warning behavior based on the predicted status result.
[0178] Optionally, the first acquisition module 31 described above is specifically used for:
[0179] Based on the user's carrier network behavior data, obtain the user's call record data;
[0180] The call record data is cleaned to determine the number of calls between the user and the contact, the call duration between the user and the contact, the number of valid call days between the user and the contact, and the maximum number of calls, the maximum call duration, and the maximum number of call days between the user and all contacts.
[0181] The connection strength calculation model between the user and the contact is determined based on the first ratio of the number of calls to the maximum number of calls, the second ratio of the call duration to the maximum call duration, the third ratio of the effective call days to the maximum call days, and the preset weight value corresponding to each ratio.
[0182] Based on the aforementioned connection strength calculation model, data ranked within the first range are selected as the social circle baseline;
[0183] Based on the call record data and the preset seasonal autoregressive integral moving average model, a communication frequency baseline is determined;
[0184] The social circle baseline and the communication frequency baseline are determined as the user's communication behavior data.
[0185] Optionally, the first acquisition module 31 described above is also specifically used for:
[0186] Based on the user's operator network behavior data, obtain the user's user equipment location data and External Data Identification Protocol (XDR) signaling data;
[0187] Based on the user equipment location data, location areas are divided according to the user dwell time, and wireless cells whose dwell time ranks in the second range within each location area are extracted to form location data.
[0188] Based on application usage habits, XDR signaling data and preset application categories are used to construct application feature vectors for corresponding categories as application usage data; the preset application categories are used to indicate financial applications and non-financial applications.
[0189] Optionally, the first processing module 32 described above is specifically used for:
[0190] Based on the user's historical communication frequency data, a time series prediction model is constructed. The communication behavior data is input into the time series prediction model to determine communication dimension anomaly features. The communication dimension anomaly features are used to quantify the degree of deviation between the user's actual communication behavior on the current day and the historical normal pattern.
[0191] The system obtains the user's current actual social group and determines the first call duration between the current actual social group and the contact person; it also parses the communication behavior data to determine the social circle baseline in the communication behavior data and obtains the second call duration between the social circle baseline and the contact person; it determines social dimension anomaly features based on the first call duration and the second call duration; the social dimension anomaly features are used to indicate the degree of social circle connection decay.
[0192] Based on the application usage data, first-time use marker information is obtained, and based on the usage duration deviation corresponding to the current usage duration and the first-time use marker information, the abnormal features of the financial application usage dimension are determined; the abnormal features of the financial application usage dimension are used to indicate the user's financial application usage habits.
[0193] Based on the application usage data and a preset database, abnormal features of the user's device security dimension are obtained; these abnormal features of the device security dimension are used to indicate the user's non-baseline application usage.
[0194] Based on the location data, determine the spatial dimension anomaly characteristics of the user's indicated location change;
[0195] Based on the application usage data, identify the time-dimensional anomalies in user nighttime data activity.
[0196] The abnormal features in the communication dimension, the abnormal features in the social dimension, the abnormal features in the financial application usage dimension, the abnormal features in the device security dimension, the abnormal features in the spatial dimension, and the abnormal features in the time dimension are taken as target features;
[0197] Based on the target features, an anomaly feature matrix with a preset matrix dimension is constructed.
[0198] Optionally, the user risk status is a discrete risk level, including:
[0199] The first security state indicates that the communication behavior data, the location data, and the application usage data all conform to the historical behavior baseline, and all identified abnormal features are below a preset risk threshold.
[0200] The second security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, one or two features are abnormal, and the degree of abnormality does not reach the threshold for triggering a preset intervention standard; the second security state requires continuous monitoring.
[0201] The third security state indicates that among all the abnormal features determined based on the communication behavior data, the location data, and the application usage data, at least three features are abnormal, or there are abnormal features in the financial application usage features or location anomalies; the third security state requires immediate intervention.
[0202] The fourth security state indicates that the user exhibits abnormalities across multiple dimensions of the communication behavior data, the location data, and the application usage data; the fourth security state requires emergency intervention.
[0203] Optionally, the second processing module 33 described above is specifically used for:
[0204] When the preset statistical generative model uses a pre-trained Hidden Markov Model, the abnormal feature matrix is input into the Hidden Markov Model, and the Hidden Markov Model is used to analyze the deviation of the current behavior pattern from the baseline to generate a first generation result; the Hidden Markov Model is trained and learned using a preset normal behavior baseline; the first generation result is used to determine whether it is in the fourth safe state;
[0205] When the deep learning discriminative model uses a pre-trained long short-term memory network sequence classifier, an anomaly feature matrix with preset time features is input into the long short-term memory network sequence classifier to generate a second generation result; the long short-term memory network sequence classifier includes an attention mechanism to assign weights to the anomaly features at different time steps; the second generation result is used to determine whether it is in a third safe state;
[0206] The abnormal feature matrix is input into the knowledge-driven rule system to generate a third generation result; the third generation result is used to determine whether the system is in a second safe state.
[0207] Optionally, the third processing module 34 described above is specifically used for:
[0208] If the first generated result, the second generated result, and the third generated result are all greater than the confidence threshold, the predicted state result of the user is determined by a multi-voting method based on the first generated result, the second generated result, and the third generated result.
[0209] If the predicted state indicates the first security state, the property security early warning action will not be issued;
[0210] If the predicted state result indicates the second security state, the property security early warning action will be logged.
[0211] If the predicted state indicates the third security state, the property security early warning action will automatically execute an early warning operation.
[0212] When the predicted state indicates the fourth security state, the property security early warning action adopts an active intervention strategy.
[0213] It should be noted that the device in this embodiment corresponds to the device used in the above-described method for early warning of user property security risks. The implementation methods in the above embodiments are all applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this application embodiment can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Therefore, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail here.
[0214] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described user property security risk warning method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may include, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0215] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described user property security risk warning method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0216] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.
[0217] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0218] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0219] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for early warning of user property security risks, characterized in that, include: Based on the user's carrier network behavior data, obtain the user's communication behavior data, location data, and application usage data; Based on the communication behavior data, the location data, and the application usage data, target features are generated, and an anomaly feature matrix is constructed based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features; Based on the anomaly feature matrix, using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system, a first generation result, a second generation result, and a third generation result containing the user's risk state are generated, respectively. This includes: when the preset statistical generative model uses a pre-trained Hidden Markov Model (HMM), the anomaly feature matrix is input into the HMM, and the HMM is used to analyze the deviation of the current behavior pattern from the baseline to generate the first generation result; the HMM is trained using a preset normal behavior baseline; the first generation result is used to determine whether the user is in a fourth safe state; when the deep learning discriminative model uses a pre-trained Long Short-Term Memory (LSTM) sequence classifier, an anomaly feature matrix with added preset time features is input into the LSM sequence classifier to generate the second generation result; the LSM sequence classifier includes an attention mechanism to assign weights to anomaly features at different time steps; the second generation result is used to determine whether the user is in a third safe state; the anomaly feature matrix is input into the knowledge-driven rule system to generate the third generation result; the third generation result is used to determine whether the user is in a second safe state. Based on the first generation result, the second generation result, and the third generation result, the predicted status result of the user is determined, and the corresponding property security warning behavior is executed based on the predicted status result; The user risk status is a discrete risk level, including: a first security status, indicating that the communication behavior data, location data, and application usage data all conform to historical behavior baselines, and all identified abnormal features are below a preset risk threshold; a second security status, indicating that among all abnormal features determined based on the communication behavior data, location data, and application usage data, one or two features are abnormal, and the degree of abnormality does not reach the triggering preset intervention standard; the second security status requires continuous monitoring; a third security status, indicating that among all abnormal features determined based on the communication behavior data, location data, and application usage data, at least three features are abnormal, or there are abnormal financial application usage features or location anomaly features; the third security status requires immediate intervention; and a fourth security status, indicating that the user exhibits abnormalities in multiple dimensions of the communication behavior data, location data, and application usage data; the fourth security status requires emergency intervention.
2. The method according to claim 1, characterized in that, Based on the user's carrier network behavior data, the user's communication behavior data is obtained, including: Based on the user's carrier network behavior data, obtain the user's call record data; The call record data is cleaned to determine the number of calls between the user and the contact, the call duration between the user and the contact, the number of valid call days between the user and the contact, and the maximum number of calls, the maximum call duration, and the maximum number of call days between the user and all contacts. The connection strength calculation model between the user and the contact is determined based on the first ratio of the number of calls to the maximum number of calls, the second ratio of the call duration to the maximum call duration, the third ratio of the effective call days to the maximum call days, and the preset weight value corresponding to each ratio. Based on the aforementioned connection strength calculation model, data ranked within the first range are selected as the social circle baseline; Based on the call record data and the preset seasonal autoregressive integral moving average model, a communication frequency baseline is determined; The social circle baseline and the communication frequency baseline are determined as the user's communication behavior data.
3. The method according to claim 1, characterized in that, Based on the user's carrier network behavior data, the user's location data and application usage data are obtained, including: Based on the user's operator network behavior data, obtain the user's user equipment location data and External Data Identification Protocol (XDR) signaling data; Based on the user equipment location data, location areas are divided according to the user dwell time, and wireless cells whose dwell time ranks in the second range within each location area are extracted to form location data. Based on application usage habits, XDR signaling data and preset application categories are used to construct application feature vectors for corresponding categories as application usage data; the preset application categories are used to indicate financial applications and non-financial applications.
4. The method according to claim 1, characterized in that, Based on the communication behavior data, the location data, and the application usage data, target features are generated, and based on the target features, an anomaly feature matrix is constructed, including: Based on the user's historical communication frequency data, a time series prediction model is constructed. The communication behavior data is input into the time series prediction model to determine communication dimension anomaly features. The communication dimension anomaly features are used to quantify the degree of deviation between the user's actual communication behavior on the current day and the historical normal pattern. The system obtains the user's current actual social group and determines the first call duration between the current actual social group and the contact person; it also parses the communication behavior data to determine the social circle baseline in the communication behavior data and obtains the second call duration between the social circle baseline and the contact person; it determines social dimension anomaly features based on the first call duration and the second call duration; the social dimension anomaly features are used to indicate the degree of social circle connection decay. Based on the application usage data, first-time use marker information is obtained, and based on the usage duration deviation corresponding to the current usage duration and the first-time use marker information, the abnormal features of the financial application usage dimension are determined; the abnormal features of the financial application usage dimension are used to indicate the user's financial application usage habits. Based on the application usage data and a preset database, abnormal features of the user's device security dimension are obtained; these abnormal features of the device security dimension are used to indicate the user's non-baseline application usage. Based on the location data, determine the spatial dimension anomaly characteristics of the user's indicated location change; Based on the application usage data, identify the time-dimensional anomalies in user nighttime data activity. The abnormal features in the communication dimension, the abnormal features in the social dimension, the abnormal features in the financial application usage dimension, the abnormal features in the device security dimension, the abnormal features in the spatial dimension, and the abnormal features in the time dimension are taken as target features; Based on the target features, an anomaly feature matrix with a preset matrix dimension is constructed.
5. The method according to claim 1, characterized in that, Based on the first generation result, the second generation result, and the third generation result, the predicted status result of the user is determined, and corresponding property security early warning actions are executed based on the predicted status result, including: If the first generated result, the second generated result, and the third generated result are all greater than the confidence threshold, the predicted state result of the user is determined by a multi-voting method based on the first generated result, the second generated result, and the third generated result. If the predicted state indicates the first security state, the property security early warning action will not be issued; If the predicted state result indicates the second security state, the property security early warning action will be logged. If the predicted state indicates the third security state, the property security early warning action will automatically execute an early warning operation. When the predicted state indicates the fourth security state, the property security early warning action adopts an active intervention strategy.
6. A user property security risk early warning device, characterized in that, include: The first acquisition module is used to acquire the user's communication behavior data, location data, and application usage data based on the user's operator network behavior data; The first processing module is used to generate target features based on the communication behavior data, the location data, and the application usage data, and to construct an anomaly feature matrix based on the target features; the target features include communication dimension anomaly features, social dimension anomaly features, financial application usage dimension anomaly features, device security dimension anomaly features, spatial dimension anomaly features, and time dimension anomaly features; The second processing module is used to generate a first generation result, a second generation result, and a third generation result containing the user's risk status based on the abnormal feature matrix using a preset statistical generative model, a deep learning discriminative model, and a knowledge-driven rule system. The user risk status is a discrete risk level, including: a first safety status, indicating that the communication behavior data, location data, and application usage data all conform to historical behavior baselines, and all identified abnormal features are below a preset risk threshold; a second safety status, indicating that among all abnormal features determined based on the communication behavior data, location data, and application usage data, one or two features are abnormal, and the degree of abnormality does not reach the triggering preset intervention standard; the second safety status requires continuous monitoring; a third safety status, indicating that among all abnormal features determined based on the communication behavior data, location data, and application usage data, at least three features are abnormal, or there are abnormal financial application usage features or location anomaly features; the third safety status requires immediate intervention; and a fourth safety status, indicating that the user exhibits abnormalities in multiple dimensions of the communication behavior data, location data, and application usage data; the fourth safety status requires emergency intervention. The third processing module is used to determine the user's predicted status result based on the first generation result, the second generation result, and the third generation result, and to execute the corresponding property security warning behavior based on the predicted status result; The second processing module is specifically used for: When the preset statistical generative model uses a pre-trained Hidden Markov Model, the abnormal feature matrix is input into the Hidden Markov Model, and the Hidden Markov Model is used to analyze the deviation of the current behavior pattern from the baseline to generate a first generation result; the Hidden Markov Model is trained and learned using a preset normal behavior baseline; the first generation result is used to determine whether it is in the fourth safe state; When the deep learning discriminative model uses a pre-trained long short-term memory network sequence classifier, an anomaly feature matrix with preset time features is input into the long short-term memory network sequence classifier to generate a second generation result; the long short-term memory network sequence classifier includes an attention mechanism to assign weights to the anomaly features at different time steps; the second generation result is used to determine whether it is in a third safe state; The abnormal feature matrix is input into the knowledge-driven rule system to generate a third generation result; the third generation result is used to determine whether the system is in a second safe state.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 5.
8. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 5.