A method and system for tracing abnormal rights data of a communication behavior portrait
By generating a unified event sequence and constructing a communication behavior profile, the problem of data closure between communication links and business links in rights and interests business was solved, realizing the interpretability and verifiability of rights and interests data anomaly tracing, and improving the accuracy of risk assessment and the reliability of the tracing process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN ZHONGCHENG XINFU COMM TECH CO LTD
- Filing Date
- 2026-03-04
- Publication Date
- 2026-06-02
AI Technical Summary
In membership benefits business, it is difficult to form a complete closed loop of event data between communication links and business links, which makes it challenging to analyze the causes of abnormal events, and the credibility and completeness of the evidence chain are difficult to fully verify when collaborating across departments.
By acquiring the event stream of rights and interests business and the event stream of communication links, we perform field standardization, subject association and time sequence alignment to generate a unified event sequence, build a communication behavior profile, calculate the integrity score of the verification code link, identify the attack chain stage, and perform graded loss prevention and control to generate a verifiable summary.
It establishes a unified data foundation for cross-domain analysis, possesses time sensitivity and multi-dimensional characterization capabilities, supports the interpretability of risk assessment and the traceability of causal relationships, and ensures the real-time, auditable, and non-repudiation nature of the tracing process.
Smart Images

Figure CN122133202A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data governance technology, and in particular to a method and system for tracing the source of abnormal rights data based on integrated communication behavior profiling. Background Technology
[0002] In business scenarios involving membership benefits, points, coupons, subsidy eligibility, and virtual gift cards, the redemption, binding, modification, and cancellation of these benefits generally rely on communication links such as SMS verification codes, notification SMS, telephone verification, and carrier authentication to complete the identity verification and transaction confirmation process. Related systems collect user communication behavior data and business operation logs to build risk control models to identify abnormal behavior. They primarily use rule engines or machine learning models to monitor and judge indicators such as login behavior, cancellation behavior, and verification code request frequency, generating risk scores and triggering corresponding handling processes. These systems typically store communication link logs and business operation logs in different data domains, establishing event correlations through timestamp alignment to achieve cross-domain monitoring and risk assessment of user behavior.
[0003] However, in the process of handling abnormal events in rights and interests data, it is difficult for the event data in the communication link and the business link to form a complete closed loop, which makes it challenging to analyze the causes of abnormal events, and the credibility and completeness of the evidence chain are difficult to be fully verified when collaborating across departments. Summary of the Invention
[0004] This application provides a method and system for tracing the source of abnormal rights and interests data in converged communication behavior profiling to solve the above problems.
[0005] Firstly, this application provides a method for tracing the source of abnormal rights data in a converged communication behavior profile, the method comprising: S1. Obtain the rights and interests service event stream and communication link event stream related to the target entity, and perform field standardization, entity association, and time sequence alignment on the rights and interests service event stream and the communication link event stream to generate a unified event sequence; S2. Within a preset sliding time window, extract number lifecycle features, SMS link features, call forwarding / call transfer features, and behavioral response features based on the unified event sequence to construct a communication behavior profile of the target entity; S3. Calculate a verification code link integrity score based on the communication behavior profile. The verification code link integrity score is used to characterize the possibility that the verification code link is controlled by a third party, and outputs the verification code link... S4. Using the key event as an anchor point, based on the CAPTCHA link integrity score, the evidence factor, and the sequential logical constraints of the unified event sequence, construct a causal relationship and identify the attack chain stage, and backtrack to generate the attack chain path corresponding to the key event; S5. When the CAPTCHA link integrity score or the attack chain stage meets the preset triggering conditions, execute the graded loss prevention and control action, and simultaneously generate an evidence package set containing an event summary, a profile summary, and a control summary. Construct a verifiable summary for the evidence package set in chronological order and output the tracing result corresponding to the attack chain path.
[0006] Through the above technical solution, the event flow of rights and interests business and the event flow of communication links are obtained and the fields are standardized, the subject is associated and the time sequence is aligned to generate a unified event sequence, providing a consistent data foundation for cross-domain analysis. By using a sliding time window mechanism, four types of features are extracted from this sequence: number life cycle, SMS link, call forwarding / call transfer and behavioral response, to construct a communication behavior profile with time sensitivity and multi-dimensional representation capabilities. Based on the profile, the integrity score of the verification code link is calculated and a high-contribution evidence factor is output to make the risk assessment interpretable. Using key rights and interests events as anchors, a causal relationship graph is constructed by combining the score, evidence factor and business logic constraints and backtracking to identify the attack chain stages and paths, realizing a structured attribution from phenomenon to essence. When the risk reaches the threshold, graded stop loss and evidence package solidification are executed simultaneously, and a verifiable digest is constructed through a hash chain to ensure that the entire tracing process is real-time, auditable and non-repudiable, thereby effectively solving the problem in the existing technology that the verification code link is attacked but the tracing cannot be closed.
[0007] Optionally, generating a unified event sequence includes: mapping the rights and interests service event stream and the communication link event stream to unified event objects, wherein the unified event object includes a subject identifier, an event type, an event timestamp, and a parameter summary field; wherein the parameter summary field is obtained by desensitizing preset key fields in the event parameters, concatenating them, and hashing them; aggregating the unified event objects based on the subject identifier, and sorting them according to the event timestamp within the same subject; when two adjacent unified event objects have the same event type, the same parameter summary field, and an event timestamp interval less than a preset deduplication threshold, performing deduplication and merging to obtain the unified event sequence.
[0008] Through the above technical solution, the event flow of rights and interests business and the event flow of communication links are mapped into a unified event object containing subject identifier, event type, event timestamp and parameter summary fields, realizing the data semantic alignment of heterogeneous events; the parameter summary field is used to perform de-identified hashing of key parameters, which supports content consistency judgment while ensuring privacy compliance; through subject identifier aggregation and time sequence sorting, an individual-level traceable event flow is constructed; and deduplication and merging are performed based on the triple constraints of event type, parameter summary and time proximity, which significantly improves the accuracy and lightweight nature of the unified event sequence, thereby providing a stable, clean and auditable data foundation for the field standardization, subject association and time sequence alignment required by step S1 in this application.
[0009] Optionally, constructing the communication behavior profile of the target subject includes: constructing an event capture interval centered on the event timestamp of the key rights event, including a forward window and a backward window, and performing bucket statistics on communication link events within the event capture interval according to a preset statistical granularity; wherein the statistical granularity is at least one of second-level, minute-level, or hour-level; concatenating the counting features, latency features, and switching features within each bucket to form a profile vector, and attaching a missing mask to the profile vector to identify missing dimensions.
[0010] The above technical solution achieves precise anchoring of the spatiotemporal context of the attack chain by constructing forward and backward event extraction intervals centered on the timestamps of key rights events. It also addresses the representation needs of both short-term mutations and long-term stability by supporting a bucketing mechanism with at least one statistical granularity at the second, minute, and hour levels. Furthermore, it ensures robustness and interpretability of the profile even in scenarios with incomplete data by concatenating the counting features, latency features, and switching features within each bucket into a structured profile vector and adding a missing mask. Ultimately, this supports the communication behavior profile construction task defined in step S2 of this application, providing a high-quality, multi-granular, confidence-labeled behavioral input foundation for CAPTCHA link integrity scoring and attack chain stage identification.
[0011] Optionally, the number lifecycle characteristics include at least two of the following: number of card replacement / supplementation status changes, number portability flag, number of suspension / reactivation status changes, and number of roaming status changes; the number lifecycle characteristics include a mutation intensity value, which is calculated based on the number of status changes within a preset time window before and after the key rights event and the time interval from the key rights event according to a preset decay function.
[0012] Through the above technical solution, at least two of the following are considered as the basic components of the number lifecycle characteristics: the number of SIM card replacement / replacement status changes, number portability markers, number of suspension / reactivation status changes, and number of roaming status changes. Furthermore, a time-weighted indicator, the mutation intensity value, is introduced to achieve a dual representation of the stability changes in number control: on the one hand, discrete counting is used to capture the frequency of status changes, and on the other hand, a decay function is used to model the influence weight of time proximity on risk assessment. Based on this, the mutation intensity value, as a key component of the communication behavior profile in step S2, provides an interpretable and quantifiable input basis for the status mutation sub-item of the verification code link integrity score in step S3. This supports the system in identifying the precursor signals of SIM card hijacking attacks before critical events occur, improving the attribution accuracy and timeliness of the overall tracing path.
[0013] Optionally, the SMS link characteristics include at least two of the following: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters; wherein, the receipt delay distribution parameters include the mean and quantiles; the SMS link characteristics include routing anomalies, which are obtained by calculating the difference between the receipt delay distribution of the current window and the historical baseline distribution, and the difference includes at least one of KL divergence, JS divergence, or quantile difference.
[0014] Through the above technical solution, this application uses the code sending request frequency, channel / route identifier switching times, and receipt delay distribution parameters to collaboratively characterize the behavioral stability and path consistency of the SMS link. It also uses route outliers to quantify the systematic deviation of the current distribution from the historical baseline, enabling the verification code link integrity score to accurately capture the link loss of control risk caused by channel detours, route hijacking, or network layer interference. This supports the credible assessment of the possibility of third-party control in step S3 and provides an interpretable source of evidence for building causal relationships in step S4.
[0015] Optionally, the call forwarding / call transfer feature includes at least two of the following: the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of call forwarding; the call forwarding / call transfer feature includes a time coupling value, which is calculated based on the number of times the call forwarding status change event and the code request event co-occur within a preset coupling window and their time interval.
[0016] By employing the aforementioned technical solution, at least two of the following basic features—the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of the call forwarding—are used as fundamental features. Furthermore, a temporal coupling value is introduced to characterize the temporal coordination relationship between these features and the code issuance request event. This enables fine-grained modeling of composite attack behaviors involving call forwarding and CAPTCHA interception. Based on this, the temporal coupling value, as a newly added limiting feature, forms a hierarchical expression with the aforementioned three basic features. The basic features reflect the degree of abnormality in the call forwarding behavior itself, while the temporal coupling value focuses on its temporal logical association with key rights actions. The synergistic effect of these two features allows the CAPTCHA link integrity score to more accurately distinguish between normal business call forwarding (such as do-not-disturb scenarios in office settings) and malicious link takeover (such as targeted interception after SIM card hijacking). This improves the accuracy of the scoring in step S3 and provides interpretable and verifiable key support for attack chain stage identification (especially the A→B stage transition) in step S4.
[0017] Optionally, the verification code link integrity score is obtained by: normalizing and then weighting the status mutation sub-items, route suspicion sub-items, call forwarding suspicion sub-items, response anomaly sub-items, and consistency conflict sub-items, where the weight of each sub-item is a preset weight or an adaptive weight obtained based on historical normal sample statistics; the consistency conflict sub-items are calculated based on the degree of conflict between the number lifecycle mutation intensity value and the device fingerprint stability value, where the device fingerprint stability value is determined by at least two of the following: whether the device identifier has changed, whether the IP / ASN has changed, and whether the geographical area identifier has changed; the output of the evidence factors includes: sorting the contribution of each sub-item corresponding to the verification code link integrity score, selecting the top N items with the largest contribution as the evidence factor set, where the size of N is proportional to the total data size of the sub-items; generating a verifiable evidence fragment index for each evidence factor, where the evidence fragment index is used to indicate the corresponding event type, event timestamp interval, and feature name involved in the calculation.
[0018] Through the above technical solution, the sub-items of state mutation, suspicious routing, suspicious call forwarding, abnormal response, and consistency conflict are normalized and weighted, achieving a multi-dimensional comprehensive assessment of the controllability risk of the verification code link. By using the device fingerprint stability value composed of device identifier, IP / ASN, and geographic region identifier, and conflict modeling with the number lifecycle mutation intensity value, the scoring ability in the scenario of person-number separation attack is enhanced. By quantifying and ranking the contribution of each sub-item and extracting the top N items to form an evidence factor set, the causes of high risk become explicit, sortable, and attributable. Furthermore, a structured index containing event type, timestamp interval, and feature name is generated for each evidence factor to ensure that the tracing conclusions are verifiable and verifiable. Thus, without introducing new inventive points, it fully supports the technical functions defined in step S3 of this application, effectively solving the problems of lack of interpretability and difficulty in supporting accountability and review in the background art.
[0019] Optionally, the construction of causal relationships includes: calculating the causal relationship strength for event pairs in a unified event sequence, wherein the causal relationship strength is determined by weighted factors such as temporal proximity, subject consistency, stage template sequence constraints, and evidence factor support; constructing a directed graph with events as nodes and causal relationship strength as edge weights, and performing a reverse search from largest to smallest edge weights, starting from the node corresponding to the key event of the rights and interests, to output the attack chain path; wherein the search is terminated when the cumulative edge weight of the path is lower than a preset pruning threshold.
[0020] The above technical solution characterizes the physical temporal rationality of events through time proximity, anchors the scope of real attack subjects through subject consistency, filters illegal jumps by embedding domain knowledge through stage template sequence constraints, and empowers causal reasoning by using evidence factor support to reverse the risk scoring results. On this basis, taking key rights events as the starting point for tracing, the solution organizes discrete events into attack chain paths with logical coherence and evidence support by using weighted directed graph modeling and greedy reverse search mechanisms. Finally, by using a pruning threshold mechanism, the solution effectively controls computational overhead while ensuring attribution accuracy, enabling the solution to complete the real-time identification and visualization output of complex multi-hop attack chains within millisecond response time.
[0021] Optionally, the step of executing tiered loss mitigation actions and simultaneously generating an evidence package set includes: performing soft mitigation when the verification code link integrity score exceeds a first threshold, wherein the soft mitigation includes at least one of secondary verification, frequency and limit limiting, and delayed activation; performing hard mitigation when the attack chain stage confidence exceeds a second threshold and the key event of the rights is any of the events of rebinding / retrieval / cancellation, wherein the hard mitigation includes at least one of freezing rights, locking sensitive account operations, isolating cancellation, and isolating SMS channels; generating an evidence package for each event that triggers the mitigation, wherein the evidence package contains an event summary, a profile summary, and a mitigation summary; wherein the profile summary contains the hash value of the profile vector, a missing mask, and a set of evidence factors, and constructs a hash chain for the evidence package summary in chronological order and outputs the final hash as a verifiable summary.
[0022] By employing the aforementioned technical solutions, and using CAPTCHA link integrity scoring and attack chain stage confidence as dual criteria to drive decision-making, differentiated responses to events with different risk levels are achieved. Through the coordinated configuration of soft and hard handling, excessive blocking that interferes with normal users is avoided, while ensuring immediate blocking of high-risk attack chains. By generating structured evidence packages for each action and constructing hash chains from their digests, a closed loop is formed from risk identification, action execution, to evidence solidification. The final hash, as a verifiable digest, provides irrefutable technical evidence for judicial evidence, internal audits, and cross-agency investigations.
[0023] Secondly, this application provides a system comprising: a sequence generation module, used to acquire a rights and interests service event stream and a communication link event stream related to a target entity, and to perform field standardization, entity association, and time sequence alignment on the rights and interests service event stream and the communication link event stream to generate a unified event sequence; a profiling module, used to extract number lifecycle features, SMS link features, call forwarding / call transfer features, and behavioral response features based on the unified event sequence within a preset sliding time window to construct a communication behavior profile of the target entity; and an evidence factor query module, used to calculate a verification code link integrity score based on the communication behavior profile, wherein the verification code link integrity score is used to characterize the possibility that the verification code link can be controlled by a third party. The system is configured to: 1) determine the probability of success of the verification code link integrity score and output at least one evidence factor that contributes most to the score; 2) analyze the attack path, using key events as anchors, constructing causal relationships and identifying attack chain stages based on the verification code link integrity score, the evidence factors, and the sequential logical constraints of the unified event sequence; and 3) generate an attack chain path corresponding to the key events, backtracking to generate the attack chain path corresponding to the key events. 4) trace the source, executing tiered loss mitigation actions when the verification code link integrity score or the attack chain stage meets preset trigger conditions, and simultaneously generating an evidence package set containing an event summary, a profile summary, and a handling summary. The system then constructs a verifiable summary of the evidence package set in chronological order and outputs the source tracing result corresponding to the attack chain path. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart illustrating a method for tracing abnormal rights data in a converged communication behavior profile, as provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a rights data anomaly tracing system for integrated communication behavior profiling, provided as an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0027] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0028] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0029] Example 1 Figure 1 A flowchart of a method provided in one embodiment of this application is shown below. Figure 1 As shown, the method includes: S1: Obtain the rights and interests business event stream and communication link event stream related to the target entity, perform field standardization, entity association and time sequence alignment on the rights and interests business event stream and communication link event stream, and generate a unified event sequence; Among them, the rights and benefits business event stream can refer to the structured log stream from the rights and benefits business system side, including but not limited to events such as requesting code issuance, verifying verification codes, login / authentication, rights and benefits redemption, rights and benefits binding / unbinding, account information changes, cancellation / redemption, refund / revocation, blacklist / whitelist hits, and policy version numbers; the communication link event stream can refer to the structured log stream from the operator network element, SMS platform, signaling monitoring system, terminal SDK, or base station side, including but not limited to SMS delivery results and receipts, SMS channel / routing identifiers, short-term high-frequency code issuance, number status changes (service suspension and resumption, SIM card replacement, number portability, roaming, SIM card replacement), call forwarding / call forwarding status, abnormal call behavior, coarse-grained changes in terminal base station / region, and changes in communication permissions on the terminal / SDK side. Field standardization can refer to mapping semantically identical but inconsistently named fields from different sources to a unified field name. For example, msisdn, phone_no, and mobile can be uniformly mapped to the subject identifier, event_time, timestamp, and occur_time can be uniformly mapped to the event timestamp, and sms_channel, route_id, and gateway can be uniformly mapped to the SMS channel identifier. Subject association can refer to establishing a cross-domain subject mapping relationship based on a preset association key. The association key includes one or more of the following: mobile phone number, user ID, device fingerprint, rights ID, and verification ID, which is used to merge rights-side and communication-side events into the same logical subject. Time alignment can refer to arranging all events under the same subject in ascending order of event timestamps under a unified time base (such as UTC millisecond-level timestamps), and interpolating or truncating events with inconsistent timestamp precision to ensure that subsequent sliding window calculations are time-comparable. This application, for example, can perform name mapping and type casting on fields in the rights and interests business event stream according to a field standardization rule table, and simultaneously perform corresponding mapping on the communication link event stream to obtain an intermediate event stream with a consistent format. Alternatively, this application can construct an inverted index based on the subject identifier hash value, perform bucket loading on rights and interests events and communication events respectively, and then perform in-memory association aggregation. Furthermore, this application can construct a time-ordered skip list based on event timestamps, receive dual-stream events online, and perform deduplication, sorting, and alignment in real time. Based on any of the above methods, this application obtains a unified event sequence with consistent structure, traceable subjects, and comparable times, providing a basic data view for subsequent profile construction and causal modeling.
[0030] S2: Within a preset sliding time window, extract number lifecycle features, SMS link features, call forwarding / call transfer features, and behavioral response features based on a unified event sequence to construct a communication behavior profile of the target subject; Among them, the number lifecycle feature can refer to a structured indicator that reflects the stability of the lifecycle status of a mobile number in a communication network. It can be at least two of the following: the number of times the SIM card is replaced / replaced, the number portability flag, the number of times the number is suspended and resumed, and the number of times the roaming status changes. In this embodiment, this feature is used to characterize whether the ownership and control of the number have been transferred unexpectedly. Its numerical change and the temporal proximity of key rights events together constitute an important criterion for the controllability of the link. SMS link characteristics can refer to structured indicators that reflect the stability and reliability of SMS transmission paths. These can be at least two of the following: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters. Among them, the receipt delay distribution parameters can be the mean and 90th quantile of the receipt delay, which are used to characterize the central tendency and tail risk of SMS link response delay. In this embodiment, this feature is used to identify abnormal link behaviors such as SMS hijacking, channel pollution, and route detours. Call forwarding / call transfer characteristics can refer to structured indicators that reflect the activation status of the call forwarding function and its coupling relationship with key operations. These can be at least two of the following: the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of the call forwarding. In this embodiment, this characteristic is used to identify the preparatory behavior of attackers to hijack notification voice / SMS messages through call forwarding. Behavioral response characteristics can refer to structured indicators that reflect the user's response pattern to CAPTCHA-type interactive behaviors. These can include the time distribution of CAPTCHA submission after issuance, the number of submission errors, and whether there were any abnormal login / binding attempts before submission. In this embodiment, these characteristics are used to identify abnormal response patterns such as batch submission by automated tools and bypassing manual confirmation. This application can, for example, extract subsequences from a unified event sequence based on a sliding time window centered on the timestamp of key rights events (e.g., from the first 15 minutes to the last 30 minutes), and count SMS delivery events at a second-level statistical granularity, call forwarding switch change events at a minute-level statistical granularity, and call suspension / resumption events at an hour-level statistical granularity. The extracted counting features, delay features, and switching features at each granularity are then concatenated into a fixed-dimensional vector. Alternatively, this application can group all communication link events within the extracted interval by event type, calculate the variance, maximum consecutive interval, and first-to-last time difference of each group's time interval sequence, and add these time-series statistics as supplementary features to the profile vector. Furthermore, this application can attach a missing mask to each dimension of the profile vector; when a certain type of event does not occur within the current window, the corresponding dimension is set to zero, and a 1 is written to the corresponding position of the missing mask to explicitly indicate that the dimension is unreliable. Based on any of the above methods, this application obtains a communication behavior profile with time sensitivity, multi-granularity coverage, and missing robustness, providing interpretable input for subsequent integrity scoring.
[0031] S3: Calculate the CAPTCHA link integrity score based on the communication behavior profile. The CAPTCHA link integrity score is used to characterize the possibility that the CAPTCHA link is controlled by a third party, and outputs at least one evidence factor that contributes the most to the CAPTCHA link integrity score. Among them, the verification code link integrity score can be a quantitative indicator that comprehensively evaluates whether the communication link on which the SMS / voice verification code depends is in a state of user autonomy and control. The higher the score, the greater the possibility that the link is hijacked, forwarded, replaced or bypassed by a third party. Evidence factors can refer to a single feature or combination of features that has the highest local contribution to the calculation result of S_v. They can be directly mapped to the specific event type, timestamp interval and original feature name involved in the calculation in the unified event sequence, in order to support the source tracing interpretation and manual review. This application can, for example, convert each feature in the communication behavior profile into a sub-item score according to a preset mapping relationship: for example, mapping the number of SIM card changes to a status change sub-item score, mapping the number of SMS channel switching times and the quantile difference of receipt delay to a route suspicion sub-item score, mapping the number of call forwarding switch changes and the time interval of code issuance events to a call forwarding suspicion sub-item score, mapping the standard deviation of verification code submission time to a response anomaly sub-item score, and mapping the difference between the strength of number status changes and the stability of device fingerprints to a consistency conflict sub-item score; then, after normalizing each sub-item score, it is weighted and summed according to a preset weight to obtain S_v; this application can also, for example, use an interpretable machine learning model (such as piecewise logistic regression or shallow decision tree) to score the profile vector, and the model output simultaneously provides the local contribution of each feature (such as SHAP value), and select the top N items as evidence factors according to this ranking; further, this application can also encapsulate the S_v calculation process into an auditable function module, and record the input profile, sub-item scores, weight configuration and output results for each call to ensure that the scoring process is reproducible. This application obtains a captcha link integrity score and corresponding evidence factor with interpretability, verifiability and auditability based on any of the above methods, providing risk criteria for subsequent attack chain identification.
[0032] S4: Using key rights and interests events as anchor points, based on the CAPTCHA link integrity score, evidence factors, and the sequential logical constraints of the unified event sequence, construct causal relationships and identify attack chain stages, and backtrack to generate attack chain paths corresponding to key rights and interests events. Among them, the key event of rights and interests can refer to the operation event with high risk attributes in the rights and interests business process, which can be any event among rebinding, retrieval, cancellation, large-amount exchange, and change of sensitive information; in this embodiment, the event is used as the endpoint node of causal reasoning to start the reverse search; Sequential logic constraints can refer to the temporal dependencies and business rationality constraints between stages in a predefined attack chain stage template. For example, changing a SIM card / porting a number must precede issuing a verification code, issuing a verification code must precede submitting a verification code, and submitting a verification code must precede the cancellation of benefits. This constraint is used to prune causal edges that do not conform to business logic. The attack chain phase can refer to dividing a typical CAPTCHA attack process into several semantically clear phases, including but not limited to: A) Communication link controllability preparation (card replacement / number portability / service suspension / call forwarding), B) CAPTCHA acquisition / hijacking (high-frequency code issuance / routing anomaly / receipt anomaly / response anomaly), C) Account control establishment (login / change binding / retrieval / unbinding device), D) Rights operation (claiming / transferring / cancellation / bulk redemption), E) Clearing traces (cancellation / reset / rebinding); This application can, for example, calculate the causal relationship strength between any two event pairs in a unified event sequence. This strength is determined by a weighted average of temporal proximity (reciprocal decay function), subject consistency (matching degree of the same subject identifier), stage template sequence constraints (whether it conforms to the A→B→C→D logical chain), and evidence factor support (whether the event belongs to the event type and time interval involved in the current evidence factor). Alternatively, this application can construct a sparse directed graph by treating all events as nodes and the causal relationship strength that meets a threshold as directed edge weights. Starting from the node corresponding to the key event, a depth-first reverse search is performed along the incoming edge direction according to the edge weight from largest to smallest. The branch is terminated when the cumulative edge weight of the path is lower than a preset pruning threshold. Further, this application can calculate the stage coverage rate (the proportion of stages covered in the path to the total number of stages in the complete template) and the evidence support rate (the proportion of events in the path belonging to the evidence factor index range) for each candidate path, and select the path with the highest comprehensive score as the final attack chain path. Based on any of the above methods, this application obtains an attack chain path with business readability, logical verifiability, and evidence traceability, achieving structured attribution from result to cause.
[0033] S5: When the verification code link integrity score or attack chain stage meets the preset triggering conditions, execute the hierarchical loss prevention and control action, and simultaneously generate an evidence package set containing event summary, profile summary and handling summary. Construct a verifiable summary for the evidence package set in chronological order and output the source tracing result corresponding to the attack chain path. The preset trigger conditions include two categories: the first category is that S_v exceeds the first threshold (e.g., 0.75), and the second category is that the attack chain stage identification reaches stage C (account control establishment) or stage D (equity operation) and the path confidence exceeds the second threshold (e.g., 0.8). Tiered loss prevention measures include two types: soft measures and hard measures. Soft measures can include secondary verification (device binding verification, risk Q&A, delayed effect), frequency and limit limits, and gray-scale blocking. Hard measures can include freezing equity IDs, locking sensitive account operations, isolating write-offs, and isolating SMS channels. An evidence package (E_i) can refer to a structured evidence unit generated for each event that triggers a disposition, including an event summary (standardized fields + anonymized summary of key parameters), a profile summary (hash value of the profile vector, missing mask, and set of evidence factors), and a disposition summary (disposition type, trigger threshold, and policy version number). A verifiable digest can be the end hash value obtained by constructing a hash chain of evidence packages E_i in chronological order, i.e., H_n = Hash(H_{n−1} || E_n), where H_0 is the initial hash constant; this hash chain structure ensures the temporal integrity and content immutability of the evidence package set, and supports subsequent cross-departmental audit verification; For example, this application could trigger a hard-handling action to freeze the rights ID=Q20240601001 and isolate the SMS channel=CMCC-SMS-GW-07 when S_v=0.87 > 0.75, and simultaneously generate the corresponding evidence package E_1; this application could also trigger a hard-handling action to lock sensitive account operations and isolate write-offs when the attack chain path is identified as A→B→C→D and the confidence level is 0.92 > 0.8, and generate the corresponding evidence package E_2; furthermore, this application could also construct a hash chain by ordering E_1 and E_2 in ascending order of timestamps: H_1 = Hash(H_0 || E_1), H_2 = Hash(H_1 || E_2), and finally output H_2 as a verifiable digest, along with a source tracing report, the report content of which includes the attack chain path, the responsibility points at each stage, the evidence factor list and the H_2 value. This application obtains traceability results that are real-time, interpretable, and verifiable based on any of the above methods, supporting risk control closed loop and judicial audit.
[0034] Example 2: In another optional embodiment, this application also provides a process for generating a uniform event sequence, including: Step 1: Map the rights and interests business event flow and the communication link event flow to a unified event object. The unified event object includes a subject identifier, event type, event timestamp, and parameter summary field. The subject identifier can be a logical identifier used to uniquely represent the entity to which the event belongs. It can be at least one of a mobile phone number, user ID, device fingerprint hash value, or rights account ID. It achieves cross-domain alignment in the rights business event flow and the communication link event flow through preset mapping rules. Event type can refer to the enumerated value after normalizing and classifying the event semantics, such as code sending request, SMS receipt, suspension / reactivation change, rights and benefits cancellation, device login, call forwarding switch change. This classification system covers two major areas: rights and benefits services and communication links, and supports template matching required for subsequent causal stage identification. Event timestamps can refer to the standardized time representation of the moment an event occurs, uniformly converted to millisecond-precision timestamps in the UTC+0 time zone to support cross-system timing alignment and sliding window calculations; The parameter summary field can refer to a fixed-length summary value obtained by concatenating and hashing preset key fields in the original event after desensitization processing; The parameter summary field can be obtained by desensitizing the preset key fields in the event parameters, concatenating them, and then taking the hash. Preset key fields include, but are not limited to: mobile phone number, verification code content, SMS content summary, terminal IP address, ASN number, base station coarse-grained location code, device fingerprint fragment, request source APP package name, and risk control strategy version number; De-identification can be performed by performing irreversible hashing on the field value (such as SHA-256), truncation masking (such as keeping only the last 6 bits of MD5), or field-level k-anonymization (such as generalizing the IP address to the C segment); The concatenation operation can connect the desensitized field values with a separator (such as |) to form a string according to a preset field order; Hash can be performed on the concatenated string to perform a one-way hash operation and output a fixed-length digest value (such as a 32-byte hexadecimal string), which is used to support event content consistency comparison without exposing the original parameters; This application obtains the parameter summary field based on any of the above methods, so that it satisfies privacy compliance requirements while retaining the ability to distinguish the content equivalence between events.
[0035] This application obtains the parameter summary field based on any of the above methods, which is used to support subsequent deduplication and merging and evidence fragment index construction.
[0036] Step 2: Aggregate unified event objects based on the subject identifier, and sort them according to the event timestamp within the same subject; Among them, the subject identifier serves as the aggregation key, used to merge events scattered across different log sources into the same logical subject dimension; Aggregation operations can group all event objects with the same subject identifier into the same event set, which constitutes the complete behavioral trajectory of that subject. Event timestamp sorting can be performed within the same main event set, arranging all unified event objects in ascending order based on event timestamps to ensure consistency between timing logic and sliding window extraction; This application obtains a set of events organized by subject and arranged chronologically based on any of the above methods, providing a structured input basis for subsequent profile construction and causal path backtracking.
[0037] Step 3: When two adjacent unified event objects have the same event type, the same parameter summary field, and an event timestamp interval less than the preset deduplication threshold, perform deduplication and merging to obtain a unified event sequence; Among them, "same event type" can mean that the event type field of two identical event objects has completely the same value. Having identical parameter digest fields can mean that the parameter digest field values of two identical event objects are completely equal after a byte-level comparison. The event timestamp interval can be the time difference between the timestamp of the next event and the timestamp of the previous event, in milliseconds. The preset deduplication threshold can be any value among 1000 milliseconds, 5000 milliseconds, or 30000 milliseconds, and its setting is based on the typical communication link event response delay and the sampling period of the rights and interests business system log. Deduplication and merging can either retain the event object with the earlier timestamp and expand its parameter summary field to record the original parameter summary | number of repetitions | first timestamp | last timestamp, or retain only a single representative event and attach a repetition count flag; This application obtains a unified event sequence after deduplication and merging based on any of the above methods, effectively filtering out duplicate event noise caused by system retries, redundant log reporting, and multi-channel concurrent triggering.
[0038] Example 3: In another optional embodiment, this application also provides for constructing a communication behavior profile of the target subject, including: Step 1: Using the event timestamp of the key rights and interests event as the center, construct an event interception interval containing a forward window and a backward window, and perform bucket statistics on the communication link events within the event interception interval according to the preset statistical granularity; Among them, a Key Event (KE) can refer to a rights and interests business event as defined in this application that triggers the initiation of the tracing process by this method, such as a change of binding, retrieval, cancellation, large-amount exchange, or change of sensitive information event, which has a clear event type identifier and event timestamp; The forward window is any continuous time interval between 0.5 hours and 72 hours before the occurrence of the Key Event, used to capture behavioral signals in the attack preparation phase; The backward window is any continuous time interval between 0 minutes and 24 hours after the occurrence of the Key Event, used to capture behavioral signals in the attack implementation and subsequent response phases; The event interception interval is the union of the forward window and the backward window, and its overall time span can be dynamically configured according to the business risk level, for example, extended to ±48 hours in high-risk scenarios and narrowed to ±15 minutes in low-risk scenarios; The preset statistical granularity is at least one of the following: second-level, minute-level, or hour-level. For example, second-level granularity is used when detecting high-frequency code sending behavior, hour-level granularity is used when analyzing the stability of number status, and minute-level granularity is used when evaluating the rhythm of SMS routing. Bucket statistics can refer to dividing the event interception interval into several equal-length buckets according to the selected statistical granularity, and performing aggregation calculations on communication link events falling into each time bucket. This application can, for example, divide a unified event sequence aligned with timestamps into time buckets according to a selected statistical granularity, and within each bucket, count the occurrence number of communication link events, average retrieval delay, maximum delay, minimum delay, number of channel identifier changes, number of route identifier changes, and number of base station area handovers. Alternatively, this application can directly map the original timestamps of the communication link event stream to the corresponding time buckets, and then perform aggregation of events within each bucket on the same dimension. Furthermore, this application can also use a sliding window mechanism to perform overlapping statistics on time buckets of fixed length to enhance sensitivity to short-term abrupt changes. Based on any of the above methods, this application obtains structured communication behavior sampling results covering multiple time scales, supporting the subsequent construction of profile vectors.
[0039] Step 2: The statistical granularity is at least one of the following: second-level, minute-level, or hour-level. Among them, the second level is a time granularity of 1 second to 60 seconds, which is suitable for capturing instantaneous high-frequency behaviors, such as short-term burst code sending requests, millisecond-level switching of call forwarding switches, and concentrated distribution of receipt timeouts; the minute level is a time granularity of 1 minute to 60 minutes, which is suitable for characterizing the rhythm of medium-frequency behaviors, such as SMS channel rotation cycle, changes in receipt delay trends, and roaming status dwell time; the hour level is a time granularity of 1 hour to 24 hours, which is suitable for reflecting the stability of long-term behaviors, such as the persistence of number lifecycle status (service suspension / SIM card replacement), the stability of geographic area identification, and the consistency cycle of device fingerprint and communication behavior; at least one of these indicates that the system supports single-granularity or multi-granularity parallel statistics, and the features output by different granularities can participate in the input of different sub-models, or be weighted and fused into a unified profile vector.
[0040] For example, this application can select the main statistical granularity based on the risk type of key rights events: for cancellation events, use a dual granularity of second-level + minute-level, and for rebinding events, use a dual granularity of minute-level + hour-level; this application can also automatically recommend statistical granularity based on historical baseline volatility—if the standard deviation of SMS delivery delay for a certain number in the past 7 days is greater than the threshold, then increase its second-level statistical weight.
[0041] Step 3: Concatenate the counting features, latency features, and switching features in each bucket to form a profile vector, and add a missing mask to the profile vector to identify the missing dimensions.
[0042] The counting features include the total number of communication link events, the number of code sending requests, the number of successful receipts, the number of call forwarding switch changes, and the number of suspension / reactivation status changes within a specified time bucket; the latency features include the average latency, 90th percentile latency, maximum latency, latency variance, and latency abrupt change amplitude of SMS receipts within a specified time bucket; the handover features include the number of handovers of SMS channel identifiers, SMS route identifiers, base station location coarse-grained identifiers, and IP address ASN segments within a specified time bucket; the concatenation to form a profile vector involves connecting all the counting features, latency features, and handover features extracted from the same time bucket in a preset order into a one-dimensional numerical vector, such as [count_1, latency_1, handover_1, count_2, latency_2, handover_2, …]; this vector can be stacked into a two-dimensional matrix according to the time bucket order, or flattened into a high-dimensional sparse vector; The missing mask is a binary vector with the same dimension as the profile vector. Its elements take the values 1 or 0, which respectively indicate whether there are valid statistical values in the corresponding dimension. When no communication link event occurs in a certain time bucket, or when a certain feature cannot be calculated due to missing data, the mask at the corresponding position is set to 0. This application can, for example, normalize and concatenate the local profile vectors generated by the second-level, minute-level, and hour-level buckets respectively, and then uniformly attach a global missing mask; it can also generate a mask sub-vector separately for each type of feature (count / latency / switching), and finally merge them into a composite mask; furthermore, it can encode the missing mask as part of the profile vector, for example, by filling missing values with -1 and updating the mask bits synchronously. Based on any of the above methods, this application obtains a structured communication behavior representation with missing value awareness, enabling the subsequent model to maintain inference stability even when some fields are missing.
[0043] Example 4: In one possible implementation, this application also provides number lifecycle characteristics including at least two of the following: number of SIM card replacement / replacement status changes, number portability flag, number of suspension / reactivation status changes, and number of roaming status changes; the number lifecycle characteristics include a mutation intensity value, which is calculated based on the number of status changes within a preset time window before and after a key rights event and the time interval from the key rights event according to a preset decay function, including: Step 1: Number lifecycle characteristics include at least two of the following: number of SIM card replacement / replacement status changes, number portability status changes, number of suspension / reactivation status changes, and number of roaming status changes; Among them, the number of SIM card replacement / replacement status changes can refer to the number of events in the unified event sequence where the number corresponding to the subject identifier undergoes SIM card replacement or reissue operations; the number portability flag can refer to the status flag of the number recorded in the communication link event stream as having completed or being executed the number portability process; the number of suspension / reinstatement status changes can refer to the total number of events where the number's service status switches between suspension and resumption; and the number of roaming status changes can refer to the number of events where the number switches between the home network and the non-home network. The above four features are all key state-related indicators that reflect the stability of number control and the reliability of communication links, and together they constitute the behavioral representation basis of the number lifecycle dimension. In this embodiment, each feature exists in the form of discrete counts and is accumulated within a sliding time window; when any feature is missing, the corresponding field is set to null and identified by a subsequent missing mask; This application can, for example, match and count event entries in a unified event sequence whose event types are SIM card replacement application, SIM card activation, number portability acceptance, service suspension, service resumption, entering the roaming area, and leaving the roaming area. Alternatively, it can parse the aforementioned status change events from signaling logs carrying standard status codes (such as the service status code defined in 3GPP TS 23.003) in the communication link event stream and complete the counting. Furthermore, this application can combine the real-time status snapshot returned by the operator's side interface with the incremental update results of the event stream for fusion verification and output the final count value. This application obtains a quantitative expression of the stability change trend of the number's lifecycle based on any of the above methods.
[0044] Step 2: The number lifecycle feature includes a mutation intensity value, which is calculated based on the number of state changes within a preset time window before and after a key equity event and the time interval from the key equity event, according to a preset decay function. Among them, the mutation intensity value is a weighted quantitative indicator used to characterize the degree of concentrated change in the life cycle status of a number near key events of rights and interests; The mutation intensity value can be obtained by recording the timestamp of the key event as t_0, extracting all state change events that meet the above definition within the time window [t_0-Δt, t_0+Δt], assigning a weight to each event based on the time difference |t_i-t_0| between its occurrence time t_i and t_0, using a preset decay function f(⋅), and then multiplying it by the weight coefficient of the state change type corresponding to the event and summing the results. The preset decay function can be the exponential decay function f(d)=e^(-λd), where d=|t_i-t_0| and λ is the decay coefficient, which is used to adjust the intensity of the influence of time proximity on risk contribution. In this embodiment, the role of the mutation intensity value is to transform the original static frequency statistics into a dynamic risk focus indicator, so that the state change behavior closer to the key event of rights and interests receives a higher discrimination weight, thereby more accurately depicting the rhythm of the attacker's concentrated manipulation of the communication link before implementing rights and interests tampering; this indicator is directly used as one of the input elements for constructing the communication behavior profile in step S2, and participates in the calculation of the state mutation sub-item in the verification code link integrity score in step S3. This application can, for example, calculate the weight of a single event by substituting the absolute difference between the timestamp of the state change event and the timestamp of the key rights event into an exponential decay function, and then multiply it by the basic risk coefficient of the event type (e.g., 1.5 for card replacement, 1.2 for number portability, 0.8 for suspension / reactivation, and 0.6 for roaming) and summing the results to obtain the mutation intensity value. Alternatively, this application can map the relative position of the event within the time window to a predefined segmented weight table to obtain the weight of a single event, and then combine this with the risk level of the event type for weighted summation. Furthermore, this application can normalize the time difference d to the [0,1] interval, input a Sigmoid smoothing function g(d)=1 / (1+e^(k(dc))) to generate continuously decaying weights, and then complete the weighted aggregation. This application obtains an enhanced sensitivity expression for abnormal changes in the lifecycle of a number in the time dimension based on any of the above methods.
[0045] Example 5: In one possible implementation, this application further provides SMS link characteristics including at least two of the following: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters; wherein, the receipt delay distribution parameters include the mean and quantiles; the SMS link characteristics include routing outliers, including: Step 1: SMS link characteristics include at least two of the following: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters; Among them, the code sending request frequency can refer to the number of verification code sending requests triggered by the target subject per unit time; the number of SMS channel identifier switching times can refer to the number of times the rights system calls different SMS service providers or internal channel numbers change within a sliding time window; the number of SMS route identifier switching times can refer to the number of times the underlying SMS route path identifier changes due to operator policies, regional scheduling, or load balancing under the same channel; and the receipt delay distribution parameter can refer to the statistical representation of the delay experienced by the network side in returning a successful receipt after the SMS is sent to the terminal, including the mean and quantiles of the delay sequence (e.g., the 25th, 50th, 75th, and 90th percentile values). The frequency of sending codes can be used to reflect whether the target entity has short-term high-frequency probing behavior; the number of times the SMS channel identifier is switched can be used to reflect whether the communication link is actively bypassed by the established risk control channel; the number of times the SMS route identifier is switched can reflect whether there are unexpected jumps in the underlying transmission path; the receipt delay distribution parameter can be used to reflect the stability and controllability of the SMS end-to-end path. This application can, for example, calculate and normalize the timestamp density of code-issuing events in the rights and interests service event stream within a unit time window into a code-issuing request frequency; it can also identify channel switching actions and accumulate the number of switching based on the change in the number of consecutively occurring different SMS channel identifier field values in a unified event sequence; further, it can determine whether a valid switch is constituted based on the number of changes in the SMS routing identifier field value between adjacent events in a unified event sequence, combined with a preset time neighborhood constraint; this application obtains a quantitative characterization of SMS link stability and path controllability based on any of the above methods.
[0046] Step 2: The receipt delay distribution parameters include the mean and quantiles; Among them, the mean is a central tendency indicator obtained by taking the arithmetic mean of all SMS delivery delay values within the current sliding time window; the quantile is the delay value corresponding to the specified cumulative probability position after sorting all SMS delivery delay values within the current sliding time window in ascending order. The mean can be used to characterize the offset trend of the overall link response level; the quantiles can be used to characterize the tail characteristics and dispersion of the delay distribution, especially sensitive to abnormal long-tail delays. This application could, for example, assemble a set L = {l1, l2, ..., l...} of all receipt delay values within the current window. n}, perform sorting and linear interpolation calculations on it to obtain the p-th percentile Q. p (L); This application may also employ the truncated mean method, removing the highest and lowest 5% of time delay samples and then calculating the mean of the remaining samples; further, this application may also fit a gamma distribution or a log-normal distribution to the histogram of the time delay sequence within a sliding window, and extract the distribution parameters as alternative distribution representations; This application obtains a robust characterization of the receipt time delay distribution pattern based on any of the above methods.
[0047] Step 3: SMS link characteristics include routing anomalies. Routing anomalies are obtained by calculating the difference between the current window's receipt delay distribution and the historical baseline distribution. The difference includes at least one of KL divergence, JS divergence, or quantile difference. KL divergence is an information theory metric that measures the asymmetric difference between two probability distributions; JS divergence is a symmetric smoothed version of KL divergence, with a value range of [0,1]; quantile difference can refer to the absolute difference between the corresponding values of the same quantile (such as the 50th and 90th percentiles) in the current distribution and the baseline distribution. KL divergence can be used to reflect the information gain or distortion of the current distribution relative to the baseline distribution; JS divergence can be used to provide a symmetric, bounded, and robust measure of distribution deviation; quantile difference can be used to locate significant shifts in a specific delay segment, avoiding interference from overall distribution fitting errors. For example, this application can model the receipt delays of the current window and the historical baseline window as discrete probability distributions P and Q, respectively, and calculate KL(P∥Q) as the routing outlier; this application can also input P and Q into the JS divergence formula JS(P, Q) = ½KL(P∥M) + ½KL(Q∥M), where M = ½(P + Q), and output the normalized routing outlier; further, this application can also extract the delay values of the current and baseline windows at the 50th, 75th, and 90th percentiles, respectively, and calculate the weighted sum of the three sets of differences as the routing outlier; this application obtains an objective deviation criterion for the stability of SMS routing based on any of the above methods.
[0048] This application characterizes the behavioral stability and path consistency of the SMS link by using the frequency of code issuance requests, the number of channel / route identifier switching, and the receipt delay distribution parameters. It also uses route outliers to quantify the systematic deviation of the current distribution from the historical baseline, enabling the verification code link integrity score to accurately capture the risk of link loss of control caused by channel detours, route hijacking, or network layer interference. This supports the credible assessment of the possibility of third-party control in step S3 and provides an interpretable source of evidence for building causal relationships in step S4.
[0049] Example 6: In an optional embodiment, this application further provides call forwarding / call transfer features including at least two of the following: the number of times the call forwarding switch state changes, the number of times the call forwarding target changes, and the call forwarding duration; the call forwarding / call transfer features include a time coupling value, which is calculated based on the number of times the call forwarding state change event and the code request event co-occur within a preset coupling window and their time interval, including: Step 1: Call forwarding / call transfer characteristics include at least two of the following: the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of call forwarding; The number of times the call forwarding switch status changes refers to the frequency of the call forwarding function being turned on or off in the communication link of the target subject within a preset sliding time window. This number can reflect the intensity of the user's active use of the call forwarding function or abnormal activation behavior. Its technical role is to characterize whether the control of the communication link has changed unexpectedly. In this embodiment, this number is used as an independent input to participate in the subsequent time coupling degree calculation and provides a basic counting basis for the suspicious call forwarding sub-item in the verification code link integrity score. The number of times the call forwarding target changes can refer to the number of times the target number to which the call forwarding is directed changes within the same time window. This number is a stability indicator of the forwarded target in the communication link. Its technical function is to identify whether there is a short-term, high-frequency change of the interception endpoint, thereby helping to determine whether there is a systematic SMS hijacking intention. In this embodiment, this number and the number of switch state changes together constitute a two-dimensional characterization of the activity of call forwarding behavior, which is used to enhance the robustness of identifying covert call forwarding attacks. The duration of call forwarding can refer to the length of time that the call forwarding function is enabled in a single instance. This duration is the time span from when the call forwarding switch is turned on to when it is turned off. Its technical function is to distinguish between temporary business needs (such as meeting do-not-disturb) and long-term link takeover behavior. In this embodiment, this duration, together with the aforementioned two counts, participates in the weighted fusion of suspicious call forwarding sub-items to avoid misjudging occasional and brief call forwarding operations as attack signals.
[0050] Step 2: Call forwarding / call transfer features include a time coupling value, which is calculated based on the number of times the call transfer status change event and the code request event co-occur within a preset coupling window and their time interval. The time coupling value can be a dimensionless indicator used to quantify the tendency of call transfer status change events and code issuance request events to occur together in the time dimension; this indicator does not depend on absolute timestamps, but focuses on the relative distribution relationship between the two in the local time neighborhood. The time coupling value can be obtained as follows: taking each code issuance request event as the anchor point, setting a preset coupling window (e.g., ±300 seconds) before and after it, counting the number of all call transfer status change events that occur within the window, and attenuating and weighting each co-occurring event pair according to its time interval, and finally normalizing and outputting a single value; the technical role of this method is to highlight the high-risk behavior mode of initiating call transfer immediately after code issuance, and suppress noise interference caused by long-distance, isolated events; The temporal coupling value can also be obtained by: constructing a histogram of event pair time differences, dividing the time into multiple time sub-intervals (e.g., [0,10), [10,60), [60,300) seconds) within a preset coupling window, statistically analyzing the co-occurrence frequency of call transfer status change events and code issuance request events within each sub-interval, and summing the results after weighting each sub-interval using a piecewise weighting function; the technical advantage of this method is to differentiate the coupling strength at different time scales, enabling the model to respond to both instant hijacking and pre-set hijacking attack paths; furthermore, the temporal coupling value can also be obtained by: encoding call transfer status change events and code issuance request events as time series points, calculating the peak position and amplitude of their cross-correlation function within the coupling window, and taking the amplitude as a proxy for coupling strength; the technical advantage of this method is to capture the periodic or trend-based collaborative relationship between event sequences, which is suitable for advanced attack scenarios with multiple rounds of tentative call transfer and code issuance interactions.
[0051] This application obtains a structured representation of the temporal correlation between call forwarding behavior and CAPTCHA acquisition behavior based on any of the above methods, supporting the accurate assignment of the suspicious call forwarding sub-item in the CAPTCHA link integrity score.
[0052] For example, this application may detect a call forwarding switch activation operation within 2 minutes before a critical event (such as an account recovery request) occurs, and the time interval between this operation and the code issuance request event is 8 seconds. This is then counted as a high-weighted coupling event. Subsequently, the time interval, event type combination, and subject identifier consistency of this event pair are extracted from the unified event sequence and sent to the time coupling degree calculation module to output 0.87. This value is two standard deviations higher than the historical baseline mean and is judged as a strong coupling signal, triggering a significant increase in the suspicious call forwarding sub-item.
[0053] Example 7: In another optional embodiment, this application also provides that the verification code link integrity score is obtained in the following way: The status mutation sub-item, route suspicion sub-item, call forwarding suspicion sub-item, response anomaly sub-item, and consistency conflict sub-item are normalized and then weighted and summed, wherein the weight of each sub-item is a preset weight or an adaptive weight obtained based on historical normal sample statistics; the consistency conflict sub-item is calculated based on the degree of conflict between the number lifecycle mutation intensity value and the device fingerprint stability value, wherein the device fingerprint stability value is determined by at least two of the following: whether the device identifier has changed, whether the IP / ASN has changed, and whether the geographical area identifier has changed; the output of the evidence factor includes: Step 1: Normalize the status mutation sub-items, route suspicion sub-items, call forwarding suspicion sub-items, response anomaly sub-items, and consistency conflict sub-items, and then sum them by weight. The weight of each sub-item is a preset weight or an adaptive weight obtained based on the statistics of historical normal samples. Among them, the status mutation sub-item can refer to a quantitative indicator reflecting the change in the life cycle status of the communication number of the target subject within a preset time window before and after the critical event of rights and interests. It is the result of aggregating at least two of the following: the number of card replacement / replacement status changes, number portability marking, number of suspension and resumption status changes, and number of roaming status changes, after obtaining the mutation intensity value. The mutation intensity value is calculated according to the number of status changes and the time interval from the critical event of rights and interests using a preset decay function. Its function is to characterize the potential impact of number status disturbance on the controllability of the verification code link and to serve as one of the input dimensions for the verification code link integrity score. Suspicious routing sub-items can refer to quantitative indicators reflecting the degree of abnormality in SMS delivery paths. They are the result of aggregating at least two of the following parameters: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters, after considering routing anomaly values. Routing anomaly values are obtained by calculating the difference between the receipt delay distribution of the current window and the historical baseline distribution. The difference includes at least one of KL divergence, JS divergence, or quantile difference. Its function is to characterize whether the SMS link has been hijacked or bypassed by a third party, thereby weakening the effectiveness of the verification code. Suspicious call forwarding items can refer to quantitative indicators reflecting the temporal coupling relationship between call forwarding behavior and CAPTCHA interaction behavior. It is the result of aggregating at least two of the following: the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of the call forwarding, by a time coupling degree value. The time coupling degree value is calculated based on the number of times the call forwarding status change event and the CAPTCHA sending request event co-occur within a preset coupling window and their time interval. Its function is to characterize whether an attacker has achieved SMS interception through call forwarding control, thus constituting a key link in the risk of CAPTCHA link integrity. The response anomaly sub-item can refer to a quantitative indicator that reflects the deviation of the user's response behavior to the verification code from the normal pattern. It is the result of at least two of the following: the distribution of verification code submission time, the number of submission errors, the retry mode, and the number of abnormal login attempts before submission, after being aggregated by the anomaly score. Its function is to characterize whether the terminal side is in a state of non-personal control, thereby affecting the authenticity of the verification code verification process. The consistency conflict sub-item can refer to a quantitative indicator reflecting the degree of logical contradiction between the number lifecycle mutation intensity value and the device fingerprint stability value. It is the conflict score obtained after determining the conflict between the device fingerprint stability value, which is composed of at least two of the following: whether the device identifier changes, whether the IP / ASN changes, and whether the geographic area identifier changes, and the number lifecycle mutation intensity value. The device fingerprint stability value is used to characterize the continuity and credibility of the terminal-side identity identifier. Its role is to form cross-validation with the number-side mutation behavior. When the two show significant inconsistency (e.g., the number is frequently changed but the device fingerprint is stable for a long time, or the device is frequently changed but the number status does not change), it indicates the risk of person-number separation or device impersonation, thereby enhancing the discrimination robustness of the verification code link integrity score. This application could, for example, perform Min-Max or Z-score normalization on the five sub-items respectively, and then linearly sum them according to preset weights to obtain the CAPTCHA link integrity score; alternatively, it could construct an empirical distribution of each sub-item based on historical normal samples, complete normalization using quantile mapping, and learn the dynamic weights of each sub-item under different business scenarios through a gradient boosting tree model to achieve adaptive weighting; furthermore, it could employ an attention mechanism to perform context-aware weighting on each sub-item, making the score more consistent with the semantic context of the current key rights event. Based on any of the above methods, this application obtains a comprehensive and interpretable quantitative assessment result of whether the CAPTCHA link is controlled by a third party.
[0054] Step 2: The consistency conflict sub-item is calculated based on the degree of conflict between the number lifecycle mutation intensity value and the device fingerprint stability value. The device fingerprint stability value is determined by at least two of the following: whether the device identifier has changed, whether the IP / ASN has changed, and whether the geographic area identifier has changed. Among them, the device identifier can be a field that uniquely identifies the terminal device, which can be any one of Android ID, OAID, IDFA, IMEI hash value or device fingerprint digest. Its function is to provide a long-term stable anchor point for the terminal identity. IP / ASN can refer to the IP address range and autonomous system number to which the terminal accesses the network, and its function is to characterize the macroscopic stability of the network access environment. Geographic region identifiers can refer to the coarse-grained geographical location information of the terminal, which is the geographical unit mapped by the city, provincial administrative division code, or operator LAC / CI code of the base station. Its function is to help determine the consistency of the device's activity space. This application can, for example, combine the changes of any two of the three elements—device identifier, IP / ASN, and geographic region identifier—within a preset time window before and after a critical event to generate a conflict pattern code (e.g., 011 represents unchanged device identifier, changed IP / ASN, and changed geographic region), and then match a preset conflict intensity level by looking up a table. Alternatively, this application can calculate the frequency of device identifier changes, IP / ASN switching entropy, and geographic region transition distance separately, weight and fuse them, and perform a Pearson correlation test with the number lifecycle mutation intensity value, using the absolute value of the negative correlation coefficient as the conflict degree score. Furthermore, this application can construct a dual-stream contrastive neural network to encode the number lifecycle sequence and the device fingerprint sequence separately, outputting the semantic distance between the two as the conflict degree. Based on any of the above methods, this application obtains a quantitative conflict criterion for whether the relationship between the person, number, and device is consistent, supporting the generation of consistency conflict sub-items.
[0055] Step 3: Sort the contribution of each sub-item corresponding to the CAPTCHA link integrity score, and select the top N items with the largest contribution as the evidence factor set. The size of N is proportional to the total data size of the sub-items. The contribution of each sub-item can refer to the marginal influence of each sub-item on the final score during the weighted summation of the CAPTCHA link integrity score. It is the absolute value of the product of the original value of each sub-item and its corresponding weight, or the attribution score calculated using Shapley value, LIME local interpretation, or gradient backpropagation. Its role is to identify the dominant risk source that leads to an abnormal increase in the score, and to provide traceable evidence for manual review and responsibility attribution. The evidence factor set can refer to the risk evidence set consisting of the top N sub-items ranked by contribution, where N is a positive integer and its value satisfies N = ⌊α × K⌋, where K is the total number of sub-items participating in the scoring (K=5 in this example), and α is the proportionality coefficient, ranging from 0.4 to 0.8, with 0.6 being an option. Its role is to balance the granularity of interpretation with operability, ensuring that key risks are not overlooked while avoiding redundant interference. This application can, for example, determine the contribution of each sub-item based on the partial derivative value in the weighted summation formula, and then truncate the top N items after sorting them in descending order; this application can also, for example, use the permutation test method to sequentially mask individual sub-items and observe the magnitude of score changes, using the magnitude of change as a contribution index; further, this application can also utilize the feature importance output built into the ensemble model to map to the corresponding sub-items for sorting. Based on any of the above methods, this application obtains a set of verifiable, sortable, and truncationable evidence factors, supporting the implementation in step S3 of outputting at least one evidence factor with the highest contribution to the CAPTCHA link integrity score.
[0056] Step 4: Generate a verifiable evidence fragment index for each evidence factor. The evidence fragment index is used to indicate the corresponding event type, event timestamp range, and feature name involved in the calculation. Among them, the event type can refer to the standard event category defined in the unified event sequence, such as card replacement event, call forwarding switch change event, SMS delivery failure event, verification code submission event, device identification change event, etc. Its function is to clarify the evidence source domain. The event timestamp interval can refer to the time range of the event on which the calculation of the evidence factor depends. It is the time window from Δt1 before the key event to Δt2 after the event. Δt1 and Δt2 are set independently according to the sub-item type (for example, the sub-item of state change corresponds to Δt1=7 days and Δt2=1 hour, and the sub-item of response abnormality corresponds to Δt1=5 minutes and Δt2=30 seconds). Its function is to limit the time validity boundary of the evidence. The feature name involved in the calculation can refer to the underlying feature field actually called in the calculation process of this sub-item, such as the number of card replacements, call forwarding switch status, average receipt delay, standard deviation of verification code submission time, device identifier hash value, etc. Its role is to anchor the technical implementation path of evidence; This application can, for example, concatenate the event type, timestamp range, and feature name into a structured string index (e.g., [card replacement event][2024-03-01T08:00:00Z–2024-03-01T09:00:00Z][number of card replacements]), and attach an MD5 hash value to ensure immutability; this application can also encode the three as independent fields and store them in a key-value pair structure, supporting fast retrieval by any dimension; furthermore, this application can also embed the index content into a lightweight Protobuf message body, sign it, and write it into the evidence package metadata area. Based on any of the above methods, this application obtains a locationable, verifiable, and auditable evidence fragment index, satisfying the collaborative requirements of outputting at least one evidence factor in step S3 and generating an evidence package set containing an event summary, a profile summary, and a disposal summary in step S5.
[0057] Example 8: In another optional embodiment, this application also provides a process for constructing causal relationships, including: Step 1: Calculate the causal relationship strength for event pairs in the unified event sequence. The causal relationship strength is determined by a weighted average of temporal proximity, subject consistency, stage template sequence constraints, and evidence factor support. Among them, temporal proximity refers to the degree of temporal tightness reflected by the time interval between adjacent or near-neighbor events of the same subject in a unified event sequence. Its value decreases monotonically as the difference in event timestamps increases, and it is used to characterize whether a preceding event has the physical feasibility to trigger a subsequent event within a reasonable time window. Subject consistency refers to whether the subject identifiers carried by the two events that constitute an event pair belong to the same user, the same mobile phone number, the same device fingerprint, or the same set of rights IDs. It is used to exclude false cross-subject associations and ensure that causal inferences fall within the actual attacker's behavior trajectory. Stage template sequence constraint refers to the legal logical order relationship between each stage in a predefined attack chain stage template. For example, the communication link controllability preparation stage precedes the verification code acquisition / hijacking stage, and the verification code acquisition / hijacking stage precedes the account control establishment stage. This constraint is embedded in the calculation process in the form of a Boolean mask or a stage transition matrix to filter out pseudo-causal edges that violate the attack evolution law. Evidence factor support can refer to the degree to which at least one event in the current event pair matches the set of evidence factors output in Example 7. Matching methods include, but are not limited to: the event type falling within the range of event types indicated by the evidence factor, the event timestamp falling within the timestamp interval indicated by the evidence factor, and the semantic mapping relationship between the event parameter summary field and the feature name associated with the evidence factor. This support is used to back-inject interpretable risk discrimination results into the causal modeling process, so that event pairs covered by high-contribution evidence can obtain higher association weights.
[0058] This application determines whether a potential causal relationship exists between any two events and quantifies the strength of this relationship by weighting factors such as temporal proximity, subject consistency, stage template sequence constraints, and evidence factor support. It also applies hard logical constraints to event type combinations based on a preset stage transition matrix and dynamically adjusts edge weights by combining a time decay function with subject matching results. Furthermore, it uses evidence factors as attention weights to guide causal edge generation, assigning higher initial correlation scores to event pairs containing high-contribution evidence factors, which are then normalized to represent the final causal correlation strength. Based on any of the above methods, this application obtains a structured representation of the logical dependencies between events in a unified event sequence, supporting interpretable backtracking of subsequent attack chain paths.
[0059] Step 2: Construct a directed graph with events as nodes and causal relationship strength as edge weights, and start from the node corresponding to the key event of the stake, and perform a reverse search according to the edge weight from large to small to output the attack chain path; Among these strategies, using events as nodes can refer to abstracting each standardized event object in a unified event sequence as a vertex in a graph. Each vertex carries its original attribute fields (event type, subject identifier, event timestamp, parameter summary field) and profile context information (such as the hash value of the communication behavior profile P(u,t) to which it belongs). Using causal correlation strength as edge weight can refer to establishing a directed edge between corresponding nodes only when the calculated causal correlation strength between two events is greater than zero, and using this strength value as the weight of the edge. The direction of the edge is determined by the chronological relationship (from the preceding event to the subsequent event). Starting from the node corresponding to the key rights event, performing a reverse search according to the edge weight from largest to smallest can refer to setting the key rights event node as the root node of the graph search, traversing its predecessor nodes in reverse along all incoming edges (i.e., edges pointing to the node), and at each level, prioritizing the expansion of the predecessor nodes connected by the incoming edge with the largest edge weight, thereby forming a high-confidence path tracing back from the result to the source. This strategy ensures that the search process focuses on the combination of antecedent actions most likely to drive the occurrence of rights anomalies, rather than exhaustively listing all historical events.
[0060] This application employs a depth-first search with greedy pruning to perform a reverse search: each time, it selects the path with the largest weight from all incoming edges of the current node and records the cumulative edge weight; when the cumulative edge weight of a path is lower than a preset pruning threshold, the expansion of that branch is immediately terminated. This application also employs a breadth-first search with a priority queue (such as a Dijkstra variant), sorting each node according to its largest incoming edge weight and prioritizing high-weight predecessor paths. Furthermore, this application introduces path length constraints and stage diversity constraints to ensure cumulative edge weight while avoiding the repeated selection of multiple events of the same stage type, thus improving the rationality of the attack chain stage distribution. Based on any of the above methods, this application obtains one or more attack chain paths that conform to logical constraints and evidence support. Each path consists of ordered event nodes and includes explanations of the causal strength and evidence attribution of each jump link.
[0061] Step 3: Terminate the search when the cumulative edge weight of the path is lower than the preset pruning threshold.
[0062] The cumulative edge weight of a path refers to the sum of the causal correlation strengths of all directed edges traversed along the reverse search path, starting from the key event node of the rights and interests. It is used to comprehensively measure the overall credibility level of the entire path. The preset pruning threshold is a scalar parameter pre-configured by the system, with a value range of (0, 1], used to control the balance between search depth and path quality. This threshold can be dynamically adjusted according to the business risk level. For example, it is set to 0.75 for high-sensitivity rights and interests scenarios and 0.55 for regular points redemption scenarios. When the cumulative edge weight of a search path is lower than this threshold, it indicates that the overall support of each jump link in the path is insufficient. Continuing to extend it will lead to unreliable attribution results, so the exploration of this path is terminated.
[0063] This application, for example, sets a preset pruning threshold to 0.6 and updates the cumulative edge weights in real time each time a new node is added to the current path. Once the cumulative value is detected to fall below the threshold, it reverts to the previous node and switches to the next highest weight edge. This application, for example, uses a sliding window-style cumulative mechanism, only performing threshold judgment on the sum of the edge weights of the most recent 5 hops to balance long-range dependencies and strong local causality. Furthermore, this application, for example, combines path stage integrity with composite pruning; that is, when the path has covered the three stages A→B→C but the cumulative edge weight is still below the threshold, the threshold is allowed to be appropriately relaxed to ensure stage closure. This application, based on any of the above methods, achieves automatic truncation of invalid or low-confidence paths, ensuring that the output attack chain path has interpretability, verifiability, and engineering practicality.
[0064] Example 9: In another optional embodiment, this application also provides the execution of tiered stop-loss actions and the simultaneous generation of an evidence package set, including: Step 1: When the verification code link integrity score exceeds the first threshold, soft handling is performed. Soft handling includes at least one of the following: secondary verification, frequency limit, and delayed implementation. Among them, the verification code link integrity score (S_v) is used to characterize the possibility that the verification code link is controlled by a third party. Its numerical range is [0,1] or a continuous real number interval after linear / nonlinear mapping. This score has been defined and calculated in Example 7. Its components include state mutation sub-items, route suspicious sub-items, call forwarding suspicious sub-items, response abnormal sub-items and consistency conflict sub-items. Each sub-item is synthesized after normalization and weighting. The first threshold is a preset constant or a boundary value dynamically determined based on the historical normal sample distribution, such as the 95th percentile of the historical S_v distribution. Secondary verification can be an enhanced verification method based on device fingerprints, biometrics, hardware keys, or a Trusted Execution Environment (TEE); Frequency and quota limits can refer to imposing an upper limit on the number of verification code requests, rights and interests revocations, or sensitive operations performed by the same entity within a unit of time. Delayed effectiveness can mean that the results of rights binding, account rebinding, or cancellation are not immediately effective, but instead enter a TTL buffer period, during which manual review or supplementary verification is accepted. This application, for example, could trigger a secondary verification module to perform device binding status verification on the current user session based on whether the verification code link integrity score exceeds a first threshold; it could also call a rate-limiting strategy engine to limit the SMS channel request queue of the target mobile number based on the judgment result; further, it could write the current rights revocation operation into a delayed execution queue based on the judgment result, and attach automatic rollback logic upon timeout. This application, based on any of the above methods, achieves flexible intervention in communication link behavior that is high-risk but has not yet been confirmed as an attack, balancing risk control effectiveness with user experience continuity.
[0065] Step 2: When the confidence level of the attack chain stage exceeds the second threshold and the key event for the rights and interests is any of the events of rebinding / retrieval / cancellation, hard measures are taken. Hard measures include freezing rights and interests, locking sensitive account operations, isolating cancellation and isolating SMS channels. The attack chain stage confidence is the weighted cumulative value of each edge weight on the reverse search path of causal association strength in Example 8. It reflects the degree to which the attack chain path identified by tracing back from the key event of rights and interests conforms to the preset stage template (A→B→C→D). The confidence is a dimensionless value with a value range of [0,1]. The second threshold is an independently set security decision boundary, for example, a value of 0.6, which is higher than the first threshold and reflects a higher degree of risk confidence; Key events for rights and interests are operational event types with high business sensitivity and high asset transfer risk as defined in Example 1, including but not limited to account rebinding, password retrieval, rights and interests cancellation, large-amount exchange, and modification of sensitive information; such events have been used as anchor points in step S4 of Example 1; Freezing rights can refer to marking the target rights ID (such as coupon code, points account, virtual card number) as unusable, but retaining its ownership and basic attributes; Locking sensitive account operations can refer to prohibiting operations that require strong identity verification, such as changing account binding, retrieving account information, canceling account, unbinding device, and resetting payment password. Isolation revocation can refer to removing the target equity ID from the revocation channel or routing its revocation request to a sandbox environment for isolation processing; Isolating the SMS channel can refer to switching the SMS delivery request of the target mobile number from the main channel to the audit-dedicated channel, or prohibiting it from being sent through specific high-risk routes (such as overseas gateways or virtual operator channels); This application, for example, can call the rights management service interface to freeze the target coupon ID based on the joint judgment result of the attack chain stage confidence level and the type of rights key event, and simultaneously update its status field to FROZEN_BY_RISK; this application can also send a lock command to the account center service based on the joint judgment result, causing the two API interfaces for modifying the bound mobile phone and resetting the login password of the target account to return a unified rejection response; furthermore, this application can also forward all SMS requests for the mobile phone number within the next 30 minutes to the isolation channel based on the joint judgment result, and mark RISK_ISOLATION_FLAG=TRUE in the channel log. This application, based on any of the above methods, can achieve strong intervention against high-risk operations that have strong attack chain evidence, effectively blocking the path of rights asset loss.
[0066] Step 3: Generate an evidence package for each event that triggers the action. The evidence package includes an event summary, a profile summary, and a action summary. The event summary can refer to a standardized description formed by structured extraction of key rights events that trigger the disposal, including event type, subject identifier, event timestamp, source system, and key parameter summary (such as rights ID, write-off amount, mobile phone number before binding, and mobile phone number after binding); the key parameter summary has been defined in Example 2 as the hash obtained by desensitizing and concatenating preset key fields; A profile summary can refer to a lightweight encapsulation of a snapshot of the communication behavior profile relied upon during the execution of the action, including the hash value of the profile vector, a missing mask, and a set of evidence factors. The profile vector, as defined in Example 3, is composed of the counting features, latency features, and switching features obtained from bucket statistics. The missing mask is used to identify dimensions in the profile vector that cannot be calculated due to missing data. The set of evidence factors, as defined in Example 7, is the top N items selected after sorting the contribution of each sub-item of the CAPTCHA link integrity score, with each item corresponding to a verifiable evidence fragment index. The action summary can refer to the semantic record of the stop-loss action executed this time, including the action type (soft action / hard action), specific action (such as performing secondary verification, freezing equity ID=EQ20240615XXXX), triggering conditions (such as S_v=0.83>0.75, confidence level=0.87>0.6 and event type=write-off), strategy version number, and execution timestamp; This application, for example, may invoke an evidence package generator service based on an event instance that triggers the action, serializing the event digest JSON, the profile digest structure, and the action digest key-value pairs into a unified format byte stream, and attaching a digital signature. Alternatively, based on the same event instance, this application may combine the profile vector hash value, the missing mask bitmap, and the evidence factor index list into a profile digest block, and generate its digest using the SHA-256 algorithm. Furthermore, this application may, based on the event instance, parse the frozen rights ID=EQ20240615XXXX in the action digest into a standard resource identifier (URI) and embed it into the evidence package metadata field. This application obtains a clearly structured, complete, and independently parsable minimum evidence unit based on any of the above methods, supporting subsequent auditing and review.
[0067] Step 4: The profile summary includes the hash value of the profile vector, the missing mask, and the set of evidence factors. A hash chain is constructed on the evidence package summary in chronological order, and the final hash is output as a verifiable summary. The hash value of the image vector can be calculated from the complete image vector byte sequence using a deterministic hash algorithm (such as SHA-256) to ensure that the same input always yields the same output. The missing mask can be a binary bitmap with the same length as the dimension of the image vector, where each bit corresponds to a feature dimension. A value of 1 indicates that the data in that dimension is missing, and a value of 0 indicates that the dimension is valid. The set of evidence factors can be a list of evidence factor names arranged in descending order of contribution, with each name strictly corresponding to the sub-item name defined in Implementation 7 (such as the state mutation sub-item and the route suspicion sub-item). An evidence package summary can refer to a summary string generated by standardizing and concatenating the event summary, profile summary, and disposal summary in the evidence package. For example, it can be generated by concatenating a JSON string in a fixed field order and then taking the hash. A hash chain can refer to a singly linked structure formed by sequentially performing the operation H_i = Hash(H_{i−1} || E_i) with the first evidence packet digest as the initial input, where E_i is the i-th evidence packet digest and H_0 is a preset initial vector (such as an all-zero hash). The end hash can refer to the last hash value H_n calculated in the hash chain, which is persistently stored and provided as the unique fingerprint of the entire evidence sequence. This application could, for example, sort all evidence packet digests generated on a given day in ascending order of timestamps, and then sequentially call the hash chain construction module to perform concatenated hash operations. Alternatively, it could encode each evidence packet digest in TLV (Type-Length-Value) format and send it to the hash engine based on the same sorting result. Furthermore, it could, based on the hash chain result, synchronously update the terminal hash field stored in the database after each new E_i, and push an incremental notification to the audit system. Based on any of the above methods, this application obtains an evidence chain expression form with tamper-resistance, verifiability, and temporal integrity, enabling any third party to verify that the evidence packet sequence has not been added, deleted, or tampered with by replaying the hash chain process.
[0068] Figure 2 This application provides a schematic diagram of the structure of a rights data anomaly tracing system for integrated communication behavior profiling, as shown in one embodiment. Figure 2 As shown, the rights and interests data anomaly tracing system 300 of the integrated communication behavior profile in this embodiment includes: sequence generation module 301, profile module 302, evidence factor query module 303, attack path analysis module 304, and tracing module 305.
[0069] The sequence generation module 301 is used to acquire the rights and interests service event stream and the communication link event stream related to the target subject, and to perform field standardization, subject association, and time sequence alignment on the rights and interests service event stream and the communication link event stream to generate a unified event sequence; the profiling module 302 is used to extract number lifecycle features, SMS link features, call forwarding / call transfer features, and behavioral response features based on the unified event sequence within a preset sliding time window to construct a communication behavior profile of the target subject; the evidence factor query module 303 is used to calculate the verification code link integrity score based on the communication behavior profile, the verification code link integrity score is used to characterize the possibility that the verification code link is controlled by a third party, and outputs the verification code link integrity score. The system includes at least one evidence factor that contributes the most to the verification code link integrity score; an attack path analysis module 304, which uses a key event as an anchor point to construct a causal relationship and identify attack chain stages based on the verification code link integrity score, the evidence factor, and the sequential logical constraints of the unified event sequence, and backtracks to generate an attack chain path corresponding to the key event; and a source tracing module 305, which executes a graded loss-stopping action when the verification code link integrity score or the attack chain stage meets a preset trigger condition, and simultaneously generates an evidence package set containing an event summary, a profile summary, and a handling summary, constructs a verifiable summary for the evidence package set in chronological order, and outputs a source tracing result corresponding to the attack chain path.
[0070] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
Claims
1. A method for tracing the source of abnormal rights data based on integrated communication behavior profiling, characterized in that, include: S1. Obtain the rights and interests business event stream and communication link event stream related to the target entity, and perform field standardization, entity association and time sequence alignment on the rights and interests business event stream and the communication link event stream to generate a unified event sequence; S2. Within a preset sliding time window, extract number lifecycle features, SMS link features, call forwarding / call forwarding features, and behavioral response features based on the unified event sequence to construct a communication behavior profile of the target subject; S3. Calculate the CAPTCHA link integrity score based on the communication behavior profile. The CAPTCHA link integrity score is used to characterize the possibility that the CAPTCHA link is controlled by a third party, and output at least one evidence factor that contributes the most to the CAPTCHA link integrity score. S4. Using key rights and interests events as anchor points, based on the verification code link integrity score, the evidence factor, and the sequential logical constraints of the unified event sequence, construct causal relationships and identify attack chain stages, and backtrack to generate attack chain paths corresponding to the key rights and interests events. S5. When the verification code link integrity score or the attack chain stage meets the preset triggering conditions, a graded loss prevention and control action is executed, and an evidence package set containing an event summary, a profile summary and a control summary is generated simultaneously. A verifiable summary is constructed on the evidence package set in chronological order, and the source tracing result corresponding to the attack chain path is output.
2. The method according to claim 1, characterized in that, The generation of the unified event sequence includes: The rights and interests service event stream and the communication link event stream are respectively mapped to a unified event object, and the unified event object includes a subject identifier, event type, event timestamp, and parameter summary field; The parameter summary field is obtained by desensitizing the preset key fields in the event parameters, concatenating them, and then taking their hash. Aggregate unified event objects based on the subject identifier, and sort them according to the event timestamp within the same subject; When two adjacent unified event objects have the same event type, the same parameter summary field, and an event timestamp interval less than a preset deduplication threshold, deduplication and merging are performed to obtain the unified event sequence.
3. The method according to claim 1, characterized in that, The construction of the communication behavior profile of the target subject includes: Centered on the event timestamp of the aforementioned key rights event, an event interception interval containing a forward window and a backward window is constructed, and within the event interception interval, communication link events are bucketed and statistically analyzed according to a preset statistical granularity. The statistical granularity is at least one of the following: second-level, minute-level, or hour-level. The counting features, delay features, and switching features in each bucket are concatenated to form a profile vector, and a missing mask is added to the profile vector to identify the missing dimensions.
4. The method according to claim 1, characterized in that, The number lifecycle characteristics include at least two of the following: number of times the SIM card replacement / supplementation status changes, number portability flag, number of times the suspension / reactivation status changes, and number of times the roaming status changes; The number lifecycle feature includes a mutation intensity value, which is calculated based on the number of state changes within a preset time window before and after the key event and the time interval from the key event, according to a preset decay function.
5. The method according to claim 1, characterized in that, The SMS link characteristics include at least two of the following: code sending request frequency, number of SMS channel identifier switching times, number of SMS route identifier switching times, and receipt delay distribution parameters; The receipt delay distribution parameters include the mean and quantiles; The SMS link features include routing anomalies, which are obtained by calculating the difference between the current window's receipt delay distribution and the historical baseline distribution. The difference includes at least one of KL divergence, JS divergence, or quantile difference.
6. The method according to claim 1, characterized in that, The call forwarding / call transfer characteristics include at least two of the following: the number of times the call forwarding switch status changes, the number of times the call forwarding target changes, and the duration of the call forwarding; The call transfer / call forwarding feature includes a time coupling value, which is calculated based on the number of times the call transfer status change event and the code request event co-occur within a preset coupling window and their time interval.
7. The method according to claim 1, characterized in that, The verification code link integrity score is obtained through the following method: After normalizing the sub-items of state mutation, route suspicion, call forwarding suspicion, response anomaly, and consistency conflict, the sub-items are weighted and summed. The weight of each sub-item is either a preset weight or an adaptive weight obtained based on the statistics of historical normal samples. The consistency conflict sub-item is calculated based on the degree of conflict between the number lifecycle mutation intensity value and the device fingerprint stability value. The device fingerprint stability value is determined by at least two of the following: whether the device identifier changes, whether the IP / ASN changes, and whether the geographical area identifier changes. The output of the evidence factor includes: The contribution of each sub-item corresponding to the verification code link integrity score is sorted, and the top N items with the largest contribution are selected as the evidence factor set. The size of N is proportional to the total data size of the sub-items. For each evidence factor, a verifiable evidence fragment index is generated, which is used to indicate the corresponding event type, event timestamp range, and feature name involved in the calculation.
8. The method according to claim 1, characterized in that, The construction of causal relationships includes: The causal association strength is calculated for event pairs in a unified event sequence. The causal association strength is determined by a weighted average of temporal proximity, subject consistency, stage template sequence constraints, and evidence factor support. A directed graph is constructed using events as nodes and the strength of causal relationships as edge weights. Starting from the nodes corresponding to the key events of the rights and interests, a reverse search is performed according to the edge weights from large to small to output the attack chain path. The search is terminated when the cumulative edge weight of the path is lower than the preset pruning threshold.
9. The method according to claim 1, characterized in that, The execution of tiered loss mitigation actions and the simultaneous generation of an evidence package include: When the verification code link integrity score exceeds the first threshold, soft processing is performed. The soft processing includes at least one of secondary verification, frequency limiting, and delayed activation. When the confidence level of the attack chain stage exceeds the second threshold and the key event of the rights and interests is any one of the events of rebinding / retrieval / cancellation, hard handling is performed. The hard handling includes at least one of freezing rights and interests, locking sensitive account operations, isolating cancellation and isolating SMS channels. For each event that triggers a response, an evidence package is generated, which includes an event summary, a profile summary, and a response summary. The profile summary includes the hash value of the profile vector, the missing mask, and the set of evidence factors. The hash chain of the evidence package summary is constructed in chronological order, and the final hash is output as a verifiable summary.
10. A rights data anomaly tracing system integrating communication behavior profiling, characterized in that, The method applied to any one of claims 1-9 includes: The sequence generation module is used to acquire the rights and interests business event stream and the communication link event stream related to the target subject, and to perform field standardization, subject association and time sequence alignment on the rights and interests business event stream and the communication link event stream to generate a unified event sequence. The profiling module is used to extract number lifecycle features, SMS link features, call forwarding / call transfer features, and behavioral response features based on the unified event sequence within a preset sliding time window, and construct a communication behavior profile of the target subject. The evidence factor query module is used to calculate the verification code link integrity score based on the communication behavior profile. The verification code link integrity score is used to characterize the possibility that the verification code link is controlled by a third party, and outputs at least one evidence factor that contributes the most to the verification code link integrity score. The attack path analysis module is used to construct causal relationships and identify attack chain stages based on the verification code link integrity score, the evidence factor, and the sequential logical constraints of the unified event sequence, using key rights events as anchor points, and to backtrack and generate attack chain paths corresponding to the key rights events. The tracing module is used to perform graded loss prevention and control actions when the verification code link integrity score or the attack chain stage meets the preset trigger conditions, and simultaneously generate an evidence package set containing an event summary, a profile summary and a control summary. The module constructs a verifiable summary for the evidence package set in chronological order and outputs the tracing result corresponding to the attack chain path.