Abnormity early warning method, device and equipment based on stage permeation and medium
By constructing a continuous data chain and multi-dimensional dynamic profiles, a phased penetration model is established, solving the problem of the inability to identify potential risk groups in existing technologies, and realizing accurate analysis and early warning of customer behavior.
Patent Information
- Application Number
- CN202511246078.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing technologies lack progressive, layered penetration analysis of customers at different service stages, making it impossible to establish a phased relationship from interaction to use to anomalies, resulting in difficulty in accurately identifying potential risk groups.
The collected behavioral data of the target audience throughout the entire interaction process forms a continuous data chain, constructs a multi-dimensional dynamic profile, and establishes a phased penetration model. It defines the progressive migration relationship of the group in the service interaction, usage and anomaly occurrence stages, extracts feature identifiers for overlap and difference through penetration analysis, generates an early warning model and outputs the anomaly level.
It enables precise identification of potential risk groups, improves the accuracy and real-time nature of early warnings, and allows for early identification and warning before abnormal risks become apparent.
Smart Images

Figure CN120951097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an anomaly early warning method, apparatus, device, and storage medium based on phased penetration. Background Technology
[0002] In the fintech sector, existing customer complaint early warning technologies largely remain at the reactive stage after a complaint occurs, lacking a systematic pre-emptive warning mechanism. Most methods only conduct macro-level statistical analysis of the overall customer base, failing to delve deeper into customer behavior characteristics at different stages such as browsing, inquiries, transactions, and claims. This crude approach makes it difficult to identify potentially high-risk complaint groups, leading companies to intervene only after a complaint has occurred, missing opportunities for early intervention and service improvement. Furthermore, while some systems attempt to incorporate customer purchase information or basic data for early warning, the data dimensions used are relatively singular, neglecting the interaction data and claims-related data accumulated by customers during service use. This results in an incomplete risk assessment and significant biases in the warning results. Simultaneously, existing technologies, when utilizing historical complaint data, only reach a simple summary and statistical level, failing to delve into the differences between different customer categories or to trace the logical chain from purchase to claims to complaints, thus lacking sufficient scientific support for risk modeling.
[0003] In the healthcare sector, existing technologies also have significant shortcomings in addressing service complaints from patients or users. Most methods simply analyze basic patient registration information and partial medical records as input data, failing to establish a dynamic analysis model covering the entire process from registration, consultation, treatment, follow-up, to service feedback. This results in complaint alerts often being limited to data from a single point in time, failing to reflect the accumulated behavioral characteristics and service experiences of patients across multiple stages. Due to the lack of phased penetration analysis, healthcare institutions struggle to effectively identify high-risk patient groups who gradually accumulate dissatisfaction and may file complaints during follow-up or long-term treatment. Furthermore, existing complaint management systems focus more on past anomalies, failing to progressively model patient groups at different stages and reveal the potential causal relationship between treatment experience and complaint occurrence, leading to insufficient accuracy and targeting of alerts. Particularly in utilizing historical complaint data, existing methods primarily rely on statistical quantity, lacking the decomposition and analysis of complaint probabilities for different characteristic groups, thus limiting the proactiveness and scientific rigor of healthcare services in risk prevention and control. Summary of the Invention
[0004] The main objective of this invention is to provide an anomaly early warning method, device, equipment, and storage medium based on phased penetration, aiming to solve the technical problem of lacking progressive layered penetration analysis of customers at different service stages, failing to establish a phased correlation from interaction to use to anomalies, and thus making it difficult to accurately identify potential risk groups.
[0005] To achieve the above objectives, the present invention provides an anomaly early warning method based on phased penetration, comprising:
[0006] The collected behavioral data of the object throughout the entire interaction process forms a continuous data chain, and a multi-dimensional dynamic profile is constructed based on the continuous data chain;
[0007] Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model;
[0008] Based on the staged penetration model and the multi-dimensional dynamic profile, perform the first-level penetration analysis to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group;
[0009] Based on the staged penetration model and the multi-dimensional dynamic profile, a second-level penetration analysis is performed to extract the feature identification differences between the service usage stage group and the anomaly occurrence stage group.
[0010] An early warning model is generated based on the overlap of the feature identifiers, the differences between the feature identifiers, and historical data.
[0011] The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output.
[0012] An early warning signal is triggered based on the aforementioned anomaly level.
[0013] Furthermore, to achieve the above objectives, the present invention provides an anomaly early warning device based on phased penetration, comprising:
[0014] The data acquisition module is used to collect behavioral data of the object throughout the entire interaction process to form a continuous data chain, and to build a multi-dimensional dynamic profile based on the continuous data chain;
[0015] The model building module is used to establish a phased penetration model and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model.
[0016] The first-level analysis module is used to perform first-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group.
[0017] The second-level analysis module is used to perform second-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the feature identification difference between the service usage stage group and the anomaly occurrence stage group.
[0018] The early warning model generation module is used to generate an early warning model based on the overlap of the feature identifiers, the difference of the feature identifiers, and historical data.
[0019] The anomaly detection module is used to input the current state of the multi-dimensional dynamic profile into the early warning model and output the anomaly level.
[0020] The early warning triggering module is used to trigger an early warning signal based on the anomaly level.
[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an anomaly warning program based on phased penetration stored in the memory and executable on the processor, wherein when the anomaly warning program based on phased penetration is executed by the processor, it implements the steps of the anomaly warning method based on phased penetration as described above.
[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an anomaly warning program based on phased penetration, wherein when the anomaly warning program based on phased penetration is executed by a processor, it implements the steps of the anomaly warning method based on phased penetration as described above.
[0023] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses an anomaly early warning method, device, equipment, and medium based on phased penetration, comprising: collecting behavioral data of an object throughout the entire interaction process to form a continuous data chain, and constructing a multi-dimensional dynamic profile based on the continuous data chain; establishing a phased penetration model, defining the progressive migration relationship between service interaction phase groups, service usage phase groups, and anomaly occurrence phase groups in the phased penetration model; performing a first-level penetration analysis based on the phased penetration model and the multi-dimensional dynamic profile to extract the overlap of feature identifiers between the service interaction phase group and the service usage phase group; performing a second-level penetration analysis based on the phased penetration model and the multi-dimensional dynamic profile to extract the difference in feature identifiers between the service usage phase group and the anomaly occurrence phase group; generating an early warning model based on the feature identifier overlap, feature identifier difference, and historical data; inputting the current state of the multi-dimensional dynamic profile into the early warning model and outputting the anomaly level; and triggering an early warning signal based on the anomaly level. This invention establishes a progressive migration relationship across interaction, usage, and anomaly stages, combined with multi-dimensional dynamic profiling, to extract key indicators from feature overlap and difference. It then utilizes historical data to construct an early warning model, enabling accurate identification of anomaly levels. When the anomaly level reaches a preset threshold, an early warning signal is triggered, allowing for early identification and warning before anomaly risks become apparent, thus improving the accuracy and real-time nature of early warnings. Attached Figure Description
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0025] Figure 1 This is a schematic diagram of an application environment for an anomaly early warning method based on staged penetration in one embodiment of the present invention;
[0026] Figure 2 This is a flowchart illustrating an embodiment of the anomaly early warning method based on phased penetration according to the present invention;
[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly early warning device based on staged penetration of the present invention.
[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0029] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0030] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0031] The anomaly early warning method based on phased penetration provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can collect behavioral data of the user terminal throughout the entire interaction process to form a continuous data chain, and construct a multi-dimensional dynamic profile based on this continuous data chain. A phased penetration model is established, defining the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group. Based on the phased penetration model and the multi-dimensional dynamic profile, a first-level penetration analysis is performed to extract the feature overlap between the service interaction phase group and the service usage phase group. A second-level penetration analysis is performed based on the phased penetration model and the multi-dimensional dynamic profile to extract the feature difference between the service usage phase group and the anomaly occurrence phase group. An early warning model is generated based on the feature overlap, feature difference, and historical data. The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output. An early warning signal is triggered based on the anomaly level. This invention, by establishing a progressive migration relationship between the interaction, usage, and anomaly phases and combining it with a multi-dimensional dynamic profile, achieves the extraction of key indicators from feature overlap and difference, and utilizes historical data to construct an early warning model, enabling accurate identification of anomaly levels. When the anomaly level reaches a preset threshold, an early warning signal is triggered, enabling early identification and warning before the anomaly risk becomes apparent, thus improving the accuracy and real-time nature of the warning. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the anomaly warning method based on phased penetration provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0033] like Figure 2 As shown, the anomaly early warning method based on phased penetration proposed in this invention includes the following steps:
[0034] S10, collect the behavioral data of the object throughout the entire interaction process to form a continuous data chain, and construct a multi-dimensional dynamic profile based on the continuous data chain;
[0035] In this embodiment, the scope and identification of the data collection objects are first clearly defined. The data collection objects include individual users, enterprise accounts, device terminals, channel agents, and organizational entities that can be uniquely identified in specific business scenarios. Identification sources can include user identifiers from the account system, device fingerprints, compliant document hashes, contract numbers, policy numbers, session markers, etc. To avoid the same object being split across multiple channels, an identifier mapping table and relationship graph are established. The mappings from account to device, device to session, and session to transaction are maintained as incrementally updatable association edges, supporting one-to-many and many-to-one relationships. The relationship edges record the credibility score and effective time interval for subsequent conflict decision-making.
[0036] The entire interaction process refers to the sequence of online and offline touchpoints around an object throughout its business lifecycle, not limited to a single business link. It originates from event ontology modeling of the business journey, with typical touchpoints including browsing, consultation, favorites, adding to cart, transaction decision, payment execution, service usage, service feedback, exception reporting, and exception handling. Each touchpoint type defines an event name, event key, event time, processing time, source channel, participants, and context extension fields to ensure consistent coding across systems. To accommodate new touchpoint types, the event ontology uses versioning management, with newly added fields accessed as optional fields without disrupting existing parsing.
[0037] Behavioral data refers to structured and semi-structured records generated at each touchpoint throughout the entire interaction process. Sources include front-end event logs, server-side business logs, customer service interaction records, work order processing records, payment gateway receipts, IoT device reports, and third-party risk control feedback. Each record must contain at least the event time and object key; to support reliable sorting, a monotonically increasing sequence number or causal vector is additionally introduced; when this cannot be directly provided, it is inferred using a gateway access time and channel latency model. Semi-structured text can be segmented, entity extracted, and labeled using a classifier to extract discrete features such as problem category, sentiment, and handling instructions, while retaining the original text index for traceability.
[0038] A continuous data chain is an incremental directed chain that concatenates multi-source behavioral data of an object according to event time and causal order. During construction, alignment, denoising, deduplication, and reordering are performed first. Alignment refers to correcting deviations in time bases across different systems, using time synchronization logs and reference heartbeat streams to verify deviation curves, achieving dual-scale calibration at both daily and millisecond levels. Denoising filters records with format defects, missing fields, abnormal timestamps, duplicate reports, etc., using rule-based validation and anomaly detection models and providing reasons for discarding them. Deduplication involves calculating hash fingerprints for the same object and event key within a sliding time window, retaining the record with the highest confidence within a matching threshold. Reordering rearranges out-of-order records according to event time, employing a delayed merging strategy for severely late records, filling them back into the chain and marking them as filled. The chain's data structure uses a combination of append-only logs and segment indexes, providing both sequential scanning and random access by event key; each append generates a chain version number for incremental recalculation of the profile.
[0039] Multi-dimensional dynamic profiles are updatable feature sets built on a continuous data chain. The dimension system originates from business analysis and risk modeling needs and generally includes three categories: basic attribute dimensions, economic attribute dimensions, and service product association dimensions. Basic attribute dimensions extract gender, age group, region, channel preference, etc., from registration information, compliant external data, and high-confidence signals, establishing credibility and update time. Economic attribute dimensions extract amount distribution, periodicity, growth rate, and stability indicators from transactions, payments, balances, credit limits, income certificates, and consumption records, using time windows and exponential decay weights to form both real-time and trend views. Service product association dimensions extract trigger frequency, processing time, cross-departmental transfer frequency, product mix, and lifecycle stage from service usage and after-sales stages. Each dimension's feature definition includes a measurement function, windowing strategy, missing value handling rules, and legal and abnormal value ranges. Profile updates follow a parallel event-driven and timed-driven mechanism. Event arrival triggers incremental recalculation of local features, while timed tasks perform window rolling and baseline re-estimation at daily or monthly boundaries. In case of conflicts, the latest event time takes priority, while historical snapshots and audit trails are retained.
[0040] To ensure the availability of the aforementioned chains and profiles, a quality measurement system and backtracking mechanism are established. Quality metrics track coverage, timeliness, accuracy, consistency, and traceability, outputting scores and alerts for each dimension. The backtracking mechanism allows rollback to any version number for recalculation at the object or dimension level, supporting gray-scale replay and comparative verification. Data protection is achieved through anonymization, minimal data collection, access auditing, and usage restrictions. Sensitive fields are encrypted using irreversible hashing or specific formats, and during feature processing, only necessary statistics are retained while blocking access to the original plaintext.
[0041] Data acquisition and chain construction can adopt a streaming architecture. Events enter the message queue from multi-channel access gateways, are partitioned by object key, and window aggregation is performed using event time semantics. Identity resolution uses both deterministic matching and learning models. The former generates strong matches based on account binding relationships, contract relationships, and rules of same device, same location, and same fingerprint. The latter calculates similarity using graph embedding and contrastive learning and sets a threshold. New edges exceeding the threshold require manual review or are delayed in taking effect. Out-of-order control adopts a watermark strategy, with a maximum tolerable delay time for events lagging behind the watermark. Late events within the acceptable range are still merged into the current window, while those exceeding the range are entered into the compensation channel and backfilled.
[0042] A batch processing architecture can also be adopted. Incremental data is extracted daily from the log repository and business database, fully sorted according to event time, and object shards are generated. Within each shard, hash deduplication, format standardization, and field completion are performed. The new segments are then appended to the chain log, generating incremental profile data. To reduce peak costs, a micro-batch strategy can be used, rolling out blocks every ten minutes or hour, performing local sorting and merging within each block, and reconciling cross-block conflicts at the end of the day.
[0043] A hybrid architecture can also be adopted. High-value events such as transactions and exception handling use a low-latency streaming link, while low-sensitivity events such as browsing and favorites use a batch processing link. The two links are aligned and merged by time at the profile layer. Feature calculation uses pluggable operators. The streaming link updates near-window indicators in real time, while the batch processing link backfills long-window indicators asynchronously.
[0044] Feature engineering is configurable. Time windows can be set to either rolling windows or sliding windows. Rolling windows are suitable for accounts receivable metrics, while sliding windows are suitable for metrics with strong correlation to recent data. Decay weights can use exponential decay or piecewise linear decay. Exponential decay can adjust for historical influence through the half-life parameter, while piecewise linear decay facilitates alignment with business thresholds. Missing data handling can choose zero-filling, mean-filling, recent observation forwarding, or model estimation, with the choice based on the meaning of the dimension and its sensitivity to subsequent modeling. Outlier handling can employ quantile truncation, robust tail reduction, or isolated forest detection. Detected points are neither deleted nor directly included in the mean; instead, they are added to the outlier count feature and reflected in the quality report.
[0045] Data consistency is achieved through a two-level verification process. The first level performs field-level verification and enumeration validity checks at the access gateway, with failed records entering a repair queue. The second level performs causal consistency checks during chain merging; if time reversal or impossible states occur, replay and reconciliation with the data provider are triggered. To support auditing, the chain log uses immutable storage with checksums, and any backfilling is implemented by appending new segments, while old segments remain read-only.
[0046] In terms of computing resource adaptation, when resources are scarce, a tiered strategy can be adopted to prioritize updates of important dimensions and postpone updates of low-frequency dimensions. Lightweight operators and in-memory indexes are used for real-time paths, while comprehensive computation is performed on offline paths. When deploying across regions, local chain construction and feature pre-computation can be completed at edge nodes, while the central node only performs merging and global alignment, reducing cross-domain latency and data cross-border risks.
[0047] Example Description: In the healthcare business, for wearable device platforms and online health service platforms, devices report physiological events such as heart rate, steps, and sleep segments, while the platform records service events such as course browsing, medical consultations, medication collection, and after-sales processing. These two types of events are aggregated into a continuous data chain through object identifiers and device binding relationships. The profile records age group and region at the basic attribute level, payment cycle and order amount fluctuations at the economic attribute level, and course completion rate, number of consultations, prescription collection intervals, and after-sales processing delays at the service product association level. Update rules are triggered by event arrival to recalculate near-window features, and long-term trends are refreshed weekly, providing stable input for subsequent risk identification.
[0048] In the fintech business, targeting digital banks and internet wealth management platforms, front-end logs generate page visits and function click events, while the back-end system generates events such as account opening, credit granting, transactions, repayments, rights usage, risk verification, and dispute resolution. A continuous data chain is constructed through multiple mappings from accounts to devices and contracts to transactions. The profile records customer segmentation and channel preferences at the basic attribute level, deposit and withdrawal rhythm, credit occupancy rate, overdue days distribution, and asset volatility at the economic attribute level, and function usage frequency, customer service interaction duration, and work order closure time at the service product association level. A hybrid architecture is adopted to maintain near real-time updates, and downstream use of the profile is suspended when the data quality score falls below a threshold to avoid misjudgment.
[0049] This embodiment unifies discrete events scattered across multiple channels and systems into an object-level time-series structure through the collaborative construction of a continuous data chain and multi-dimensional dynamic profiles. Profile features are continuously updated according to windowing and decay rules, forming a computational foundation that reflects the current state while preserving historical trajectories. This enables the stable extraction of overlapping and discrepancy indicators across stages in subsequent processes, avoiding biases caused by data fragmentation, reducing identification lag and the probability of misjudgment, while providing traceable and replayable verification capabilities, facilitating rapid reuse and expansion in different business environments.
[0050] S20, Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model;
[0051] In this embodiment, to address the problem of effectively associating the behavior of objects at different stages, a stage penetration model needs to be established to enable data to exhibit a progressive migration relationship between different stages. Establishing a stage penetration model first requires building a hierarchical framework. This framework, based on stage division, categorizes the object groups in the business into service interaction stage groups, service usage stage groups, and anomaly occurrence stage groups. The service interaction stage group refers to the collection of objects formed at the front-end interaction or business entry point, such as those browsing, consulting, registering, or making initial transactions. Its characteristics mainly reflect the initial interest and potential intent of the objects. The service usage stage group refers to the collection of objects that have entered the business usage stage, such as those completing transactions, actually using services, or participating in subsequent business functions. Its characteristics mainly reflect the continuous behavior and stability of the objects. The anomaly occurrence stage group refers to the collection of objects that exhibit abnormal behavior or trigger abnormal events during interaction and service usage, such as those generating complaints, failing transactions, or experiencing service not meeting expectations. Its characteristics mainly reflect the abnormal risk of the objects.
[0052] Defining progressive migration relationships in this model requires describing the migration process from the service interaction stage group to the service usage stage group, and from the service usage stage group to the anomaly occurrence stage group. This progressive migration relationship is not a simple stage division, but rather based on quantifiable transformation conditions and feature mapping rules. For example, the migration from the service interaction stage group to the service usage stage group can be achieved by setting a transformation ratio threshold and behavioral feature matching conditions. Only when an object meets certain behavioral frequencies or intent indicators will it be mapped to the next stage group. Similarly, the migration from the service usage stage group to the anomaly occurrence stage group relies on historical data statistics and the triggering of characteristic anomaly conditions, such as service response time exceeding a threshold, increased operation failure rate, and the appearance of negative interaction tags.
[0053] To achieve dynamic model adaptation, a strategy library needs to be embedded in the phased penetration model. The strategy library contains various ways to define transfer relationships, such as proportional conditions set through statistical methods, feature combination conditions generated by machine learning models, or rule conditions extracted from expert knowledge. The role of the strategy library is to support the progressive transfer between different phases and to flexibly adjust it in different business scenarios, ensuring that the model can be continuously optimized as data accumulates.
[0054] In this way, the phased penetration model can not only establish progressive mappings between groups, but also generate weight parameters based on the feature differences extracted during the migration process. These weight parameters are used for subsequent penetration analysis to ensure that the model can capture key behavioral trajectories and risk signals during the group's evolution.
[0055] The establishment of a phased penetration model can be achieved in several ways. In one implementation, a rule-based model can define different phase groups and their migration conditions. For example, a migration from the service interaction phase group to the service usage phase group can be initiated when the number of page views exceeds a certain threshold. In another implementation, statistical modeling can be used to compare the actual conversion ratio between different phase groups with a preset threshold, dynamically adjusting the migration relationship when the actual ratio exceeds the threshold. Alternatively, machine learning can be introduced, inputting the behavioral characteristics of the service interaction phase group into a classification model, outputting the probability of migration to the service usage phase group, and using a probability threshold as the migration condition. For migration from the service usage phase group to the anomaly occurrence phase group, anomaly detection models, such as isolated forests or cluster deviation detection, can be used to automatically identify high-risk objects and complete the migration.
[0056] Example Explanation: In the healthcare business, the service interaction stage group can include users who have registered on online health service platforms, browsed health courses, or participated in initial consultations; the service usage stage group includes users who have purchased health management services and continue to use fitness guidance or remote consultation functions; the anomaly occurrence stage group includes users who encounter errors in health data synchronization, unresponsive remote services, or submit negative feedback on their service experience. By analyzing the progressive migration relationships of the stage penetration model, it's possible to identify which users are more likely to enter the anomaly stage due to data delays or service instability.
[0057] In the fintech business, the service interaction stage can include customers who register for an account, browse wealth management products, and undergo risk assessments; the service usage stage includes customers who have completed fund inflows, purchased wealth management products, and participated in credit and payment transactions; and the anomaly occurrence stage includes customers who experience transaction failures, delayed fund inflows, or overdue repayments. By using the progressive migration relationship of the stage penetration model, the entire process of a customer's journey from normal interaction to service usage and then to anomaly occurrence can be tracked, and the migration conditions and differential characteristics can be extracted to support subsequent risk warnings.
[0058] This embodiment establishes a phased penetration model and defines progressive migration relationships, enabling the systematic stratification of dispersed behavioral objects according to stages, thereby forming a progressively advancing group structure. This approach not only avoids a general analysis of the entire object but also allows groups at different stages to be dynamically mapped through transformation conditions, ensuring the model's adaptability and scalability.
[0059] S30, based on the staged penetration model and the multi-dimensional dynamic profile, perform the first-level penetration analysis to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group;
[0060] In this embodiment, focusing on the behavioral transmission relationship between target groups across stages, the first-level penetration analysis uses a stage penetration model and a multi-dimensional dynamic profile as common constraints to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group. The stage penetration model provides stage definitions, time boundaries, migration paths, and sampling criteria to determine the data range and alignment method for the two groups. The multi-dimensional dynamic profile provides a computable set of feature identifiers derived from dimensions such as basic identity attributes, economic capacity and expenditure intensity, service and product interaction records, channel and touchpoint behavior, and time-series derivation. Each feature identifier has a value range, sampling window, update frequency, and missing value handling strategy. The service interaction stage group refers to a group of objects that generate observable behavior at the interaction entry point or pre-processing stage, while the service usage stage group refers to a group of objects that generate effective usage records during the business usage stage. Both groups extract features using a unified profile dictionary and the same observation window. The first-level penetration analysis refers to the quantitative alignment calculation carried out along the migration path from service interaction to service usage, aiming to identify the similarity in the distribution of the same-named feature identifiers between the two groups. Feature overlap refers to the similarity between the statistical distributions or representation vectors of a particular feature or a group of feature identifiers across two populations. It can be implemented as discrete distributions using Jaccard similarity, Bhattacharyya coefficient, Jensen-Shannon similarity, or continuous variables using binned similarity, cosine similarity, or the non-negative normalized form of Pearson correlation. To avoid bias, the two populations are first deduplicated and time-aligned based on the observation window provided by the staged penetration model. Then, a unified frequency histogram or density estimate is generated for each feature identifier according to the profile dictionary, followed by zero-frequency smoothing and total normalization to obtain comparable distribution representations. To accommodate multi-feature linkage, this can be extended to the joint overlap calculation of feature identifier groups, aggregating multiple single-feature similarities using weighted or minimum value rules. The weights are derived from migration weights or historical stability measures defined in the staged penetration model. For cases of extreme class imbalance or large differences in sample size, stratified resampling or importance weighting is used to correct for cardinal effects, ensuring that the overlap is not amplified by population size or diluted by sparse noise. For time-sensitive features, a rolling window is used to generate a multi-timepoint overlap sequence, and median filtering or exponential smoothing is applied to abrupt changes to avoid short-term abnormal disturbances affecting structural judgments. The final output is a feature identifier overlap vector and its confidence interval, covering both single feature and feature group granularities, and the calculation caliber and window identifier are recorded for stable reuse in subsequent model components.
[0061] This can be achieved using discrete distribution similarity calculation. A unified binning or category set is constructed for each discrete feature identifier. Normalized frequency vectors are statistically calculated for the service interaction stage group and the service usage stage group respectively. The overlap is then calculated as the sum of the minimum values of each component of the vector, while the confidence interval is saved and can be obtained through bootstrap sampling. Alternatively, a kernel density method for continuous variables can be used. The probability density function is estimated for each numerical feature identifier, and the Bhattacharyya coefficient is calculated by numerical integration. The results are mapped to a zero-to-one interval, and the bandwidth is optimized through cross-validation to balance bias and variance. Representation learning can also be used. Multidimensional feature identifiers are encoded into low-dimensional embeddings. An encoder is trained under stage label supervision through contrastive learning, making the representations of the same feature identifier comparable across different stages. The overlap is calculated using cosine similarity of the embeddings and temperature calibration. For environments with large differences in sample size or high-frequency repeated records within a stage, importance weighting is introduced, with the weight equal to the ratio of the target population distribution to the sampling distribution, or a weighting based on propensity scores is used to reduce selection bias. Time alignment can be achieved through event alignment windows, such as constructing a symmetrical window with the first strong interaction event as the zero point, or tracing back a fixed length from the most recent service usage event to ensure that the two groups are compared within the same lifecycle. Data quality assurance can be achieved by performing missing value identification and imputation before overlapping degree calculation, setting a low-frequency merging threshold for categories, and using Winsorize or quantile scaling for outliers to make the output more robust. To avoid aggregation bias caused by collinearity between features, a correlation penalty or Shapley weighting can be introduced when aggregating overlapping degree to make the contribution more consistent with the marginal information content.
[0062] Example Explanation: In an online health management platform within the healthcare sector, the service interaction stage includes users who have registered and browsed nutrition courses or submitted an initial questionnaire, while the service usage stage includes users who have purchased personalized nutrition plans and submitted periodic health data. Overlap is calculated based on features such as nutrition expenditure intensity, course completion rate tiers, and mobile interaction time periods. If the similarity of the distribution of high-to-medium course completion rates between the two groups is close to one, it indicates that the feature has stable consistency in the transition from interaction to usage, and can serve as an important input for subsequent weight learning.
[0063] In digital wealth platforms within the fintech business, the service interaction stage includes customers who have completed risk assessments and browsed product descriptions, while the service usage stage includes customers who have completed subscriptions and continue to hold their investments. Overlap is calculated based on characteristics such as disposable income percentiles, risk preference tags, and click depth across reach channels. If the joint overlap between risk preference and click depth remains consistently high, it indicates a stable and consistent relationship between early behavior and subsequent usage, facilitating the increase of weights for relevant identifiers and enhancing early identification capabilities in subsequent modeling.
[0064] This embodiment uses the feature overlap degree obtained based on stage constraints and unified profile standards to quantify the structural similarity of the service interaction stage group and the service usage stage group on the same-name features. It outputs a stable, verifiable overlap degree vector with windowed tracking capability, which reduces the misleading effect caused by group size and sampling bias, and provides directly connectable numerical input for subsequent threshold-driven and weight learning, thereby improving the interpretability and forward identification capability of subsequent analysis and early warning links.
[0065] S40, based on the staged penetration model and the multi-dimensional dynamic profile, perform a second-level penetration analysis to extract the feature identification differences between the service usage stage group and the anomaly occurrence stage group;
[0066] In this embodiment, the second-level penetration analysis takes a phased penetration model and a multi-dimensional dynamic profile as input, aiming to measure the characteristic differences between the service usage phase group and the anomaly occurrence phase group in terms of service product association dimensions. The phased penetration model provides the migration relationship from service usage to anomaly occurrence, including conversion ratio thresholds, activation conditions, sample caliber, and migration weights, ensuring comparability and logical continuity between the two groups. The multi-dimensional dynamic profile provides a complete feature set for the service product association dimension, covering service call counts, service type distribution, response time, processing flow path, historical feedback records, usage intensity of associated products, cross-channel touchpoints, etc. These features are defined using a unified dictionary and sampling rules. The service usage phase group refers to the set of objects in the normal usage stage of the business, while the anomaly occurrence phase group refers to the set of objects that have experienced an anomaly or entered the handling stage. The two are demarcated and data extracted through the second migration relationship of the phased penetration model. The second-level penetration analysis calculates the distribution difference of the same-named features in the two groups along the migration path from service usage to anomaly occurrence, obtaining the feature identification difference degree. The dissimilarity can be defined as the Kullback-Leibler divergence, Chi-square distance, or Hellinger distance of a discrete distribution, or the mean difference, variance ratio, or Wasserstein distance of continuous features, ultimately normalized to a range of zero to one. During the calculation, the two groups must first undergo sample deduplication, time alignment, and consistency checks within the same window. Then, a distribution representation is constructed for each feature of the service product association dimension. Missing values can be handled using mode imputation or multiple imputation, and extreme values can be handled using quantile scaling or Winsorize truncation. The dissimilarity statistics must include confidence intervals, which can be generated through bootstrapping resampling or Bayesian inference to prevent overfitting due to sample sparsity. Aggregation of dissimilarity for multidimensional features can be achieved using weighted averages, maximum criterion, or information gain weighting, with weights derived from migration weights in the stage penetration model or historical anomaly prediction capabilities. The dissimilarity of time-sensitive features needs to be calculated at multiple time points to form a time series, with short-term anomalies smoothed to ensure that the dissimilarity represents structural trends rather than random fluctuations. The final output is a feature identifier difference vector, covering single-dimensional features and feature groups, and is bound to the second transfer relation for use in subsequent model calculations.
[0067] This can be achieved using statistical distribution differences. A unified category set is constructed for each discrete service product feature, and the frequency distributions of the service usage stage group and the anomaly occurrence stage group are statistically analyzed. The Chi-square distance or Jensen-Shannon divergence is calculated and normalized to a difference value. Alternatively, it can be achieved by comparing the distributions of continuous features. Kernel density estimates are constructed for numerical features such as service response time and processing cycle, and the Wasserstein distance is calculated as the difference value. The bandwidth parameter is optimized through cross-validation to improve robustness. Representation learning can also be used. Multidimensional service product features are mapped to a low-dimensional embedding space, and a discriminative model is trained under stage label supervision. The cosine distance or Mahalanobis distance of the embedding vectors is used to measure the differences between groups. For scenarios with significant differences in sample size, oversampling, undersampling, or propensity score weighting can be used to balance the sizes of the two groups to ensure that the difference value is not affected by quantity effects. For scenarios with few anomalous samples, a generative model can be used to expand the feature distribution of the anomalous group, and then the stability can be improved by calculating the difference value.
[0068] Example Explanation: In online health management platforms within the healthcare sector, the service usage phase includes users actively using health guidance services and uploading vital sign data, while the anomaly occurrence phase includes users reporting service anomalies or data upload failures. By comparing the distribution of service response time and data upload frequency, a significantly increased difference in the degree of variation between prolonged response time and a sharp drop in upload frequency within the anomaly group suggests that these characteristics may be precursors to anomalies.
[0069] In the fintech sector, digital wealth management platforms serve two groups: those in the active usage phase (customers who have purchased wealth management products and are currently experiencing normal returns) and those in the abnormal phase (customers whose withdrawals have failed or whose funds have been frozen). By comparing the distribution of transaction channel usage and the differences in operation delay times, if the usage ratio of a certain type of channel suddenly increases and the operation delay time significantly lengthens in the abnormal group, it indicates that these characteristics had already deviated before the anomaly occurred, providing a basis for the risk control system to trigger early warnings.
[0070] This embodiment calculates the difference in feature identifiers between the service usage stage group and the anomaly occurrence stage group under stage migration constraints. This reveals the shift pattern of specific service product features before entering the anomaly stage, presenting early signs before the anomaly formation with quantitative indicators. This avoids missing subtle differences by relying solely on overall group statistics, thus providing more interpretable and forward-looking input for subsequent early warning models.
[0071] S50, generate an early warning model based on the overlap of the feature identifiers, the difference of the feature identifiers, and historical data.
[0072] In this embodiment, the early warning model uses feature overlap, feature difference, and historical data as input elements, and is oriented towards a unified calculation framework for risk scoring and level determination. Feature overlap originates from the first-level penetration analysis, reflecting the similarity in distribution between the service interaction stage group and the service usage stage group in the economic attribute dimension, and includes four types of information: dimension name, overlap value, time label, and extraction caliber. Feature difference originates from the second-level penetration analysis, reflecting the distribution offset between the service usage stage group and the anomaly occurrence stage group in the service product association dimension, and includes four types of information: dimension name, difference value, time label, and extraction caliber. Historical data refers to behavioral records and status change records collected in past time windows, used to mark whether abnormal events have occurred and provide statistical benchmarks, including object identifiers, timestamps, stage labels, anomaly labels, and dimension feature snapshots. Before entering the modeling stage, the input elements undergo consistency checks, including time window alignment, consistency of the dictionary of dimensions with the same name, consistency of missing and extreme value handling strategies, and consistency of sampling rules, ensuring that quantitative indicators from different sources are comparable and can be superimposed.
[0073] The weighting design reflects the relative contributions of the two types of indicators. For feature overlap, correlation and conditional occurrence measures with anomaly labels within the historical window are calculated, and weight coefficients for overlap are given. Smoothing terms and confidence interval constraints are set to suppress sample fluctuations. For feature difference, the same two measures are calculated, and weight coefficients for difference are given. A time decay factor is introduced to highlight recent structural changes. Normalization and upper limit pruning rules are introduced in the weight space to avoid overloading of a single dimension. The scoring structure adopts an interpretable weighted combination form. The contribution of economic attributes is determined by feature overlap and corresponding weight coefficients, while the contribution of service products is determined by feature difference and corresponding weight coefficients. Basic attributes or other non-penetrating source variables are used as baseline terms in the superposition to accommodate population size and long-term steady-state differences. To improve cross-time stability, a hierarchical calibration layer is constructed, and the grouped scores are monotonically mapped to the historical anomaly benchmark to ensure consistency between the ranking of scores and occurrence probabilities.
[0074] The risk level threshold system comprises three tiers: high-risk, medium-risk, and low-risk. Thresholds are derived from a combination of historical data distribution and business tolerance settings. Interval division is achieved through quantile calibration or Bayesian calibration, and a drift monitoring mechanism maintains threshold validity. After the threshold and score are correlated, an anomaly level is output: values above the high-risk threshold output a high-risk anomaly level, values in the middle range output a medium-risk anomaly level, and values below the low-risk threshold output a low-risk anomaly level. To support variations in deployment environments, a feature availability awareness mechanism and degradation strategy are introduced. When some dimensions become temporarily unavailable, the system automatically reconstructs a subset of effective features and weight scaling factors, maintaining computable scores and output levels. The entire model workflow includes six stages: input validation, weight estimation, score synthesis, hierarchical calibration, threshold determination, and output write-back. Version numbers, timestamps, and descriptions are provided for easy traceability and auditing.
[0075] This can be achieved using a statistically weighted scoring method. A historical window is constructed and anomaly labels are generated. The conditional occurrence rate and mutual information of feature identifier overlap and difference are summarized by dimension and converted into weight coefficients on both sides. After normalization, the weights are multiplied by their corresponding indicators one by one and summed to obtain the original score. A piecewise linear monotonic mapping is introduced for probability calibration, and a quantile threshold is used to generate three levels. This implementation has low computational cost and is suitable for real-time operation in high-concurrency environments.
[0076] Alternatively, generalized linear modeling can be used. Feature overlap and feature difference are used as explanatory variables, with interaction and time decay terms added. Log-odds regression is used to fit the outlier probabilities, and L1 or group regularization constraints are employed to control multidimensional collinearity and overfitting. Finally, three threshold levels are determined using equal density or equal error rate strategies. This implementation provides parameter estimation and significance testing, facilitating interpretation and auditing.
[0077] Alternatively, a two-stage implementation using gradient boosting trees and a calibration layer can be employed. The first stage involves a base learner learning nonlinear relationships, with inputs including overlap and difference vectors, as well as stable features derived from historical data such as recent fluctuation amplitudes and long-term mean differences. The second stage uses isoregression or temperature scaling for probabilistic calibration, and finally generates a level based on business requirements. This implementation adapts to strong interactions and nonlinear boundaries between features, making it suitable for scenarios with large feature sizes or complex distributions.
[0078] In terms of adaptation, when data is sparse, hierarchical smoothing and information sharing are introduced. Dimensions of the same family are aggregated hierarchically to form parent indicators, and then parent information is fed back into child weights to improve robustness. During data drift, a sliding window is used to re-evaluate weights and thresholds, with the window length automatically adjusted based on statistical uncertainty and response speed. When migrating across environments, a unified dimension dictionary is used through a feature mapping table, and alignment loss is used to constrain weight differences, ensuring consistent level meanings across different business lines. For performance optimization, throughput is improved through vectorized weighting and batch processing comparisons; numerical stabilization techniques are used to control gradient bursting caused by extreme weights; and online calibration and fine-tuning of thresholds are used to address short-term distribution disturbances.
[0079] Example Explanation: In a remote health management platform within the healthcare sector, the overlap of feature identifiers reflects the stable consistency of a certain population group in terms of economic attributes, such as the similarity of their long-term payment capacity range. The difference in feature identifiers reflects the offset in terms of service product association, such as changes in device upload latency distribution. Historical data records device anomalies and service interruption events. After statistically obtaining the weighting coefficients on both sides, a score is generated and the anomaly level is output. Based on this, the platform arranges maintenance intervention and user guidance to proactively reduce the risk of service interruptions.
[0080] In online wealth management platforms within the fintech business, the overlap of feature identifiers reflects stable consumption and similar asset ranges in terms of economic attributes, while the difference in feature identifiers reflects behavioral deviations in terms of service product relevance, such as increased usage of specific transaction channels and longer operation response times. Historical data records abnormal fund operations and transaction failures. Weights are calculated and scores are synthesized according to the above process. After being categorized by percentile thresholds, anomaly levels are formed. Based on these levels, the system triggers risk control verification and customer communication strategies of varying intensities, improving the proactiveness of risk handling and the efficiency of resource allocation.
[0081] This embodiment utilizes both feature overlap and feature difference within a unified scoring structure. The similarity strength on the economic attribute side and the offset strength on the service product side are converted into weighted quantitative contributions. Combined with historical data, weight learning and threshold calibration are completed, directly translating the structural information obtained through layered penetration into executable level outputs. Through calibration and drift adaptation mechanisms, the output anomaly levels remain stable and interpretable across different time windows and operating environments, improving recognition rate and usability without sacrificing real-time performance.
[0082] S60, input the current state of the multi-dimensional dynamic profile into the early warning model and output the anomaly level;
[0083] In this embodiment, the current state of the multi-dimensional dynamic profile includes a feature snapshot associated with the object's most recent time or current window, located by timestamp and organized in the form of feature key values. To enter the early warning model, three types of real-time data need to be decompressed from this profile. The first type is real-time feature values of the basic attribute dimension, covering stable or slowly changing identifiers such as gender, age group, regional stratification, and channel activity level. These are converted into computable vectors using enumeration mapping and one-hot or target encoding methods, and dictionary consistency checks and missing placeholder strategies are performed before input. The second type is real-time feature values of the economic attribute dimension, corresponding to economic strength indicators such as income range, monthly consumption range, and payment ability stratification. These are aggregated by the most recent billing cycle or observation window, using bin numbering or quantile numbering to unify the scale, and a sliding window statistical correction is used to correct abnormal jumps. The third category consists of real-time feature values related to service products, covering process quantities coupled with products and services, such as service trigger count, service response cycle, product holding cycle, interface failure rate, and processing time distribution quantile. It adopts three types of operators: in-window counting, average duration and high quantile, and proportional indicators, unifies the units, and supplements the logarithmic transformation of latency to reduce the long tail effect.
[0084] In the preliminary modeling stage, the early warning model has generated feature overlap weight coefficients and feature difference weight coefficients, and configured high-risk, medium-risk, and low-risk thresholds. In the current execution stage, the real-time feature values of the economic attribute dimension are matched item by item with the feature overlap weight coefficients, performing a vector multiplication and addition to output the first weighted feature value, which expresses the positive and negative contribution and intensity of the economic attribute side to the risk. Subsequently, the real-time feature values of the service product association dimension are matched item by item with the feature difference weight coefficients, also performing a multiplication and addition to output the second weighted feature value, which expresses the marginal impact of structural shifts in the product service side. The real-time feature values of the basic attribute dimension participate in the score synthesis as baseline terms, and their weights are derived from stable terms or mapping terms in the model calibration layer, ensuring that differences in population size are accounted for. The three categories are combined to form a real-time comprehensive score, which can be expressed as S = f0 (real-time feature values of the basic attribute dimension) + f1 (first weighted feature value) + f2 (second weighted feature value), where f0, f1, and f2 are fixed monotonic mappings or linear terms during the training phase. To ensure interpretability and stability across time periods, the real-time comprehensive score enters a hierarchical calibration module, utilizing monotonic mappings learned from historical data to achieve probabilistic or interval-based output. Finally, the real-time comprehensive score is compared with high-risk, medium-risk, and low-risk thresholds to determine the anomaly level. If the real-time comprehensive score is greater than or equal to the high-risk threshold, a high-risk anomaly level is output; if it falls between the medium-risk and high-risk thresholds, a medium-risk anomaly level is output; if it is below the medium-risk threshold but greater than or equal to the low-risk threshold, a low-risk anomaly level is output; if it is below the low-risk threshold, the system can maintain the low-risk anomaly level or output a lower-level placeholder label to support compatibility strategies. To address online feature loss and drift, the execution path incorporates feature availability awareness and adaptive scaling. When real-time feature values for economic attributes or service product association dimensions are missing or outdated, the system automatically removes invalid components and scales the effective components using weight recalibration coefficients to maintain comparability of the real-time comprehensive score on the same scale. All inputs and outputs are written to audit logs with version numbers and timestamps, including the weight version, threshold version, and profile version used, facilitating backtracking.
[0085] This can be implemented using a lightweight vector weighting and quantile threshold determination mechanism. At runtime, the current state of the multi-dimensional dynamic profile is read from the feature storage. The real-time feature values of the basic attribute dimensions are mapped to baseline sub-items B according to the mapping table. The dot product of the real-time feature values of the economic attribute dimension and the feature identifier overlap weight coefficient is used to obtain sub-item E. The dot product of the real-time feature values of the service product association dimension and the feature identifier difference weight coefficient is used to obtain sub-item P. A real-time comprehensive score is formed as S = αB + βE + γP, where the coefficients α, β, and γ are determined and frozen during the offline phase through calibration. The anomaly level is output by comparing S with three threshold levels. To improve throughput, batch vectorized calculation is used in the implementation, concatenating multiple objects in the same time slice into a matrix, and broadcasting the weight coefficients as column vectors to reduce redundant loading.
[0086] A two-stage calibration pipeline can also be used. The first stage outputs the raw score Sraw, and the second stage uses isoregression to map Sraw to probabilities. Thresholds are then determined using the probability space, with the threshold derived from historical data distribution or optimization of misjudgment costs. Two-stage calibration can quickly correct drift after deployment using a small amount of labeled data, avoiding frequent retraining.
[0087] A latency tolerance and fault compensation mechanism can also be employed. For objects with delayed data in the real-time channel, to ensure timely response, a temporary score Stemp is first calculated using a subset of available features, and a temporary anomaly level is output, while a compensation flag is registered. When the delayed data arrives, a recalculation is triggered. If the level changes, a correction record is issued, and priority adjustment is performed in the subsequent processing module. This mechanism ensures stable output even when online execution is affected by network jitter or inter-regional aggregation delays.
[0088] For parameter optimization, α, β, and γ can be optimized offline according to the objective function, with the objective being a weighted combination of level stability, recall, and precision. Thresholds can use cost-sensitive quantiles or equal error rate strategies to address the varying tolerances of different business lines for missed and false alarms at high-risk anomaly levels. When deployed across environments, different data domains are aligned using dimensional dictionary alignment and standardization centers to maintain consistency in feature scale and weight meaning; temperature scaling adjusts the probability scale so that anomaly levels reflect similar handling implications across different business lines.
[0089] Example Description: In a remote health management platform within the healthcare business, the current status profile includes indicators such as device upload latency percentile, interaction frequency, and payment capability tiers over the past seven days. The system calculates a first weighted feature value reflecting economic stability and a second weighted feature value reflecting the offset between the device and service links. A real-time comprehensive score is synthesized and compared with three threshold levels. The resulting anomaly level is used to trigger maintenance inspections or user guidance, allowing for intervention before problems escalate.
[0090] In online wealth management platforms within the fintech sector, the current profile includes indicators such as recent fund inflow / outflow percentiles, transaction channel response times, product holding periods, and asset ranges. The system weights real-time feature values for economic attributes using a feature identifier overlap weighting coefficient, and real-time feature values for service product association using a feature identifier difference weighting coefficient. This results in a comprehensive real-time score, which is then mapped to an anomaly level. This score drives risk control verification and customer communication strategies of varying intensities, enabling proactive identification and tiered response to potential anomalies.
[0091] This embodiment uses a multi-dimensional dynamic profile to map the current state into an early warning model. The similarity strength of economic attributes and the offset strength of service products are simultaneously included in the real-time comprehensive score. Combined with hierarchical calibration and three threshold levels, a stable anomaly level output is formed. Online adaptive and compensation mechanisms reduce the impact of feature loss and data drift on the judgment, making the level both interpretable and real-time, facilitating subsequent graded handling and resource allocation.
[0092] S70, trigger an early warning signal based on the aforementioned anomaly level.
[0093] In this embodiment, the anomaly level originates from the interval judgment output of the early warning model, including three categories of labels: high-risk anomaly level, medium-risk anomaly level, and low-risk anomaly level, along with corresponding scores, threshold versions, and timestamps. Triggering an early warning signal requires three consecutive steps. The first step is level resolution and routing mapping, which reads the anomaly level and retrieves the correspondence between signal type, sending channel, and handling priority identifier in the mapping table. High-risk anomaly levels are mapped to high-risk early warning signals and associated with the first handling priority identifier; medium-risk anomaly levels are mapped to medium-risk early warning signals and associated with the second handling priority identifier; and low-risk anomaly levels are mapped to low-risk early warning signals and associated with the third handling priority identifier. The mapping table is managed by version, including the effective time and grayscale ratio, and supports smooth switching. The second step is trigger condition determination and jitter suppression, which establishes a silent window and growth criteria for the same object and the same anomaly level. Within the silent window, only when the score increases by more than a set increment compared to the previous event is the trigger allowed to avoid repeated triggering; after exceeding the window, the latest score is used to re-trigger. To combat score fluctuations, a dual-threshold mechanism is introduced, separating the entry and exit thresholds. An upward signal is only generated when the score exceeds the entry threshold, and the status is cleared only when it falls below the exit threshold. The third stage involves signal instantiation and idempotency protection. An early warning signal payload is generated, including object identifier, anomaly level, score, threshold hit information, handling priority identifier, trigger time, version number, and tracing pointer. The idempotency key is calculated using the object identifier, level, threshold version, and discrete time slice, and duplicates are removed. To ensure traceability, a globally unique alarm number is assigned, and parent-child relationships are recorded to support escalation or merging.
[0094] The issuance of early warning signals utilizes an asynchronous message channel, with default routing to a message queue to handle high concurrency. A rate shaper and priority scheduling are configured before the queue; high-risk early warning signals occupy the highest-weighted channel and are dequeued first, while medium- and low-risk signals are processed in a weighted rotation. To adapt to different receivers, the signal transcoding layer renders the payload into a unified event pattern, simultaneously assembling SMS, in-app push, email, or API call messages according to channel templates. Channel selection is controlled by routing mapping and can be overlaid with time-based strategies and object preferences. To avoid noise, alarm merging and clustering capabilities are provided, aggregating alarms into group-level signals within a window based on object, anomaly level, feature cluster, or service product association. Aggregation records retain member lists and representative scores, and subsequent escalation or degrade is processed on a group-by-group basis.
[0095] Priority flags drive downstream timing and dispatching. The first priority flag starts the shortest response timer and is submitted to the immediate processing queue; the second priority flag starts a longer timer and is submitted to the fast processing queue; and the third priority flag enters the regular processing queue. All timers and queues record timeout callbacks and escalation rules. If a task is not signed off or completed within the time limit, the warning signal corresponding to the exception level is automatically escalated or the priority flag is raised. To ensure consistency, a state machine is introduced to manage the signal lifecycle. States include New, Routed, Notified, Signed Off, Processing, Closed, and Escalated. State transitions are event-driven and written to the audit log.
[0096] An event-driven architecture can be used to implement the triggering chain. The level resolution module receives anomaly level events, retrieves the routing configuration from the memory-mapped table, constructs the signal payload, and writes it to a high-performance message queue. The deduplication unit calculates the idempotent key; if the cache is hit, the event is discarded; otherwise, it is allowed and written to the short-term key cache. Jitter suppression is achieved through a combination of dual thresholds and a silent window; the silent window parameter and the ratio of the entry / exit threshold can be tuned online. The priority scheduler uses a multi-queue weighted round-robin approach, with high-risk warning signals having the highest weight, followed by medium-risk signals, and low-risk signals having the lowest weight. When the same object exhibits multiple anomaly levels within a short period, a stack-top retention strategy is adopted, retaining only the highest level and writing the rest as annotations to the same alarm number to avoid multiple dispatches. The channel adapter selects the sender according to the routing configuration: SMS and email use the external gateway, while in-application push and API calls use the internal service gateway. Failure retries employ exponential backoff and switch to a backup channel after reaching the maximum number of attempts, while simultaneously writing failure details back to the audit database for subsequent parameter tuning. Priority flags and queue mappings are centrally managed by the strategy center, which provides versioning interfaces and grayscale parameters. During runtime, the latest version is periodically retrieved and the priority is adjusted at the window boundary without interrupting tasks in transit. Upgrade rules are triggered by timer events. If a task with the highest priority flag is not acknowledged within the target timeframe, the system automatically dispatches it to the monitoring channel and adds an upgrade flag to the original signal. Medium and low-risk tasks can be configured for cross-level upgrades or simply have their queue weight increased to avoid excessive disruption.
[0097] A clustering and merging module can be added before triggering. For adjacent objects of the same service product related dimensions, when a large number of medium-risk anomalies occur in a short period of time, they are merged into a scenario-level early warning signal and the handling priority is increased. The signal is then dispatched to the operations or maintenance team for batch processing. High-risk anomalies of individual objects are still dispatched independently. This clustering threshold and time window can be adaptively adjusted according to the business fluctuation cycle, and a reasonable window is learned by using historical peaks to reduce erroneous merging.
[0098] A feedback loop can also be introduced. Each warning signal, upon closure, collects the handling results, time consumption, and accuracy tags, writing them back to the evaluation database. It periodically calculates the arrival rate, acceptance rate, response time, and the percentage of signals converted into actual anomalies for different levels and handling priorities, providing recalibration data for routing mapping and threshold settings. For dimensions with high false alarm rates, it automatically lowers the routing strength or extends the silence window; for scenarios with high false alarm rates, it shortens the silence window and lowers the entry threshold. The online system loads new mappings and parameters through hot updates, without requiring system downtime.
[0099] Example Description: In a remote health service platform in the healthcare business field, if an object shows a high-risk abnormality level twice consecutively, the level resolution module maps it to a high-risk warning signal and associates it with the first priority identifier. The deduplication device suppresses duplicate events within the same time slice, retaining only one valid trigger. The signal enters the instant processing queue and simultaneously triggers two channels: in-application push and API call. If the API call fails, the SMS gateway is switched. The timer is set to automatically upgrade if not signed for within fifteen minutes. The maintenance channel prioritizes dispatching orders and records closed-loop data after the upgrade event arrives.
[0100] In online trading platforms within the fintech business sector, if multiple objects generate medium-risk anomalies within the same service product association dimension during a certain period, the grouping and merging module aggregates them into a scenario-level early warning signal and elevates it to the second priority handling identifier, routing it to the operations batch processing queue. If individual objects simultaneously exhibit high-risk anomalies, the system retains the individual high-risk early warning signal and pushes it to the risk control channel via the real-time processing queue to avoid it being masked by merging rules. The overall trigger link records the idempotent key, alarm number, and state transition for subsequent parameter recalibration.
[0101] This embodiment achieves a one-to-one mapping between anomaly levels and warning signals, and reduces duplication and noise through jitter suppression, deduplication, and clustering. Priority flags directly connect triggering results to queues and timers, ensuring high-risk paths receive the shortest latency and highest resource weight. Asynchronous and idempotent design maintains stability under high concurrency, and upgrade and feedback loops allow for continuous correction of the triggering strategy, improving overall response time and resource utilization while reducing operational costs from false triggers.
[0102] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, device, equipment, and medium for anomaly early warning based on phased penetration, comprising: collecting behavioral data of an object throughout the entire interaction process to form a continuous data chain, and constructing a multi-dimensional dynamic profile based on the continuous data chain; establishing a phased penetration model, defining a progressive migration relationship between service interaction phase groups, service usage phase groups, and anomaly occurrence phase groups in the phased penetration model; performing a first-level penetration analysis based on the phased penetration model and the multi-dimensional dynamic profile to extract the feature overlap between the service interaction phase group and the service usage phase group; performing a second-level penetration analysis based on the phased penetration model and the multi-dimensional dynamic profile to extract the feature difference between the service usage phase group and the anomaly occurrence phase group; generating an early warning model based on the feature overlap, feature difference, and historical data; inputting the current state of the multi-dimensional dynamic profile into the early warning model and outputting the anomaly level; and triggering an early warning signal based on the anomaly level. This invention, by establishing a progressive migration relationship between the interaction, usage, and anomaly phases and combining it with a multi-dimensional dynamic profile, achieves the extraction of key indicators from feature overlap and difference, and utilizes historical data to construct an early warning model, enabling accurate identification of anomaly levels. When the anomaly level reaches a preset threshold, an early warning signal is triggered, thereby enabling early identification and warning before the anomaly risk becomes apparent, improving the accuracy and real-time nature of the warning.
[0103] In one embodiment, step S10 above includes:
[0104] S101, set browsing stage nodes, consultation stage nodes, transaction decision stage nodes, service usage stage nodes, and exception stage nodes in the entire interaction process;
[0105] S102, in the browsing node, record the content access type characteristics and access frequency characteristics;
[0106] S103, In the consultation process node, record the characteristics of the interaction channel type and the characteristics of the question classification;
[0107] S104, In the transaction decision-making stage node, record the product selection type characteristics and decision cycle characteristics;
[0108] S105, In the service usage stage node, record the service trigger count characteristics and service response cycle characteristics;
[0109] S106, In the abnormal process node, record the characteristics of the cause of the abnormality and the characteristics of the handling result;
[0110] S107, integrate the features of all node records in the time series to form a continuous data chain;
[0111] S108. Based on the continuous data chain, extract basic attribute dimension features, economic attribute dimension features, and service product association dimension features, and construct a multi-dimensional dynamic profile based on the basic attribute dimension features, economic attribute dimension features, and service product association dimension features.
[0112] In this embodiment, the entire interaction process covers the complete cycle from the first touchpoint to the anomaly closure. It sets nodes for browsing, consultation, transaction decision-making, service usage, and anomaly stages, unifying event identifiers, object identifiers, and timestamp precision. A single event pattern carries node type, feature key value, source channel, and traceability pointer. Browsing nodes are oriented towards the page and content system. Content access type characteristics are derived from the page tag system and content classification table, such as product introduction pages, policy interpretation pages, and fee explanation pages. Address pattern matching and page metadata mapping are used to map these to discrete categories or multi-tag sparse vectors. Access frequency characteristics are based on a sliding time window to count unique access counts and cumulative dwell time. Duplicate reporting caused by refreshes is handled using event hashing for deduplication and session-level aggregation. Cross-device access is merged using object identifiers. The consultation stage nodes are oriented towards interactive channels. Channel type characteristics are identified from the session entry point, covering online consultation, telephone agents, intelligent Q&A, email processing, etc., with unified coding to support cross-channel alignment. Question classification characteristics are mapped from inquiry text or processing tags to a hierarchical question database. Rule templates and a lightweight text classifier can collaboratively output first- to third-level classifications, recording confidence levels to support subsequent quality control. The transaction decision stage nodes capture the behavior from intention to decision. Product selection type characteristics are directly taken from the purchase list and trial calculation records, coded by product family and configuration item, supporting one-to-many scenarios. The decision cycle characteristics are calculated using time difference, representing the continuous duration from the initial attention or trial calculation to the confirmation or abandonment time. Abnormal interruptions are addressed through a session recovery mechanism and re-entry markers to ensure cycle continuity. The service usage node quantifies actual service triggers and responses. The service trigger count feature counts the number of triggers for the same object within a window and categorizes them by service type. The service response cycle feature is calculated based on the trigger acceptance time and the first response time; when manual and automatic responses are concurrent, the earliest valid time is selected, and timeouts and multiple reopenings are corrected using the work order master-slave relationship. The anomaly node focuses on anomaly attribution and closure. The anomaly cause feature is aligned with the anomaly classification system, distinguishing between rule validation failures, service quality deviations, billing disputes, etc. When multiple tags coexist, primary and secondary causes are synthesized according to weight. The handling result feature records the processing actions and result status, including whether the issue was resolved, whether compensation was provided, whether it was escalated, the resolution time, and the attribution of responsibility, providing an endpoint marker for time series evaluation.
[0113] Time series integration sorts events by event time and employs a watermarking strategy for late data to ensure stable readability of the continuous data chain after a fixed delay. Each chain includes a version number and source path for easy backtracking and recalculation. Data across time zones is uniformly converted to a single time base, and session boundary markers are inserted into the chain to segment long-cycle behaviors. To improve data reliability, integrity checks and outlier management are introduced. Integrity is evaluated based on field coverage and constraint rules; data that fails to meet the standards enters a repair queue. Outliers are screened out using quantile fencing and business constraints, and correction strategies are recorded to maintain interpretability. Chain storage adopts a dual-track structure of wide tables and logs. Wide tables aggregate key features for high-frequency retrieval, while logs retain fine-grained events for retraining and auditing.
[0114] Feature extraction revolves around three dimensions. The basic attribute dimension covers age range, region, identity category, and device profile, derived from registration information, authentication information, and device fingerprints. Uniqueness is ensured through primary key association and conflict resolution. Missing and anomaly detection employs a distribution-aware imputation strategy with missing indicator bits. During the encoding stage, one-hot encoding or target encoding is selected based on model needs, while sensitive fields are anonymized and access controlled. The economic attribute dimension reflects income level, consumption capacity, and payment preferences. Income level can be inferred from salary slips, tax returns, or self-filled ranges. Monthly consumption level is constructed from payment details and bill summaries. Payment preferences are based on payment tools and installment payment behavior statistics. Continuous features are binned and standardized, and time decay weights are introduced to improve timeliness. The service product association dimension reflects the interaction between the object and the product and service, including product holding period, service trigger frequency, function usage coverage, response experience indicators and anomaly density. The holding period is measured in calendar days from the effective date to the current date and is statistically segmented in segments during the change period. The trigger frequency is divided into short-term and medium-to-long-term scales according to service type and window. The coverage is the ratio of function usage count to the total number of functions. The response experience is the quantile of the response period and the default ratio. The anomaly density is the proportion of abnormal events within the window to all service events.
[0115] Multi-dimensional dynamic profiles are represented by time-sensitive vectors. These vectors consist of basic attribute blocks, economic attribute blocks, and service product association blocks. Each block contains a fixed-length or sparse sub-vector and a time decay parameter, recording the generation time and update factor. Profile updates employ an incremental mechanism; real-time events trigger local recalculation of the same object upon entering the feature pipeline. Recalculation rules follow a dependency graph, avoiding full write-back. To ensure cross-system consistency, profiles are exposed externally through versioned interfaces, including current values, window statistics, recent periodic change rates, and confidence levels. Consumers consume profiles by version and read compatible mappings during incompatible upgrades. Profile quality is monitored using four metrics: coverage measures the usable proportion of fields; accuracy is verified through sampling and reconciliation; timeliness is measured by the end-to-end latency from event to profile availability; and stability is monitored by intraday drift and weekly comparison. Privacy compliance is achieved through anonymization, minimal data collection, and usage binding. Differential privacy noise is used to control the risk of information leakage in the aggregated output of sensitive groups. Scalability is achieved through pluggable feature operators and unified metric specifications. When encountering new business nodes, only event mappings and operator templates need to be added to integrate them into the profile.
[0116] This embodiment transforms scattered records into a continuous data chain through unified node-based data collection and time-series integration. Three-dimensional extraction normalizes behavior, attributes, and service relationships into a single profile vector. Incremental updates and time decay ensure that recent behaviors contribute more significantly to the profile, while quality and privacy controls reduce data noise and compliance risks. This results in a computable, traceable, and alignable profile representation at the same object level, providing stable input for subsequent layered penetration testing, threshold determination, and risk classification. It also reduces information gaps caused by cross-system integration, improves real-time processing efficiency, and enhances the accuracy of subsequent analyses.
[0117] In one embodiment, step S20 above includes:
[0118] S201, Create a phased penetration model framework, and define the initial scope of the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model framework;
[0119] S202, Create a conversion strategy library, and define a first migration relationship conversion strategy in the conversion strategy library. The first migration relationship conversion strategy includes a conversion ratio threshold from the service interaction stage group to the service usage stage group and a first migration relationship feature matching condition.
[0120] S203, Define a second migration relationship conversion strategy in the conversion strategy library. The second migration relationship conversion strategy includes a conversion ratio threshold from the service usage stage group to the anomaly occurrence stage group and a second migration relationship feature matching condition.
[0121] S204. Based on historical data, calculate the actual conversion rate from the service interaction stage group to the service use stage group, and the actual conversion rate from the service use stage group to the anomaly occurrence stage group.
[0122] S205, compare the actual conversion ratio with the corresponding conversion ratio threshold;
[0123] S206, when the actual conversion ratio from the service interaction stage group to the service use stage group exceeds the corresponding conversion ratio threshold, the first migration relationship feature matching condition is activated.
[0124] S207, when the actual conversion ratio from the service usage phase group to the anomaly occurrence phase group exceeds the corresponding conversion ratio threshold, the second migration relationship feature matching condition is activated.
[0125] S208, Generate conversion strategy weight parameters based on activated feature matching conditions;
[0126] S209, Based on the initial scope of the service interaction stage group, the conversion strategy weight parameters, and the activated first migration relationship feature matching conditions, construct the first migration relationship;
[0127] S210, construct the second migration relationship based on the initial scope of the service usage stage group, the conversion strategy weight parameters, and the activated second migration relationship feature matching conditions;
[0128] S211, Generate a service usage phase group set based on the first migration relationship;
[0129] S212, Generate a set of groups for the anomaly occurrence stage based on the second migration relationship;
[0130] S213, Generate a service interaction phase group set based on the initial range of the service interaction phase group.
[0131] In this embodiment, the phased penetration model uses an independent model framework as its carrier, within which three types of configurable objects and two types of computational pipelines are reserved. The three types of objects include group definitions, migration relationships, and a strategy library; the two types of computational pipelines include historical evaluation and relationship generation. First, a set of group definitions is established in the model framework, defining the initial scope of groups in the service interaction phase, service usage phase, and anomaly occurrence phase. The initial scope is expressed using Boolean expressions or query templates. The inputs are dynamic profiles and event chains, and the outputs are a set of object identifiers and a time window. Multiple conditions can be applied in parallel, and mutually exclusive conditions can be prioritized. Version numbers and effective intervals are provided for backtracking.
[0132] The strategy library centrally manages conversion strategies and matching conditions, employing a hybrid structure of key-value tables and rule graphs. The first migration relationship conversion strategy addresses the flow from service interaction to service use, including a conversion ratio threshold and matching conditions for first migration relationship features. The second migration relationship conversion strategy addresses the flow from service use to anomaly occurrence, also including a conversion ratio threshold and matching conditions for second migration relationship features. Thresholds are recorded in a hierarchical configuration, supporting basic thresholds, industry-specific adjustment values, and time decay factors. Matching conditions are represented by feature selection expressions and sets of comparison operators, supporting discrete feature sets, continuous feature binning intervals, and combined feature interaction items. The strategy library records the risk level, applicable data period, confidence interval, and reviewer for each strategy to ensure traceability.
[0133] The historical evaluation pipeline focuses on the statistics and comparison of conversion rates. Inputs include object flows and profile snapshots of three groups within a historical window. When calculating the actual conversion rate from service interaction to service use, a baseline set is first generated for the service interaction group at the start of the window. Then, service use events generated within the window are associated based on the first migration relationship candidate condition, using object identification and time sequence as association criteria and excluding duplicate triggers. The numerator is the number of objects that have completed the flow, and the denominator is the size of the baseline set. The actual conversion rate from service use to anomaly occurrence is calculated using the same method. To reduce random fluctuations, beta prior or sliding weighting is introduced for rate estimation, and the window length can be optimized according to business seasonality. The comparison between the rate and the threshold uses a combination of significance testing and confidence interval overlap determination, providing three states: exceeding, approaching, or not reaching. The state is accompanied by records of the comparison time point, the window used, and the sample size.
[0134] When the proportion exceeds a threshold, the matching condition enters an active state. The active state is described by a flag and an expiration date. If necessary, a condition snapshot is generated to lock the condition expression and feature binning version, preventing drift caused by subsequent feature dictionary updates. The activated condition then enters the weight generation stage. Weight generation feeds the explanatory variables of the activated condition and historical transformation labels into a weighted learner. The learner can use log-odds regression, gradient boosting, or hierarchical Bayesian modeling, or it can use hierarchical scoring rules and entropy weighting. To ensure interpretability and stability, the weight vector undergoes monotonicity constraints, regularization, and confidence discounting. The output is the transformation strategy weight parameters, which include timestamps and applicable relationship type labels.
[0135] The relationship generation pipeline constructs a migration relationship graph based on three elements. The first migration relationship uses the initial scope of the service interaction stage group, the conversion strategy weight parameters, and the activated first migration relationship feature matching conditions to jointly determine the edge set from service interaction to service use. The second migration relationship uses the initial scope of the service use stage group, the sub-vectors of the same set of weight parameters under the second relationship, and the activated second migration relationship feature matching conditions to determine the edge set from service use to anomaly occurrence. The generation process first calculates the object-level matching score within the initial scope. The score is a weighted linear or non-linear combination of the matching condition trigger items. Then, the score is thresholded and truncated by ranking. The threshold is taken from the activated strategy, and the ranking ratio can be set in combination with the sample size and expansion strategy. The set output is obtained by projecting the edge set. The service use stage group set is obtained according to the first migration relationship, and the anomaly occurrence stage group set is obtained according to the second migration relationship. The service interaction stage group set comes from the valid instances in the current batch of the initial scope and is used to provide a comparison and supply boundary. All set outputs include member identifiers, scores, source strategies, and validity periods to facilitate subsequent hierarchical penetration and early warning model consumption.
[0136] To control cross-contamination and circular dependencies, relation generation is performed topologically within a single batch, calculating the first migration relation first, then the second. Conflict resolution is performed on multiple edges of the same object when necessary, following the rules of high score priority and relation priority to prevent the same object from falling into unreasonable multiple sets within the same batch. The model framework provides a replay interface, allowing the loading of any historical version of the strategy library, feature dictionary, and initial population range to reproduce experimental results. A shadow mode is also provided, calculating offline outputs of new weights or thresholds without affecting online results, for safe switching. To ensure cross-domain consistency, all thresholds, conditions, and weights are named using relation types as namespaces, and the configurations of the first and second migration relations are stored separately to prevent naming confusion.
[0137] This embodiment links the initial population scope, conversion strategy threshold, activation matching conditions, and data-driven weight parameters within a unified framework. Historical evaluation provides a stable actual conversion ratio, activation actions transform static conditions into time-sensitive rules, weight generation quantifies empirical rules into computable parameters, and relationship generation maps individual-level scores to explicit migration edges and stage sets. This forms a replayable, configurable, and interpretable staged penetration structure, avoiding biases caused by coarse-grained processing of the overall sample, improving the identification resolution between different stages, and providing a highly consistent input set and clear cross-stage flow relationships for subsequent hierarchical penetration analysis and early warning modeling.
[0138] In one embodiment, step S30 above includes:
[0139] S301, invoke the first migration relationship in the phased penetration model;
[0140] S302, based on the first migration relationship, extract the economic attribute dimension feature dataset of the service interaction stage group and the economic attribute dimension feature dataset of the service usage stage group from the multi-dimensional dynamic profile.
[0141] S303, compare the economic attribute dimension feature dataset of the service interaction stage group with the economic attribute dimension feature dataset of the service usage stage group, and generate a first feature distribution difference map;
[0142] S304, determine the feature overlap value of each economic attribute dimension based on the first feature distribution difference map;
[0143] S305, when the feature overlap value of a certain economic attribute dimension exceeds a preset threshold, the economic attribute dimension is marked as a high overlap feature identifier.
[0144] S306, associate the high overlap feature identifier with the first migration relationship, and summarize the set of high overlap feature identifiers to form the feature identifier overlap degree.
[0145] In this embodiment, the first migration relationship serves as a relationship instance in the phased penetration model, defining the object flow and matching boundaries between the service interaction phase group and the service usage phase group. It includes executable conditional expressions, value dictionary versions, and time-effectiveness information. The multi-dimensional dynamic profile provides attribute slices of objects at the same time baseline. The economic attribute dimension includes indicators such as annual income, disposable income, consumption limit, payment ability index, credit limit utilization rate, and installment payment ratio, maintained using a dictionary that combines discrete binning and continuous intervals. The service interaction phase group and the service usage phase group are obtained by concatenating the profile table and event table using a group filtering statement. This statement is instantiated from the conditional template of the first migration relationship, and the output is a set of object identifiers and a feature vector of the economic attribute dimension, time-aligned to a unified observation window.
[0146] The construction of the economic attribute dimension feature dataset employs field-level validation and missing value handling. Discrete features are unified to a standard coding space, continuous features are segmented according to distribution and bin boundary versions are recorded, outliers are handled through double-sided truncation and quantile pull-back, and missing values are imputed using the most recent observation within the same group or distribution inference, retaining the marker bits from the imputation process for subsequent weight discounting. Two datasets are created for each group, with columns aligned and rows deduplicated, each row uniquely identifying the corresponding object and binding a timestamp. To avoid sampling bias, group size alignment is performed during dataset construction, using resampling or inverse probability weights, with weights written to the sample attribute columns for subsequent statistical use.
[0147] The first feature distribution difference map is generated by comparing the distribution of each economic attribute dimension in the two datasets. Discrete features calculate the category frequency distribution and smooth it, while continuous features generate density curves based on kernel density or binning frequency. The core output of the difference map is dimension-level distribution pairs, comparison indicators, and visualization coordinates. The comparison indicators include two types: symmetric distance and similarity. The former indicates the magnitude of the distribution difference, and the latter indicates the degree of distribution overlap. Symmetric distance can be calculated using Jason-Shannon divergence and total variation distance, while similarity can be calculated using overlap coefficient and Hellinger similarity. All indicators are adjusted with group weights and include confidence intervals. The calculation uses bootstrapping or Delta approximation, and the confidence intervals and sample sizes are written into the map's metadata. The map stores the density sequences of the two groups, similarity scores, and traceable feature dictionary versions using the dimension as the key.
[0148] Feature overlap values are determined based on the difference map. For discrete features, the overlap is calculated by summing the minimum values of class frequencies for each category. For continuous features, the overlap is calculated by the area under the two kernel density curves or by using the squared form of Hellinger similarity. To stabilize the estimation of small sample dimensions, a sample size correction term and a smoothing prior are added to the overlap value. The correction term is adjusted based on the effective sample size and the number of categories, and the prior is derived from the hierarchical summation of historical distributions. The calculation result falls within the range of zero to one, and three additional pieces of information are recorded: sample size, smoothing coefficient, and calculation method identifier. Preset thresholds are provided through the strategy parameter service and differentiated by dimension type. Discrete and continuous dimensions can use different threshold ranges. The thresholds have decay or rolling update attributes over time, and the historical window for threshold generation and parameter human-machine calibration records are also recorded.
[0149] High overlap feature identifiers are generated after comparing the overlap value with a preset threshold. The comparison logic first performs a significance determination. If the overlap value minus the lower confidence bound of the threshold is still regular, it is directly labeled; otherwise, it enters gray zone processing. Gray zone processing makes a conservative or aggressive labeling decision based on the sample size and business priority. The decision strategy is written into the source description of the label entry. The identifier adopts a standardized naming method of dimension name plus value range or value enumeration and is bound to the portrait version, feature dictionary version, and first migration relationship version at the time of calculation to ensure consistency across periods. The labeled high overlap feature identifier is structurally associated with the first migration relationship. The associated object is the condition set of the relationship, and the writing method is either rule strengthening or weight enhancement. Rule strengthening transforms the identifier into an explicit mandatory or preferred condition, assigning a threshold allowable interval and expiration period. Weight enhancement maps the identifier to a positive gain of feature coefficients or a dynamic reduction of the threshold. The gain magnitude is jointly determined by the overlap and sample size and is limited by monotonicity constraints.
[0150] The feature overlap set is formed by summing all labeled economic attribute dimension entries. Elements within the set include dimension code, value range, overlap value, threshold snapshot, significance conclusion, and version number. The set maintains a deduplication and conflict resolution mechanism. If multiple value ranges overlap, they are merged or multiple granular versions are retained after sorting by overlap value and sample size. The set provides two types of interfaces: a query interface for assembling and interpreting the conditions of the first migration relation, and a subscription interface for initializing weights or setting boundaries for the subsequent early warning model along the economic attribute path. To avoid semantic drift across modules, set naming and field naming strictly reuse the namespace of the economic attribute dimension dictionary and relation configuration. The entire processing chain executes within a unified time window. The window length, sliding step size, and minimum sample size are configurable and protected by monitoring indicators, including distribution drift score, estimated variance, and outlier percentage. If a protection threshold is reached, labeling is paused and a calibration prompt is issued.
[0151] This embodiment, by using the first migration relationship as a guide to perform distribution alignment and overlap measurement on the economic attribute dimension, can express the stable commonalities between the service interaction stage group and the service use stage group as calculable identifiers and deposit them into a structured set. The bidirectional binding of identifiers and relationships directly influences subsequent relationship assembly and scoring thresholds, making the identification boundary from interaction to use adaptive and interpretable. Distribution difference maps and significance judgments jointly suppress misjudgments caused by accidental consistency. Sample size correction and version management ensure cross-period consistency and traceability. The aggregated feature identifier overlap provides a robust economic attribute feature subspace for subsequent early warning modeling, reducing the interference of noise dimensions on downstream judgments and forming verifiable causal guidance in the stage penetration chain.
[0152] In one embodiment, step S40 above includes:
[0153] S401, invoke the second migration relationship in the staged penetration model;
[0154] S402, based on the second migration relationship, extract the service product association dimension feature dataset of the service usage stage group and the service product association dimension feature dataset of the anomaly occurrence stage group from the multi-dimensional dynamic profile.
[0155] S403, compare the service product association dimension feature dataset of the service usage stage group with the service product association dimension feature dataset of the anomaly occurrence stage group, and generate a second feature distribution difference map;
[0156] S404, determine the feature difference value of each service product association dimension based on the second feature distribution difference map;
[0157] S405, when the feature difference value of a certain service product association dimension exceeds a preset threshold, the service product association dimension is marked as a high difference feature identifier;
[0158] S406, associate the high-discrepancy feature identifier with the second migration relationship, and summarize the set of high-discrepancy feature identifiers to form the feature identifier discrepancy.
[0159] In this embodiment, the second migration relationship defines the migration path, triggering conditions, and applicable time window for the service usage phase group to the anomaly occurrence phase group. The relationship entity includes a set of conditional expressions, a weighted version, a time effective interval, and a grayscale switch, ensuring that the same object receives only one relationship determination within the same observation window. The multi-dimensional dynamic profile provides attribute slices of objects under a unified time benchmark. The service product association dimension covers indicators such as service trigger count, service response cycle, number of handling stages, processing closure time, problem type distribution, product holding period, add-on / unsubscribe events, channel switching frequency, and cross-channel upgrade rate. Discrete and continuous parallel modeling is used, and dictionary versions and value boundaries are bound. The service usage phase group and the anomaly occurrence phase group are instantiated on the profile and event details using the conditional template of the second migration relationship, forming an object set and a service product association dimension feature vector. All vectors are aligned to a unified observation window, such as a fixed-length window that traces back from the latest service completion time or the latest anomaly closure time, and the window length and anchor point time are recorded.
[0160] The service product association dimension feature dataset was constructed separately for the two groups. The construction process included field validity verification, time alignment, outlier and missing value handling, weighting, and sampling consistency control. Discrete features were unified into a standard coding space and low-frequency features were merged. Continuous features were recorded at boundaries using empirical quantiles or adaptive binning based on distribution. Outliers were pulled back using bilateral quantiles, and missing values were imputed using the most recent observation within the object. If imputed could not be imputed, multiple imputation was performed based on the conditional distribution within the same group, and the imputed position was marked on the sample for subsequent weight discounting. To mitigate the bias caused by differences in group size, inverse probability weights or resampling were introduced to achieve size alignment. The weight column was saved with the sample and participated in subsequent statistical calculations. Finally, two column-aligned, row-deduplicated datasets were formed, with each row corresponding to a unique object and containing service product association dimension features, sample weights, observation timestamps, and dictionary version numbers.
[0161] The second feature distribution difference map is generated by comparing the distributions of the two datasets across each service product association dimension. For discrete dimensions, weighted class frequencies are calculated and smoothed using Laplace or Dirichlet methods; for continuous dimensions, kernel density estimation or binning frequencies are used to generate weighted density curves, and confidence bands are output. The comparison metrics provide both difference and similarity metrics. Difference metrics include total variation distance, Jason-Shannon divergence, and first-order Wasserstein distance; similarity metrics include overlap coefficient and Hellinger similarity. All metrics are adjusted for sample weights, and confidence intervals are provided using bootstrapping or Delta methods. The map stores the density sequences of the two groups, comparison metric values, confidence intervals, sample sizes, and version metadata at the dimensional granularity, and provides recalculated random seeds and kernel bandwidth logs to ensure reproducibility.
[0162] The feature difference value is uniformly defined on the comparison index of the second feature distribution difference map. Discrete dimensions use a weighted form of total variation distance or Jason-Shannon divergence, while continuous dimensions use a normalized form of the overlap area complement or Wasserstein distance. Different indices are mapped to a unified interval [0,1] based on commutativity, with larger values representing stronger differences. To suppress fluctuations in small sample dimensions, a sample size correction term and a stable prior are introduced, with correction coefficients as follows:
[0163]
[0164] Where: n eff n represents the weighted effective sample size, used to reflect the actual number of independent contributing samples under the weighted distribution; ref This represents the minimum effective sample threshold for a dimension, used to ensure that the variance calculation for that dimension is adjusted by shrinking when the sample size is insufficient. The prior is derived from the hierarchical summary of historical periods, and is fused with the current estimate using convex combination. The final output is a mapping from dimension to feature variance value, along with the effective sample size, stability coefficient, and calculation method identifier.
[0165] Preset thresholds are configured independently by dimension type, business period, and version management. Thresholds are derived from either a quantile strategy based on historical distributions or a Bayesian risk minimization strategy; these two strategies are interchangeable: the quantile strategy uses the upper α quantile of feature difference values within a historical period as the threshold; the risk minimization strategy finds the cutoff point that minimizes the expected sum of weighted false alarms and missed alarms given the cost matrix. Thresholds have a rolling update mechanism and an expiration period, and retain human-machine calibration records and effective snapshots.
[0166] High-discretion feature identifiers are generated after comparing the feature discretion value with a preset threshold. The comparison first determines significance; if the discretion value minus the lower confidence bound of the threshold is still greater than zero, it is directly labeled. When in the gray zone, a conservative or aggressive strategy is selected based on the effective sample size, historical stability coefficient, and business priority. The identifier uses a standardized naming convention of dimension name plus value range or category enumeration, binding the portrait version, dictionary version, and second migration relation version used in the calculation, and recording the threshold snapshot and significance level at the time of generation. The association between high-discretion feature identifiers and the second migration relation is achieved through two paths: the rule-strengthening path writes the identifier into the relation condition set as an interception or priority monitoring condition, setting weights, tolerances, and expiration periods; the weight-enhancing path maps the identifier to a positive gain of the feature coefficient or a threshold increase. The gain monotonically increases with the discretion value and the effective sample size and is subject to a top-hat constraint to avoid over-amplification.
[0167] Feature identifier dissimilarity is a aggregated representation of highly dissimilarity feature identifiers. The set manages conflicts and overlaps; if multiple intervals overlap within the same dimension, they are sorted based on dissimilarity values and sample size before interval merging or multi-granularity versions are retained. If co-occurrence relationships exist across different dimensions, connection references are provided through mutual information or conditional lift calculations, without altering the definition of individual identifiers. The set provides two types of interfaces: query interface service for strategy assembly and interpretation output of the second migration relationship; and subscription interface service for weight initialization, dynamic threshold adjustment, and rule whitelist / blacklist synchronization of the downstream early warning model on the service product association path. The entire process is under unified monitoring, with monitoring items including distribution drift score, indicator variance, imputation ratio, and consistency with comparison indicators. When drift or uncertainty exceeds the protection threshold, new identifier writing is paused and a calibration reminder is triggered, while existing identifiers are retained until a new version replaces them.
[0168] This embodiment constructs a second feature distribution difference map on the service product association dimension by using the second migration relationship as a guide and quantifies the feature difference value. This allows for the stable deposition of key discrepancies between the service usage stage group and the anomaly occurrence stage group in the form of highly differentiated feature identifiers. These identifiers are then dually bound to the second migration relationship through rule strengthening and weight enhancement, thereby forming an executable, traceable, and adjustable differentiated discrimination boundary on the usage → anomaly migration path. Sample size correction, confidence intervals, and rolling thresholds reduce false alarms and missed alarms caused by occasional fluctuations. Aggregated management and versioned metadata ensure cross-period consistency and rapid backtracking. It provides a focused service product association feature subspace and a clear gain direction for downstream early warning modeling, improving the sensitivity and specificity of anomaly identification, and providing a verifiable chain of evidence for strategy interpretation and compliance modules.
[0169] In one embodiment, step S50 above includes:
[0170] S501, Based on historical data, calculate the overlap rate of complaints corresponding to the overlap of the aforementioned feature identifiers;
[0171] S502, Based on historical data, calculate the complaint rate corresponding to the difference degree of the feature identifier difference degree;
[0172] S503, determine the feature identification overlap weight coefficient based on the overlap complaint occurrence rate;
[0173] S504, determine the feature identification difference weight coefficient based on the difference complaint incidence rate;
[0174] S505 sets high-risk, medium-risk, and low-risk thresholds based on historical data distribution.
[0175] S506, combine the feature identifier overlap degree, feature identifier overlap degree weight coefficient, feature identifier difference degree, feature identifier difference degree weight coefficient, historical data, high-risk threshold, medium-risk threshold and low-risk threshold to generate an early warning model.
[0176] In this embodiment, the process of generating an early warning model based on feature overlap, feature difference, and historical data first requires organizing and decomposing the historical data. Historical data is not a single-source record, but rather a collection of long-term accumulated object interaction records, anomaly occurrence records, and handling feedback information. Through a unified time window and object identifier, all historical samples are standardized to obtain a feature dataset reflecting the state of objects at different stages. Based on this, feature overlap is statistically analyzed, calculating the probability of anomalies occurring in groups with high overlap features within the historical samples, and this probability is defined as the overlap complaint rate. Similarly, feature difference is statistically analyzed, calculating the probability of anomalies occurring in groups with high difference features within the historical samples, and this probability is defined as the difference complaint rate.
[0177] After statistical analysis, the two occurrence rates need to be mapped to weight coefficients applicable to the early warning model. These weight coefficients reflect the degree of influence of different feature identifiers in anomaly prediction. Smoothing and effective sample size correction are introduced during the mapping process to avoid fluctuations caused by a small number of samples affecting the overall weights. When the historical sample size of a certain feature is insufficient to meet a set threshold, the contribution of that feature is proportionally reduced, preventing the model from being excessively biased by random data. The resulting feature overlap and feature difference weight coefficients reflect both statistical regularity and data stability.
[0178] Next, risk thresholds need to be set. These thresholds are based on historical data distribution, and by analyzing the probability of abnormal events in different groups, they are divided into high-risk, medium-risk, and low-risk ranges. The thresholds serve not only as boundary conditions for model judgment but also as the decision-making benchmark for triggering early warnings. By adjusting the threshold positions, a trade-off can be struck between false negatives and false alarms, thus adapting to the risk tolerance requirements of different business scenarios.
[0179] Ultimately, feature overlap, feature difference, two-class weighting coefficients, and risk thresholds are combined into a complete early warning model. In actual operation, the dynamic profile of an object is processed by this model to obtain a risk score, and the corresponding risk level is output based on a comparison of the score and the threshold. The generation process of this model not only incorporates statistical regularities but also dynamically adapts to different risk level distributions, ensuring stable performance across different time periods and business scenarios.
[0180] This embodiment statistically analyzes historical data to determine the probability of anomalies corresponding to the overlap and difference of feature identifiers, converting these probabilities into correctable weighting coefficients. These coefficients are then combined with high-risk, medium-risk, and low-risk thresholds set based on data distribution to form an interpretable, stable, and dynamically adjustable early warning model. This model can score the status of objects in real time during operation and achieve differentiated risk output through tiered threshold judgments. This not only avoids excessive bias caused by data sparsity but also ensures consistency between risk assessment and historical patterns, thereby improving the accuracy and reliability of anomaly early warnings and providing precise support for subsequent tiered handling.
[0181] In one embodiment, step S60 above includes:
[0182] S601, extract the real-time feature values of the basic attribute dimension, the real-time feature values of the economic attribute dimension, and the real-time feature values of the service product association dimension from the current state of the multi-dimensional dynamic profile;
[0183] S602, the feature overlap coefficient in the early warning model is used to weight the real-time feature value of the economic attribute dimension to generate a first weighted feature value;
[0184] S603, apply the feature identifier difference weighting coefficient in the early warning model to weight the real-time feature value of the service product association dimension to generate a second weighted feature value;
[0185] S604, combine the real-time feature values of the basic attribute dimensions, the first weighted feature value, and the second weighted feature value to generate a real-time comprehensive score;
[0186] S605, compare the real-time comprehensive score with the high-risk threshold, medium-risk threshold and low-risk threshold in the early warning model;
[0187] S606 outputs a high-risk anomaly level, a medium-risk anomaly level, or a low-risk anomaly level based on the comparison results.
[0188] In this embodiment, before the current state of the multi-dimensional dynamic profile enters the early warning model, a time alignment and data integrity verification are performed. The profile uses the object's unique identifier and the latest timestamp as the retrieval key to pull a snapshot of the profile consistent with that time point, ensuring that the real-time feature values of the basic attribute dimension, the economic attribute dimension, and the service product association dimension all exist simultaneously and are within the same time window. Missing items are handled using a constraint method: discrete fields that cannot be interpolated are placed with an unknown category, while continuous fields that can be interpolated are stably filled using adjacent windows and an integrity marker is added to ensure that subsequent calculations are not interrupted by null values.
[0189] Real-time feature values for basic attribute dimensions are directly extracted from fields such as age group, geographic level, and user type tags in the user profile. To avoid combination bias caused by scale inconsistencies, categorical fields are uniformly encoded, and numerical fields are mapped to fixed intervals using linear mapping to ensure they fall within an overlayable range. The original distribution quantile labels are preserved for traceable interpretation when needed. This part is directly incorporated in subsequent combination stages and does not participate in weight transformation.
[0190] Real-time feature values for the economic attribute dimension are derived from fields such as income range, monthly expenditure intensity, average order value range, and payment fulfillment behavior. The early warning model has a built-in feature overlap weighting coefficient, representing the stable commonalities related to service usage transfer in the first-level penetration analysis. When calculating the first weighted feature value, a single summary quantity of the economic dimension is generated by weighting each field item by item. Three types of constraints are incorporated in the process: first, the weight value range is limited to prevent a single economic indicator from excessively amplifying its contribution when it rises in the short term; second, field-level confidence flag control is implemented, which weakens the contribution of a field if the sampling confidence of the field in the current window decreases; and third, outlier suppression is implemented, which truncates out-of-bounds inputs to a preset upper limit and records the truncation flag for easy auditing.
[0191] The real-time feature values for the service product association dimension are derived from fields such as service trigger count, processing time distribution, product holding period, function call structure, and after-sales interaction sequence. The early warning model has a built-in feature identification difference weight coefficient, reflecting the difference signals that distinguish the user group from the abnormal group in the second-level penetration analysis. When calculating the second weighted feature value, a field-level item-by-item weighting process consistent with the economic dimension is adopted, while introducing sequence smoothing: for counts and durations representing processes, incremental quantization with a fixed-length sliding window is used to limit the impact of sudden spikes within a window period, thereby avoiding instantaneous fluctuations on the service side from directly causing level jumps.
[0192] The real-time comprehensive score is obtained by aggregating three types of metrics: the base metrics formed by real-time feature values of basic attribute dimensions, the first weighted feature value, and the second weighted feature value. The aggregation process consists of two steps. First, scale alignment is performed to ensure that the three metrics are within the same decision range; then, monotonic calibration is performed to ensure that the contributions of the economic dimension and the service product dimension maintain a monotonically interpretable relationship relative to the base metrics. The entire process does not introduce new training parameters, relying entirely on the generated weight coefficients and the established calibration table to ensure consistency with the historical threshold system.
[0193] The threshold comparison uses the real-time comprehensive score as input and performs interval determination against high-risk, medium-risk, and low-risk thresholds respectively. To reduce boundary jitter, a lightweight hysteresis is added: when the score is within a small interval near adjacent thresholds, the level from the previous time step is maintained until the score continuously crosses the interval. The comparison logic outputs a high-risk anomaly level, a medium-risk anomaly level, or a low-risk anomaly level, and also returns detailed contribution information for interpretation, including the ratio of the first weighted feature value to the second weighted feature value and a list of truncated fields, facilitating auditing and subsequent processing.
[0194] Before data enters the early warning model, the exchange method can be synchronous retrieval, which is suitable for business flows with high real-time requirements; it can also be asynchronous event triggering, where the profile service pushes the data to the model service when the snapshot is updated, which is suitable for scenarios where throughput is prioritized; or it can be micro-batch aggregation, which merges the profile snapshots of several objects according to fixed time slices and then evaluates them in batches to reduce peak pressure and improve computing power utilization.
[0195] In processing real-time feature values of economic attributes, fixed interval mapping can be used to facilitate cross-cycle stability; quantile mapping can be used to maintain the robustness of relative rank when the distribution shifts; or a business-defined discrete tier table can be used to ensure consistency with external strategy caliber.
[0196] For the sequence smoothing of real-time feature values in the service product association dimension, a sliding window balancing strategy can be used to balance sensitivity and stability; an exponential decay strategy can be used to give higher weight to recent fluctuations; and the maximum value retention within the window can be configured for strong time series fields to specifically capture extreme congestion or abnormal backlog.
[0197] For monotonic calibration of real-time comprehensive scoring, an offline-generated segmented mapping table can be used to ensure interpretability and replayability; a lightweight online calibration cache can be used, refreshed daily or by shift to adapt to business seasonality; or a group calibration can be used, using key tags in the profile as group keys to maintain separate calibration tables and threshold sets for different customer groups, thereby achieving differentiated judgment without changing the core process.
[0198] In implementing hysteresis for threshold comparison, one approach is to define a neighborhood dwell length, requiring the score to remain outside the neighborhood for several evaluation cycles before allowing a level switch; alternatively, one can increase a small safety boundary to expand the non-switchable zone of the threshold interval; or one can combine the object's historical level, adjusting the threshold in advance if a trend of higher levels is given for several consecutive cycles, in order to compress response latency.
[0199] In terms of output interpretation, contribution details and anomaly levels can be written together into the event stream or alarm table for downstream processing modules to read; intermediate quantities for playback can also be recorded at the same time, allowing the audit system to recalculate and score the same batch of inputs to verify consistency; confidence indicators can also be attached, using the proportion of data from missing data or truncation as a confidence deduction signal to guide whether manual review is needed later.
[0200] Example Explanation: In remote health management platforms within the healthcare field, a comprehensive risk warning mechanism can be built through interactive full-process data collection and phased penetration models. Throughout the entire interaction process between the user and the health service system, the platform sets up nodes for browsing, consultation, decision-making, service usage, and anomaly detection. During the browsing stage, users may view nutrition plans, exercise prescriptions, health product information, etc., and the system records the access type and frequency. During the consultation stage, users may ask questions through online Q&A, voice interaction, or mobile windows, and the system records the interaction channel and question category. During the decision-making stage, the system records the type of health service selected by the user and the time required for the decision. During the service usage stage, the platform records the number of health service calls and response cycles, such as the duration and frequency of remote nutritionist guidance. In the anomaly detection stage, the system records the reasons and processing results for health advice not taking effect, abnormal remote device data, or unresolved user feedback issues. All node data is integrated according to time series to generate a continuous data chain, and then basic attribute dimensions (such as user age group, regional tags), economic attribute dimensions (such as health subscription fees, health product consumption level), and service product related dimensions (such as exercise prescription execution frequency, remote follow-up frequency) are extracted to form a multi-dimensional dynamic profile.
[0201] Based on dynamic user profiling, a phased penetration model is established, dividing users into service interaction, service usage, and anomaly occurrence phases. A progressive migration relationship is established through a conversion strategy library. For example, the proportion of users in the health consultation phase who actually use health management services is statistically analyzed. If this proportion exceeds a set threshold, it indicates that the characteristics of this group have triggered migration conditions, and they enter the service usage phase. Similarly, if users in the service usage phase trigger an anomaly due to prolonged failure to achieve health goals or poor service experience, they enter the anomaly occurrence phase. By analyzing historical conversion rates and comparing thresholds, the weight parameters for group migration can be dynamically adjusted, enabling the model to adapt to changes in the user base under different health service scenarios.
[0202] In the first-level penetration analysis, the system, based on the phased penetration model, invokes the first migration relationship to extract economic attribute dimension feature datasets from the multi-dimensional dynamic profile of the service interaction stage group and the service usage stage group, such as willingness to pay for nutritional guidance and preference for exercise prescriptions. It then compares the distribution characteristics of the two groups and generates a feature distribution difference map. When the overlap of a certain economic attribute exceeds a threshold, for example, if a user shows a high willingness to pay during consultation and this is verified in the actual usage stage, that attribute is marked as a high overlap feature, ultimately forming a feature identification overlap set. The second-level penetration analysis, based on the second migration relationship, extracts service product association dimension features from the multi-dimensional dynamic profile of the service usage stage group and the anomaly occurrence stage group, such as response time for remote follow-up and alarm frequency of user health devices. It generates a second feature distribution difference map and calculates the difference degree. When a service feature deviates significantly from the distribution of the anomaly group, it is marked as a high difference feature, forming a feature identification difference set.
[0203] By combining feature overlap and difference with historical health data, the system generates an early warning model. Historical data includes statistics on health service satisfaction and repurchase rates corresponding to highly overlapping features, and abnormal feedback rates or health plan abandonment rates corresponding to highly differing features. By comparing these rates, the weight of each feature in the early warning model is determined, and high-risk, medium-risk, and low-risk thresholds are set. For example, a high-risk threshold corresponds to health service suspension or long-term abnormalities, a medium-risk threshold corresponds to periodic execution deviations, and a low-risk threshold corresponds to slight fluctuations. Finally, the overlap, difference, weight coefficients, and thresholds are combined to generate an early warning model that can be used for real-time judgment.
[0204] When new health interaction data enters the system, the current status of the dynamic profile is input into the early warning model. The system extracts real-time feature values for basic attribute dimensions (such as recent age group classification and regional differences), economic attribute dimensions (such as current health product consumption levels), and service product association dimensions (such as the number of recent follow-ups). The early warning model applies overlap weighting coefficients to economic dimension features and difference weighting coefficients to service dimension features, obtaining first and second weighted feature values respectively. These are then combined with the basic attribute feature values to generate a real-time comprehensive score. The comprehensive score is compared with high, medium, and low risk thresholds, outputting the corresponding anomaly level. If a high-risk anomaly level is output, it indicates that the user may be about to exhibit serious abnormal behavior in health management, such as persistent failure to achieve health goals or frequent triggering of health alerts; the system will trigger a high-risk warning signal. If the level is medium-risk, it may indicate that the user is experiencing periodic fluctuations, such as a decrease in the frequency of exercise plan execution. If the level is low-risk, only minor attention is required.
[0205] In this way, healthcare platforms can integrate the entire data chain from user browsing, consultation, use to anomalies. Combined with the hierarchical progression of the stage penetration model and the feature analysis of multi-dimensional dynamic profiles, they can accurately identify and warn of risks to user groups at different stages. This not only provides alerts before potential anomalies occur, but also provides quantifiable reference data for the allocation of health service resources.
[0206] In the fintech business, users generate a continuous data chain throughout their interaction with financial service platforms. This includes browsing financial product pages, consulting on investment or loan plans, making transaction decisions, using financial services such as payment or claims, and providing feedback or appeals when anomalies occur. The platform collects features at each stage to construct a continuous data chain covering browsing nodes (e.g., product browsing frequency, type of financial information access), consultation nodes (e.g., online customer service channel selection, type of question), transaction decision nodes (e.g., wealth management product selection cycle, loan application duration), service usage nodes (e.g., payment success rate, claims duration, account call frequency), and anomaly nodes (e.g., reasons for transaction failure, risk control interception results). The system then integrates this data into a time series, extracting basic attribute dimensions (customer age, region, credit rating), economic attribute dimensions (spending capacity, liquidity), and service product association dimensions (payment channel usage frequency, number of claims triggered), and constructing a multi-dimensional dynamic profile.
[0207] Based on dynamic customer profiling, a phased penetration model is established, dividing customers into three groups: service interaction stage, service usage stage, and anomaly occurrence stage. Historical transaction data is used to statistically analyze the migration ratios between these groups. The system defines migration strategies from service interaction to service usage in a conversion strategy library. For example, when a customer exhibits high activity during the browsing and consultation stage and ultimately converts into a transaction, the first migration relationship is satisfied. Similarly, migration strategies from service usage to anomaly occurrence are defined, such as when repeated delays occur in the claims process or transactions are blocked, satisfying the second migration relationship. By comparing actual conversion ratios with set thresholds, the model automatically activates relevant migration conditions and generates migration relationships, forming group sets, thereby revealing the progressive trajectory of customers from normal interaction to potential risk.
[0208] In the first-level penetration analysis, the system invokes the first migration relationship to extract economic attribute dimensions of the service interaction group and the service user group from the dynamic profile, such as fund activity and average monthly consumption level. It compares the distribution of fund flow characteristics between the two groups, generates a feature distribution difference map, and calculates the overlap value. When a certain economic attribute is highly consistent between the two groups, such as high-spending users being more likely to enter the actual transaction stage, this attribute is marked as a high overlap feature, ultimately forming a feature identification overlap set. In the second-level penetration analysis, the system invokes the second migration relationship to extract service product association characteristics between the service user group and the abnormal group, such as the number of cross-border payments, the proportion of failed claims, and the frequency of transaction freezes. It compares the distribution of the two groups, generates a second feature distribution difference map, and calculates the difference. When a certain dimension deviates significantly in the abnormal group, such as a significant increase in the abnormality rate of the high-frequency cross-border payment group, this attribute is marked as a high difference feature, forming a feature identification difference set.
[0209] The system then combines feature overlap, feature difference, and historical financial service data to generate an early warning model. In the historical data, the system calculates the complaint rate or transaction success rate of low-risk customers corresponding to high overlap features, and the anomaly rate or risk event trigger probability corresponding to high difference features. Feature weights are determined by the occurrence rate, and high, medium, and low risk thresholds are set based on the data distribution. For example, a high-risk threshold corresponds to a customer group with a transaction failure rate exceeding 70%, a medium-risk threshold corresponds to customers with a failure rate between 30% and 70%, and a low-risk threshold corresponds to customers with a failure rate below 30%. The resulting early warning model can make real-time judgments on newly input profile data.
[0210] When new data enters the system, the current status of the dynamic profile is input into the early warning model. The system extracts real-time values of basic attribute dimensions (such as customer credit score and active region), economic attribute dimensions (such as monthly fund inflows and outflows), and service product association dimensions (such as recent payment failure rate). The model combines the economic attribute dimension values with the overlap weight of feature identifiers, and the service product association dimension values with the difference weight of feature identifiers, generating a weighted result. This result is then combined with the real-time values of the basic attributes to form a real-time comprehensive score. This score is compared with a threshold to output an anomaly level, which may be high-risk, medium-risk, or low-risk. If a high-risk level is output, the platform immediately triggers an early warning signal, such as temporarily restricting the customer account, conducting manual review, or notifying the risk control team in real time. If a medium-risk level is output, a medium-priority early warning signal is triggered, prompting secondary verification or delayed settlement. If a low-risk level is output, a low-risk early warning signal is output, recorded only in the background for subsequent continuous tracking.
[0211] This embodiment extracts the current state of a multi-dimensional dynamic profile in a consistent manner and directly interfaces it with the weighting systems from the first and second-level penetration analyses. Controlled weighting is performed on the economic and service product dimensions respectively, and the results are aggregated into a real-time comprehensive score in a replayable manner. This score is then combined with a three-level threshold and boundary hysteresis set based on historical distribution to form a stable and sensitive grade output. This process maps the differences between static and dynamic profiles, and between usage and abnormal stages, into interpretable contributions, avoiding grade fluctuations caused by short-term noise while retaining the ability to respond to sudden anomalies. This improves the early identification rate at the same level of false alarms and provides traceable quantitative evidence for subsequent graded handling.
[0212] In one embodiment, an anomaly early warning device based on phased penetration is provided, which corresponds one-to-one with the anomaly early warning method based on phased penetration in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the anomaly early warning device based on phased penetration of the present invention. The modules include a data acquisition module 10, a model construction module 20, a first-level analysis module 30, a second-level analysis module 40, an early warning model generation module 50, an anomaly determination module 60, and an early warning triggering module 70. Detailed descriptions of each functional module are as follows:
[0213] The data acquisition module 10 is used to collect behavioral data of the object during the entire interaction process to form a continuous data chain, and to construct a multi-dimensional dynamic profile based on the continuous data chain.
[0214] The model building module 20 is used to establish a phased penetration model and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model.
[0215] The first-level analysis module 30 is used to perform first-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group.
[0216] The second-level analysis module 40 is used to perform second-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the feature identification difference between the service usage stage group and the anomaly occurrence stage group.
[0217] The early warning model generation module 50 is used to generate an early warning model based on the feature identifier overlap, the feature identifier difference, and historical data.
[0218] The anomaly determination module 60 is used to input the current state of the multi-dimensional dynamic profile into the early warning model and output the anomaly level;
[0219] The early warning triggering module 70 is used to trigger an early warning signal based on the abnormality level.
[0220] In one embodiment, the data acquisition module 10 is specifically used for:
[0221] Set up browsing, consultation, transaction decision-making, service usage, and error handling nodes throughout the entire interaction process;
[0222] In the browsing process nodes, the content access type characteristics and access frequency characteristics are recorded;
[0223] In the consultation process nodes, the characteristics of the interaction channel type and the characteristics of the question classification are recorded;
[0224] In the transaction decision-making stage, the product selection type characteristics and decision cycle characteristics are recorded;
[0225] In the service usage stage, record the service trigger count characteristics and service response cycle characteristics;
[0226] In the abnormal process node, the characteristics of the cause of the abnormality and the characteristics of the handling result are recorded;
[0227] Integrate the features of all node records in the time series to form a continuous data chain;
[0228] Based on the continuous data chain, basic attribute dimension features, economic attribute dimension features, and service product association dimension features are extracted, and a multi-dimensional dynamic profile is constructed based on the basic attribute dimension features, economic attribute dimension features, and service product association dimension features.
[0229] In one embodiment, the model building module 20 is specifically used for:
[0230] Create a phased penetration model framework, and define the initial scope of the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model framework;
[0231] Create a conversion strategy library, and define a first migration relationship conversion strategy in the conversion strategy library. The first migration relationship conversion strategy includes a conversion ratio threshold from the service interaction stage group to the service usage stage group and a first migration relationship feature matching condition.
[0232] A second migration relationship conversion strategy is defined in the conversion strategy library. The second migration relationship conversion strategy includes a conversion ratio threshold from the service usage stage group to the anomaly occurrence stage group and a second migration relationship feature matching condition.
[0233] Based on historical data, we statistically analyzed the actual conversion rate from the service interaction stage group to the service use stage group, and the actual conversion rate from the service use stage group to the anomaly occurrence stage group.
[0234] Compare the actual conversion rate with the corresponding conversion rate threshold;
[0235] When the actual conversion ratio from the service interaction stage group to the service usage stage group exceeds the corresponding conversion ratio threshold, the first migration relationship feature matching condition is activated.
[0236] When the actual conversion ratio from the service usage phase group to the anomaly occurrence phase group exceeds the corresponding conversion ratio threshold, the second migration relationship feature matching condition is activated.
[0237] Generate conversion strategy weight parameters based on activated feature matching conditions;
[0238] Based on the initial scope of the group in the service interaction stage, the conversion strategy weight parameters, and the feature matching conditions of the activated first migration relationship, a first migration relationship is constructed.
[0239] Based on the initial scope of the service usage stage group, the conversion strategy weight parameters, and the activated second migration relationship feature matching conditions, a second migration relationship is constructed;
[0240] Generate a service usage phase group set based on the first migration relationship;
[0241] Generate a set of groups at the anomaly occurrence stage based on the second migration relationship;
[0242] A service interaction phase group set is generated based on the initial scope of the service interaction phase group.
[0243] In one embodiment, the first-level analysis module 30 is specifically used for:
[0244] Invoke the first migration relationship in the phased penetration model;
[0245] Based on the first migration relationship, the economic attribute dimension feature dataset of the service interaction stage group and the economic attribute dimension feature dataset of the service usage stage group are extracted from the multi-dimensional dynamic profile.
[0246] Compare the economic attribute dimension feature dataset of the service interaction stage group with the economic attribute dimension feature dataset of the service usage stage group to generate a first feature distribution difference map.
[0247] Based on the first feature distribution difference map, determine the feature overlap value of each economic attribute dimension;
[0248] When the feature overlap value of a certain economic attribute dimension exceeds a preset threshold, the economic attribute dimension is marked as a high overlap feature identifier.
[0249] The high overlap feature identifiers are associated with the first migration relationship, and the set of high overlap feature identifiers is summarized to form the feature identifier overlap degree.
[0250] In one embodiment, the second-level analysis module 40 is specifically used for:
[0251] Invoke the second migration relationship in the phased penetration model;
[0252] Based on the second migration relationship, the service product association dimension feature dataset of the service usage stage group and the service product association dimension feature dataset of the anomaly occurrence stage group are extracted from the multi-dimensional dynamic profile.
[0253] Compare the service product association dimension feature dataset of the service usage stage group with the service product association dimension feature dataset of the anomaly occurrence stage group to generate a second feature distribution difference map.
[0254] Based on the second feature distribution difference map, determine the feature difference value of each service product association dimension;
[0255] When the feature difference value of a certain service product association dimension exceeds a preset threshold, the service product association dimension is marked as a high difference feature identifier;
[0256] The high-discrepancy feature identifiers are associated with the second migration relationship, and the set of high-discrepancy feature identifiers is summarized to form the feature identifier discrepancy.
[0257] In one embodiment, the early warning model generation module 50 is specifically used for:
[0258] Based on historical data, the occurrence rate of overlap complaints corresponding to the overlap of the aforementioned feature identifiers is statistically analyzed.
[0259] Based on historical data, the occurrence rate of complaints corresponding to the difference in the aforementioned feature identifiers is statistically analyzed.
[0260] The overlap ratio weighting coefficient of the feature identifier is determined based on the overlap complaint incidence rate.
[0261] The feature identification difference weight coefficient is determined based on the difference complaint incidence rate.
[0262] High-risk, medium-risk, and low-risk thresholds are set based on historical data distribution.
[0263] The early warning model is generated by combining the feature overlap degree, feature overlap degree weight coefficient, feature difference degree, feature difference degree weight coefficient, historical data, high-risk threshold, medium-risk threshold, and low-risk threshold.
[0264] In one embodiment, the anomaly determination module 60 is specifically used for:
[0265] Extract real-time feature values of basic attribute dimension, economic attribute dimension, and service product association dimension from the current state of the multi-dimensional dynamic profile;
[0266] The feature overlap coefficient in the early warning model is used to weight the real-time feature value of the economic attribute dimension to generate a first weighted feature value.
[0267] The feature identification difference weighting coefficient in the early warning model is used to weight the real-time feature values of the service product association dimension to generate a second weighted feature value.
[0268] The real-time feature values of the basic attribute dimensions, the first weighted feature value, and the second weighted feature value are combined to generate a real-time comprehensive score;
[0269] The real-time comprehensive score is compared with the high-risk threshold, medium-risk threshold, and low-risk threshold in the early warning model;
[0270] The output will indicate whether the anomaly is high-risk, medium-risk, or low-risk.
[0271] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a phased penetration-based anomaly early warning method on the server side.
[0272] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a phased penetration-based anomaly warning method on the user side.
[0273] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0274] The collected behavioral data of the object throughout the entire interaction process forms a continuous data chain, and a multi-dimensional dynamic profile is constructed based on the continuous data chain;
[0275] Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model;
[0276] Based on the staged penetration model and the multi-dimensional dynamic profile, perform the first-level penetration analysis to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group;
[0277] Based on the staged penetration model and the multi-dimensional dynamic profile, a second-level penetration analysis is performed to extract the feature identification differences between the service usage stage group and the anomaly occurrence stage group.
[0278] An early warning model is generated based on the overlap of the feature identifiers, the differences between the feature identifiers, and historical data.
[0279] The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output.
[0280] An early warning signal is triggered based on the aforementioned anomaly level.
[0281] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0282] The collected behavioral data of the object throughout the entire interaction process forms a continuous data chain, and a multi-dimensional dynamic profile is constructed based on the continuous data chain;
[0283] Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model;
[0284] Based on the staged penetration model and the multi-dimensional dynamic profile, perform the first-level penetration analysis to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group;
[0285] Based on the staged penetration model and the multi-dimensional dynamic profile, a second-level penetration analysis is performed to extract the feature identification differences between the service usage stage group and the anomaly occurrence stage group.
[0286] An early warning model is generated based on the overlap of the feature identifiers, the differences between the feature identifiers, and historical data.
[0287] The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output.
[0288] An early warning signal is triggered based on the aforementioned anomaly level.
[0289] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0290] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0291] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0292] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An anomaly early warning method based on phased penetration, characterized in that, Includes the following steps: The collected behavioral data of the object throughout the entire interaction process forms a continuous data chain, and a multi-dimensional dynamic profile is constructed based on the continuous data chain; Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model; Based on the staged penetration model and the multi-dimensional dynamic profile, perform the first-level penetration analysis to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group; Based on the staged penetration model and the multi-dimensional dynamic profile, a second-level penetration analysis is performed to extract the feature identification differences between the service usage stage group and the anomaly occurrence stage group. An early warning model is generated based on the overlap of the feature identifiers, the differences between the feature identifiers, and historical data. The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output. An early warning signal is triggered based on the aforementioned anomaly level.
2. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, The collected behavioral data of the subject throughout the entire interaction process forms a continuous data chain, and a multi-dimensional dynamic profile is constructed based on the continuous data chain, including: Set up browsing, consultation, transaction decision-making, service usage, and error handling nodes throughout the entire interaction process; In the browsing process nodes, the content access type characteristics and access frequency characteristics are recorded; In the consultation process nodes, the characteristics of the interaction channel type and the characteristics of the question classification are recorded; In the transaction decision-making stage, the product selection type characteristics and decision cycle characteristics are recorded; In the service usage stage, record the service trigger count characteristics and service response cycle characteristics; In the abnormal process node, the characteristics of the cause of the abnormality and the characteristics of the handling result are recorded; Integrate the features of all node records in the time series to form a continuous data chain; Based on the continuous data chain, basic attribute dimension features, economic attribute dimension features, and service product association dimension features are extracted, and a multi-dimensional dynamic profile is constructed based on the basic attribute dimension features, economic attribute dimension features, and service product association dimension features.
3. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, Establish a phased penetration model, and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model, including: Create a phased penetration model framework, and define the initial scope of the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model framework; Create a conversion strategy library, and define a first migration relationship conversion strategy in the conversion strategy library. The first migration relationship conversion strategy includes a conversion ratio threshold from the service interaction stage group to the service usage stage group and a first migration relationship feature matching condition. A second migration relationship conversion strategy is defined in the conversion strategy library. The second migration relationship conversion strategy includes a conversion ratio threshold from the service usage stage group to the anomaly occurrence stage group and a second migration relationship feature matching condition. Based on historical data, we statistically analyzed the actual conversion rate from the service interaction stage group to the service use stage group, and the actual conversion rate from the service use stage group to the anomaly occurrence stage group. Compare the actual conversion rate with the corresponding conversion rate threshold; When the actual conversion ratio from the service interaction stage group to the service usage stage group exceeds the corresponding conversion ratio threshold, the first migration relationship feature matching condition is activated. When the actual conversion ratio from the service usage phase group to the anomaly occurrence phase group exceeds the corresponding conversion ratio threshold, the second migration relationship feature matching condition is activated. Generate conversion strategy weight parameters based on activated feature matching conditions; Based on the initial scope of the group in the service interaction stage, the conversion strategy weight parameters, and the feature matching conditions of the activated first migration relationship, a first migration relationship is constructed. Based on the initial scope of the service usage stage group, the conversion strategy weight parameters, and the activated second migration relationship feature matching conditions, a second migration relationship is constructed; Generate a service usage phase group set based on the first migration relationship; Generate a set of groups at the anomaly occurrence stage based on the second migration relationship; A service interaction phase group set is generated based on the initial scope of the service interaction phase group.
4. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, Based on the staged penetration model and the multi-dimensional dynamic profile, a first-level penetration analysis is performed to extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group, including: Invoke the first migration relationship in the phased penetration model; Based on the first migration relationship, the economic attribute dimension feature dataset of the service interaction stage group and the economic attribute dimension feature dataset of the service usage stage group are extracted from the multi-dimensional dynamic profile. Compare the economic attribute dimension feature dataset of the service interaction stage group with the economic attribute dimension feature dataset of the service usage stage group to generate a first feature distribution difference map. Based on the first feature distribution difference map, determine the feature overlap value of each economic attribute dimension; When the feature overlap value of a certain economic attribute dimension exceeds a preset threshold, the economic attribute dimension is marked as a high overlap feature identifier. The high overlap feature identifiers are associated with the first migration relationship, and the set of high overlap feature identifiers is summarized to form the feature identifier overlap degree.
5. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, Based on the staged penetration model and the multi-dimensional dynamic profile, a second-level penetration analysis is performed to extract the feature differences between the service usage stage group and the anomaly occurrence stage group, including: Invoke the second migration relationship in the phased penetration model; Based on the second migration relationship, the service product association dimension feature dataset of the service usage stage group and the service product association dimension feature dataset of the anomaly occurrence stage group are extracted from the multi-dimensional dynamic profile. Compare the service product association dimension feature dataset of the service usage stage group with the service product association dimension feature dataset of the anomaly occurrence stage group to generate a second feature distribution difference map. Based on the second feature distribution difference map, determine the feature difference value of each service product association dimension; When the feature difference value of a certain service product association dimension exceeds a preset threshold, the service product association dimension is marked as a high difference feature identifier; The high-discrepancy feature identifiers are associated with the second migration relationship, and the set of high-discrepancy feature identifiers is summarized to form the feature identifier discrepancy.
6. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, A warning model is generated based on the overlap of the feature identifiers, the difference of the feature identifiers, and historical data, including: Based on historical data, the occurrence rate of overlap complaints corresponding to the overlap of the aforementioned feature identifiers is statistically analyzed. Based on historical data, the occurrence rate of complaints corresponding to the difference in the aforementioned feature identifiers is statistically analyzed. The overlap ratio weighting coefficient of the feature identifier is determined based on the overlap complaint incidence rate. The feature identification difference weight coefficient is determined based on the difference complaint incidence rate. High-risk, medium-risk, and low-risk thresholds are set based on historical data distribution. The early warning model is generated by combining the feature overlap degree, feature overlap degree weight coefficient, feature difference degree, feature difference degree weight coefficient, historical data, high-risk threshold, medium-risk threshold, and low-risk threshold.
7. The anomaly early warning method based on phased penetration as described in claim 1, characterized in that, The current state of the multi-dimensional dynamic profile is input into the early warning model, and the anomaly level is output, including: Extract real-time feature values of basic attribute dimension, economic attribute dimension, and service product association dimension from the current state of the multi-dimensional dynamic profile; The feature overlap coefficient in the early warning model is used to weight the real-time feature value of the economic attribute dimension to generate a first weighted feature value. The feature identification difference weighting coefficient in the early warning model is used to weight the real-time feature values of the service product association dimension to generate a second weighted feature value. The real-time feature values of the basic attribute dimensions, the first weighted feature value, and the second weighted feature value are combined to generate a real-time comprehensive score; The real-time comprehensive score is compared with the high-risk threshold, medium-risk threshold, and low-risk threshold in the early warning model; The output will indicate whether the anomaly is high-risk, medium-risk, or low-risk.
8. An anomaly early warning device based on phased penetration, characterized in that, The anomaly early warning device based on phased penetration includes: The data acquisition module is used to collect behavioral data of the object throughout the entire interaction process to form a continuous data chain, and to build a multi-dimensional dynamic profile based on the continuous data chain; The model building module is used to establish a phased penetration model and define the progressive migration relationship between the service interaction phase group, the service usage phase group, and the anomaly occurrence phase group in the phased penetration model. The first-level analysis module is used to perform first-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the overlap of feature identifiers between the service interaction stage group and the service usage stage group. The second-level analysis module is used to perform second-level penetration analysis based on the stage penetration model and the multi-dimensional dynamic profile, and extract the feature identification difference between the service usage stage group and the anomaly occurrence stage group. The early warning model generation module is used to generate an early warning model based on the overlap of the feature identifiers, the difference of the feature identifiers, and historical data. The anomaly detection module is used to input the current state of the multi-dimensional dynamic profile into the early warning model and output the anomaly level. The early warning triggering module is used to trigger an early warning signal based on the anomaly level.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a phased penetration-based anomaly warning program stored in the memory and executable on the processor, wherein the phased penetration-based anomaly warning program, when executed by the processor, implements the steps of the phased penetration-based anomaly warning method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an anomaly warning program based on phased penetration, which, when executed by a processor, implements the steps of the anomaly warning method based on phased penetration as described in any one of claims 1-7.
Citation Information
Patent Citations
User equipment identification method and device and computer equipment
CN113570222A
Financial Internet of Things platform equipment early warning method and device
CN115688110A
Cited By
Group abnormal transaction behavior identification method and device, medium and electronic equipment
CN121685128A
Intelligent supplier portrait generation system
CN121745984A