A deep learning agent collaborative decision method and system for operator risk control

CN122549929APending Publication Date: 2026-08-11BEIJING TIANYUAN DIKE NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0010]为此,本发明提供一种面向运营商风控的深度学习代理协同决策方法及系统,解决现有技术存在的联合表征不足、处置闭环缺失、约束适配能力弱、迭代优化机制缺失的问题

Benefits of technology

[0067]第一,构建多层级风险表征体系,实现运营商多源异构风控数据的统一联合编码;通过风险关系图与图传播机制,突破传统单点异常识别局限,可识别团伙协同欺诈等跨实体复杂风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549929A_ABST
    Figure CN122549929A_ABST
Patent Text Reader

Abstract

The application discloses a deep learning agent collaborative decision method and system for operator risk control, and belongs to the technical field of communication risk control and artificial intelligence. In view of the defects of the prior art, such as insufficient representation of multi-source heterogeneous data, disconnection of risk identification and disposal, lack of business constraint correction and feedback optimization, a multi-level risk representation is generated through joint coding of multi-source data, and single-point and multi-entity collaborative anomaly is identified in combination with a risk relationship graph. A serialized candidate action sequence is generated through an intelligent agent, and an executable risk control strategy is output after multi-dimensional constraint correction. The risk representation model and the decision strategy are updated based on execution feedback. The application realizes full-process processing of risk control, takes into account risk control effect, compliance requirements and customer experience, and is suitable for high-concurrency real-time risk control scenarios of operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of communication network security, operator business risk control and artificial intelligence technology, and in particular relates to a deep learning agent collaborative decision-making method and system for operator risk control. Background Technology

[0002] With the deepening of the digital transformation of telecommunications services, operators have become comprehensive business service providers covering individual users, government and enterprise customers, and channel agents. Their risk control scenarios are no longer limited to traditional network security protection, but extend to the entire business chain, including business processing, account security, channel compliance, and anti-fraud. Unlike traditional internet single-point account risk control and network security alarm scenarios, operator risk control covers multiple entities such as users, numbers, devices, business orders, sessions, and contacts. The risk forms include both abnormal operations of a single entity and group-style collaborative risks formed by multiple entities around the same device, the same geographical area, and the same business chain. At the same time, the output of operator risk control is not limited to the identification result of "whether it is abnormal", but must further generate executable actions that meet regulatory requirements, business continuity requirements, and customer service level requirements. This places far higher demands on the decision-making capabilities of the risk control system than traditional risk control solutions.

[0003] Currently, existing risk control technologies have formed relatively mature technical routes in the fields of behavior vectorization, anomaly identification, and risk scoring: The first type of technical route revolves around event vectorization and composite anomaly scoring. By extracting event vectors from network behavior logs and business events, composite anomaly scores are obtained based on time window statistics, context comparison, and hierarchical weighted fusion. These scores are used to determine whether behavior deviates from historical baselines. This type of solution provides a standardized event representation and scoring framework for behavior anomaly identification and has been widely used in network security and single-point behavior analysis. The second type of technical route focuses on the aggregation and generation of user behavior vectors. By encoding and aggregating discrete, cross-time period user interaction behaviors, access trajectories, and operation sequences, a unified user-level behavior vector is generated for subsequent classification, clustering, and anomaly identification. This provides a feasible aggregation approach for long-term user-level behavior profiling and risk modeling. In addition, existing technologies have also proposed a unified vectorization scheme for multiple types of interactive behaviors. By structurally encoding browsing trajectories and operation sequences, a unified vector expression of heterogeneous interactive data is achieved, providing a reference basis for the joint analysis of multi-source data.

[0004] However, when the aforementioned existing technologies are directly applied to operator risk control scenarios, the following insurmountable technical defects still exist:

[0005] First, existing technologies cannot meet the joint representation needs of operators' multi-source heterogeneous business data. The data sources for operators' risk control not only include communication behavior data, but also cover a variety of heterogeneous data such as account operations, business processing, device association, historical interactions, historical handling and response feedback. These various types of data have problems such as inconsistent time granularity, large differences in field distribution, coexistence of structured and unstructured data, and multiple perspectives on the same entity. Existing vectorization solutions are mostly based on single-type behavior flow or event flow, and cannot form a unified, closed-loop, multi-level risk representation that can be used for real-time decision-making at the entity level, behavior level, and event level, let alone achieve unified modeling of cross-entity related risks.

[0006] Second, existing technologies only go as far as anomaly identification and scoring output, failing to form a complete decision-making loop for risk control. The core requirement of operator risk control is not merely outputting risk scores, but generating executable actions within a very short latency. For example, different actions such as secondary verification, business restrictions, delayed execution, manual review, and observation and release may be adopted for different risk scenarios. It may even be necessary to generate a sequence of actions with sequential dependencies. However, existing technologies only output risk labels or anomaly scores, failing to transform the risk characterization results into a sequence of candidate actions adapted to the business context. This results in the identification results not being directly implemented into executable risk control strategies.

[0007] Third, existing technologies do not consider the multiple business constraints under the risk control scenarios of operators, and cannot achieve dynamic correction and compliance adaptation of actions. The execution of risk control actions by operators is subject to multiple constraints such as regulatory rules, business continuity requirements, and customer service level requirements. For example, regulatory requirements prohibit the direct approval of high-risk operations for specific businesses, core communication capabilities cannot be arbitrarily blocked, and high-value customers require more cautious intervention methods. However, existing technologies focus on anomaly detection as their core objective and have not established a constraint correction mechanism after candidate actions are generated. This results in the output decision results being unable to adapt to the compliance and service requirements of the real operating environment, which can easily lead to compliance risks, business interruptions, or customer complaints.

[0008] Fourth, existing technologies lack a joint optimization loop based on execution feedback, making it impossible to achieve continuous iteration of risk control capabilities. The ultimate effectiveness of operator risk control depends on the combined effect of the accuracy of risk identification and the rationality of handling actions. Accurate identification but aggressive actions can cause false positives, while conservative actions can lead to missed risks. However, existing technologies only optimize the anomaly identification model and do not establish a joint update mechanism that simultaneously feeds back user responses, appeal results, misjudgment corrections, and subsequent risk changes to the risk representation model and action decision-making strategies. This fails to achieve synergistic optimization of identification and handling capabilities, resulting in the system being unable to adapt to constantly evolving risk patterns after long-term operation.

[0009] In summary, existing technologies cannot meet the complex needs of operators' risk control scenarios. There is an urgent need for a deep learning agent collaborative decision-making scheme that can achieve unified representation of multi-source data, collaborative risk identification, intelligent action decision-making, dynamic correction of constraints, and joint optimization of feedback. This would solve the problems of insufficient joint representation, lack of closed-loop processing, weak constraint adaptation capability, and lack of iterative optimization mechanism in existing technologies. Summary of the Invention

[0010] To address these issues, this invention provides a deep learning agent collaborative decision-making method and system for operator risk control, solving the problems of insufficient joint representation, lack of closed-loop processing, weak constraint adaptation capability, and lack of iterative optimization mechanism in existing technologies.

[0011] To achieve the above objectives, the present invention provides the following technical solution: a deep learning agent collaborative decision-making method for operator risk control, applied to an operator risk control platform, comprising the following steps:

[0012] S1. Acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event;

[0013] S2. Based on the preprocessed multi-source heterogeneous risk control data, generate multi-level risk representation vectors at the entity level, behavior level, and event level. Construct a risk relationship graph based on the relationship between entities. Identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation. Output multi-modal fusion risk representation and local risk subgraph.

[0014] S3. Input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module to generate several candidate action sequences with action dependencies in the preset action space.

[0015] S4. Perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution;

[0016] S5. Collect feedback data after the risk control strategy is executed, construct monitoring signals and reinforcement rewards based on the feedback data, and jointly update the risk characterization model and intelligent agent decision-making strategy.

[0017] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, step S1 involves performing standardized preprocessing on the multi-source heterogeneous risk control data, specifically including:

[0018] The numerical features in the multi-source heterogeneous risk control data are subjected to unified standardization processing. Normalization calculation is performed based on the mean and variance of each field within the training statistical window to eliminate the differences in the units and ranges of values ​​of different fields.

[0019] Time alignment weights are calculated for historical records associated with risk events, and an exponential decay mechanism is adopted, with higher weights for historical records that are closer to the current risk event time interval, thus weakening the impact of weakly correlated behaviors on the current risk assessment.

[0020] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, step S2 generates multi-level risk representation vectors at the entity, behavior, and event levels based on the preprocessed multi-source heterogeneous risk control data. Specifically, this includes:

[0021] Joint encoding is performed on the entity's identity features, statistical features, and context features. The concatenated identity features, statistical features, and context features are then input into the entity-level encoding network and mapped to a unified vector space to generate an entity-level risk representation vector.

[0022] The behavioral sequences corresponding to risk events are time-weighted and encoded. The basic encoding results of each record in the behavioral sequence are aggregated according to time alignment weights and then input into the behavioral-level temporal coding network to generate a behavioral-level risk representation vector.

[0023] The aggregated entity vector, behavioral risk representation, and context vector of all entities involved in the event are fused and encoded. The aggregated entity vector, behavioral risk representation, and context vector are concatenated and input into the event-level encoding network to generate the event-level risk representation vector.

[0024] Weighted fusion is performed on the historical behavior vector and historical feedback statistical vector obtained from the encoding of event-level risk representation, historical interaction and disposal records, to generate a multimodal fusion risk representation vector.

[0025] As a preferred solution for deep learning-based agent collaborative decision-making methods for operator risk control, step S2 involves constructing a risk relationship graph based on the relationships between entities, specifically including:

[0026] The edge weights between entity nodes in the risk relationship graph are calculated by combining the basic association strength of the corresponding relationship types between nodes with the time decay factor. The closer the occurrence time of the association behavior between nodes, the higher the weight of the corresponding edge.

[0027] The node representation is updated by performing multi-level graph propagation. In each layer of propagation, the features of all neighboring nodes of the current node are aggregated, and after being combined with the node's own features and processed by the activation function, the hidden representation of the node in the current layer is generated.

[0028] Extract local risk subgraphs corresponding to risk events from the global risk relationship graph, and retain nodes whose graph distance from the core node of the event does not exceed the preset maximum number of hops and whose overall node connection strength is not lower than the preset screening threshold to form local risk subgraphs that are strongly associated with the current event.

[0029] The construction of the risk relationship graph also includes constructing a multi-relationship adjacency matrix, summing the edge weights of each pair of nodes under all relationship types to obtain the comprehensive adjacency weight between nodes, which serves as the input for graph propagation;

[0030] In step S2, single-point anomalies and multi-entity collaborative anomalies are identified through graph propagation calculations, specifically including:

[0031] A single-point anomaly score is calculated based on a multimodal fusion risk representation vector, and the score network is input to output the degree of deviation of the event itself from the historical baseline.

[0032] The collaboration anomaly score is calculated based on the node representation of the local risk subgraph. After pooling and aggregating the final hidden representation of all nodes in the subgraph, it is input into the collaboration scoring network and the anomaly degree of multi-entity collaboration behavior is output.

[0033] A comprehensive anomaly score is obtained by fusing single-point anomaly scores and collaborative anomaly scores. The fusion weights of the two types of scores are all non-negative and the sum of the weights is 1.

[0034] Based on the comprehensive anomaly score, the completeness of the event input data, and the structural stability of the local risk subgraph, the risk confidence level of the current risk assessment is calculated, reflecting the credibility of the risk evidence.

[0035] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, in step S3, the intelligent agent decision-making module constructs the decision state of the intelligent agent during the process of generating several candidate action sequences with action dependencies in the preset action space, including:

[0036] The multimodal fusion risk representation vector, the pooled representation of the local risk subgraph, the historical strategy effect encoding results, the business constraint encoding results, and the risk confidence are concatenated to form a unified decision state vector;

[0037] The policy probability of a single action is calculated based on the decision state vector. After the input policy network is normalized, the selection probability of each possible action is output.

[0038] The executableness of single-action strategy probability execution under regulatory and business constraints is masked. Based on the set of regulatory rules and the set of business constraints, the probability of actions that do not meet the constraints is suppressed, and compliant actions are retained to participate in the generation of subsequent candidate sequences.

[0039] Perform executability constraint projection on the candidate action sequence, and modify the original candidate action sequence according to regulatory rules, business constraints, manual review capacity limit and resource budget limit, eliminate unexecutable action combinations, and output the valid action sequence that meets the execution conditions;

[0040] After an action is executed, a decision state transition is performed. Based on the current decision state, the actions already executed, the intermediate results of the action execution, or observations, the decision state for the next step is updated to achieve sequential dynamic decision-making.

[0041] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, step S3 generates candidate action sequences with action dependencies, specifically including:

[0042] The sequenced candidate action sequence is generated based on the single-action strategy probability. The conditional selection probability of each action is multiplied according to the execution order of the actions to obtain the joint generation probability of the entire candidate action sequence.

[0043] The overall utility of each candidate action sequence is calculated. The overall value score of the sequence is obtained by combining the estimated risk reduction benefit, total cost of action execution, user impact range coefficient, and business importance coefficient with preset weights.

[0044] The estimated risk reduction benefit of the candidate action sequence is calculated by subtracting the expected posterior risk score after executing the action sequence from the comprehensive risk score before the action is executed. The expected posterior risk score is obtained based on the statistical analysis of the execution effects of similar strategies in the past.

[0045] The total execution cost of the candidate action sequence is obtained by summing the system execution cost, manual processing cost, and customer service cost of each action in the sequence.

[0046] The user influence range coefficient of the candidate action sequence is the sum of the basic influence intensity of each action in the sequence and the corresponding customer service level weight. The higher the customer service level, the greater the influence coefficient of the same action.

[0047] As a preferred solution for deep learning agent collaborative decision-making method for operator risk control, in step S4, the candidate action sequence is ranked by cost and benefit. Based on the comprehensive utility of the candidate action sequence, the penalty items that violate regulatory preferences, the penalty items that affect business continuity, and the penalty items that affect customer service experience are subtracted to obtain the final ranking score.

[0048] The admission criteria for directly releasing candidate action sequences are as follows: only when the overall anomaly score of an event is lower than the preset release risk threshold and the risk confidence level is lower than the preset high confidence risk threshold, the candidate action sequence that is directly released can enter the final ranking; otherwise, it will be blocked or downweighted.

[0049] The capacity constraint rule for candidate action sequences that include manual review is: a candidate action sequence that includes manual review can only enter the final sorting if the sum of the number of tasks in the current manual review queue and the new manual review task requirements of the candidate action sequence does not exceed the upper limit of manual review capacity.

[0050] The execution delay constraint rule for candidate action sequences is: a candidate action sequence can only enter the final sorting if the estimated execution delay does not exceed the system's preset real-time decision delay threshold.

[0051] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, step S4 also includes performing hierarchical constraint correction for different business scenarios. For different business scenarios, different scenario mapping weights are configured for regulatory constraints, business continuity constraints, and customer service level constraints, and the total constraint correction intensity under the current scenario is calculated. Among them, the weight of customer service level constraints is increased for high-value customer scenarios, the weight of regulatory constraints is increased for key regulatory business scenarios, and the default weight configuration with the best cost benefit is adopted for ordinary business scenarios.

[0052] As a preferred scheme for deep learning agent collaborative decision-making method for operator risk control, step S5 involves constructing a supervision signal and reinforcement reward based on the feedback data, and jointly updating the risk representation model and intelligent agent decision-making strategy, specifically including:

[0053] The combined reward value is calculated based on the execution feedback data, taking into account the risk prevention effect, misjudgment, changes in business losses, and additional delays introduced by the handling, and is calculated by weighting according to preset weights.

[0054] A loss function for the risk characterization model is constructed, which combines risk category supervision loss and anomaly scoring calibration loss to optimize the accuracy of risk identification.

[0055] A loss function for intelligent agent strategy is constructed, which combines execution reward and the generation probability of action sequence to improve the selection probability of high-yield action sequence;

[0056] A joint loss function is constructed, which sums the risk representation loss, policy loss, and model regularization term according to preset weights to achieve joint optimization of the risk representation model and the intelligent agent strategy.

[0057] The joint update adopts a controlled incremental update rule. When the confidence interval width of the feedback estimate does not exceed a preset threshold, the model parameters are updated based on the gradient of the joint loss at a preset step size to filter out the interference of low-quality and unstable feedback on the model.

[0058] As a preferred solution for deep learning agent collaborative decision-making methods for operator risk control, it also includes a full-link anomaly degradation handling mechanism:

[0059] When some data sources are missing in multi-source heterogeneous risk control data, a missing identifier vector and confidence reduction mechanism are used to perform downgraded inference; when the estimated execution delay of a candidate action sequence exceeds the preset real-time decision threshold, the length of the action sequence is pruned or switched to a low-resource alternative action; when the manual review capacity reaches the preset limit, the cost weight of the manual review action is dynamically increased, and candidate action sequences with low resource consumption are output first.

[0060] This invention also provides a deep learning agent collaborative decision-making system for operator risk control, used to execute the above-mentioned deep learning agent collaborative decision-making method for operator risk control, including:

[0061] The multi-source data access and preprocessing unit is used to acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event.

[0062] The multi-level risk characterization construction unit is used to generate multi-level risk characterization vectors at the entity level, behavior level, and event level based on the preprocessed multi-source heterogeneous risk control data, construct a risk relationship graph based on the relationship between entities, identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation, and output multi-modal fusion risk characterization and local risk subgraphs.

[0063] The intelligent agent action generation and collaborative decision-making unit is used to input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module, and generate a number of candidate action sequences with action dependencies in the preset action space.

[0064] The action orchestration and execution unit is used to perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution;

[0065] The feedback learning and joint update unit is used to collect feedback data after the risk control strategy is executed, construct supervision signals and reinforcement rewards based on the feedback data, and jointly update the risk representation model and the intelligent agent decision-making strategy.

[0066] The present invention has the following advantages:

[0067] First, a multi-level risk characterization system is constructed to achieve unified joint coding of multi-source heterogeneous risk control data of operators; through risk relationship diagrams and graph propagation mechanisms, the limitations of traditional single-point anomaly identification are broken, and complex cross-entity risks such as gang-related fraud can be identified.

[0068] Second, by outputting a sequence of candidate actions with dependencies through the intelligent agent module, the traditional single risk score output is replaced, thus connecting the risk identification-strategy generation-action execution chain, solving the pain point of disconnect between identification and handling, and improving the feasibility and flexibility of risk control strategies.

[0069] Third, regulatory rules, business continuity requirements, and customer service level requirements will be incorporated into the action constraint correction framework. Differentiated strategies will be implemented for different business scenarios to avoid a "one-size-fits-all" approach and to balance risk control effectiveness, business stability, and customer experience.

[0070] Fourth, establish a feedback-driven joint update mechanism to simultaneously optimize risk characterization models and decision-making strategies based on feedback from the entire risk management process. This breaks through the limitations of traditional methods that only optimize identification accuracy, enabling a synergistic improvement in identification and management capabilities and allowing for adaptive adaptation to the evolution of risk patterns.

[0071] Fifth, a full-link anomaly degradation and real-time guarantee mechanism is designed, which can meet the millisecond-level real-time decision-making requirements in high-concurrency scenarios and can still run stably in abnormal scenarios such as missing data sources and resource overload, making the system highly robust. Attached Figure Description

[0072] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0073] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0074] Figure 1 This is a schematic diagram of the deep learning agent collaborative decision-making method for operator risk control provided in an embodiment of the present invention;

[0075] Figure 2 This is a schematic diagram of the multi-source data access and preprocessing process provided in an embodiment of the present invention;

[0076] Figure 3 This is a schematic diagram of the multi-level risk characterization and risk relationship diagram construction process provided in the embodiments of the present invention;

[0077] Figure 4 This is a schematic diagram of the intelligent agent candidate action generation and collaborative decision-making process provided in an embodiment of the present invention;

[0078] Figure 5 This is a schematic diagram of the action orchestration, constraint correction, and execution process provided in the embodiments of the present invention;

[0079] Figure 6 This is a schematic diagram of the feedback learning and joint update process provided in the embodiments of the present invention;

[0080] Figure 7 This is a schematic diagram of the architecture of a deep learning agent collaborative decision-making system for operator risk control provided in an embodiment of the present invention. Detailed Implementation

[0081] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0082] Example 1

[0083] See Figure 1 Embodiment 1 of the present invention provides a deep learning agent collaborative decision-making method for operator risk control, applied to an operator risk control platform, including the following steps:

[0084] S1. Acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event;

[0085] S2. Based on the preprocessed multi-source heterogeneous risk control data, generate multi-level risk representation vectors at the entity level, behavior level, and event level. Construct a risk relationship graph based on the relationship between entities. Identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation. Output multi-modal fusion risk representation and local risk subgraph.

[0086] S3. Input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module to generate several candidate action sequences with action dependencies in the preset action space.

[0087] S4. Perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution;

[0088] S5. Collect feedback data after the risk control strategy is executed, construct monitoring signals and reinforcement rewards based on the feedback data, and jointly update the risk characterization model and intelligent agent decision-making strategy.

[0089] In this embodiment, step S1 involves performing standardized preprocessing on the multi-source heterogeneous risk control data, specifically including:

[0090] The numerical features in the multi-source heterogeneous risk control data are subjected to unified standardization processing. Normalization calculation is performed based on the mean and variance of each field within the training statistical window to eliminate the differences in the units and ranges of values ​​of different fields.

[0091] Time alignment weights are calculated for historical records associated with risk events, and an exponential decay mechanism is adopted, with higher weights for historical records that are closer to the current risk event time interval, thus weakening the impact of weakly correlated behaviors on the current risk assessment.

[0092] See Figure 2 Specifically, in step S1, the multi-source data is standardized to address issues such as inconsistent field units, large differences in value ranges, and different encodings of similar fields across systems. The standardized result serves as the basic input for joint encoding.

[0093] Formula (1) is used to represent the unified standardization process of multi-source features:

[0094]

[0095] in, Indicates an event In the The original feature values ​​on a numeric field Represents the standardized eigenvalues. and They represent the first The mean and variance of each field within the training statistics window. To prevent smoothing terms with a denominator of zero, the upstream input of this formula is the original multi-source field, and the downstream is used for the multi-source joint coding function. The numerical input portion. For categorical features, hash features, and text features, the corresponding ideas in this formula are reflected through vocabulary mapping, embedding lookup tables, and text encoders, which do not conflict with the numerical standardization formula.

[0096] When multiple events have time skew, it is necessary to establish time alignment weights so that data closer to the current event contributes more to risk characterization. This is especially true in carrier risk control scenarios, where switching multiple numbers on the same device within a short period or repeatedly performing sensitive operations on the same account within a short period carries higher risk significance.

[0097] Formula (2) is used to express how the time alignment weight is calculated:

[0098]

[0099] in, Representation and event Related history Time weighting and These represent the timestamps of the current event and the historical records, respectively. This represents the time decay factor. The upstream input of this formula is a timestamp and the time decay factor, and the downstream is used for weighted aggregation of behavioral sequences. And construct the edge weights of the risk relationship graph. Through this formula, short-term, densely occurring abnormal behaviors will receive higher weights during aggregation, while weakly correlated behaviors with long time intervals will be naturally attenuated.

[0100] In this embodiment, step S2 involves generating multi-level risk representation vectors at the entity, behavior, and event levels based on the preprocessed multi-source heterogeneous risk control data. Specifically, this includes:

[0101] Joint encoding is performed on the entity's identity features, statistical features, and context features. The concatenated identity features, statistical features, and context features are then input into the entity-level encoding network and mapped to a unified vector space to generate an entity-level risk representation vector.

[0102] The behavioral sequences corresponding to risk events are time-weighted and encoded. The basic encoding results of each record in the behavioral sequence are aggregated according to time alignment weights and then input into the behavioral-level temporal coding network to generate a behavioral-level risk representation vector.

[0103] The aggregated entity vector, behavioral risk representation, and context vector of all entities involved in the event are fused and encoded. The aggregated entity vector, behavioral risk representation, and context vector are concatenated and input into the event-level encoding network to generate the event-level risk representation vector.

[0104] Weighted fusion is performed on the historical behavior vector and historical feedback statistical vector obtained from the encoding of event-level risk representation, historical interaction and disposal records, to generate a multimodal fusion risk representation vector.

[0105] See Figure 3Specifically, multi-source joint encoding first performs representation learning on the entity dimension. For different entities such as users, numbers, devices, orders, sessions, and contacts, their original features come from different sources and need to be mapped to the same vector space through a unified encoding function.

[0106] Formula (3) is used to represent the generation of entity-level risk representation:

[0107]

[0108] in, Representing entities Identity-type feature vectors Represents statistical feature vectors. Represents the context class feature vector. Represents entity-level coding networks. Representing entities The entity-level risk representation vector. The upstream input of this formula comes from standardized structured features and discrete embeddings, while the downstream is used for event-level aggregation and initial node representation of the risk relationship graph. In engineering implementation, Multilayer perceptrons, gated fusion networks, or deep coding networks with cross terms can be used.

[0109] Behavioral-level risk characterization is used to capture sequential patterns within an event time window, such as multiple authentication failures within a short period, rapid login switching across devices, and batch triggering of business transactions. Static entity vectors alone are insufficient to represent these temporal behaviors; behavioral sequences are required. Perform sequence modeling.

[0110] Formula (4) is used to represent the generation of behavioral-level risk representations:

[0111]

[0112] in, Indicates the recording of behavior The basic encoding result, The time weights obtained from formula (2) This represents a behavioral-level temporal coding network. Indicates an event The behavioral-level risk representation vector is calculated. The upstream input to this formula is the behavioral sequence and temporal weights, while the downstream is used for event-level fusion and single-point anomaly identification. This formula shows that behavioral segments that are closer in time and have more concentrated risk will have a more significant impact on the behavioral-level representation.

[0113] Event-level risk representation is used to integrate entities, behaviors, and context. Since operator risk control typically revolves around "events" rather than "single entities," a unified representation for event-oriented decision-making is needed.

[0114] Formula (5) is used to represent the generation of event-level risk representations:

[0115]

[0116] in, Indicates an event Aggregated entity vectors involving entities, This represents a behavioral-level risk characterization. Indicates from the context set Extracted context vector, This indicates an event-level coding network. This represents an event-level risk representation. The upstream inputs to this formula are entity-level aggregation results, behavioral-level vectors, and context vectors, while the downstream inputs are used for multimodal fusion and risk confidence estimation.

[0117] To accommodate text interaction features, historical processing features, and feedback statistical features, this invention further employs a multimodal fusion strategy. Multimodal fusion prevents the model from becoming biased due to relying solely on one type of signal, such as ignoring appeal history based only on communication statistical features, or ignoring device sharing relationships based only on account behavior.

[0118] Formula (6) is used to represent the construction of the multimodal fusion vector:

[0119]

[0120] in, This represents a historical behavior vector encoded from historical interaction records and historical action records. Represents the historical feedback statistical vector. The weights are non-negative fusion weights and satisfy the following conditions: , Indicates an event The multimodal fusion vector is used. The upstream inputs to this formula are event-level risk representations, historical vectors, and feedback vectors, while the downstream input is directly to the intelligent agent decision-making module. Through this formula, the model can dynamically balance current risk evidence with historical handling experience according to the scenario.

[0121] In this embodiment, step S2, which involves constructing a risk relationship diagram based on the relationships between entities, specifically includes:

[0122] The edge weights between entity nodes in the risk relationship graph are calculated by combining the basic association strength of the corresponding relationship types between nodes with the time decay factor. The closer the occurrence time of the association behavior between nodes, the higher the weight of the corresponding edge.

[0123] The node representation is updated by performing multi-level graph propagation. In each layer of propagation, the features of all neighboring nodes of the current node are aggregated, and after being combined with the node's own features and processed by the activation function, the hidden representation of the node in the current layer is generated.

[0124] Extract local risk subgraphs corresponding to risk events from the global risk relationship graph, and retain nodes whose graph distance from the core node of the event does not exceed the preset maximum number of hops and whose overall node connection strength is not lower than the preset screening threshold to form local risk subgraphs that are strongly associated with the current event.

[0125] The construction of the risk relationship graph also includes constructing a multi-relationship adjacency matrix, summing the edge weights of each pair of nodes under all relationship types to obtain the comprehensive adjacency weight between nodes, which serves as the input for graph propagation.

[0126] Specifically, during the risk relationship graph construction phase, the system needs to define the edge weights between different types of entities. Edge weights reflect not only the existence of a connection, but also the strength of the connection, the degree of sharing, and its recentity. For example, the edge weight between multiple numbers on a shared device is typically higher than the edge weight between occasionally associated contacts.

[0127] Formula (7) is used to represent the calculation of edge weights in the risk relationship graph:

[0128]

[0129] in, Represents a node With nodes In relation types The underlying correlation strength, and The timestamp representing the most recent associated behavior between two nodes. Indicates the time decay factor. This represents the edge weight under this relationship type. The upstream input of this formula is the relationship extraction result and timestamp, and the downstream is used to generate the adjacency matrix. It also drives graph propagation. This formula allows the system to uniformly weight relationships such as shared devices, shared contacts, shared order links, and shared session transfers.

[0130] Based on edge weights, an adjacency matrix with relational types can be constructed as the core input of the graph propagation model.

[0131] Formula (8) is used to represent the construction of the multi-relation adjacency matrix:

[0132]

[0133] in, Represents a node With nodes Adjacency weights in the comprehensive risk relationship diagram Represents a set of edge types. This represents the edge weights under a specific relation type. The upstream input of this formula is the edge weights calculated by formula (7), and the downstream input is used for the graph propagation function. In practical implementations, the adjacency matrix can also be retained, and relation-aware aggregation can be performed during propagation. This formula corresponds to a comprehensive expression.

[0134] The graph propagation phase requires the aggregation of risk information within a local neighborhood to perceive group-based resource sharing and chain-like transmission paths. Single-layer aggregation is often insufficient to reveal multi-hop collaborative risks; therefore, this invention employs multi-layer propagation.

[0135] Formula (9) is used to represent the graph propagation update rule:

[0136]

[0137] in, Represents a node In the Hidden representation of layers, Represents a node The neighborhood group, Indicates the first Layer neighbor message transformation matrix, Indicates the first Layer self-loop transformation matrix, This represents the activation function. The upstream input of this formula is the adjacency matrix. and initial node representation Downstream, it is used to construct local risk subgraph representations and collaborative anomaly scoring. Through multi-layer propagation, the system can identify shared device links, number relay processing links, and group synchronization behaviors.

[0138] To balance real-time performance and local interpretability, the system extracts event-related local risk subgraphs from the full graph. This avoids latency and noise issues caused by full-image propagation in large-scale scenarios.

[0139] Formula (10) is used to represent the extraction criteria for local risk subgraphs:

[0140]

[0141] in, Represents a node Relative to the event Graph distance of core nodes, This represents the maximum number of hops in the local subgraph. Represents a node The overall connection strength, Representation and event The relevant local screening thresholds are used. The upstream inputs to this formula are a global risk relationship graph and core event entities, while the downstream inputs are used for collaborative anomaly scoring and intelligent agent context awareness. Through this formula, the system retains high-risk neighborhoods closely related to the current event, improving the focus of subsequent decisions.

[0142] In step S2, the identification of single-point anomalies and multi-entity collaborative anomalies through graph propagation calculations specifically includes:

[0143] A single-point anomaly score is calculated based on a multimodal fusion risk representation vector, and the score network is input to output the degree of deviation of the event itself from the historical baseline.

[0144] The collaboration anomaly score is calculated based on the node representation of the local risk subgraph. After pooling and aggregating the final hidden representation of all nodes in the subgraph, it is input into the collaboration scoring network and the anomaly degree of multi-entity collaboration behavior is output.

[0145] A comprehensive anomaly score is obtained by fusing single-point anomaly scores and collaborative anomaly scores. The fusion weights of the two types of scores are all non-negative and the sum of the weights is 1.

[0146] Based on the comprehensive anomaly score, the completeness of the event input data, and the structural stability of the local risk subgraph, the risk confidence level of the current risk assessment is calculated, reflecting the credibility of the risk evidence.

[0147] Specifically, at the anomaly identification level, this invention distinguishes between single-point anomalies and collaborative anomalies. Single-point anomalies focus more on the deviation of the event itself from the baseline, while collaborative anomalies focus on the propagation of the event in the graph structure and its peer consistency.

[0148] Formula (11) is used to represent the single-point anomaly score:

[0149]

[0150] in, This represents a single-point anomaly scoring function. and These are the scoring parameters, Indicates an event The single-point anomaly scoring. The upstream input of this formula is a multimodal fusion vector. Downstream and collaborative anomaly scoring together constitute a comprehensive risk assessment. Single-point anomaly scoring is suitable for identifying patterns such as a large number of failed verification codes in a short period of time, abnormal logins across different locations, or frequent changes in service plans.

[0151] Formula (12) is used to represent the collaborative anomaly score:

[0152]

[0153] in, This represents the collaborative anomaly scoring function. This represents a function that performs pooling aggregation on the node representations of the subgraph. Indicates the first The node representation after layer propagation, and These are the collaborative scoring parameters, This represents the collaborative anomaly score. The upstream input of this formula is a local risk subgraph. The results of its graph propagation are used downstream for comprehensive risk scoring and risk confidence estimation. Collaborative anomaly scoring helps identify collaborative risks such as group-based shared device processing, chain transfers, and bulk number manipulation.

[0154] In practical systems, it is necessary to fuse single-point anomalies with collaborative anomalies to obtain the final risk score and risk confidence estimate.

[0155] Formula (13) is used to represent the comprehensive risk score:

[0156]

[0157] in, and To integrate weights, satisfy , Indicates an event The comprehensive anomaly score is calculated. The upstream inputs of this formula are single-point anomaly scores and collaborative anomaly scores, while the downstream inputs are used for risk confidence calculation and the risk input of the intelligent agent decision-making module.

[0158] To avoid taking overly aggressive actions when evidence is insufficient, this invention further estimates the risk confidence level. And its confidence interval. Risk confidence depends on both the score values ​​and the completeness of the inputs, historical consistency, and graph structure stability.

[0159] Formula (14) is used to express the risk confidence level estimate:

[0160]

[0161] in, Indicates an event Input integrity metrics reflect whether key data sources are missing; This represents a structural stability index, reflecting the consistency of related evidence in local risk subplots; These are weight parameters; This represents the risk confidence level. The upstream inputs to this formula are the comprehensive anomaly score, input integrity index, and structural stability index; the downstream uses are for candidate action sequence generation and constraint correction. If... lower or If the fluctuations are large, then even If the value is high, the system will also respond to a lower value. Suppress overexcitation.

[0162] Regarding exception handling, this module supports the following degradation mechanism. Firstly, if If missing, retain. , and The generated representation, while simultaneously lowering the reliability of the collaborative anomaly score; secondly, if historical handling records... Delay, then the history vector Third, if a local risk subgraph is too large and may affect real-time performance, it can be reduced. or improve Implement subgraph clipping to meet Require.

[0163] Regarding the data interface with downstream modules, this module outputs... , , and This serves as the core input for the intelligent agent action generation and collaborative decision-making module. Used to represent risk semantics Used to represent structural context, Indicates the intensity of risk. This indicates the credibility of the evidence. The interface design ensures that downstream decision-makers not only see the "level of risk," but also the "sources and structural relationships of risk."

[0164] In this embodiment, in step S3, during the process of generating several candidate action sequences with action dependencies in a preset action space, the intelligent agent decision-making module constructs the decision state of the intelligent agent, including:

[0165] The multimodal fusion risk representation vector, the pooled representation of the local risk subgraph, the historical strategy effect encoding results, the business constraint encoding results, and the risk confidence are concatenated to form a unified decision state vector;

[0166] The policy probability of a single action is calculated based on the decision state vector. After the input policy network is normalized, the selection probability of each possible action is output.

[0167] The executableness of single-action strategy probability execution under regulatory and business constraints is masked. Based on the set of regulatory rules and the set of business constraints, the probability of actions that do not meet the constraints is suppressed, and compliant actions are retained to participate in the generation of subsequent candidate sequences.

[0168] Perform executability constraint projection on the candidate action sequence, and modify the original candidate action sequence according to regulatory rules, business constraints, manual review capacity limit and resource budget limit, eliminate unexecutable action combinations, and output the valid action sequence that meets the execution conditions;

[0169] After an action is executed, a decision state transition is performed. Based on the current decision state, the actions already executed, the intermediate results of the action execution, or observations, the decision state for the next step is updated to achieve sequential dynamic decision-making.

[0170] In step S3, generating a candidate action sequence with action dependencies specifically includes:

[0171] The sequenced candidate action sequence is generated based on the single-action strategy probability. The conditional selection probability of each action is multiplied according to the execution order of the actions to obtain the joint generation probability of the entire candidate action sequence.

[0172] The overall utility of each candidate action sequence is calculated. The overall value score of the sequence is obtained by combining the estimated risk reduction benefit, total cost of action execution, user impact range coefficient, and business importance coefficient with preset weights.

[0173] The estimated risk reduction benefit of the candidate action sequence is calculated by subtracting the expected posterior risk score after executing the action sequence from the comprehensive risk score before the action is executed. The expected posterior risk score is obtained based on the statistical analysis of the execution effects of similar strategies in the past.

[0174] The total execution cost of the candidate action sequence is obtained by summing the system execution cost, manual processing cost, and customer service cost of each action in the sequence.

[0175] The user influence range coefficient of the candidate action sequence is the sum of the basic influence intensity of each action in the sequence and the corresponding customer service level weight. The higher the customer service level, the greater the influence coefficient of the same action.

[0176] See Figure 4 Specifically, the goal of step S3 is to [do something] in the action space. Generate multiple candidate action sequences To achieve this goal, this module maps risk semantics, graph structure context, historical strategy effects, and business constraints to a unified decision state.

[0177] Formula (15) is used to represent the construction of the decision state of the intelligent agent:

[0178]

[0179] in, Indicates an event The decision state vector, This represents a pooling representation of a local risk subgraph. This represents the encoded results of the historical strategy effect. This indicates the result of the business constraint coding. This represents the risk confidence level. The upstream input of this formula is upstream risk characterization and business constraint information, while the downstream is used for action probability calculation and candidate action sequence search. Through the unified construction of decision states, intelligent agents can simultaneously utilize risk characteristics, structural relationships, and business constraints.

[0180] Given a decision state Then, the system first evaluates the selection probability of a single action to form the basic distribution for action generation. This distribution is not the final policy, but rather a priori for candidate sequence search.

[0181] Formula (16) is used to represent the probability of a single-action policy:

[0182]

[0183] in, Indicates in the event Selecting an action in context The strategy probability, and For parameters used to calculate the probability of an action, the subscript is... Indicates the first One action. The upstream input of this formula is the decision state. Downstream is used to construct candidate action sequences. If an action is prohibited by current regulatory rules, it must be approved before entering Softmax. Suppress its logarithmic score to an extremely low value.

[0184] To express the dependencies between actions, this invention employs sequential decision-making. That is, the generation of a subsequent action depends on whether the previous action was executed successfully, whether a termination condition was triggered, and the state transition following the previous action.

[0185] Formula (17) is used to represent the joint probability of the candidate action sequence:

[0186]

[0187] in, Indicates the first A candidate action sequence, This indicates the length of the sequence. Indicates the first The state before step selection, Indicates the first The conditional probability of each action. The upstream input of this formula is the conditional probability of a single action, and the downstream is used for searching and ranking candidate action sequences. This design ensures that "verify first, then restrict" and "restrict first, then verify" will be distinguished as different sequences.

[0188] Action execution triggers state transitions. For example, if a verification interaction succeeds, subsequent actions may be allowed immediately; if it fails, the priority of the restriction action may be increased; if there is no response after a timeout, it may be transferred to observation or manual review. Therefore, it is necessary to explicitly define the action state transition function.

[0189] Formula (18) is used to represent the state transition after the action is performed:

[0190]

[0191] in, Represents the state transition function. Indicates the first The state before stepping, Indicates the first The action to be performed step by step This indicates the result or intermediate observation of the action execution, such as verification passed, verification failed, user unresponsive, or resource exhaustion. This represents the updated state. The upstream input of this formula is the current state, action, and action result, while the downstream is used for the next action decision. Through this formula, the action sequence generation mechanism can explicitly express dynamic collaborative decision-making, rather than a one-time static output.

[0192] In operator risk control, action selection must consider both benefits and costs. Choosing only the action with the highest probability may lead to overly conservative approaches in high-risk scenarios, or overly aggressive approaches when evidence is insufficient. Therefore, this invention defines the comprehensive utility of candidate action sequences.

[0193] Formula (19) is used to represent the overall utility of candidate action sequences:

[0194]

[0195] in, This indicates the projected risk reduction benefit of the candidate action sequence. Indicates the cost of the action. This represents the coefficient indicating the scope of user influence. Indicates the importance coefficient of the business. For weight parameters, This represents the overall utility. The upstream input of this formula consists of estimates of the benefits, costs, and impacts corresponding to each action sequence, while the downstream is used for ranking candidate actions and correcting downstream constraints. In engineering, It can be provided by statistics of similar historical events, offline counterfactual estimation, or reinforced value networks.

[0196] To characterize the risk mitigation effect of action sequences, this invention explicitly models the risk reduction benefit. This benefit depends not only on the current risk score but also on the average effect of similar historical actions.

[0197] Formula (20) is used to express the risk reduction return estimate:

[0198]

[0199] in, This indicates the comprehensive risk score before implementation. Indicates the execution of candidate action sequences Expected posterior risk score of subsequent events Expressing conditional expectation, This represents the set of historical strategy effects. The upstream input to this formula is the current risk score and historical strategy effects, while the downstream is used for comprehensive utility calculation. If an action, although historically costly, significantly reduces posterior risk, then... It will increase accordingly.

[0200] Action cost This includes not only direct system resource costs, but also manual review costs, customer service costs, and potential complaint handling costs. In carrier scenarios, manual review is particularly limited by capacity.

[0201] Formula (21) is used to express the estimation of motion costs:

[0202]

[0203] in, Indicates action The system execution cost, This indicates the cost of manual processing. This indicates customer service costs. This represents the total cost of the action sequence. The upstream inputs to this formula are action cost configuration and real-time resource status, while the downstream is used for comprehensive utility and resource constraint verification. If the manual review pool is close to... Then, the relevant actions are manually reviewed. Increase the dynamic.

[0204] User Influence Scope Coefficient This is used to reflect the intensity of an action's impact on user perception and the potential scope of its functionalities. For example, restricting all core business capabilities has a significantly greater impact than delaying a single high-risk operation; the experience impact of restricting high-value customers is also greater than that of ordinary customers.

[0205] Formula (22) is used to represent the user's influence range coefficient:

[0206]

[0207] in, Indicates action The basic influence intensity, Indicates the weight of customer service level. This represents the impact range coefficient of the action sequence on the user. The upstream inputs to this formula are the action template and customer level configuration, while the downstream is used for comprehensive utility and constraint correction. For high-value customers, A larger value usually means that the system will be more cautious in using high-intervention actions.

[0208] After candidate actions are generated, they need to be masked and modified according to regulatory rules and business constraints. Some actions are not executable under specific business conditions, some actions must be verified before being restricted, and some actions need to be downgraded and replaced for certain types of customers.

[0209] Formula (23) is used to represent the masking of action executability:

[0210]

[0211] in, This represents the probability of an action after constraint masking. Indicates according to the set of regulatory rules and business constraint set This function masks actions. The upstream input to the formula is the original action probabilities and constraint set, and the downstream is used for candidate action sequence search. If a regulatory requirement stipulates that key businesses cannot be directly approved without sufficient evidence, then the "direct approval" action will be masked or significantly downweighted.

[0212] To generate candidate sequences that will eventually be sent to the downstream orchestration module, this invention performs constrained projection on the action sequences to ensure that they belong to the executable space.

[0213] Formula (24) is used to represent the constrained projection of the candidate action sequence:

[0214]

[0215] in, This represents the sequence of candidate actions that can be executed after projection. Represents the constraint projection function. This indicates the maximum capacity for manual review. This represents the upper limit of the resource budget. The upstream input of this formula is the original candidate action sequence and constraint set, and the downstream is used by the action orchestration and execution module. If a candidate sequence contains actions that exceed the capacity for manual review, the projection function will replace it with a low-resource version that "verifies the interaction first, then observes," or trigger a delayed queuing mechanism if necessary.

[0216] Regarding action triggering conditions, this module generally adopts the following strategy: when High and When the time is high, prioritize generating verification, restriction, and delay actions; when High but When the time is moderate, prioritize generating verification and manual review actions; when When approaching the threshold and with a high customer service level, prioritize generating observation and lightweight validation actions; when < If there are no obvious coordination anomalies, a direct release or minimal observation action can be generated. The termination conditions for the action include successful verification, the risk post-hoc score falling below the threshold, a clear conclusion from manual review, exceeding the maximum allowable processing time, or the business window being closed.

[0217] Regarding exception handling, this module supports the following mechanisms: if an external execution interface that an action depends on becomes unavailable, it is immediately downgraded to an alternative action of the same level; if the action sequence search is not completed within the latency budget, the currently generated sequence is returned. A high-efficiency sequence; if the policy probability distribution is too flat, it indicates that the model uncertainty is high, so the priority of manual review or observation should be increased to control the risk of misjudgment.

[0218] Regarding the interface with upstream and downstream components, this module receives risk representations and graph context from upstream and outputs structured candidate action sequences from downstream. Each sequence includes at least an action list, execution order, action parameters, estimated utility, constraint description, termination conditions, and a summary of key evidence. This interface design allows the action orchestration and execution module to perform secondary sorting and constraint correction without having to reinterpret the original risk characteristics.

[0219] In this embodiment, in step S4, the candidate action sequences are ranked by cost and benefit. Based on the comprehensive utility of the candidate action sequences, penalties for violating regulatory preferences, penalties for affecting business continuity, and penalties for affecting customer service experience are subtracted to obtain the final arrangement and ranking score.

[0220] The admission criteria for directly releasing candidate action sequences are as follows: only when the overall anomaly score of an event is lower than the preset release risk threshold and the risk confidence level is lower than the preset high confidence risk threshold, the candidate action sequence that is directly released can enter the final ranking; otherwise, it will be blocked or downweighted.

[0221] The capacity constraint rule for candidate action sequences that include manual review is: a candidate action sequence that includes manual review can only enter the final sorting if the sum of the number of tasks in the current manual review queue and the new manual review task requirements of the candidate action sequence does not exceed the upper limit of manual review capacity.

[0222] The execution delay constraint rule for candidate action sequences is: a candidate action sequence can only enter the final sorting if the estimated execution delay does not exceed the system's preset real-time decision delay threshold.

[0223] Step S4 further includes performing hierarchical constraint correction for different business scenarios. For different business scenarios, different scenario mapping weights are configured for regulatory constraints, business continuity constraints, and customer service level constraints, and the total constraint correction intensity under the current scenario is calculated. Specifically, the weight of customer service level constraints is increased for high-value customer scenarios, the weight of regulatory constraints is increased for key regulatory business scenarios, and the default weight configuration with the best cost-benefit is adopted for ordinary business scenarios.

[0224] See Figure 5 Specifically, the sorting not only takes into account the combined utility from the upstream modules. Furthermore, regulatory amendments, business continuity amendments, and service level amendments are added to form the final ranking score under a truly executable environment.

[0225] Formula (25) is used to represent the ranking score of candidate action sequences:

[0226]

[0227] in, This represents the ranking score of the candidate action sequence. This indicates penalties for violating regulatory preferences. This indicates penalties that are detrimental to business continuity. This indicates penalties that negatively impact the customer service experience. , and These represent the corresponding constraint weights. The upstream input of this formula consists of the overall utility and various penalty terms, while the downstream input is used for candidate sequence ranking and final policy selection.

[0228] After sorting, this module needs to determine whether the candidate action sequence meets different execution thresholds, such as release, observation, restriction, or manual review. In particular, for actions that are released directly, additional risk threshold conditions need to be met.

[0229] Formula (26) is used to express the conditions for direct release:

[0230]

[0231] in, Indicates an event An indicator function to determine whether the conditions for direct release are met. This indicates the risk threshold for allowing passage. This represents the high-confidence risk threshold. The upstream inputs to this formula are the risk score and risk confidence level, while the downstream is used for constraint correction. If an event score is low and there is no high-confidence risk, the candidate sequence is allowed to enter the final ranking; otherwise, even if its overall utility is high, it will be blocked or downweighted.

[0232] For candidate action sequences that require manual review, the system must check the upper limit of manual review capacity. If the capacity is close to the limit, an alternative strategy or delayed processing should be selected to prevent the manual process from blocking and causing a system collapse.

[0233] Formula (27) is used to represent the manual review capacity constraint:

[0234]

[0235] in, This indicates the number of tasks currently in the manual review queue. Represents candidate action sequences The newly added requirement for manual review tasks, This indicates the maximum capacity for manual review. This indicates whether the sequence meets the manual review capacity constraint. The upstream input of this formula is the real-time queue status and candidate action requirements, and the downstream is used to perform filtering and constraint projection.

[0236] Regarding the tiered strategy execution rules, this invention sets different correction strategies for different business scenarios. For high-value customers, the system prefers "verification + observation" or "verification + manual review" rather than direct restriction; for key monitored businesses, the system prefers "restricting key high-risk operations first, then verifying"; for ordinary businesses, the system can emphasize optimal cost-benefit.

[0237] Formula (28) is used to represent the hierarchical constraint correction coefficient:

[0238]

[0239] in, Indicates an event The total constraint correction strength in the current business scenario. This represents the scenario mapping weight, used to express the degree of emphasis placed on the three types of constraints—regulation, continuity, and service level—by different scenarios. The upstream input to this formula is the scenario category and basic constraint coefficients, while the downstream is used to update the ranking score and action replacement strategy. For key regulatory businesses, A higher value can be chosen; for high-value customers, A higher value can be taken.

[0240] After the candidate action sequences are sorted and constrained, the system selects the final strategy and issues it for execution. The action status during execution generally includes pending execution, executing, successful execution, failed execution, timeout, and rollback completion. If a preceding action fails due to an external interface exception, the system can call an equivalent alternative action; if the failure is due to regulatory verification failure, the sequence is terminated and switched to a backup sequence.

[0241] In this embodiment, step S5 involves constructing a monitoring signal and reinforcement reward based on the feedback data, and jointly updating the risk representation model and the intelligent agent decision-making strategy, specifically including:

[0242] The combined reward value is calculated based on the execution feedback data, taking into account the risk prevention effect, misjudgment, changes in business losses, and additional delays introduced by the handling, and is calculated by weighting according to preset weights.

[0243] A loss function for the risk characterization model is constructed, which combines risk category supervision loss and anomaly scoring calibration loss to optimize the accuracy of risk identification.

[0244] A loss function for intelligent agent strategy is constructed, which combines execution reward and the generation probability of action sequence to improve the selection probability of high-yield action sequence;

[0245] A joint loss function is constructed, which sums the risk representation loss, policy loss, and model regularization term according to preset weights to achieve joint optimization of the risk representation model and the intelligent agent strategy.

[0246] The joint update adopts a controlled incremental update rule. When the confidence interval width of the feedback estimate does not exceed a preset threshold, the model parameters are updated based on the gradient of the joint loss at a preset step size to filter out the interference of low-quality and unstable feedback on the model.

[0247] See Figure 6 Specifically, after an action is executed, feedback needs to be collected and rewards generated. These rewards not only reflect whether risks have been mitigated, but also reflect changes in the cost of misjudgment and business losses. To integrate "risk control effectiveness" and "business impact" into the learning objectives, this invention employs a combined reward system.

[0248] Formula (29) is used to represent the feedback reward function:

[0249]

[0250] in, An indicator function that shows whether the risk has been effectively contained. An indicator function that indicates whether a false positive has occurred. Indicates the change in business losses. This indicates the additional delay introduced by the processing. As a reward weight, Indicates an event The final reward value. The upstream inputs to this formula are the execution result, misjudgment status, changes in business losses, and latency metrics; the downstream inputs are used for the strategy loss function and joint updates. If the risk is effectively blocked and there are no misjudgments, then... The impact could be significant; if accidental injury is severe or business losses increase, then... decline.

[0251] To update the risk characterization model with feedback data, supervisory labels need to be constructed from the feedback samples. For samples that are appealed and verified as misjudged, the system corrects their risk labels to the low-risk category; for samples that continue to pose a risk after intervention, the system marks them as "insufficiently addressed" samples and increases vigilance for subsequent patterns.

[0252] Formula (30) is used to represent the loss function of the risk characterization model:

[0253]

[0254] in, Indicates an event Supervisory labels, This indicates that the risk characterization model represents the event. The predicted output, This represents the target risk score after feedback correction. For the weights of the scoring regression terms, This represents the loss function of the risk characterization model. The upstream inputs to this formula are the feedback-constructed labels and the model predictions, while the downstream inputs are used for joint loss calculation. The first term is used for risk category supervision, and the second term is used to calibrate the anomaly scores to the feedback-driven target value.

[0255] For intelligent agent strategies, this invention employs a reward-based policy optimization objective to improve the quality of long-term action selection. This objective can be implemented as a reinforcement learning-based policy gradient or as an offline weighted behavior cloning. This invention is described below in the form of a policy gradient.

[0256] Formula (31) is used to represent the policy loss function:

[0257]

[0258] in, Indicates an event The final sequence of actions that is executed and evaluated. This represents the probability that the intelligent agent assigns to this action sequence. This represents the reward value calculated using formula (29). This represents the policy loss function. The upstream input to this formula is the execution sequence probability and the feedback reward, and the downstream is used for joint updates. Action sequences with high rewards will receive a greater probability boost during training, while action sequences with low or negative rewards will be suppressed.

[0259] To simultaneously optimize risk representation and action strategy, this invention employs a joint loss function to coordinate and update the two parts within a unified training framework.

[0260] Formula (32) is used to represent the joint loss function:

[0261]

[0262] in, These are the weighting coefficients. Represents the set of parameters for the joint model. Represents a regular term. This represents the joint loss function. The upstream inputs to this formula are the risk representation loss and the policy loss, while the downstream is used for gradient updates. Through this joint objective, the system improves its recognition capabilities while also considering the effectiveness of action decision-making, avoiding local optima caused by the split optimization of the two parts.

[0263] After obtaining the joint loss, it needs to be updated online or incrementally. Considering the real-time requirements and model stability requirements of operator risk control systems, this invention adopts controlled step-size updates and combines confidence intervals to filter low-quality feedback.

[0264] Formula (33) is used to represent the incremental update rule:

[0265]

[0266] in, and These represent the parameters before and after the update, respectively. Indicates the update step size. This represents the gradient of the joint loss with respect to the parameters. Indicates an event The width of the confidence interval for the feedback estimate. This indicates the threshold of the confidence interval that allows updates. This represents the indicator function. The upstream input to this formula is the joint loss and feedback confidence assessment, while the downstream is used for model parameter updates. If the feedback evidence is unstable, such as when the appeal outcome has not yet been finalized, the system will not perform a strong update on that sample.

[0267] One possible embodiment also includes a full-link anomaly degradation handling mechanism:

[0268] When some data sources are missing in multi-source heterogeneous risk control data, a missing identifier vector and confidence reduction mechanism are used to perform downgraded inference; when the estimated execution delay of a candidate action sequence exceeds the preset real-time decision threshold, the length of the action sequence is pruned or switched to a low-resource alternative action; when the manual review capacity reaches the preset limit, the cost weight of the manual review action is dynamically increased, and candidate action sequences with low resource consumption are output first.

[0269] Specifically, to meet real-time requirements, this module also verifies execution latency. If a candidate action sequence requires multiple levels of external system calls, the total processing latency may exceed [a certain threshold]. In this case, it is necessary to downgrade it in advance during the scheduling stage.

[0270] Formula (34) is used to represent the execution delay constraint:

[0271]

[0272] in, Represents candidate action sequences The estimated execution delay, This indicates the threshold for the system's allowed real-time decision-making delay. This indicates whether the sequence meets the latency constraint. The upstream input of this formula is the action execution path estimate and the system latency budget, and the downstream is used for final execution filtering. Through this constraint, the system avoids slowing down the main process due to excessively long action chains in high-concurrency environments.

[0273] In summary, this invention completes a closed loop from "candidate action sequences" to "executable risk control strategies and then to model updates" through sorting, constraint correction, execution, and feedback learning. Sorting addresses the balance between benefits and costs, constraint correction addresses compliance and service requirements in real-world business environments, and feedback learning addresses long-term adaptive optimization. At the system implementation level, this invention can be deployed via a service-oriented architecture. The multi-source risk characterization construction module can be deployed as a feature and graph inference service; the intelligent agent action generation and collaborative decision-making module can be deployed as a low-latency strategy generation service; and the action orchestration execution and feedback learning module can be deployed as an execution learning service with idempotent control and state recursion capabilities. These services can interact via message bus, feature cache, and online model service interfaces. For electronic device implementation, the processor executes instructions from memory to complete the above steps and communicates with external business systems, customer service systems, SMS gateways, manual review systems, and log systems through network interfaces.

[0274] In a specific carrier scenario, if a number logs in across devices within a short period, initiates high-risk SIM card replacements, and is associated with multiple recently penalized contact nodes, and similar combined actions have historically demonstrated high fraudulent profits, then the upstream module will generate a high [percentage missing]. and Meanwhile, a clear collaborative structure was observed in the local risk subgraph. The intelligent agent module may generate multiple candidate action sequences, such as "SMS verification + restriction on card replacement," "APP identity verification + delayed execution + manual review," and "direct restriction of key business capabilities + manual review." After considering customer level, regulatory rules, and manual review capacity, the action orchestration module may choose the sequence of "initiating APP identity verification first, and delaying execution and transferring to manual review if it fails," to balance risk control benefits and customer experience. After execution, if the user passes the identity verification normally and no further risks occur, the feedback learning module will reduce the strong restriction tendency of this mode in similar contexts; conversely, if fraud still occurs later, the system will increase the risk weight of the relevant collaborative mode and the value assessment of the restriction action.

[0275] Furthermore, in high-value customer scenarios, the tiered strategy execution rules can be specified as follows: High but When the high confidence level is not reached, direct long-term restrictions on core communication capabilities should be prohibited. Instead, a combination of high-confidence verification and manual review should be prioritized. In key regulatory business scenarios, it can be stipulated that once a specific group-style collaborative structure threshold is met, sensitive business processing links can be frozen first, followed by subsequent verification. This differentiated execution rule does not require changes to the upstream risk characterization module; it can be implemented solely through constraint modifications during the action orchestration phase, thereby maintaining system architecture decoupling and real-time performance.

[0276] Example 2

[0277] See Figure 7 Embodiment 2 of the present invention also provides a deep learning agent collaborative decision-making system for operator risk control, used to execute the deep learning agent collaborative decision-making method for operator risk control described in the above embodiments, including:

[0278] The multi-source data access and preprocessing unit 100 is used to acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event.

[0279] The multi-level risk characterization construction unit 200 is used to generate multi-level risk characterization vectors at the entity level, behavior level, and event level based on the preprocessed multi-source heterogeneous risk control data, construct a risk relationship graph based on the relationship between entities, identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation, and output multi-modal fusion risk characterization and local risk subgraphs.

[0280] The intelligent agent action generation and collaborative decision-making unit 300 is used to input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module, and generate a number of candidate action sequences with action dependencies in the preset action space.

[0281] The action orchestration and execution unit 400 is used to perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution;

[0282] The feedback learning and joint update unit 500 is used to collect feedback data after the risk control strategy is executed, construct supervision signals and reinforcement rewards based on the feedback data, and jointly update the risk representation model and intelligent agent decision-making strategy.

[0283] It should be noted that the information interaction and execution process between the various units of the above system are based on the same concept as the method embodiment in Embodiment 1 of this application, and the resulting technical effects are the same as those in the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in this application, and it will not be repeated here.

[0284] Example 3

[0285] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium storing program code for a deep learning agent collaborative decision-making method for operator risk control. The program code includes instructions for executing the deep learning agent collaborative decision-making method for operator risk control as described in Embodiment 1 or any possible implementation thereof.

[0286] Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0287] Example 4

[0288] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0289] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor can execute the deep learning agent collaborative decision-making method for operator risk control according to Embodiment 1 or any possible implementation thereof by calling the program instructions.

[0290] Specifically, a processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.

[0291] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0292] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0293] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A deep learning agent collaborative decision-making method for operator risk control, applied to an operator risk control platform, characterized in that, Includes the following steps: S1. Acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event; S2. Based on the preprocessed multi-source heterogeneous risk control data, generate multi-level risk representation vectors at the entity level, behavior level, and event level. Construct a risk relationship graph based on the relationship between entities. Identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation. Output multi-modal fusion risk representation and local risk subgraph. S3. Input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module to generate several candidate action sequences with action dependencies in the preset action space. S4. Perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution; S5. Collect feedback data after the risk control strategy is executed, construct monitoring signals and reinforcement rewards based on the feedback data, and jointly update the risk characterization model and intelligent agent decision-making strategy.

2. The method of claim 1, wherein, In step S1, standardization preprocessing is performed on the multi-source heterogeneous risk control data, specifically including: The numerical features in the multi-source heterogeneous risk control data are subjected to unified standardization processing. Normalization calculation is performed based on the mean and variance of each field within the training statistical window to eliminate the differences in the units and ranges of values ​​of different fields. Time alignment weights are calculated for historical records associated with risk events, and an exponential decay mechanism is adopted, with higher weights for historical records that are closer to the current risk event time interval, thus weakening the impact of weakly correlated behaviors on the current risk assessment.

3. The method of claim 2, wherein, In step S2, based on the preprocessed multi-source heterogeneous risk control data, a multi-level risk representation vector is generated at the entity level, behavior level, and event level, specifically including: Joint encoding is performed on the entity's identity features, statistical features, and context features. The concatenated identity features, statistical features, and context features are then input into the entity-level encoding network and mapped to a unified vector space to generate an entity-level risk representation vector. The behavioral sequences corresponding to risk events are time-weighted and encoded. The basic encoding results of each record in the behavioral sequence are aggregated according to time alignment weights and then input into the behavioral-level temporal coding network to generate a behavioral-level risk representation vector. The aggregated entity vector, behavioral risk representation, and context vector of all entities involved in the event are fused and encoded. The aggregated entity vector, behavioral risk representation, and context vector are concatenated and input into the event-level encoding network to generate the event-level risk representation vector. Weighted fusion is performed on the historical behavior vector and historical feedback statistical vector obtained from the encoding of event-level risk representation, historical interaction and disposal records, to generate a multimodal fusion risk representation vector.

4. The method of claim 3, wherein, In step S2, a risk relationship diagram is constructed based on the relationships between entities, specifically including: The edge weights between entity nodes in the risk relationship graph are calculated by combining the basic association strength of the corresponding relationship types between nodes with the time decay factor. The closer the occurrence time of the association behavior between nodes, the higher the weight of the corresponding edge. The node representation is updated by performing multi-level graph propagation. In each layer of propagation, the features of all neighboring nodes of the current node are aggregated, and after being combined with the node's own features and processed by the activation function, the hidden representation of the node in the current layer is generated. Extract local risk subgraphs corresponding to risk events from the global risk relationship graph, and retain nodes whose graph distance from the core node of the event does not exceed the preset maximum number of hops and whose overall node connection strength is not lower than the preset screening threshold to form local risk subgraphs that are strongly associated with the current event. The construction of the risk relationship graph also includes constructing a multi-relationship adjacency matrix, summing the edge weights of each pair of nodes under all relationship types to obtain the comprehensive adjacency weight between nodes, which serves as the input for graph propagation; In step S2, single-point anomalies and multi-entity collaborative anomalies are identified through graph propagation calculations, specifically including: A single-point anomaly score is calculated based on a multimodal fusion risk representation vector, and the score network is input to output the degree of deviation of the event itself from the historical baseline. The collaboration anomaly score is calculated based on the node representation of the local risk subgraph. After pooling and aggregating the final hidden representation of all nodes in the subgraph, it is input into the collaboration scoring network and the anomaly degree of multi-entity collaboration behavior is output. A comprehensive anomaly score is obtained by fusing single-point anomaly scores and collaborative anomaly scores. The fusion weights of the two types of scores are all non-negative and the sum of the weights is 1. Based on the comprehensive anomaly score, the completeness of the event input data, and the structural stability of the local risk subgraph, the risk confidence level of the current risk assessment is calculated, reflecting the credibility of the risk evidence.

5. The method of claim 4, wherein, In step S3, during the process of generating several candidate action sequences with action dependencies in the preset action space, the intelligent agent decision-making module constructs the decision state of the intelligent agent, including: The multimodal fusion risk representation vector, the pooled representation of the local risk subgraph, the historical strategy effect encoding results, the business constraint encoding results, and the risk confidence are concatenated to form a unified decision state vector; The policy probability of a single action is calculated based on the decision state vector. After the input policy network is normalized, the selection probability of each optional action is output. The executableness of single-action strategy probability execution under regulatory and business constraints is masked. Based on the set of regulatory rules and the set of business constraints, the probability of actions that do not meet the constraints is suppressed, and compliant actions are retained to participate in the generation of subsequent candidate sequences. Perform executability constraint projection on the candidate action sequence, and modify the original candidate action sequence according to regulatory rules, business constraints, manual review capacity limit and resource budget limit, eliminate unexecutable action combinations, and output the valid action sequence that meets the execution conditions; After an action is executed, a decision state transition is performed. Based on the current decision state, the actions already executed, the intermediate results of the action execution, or observations, the decision state for the next step is updated to achieve sequential dynamic decision-making.

6. The method of claim 5, wherein the method further comprises: In step S3, candidate action sequences with action dependencies are generated, specifically including: The sequenced candidate action sequence is generated based on the single-action strategy probability. The conditional selection probability of each action is multiplied according to the execution order of the actions to obtain the joint generation probability of the entire candidate action sequence. The overall utility of each candidate action sequence is calculated. The overall value score of the sequence is obtained by combining the estimated risk reduction benefit, total cost of action execution, user impact range coefficient, and business importance coefficient with preset weights. The estimated risk reduction benefit of the candidate action sequence is calculated by subtracting the expected posterior risk score after executing the action sequence from the comprehensive risk score before the action is executed. The expected posterior risk score is obtained based on the statistical analysis of the execution effects of similar strategies in the past. The total execution cost of the candidate action sequence is obtained by summing the system execution cost, manual processing cost, and customer service cost of each action in the sequence. The user influence range coefficient of the candidate action sequence is the sum of the basic influence intensity of each action in the sequence and the corresponding customer service level weight. The higher the customer service level, the greater the influence coefficient of the same action.

7. The method of claim 1, wherein the method further comprises: In step S4, the candidate action sequences are ranked by cost and benefit. Based on the overall utility of the candidate action sequences, penalties for violating regulatory preferences, penalties for affecting business continuity, and penalties for affecting customer service experience are subtracted to obtain the final ranking score. The admission criteria for directly releasing candidate action sequences are as follows: only when the overall anomaly score of an event is lower than the preset release risk threshold and the risk confidence level is lower than the preset high confidence risk threshold, the candidate action sequence that is directly released can enter the final ranking; otherwise, it will be blocked or downweighted. The capacity constraint rule for candidate action sequences that include manual review is: a candidate action sequence that includes manual review can only enter the final sorting if the sum of the number of tasks in the current manual review queue and the new manual review task requirements of the candidate action sequence does not exceed the upper limit of manual review capacity. The execution delay constraint rule for candidate action sequences is: a candidate action sequence can only enter the final sorting if the estimated execution delay does not exceed the system's preset real-time decision delay threshold. 8.The method of claim 7, wherein, Step S4 also includes performing hierarchical constraint correction for different business scenarios. For different business scenarios, different scenario mapping weights are configured for regulatory constraints, business continuity constraints, and customer service level constraints, and the total constraint correction strength under the current scenario is calculated. In high-value customer scenarios, the weight of customer service level constraints is increased; in key regulatory business scenarios, the weight of regulatory constraints is increased; and in ordinary business scenarios, the default weight configuration with the best cost-benefit ratio is adopted. 9.The method of claim 8, wherein, In step S5, a supervision signal and reinforcement reward are constructed based on the feedback data, and a joint update is performed on the risk representation model and the intelligent agent decision-making strategy, specifically including: The combined reward value is calculated based on the execution feedback data, taking into account the risk prevention effect, misjudgment, changes in business losses, and additional delays introduced by the handling, and is calculated according to preset weights. A loss function for the risk characterization model is constructed, which combines risk category supervision loss and anomaly scoring calibration loss to optimize the accuracy of risk identification. A loss function for intelligent agent strategy is constructed, which combines execution reward and the generation probability of action sequence to improve the selection probability of high-yield action sequence; A joint loss function is constructed, which sums the risk representation loss, policy loss, and model regularization term according to preset weights to achieve joint optimization of the risk representation model and the intelligent agent strategy. The joint update adopts a controlled incremental update rule. When the confidence interval width of the feedback estimate does not exceed a preset threshold, the model parameters are updated based on the gradient of the joint loss with a preset step size to filter out the interference of low-quality and unstable feedback on the model. It also includes a full-link anomaly degradation handling mechanism: When some data sources are missing in multi-source heterogeneous risk control data, a missing identifier vector and confidence reduction mechanism are used to perform downgraded inference; when the estimated execution delay of a candidate action sequence exceeds the preset real-time decision threshold, the length of the action sequence is pruned or switched to a low-resource alternative action; when the manual review capacity reaches the preset limit, the cost weight of the manual review action is dynamically increased, and candidate action sequences with low resource consumption are output first.

10. A deep learning agent collaborative decision system for operator risk control, configured to perform the deep learning agent collaborative decision method for operator risk control according to any one of claims 1-9, characterized in that, include: The multi-source data access and preprocessing unit is used to acquire multi-source heterogeneous risk control data related to the risk event to be analyzed, perform standardized preprocessing and cross-entity association on the multi-source heterogeneous risk control data, and construct the entity association set and observation sample window corresponding to the risk event. The multi-level risk characterization construction unit is used to generate multi-level risk characterization vectors at the entity level, behavior level, and event level based on the preprocessed multi-source heterogeneous risk control data, construct a risk relationship graph based on the relationship between entities, identify single-point anomalies and multi-entity collaborative anomalies through graph propagation calculation, and output multi-modal fusion risk characterization and local risk subgraphs. The intelligent agent action generation and collaborative decision-making unit is used to input the multimodal fusion risk representation, the local risk subgraph, the risk confidence, the historical strategy effect and business constraints into the intelligent agent decision-making module, and generate a number of candidate action sequences with action dependencies in the preset action space. The action orchestration and execution unit is used to perform cost-benefit sorting, multi-dimensional constraint correction and executability projection on several candidate action sequences, filter to obtain the final executable risk control strategy and issue it for execution; The feedback learning and joint update unit is used to collect feedback data after the risk control strategy is executed, construct supervision signals and reinforcement rewards based on the feedback data, and jointly update the risk representation model and the intelligent agent decision-making strategy.