A method, device and related product for API risk identification

CN122818375APending Publication Date: 2026-09-25HUA XIA BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611155493.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有技术中,仅使用用户实体行为分析只能识别异地、高频等表层访问异常,无法识别支付绕过、金额篡改等业务逻辑漏洞,易产生漏报

Benefits of technology

[0019]基于上述一种API风险识别的方法,本申请还公开了一种计算机程序产品,所述计算机程序产品包括计算机程序,所述计算机程序被处理器运行时,用于实现上述任一项方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818375A_ABST
    Figure CN122818375A_ABST
Patent Text Reader

Abstract

The application discloses an API risk identification method and device and related products, and relates to the technical field of data processing. UEBA features and business sequence features are extracted from a target API call request to support two-dimensional detection of user behavior and business process, behavior anomaly detection is performed based on the UEBA features, and user behavior anomalies such as out-of-place login, batch crawler, low-frequency slow attack and gang call are accurately identified. Sequence violation detection is performed based on the business sequence features, and business logic vulnerabilities such as payment bypass and step omission are identified. The two-dimensional detection results are fused to obtain a comprehensive risk detection result, which makes up for the short board of pure process detection that cannot distinguish between normal and abnormal operations, solves the defect of missing business attack relying only on behavior analysis, reduces false positives and false negatives, and realizes accurate identification of complex attacks. An entity association graph is constructed based on the target API call request and the comprehensive risk detection result, so that the attack link can be quickly and accurately restored, and the attack source can be located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus and related products for API risk identification. Background Technology

[0002] Application Programming Interface (API) is the core channel for internal and external business interaction within an enterprise. Currently, mainstream API security detection is divided into two types of technologies: user entity behavior analysis and business call sequence detection. The former relies on users, devices, and historical behavior to identify abnormal access, while the latter relies on sequence models to detect the compliance of business call processes. These two technologies can be deployed separately for risk identification, or there is a simple overlay approach to integrated detection. Graph databases and adaptive machine learning models are also used for attack tracing and dynamic rule learning.

[0003] In existing technologies, user entity behavior analysis alone can only identify superficial access anomalies such as out-of-town or high-frequency access, but cannot identify business logic vulnerabilities such as payment bypass and amount tampering, easily leading to false negatives. Relying solely on business sequence detection cannot distinguish between normal user actions and malicious attacks, resulting in a high number of false positives. Existing fusion methods simply superimpose two types of detection results without unified feature extraction or collaborative risk score calculation, still resulting in false negatives and false negatives in scenarios involving minor behaviors combined with serious violations. Furthermore, traditional detection relies on manual maintenance of business rules, leaving security monitoring gaps after business iterations and lacking automated attack tracing and group identification capabilities. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a method, apparatus, and related products for API risk identification. These methods simultaneously extract two types of features: entity behavior and business sequence, and perform parallel detection using dual engines. A weighted algorithm is used to fuse the scores of the two types of anomalies to determine risk levels. An adaptive model is employed to automatically update the behavioral baseline and business process rules. Based on entity association graphs, the methods achieve full-chain attack tracing and group attack identification, balancing access behavior and business process detection. This reduces false positives and false negatives and adapts to continuously changing business scenarios.

[0005] This application discloses a method for API risk identification, the method comprising: Extract user entity behavior analysis (UEBA) features and business sequence features from the target application programming interface (API) call requests; The UEBA features are subjected to behavior anomaly detection to obtain behavior anomaly detection results; The business sequence features are subjected to sequence violation detection to obtain the sequence violation detection results; By integrating the abnormal behavior detection results and the sequence violation detection results, a comprehensive risk detection result is obtained; Based on the target API call request and the comprehensive risk detection results, an entity association graph is constructed to reconstruct the attack path and identify the attack behavior.

[0006] Optionally, the step of extracting user entity behavior analysis (UEBA) features and business sequence features from the target application programming interface (API) call request includes: Extract user entity information, call feature information, and business operation information from the target API call request; Invalid, duplicate, and incorrectly formatted data are removed from the user entity information, the call feature information, and the business operation information, and the data format is standardized to obtain the processed data; The UEBA feature and the service sequence feature are extracted from the processed data; the UEBA feature comes from the user entity information and the call feature information; the service sequence feature comes from the service operation information.

[0007] Optionally, the step of performing behavioral anomaly detection on the UEBA features to obtain behavioral anomaly detection results includes: The real-time UEBA features are compared with the behavioral baseline model to identify abnormal behaviors, including low-frequency slow attacks and gang-style low-frequency slow attacks. The abnormal behavior score is calculated based on the abnormal behavior to obtain the abnormal behavior detection result; the behavior baseline model is constructed based on historical UEBA features.

[0008] Optionally, the step of performing sequence violation detection on the business sequence features to obtain the sequence violation detection result includes: The real-time characteristics of the business sequence are compared with the business process compliance model to identify non-compliant businesses. The non-compliant businesses include payment process bypass, order amount tampering, verification code skipping, repeated coupon redemption, and overselling and duplicate delivery caused by condition competition. The business process compliance model is built by artificial intelligence (AI) by learning historical legitimate API call sequence data and compliant business process data. The sequence violation score is calculated based on the violation, and the sequence violation detection result is obtained.

[0009] Optionally, the fusion of the behavioral anomaly detection results and the sequence violation detection results to obtain a comprehensive risk detection result includes: The abnormal behavior detection results and the sequence violation detection results are weighted and fused to calculate the total fusion score, thus obtaining the comprehensive risk detection result.

[0010] Optionally, the calculation of the total score to obtain the comprehensive risk detection result includes: The risk level is determined based on the range of values ​​for the total fusion score; Differentiated strategies are configured based on different risk levels to obtain the comprehensive risk detection results.

[0011] Optionally, constructing the entity association graph based on the target API call request and the comprehensive risk detection result includes: The user account, login address, device fingerprint, API interface, business order, and operation sequence in the target API call request are constructed into entity nodes; Obtain the information related to each entity node in the target API call request and the comprehensive risk detection result to obtain the node attributes of each entity node; Based on the association relationships of each entity node in the target API call request, directed association edges are generated to generate the entity association graph.

[0012] Based on the above-mentioned method for API risk identification, this application also discloses an apparatus for API risk identification, comprising: an extraction unit, a behavior anomaly detection unit, a sequence violation detection unit, a fusion unit, and a construction unit; The extraction unit is used to extract User Entity Behavior Analysis (UEBA) features and business sequence features from the target application programming interface (API) call request; The behavior anomaly detection unit is used to perform behavior anomaly detection on the UEBA features and obtain behavior anomaly detection results; The sequence violation detection unit is used to perform sequence violation detection on the business sequence features and obtain the sequence violation detection result; The fusion unit is used to fuse the abnormal behavior detection results and the sequence violation detection results to obtain a comprehensive risk detection result; The construction unit is used to construct an entity association graph based on the target API call request and the comprehensive risk detection results, so as to reconstruct the attack path and identify the attack behavior.

[0013] Optionally, the extraction unit includes: The information extraction subunit is used to extract user entity information, call feature information, and business operation information from the target API call request. The processing subunit is used to remove invalid, duplicate, and incorrectly formatted data from the user entity information, the call feature information, and the business operation information, and to unify the data format to obtain processed data. The feature extraction subunit is used to extract the UEBA feature and the service sequence feature from the processed data; the UEBA feature comes from the user entity information and the call feature information; the service sequence feature comes from the service operation information.

[0014] Optionally, the behavior anomaly detection unit includes: The behavior comparison subunit is used to compare the real-time UEBA features with the behavior baseline model to identify abnormal behaviors; the abnormal behaviors include low-frequency slow attacks and gang-style low-frequency slow attacks. The behavior detection subunit is used to calculate the behavior anomaly score based on the abnormal behavior to obtain the behavior anomaly detection result; the behavior baseline model is constructed based on historical UEBA features.

[0015] Optionally, the sequence violation detection unit includes: The business comparison subunit is used to compare the real-time business sequence characteristics with the business process compliance model to identify non-compliant businesses. The non-compliant businesses include payment process bypass, order amount tampering, verification code skipping, duplicate coupon redemption, and overselling and duplicate delivery caused by condition competition. The business process compliance model is constructed by artificial intelligence (AI) by learning historical legitimate API call sequence data and compliant business process data. The business detection subunit is used to calculate the sequence violation score based on the violation business and obtain the sequence violation detection result.

[0016] Optionally, the fusion unit includes: The fusion calculation subunit is used to weight and fuse the behavior anomaly detection results and the sequence violation detection results to calculate the fusion total score and obtain the comprehensive risk detection result.

[0017] Optionally, the fusion computing subunit includes: The risk level determination subunit is used to determine the risk level based on the range of values ​​for the total fusion score; The strategy configuration subunit is used to configure differentiated strategies according to different risk levels to obtain the comprehensive risk detection results.

[0018] Optionally, the building unit includes: The node construction subunit is used to construct the user account, login address, device fingerprint, API interface, business order and operation sequence in the target API call request into an entity node; The attribute acquisition subunit is used to acquire information related to each entity node in the target API call request and the comprehensive risk detection result to obtain the node attributes of each entity node. The graph generation subunit is used to generate directed association edges based on the association relationships of each entity node in the target API call request, thereby generating the entity association graph.

[0019] Based on the above-mentioned method for API risk identification, this application also discloses a computer program product, which includes a computer program that, when run by a processor, is used to implement any of the above-mentioned methods.

[0020] Based on the above-mentioned method for API risk identification, this application also discloses a storage medium for storing computer program instructions, which, when executed by a central processing unit, are used to implement any of the above-mentioned methods.

[0021] This application discloses a method, apparatus, and related products for API risk identification. It synchronously extracts user entity behavior analysis (UEBA) features and business sequence features from target application programming interface (API) call requests, using the same original request data to support dual-dimensional detection of user behavior and business processes. Behavioral anomaly detection based solely on UEBA features can accurately identify anomalies at the user behavior level, such as logins from different locations, batch crawling, low-frequency slow attacks, and group calls. Sequence violation detection based solely on business sequence features can identify business logic vulnerabilities such as payment bypass, amount tampering, and missing steps. Integrating the behavioral anomaly detection results and sequence violation detection results yields a comprehensive risk detection result, overcoming the limitation of simple process detection in distinguishing between normal and abnormal operations, and solving the defect of failing to detect business-related attacks relying solely on behavioral analysis. This significantly reduces false positives and false negatives, enabling accurate identification of complex attacks such as compound fraud and crawler bypass. Ultimately, based on the target API call requests and comprehensive risk detection results, an entity association graph is constructed, which fully connects various related objects such as accounts, devices, Internet Protocol (IP) addresses, interfaces, and business orders. This allows for the rapid and accurate reconstruction of the complete attack chain from the attack entry point to the target data, accurately locating the source of the attack, and providing a complete and intuitive basis for security handling and audit tracing. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating an API risk identification method disclosed in an embodiment of this application; Figure 2This is a flowchart illustrating another method for API risk identification disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an API risk identification device disclosed in an embodiment of this application. Detailed Implementation

[0024] Existing API risk detection technologies suffer from multiple shortcomings. Solutions relying solely on user entity behavior analysis can only identify account-level anomalies such as cross-regional or batch anomalies, failing to detect business logic vulnerabilities like payment bypass and amount tampering, and are prone to missing complex attacks. Solutions relying solely on business sequence detection cannot distinguish between normal user actions and malicious process tampering, resulting in a high number of false positives. Some simplified fusion solutions simply superimpose two types of alarm results without deep integration of features and decision-making, still leading to numerous missed and false positives. Furthermore, traditional detection solutions rely on manual configuration of numerous business rules, requiring manual maintenance after business process updates, creating security gaps, and lacking the ability to fully trace attack sources and identify organized attacks.

[0025] To address the aforementioned issues, this application first collects standardized API request data, simultaneously extracts user entity behavior analysis features and business sequence features, outputs behavior anomaly scores and sequence violation scores in parallel through dual detection engines, calculates the comprehensive risk level using a weighted fusion algorithm and matches it with a graded handling strategy, and finally constructs a correlation graph based on various entity data to complete attack path reconstruction and gang attack identification. The entire solution is equipped with an adaptive learning model, requiring no extensive manual rule configuration and can continuously adapt to dynamic business changes.

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Example 1: This application discloses a method for API risk identification.

[0028] For details, please refer to Figure 1 The API risk identification method disclosed in this embodiment includes the following steps: Step 101: Extract user entity behavior analysis UEBA features and business sequence features from the target application programming interface (API) call requests.

[0029] In this embodiment, this step serves as a pre-processing and feature extraction stage before the entire detection process. Its core function is to clean and standardize the original API call request, and simultaneously separate and generate two types of core detection features. This provides unified, standardized, and noise-free feature input data for subsequent dual-engine parallel detection. This step uniformly disassembles all basic information from the same original API call request, cleans and standardizes it, and then extracts two types of features separately. This ensures that the User and Entity Behavior Analytics (UEBA) features and business sequence features are time-synchronized and share the same information source. This eliminates detection errors caused by feature fragmentation from the data source and supports the subsequent dual detection mechanism for behavior and process.

[0030] As a feasible solution, the first step is to extract user entity information, call characteristic information, and business operation information from the target API call requests. A dual collection method combining side-channel monitoring and sequential data acquisition can be employed to comprehensively capture all API call requests initiated by front-end users, back-end services, and third parties, extracting basic information from the original request messages. This basic information can include user entity information, call characteristic information, and business operation information. User entity information may include user account, login IP, device fingerprint, geographical location, and account permission level. Call characteristic information may include call frequency, query rate per second, call time period, interface response time, return code, and data volume returned per request. Business operation information may include API call path, interface call order, request parameters, and current business status. These three types of basic information comprehensively cover the three major detection dimensions of entity profile, call behavior, and business process, with no critical information omitted.

[0031] After information extraction, invalid, duplicate, and incorrectly formatted data in user entity information, call feature information, and business operation information are removed, and the data format is standardized to obtain processed data. The noise reduction stage filters out empty requests, illegal or malformed requests, and repeatedly sent requests from the same source, removing error messages with missing fields or corrupted characters to prevent dirty data from interfering with subsequent feature calculations. Standardization converts all heterogeneous data into a pre-defined standardized format. For example, timestamps are standardized into standard strings of "year, month, day, hour, minute, second," and various request parameters are converted to standard string formats, eliminating data format differences caused by different businesses and callers. This ensures stable operation of subsequent feature extraction logic and outputs standardized processed data with a unified format.

[0032] In this embodiment, based on the processed standardized data, UEBA features and business sequence features are extracted simultaneously. UEBA features originate from user entity information and call feature information, while business sequence features originate from business operation information. For UEBA features, user behavior baseline features such as historical call habits, frequently used interfaces, and average call frequency can be extracted from user entity information and call feature information; IP profile features such as IP address and IP access history; device behavior features such as device model and fixed login time period; and call frequency features such as real-time query rate per second and deviation from historical baseline, comprehensively depicting the long-term, real-time behavioral patterns of entities. For business sequence features, sequence features such as sequential call relationships of interfaces, compliant business status jump rules, dependency relationships of interface pre-steps, and the rationality of request parameter meanings can be extracted from business operation information. Simultaneously, a large model pre-trained by a converter is used to complete parameter semantic analysis, identifying sensitive fields and abnormally tampered parameters. The simultaneous output of these two types of features across two dimensions enables subsequent dual-engine detection modules to perform parallel real-time computation, providing complete and synchronous feature input for calculating behavior anomaly scores and sequence violation scores.

[0033] This step involves the integrated collection and generation of UEBA features and business sequence features through source information decomposition, unified denoising and standardization, and domain-specific features. This solves the shortcomings of traditional solutions, such as fragmented feature data, asynchronous time sequence, and dirty data interfering with detection accuracy. The unified feature extraction system provides a reliable data foundation for subsequent deep integration and decision-making of behavior and processes, while taking into account real-time performance and data integrity, and adapting to continuous detection scenarios with full online API call traffic.

[0034] Step 102: Perform behavioral anomaly detection on the UEBA features to obtain behavioral anomaly detection results.

[0035] In this embodiment, this step follows the UEBA features output in step 101, relies on a dynamically adaptive behavioral baseline model to complete entity behavior anomaly identification, and outputs a quantified behavioral anomaly score, providing a basis for risk determination at the behavioral dimension for subsequent fusion decision-making. This step, through the complete logic of constructing a baseline from historical behavior, identifying anomalies through real-time feature comparison, and quantifying and outputting anomaly scores, can accurately identify various types of entity abnormal behavior, fill the blind spots in user behavior dimension identification in business process detection, achieve independent quantitative assessment of behavioral risks, and support the deep fusion and determination of subsequent behavior and process detection results.

[0036] As a feasible solution, a behavioral baseline model is pre-built based on historical UEBA features. This model can be trained using one or more unsupervised learning algorithms, such as Isolation Forest, Autoencoder, Temporal Prediction, and Density Clustering. The model has dynamic adaptive update capabilities, and the update cycle can be set by the user. When the fluctuation range of normal user behavior exceeds a preset value (e.g., 30%), a new user entity is added, or the API interface type changes, the baseline threshold is automatically updated. The updated baseline threshold is the average of the most recent normal user behavior features, continuously aligning with users' long-term stable operating habits and avoiding the inevitable failure of a fixed baseline due to changes in business and users.

[0037] After the baseline model is built, the real-time UEBA features are compared with the behavioral baseline model to identify abnormal behavior. Abnormal behavior can include low-frequency slow attacks and group-based low-frequency slow attacks. The real-time extracted multi-dimensional UEBA features of users, IPs, and devices are compared with the historical normal behavior features stored in the baseline model to determine the degree of deviation. Any feature deviating from the baseline threshold is considered to indicate the presence of abnormal behavior.

[0038] As a feasible solution, for highly covert low-frequency, slow attacks, identification can be achieved through a comprehensive analysis of call frequency, parameter patterns, and long-term behavioral consistency. For example, identification criteria can be set as a query rate of less than or equal to 5 per second, a continuous call duration of greater than or equal to 30 minutes, and a parameter continuously increasing or decreasing sequence length greater than or equal to 50. Simultaneously, density clustering algorithms are used to aggregate similar abnormal call behaviors. When the number of entities exhibiting the same abnormal behavior is no less than 3, it is identified as a group-based low-frequency, slow attack. This approach can accurately capture covert attack behaviors of black market operators that involve mass, slow scanning and enumeration of accounts and interfaces, overcoming the shortcomings of traditional methods in identifying low-frequency, slow attacks.

[0039] As a feasible solution, after identifying abnormal behavior, an anomaly score can be calculated based on the abnormal behavior to obtain the behavior anomaly detection result. Specifically, a behavior anomaly score can be generated by converting the overall deviation between real-time UEBA features and the behavior baseline model. The score range can be set by the user; the higher the score, the more serious the deviation of the current user, IP, or device's behavior from the normal baseline, and the higher the risk level of the entity's behavior. This behavior anomaly score, as a standardized and quantified result, carries all behavioral dimension risk information from this step and is directly and synchronously transmitted to the subsequent fusion decision module. It is used to calculate the comprehensive risk score by weighting and fusing with the business sequence violation score, completing the basic data output for dual verification.

[0040] This step builds a dynamically updatable behavior baseline model based on historical UEBA features. Through real-time feature comparison, it accurately identifies ordinary behavioral anomalies, low-frequency slow attacks, and gang-related low-frequency slow attacks. Finally, it outputs a unified and quantified behavioral anomaly score, which fully covers all risk identification scenarios at the user entity level. This solves the shortcomings of traditional detection methods that are difficult to identify hidden low-frequency attacks and cannot quantify behavioral risks.

[0041] Step 103: Perform sequence violation detection on the business sequence features to obtain the sequence violation detection results.

[0042] In this embodiment, this step and step 102 can be executed in parallel. Taking the business sequence features output from step 101, and relying on a business process compliance model autonomously trained by Artificial Intelligence (AI), this step completes the identification of violations at the business logic level, outputting a quantified sequence violation score. This provides a basis for risk determination at the business process dimension for subsequent fusion decisions. This step verifies the business flow logic of API calls, accurately capturing various process bypass and parameter tampering attacks, compensating for the shortcomings of behavior anomaly detection in business semantic recognition, and complementing the behavior detection in step 102 to jointly build a two-dimensional detection system.

[0043] As a feasible solution, a business process compliance model is pre-built using AI to learn historical legitimate API call sequence data and manually labeled compliant business process data. This model can be trained using one or two deep sequence models, such as Transformer and Long Short-Term Memory (LSTM) networks. The AI's automatic learning cycle can be set (e.g., 0.5 to 12 hours). When online business processes change, the model can automatically receive the change information and relearn to generate new compliance rules; the relearning time can be set to no more than 1 hour. The entire process requires no manual configuration or modification of business sequence rules, eliminating security detection gaps during business iterations and adapting to dynamically changing online business scenarios.

[0044] After constructing the business process compliance model, real-time business sequence characteristics are compared with the model to identify non-compliant business activities. Non-compliant activities can include payment process bypassing, order amount tampering, CAPTCHA skipping, duplicate coupon redemption, and overselling or duplicate shipments due to race conditions. The sequence characteristics extracted in real-time, such as the order of API calls, business status transitions, dependencies of preceding steps, and request parameter semantics, are compared item by item with the legitimate business flow rules stored in the model. Once a non-compliant transition, missing information, or abnormal parameters are found, a non-compliant business activity is identified.

[0045] Typical violation scenarios include: payment process bypass (calling the order completion interface directly without a payment callback step), order amount tampering (the amount in the request parameters deviates from the product's base price by more than 30%), CAPTCHA skipping (calling the verification interface directly without executing the CAPTCHA sending step), and abnormal call sequences of claiming, canceling, and repeatedly claiming coupons. It can also identify business risks such as overselling and duplicate order shipments caused by race conditions, comprehensively covering high-frequency business logic attack types in the industry. Furthermore, for request parameters, semantic analysis is performed using a pre-trained large model of the converter, combined with regular expressions for dual identification of sensitive fields and tampered abnormal values, further improving the accuracy of violation identification.

[0046] As a feasible solution, after identifying all non-compliant business transactions, a sequence violation score can be calculated based on these transactions to obtain the sequence violation detection result. Specifically, a sequence violation score can be generated based on the severity of the deviation of the current business sequence from the compliance process. The score range can be set by the user, with higher scores indicating more severe business logic vulnerabilities in the current API call. This standardized and quantified sequence violation score is pushed to the subsequent fusion decision module in real time, where it is weighted and fused with the behavioral anomaly score output in step 102 to achieve collaborative determination of behavioral and process risks, providing quantitative data support from a business dimension for comprehensive risk level classification.

[0047] This step relies on AI to autonomously learn and build a dynamic, adaptive business process compliance model. It eliminates the need for manual maintenance of process rules and automatically adapts to business iterations and updates. Through real-time business sequence feature comparison, it accurately identifies various typical business violations and attacks, and quantifies and outputs standardized sequence violation scores. This addresses the shortcomings of traditional single-behavior detection, which cannot identify business logic vulnerabilities, and the poor adaptability of manually configured rules. In conjunction with the behavior anomaly detection unit, it achieves comprehensive API risk identification from both entity behavior and business process dimensions.

[0048] Step 104: Combine the abnormal behavior detection results and the sequence violation detection results to obtain a comprehensive risk detection result.

[0049] In this embodiment, this step follows the behavioral anomaly score output in step 102 and the sequence violation score output in step 103, completing the deep weighted fusion and hierarchical judgment of the two types of detection results, and outputting a comprehensive risk detection result with corresponding handling strategies. This step achieves collaborative judgment of behavioral risk and process risk by uniformly quantifying the two types of risks through weighted fusion, dividing the risk level into multiple levels according to score ranges, and matching hierarchical handling strategies. This significantly reduces the overall false alarm rate and false negative rate of detection, providing a standardized decision-making basis for real-time interception, alarm, and observation and control of API requests.

[0050] As a feasible solution, the results of behavioral anomaly detection and sequence violation detection are first weighted and fused to calculate a total fusion score. For example, the weighted fusion algorithm can be configured to multiply the behavioral anomaly score by its weight, and then add the sequence violation score to its business weight to obtain the total fusion score. Simultaneously, this module includes an independent policy configuration unit, allowing operations personnel to dynamically adjust the weight coefficients corresponding to the two types of scores as needed. Weight modifications take effect immediately without requiring a restart of the entire detection system. This approach is adaptable to the differentiated security control priorities of various industries such as e-commerce, finance, and mini-programs, flexibly balancing the judgment weights of entity behavior risks and business process vulnerability risks. For example, when both the behavioral anomaly score and the sequence violation score range from 0 to 100, and the sum of their weights is 1, the total fusion score also ranges from 0 to 100, facilitating a unified risk level classification.

[0051] As a feasible solution, after calculating the total fusion score, the risk level can be determined based on the range of the total fusion score. For example, a four-stage risk grading standard can be preset, strictly dividing the corresponding risk levels according to the score range: when the total fusion score is less than 40, it is judged as a normal level. When the total fusion score is greater than or equal to 40 and less than 70, it is judged as a low-risk level. When the total fusion score is greater than or equal to 70 and less than 90, it is judged as a high-risk level. When the total fusion score is greater than or equal to 90, it is judged as an extremely high-risk level. The grading boundary thresholds also support manual customization, allowing enterprises to flexibly adjust the boundary scores for each level according to their own security control requirements, adapting to different security protection standards.

[0052] After determining the risk level, differentiated strategies can be configured based on different risk levels to obtain a comprehensive risk detection result. As an example, a specific handling and execution strategy is matched for each risk level: Normal level requests are allowed without any restrictions. Low-risk requests are observed, retaining all request logs without triggering interception or manual intervention. High-risk requests are alerted, pushing risk information to the security operations backend and triggering a manual review process. Extremely high-risk requests are blocked, directly halting the current API call and temporarily banning the corresponding account, login IP, and device fingerprint for 1 to 24 hours. The ban duration can be adjusted based on the manual review results; if the review confirms a misjudgment, the ban can be lifted immediately; if a malicious attack is confirmed, the ban can be extended or even permanently banned. Finally, the total score, corresponding risk level, and matched handling strategy are combined to form a complete comprehensive risk detection result, which is simultaneously distributed to the execution control end, entity association tracing module, and log storage module to support the entire process of subsequent risk tracing, log retention, and real-time control.

[0053] This step achieves deep collaborative judgment of both behavioral and business-related detection results through weighted fusion, unlike the outdated fusion method of simply splicing results. It also provides full-dimensional customizable configuration capabilities for weights, tiered thresholds, and handling strategies, balancing judgment accuracy with solution versatility. The tiered and differentiated handling logic enables layered risk management, avoiding indiscriminate blocking that could disrupt normal business operations, and preventing high-risk malicious requests from being allowed to pass. From a decision-making perspective, it improves the entire API risk detection system, effectively reducing the number of false positives and false negatives, and enhancing the precision of online API security management.

[0054] Step 105: Based on the target API call request and the comprehensive risk detection results, construct an entity association graph to reconstruct the attack path and identify attack behaviors.

[0055] In the API risk intelligent detection method of this embodiment, this step follows the comprehensive risk detection results output in step 104, and combines the original API call full data to build an entity association graph, completing attack path reconstruction and gang attack identification, providing complete tracing support for security operations personnel to handle online risk events. This step generates a complete graph by building standardized entity nodes, binding risk attributes, and generating association relationships, connecting scattered API call data, realizing full-link visualized tracing of risks, and supplementing the post-event analysis and gang identification capabilities of the entire detection system.

[0056] As a feasible solution, the user account, login address, device fingerprint, API interface, business order, and operation sequence in the target API call request are first constructed as entity nodes. All entity objects are extracted from the standardized request data preprocessed in step 101, and each type of independent object is defined as a separate entity node in the graph. The nodes cover all core elements involved in the entire attack chain: including the user account initiating the call, the login IP address representing the access source, the device fingerprint identifying the hardware carrier, the API interface representing the business interaction, the business order that generates financial and data risks, and the complete operation sequence recording the flow process. No key entities are omitted, fully covering all participating objects from the attack entry point to the business target, thus building a complete node foundation for the graph.

[0057] After the entity nodes are created, information related to each entity node and comprehensive risk detection results from the target API call request are obtained to obtain the node attributes of each entity node. The basic entity information, call behavior characteristics, business flow information, and comprehensive risk detection results such as the fusion score, risk level, and handling strategy output in step 104 are uniformly bound to the corresponding entity node as node attributes. Different entity nodes can carry differentiated attributes. For example, user account nodes are bound to the account's historical call records and account risk level; IP nodes are bound to the region and historical abnormal call frequency; order nodes are bound to the order amount and whether there is a tampering violation mark, etc. The risk judgment conclusion of a single request is pushed down to each associated entity, giving the graph the ability to mark entity risks, which facilitates the rapid location of high-risk nodes.

[0058] Subsequently, directed edges are generated based on the relationships between entity nodes in the target API call request, creating an entity relationship graph. The graph outlines the objective flow, subordination, and call relationships between entities in the API call behavior, generating directed edges according to the chronological order of the actions to clearly show the interaction flow between entities. For example, a directed edge represents the call chain from user account to login IP to device fingerprint to API interface to business order. The entire entity relationship graph is persistently stored in a graph database, supporting real-time synchronization of newly added API call data to dynamically update relationships. Based on the complete graph, attack paths can be reconstructed quickly, fully outlining the complete attack chain from the attack entry point to the target business interface, accurately locating the attack source IP, device, and account. Simultaneously, cluster analysis is performed based on the multi-entity relationships within the graph to quickly identify coordinated attacks involving multiple accounts, devices, and IPs, outputting group characteristics and all associated risk entities. The generated entity relationship graph is synchronously pushed to the log storage module for archiving, while the source analysis results are fed back to the fusion decision module to assist the model update module in iteratively optimizing the detection model.

[0059] This step constructs a directed association graph with risk attributes based on entities across the entire chain, connects scattered and independent API call request data, enables rapid reconstruction of attack paths and identification of group attacks, and solves the shortcomings of traditional detection methods that cannot connect multi-dimensional risk entities and are difficult to trace the full picture of attacks. It provides visualized data support for security incident review, risk impact assessment, and group tracking.

[0060] Example 2: This application discloses another method for API risk identification; please refer to [link / reference]. Figure 2 This embodiment describes the process of constructing an entity association graph.

[0061] Step 201: Perform two-dimensional feature extraction on the real-time acquired target API call requests to obtain UEBA features and business sequence features. Proceed to steps 202 and 203.

[0062] Step 202: Compare the UEBA features with the behavioral baseline model, and calculate the behavioral anomaly score based on the comparison results. Proceed to Step 204.

[0063] Step 203: Compare the business sequence characteristics with the business process compliance model, and calculate the sequence violation score based on the comparison results. Proceed to Step 204.

[0064] Step 204: The abnormal behavior score and the sequence violation score are weighted and summed according to the preset weights to obtain the total fusion score.

[0065] Step 205: Determine the risk level and the corresponding security handling strategy based on the range of the total score.

[0066] Step 206: Based on the target API call request, and the corresponding total score, risk level, and security handling strategy, construct the part corresponding to the target API call request in the entity association graph. Return to step 201.

[0067] Based on the API risk identification method disclosed in the above embodiments, this embodiment correspondingly discloses an API risk identification apparatus. Please refer to... Figure 3 The device for API risk identification includes: an extraction unit 301, a behavior anomaly detection unit 302, a sequence violation detection unit 303, a fusion unit 304, and a construction unit 305; The extraction unit 301 is used to extract User Entity Behavior Analysis (UEBA) features and business sequence features from the target application programming interface (API) call request. The behavior anomaly detection unit 302 is used to perform behavior anomaly detection on the UEBA features and obtain behavior anomaly detection results. The sequence violation detection unit 303 is used to perform sequence violation detection on the business sequence features and obtain a sequence violation detection result. The fusion unit 304 is used to fuse the behavior anomaly detection result and the sequence violation detection result to obtain a comprehensive risk detection result; The construction unit 305 is used to construct an entity association graph based on the target API call request and the comprehensive risk detection result, so as to reconstruct the attack path and identify the attack behavior.

[0068] Optionally, the extraction unit 301 includes: The information extraction subunit is used to extract user entity information, call feature information, and business operation information from the target API call request. The processing subunit is used to remove invalid, duplicate, and incorrectly formatted data from the user entity information, the call feature information, and the business operation information, and to unify the data format to obtain processed data. The feature extraction subunit is used to extract the UEBA feature and the service sequence feature from the processed data; the UEBA feature comes from the user entity information and the call feature information; the service sequence feature comes from the service operation information.

[0069] Optionally, the abnormal behavior detection unit 302 includes: The behavior comparison subunit is used to compare the real-time UEBA features with the behavior baseline model to identify abnormal behaviors; the abnormal behaviors include low-frequency slow attacks and gang-style low-frequency slow attacks. The behavior detection subunit is used to calculate the behavior anomaly score based on the abnormal behavior to obtain the behavior anomaly detection result; the behavior baseline model is constructed based on historical UEBA features.

[0070] Optionally, the sequence violation detection unit 303 includes: The business comparison subunit is used to compare the real-time business sequence characteristics with the business process compliance model to identify non-compliant businesses. The non-compliant businesses include payment process bypass, order amount tampering, verification code skipping, duplicate coupon redemption, and overselling and duplicate delivery caused by condition competition. The business process compliance model is constructed by artificial intelligence (AI) by learning historical legitimate API call sequence data and compliant business process data. The business detection subunit is used to calculate the sequence violation score based on the violation business and obtain the sequence violation detection result.

[0071] Optionally, the fusion unit 304 includes: The fusion calculation subunit is used to weight and fuse the behavior anomaly detection results and the sequence violation detection results to calculate the fusion total score and obtain the comprehensive risk detection result.

[0072] Optionally, the fusion computing subunit includes: The risk level determination subunit is used to determine the risk level based on the range of values ​​for the total fusion score; The strategy configuration subunit is used to configure differentiated strategies according to different risk levels to obtain the comprehensive risk detection results.

[0073] Optionally, the building unit 305 includes: The node construction subunit is used to construct the user account, login address, device fingerprint, API interface, business order and operation sequence in the target API call request into an entity node; The attribute acquisition subunit is used to acquire information related to each entity node in the target API call request and the comprehensive risk detection result to obtain the node attributes of each entity node. The graph generation subunit is used to generate directed association edges based on the association relationships of each entity node in the target API call request, thereby generating the entity association graph.

[0074] Based on the above-mentioned method for API risk identification, this application also discloses a computer program product, which includes a computer program that, when run by a processor, is used to implement any of the above-mentioned methods.

[0075] Based on the above-mentioned method for API risk identification, this application also discloses a storage medium for storing computer program instructions, which, when executed by a central processing unit, are used to implement any of the above-mentioned methods.

[0076] The embodiments in this specification are described in a progressive manner. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant details can be found in the method section.

[0077] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0078] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0079] The features described in the embodiments of this specification can be substituted for or combined with each other, so that those skilled in the art can implement or use this application.

[0080] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for API risk identification, characterized in that, include: Extract user entity behavior analysis (UEBA) features and business sequence features from the target application programming interface (API) call requests; The UEBA features are subjected to behavior anomaly detection to obtain behavior anomaly detection results; The business sequence features are subjected to sequence violation detection to obtain the sequence violation detection results; By integrating the abnormal behavior detection results and the sequence violation detection results, a comprehensive risk detection result is obtained; Based on the target API call request and the comprehensive risk detection results, an entity association graph is constructed to reconstruct the attack path and identify the attack behavior.

2. The method according to claim 1, characterized in that, The extraction of User Entity Behavior Analysis (UEBA) features and business sequence features from the target application programming interface (API) call requests includes: Extract user entity information, call feature information, and business operation information from the target API call request; Invalid, duplicate, and incorrectly formatted data are removed from the user entity information, the call feature information, and the business operation information, and the data format is standardized to obtain the processed data; The UEBA feature and the service sequence feature are extracted from the processed data; the UEBA feature comes from the user entity information and the call feature information; the service sequence feature comes from the service operation information.

3. The method according to claim 1, characterized in that, The step of performing behavioral anomaly detection on the UEBA features to obtain behavioral anomaly detection results includes: The real-time UEBA features are compared with the behavioral baseline model to identify abnormal behaviors, including low-frequency slow attacks and gang-style low-frequency slow attacks. The abnormal behavior score is calculated based on the abnormal behavior to obtain the abnormal behavior detection result; the behavior baseline model is constructed based on historical UEBA features.

4. The method according to claim 1, characterized in that, The step of performing sequence violation detection on the business sequence features to obtain the sequence violation detection result includes: The real-time characteristics of the business sequence are compared with the business process compliance model to identify non-compliant businesses. The non-compliant businesses include payment process bypass, order amount tampering, verification code skipping, repeated coupon redemption, and overselling and duplicate delivery caused by condition competition. The business process compliance model is built by artificial intelligence (AI) by learning historical legitimate API call sequence data and compliant business process data. The sequence violation score is calculated based on the violation, and the sequence violation detection result is obtained.

5. The method according to claim 1, characterized in that, The fusion of the abnormal behavior detection results and the sequence violation detection results yields a comprehensive risk detection result, including: The abnormal behavior detection results and the sequence violation detection results are weighted and fused to calculate the total fusion score, thus obtaining the comprehensive risk detection result.

6. The method according to claim 5, characterized in that, The calculation of the total score yields the comprehensive risk detection result, including: The risk level is determined based on the range of values ​​for the total fusion score; Differentiated strategies are configured based on different risk levels to obtain the comprehensive risk detection results.

7. The method according to claim 1, characterized in that, The construction of the entity association graph based on the target API call request and the comprehensive risk detection results includes: The user account, login address, device fingerprint, API interface, business order, and operation sequence in the target API call request are constructed into entity nodes; Obtain the information related to each entity node in the target API call request and the comprehensive risk detection result to obtain the node attributes of each entity node; Based on the association relationships of each entity node in the target API call request, directed association edges are generated to generate the entity association graph.

8. An apparatus for API risk identification, characterized in that, include: Extraction unit, behavior anomaly detection unit, sequence violation detection unit, fusion unit, and construction unit; The extraction unit is used to extract User Entity Behavior Analysis (UEBA) features and business sequence features from the target application programming interface (API) call request; The behavior anomaly detection unit is used to perform behavior anomaly detection on the UEBA features and obtain behavior anomaly detection results; The sequence violation detection unit is used to perform sequence violation detection on the business sequence features and obtain the sequence violation detection result; The fusion unit is used to fuse the abnormal behavior detection results and the sequence violation detection results to obtain a comprehensive risk detection result; The construction unit is used to construct an entity association graph based on the target API call request and the comprehensive risk detection results, so as to reconstruct the attack path and identify the attack behavior.

9. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, is used to implement the method described in any one of claims 1 to 7.

10. A storage medium, characterized in that, Used to store computer program instructions, which, when executed by a central processing unit, are used to implement the method described in any one of claims 1 to 7.