Method, device, electronic device and storage medium for processing data
Through the unaware online method of generating and loading target rule files in the risk control management system, the problem of timeliness caused by the adjustment of risk strategy indicators in the risk control business is solved, and the flexible configuration and efficient processing of risk assessment rules are achieved.
Patent Information
- Application Number
- CN202510668242.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the existing risk control management system, adjustments or changes in risk strategy indicators must be modified and recompiled/deployed before they can take effect, resulting in a decrease in the timeliness of risk control business processing and the inability to achieve smooth online launch.
Enter configuration information through the rule configuration interface to generate the target rule file, and update it to the rule file library through broadcast, realizing the unaware launch of risk assessment rules, and using common script template files and computing state machines for dynamic compilation and loading to avoid interruptions in risk control business.
It improves the timeliness of risk control business processing, reduces the cost of repeated code writing and maintenance, ensures the consistency of computing logic, supports diversified business scenario adaptation, and improves the efficiency and accuracy of rule matching.
Smart Images

Figure CN120197955B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and more specifically, to a method, device, electronic device and storage medium for processing data in the field of data processing technology. Background Art
[0002] Risk management systems typically analyze various behavioral data of platform users and, based on established risk strategy indicators, determine whether platform users are engaging in unusual behavior. For example, by analyzing data such as device fingerprints, Internet Protocol addresses, and participation time / frequency when users participate in marketing activities, and based on targeted risk strategy indicators, they can identify unusual user participation behaviors, including repeated account registrations using virtual numbers and batch coupon redemption using multiple accounts on the same device, to mitigate the risk of platform subsidies being misappropriated.
[0003] In existing risk control management systems, each risk strategy indicator is hard-coded, meaning that the logic and parameters for all risk strategy indicators are written into the code. This means that adjustments or changes to individual risk strategy indicators must be made by modifying the source code and recompiling / redeploying them before they take effect in the risk control management system. Consequently, these adjustments or changes to risk strategy indicators cannot be smoothly implemented in the risk control management system. This requires stopping the current real-time risk control service running in the risk control management system or interrupting ongoing business requests, which impacts the timeliness of risk control processing.
[0004] Therefore, a method for processing data is urgently needed to solve the above problems. Summary of the Invention
[0005] The present application provides a method, device, electronic device and storage medium for processing data. The method generates a target rule file through configuration information input through a rule configuration interface, and updates and saves the target rule file to a rule file library through broadcasting. This method can realize the imperceptible online launch of risk assessment rules (including risk strategy indicators), avoid the suspension or interruption of risk control business, and improve the timeliness of processing risk control business.
[0006] In a first aspect, a method for processing data is provided, which includes: obtaining configuration information input by a developer in a rule configuration interface, parsing the configuration information to obtain relevant rule configuration parameters; selecting a target script file according to the type of the rule configuration parameters, and generating a target rule file from the rule configuration parameters based on the format of the target script file; updating and saving the target rule file to a rule file library by broadcasting; obtaining behavior data of multiple users, and analyzing the behavior data based on multiple rule files in the rule file library to determine user behavior data that is at risk.
[0007] During the data processing process, the configuration information entered on the visual rule configuration interface is parsed to obtain rule configuration parameters. A target script file is then selected based on the type of the rule configuration parameters. The target script file is a generic script template file corresponding to the type. Based on the format of the target script file, the rule configuration parameters are used to generate a target rule file. This allows the existing rule instances in the target script file to be replaced, instantly compiling the new calculation logic and rule configuration parameters in the target rule instance into the target script file. Furthermore, the target rule file is updated and saved to the rule file repository via a broadcast mechanism, essentially loading the target rule file into the rule file repository. The direct dynamic compilation and loading capabilities of target rule files make them "pluggable" logical units. This allows risk assessment rules (including risk strategy indicators) to be seamlessly deployed without stopping or interrupting ongoing risk control operations within the risk control management system, improving the timeliness of risk control operations. Furthermore, through the broadcast mechanism, the target rule file only needs to be transmitted once to the rule file repository, and subsequent tasks can directly read the rule file from the rule file repository, avoiding wasted network bandwidth.
[0008] In combination with the first aspect, in certain implementations of the first aspect, selecting a target script file based on the type of the rule configuration parameter includes: comparing the type with multiple preset types, determining a target type that matches the type from the multiple preset types, each of the multiple preset types corresponding to a general script template file; and determining the script template file corresponding to the target type as the target script file.
[0009] During the data processing described above, pre-set types are strongly associated with script template files. This ensures consistency in calculation logic for the same type (i.e., rule calculation logic). For example, all matching rules automatically inherit the common logic for exact field matching, reducing the writing and maintenance costs of duplicate code. Furthermore, the above solution automatically identifies types and matches them to script template files, eliminating the need for developers to manually write or select the organization of script files. This reduces the risk of manual selection errors or coding errors, improving configuration efficiency and accuracy. Furthermore, the parameterized design of the universal script template files not only ensures a standardized execution process but also supports adaptation to diverse business scenarios. For example, the script template files corresponding to statistical rules can configure time windows and aggregation algorithms, while the script template files corresponding to matching rules can flexibly adjust the threshold for similarity determination. Therefore, the above technical solution balances rule standardization and business specificity within a unified script template framework, achieving both flexibility and reusability.
[0010] In combination with the first aspect and the above-mentioned implementation methods, in certain implementation methods of the first aspect, the behavior data is analyzed based on multiple rule files in the rule file library to determine the user behavior data that is at risk, including: determining multiple candidate files that match the behavior data from the multiple rule files; analyzing the behavior data through the operation state machine corresponding to each candidate file to determine the user behavior data that is at risk, the operation state machine encapsulating the calculation logic and rule configuration parameters of the candidate file.
[0011] During the data processing described above, a dynamic rule screening mechanism is used to match multiple candidate files from the rule file library. This avoids traversing the entire rule file and significantly reduces computational redundancy. Furthermore, each computational state machine encapsulates the computational logic and rule configuration parameters of the corresponding rule file to ensure that rules do not interfere with each other. At the same time, modifying a single rule file only requires replacing the corresponding computational state machine, without the need for global downtime or reloading, thus ensuring the continuity of risk control business processing. The layered processing mechanism of rule files and computational state machines can significantly improve the timeliness of risk assessment and the maintainability of the risk management system.
[0012] In combination with the first aspect and the above-mentioned implementation methods, in certain implementation methods of the first aspect, determining multiple candidate files that match the behavior data from the multiple rule files includes: determining multiple behavior types corresponding to the behavior data; comparing each behavior type with the behavior type corresponding to each rule file in the multiple rule files, and determining candidate behavior types that correspond one-to-one to the multiple behavior types from the behavior types corresponding to the multiple rule files; and determining the rule files corresponding to the multiple candidate behavior types as the multiple candidate files.
[0013] During the data processing described above, based on the multiple behavior types corresponding to the behavioral data of multiple users, candidate behavior types corresponding to these multiple behavior types are screened from the behavior types corresponding to multiple rule files. This pre-classified screening method can quickly narrow the scope of rule matching and significantly reduce computational redundancy. Furthermore, the strict one-to-one matching logic ensures that multiple candidate rule files are highly correlated with the behavioral data of multiple users (for example, only consumption expenditure rules are included in the analysis of transaction behavior), reducing false matches and improving the rule hit rate. Therefore, this technical solution can significantly optimize the efficiency and accuracy of rule screening through a precise matching mechanism between behavior types and rule files.
[0014] In combination with the first aspect and the above-mentioned implementation methods, in certain implementation methods of the first aspect, the behavior data is analyzed through the operation state machine corresponding to each candidate file to determine the user behavior data that is at risk, including: analyzing the behavior data through the operation state machine corresponding to each candidate file, generating an intermediate result, and storing the intermediate result in an external storage module, which is an independently deployed remote dictionary server cluster; and determining the user behavior data that is at risk based on the intermediate result through the operation state machine corresponding to each candidate file.
[0015] During the above data processing process, the intermediate results of the operation state machine corresponding to each candidate file in the multiple candidate files are stored in an external storage module. This can reduce the memory pressure of the electronic device and the consumption of the storage resources of the electronic device when the intermediate results are persisted in the memory of the electronic device. Replacing the full aggregation based on the memory of the electronic device with incremental aggregation based on the remote dictionary server can also significantly reduce the consumption of storage resources of the electronic device when the program is running.
[0016] In combination with the first aspect and the above-mentioned implementation methods, in some implementation methods of the first aspect, before updating and saving the target rule file to the rule file library by broadcasting, the method also includes: publishing the target rule file to the database middleware so that the database middleware routes the target rule file to a second external storage module, and the second external storage module supports distributed management.
[0017] In the above-mentioned data processing process, before the target rule file is updated and saved to the rule file library by broadcasting, the target rule file is also published to the database middleware to back up the target rule file to the second external storage module. Among them, the database middleware acts as a temporary buffer zone, which can temporarily store the target rule file before broadcasting, thereby preventing the target rule file from being lost due to direct overwriting of the rule file library due to transmission errors or incomplete reception. The second external storage module acts as an independent backup, which is physically isolated from the rule file library to prevent the target rule file from being completely lost due to hardware failure of the rule file library. The above-mentioned scheme can restore the original target rule file when the broadcast fails through double protection, avoiding the loss of the target rule file and entering an uncontrollable state.
[0018] In combination with the first aspect and the above-mentioned implementation methods, in certain implementation methods of the first aspect, the method also includes: when the first operation state machine corresponding to the first candidate file among the multiple candidate files successfully preempts the distributed lock and the behavior data is analyzed by the first operation state machine, the user behavior data at risk after the first moment is counted by the first operation state machine to obtain a first statistical result, and the first moment is the moment when the first operation state machine successfully preempts the distributed lock; when the second moment used to count the user behavior data at risk is the third moment, the first statistical result is output by the first operation state machine, and the second moment is later than the first moment.
[0019] In the process of processing data as described above, a distributed lock preemption mechanism is introduced in the process of analyzing the behavioral data of multiple users by multiple computing state machines. This can ensure that behavioral data of the same behavior type is only processed by a single computing state machine, avoid data competition caused by multi-node concurrent control, and ensure the consistency of the results of the statistical process. After the first computing state machine successfully preempts the distributed lock, it will combine the analysis of behavioral data according to the behavior type with the output of statistical results based on the dynamic time window. This can also avoid the pressure on resources caused by frequent result output. In other words, the above scheme can optimize resource utilization while ensuring the accuracy of the results through the collaborative design of the distributed lock preemption mechanism and the time window transmission mechanism, providing flexible and controllable technical support for real-time risk monitoring in high-concurrency scenarios, and effectively improving the efficiency and reliability of user behavior data analysis in a distributed environment.
[0020] In combination with the first aspect and the above-mentioned implementation methods, in certain implementation methods of the first aspect, the method for determining the third moment includes any one of the following: when the distributed lock is set with a first validity period, the third moment is determined based on the first moment and the first validity period; when the first preset duration is set by a preset timer, the third moment is determined based on the first moment and the first preset duration.
[0021] In combination with the first aspect and the above-mentioned implementation methods, in some implementation methods of the first aspect, the method also includes: determining whether the duration between the first moment and the second moment is equal to the second preset duration set by the cyclic timer; when the duration is equal to the second preset duration and the duration is not equal to the second validity period of the distributed lock, continue to analyze the behavioral data through the first operating state machine until the duration is equal to the second validity period, and output the first statistical result through the first operating state machine, and the second validity period is greater than the second preset duration.
[0022] During the data processing described above, a loop timer periodically triggers the computational state machine to analyze the behavioral data of the first category of behaviors at preset intervals, ensuring that data analysis is continuously attempted within the second validity period (the validity period of the distributed lock). Furthermore, the second validity period of the distributed lock provides a clear execution window for the analysis process, forcing processing progress through time constraints while ensuring task exclusivity. This ensures the integrity of risky user behavior data and automatically releases the distributed lock in the event of a timeout to trigger fault tolerance, effectively mitigating data latency or processing speed discrepancies and enhancing the risk management system's resilience to backpressure in high-concurrency scenarios.
[0023] In a second aspect, a device for processing data is provided, which includes: an acquisition unit for acquiring configuration information input by a developer in a rule configuration interface, and parsing the configuration information to obtain relevant rule configuration parameters; a generation unit for selecting a target script file according to the type of the rule configuration parameters, and generating a target rule file from the rule configuration parameters based on the format of the target script file; a storage unit for updating and saving the target rule file to a rule file library by broadcasting; a processing unit for acquiring behavior data of multiple users, and analyzing the behavior data based on multiple rule files in the rule file library to determine user behavior data that is at risk.
[0024] In combination with the second aspect, in certain implementations of the second aspect, the generation unit is specifically used to: compare the type with multiple preset types, determine a target type that matches the type from the multiple preset types, each of the multiple preset types corresponds to a general script template file; and determine the script template file corresponding to the target type as the target script file.
[0025] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the processing unit is specifically used to: determine multiple candidate files that match the behavior data from the multiple rule files; analyze the behavior data through the operation state machine corresponding to each candidate file to determine the user behavior data that is risky, and the operation state machine encapsulates the calculation logic and rule configuration parameters of the candidate file.
[0026] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the device also includes: a determination unit, used to: determine multiple behavior types corresponding to the behavior data; compare each behavior type with the behavior type corresponding to each rule file in the multiple rule files, and determine candidate behavior types that correspond one-to-one to the multiple behavior types from the behavior types corresponding to the multiple rule files; and determine the rule files corresponding to the multiple candidate behavior types as the multiple candidate files.
[0027] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the processing unit is further specifically used to: analyze the behavior data through the operation state machine corresponding to each candidate file, generate an intermediate result, and store the intermediate result in an external storage module, which is an independently deployed remote dictionary server cluster; determine the user behavior data that is risky based on the intermediate result through the operation state machine corresponding to each candidate file.
[0028] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, before updating and saving the target rule file to the rule file library by broadcasting, the device also includes: a publishing unit, used to publish the target rule file to the database middleware, so that the database middleware routes the target rule file to a second external storage module, and the second external storage module supports distributed management.
[0029] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the processing unit is also used to, when the first operation state machine corresponding to the first candidate file among the multiple candidate files successfully preempts the distributed lock and the behavior data is analyzed by the first operation state machine, count the user behavior data at risk after the first moment through the first operation state machine to obtain a first statistical result, and the first moment is the moment when the first operation state machine successfully preempts the distributed lock; the device also includes: an output unit, used to output the first statistical result through the first operation state machine when the second moment used for counting the user behavior data at risk is a third moment, and the second moment is later than the first moment.
[0030] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the determination unit is also used to: when the distributed lock is set with a first validity period, determine the third moment based on the first moment and the first validity period; when the first preset duration is set by a preset timer, determine the third moment based on the first moment and the first preset duration.
[0031] In combination with the second aspect and the above-mentioned implementation methods, in some implementation methods of the second aspect, the determination unit is also used to determine whether the duration between the first moment and the second moment is equal to the second preset duration set by the cyclic timer; the storage unit is also used to continue to analyze the behavioral data through the first operating state machine until the duration is equal to the second preset duration and the duration is not equal to the second validity period of the distributed lock, and output the first statistical result through the first operating state machine, and the second validity period is greater than the second preset duration.
[0032] In a third aspect, an electronic device is provided, comprising a memory and a processor. The memory is configured to store executable program code, and the processor is configured to retrieve and execute the executable program code from the memory, so that the electronic device executes the method of the first aspect or any possible implementation of the first aspect.
[0033] In a fourth aspect, a computer-readable storage medium is provided, which stores an executable program code. When the executable program code runs on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a functional architecture diagram of a data processing system provided in an embodiment of the present application;
[0035] Figure 2 This is a schematic diagram of the structure of a data processing system provided in an embodiment of the present application;
[0036] Figure 3 is a schematic flow chart of a method for processing data provided in an embodiment of the present application;
[0037] Figure 4 This is an architectural design diagram of a rule dynamic perception module provided in an embodiment of the present application;
[0038] Figure 5 This is a schematic diagram of dividing original behavior data into multiple message queue data in a message queue format provided by an embodiment of the present application;
[0039] Figure 6 is a structural diagram of another data processing system provided in an embodiment of the present application;
[0040] Figure 7 This is an architectural design diagram of a risk analysis and calculation module provided in an embodiment of the present application;
[0041] Figure 8 is a schematic flow chart of another method for processing data provided in an embodiment of the present application;
[0042] Figure 9 is a schematic flow chart of another method for processing data provided in an embodiment of the present application;
[0043] Figure 10 This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0044] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will clearly and thoroughly describe the technical solutions in this application in conjunction with the accompanying drawings. In the description of the embodiments of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more than two.
[0046] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0047] At present, risk control management systems usually analyze various behavioral data of platform users and determine whether platform users engage in abnormal behavior based on risk strategy indicators set by developers, that is, determine whether user behavior data poses risks.
[0048] In some embodiments, N platform users register accounts under the same Internet Protocol address within a first time period. After registering their accounts, M platform users conduct transactions of a first amount under the Internet Protocol address, where M is less than N and a positive integer greater than 1. The developer sets a risk assessment rule using risk strategy indicators as follows: if N is greater than 50, the first time period is one hour, M is greater than 40, and the first amount is greater than 1000, the registration and transaction behaviors of the M platform users are considered abnormal (i.e., if greater than 80% of the platform users register accounts under the same Internet Protocol address within one hour, and after registering their accounts, conduct more than 1000 transactions under the Internet Protocol address, then the registration and transaction behaviors of greater than 80% of the platform users are considered abnormal). For the above embodiments, the risk strategy indicators are the number of participants, the address of the event, the duration of time that multiple participants participate in the same event, and the consumption indicators corresponding to the participants.
[0049] In existing risk control management systems, each risk strategy indicator is hard-coded, with the logic and parameters of the risk strategy indicator written into the code. This means that adjustments or changes to individual risk strategy indicators must be made by modifying the source code and recompiling / redeploying them before they take effect in the risk control management system. Consequently, these adjustments or changes to risk strategy indicators cannot be smoothly implemented in the risk control management system. This requires stopping the current real-time risk control service running in the risk control management system or interrupting ongoing business requests, which impacts the timeliness of risk control processing.
[0050] Therefore, in response to the above problems, this application provides a method for processing data, which can realize the imperceptible online launch of risk assessment rules (including risk strategy indicators), avoid the suspension or interruption of risk control business, and improve the timeliness of processing risk control business.
[0051] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0052] A method for processing data provided by this application relies on Figure 1 The functional architecture shown, specifically, Figure 1The functional architecture includes the application layer, computing layer, and storage layer. The application layer includes the rule configuration center, data source management center, and result notification center. The rule configuration center is used to add, publish, and update rules. The data source management center is used to add and update data sources (which can be obtained or updated from multiple platforms). The data source is user behavior data. The result notification center is used for status monitoring, anomaly notification, and qualitative analysis. Specifically, it can continue to monitor the behavior data of at least one user, issue anomaly notifications when anomalies occur, and analyze the underlying causes of anomalies. The computing layer includes a data preprocessing module, a dynamic rule perception module, a risk analysis and calculation module, and a perception result output module. The data preprocessing module preprocesses the behavior data of multiple users. The dynamic rule perception module perceives the rules in the rule configuration center. The risk analysis and calculation module performs risk analysis on user behavior data based on the rules. The perception result output module processes and outputs risky user behavior data. The storage layer includes the TDDL (Taobao Distributed Data Layer) database middleware, the Redis cluster, the TT (Times Ten) database, and the HOLO database. The TDDL database middleware is used to route rules to an external storage module (such as MySQL). The Redis cluster is used to store the operation state machine corresponding to each rule, as well as the intermediate results after analyzing the behavior data of multiple users. The TT database can be used to store the statistical results of continuing to process the risky user behavior data. The HOLO database is a lightweight database based on the MySQL protocol, which is used to store the generated results after processing the statistical results under further trigger conditions.
[0053] It should be noted that the original behavior data preprocessed by the data preprocessing module above comes from the data source management center.
[0054] In addition, it should be noted that Figure 1 The Rule Configuration Center and Rule Dynamic Perception Module are newly added modules based on the existing functional architecture. The Rule Configuration Center is used to generate or update rule files and send them to the TDDL database middleware. The Rule Dynamic Perception Module is used to perceive rule files. The Risk Analysis and Calculation Module processes the behavioral data of multiple users based on rule files and generates risk statistics. The Perception Result Output Module is used to output risk results. The Result Notification Center monitors risk status, issues exception notifications, and performs anomaly analysis when anomalies are detected.
[0055] A method for processing data provided by this application relies on Figure 2The data processing system shown in FIG. 200 includes a rule configuration center 201 , a rule dynamic perception module 202 , a data source management center 203 , and a risk analysis and calculation module 204 . The risk analysis and calculation module 204 is connected to the data source management center 203 .
[0056] The data source management center 203 is used to obtain the behavior data of multiple users and send the behavior data of the multiple users to the risk analysis and calculation module 204;
[0057] The rule configuration center 201 is used to obtain configuration information entered by developers in the rule configuration interface, parse the configuration information to obtain relevant rule configuration parameters, select a target script file based on the type of the rule configuration parameters, generate a target rule file based on the format of the target script file, and publish the target rule file via broadcast.
[0058] The rule dynamic perception module 202 is used to update and save the target rule file to the rule file library of the task manager in the task manager when the target rule file is perceived;
[0059] The risk analysis and calculation module 204 is configured to obtain the behavior data of the multiple users and the multiple rule files in the rule file library, and analyze the behavior data based on the multiple rule files to determine the user behavior data that is risky.
[0060] Specifically, refer to Figure 3 , Figure 3 This is a schematic flow chart of a method for processing data provided in an embodiment of the present application. It should be understood that the method 300 can be applied to electronic devices with data processing capabilities, or cloud servers, etc. The embodiment of the present application does not limit the device form of the execution subject. The method 300 includes the following steps:
[0061] S301, obtaining configuration information input by a developer in a rule configuration interface, and parsing the configuration information to obtain relevant rule configuration parameters.
[0062] S302: Select a target script file according to the type of the rule configuration parameter, and generate a target rule file from the rule configuration parameter based on the format of the target script file.
[0063] That is, the specific steps executed by the rule configuration center 201 correspond to the above-mentioned steps 301 and 302.
[0064] It should be understood that the rule configuration interface in S301 is a visual interface for configuring a rule file. This rule file includes the calculation logic and rule configuration parameters for at least one rule (i.e., the risk assessment rule in the aforementioned embodiment) used to analyze platform user behavior data for risk. This rule configuration interface may include input boxes for rule type, rule name, rule logic expression, parameter type, and risk level. The rule type input box is used to configure the rule type, such as the transaction type used for risk assessment of transaction behavior data. The rule logic expression input box is used to configure the calculation logic corresponding to the rule, expressed through an expression. The parameter type input box is used to configure the type of rule configuration parameter. This type refers to the logical classification of the rule and can be considered a strategy type. These types include matching and statistical types. Matching refers to a similarity-based type that emphasizes precise matching. Its core principle is to directly perform pattern matching or judgment on data using predefined rules or keyword sets, without complex analytical calculations. Statistical refers to a probabilistic analysis type that relies on predefined data distribution patterns for judgment. The risk level input box is used to configure the risk level that exists when the corresponding rule is met.
[0065] It should also be understood that the configuration information in S301 refers to the data set input in the rule configuration interface, which includes rule basic information and dynamic parameters, usually in the form of key-value pairs. Among them, the rule basic information corresponds to the content to be configured indicated by the input box, and the dynamic parameters correspond to the content in the input box. In some embodiments, the first rule basic information is the rule name, and the corresponding dynamic parameter is the consumption expenditure rule. The second rule basic information is the total consumption expenditure, and the corresponding dynamic parameter is 50 million. The third rule basic information is the consumption expenditure duration, and the corresponding dynamic parameter is 10 minutes. The fourth rule basic information is the parameter type, and the corresponding dynamic parameter is the matching class. The fifth rule basic information is the risk level, and the corresponding dynamic parameter is high risk. That is, in consumption expenditure, a behavior with a total consumption expenditure of 50 million within 10 minutes is a high-risk behavior.
[0066] It should also be understood that this configuration information includes rule configuration parameters, which refer only to the logical parameters of the risk assessment rule, i.e., some dynamic parameters. For the aforementioned consumer spending rule, the logical parameters might be 50 million and 10 minutes. These rule configuration parameters are the specific embodiment of the risk strategy indicators, quantifying the rule conditions during risk assessment through dynamic parameters.
[0067] It should also be understood that the target script file in S302 is a generic script template file that matches the type of the rule configuration parameter. Optionally, the script template file is a Groovy script file. A Groovy script file is a text file with a .groovy extension and contains executable code written in the Groovy scripting language, a dynamic programming language based on the JVM (Java Virtual Machine). The Groovy scripting language is commonly used for dynamic logic embedding, such as in rule engines.
[0068] It should also be understood that the format of the target script file in S302 refers to the organizational form of the file content of the target script file. The organizational form of the file content involves how to lay out the calculation logic of the code module to write the risk assessment rules and involves which other libraries or tool classes to reference. Optionally, the tool class RiskUtils is referenced, and its RiskUtils class encapsulates the public methods related to risk assessment. Accordingly, it is understandable that the calculation logic is implemented by the Groovy scripting language. Moreover, based on the above-mentioned target script file, the target rule file can be generated by the rule configuration parameters, indicating that the target script file is a dynamically configurable script template file.
[0069] In some embodiments, S302 generates a target rule file from the rule configuration parameters based on the format of the target script file, including: replacing the corresponding parts in the target script file with the rule configuration parameters and corresponding rule basic information based on the format of the target script file to generate a target rule file.
[0070] It should be understood that the operation logic of the target rule file is implemented through a dynamic scripting language.
[0071] In one possible implementation, selecting a target script file according to the type of the rule configuration parameter in S302 includes: comparing the type with multiple preset types, determining a target type that matches the type from the multiple preset types, each of the multiple preset types corresponding to a general script template file; and determining the script template file corresponding to the target type as the target script file.
[0072] It should be understood that each preset type corresponds to a universal script template file, and the formats of different script template files are different. The script template files corresponding to the statistical type of rules generally include logic such as accumulation, summation, grouping, and deduplication. The script template files corresponding to the matching type of rules generally include logic such as equality, similarity, and approximately equality.
[0073] In the above technical solution, preset types are strongly associated with script template files. This ensures consistency in the calculation logic (i.e., rule calculation logic) for the same preset type. For example, all matching rules automatically inherit the common logic for exact field matching, reducing the writing and maintenance costs of duplicate code. Furthermore, the above solution automatically identifies types and matches them with script template files, eliminating the need for developers to manually write or select the organization of script files. This reduces the risk of manual selection errors or coding errors, improving configuration efficiency and accuracy. Furthermore, the parameterized design of the universal script template files not only ensures a standardized execution process but also supports adaptation to diverse business scenarios. For example, the script template files corresponding to statistical rules can configure time windows and aggregation algorithms, while the script template files corresponding to matching rules can flexibly adjust the threshold for similarity determination. Therefore, the above technical solution balances rule standardization and business specificity within a unified script template framework, achieving both flexibility and reusability.
[0074] S303: Update the target rule file by broadcasting and save it to the rule file library.
[0075] That is, the specific steps executed by the rule dynamic perception module 202 correspond to the above-mentioned step 303.
[0076] It should also be understood that the broadcast method in S303 allows multiple data receiving nodes to obtain the target rule file for the same data sending node. Its core principle is that there is no designated data receiving node; all data receiving nodes connected to the same broadcast domain can obtain the target rule file. This broadcast method is understandably suitable for scenarios requiring simultaneous notification of multiple data receiving nodes and for handling emergency services (e.g., risk control). The task manager is a data receiving node, and the rule configuration center is a data sending node.
[0077] It should also be understood that Groovy script files can be directly compiled into JVM bytecode at runtime, without the need for pre-compilation into .class files. After generating the target rule file (i.e., by replacing the existing rule instance in the target script file to obtain the target rule instance), the risk management system updates the target rule file via broadcast and saves it to the rule file library. In other words, the risk management system can instantly load the new calculation logic and rule configuration parameters in the target rule instance. This ability to directly and dynamically compile and load the target rule file makes it a "pluggable" logical unit, eliminating the need to stop the risk management system from running. This prevents the suspension or interruption of risk control operations being processed by the risk management system.
[0078] It should also be noted here that replacing the existing rule instances in the target script file is a way of replacing the existing rule instances by externally injecting rule configuration parameters. This decouples the script logic of the target script file from the above-mentioned rule configuration parameters, making the generation process of the target rule file more flexible.
[0079] Optionally, before S303, the method 300 also includes: storing the target rule file in an external storage module; encapsulating the attribute information of the target rule file as a broadcast variable, the attribute information including the file storage path, the corresponding version number and the verification code of the target rule file; sending the broadcast variable to multiple data receiving nodes via broadcast; and, S303, including: monitoring the broadcast channel through each of the multiple data receiving nodes, receiving the broadcast variable, and parsing the broadcast variable to obtain the attribute information of the target rule file; downloading the target rule file from the external storage module and verifying the target rule file based on the attribute information, and when the target rule file is complete, updating and saving the target rule file to the rule file library.
[0080] It should be understood that the version number is used to identify different versions of the rule file under the same rule name. The check code is used to verify the integrity of the target rule file.
[0081] It should also be understood that the broadcast channel refers to the broadcast channel between the data sending node and the multiple data receiving nodes. The data sending node and the multiple data receiving nodes in the same local area network belong to the same broadcast domain.
[0082] Specifically, if Figure 4 FIG2 shows an architectural design diagram of a dynamic rule perception module provided in an embodiment of the present application. Exemplarily, the dynamic rule perception module 202 is configured to perceive the target rule file through the Flink CDC (Change Data Capture) component. Upon perceiving the target rule file, the Flink CDC component updates and saves the target rule file to the rule file repository of each task manager within the Flink CDC component.
[0083] Optionally, the target rule file is updated and saved to the rule file library in a key-value pair manner, where the key is an identifier of the target rule file and the value is the target rule file.
[0084] Optionally, the identifier is a behavior type corresponding to the target rule file.
[0085] S304: Obtain behavior data of multiple users, and analyze the behavior data based on multiple rule files in the rule file library to determine user behavior data that poses a risk.
[0086] That is, the specific steps executed by the data source management center 203 correspond to the steps of obtaining the behavior data of multiple users in the above step 304 .
[0087] That is, the specific steps executed by the risk analysis and calculation module 204 correspond to the above-mentioned step 304 .
[0088] It should be understood that the behavior data in S304 is used to describe the behavior of multiple users.
[0089] In some embodiments, behavioral data may include user device access data, user transaction data, user information sharing data, and the like. Optionally, user device access data includes records of users opening or closing door locks within a first time period, the number of incorrect password entries within a second time period, and the number of times the user increased the air conditioning temperature within a third time period. User information sharing data includes records of users sharing device control permissions with a first account and users deleting first information using a temporary guest account.
[0090] In one possible implementation, S304 analyzes the behavior data based on multiple rule files in the rule file library to determine the user behavior data that is at risk, including: determining multiple candidate files that match the behavior data from the multiple rule files; analyzing the behavior data through an operation state machine corresponding to each candidate file to determine the user behavior data that is at risk, the operation state machine encapsulating the calculation logic and rule configuration parameters of the candidate file.
[0091] It should be understood that the aforementioned computational state machine refers to a data processing model that encapsulates the computational logic and rule configuration parameters of a rule file. The core function of this computational state machine is to dynamically apply the rule file based on input data (in this application, the behavioral data of multiple users) and output risk determination results (in this application, the risky user behavior data). The rule configuration parameters are injected into the computational state machine as key-value pairs for dynamic reference by the computational logic. The key refers to the underlying rule information corresponding to the rule configuration parameter, and the value refers to the rule configuration parameter.
[0092] It should also be understood that the aforementioned multiple candidate files may or may not include the target rule file. Each candidate file corresponds to an operation state machine. Thus, each operation state machine can encapsulate the computational logic and rule configuration parameters of the corresponding rule file. In other words, each rule file has a corresponding operation state machine, and each operation state machine encapsulates its own computational logic and is responsible for maintaining its own rule configuration parameters. Therefore, each operation state machine can be considered a highly cohesive operation state machine.
[0093] It should be noted here that the above encapsulation refers to integrating the calculation logic and rule configuration parameters of the rule file into an independent, reusable execution unit, while hiding the internal implementation details from the outside and only exposing standardized input and output interfaces. The core of encapsulation is to achieve decoupling of calculation logic and rule configuration parameters.
[0094] It should also be noted here that for the newly added target rule files in the rule file library, the calculation logic and rule configuration parameters of the target rule files can be encapsulated to obtain the operation state machine corresponding to the target rule files, and the calculation logic corresponding to the operation state machine is dynamic. In this way, for evaluating the behavior data of any first behavior type, the first rule file corresponding to the first behavior type can be added, and the calculation logic and rule configuration parameters of the first rule file can be encapsulated to obtain the operation state machine corresponding to the first rule file, and then the behavior data of the first behavior type can be analyzed through the operation state machine corresponding to the first rule file. Therefore, in the process of each operation state machine in the multiple operation state machines processing the behavior data of the corresponding behavior type through the corresponding rule file, the multiple operation state machines do not affect each other. Therefore, the multiple operation state machines are low-coupled.
[0095] In the above technical solution, a dynamic rule screening mechanism is used to match multiple candidate files from the rule file library. This can avoid traversing the entire rule file and significantly reduce computational redundancy. Furthermore, each operation state machine encapsulates the calculation logic and rule configuration parameters of the corresponding rule file to ensure that the rules do not interfere with each other. At the same time, modifying a single rule file only requires replacing the corresponding operation state machine, without the need for global shutdown or reload, thus ensuring the processing continuity of risk control business. Through the layered processing mechanism of rule files and operation state machines, the timeliness of risk assessment and the maintainability of the risk management system can be significantly improved.
[0096] In one possible implementation, multiple candidate files that match the behavior data are determined from the multiple rule files, including: determining multiple behavior types corresponding to the behavior data; comparing each behavior type with the behavior type corresponding to each rule file in the multiple rule files, and determining candidate behavior types that correspond one-to-one to the multiple behavior types from the behavior types corresponding to the multiple rule files; and determining the rule files corresponding to the multiple candidate behavior types as the multiple candidate files.
[0097] It should be understood that the behavior type corresponding to the above rule file refers to the type of user behavior targeted by the rule file, including the behavior type of accessing the device, transaction type, and information sharing type.
[0098] In the above technical solution, based on the multiple behavior types corresponding to the behavioral data of multiple users, candidate behavior types corresponding to these multiple behavior types are screened from the behavior types corresponding to multiple rule files to select those candidate behavior types that correspond one-to-one with these multiple behavior types. This pre-classified screening method can quickly narrow the scope of rule matching and significantly reduce computational redundancy. Furthermore, the strict one-to-one matching logic ensures that multiple candidate rule files are highly correlated with the behavioral data of multiple users (for example, only consumption expenditure rules are included in the analysis of transaction behavior), reducing false matches and improving the rule hit rate. Therefore, this technical solution can significantly optimize the efficiency and accuracy of rule screening through a precise matching mechanism between behavior types and rule files.
[0099] Optionally, determining multiple behavior types corresponding to the behavior data includes: extracting behavior features of the behavior data; comparing the behavior features with sample behavior features of the behavior types corresponding to the multiple rule files, and determining candidate behavior features that match the behavior features from the multiple sample behavior features; and determining the behavior type corresponding to the candidate behavior features as the multiple behavior types.
[0100] Optionally, the method for determining the behavior data of multiple users in S304 includes: based on multiple behavior types corresponding to the original behavior data of multiple users, dividing the original behavior data into multiple message queue data in message queue format, and determining the multiple message queue data as the behavior data, each message queue data corresponds to a behavior type, and the original behavior data is unprocessed behavior data.
[0101] Optionally, each message queue data includes the occurrence time of the behavior, the behavior type and the behavior content.
[0102] It should be understood that the above process of dividing the original behavior data of multiple users into multiple message queue data in the message queue format can be regarded as a process of preprocessing the original behavior data.
[0103] Figure 5 This is a schematic diagram of an embodiment of the present application providing a method of dividing original behavior data into multiple message queue data in a message queue format.
[0104] For example, Figure 5 As shown, the raw behavior data of multiple users corresponds to multiple behavior types. According to the behavior types, the raw behavior data is divided into multiple message queue data in a message queue format. Specifically, the multiple message queue data include behavior data for the first type of behavior, behavior data for the second type of behavior, and so on. The behavior data for each type of behavior includes the behavior type, the time of occurrence of the behavior, and the specific behavior content.
[0105] In one possible implementation, the behavior data is analyzed through the operation state machine corresponding to each candidate file to determine the user behavior data that is at risk, including: analyzing the behavior data through the operation state machine corresponding to each candidate file, generating an intermediate result, and storing the intermediate result in an external storage module, which is an independently deployed remote dictionary server cluster; and determining the user behavior data that is at risk based on the intermediate result through the operation state machine corresponding to each candidate file.
[0106] It should be understood that the remote dictionary server cluster mentioned above is a Redis cluster, which includes multiple Redis servers, and the Redis server is a high-performance memory-based database.
[0107] In the above technical solution, the intermediate results of the operation state machine corresponding to each candidate file in the multiple candidate files are stored in an external storage module. This can reduce the memory pressure of the electronic device and the consumption of the electronic device's storage resources when the intermediate results are persisted in the electronic device's memory. Replacing the full aggregation based on the electronic device's memory with incremental aggregation based on the Redis server can also significantly reduce the consumption of the electronic device's storage resources when the program is running.
[0108] For example, for the account login behavior of the target application of multiple users, the intermediate result is the number of consecutive login failures of each user within 5 minutes. If the number is greater than or equal to 3, it is determined that the user's login behavior is risky.
[0109] In one possible implementation, before updating and saving the target rule file to the rule file library by broadcasting, the method 300 also includes: publishing the target rule file to the database middleware so that the database middleware routes the target rule file to a second external storage module, and the second external storage module supports distributed management.
[0110] In the above technical solution, before the target rule file is updated and saved to the rule file library by broadcasting, the target rule file is also published to the database middleware to back up the target rule file to the second external storage module. Among them, the database middleware acts as a temporary buffer zone, which can temporarily store the target rule file before broadcasting, thereby preventing the target rule file from being lost due to direct overwriting of the rule file library due to transmission errors or incomplete reception. The second external storage module acts as an independent backup, which is physically isolated from the rule file library to prevent the target rule file from being completely lost due to hardware failure of the rule file library. The above solution can restore the original target rule file when the broadcast fails through double protection, avoiding the loss of the target rule file and entering an uncontrollable state.
[0111] Another example, Figure 6 As shown, the data processing system 200 also includes TDDL database middleware, a FlinkCDC component, a Source Connector component, a first external storage module and a result analysis module. The rule dynamic perception module 202 includes the Flink CDC component, the risk analysis and calculation module 204 is connected to the data source management center 203 through the Source Connector component, the first external storage module is connected to the risk analysis and calculation module 204, and the SinkConnector component is connected to the first external storage module through the window trigger.
[0112] Among them, the target rule file is published to the TDDL database middleware so that the TDDL database middleware routes the target rule file to the second external storage module; the Flink CDC component is used to perceive the target rule file; the Source Connector component is used to extract the behavior data of multiple users from the data source management center 203 and publish it to the risk analysis and calculation module 204 in the form of an event stream; the first external storage module is used to store the intermediate results obtained after multiple operation state machines analyze the behavior data of multiple users; the result analysis module is used to call the final result output by the risk analysis and calculation module 204 and analyze it.
[0113] Optionally, the data processing system 200 also includes a window trigger and a Sink Connector component. The window trigger is used to read the intermediate results under a certain trigger condition and further calculate the intermediate results to obtain the calculation results; the Sink Connector component stores the calculation results to other external storage systems.
[0114] Furthermore, the rule dynamic perception module 202 and the risk analysis and calculation module 204 can be designed based on a distributed stream processing framework. Optionally, the distributed stream processing framework is the Apache Flink framework. However, the embodiment of the present application does not use the source code (native Flink code) of the Apache Flink framework, but instead generates the target rule file from the target script file.
[0115] Specifically, if Figure 7The figure shows an architectural design diagram of a risk analysis and calculation module provided by an embodiment of the present application. Exemplarily, the risk analysis and calculation module 204 is specifically used to: determine multiple behavior types corresponding to the behavior data of multiple users; compare each behavior type with the behavior type corresponding to each rule file in multiple rule files, and determine candidate behavior types corresponding to multiple behavior types from the behavior types corresponding to the multiple rule files; determine the rule files corresponding to the multiple candidate behavior types as multiple candidate files; analyze the behavior data of multiple users through the operation state machine corresponding to each candidate file, and determine the user behavior data that is at risk. Among them, the intermediate results obtained by the operation state machine from analyzing the behavior data of multiple users are stored in the first external storage module.
[0116] It should be understood that after analyzing the behavioral data of multiple users, multiple computing state machines may need to merge the risk results output by the multiple computing state machines or update shared resources. For example, multiple computing state machines may need to update the cumulative risk score of the same user, or calculate the global risk index. Therefore, the embodiment of the present application introduces concurrent control of multiple computing state machines. Figure 8 800 is a schematic flow chart of another method for processing data provided by an embodiment of the present application. Specifically, the method 800 includes:
[0117] S801: Obtain behavior data of multiple users, and determine multiple candidate files matching the behavior data from multiple rule files in a rule file library.
[0118] S802: When a first computing state machine corresponding to a first candidate file among the multiple candidate files successfully preempts the distributed lock and the behavior data is analyzed by the first computing state machine, the first computing state machine collects statistics on risky user behavior data after a first moment to obtain a first statistical result, where the first moment is the moment when the first computing state machine successfully preempts the distributed lock.
[0119] It should be understood that in the process of analyzing the behavior data of the multiple users by the first operation state machine, risky user behavior data will be determined from the behavior data of the multiple users.
[0120] It should also be understood that the first operation state machine in S802 corresponds to the first candidate file, and each rule file corresponds to a behavior type. Therefore, the first operation state machine can be considered an operation state machine for analyzing the behavior data of the first type of behavior in the behavior data and for collecting statistics on the analyzed risky user behavior data. The behavior type corresponding to the first candidate file is the first type of behavior.
[0121] It should also be understood that the distributed lock in S802 is used by the computing state machine to preempt shared resources. If the first computing state machine successfully preempts the distributed lock, the other computing state machines cannot merge the risk results or update the shared resources. In other words, this application uses the forced serialization of the distributed lock preemption mechanism to ensure that only one computing state machine can modify the merged risk results or update the shared resources at the same time. This ensures mutual exclusivity and result consistency in concurrent operations.
[0122] S803: When the second moment for counting risky user behavior data is the third moment, outputting the first statistical result through the first operation state machine, and the second moment is later than the first moment.
[0123] It should be noted that the above scheme describes a scenario where, when the first computing state machine successfully preempts the distributed lock, the first computing state machine analyzes the behavioral data for the first category of behaviors in the behavioral data to obtain risky first category user behavior data. During the analysis of the behavioral data for the first category of behaviors by the first computing state machine, the first computing state machine collects statistics on the risky first user behavior data after the first moment, obtains and outputs a first statistical result, and the statistical end time is the aforementioned third moment.
[0124] In addition, it should be noted that the above embodiment describes that between the first moment and the second moment, the first operation state machine can completely analyze the behavior data of multiple users.
[0125] In the above technical solution, a distributed lock preemption mechanism is introduced in the process of analyzing the behavioral data of multiple users by multiple operation state machines. This can ensure that the behavioral data of the same behavior type is only processed by a single operation state machine, avoid data competition caused by multi-node concurrent control, and ensure the consistency of the results of the statistical process. After the first operation state machine successfully preempts the distributed lock, it will combine the analysis of the behavioral data according to the behavior type and the output of the statistical results based on the dynamic time window. This can also avoid the pressure on resources caused by frequent result output. In other words, the above solution can optimize resource utilization while ensuring the accuracy of the results through the collaborative design of the distributed lock preemption mechanism and the time window transmission mechanism, providing flexible and controllable technical support for real-time risk monitoring in high-concurrency scenarios, and effectively improving the efficiency and reliability of user behavior data analysis in a distributed environment.
[0126] For example, when the first computing state machine successfully seizes the distributed lock, the first computing state machine analyzes the behavioral data of the first category of behavior in the behavioral data to obtain the number of users with risky first category behaviors. During the analysis of the behavioral data of the first category of behavior by the first computing state machine, the first computing state machine accumulates the number of users with risky first category behaviors after the first moment, obtains the accumulated number of users, and outputs it, until the third moment.
[0127] In one possible implementation, the method for determining the third moment includes any one of the following: when the distributed lock is set with a first validity period, the third moment is determined based on the first moment and the first validity period; when the first preset duration is set by a preset timer, the third moment is determined based on the first moment and the first preset duration.
[0128] It should be understood that the third moment is a moment after the first moment and after the first validity period or a moment after the first preset time length preset by the preset timer.
[0129] It should also be understood that the time period between the first moment and the third moment is the time period for outputting the first statistical result. In this regard, the above solution can be regarded as a process of revealing the window result based on the behavior time.
[0130] Optionally, based on the first moment and the first validity period, determining the third moment includes: determining the moment corresponding to the first validity period after the first moment as the third moment; and, based on the first moment and the first preset time length, determining the third moment includes: determining the moment corresponding to the first preset time length after the first moment as the third moment.
[0131] It should be noted that the first operation state machine simultaneously analyzes the behavior data of the first category of behavior in the behavior data, and collects statistics on the user behavior data of the first category of risky behavior. In some cases, the processing speed of analyzing the behavior data of the first category of behavior is slower than the processing speed of collecting statistics on the user behavior data of the first category of risky behavior. That is, there is a situation where the first preset duration set by the preset timer has arrived, but the analysis of the behavior data of the first category of behavior in the behavior data has not been completed. If the first statistical result is directly output at this time, it will cause the loss of some of the risky user behavior data obtained at the same time. Therefore, the following embodiment simultaneously introduces the second validity period of the cyclic timer and the distributed lock to solve the above problem, please refer to Figure 9 A schematic flow chart of another method for processing data is provided. Specifically, the method 900 includes:
[0132] S901 , obtaining behavior data of multiple users, and determining multiple candidate files matching the behavior data from multiple rule files in a rule file library.
[0133] S902, when the first computing state machine corresponding to the first candidate file among the multiple candidate files successfully seizes the distributed lock and the behavior data is analyzed by the first computing state machine, the user behavior data that is at risk after the first moment is counted by the first computing state machine to obtain a first statistical result, and the first moment is the moment when the first computing state machine successfully seizes the distributed lock.
[0134] S903: Determine whether the time between the first moment and the second moment for collecting statistics on risky user behavior data is equal to a second preset time set by a cyclic timer.
[0135] S904, when the duration is equal to the second preset duration and the duration is not equal to the second validity period of the distributed lock, continue to analyze the behavior data through the first computing state machine until the duration is equal to the second validity period, and output the first statistical result through the first computing state machine, and the second validity period is greater than the second preset duration.
[0136] It should be understood that the above solution is to complete the analysis of the behavior data of the first type of behavior by cyclically waiting within the second validity period for the first operation state machine to analyze the behavior data of the first type of behavior within the second preset time length.
[0137] In addition, it should be noted that the above embodiment describes that the first operation state machine has not finished analyzing the behavior data of multiple users within the first second preset time period.
[0138] In the above technical solution, a recurring timer periodically triggers the computational state machine to analyze the behavioral data of the first type of behavior at preset intervals, ensuring that the data analysis process is continuously attempted within the second validity period (the validity period of the distributed lock). Furthermore, the second validity period of the distributed lock provides a clear execution window for the analysis process, forcing processing progress through time constraints while ensuring task exclusivity. This ensures the integrity of risky user behavior data and automatically releases the distributed lock in the event of a timeout to trigger fault tolerance, effectively mitigating data delays or processing speed discrepancies and enhancing the risk management system's ability to withstand backpressure in high-concurrency scenarios.
[0139] Figure 10 1000 is a schematic diagram of a data processing device according to an embodiment of the present invention. The device 1000 includes an acquisition unit 1001, a generation unit 1002, a storage unit 1003, and a processing unit 1004.
[0140] The acquisition unit 1001 is used to acquire the configuration information input by the developer in the rule configuration interface and parse the configuration information to obtain relevant rule configuration parameters;
[0141] A generating unit 1002 is configured to select a target script file according to the type of the rule configuration parameter, and generate a target rule file from the rule configuration parameter based on the format of the target script file;
[0142] The storage unit 1003 is used to update the target rule file and save it to the rule file library by broadcasting;
[0143] The processing unit 1004 is configured to obtain behavior data of multiple users, and analyze the behavior data based on multiple rule files in the rule file library to determine risky user behavior data.
[0144] In one possible implementation, the generation unit 1002 is specifically used to: compare the type with multiple preset types, determine a target type that matches the type from the multiple preset types, each of the multiple preset types corresponds to a general script template file; and determine the script template file corresponding to the target type as the target script file.
[0145] In one possible implementation, the processing unit 1004 is specifically used to: determine multiple candidate files that match the behavior data from the multiple rule files; analyze the behavior data through the operation state machine corresponding to each candidate file to determine the user behavior data that is at risk, and the operation state machine encapsulates the calculation logic and rule configuration parameters of the candidate file.
[0146] In one possible implementation, the device 1000 also includes: a determination unit, used to: determine multiple behavior types corresponding to the behavior data; compare each behavior type with the behavior type corresponding to each rule file in the multiple rule files, and determine candidate behavior types that correspond one-to-one to the multiple behavior types from the behavior types corresponding to the multiple rule files; and determine the rule files corresponding to the multiple candidate behavior types as the multiple candidate files.
[0147] In one possible implementation, the processing unit 1004 is further specifically used to: analyze the behavior data through the operation state machine corresponding to each candidate file, generate an intermediate result, and store the intermediate result in an external storage module, which is an independently deployed remote dictionary server cluster; determine the user behavior data that is risky based on the intermediate result through the operation state machine corresponding to each candidate file.
[0148] In a possible implementation, the storage unit 1003 is further configured to store the rule file library in the external storage module.
[0149] In one possible implementation, the processing unit 1004 is also used to, when the first operating state machine corresponding to the first candidate file among the multiple candidate files successfully preempts the distributed lock and the behavior data is analyzed by the first operating state machine, count the user behavior data at risk after the first moment through the first operating state machine to obtain a first statistical result, and the first moment is the moment when the first operating state machine successfully preempts the distributed lock; the device 1000 also includes: an output unit, used to output the first statistical result through the first operating state machine when the second moment used for counting the user behavior data at risk is a third moment, and the second moment is later than the first moment.
[0150] In one possible implementation, the determination unit is further used to: when the distributed lock is set with a first validity period, determine the third moment based on the first moment and the first validity period; when the first preset duration is set by a preset timer, determine the third moment based on the first moment and the first preset duration.
[0151] In one possible implementation, the determination unit is also used to determine whether the duration between the first moment and the second moment is equal to a second preset duration set by a cyclic timer; the storage unit 1003 is also used to continue analyzing the behavior data through the first operating state machine when the duration is equal to the second preset duration and the duration is not equal to the second validity period of the distributed lock, and output the first statistical result through the first operating state machine, and the second validity period is greater than the second preset duration.
[0152] Figure 11 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Optionally, the electronic device 1100 may be a device with computing functions, or a server, which is not limited in the embodiment of the present application.
[0153] For example, Figure 11 As shown, the electronic device 1100 includes: a memory 1101 and a processor 1102, wherein the memory 1101 stores an executable program code 1103, and the processor 1102 is used to call and execute the executable program code 1103 to perform a method for processing data.
[0154] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a method for processing data provided in an embodiment of the present application.
[0155] In this embodiment, the device can be divided into functional modules based on the above-described method example. For example, each functional module can be mapped to a specific functional module, or two or more functions can be integrated into a single processing unit. The integrated modules can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used.
[0156] In the case of dividing the functional modules into corresponding functional modules, the device may further include an acquisition unit, a generation unit, a storage unit, a processing unit, a determination unit, and an output unit. It should be noted that all relevant contents involved in the above method embodiments can be referred to the functional description of the corresponding functional modules and will not be repeated here.
[0157] It should be understood that the device provided in this embodiment is used to execute the above-mentioned method for processing data, and thus can achieve the same effect as the above-mentioned implementation method.
[0158] In the case of an integrated unit, the device may include a processing unit and a storage module. When the device is applied to an electronic device, the processing unit may be used to control and manage the operation of the electronic device. The storage module may be used to support the electronic device in executing relevant executable program code.
[0159] The processing unit may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure herein. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processing (DSP) and a microprocessor, and the storage module may be a memory.
[0160] In addition, the device provided in the embodiments of the present application can specifically be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a method for processing data provided in the above embodiment.
[0161] This embodiment also provides a computer-readable storage medium, which stores executable program code. When the executable program code runs on a computer, the computer executes the above-mentioned related method steps to implement a method for processing data provided by the above embodiment.
[0162] This embodiment further provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement a method for processing data provided by the above embodiment.
[0163] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0164] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0165] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0166] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for processing data, characterized in that The method comprises: Obtaining configuration information entered by the developer in the rule configuration interface, parsing the configuration information to obtain relevant rule configuration parameters, which are logical parameters of the risk assessment rules; Selecting a target script file according to the type of the rule configuration parameter, and generating a target rule file from the rule configuration parameter based on a format of the target script file, wherein the format is an organization form of the file content of the target script file; Updating the target rule file and saving it to the rule file library by broadcasting; Obtaining behavior data of multiple users, and analyzing the behavior data based on multiple rule files in the rule file library to determine risky user behavior data; The analyzing the behavior data based on the multiple rule files in the rule file library to determine the risky user behavior data includes: Determining, from the plurality of rule files, a plurality of candidate files that match the plurality of behavior types corresponding to the behavior data; Analyzing the behavior data through a computing state machine corresponding to each candidate file to determine risky user behavior data, wherein the computing state machine encapsulates the calculation logic and rule configuration parameters of the candidate file and is used to dynamically apply the candidate file based on the behavior data; The step of selecting a target script file according to the type of the rule configuration parameter includes: Comparing the type with multiple preset types, and determining a target type that matches the type from the multiple preset types, each of the multiple preset types corresponds to a universal script template file, the preset types including a statistical class and a matching class, the script template file corresponding to the statistical class rules including accumulation, summation, grouping, and deduplication logic, and the script template file corresponding to the matching class rules including equality, similarity, and approximately equality logic; The script template file corresponding to the target type is determined as the target script file.
2. The method according to claim 1, characterized in that The determining, from the multiple rule files, multiple candidate files that match the multiple behavior types corresponding to the behavior data includes: Comparing each behavior type with the behavior type corresponding to each rule file in the multiple rule files, and determining candidate behavior types corresponding to the multiple behavior types from the behavior types corresponding to the multiple rule files; The rule files corresponding to the multiple candidate behavior types are determined as the multiple candidate files.
3. The method according to claim 1, characterized in that The analysis of the behavior data by the operation state machine corresponding to each candidate file to determine the risky user behavior data includes: Analyzing the behavior data through a computing state machine corresponding to each candidate file to generate an intermediate result, and storing the intermediate result in a first external storage module, where the first external storage module is an independently deployed remote dictionary server cluster; The risky user behavior data is determined based on the intermediate results through the operation state machine corresponding to each candidate file.
4. The method according to claim 1, wherein Before updating and saving the target rule file to the rule file library by broadcasting, the method further includes: The target rule file is published to a database middleware, so that the database middleware routes the target rule file to a second external storage module, where the second external storage module supports distributed management.
5. The method according to claim 1, wherein The method further comprises: When a first computing state machine corresponding to a first candidate file among the multiple candidate files successfully preempts the distributed lock and the behavior data is analyzed by the first computing state machine, the first computing state machine collects statistics on risky user behavior data after a first moment to obtain a first statistical result, where the first moment is the moment when the first computing state machine successfully preempts the distributed lock; In a case where the second moment for counting risky user behavior data is the third moment, the first statistical result is output through the first operation state machine, and the second moment is later than the first moment.
6. The method according to claim 5, characterized in that The method for determining the third moment includes any one of the following: In a case where the distributed lock is set with a first validity period, determining the third time based on the first time and the first validity period; In a case where the first preset duration is set by a preset timer, the third time is determined based on the first time and the first preset duration.
7. The method according to claim 5, characterized in that The method further comprises: determining whether a duration between the first moment and the second moment is equal to a second preset duration set by a cyclic timer; When the duration is equal to the second preset duration and the duration is not equal to the second validity period of the distributed lock, continue to analyze the behavioral data through the first operating state machine until the duration is equal to the second validity period, and output the first statistical result through the first operating state machine, and the second validity period is greater than the second preset duration.
8. A device for processing data, characterized in that: The device comprises: An acquisition unit is used to acquire configuration information input by a developer in a rule configuration interface, and parse the configuration information to obtain relevant rule configuration parameters, where the rule configuration parameters are logical parameters of the risk assessment rule; a generating unit, configured to select a target script file according to a type of the rule configuration parameter, and generate a target rule file from the rule configuration parameter based on a format of the target script file, wherein the format is an organization form of the file content of the target script file; A storage unit, configured to update the target rule file and save it to a rule file library by broadcasting; a processing unit, configured to obtain behavior data of a plurality of users, and analyze the behavior data based on a plurality of rule files in the rule file library to determine risky user behavior data; The processing unit is specifically configured to: Determining, from the plurality of rule files, a plurality of candidate files that match the plurality of behavior types corresponding to the behavior data; Analyzing the behavior data through a computing state machine corresponding to each candidate file to determine risky user behavior data, wherein the computing state machine encapsulates the calculation logic and rule configuration parameters of the candidate file and is used to dynamically apply the candidate file based on the behavior data; The generating unit is specifically configured to: Comparing the type with multiple preset types, and determining a target type that matches the type from the multiple preset types, each of the multiple preset types corresponds to a universal script template file, the preset types including a statistical class and a matching class, the script template file corresponding to the statistical class rules including accumulation, summation, grouping, and deduplication logic, and the script template file corresponding to the matching class rules including equality, similarity, and approximately equality logic; The script template file corresponding to the target type is determined as the target script file.
9. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable program code; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Terminal anomaly analysis method and device based on process, equipment and storage medium
CN112114995A
Business rule processing method and device, server and storage medium
CN114490694A
Script file generation method and device, computer equipment and storage medium
CN115756581A