Data processing method, device, equipment and storage medium

By collecting and processing transaction information in real time, combining the set indicator values ​​of the current and comparison periods to determine the risk coefficient, the problem of insufficient accuracy in existing technologies is solved and more efficient risk warning is achieved.

CN115796878BActive Publication Date: 2025-10-03WEBANK (CHINA)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211377650.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-10-03
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in providing early warning of financial transaction risks and fail to effectively consider the impact of time and spatial regions.

Method used

By collecting transaction information in real time and processing address information, the set indicator values ​​for the current cycle and comparison cycle are obtained, the risk coefficient is determined, and a risk warning is output when the risk coefficient is greater than or equal to the threshold.

Benefits of technology

The accuracy of risk warning has been improved, and more accurate risk assessment and warning have been achieved by considering the impact of time and space areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796878B_ABST
    Figure CN115796878B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, apparatus, device and storage medium, including: real-time collection of transaction information of a first business during a transaction; processing the transaction information of the first business based on address information to obtain the value of the set index of each of at least two areas of the first business in the current cycle; querying the value of the set index of each of at least two areas of the first business in at least one comparison cycle; determining the risk coefficient of the first business based on the value of the set index of each of the areas of the first business in the current cycle and the value of the set index of each of the areas of the first business in at least one comparison cycle; outputting a risk warning when the risk coefficient is greater than or equal to a risk threshold. In this way, when performing a risk warning, the influence of the time and spatial area risk coefficients are simultaneously considered, and the accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and is related to but not limited to data processing methods, devices, equipment and storage media. Background Art

[0002] With the rapid development of computer technology, more and more technologies are being applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher requirements are also placed on technology.

[0003] For real-time early warning of financial transaction risks, relevant technical solutions include: comparing the indicators of the current cycle with the indicators of the past cycles, obtaining the difference coefficient between the indicators of the current cycle and the indicators of the past cycles as the risk coefficient, and issuing a risk warning when the risk coefficient is greater than or equal to the set risk threshold.

[0004] However, the accuracy of relevant technologies in risk warning is low. Summary of the Invention

[0005] The present application provides a data processing method, apparatus, device, and storage medium, which simultaneously consider the influence of time and spatial regional risk factors when conducting risk warning, and have high accuracy.

[0006] The technical solution of this application is achieved as follows:

[0007] The present application provides a data processing method, the method comprising:

[0008] collecting transaction information of the first business in a transaction process in real time; the transaction information at least includes address information representing the area where the transaction occurs;

[0009] Based on the address information, the transaction information of the first business is processed to obtain a value of a set indicator of the first business in each of the at least two areas in a current period;

[0010] querying a value of a set indicator of the first service in each of at least two areas in at least one comparison period;

[0011] Determining a risk coefficient of the first business based on a value of a set indicator of the first business in each of the areas in a current period and a value of a set indicator of the first business in each of the areas in at least one comparison period;

[0012] When the risk coefficient is greater than or equal to the risk threshold, a risk warning is output.

[0013] The present application provides a data processing device, comprising:

[0014] a collection unit, configured to collect transaction information of the first business in real time during the transaction process; the transaction information at least including address information representing the area where the transaction occurs;

[0015] a processing unit, configured to process the transaction information of the first business based on the address information, and obtain a value of a set indicator of the first business in each of the at least two areas in a current period;

[0016] a query unit, configured to query a value of a set indicator of the first service in each of at least two areas in at least one comparison period;

[0017] a determining unit, configured to determine a risk coefficient of the first business based on a value of a set indicator of the first business in each of the areas in a current period and a value of a set indicator of the first business in each of the areas in at least one comparison period;

[0018] The output unit is used to output a risk warning when the risk coefficient is greater than or equal to the risk threshold.

[0019] The present application also provides an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the above-mentioned data processing method when executing the program.

[0020] The present application also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned data processing method is implemented.

[0021] The data processing method, apparatus, equipment and storage medium provided in the present application include: real-time collection of transaction information of a first business during a transaction process; the transaction information includes at least address information representing the area where the transaction occurs; based on the address information, processing the transaction information of the first business to obtain the value of the set indicator of each of the at least two areas of the first business in the current cycle; querying the value of the set indicator of each of the at least two areas of the first business in at least one comparison cycle; determining the risk coefficient of the first business based on the value of the set indicator of each of the areas of the first business in the current cycle and the value of the set indicator of each of the areas of the first business in at least one comparison cycle; and outputting a risk warning when the risk coefficient is greater than or equal to the risk threshold.

[0022] For the solution of this application, the collected transaction information includes the address information representing the area where the transaction occurred. Then, based on the address information, the value of the set indicator of the first business in each area in the current cycle is obtained. Then, based on the value of the set indicator of the first business in each area in the current cycle and the value of the set indicator of the first business in each area in the comparison cycle, the risk coefficient of the first business is determined, and a risk warning is output when the risk coefficient is greater than or equal to the risk threshold. It can be seen that when determining the risk coefficient, this application first considers the influence of time based on the values ​​of the set indicators of different cycles, and secondly, considers the influence of space based on the values ​​of the set indicators of different areas, so the accuracy of the risk warning is high. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 An optional structural diagram of a data processing system provided in an embodiment of the present application;

[0024] Figure 2 An optional flow chart of the data processing method provided in the embodiment of the present application

[0025] Figure 3 An optional flowchart of the data processing method provided in the embodiment of the present application;

[0026] Figure 4 An optional flowchart of the data processing method provided in the embodiment of the present application;

[0027] Figure 5 An optional flowchart of the data processing method provided in the embodiment of the present application;

[0028] Figure 6 An optional flowchart of the data processing method provided in the embodiment of the present application;

[0029] Figure 7 An optional flowchart of the data processing method provided in the embodiment of the present application;

[0030] Figure 8 A schematic diagram of an optional process for analyzing financial transaction changes provided in an embodiment of the present application;

[0031] Figure 9 A schematic diagram of an optional process for analyzing financial transaction changes provided in an embodiment of the present application;

[0032] Figure 10 An optional flowchart of the real-time indicator warning processing process provided in the embodiment of the present application;

[0033] Figure 11 A schematic diagram of an optional process for dimension grouping provided in an embodiment of the present application;

[0034] Figure 12 An optional flowchart of business rule verification provided in an embodiment of the present application;

[0035] Figure 13 A schematic diagram of an optional structure of a data processing device provided in an embodiment of the present application;

[0036] Figure 14 This is a schematic diagram of an optional structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0038] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0039] In the following description, the terms "first, second, and third" are used merely as examples to distinguish between different objects and do not represent a specific order or precedence for the objects. It is understood that the specific order or precedence of "first, second, and third" can be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0041] The embodiments of the present application may provide a data processing method and apparatus, a device, and a storage medium. In practical applications, the data processing method may be implemented by a data processing apparatus, and the functional entities in the data processing apparatus may be collaboratively implemented by hardware resources of the electronic device, such as computing resources such as a processor, and communication resources (e.g., for supporting various communication methods such as optical cables and cellular communications).

[0042] The data processing method provided in the embodiment of the present application is applied to a data processing system, which includes an electronic device (for the sake of distinction, it may also be referred to as a first electronic device).

[0043] The electronic device is used to perform: real-time collection of transaction information of the first business during the transaction process; the transaction information at least includes address information representing the area where the transaction occurs; based on the address information, the transaction information of the first business is processed to obtain the value of the set indicator of each of the at least two areas of the first business in the current cycle; query the value of the set indicator of each of the at least two areas of the first business in at least one comparison cycle; based on the value of the set indicator of each of the areas of the first business in the current cycle, and the value of the set indicator of each of the areas of the first business in at least one comparison cycle, determine the risk coefficient of the first business; if the risk coefficient is greater than or equal to the risk threshold, output a risk warning.

[0044] Optionally, the data processing system may further include a second electronic device.

[0045] The second electronic device is used to carry out the transaction process of the first business.

[0046] It should be noted that the first electronic device and the second electronic device may be integrated into the same electronic device, or may be independently deployed on different electronic devices.

[0047] As an example, the structure of the data processing system may be as follows Figure 1 As shown, the first electronic device 10 and the second electronic device 20 are included. Data can be transmitted between the first electronic device 10 and the second electronic device 20.

[0048] Here, the second electronic device 20 is used to perform a transaction process of the first business.

[0049] The second electronic device 20 may include a mobile terminal device (such as a mobile phone, a tablet computer, etc.) or a non-mobile terminal device (such as a desktop computer, a server, etc.).

[0050] The first electronic device 10 can communicate with the second electronic device 20. Specifically, the first electronic device 10 is used to perform: real-time collection of transaction information of the first business during the transaction process; the transaction information at least includes address information representing the area where the transaction occurs; based on the address information, the transaction information of the first business is processed to obtain the value of the set indicator of each of the at least two areas of the first business in the current cycle; query the value of the set indicator of each of the at least two areas of the first business in at least one comparison cycle; based on the value of the set indicator of each of the areas of the first business in the current cycle, and the value of the set indicator of each of the areas of the first business in at least one comparison cycle, determine the risk coefficient of the first business; if the risk coefficient is greater than or equal to the risk threshold, output a risk warning.

[0051] The first electronic device 10 may be a mobile terminal device (such as a mobile phone, a tablet computer, etc.) or a non-mobile terminal device (such as a desktop computer, a server, etc.).

[0052] Next, combine Figure 1 The schematic diagram of the data processing system shown illustrates various embodiments of the data processing method and apparatus, device and storage medium provided in the embodiments of the present application.

[0053] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a data processing device; wherein the data processing device can be deployed in Figure 1 The first electronic device 10 in the embodiment of the present application is described below by taking the electronic device as an example of the execution subject.

[0054] Figure 2 A flow chart showing an optional data processing method is shown in FIG. Figure 2 The data processing method may include but is not limited to the following: Figure 2 S201 to S203 shown.

[0055] S201. The electronic device collects transaction information of the first business in real time during the transaction process.

[0056] The transaction information at least includes address information representing the area where the transaction occurs.

[0057] The embodiment of the present application does not limit the specific type of the first service, which can be determined according to actual conditions. For example, the first service can be a financial service or a non-financial service.

[0058] The embodiments of the present application do not limit the specific content of the transaction information and can be configured according to actual conditions.

[0059] In a possible implementation, the transaction information may include: address information, transaction time, and related setting indicators.

[0060] The embodiments of this application do not limit the specific type of address information and can be configured according to actual circumstances. For example, address information may include, but is not limited to, an Internet Protocol (IP) address, an address represented by a landline telephone number, an address represented by an ID number, an address represented by a mobile phone number, and the like.

[0061] In an actual transaction process, if there are multiple types of address information, different priorities can be configured in advance for different types of addresses so that the address with the highest priority is determined as the address information of the transaction.

[0062] The set indicator is an indicator for measuring the risk in the first business transaction process. The embodiment of the present application does not limit the specific type and number of the set indicator, and can be configured according to actual conditions.

[0063] Exemplarily, the set indicators may include but are not limited to: loan transactions, anti-fraud cases, etc.

[0064] S201 may be implemented as follows: the electronic device collects transaction information of the first business in a transaction process in real time based on a message middleware subscription message system (KAFKA).

[0065] S202: The electronic device processes the transaction information of the first service based on the address information to obtain a value of a set indicator of each of the at least two areas of the first service in the current period.

[0066] The embodiment of the present application does not limit the specific duration of a cycle, and can be configured according to actual conditions. For example, a cycle can be 24 hours, or a cycle can be one month.

[0067] The transaction process of the first business involves at least two areas. The embodiment of the present application does not limit the specific total number of areas and can be determined based on actual conditions.

[0068] The embodiment of the present application does not specifically limit the division method of regions, and can be configured based on actual conditions. For example, a city can be divided into a region; or a province can be divided into a region; or a community can be divided into a region, etc.

[0069] S202 can be implemented as follows: the electronic device obtains regional information based on the address information, uses the region and period as statistical dimensions, performs statistical analysis on the transaction information of the first business, and obtains the value of the set indicator of each of the at least two regions of the first business in the current period.

[0070] For example, the values ​​of the set indicators of the first business in each of at least two areas in the current cycle obtained in S202 can be: the number of anti-fraud cases in city A within 30 minutes, the number of anti-fraud cases in city B within 30 minutes, and the number of anti-fraud cases in city C within 30 minutes.

[0071] S203: The electronic device queries a value of a set indicator of the first service in each of at least two areas in at least one comparison period.

[0072] The embodiment of the present application does not limit the number of query comparison cycles, which can be determined based on actual conditions.

[0073] The embodiment of the present application does not impose a sole limitation on the selection of the comparison period. For example, when the number of comparison periods is one, the comparison period can be the time period before the current period, or the comparison period can be multiple time periods before the current period.

[0074] It should be noted that the at least two regions in the comparison period are consistent with the at least two regions in the current period, for example, both are region A and region B.

[0075] S203 may be implemented as follows: for each comparison period in at least one comparison period, the electronic device first searches the database for data of the comparison period, and then searches the value of the set indicator of each area in the specific comparison period.

[0076] S204. The electronic device determines a risk coefficient of the first service based on a value of a set indicator of the first service in each area in a current period and a value of a set indicator of the first service in each area in at least one comparison period.

[0077] The risk coefficient of the first business is used to characterize the possibility of risk in the first business transaction process. The embodiment of the present application does not limit the determination method of the risk coefficient, and it can be determined according to actual conditions.

[0078] S205: When the risk coefficient is greater than or equal to the risk threshold, the electronic device outputs a risk warning.

[0079] The embodiment of the present application does not limit the specific output method of the risk warning, which can be determined according to the actual situation. For example, the output method of the risk warning can include but is not limited to: text warning, voice warning, logo warning, etc.

[0080] The specific value of the risk threshold in the embodiment of the present application is not limited and can be determined according to actual conditions.

[0081] S205 may be implemented as follows: the electronic device determines the magnitude relationship between the risk coefficient and the risk threshold, and outputs a risk warning when the risk coefficient is greater than or equal to the risk threshold.

[0082] It should be noted that when the risk coefficient is less than the risk threshold, it is considered that there is no risk, or the risk warning level is not reached, and no risk warning is output at this time.

[0083] Optionally, the electronic device may also output the risk coefficient when issuing a risk warning, so that the risk level can be intuitively perceived through the risk coefficient.

[0084] It is understandable that the warning level can also be determined based on the risk factor, and the warning level can be output together with the risk warning.

[0085] The data processing solution provided by the embodiment of the present application includes: real-time collection of transaction information of the first business during the transaction process; the transaction information at least includes address information representing the area where the transaction occurs; based on the address information, the transaction information of the first business is processed to obtain the value of the set indicator of each of the at least two areas of the first business in the current cycle; query the value of the set indicator of each of the at least two areas of the first business in at least one comparison cycle; based on the value of the set indicator of each of the areas of the first business in the current cycle, and the value of the set indicator of each of the areas of the first business in at least one comparison cycle, determine the risk coefficient of the first business; if the risk coefficient is greater than or equal to the risk threshold, output a risk warning.

[0086] For the solution of this application, the collected transaction information includes the address information representing the area where the transaction occurred. Then, based on the address information, the value of the set indicator of the first business in each area in the current cycle is obtained. Then, based on the value of the set indicator of the first business in each area in the current cycle and the value of the set indicator of the first business in each area in the comparison cycle, the risk coefficient of the first business is determined, and a risk warning is output when the risk coefficient is greater than or equal to the risk threshold. It can be seen that when determining the risk coefficient, this application first considers the influence of time based on the values ​​of the set indicators of different cycles, and secondly, considers the influence of space based on the values ​​of the set indicators of different areas, so the accuracy of the risk warning is high.

[0087] Next, the process of the electronic device processing the transaction information of the first service based on the address information in S202 to obtain the value of the set indicator of each of the at least two areas of the first service in the current cycle is described.

[0088] like Figure 3 As shown, it may specifically include but not be limited to the following S2021 to S2023.

[0089] S2021. The electronic device parses the address information in the transaction information to determine the area to which the address information belongs.

[0090] Different regional ranges are pre-divided to determine the region to which the address information belongs based on the regional ranges. S2021 can be implemented as follows: the electronic device parses the address information in the transaction information, determines the transaction address based on the parsing result, and then determines the region to which the address information belongs based on the region to which the transaction address belongs based on the pre-divided regional ranges.

[0091] The embodiment of the present application does not limit the process of parsing address information, which can be determined based on actual conditions.

[0092] For example, the IP address in the transaction process can be parsed and the transaction address can be determined based on the IP address; or, the ID number in the transaction process can be parsed and the transaction address can be determined based on the first 6 digits of the ID number; or, the landline phone number in the transaction process can be parsed and the transaction address can be determined based on the area code in the landline phone number.

[0093] For example, an electronic device analyzes the IP address during a transaction and determines that the transaction address is in cell ABC. The pre-defined areas include: Area A, which includes cells ABC, ABD, and ABE; Area B, which includes cells BBC, BBD, and BBE; and Area C, which includes cells CBC, CBD, and CBE. Therefore, the area to which the address belongs is determined to be area A.

[0094] The above-mentioned method of determining the region to which the address information belongs can be used for each transaction. It should be noted that different parsing methods are configured for different types of address information; for multiple types of address information, the parsing method corresponding to the type with the highest priority is preferentially used.

[0095] For example, the priority from high to low is: IP address, telephone number, ID number; for transaction information including IP address, telephone number, and provincial ID number, the method of resolving the IP address to obtain the transaction address is preferred.

[0096] S2022. The electronic device determines that the transaction time belongs to the current cycle.

[0097] S2022 can be implemented as follows: the electronic device determines the time range of the current cycle based on the set duration of the cycle and the set rules, and the electronic device determines whether the collected transaction time belongs to the time range of the current cycle. If the transaction time belongs to the time range of the current cycle, it is determined that the transaction time belongs to the current cycle.

[0098] It should be noted that if the transaction time does not fall within the time range of the current cycle, it is determined that the transaction time does not belong to the current cycle. Because data from the current cycle needs to be processed at this time, the data that does not belong to the current cycle is temporarily placed and not further processed in this step.

[0099] For example, a cycle corresponds to a duration of 24 hours, and the current cycle is configured as 0:00 to 24:00 on October 10, 2022. If the transaction time collected by the electronic device is 12:23 on October 10, 2022, then the transaction time is determined to belong to the current cycle.

[0100] S2023. The electronic device performs statistical analysis on the transaction information of the first business based on the region, cycle, and set index as statistical dimensions to obtain the value of the set index of the first business in each region in the current cycle.

[0101] Based on the above S2021 and S2022, the region to which each transaction belongs and whether each transaction belongs to the current cycle can be determined. The corresponding S2023 can be implemented as follows: the electronic device performs statistical analysis on the transaction information of the first business with region, cycle and set indicators as statistical dimensions to obtain the value of the set indicator of the first business in each region under the current cycle.

[0102] Exemplarily, the data obtained in S2023 may include: the loan transaction volume in area A in the current period, the anti-fraud application volume in area A in the current period, and the account opening volume in area A in the current period.

[0103] Next, the process of determining the risk coefficient of the first business by the electronic device in S204 based on the value of the set indicator of the first business in each area in the current cycle and the value of the set indicator of the first business in each area in at least one comparison cycle is described.

[0104] refer to Figure 4 The content shown, the process may include but is not limited to the following S2041 to S2044.

[0105] S2041: The electronic device determines, for each of the at least two areas, a first proportion of each area in the current period based on a value of a set indicator of the first service in each area in the current period.

[0106] The first proportion is used to represent the ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas in the current cycle.

[0107] For example, if the transaction process of the first business involves five areas (area A, area B, area C, area D, and area E), then based on the values ​​of the set indicators of area A, the values ​​of the set indicators of area B, the values ​​of the set indicators of area C, the values ​​of the set indicators of area D, and the values ​​of the set indicators of area E in the current period, the first proportion of area A, the first proportion of area B, the first proportion of area C, the first proportion of area D, and the first proportion of area E in the current period are determined respectively.

[0108] The first proportion of region A is the ratio of the value of the set indicator of region A in the current period to the sum of the values ​​of the set indicators of the five regions in the current period.

[0109] S2042. The electronic device determines, for each of the at least two areas, a second proportion of each of the areas in each comparison period included in the at least one comparison period based on a value of a set indicator of the first service in each of the areas in the at least one comparison period.

[0110] The second ratio is used to represent the ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas for the comparison period.

[0111] The specific implementation process of S2042 is similar to that of S2041, and reference may be made to the description of S2041. The difference is that S2042 is processed for the comparison cycle, while S2041 is processed for the current cycle.

[0112] It should be noted that if there are multiple comparison cycles, the processing process for each comparison cycle is similar and will not be described in detail.

[0113] S2043. The electronic device determines, for each comparison cycle included in the at least one comparison cycle, a difference coefficient between the current cycle and the comparison cycle based on the first proportion of each of the areas in the current cycle, the second proportion of each of the areas in the comparison cycle, and the first proportion threshold, to obtain at least one difference coefficient.

[0114] The embodiment of the present application does not limit the specific method for determining the difference coefficient, and it can be determined according to actual conditions.

[0115] Here, one difference coefficient can be obtained for one comparison cycle, and multiple difference coefficients can be obtained for multiple comparison cycles.

[0116] In a possible implementation, the difference coefficient between the current cycle and the comparison cycle may be determined based on the following formula (1).

[0117]

[0118] In formula (1), W j represents the overall difference coefficient between the current cycle and the jth comparison cycle, ∑ represents the sum of the calculations for the regions 1 to n, and A i % represents the ratio of the value of the set index of the i-th region in the current period to the value of the set index of all regions, B i % represents the ratio of the value of the set index of the i-th region in the comparison period to the values ​​of the set index of all regions, and m represents the first proportion threshold.

[0119] S2044. The electronic device determines a risk coefficient of the first business based on the at least one difference coefficient.

[0120] The embodiment of the present application does not limit the specific method of determining the risk coefficient of the first business, and it can be determined according to actual circumstances.

[0121] In a possible implementation, the sum of at least one difference coefficient may be determined as the risk coefficient of the first business.

[0122] In another possible implementation, an average value of at least one difference coefficient may be determined first, and the average value is used as the risk coefficient of the first business.

[0123] The process of determining the first ratio threshold is described below, which may specifically include but is not limited to the following three implementations.

[0124] Implementation method 1: Determine the first proportion threshold based on experience.

[0125] Implementation method 2: Determine the first proportion threshold based on at least the first reference value and the second reference value.

[0126] For implementation method 1, after determining the risk coefficient of the first business, a first proportion threshold is determined based on experience, and then the first proportion threshold determined based on experience is adjusted according to multiple groups of test data.

[0127] Next, a process of determining the first proportion threshold based on at least the first reference value and the second reference value in Implementation Mode 2 is described.

[0128] refer to Figure 5 As shown in the content, the process may include but is not limited to S501 to S504.

[0129] S501: The electronic device determines a first array for the current cycle and a second array for the comparison cycle.

[0130] The first array includes a first proportion of each of the at least two areas in the current cycle, and the second array includes a second proportion of each of the at least two areas in the comparison cycle.

[0131] The data in the first array and the second array satisfy a first order.

[0132] The embodiment of the present application does not specifically limit the first order, and it can be determined according to actual conditions.

[0133] In a possible implementation, the first order may include: sorting from large to small according to the value of the first proportion or the second proportion.

[0134] It should be noted that the first array includes multiple data, each of which is a region in the current cycle.

[0135] In another possible implementation, the first order may include: sorting from small to large according to the value of the first proportion or the second proportion.

[0136] S501 can be implemented as follows: for the current cycle, the electronic device sorts the first proportion of each area according to the first order to obtain a first array, and for the comparison cycle, sorts the second proportion of each area according to the first order to obtain a second array.

[0137] S502: The electronic device constructs a first curve based on the first array, and constructs a second curve based on the second array.

[0138] S502 can be implemented as follows: the electronic device uses each value in the first array as a vertical coordinate and establishes evenly spaced horizontal coordinates to obtain a first curve; the electronic device uses each value in the second array as a vertical coordinate and establishes evenly spaced horizontal coordinates to obtain a second curve.

[0139] The embodiment of the present application does not limit the method of constructing the first curve or the second curve, which can be determined according to actual conditions.

[0140] For example, the electronic device can be based on (1, a1), (2, a2), (3, a3)..(i, a i ) Construct the first curve, a i Indicates the first proportion of the i-th region in the current period.

[0141] S503: The electronic device determines a first reference value in the first array that meets a first condition based on the first curve, and determines a second reference value in the second array that meets the first condition based on the second curve.

[0142] The embodiment of the present application does not specifically limit the content of the first condition and can be configured according to actual conditions.

[0143] Exemplarily, the first condition may include: k i >2(k i-1 +k i-2 ); where k i Indicates the slope of two adjacent points in the first curve or the second curve.

[0144] Correspondingly, the first reference value may be a first proportion corresponding to the first sudden increase point in the first curve, and the second reference value may be a second proportion corresponding to the first sudden increase point in the second curve.

[0145] S504: The electronic device determines the first proportion threshold based at least on the first reference value and the second reference value.

[0146] The embodiment of the present application does not limit the specific method for determining the first proportion threshold.

[0147] In a possible implementation, the smaller value between the first reference value and the second reference value may be determined as the first proportion threshold.

[0148] In another possible implementation, other reference values ​​may be obtained, and the first proportion threshold value may be determined based on the first reference value, the second reference value, and the other reference values.

[0149] In this possible implementation, the process of S504 in which the electronic device determines the first proportion threshold based on at least the first reference value and the second reference value may include but is not limited to the following S5041 to S5043.

[0150] S5041. The electronic device determines a third reference value that meets a second condition in the first array.

[0151] The embodiment of the present application does not limit the content of the second condition, which can be determined based on actual conditions.

[0152] In a possible implementation, the second condition may include: the first point in the first array or the second array that satisfies the condition that the sum of all previous points is greater than a safety threshold.

[0153] Exemplarily, in the first array, the electronic device determines the proportion corresponding to the first point that satisfies the condition that the sum of all previous points is greater than the safety threshold as the third reference value.

[0154] S5042. The electronic device determines a fourth reference value that meets the second condition in the second array.

[0155] The implementation process of S5042 can refer to the description in S5041 and will not be repeated here.

[0156] S5043: The electronic device determines that the first proportion threshold is the minimum value among the first reference value, the second reference value, the third reference value, and the fourth reference value.

[0157] The following describes the process of determining the combination of warning risk dimensions.

[0158] In practice, after determining that a business has a warning risk, further analysis is needed to determine the risk's dimensional combinations. For example, if the business has multiple risk dimensions, including product, channel, and risky customer groups, further analysis is needed to determine the specific risk dimension combination within which the risk exists. Specifically, the specific risk exists within which product, channel, and risky customer group.

[0159] Specifically, refer to Figure 6The process may include but is not limited to the following S601 to S605.

[0160] S601: The electronic device determines n risk dimensions of a first service, and divides the n risk dimensions into m groups in descending order of risk weight.

[0161] The n is an integer greater than 1; the n is greater than the m, and a group includes at least one risk dimension.

[0162] Risk dimension refers to the subdivision dimension under the first business.

[0163] For example, the electronic device determines that there are 18 risk dimensions for the first business. After grouping, the following are 6 product dimensions: P1, P2, P3, P4, P5, and P6; 5 channel dimensions: U1, U2, U3, U4, and U5; and 7 risk customer group dimensions: R1, R2, R3, R4, R5, R6, and R7.

[0164] S602: The electronic device replays the transaction of the first business and selects m×x first-level risk dimensions from the risk dimensions of the first group based on the values ​​of the set indicators.

[0165] The x is an integer greater than 1.

[0166] Suppose we need to know the top x m combined dimensions.

[0167] S602 may be implemented as follows: the electronic device replays the transaction of the first business, and selects the first m×x first-level risk dimensions from the risk dimensions of the first group based on the order of the values ​​of the set indicators from large to small.

[0168] For example, assuming that the top two three combined dimensions are needed, the risk dimensions are ranked in descending order according to risk weight: risk customer group > product > channel, then six risk customer groups (R1, R2, R3, R4, R5, R7) can be obtained in S602.

[0169] S603. The electronic device combines the (m-i+1)×x risk dimensions of the i-th group with all the dimensions in the i+1-th group to obtain a combination of h risk dimensions; replays the transaction of the first business, and based on the value of the set indicator, selects a combination of (mi)×x risk dimensions from the combination of h risk dimensions.

[0170] The i and h are integers greater than 1.

[0171] For example, the electronic device combines the six risk dimensions in the first group with all the dimensions in the second group to obtain a combination of 36 risk dimensions; replays the transactions of the first business, and based on the order of the values ​​of the set indicators from large to small, selects a combination of four risk dimensions (R2P4, R3P1, R3P5, R5P3) from the 36 risk dimension combinations.

[0172] S604. The electronic device increases i by 1 and re-executes the combination of the (m-i+1)×x risk dimensions of the i-th group with all the dimensions in the i+1-th group to obtain h combinations of risk dimensions; replays the transaction of the first business, and based on the value of the set indicator, selects a combination of (mi)×x risk dimensions from the combinations of h risk dimensions; until x risk dimension combinations consisting of m risk dimensions are obtained.

[0173] For example, by combining the obtained combinations of four risk dimensions (R2P4, R3P1, R3P5, R5P3) with five channels, a total of 20 combination dimensions can be obtained. Then, among the 20 combination dimensions, they are sorted from large to small based on the values ​​of the set indicators to obtain the top two three-level combination dimensions (R3P1U4, R5P3U2).

[0174] S605. The electronic device determines at least one warning risk dimension combination from the m×x first-level risk dimensions, the combination of the (mi)×x risk dimensions, and the x risk dimension combinations consisting of the m risk dimensions based on the values ​​of the set indicators.

[0175] Based on the order of the values ​​of the set indicators from large to small, the warning risk dimension combination is determined from the m×x first-level risk dimensions, the combination of the (mi)×x risk dimensions, and the x risk dimension combinations consisting of m risk dimensions.

[0176] Exemplarily, the determined warning risk dimension combinations include: (R3P1U4, R5P3U2).

[0177] In this way, the detailed warning risk dimension of the first business is obtained.

[0178] The data processing method provided in the embodiment of the present application may also include a verification process for verifying the warning risk dimension.

[0179] refer to Figure 7 The process may include but is not limited to the following S701 to S704.

[0180] S701. The electronic device determines a warning risk dimension combination to be verified from the at least one warning risk dimension combination.

[0181] During the verification process, each warning risk dimension combination is generally determined in turn as the warning risk dimension combination to be verified.

[0182] Exemplarily, the warning risk dimension combination determined by the electronic device is R3P1U4.

[0183] S702. The electronic device restricts the warning risk dimension combinations other than the warning risk dimension combination to be verified in the at least one risk dimension combination, so as to perform transaction verification based on the warning risk dimension combination to be verified by using the bypass of the first business.

[0184] Exemplarily, the electronic device performs transaction verification by bypassing the first service through R3P1U4.

[0185] Here, the verification process through bypass does not affect normal business transactions and has good compatibility.

[0186] S703: The electronic device collects the values ​​of the indicators set during the transaction verification process.

[0187] The collection process here can refer to the detailed description in S201, which will not be repeated here.

[0188] S704: When the value of the set indicator satisfies the third condition, the electronic device determines the warning risk dimension combination to be verified as the target risk dimension combination.

[0189] The embodiment of the present application does not limit the specific content of the third condition, which can be determined based on actual conditions.

[0190] For example, when the electronic device determines that during the R3P1U4 verification process, the anti-fraud application volume for product A has surged, R3P1U4 is determined as the target risk dimension combination.

[0191] It should be noted that if the warning risk dimension combination to be verified does not meet the third condition, it is determined that the warning risk dimension combination to be verified is not the target risk dimension combination; then a new warning risk dimension combination to be verified is re-determined and a new verification process is started.

[0192] The following uses financial transactions as an example to illustrate the data processing method provided in the embodiments of the present application through an embodiment.

[0193] To facilitate understanding, some technical terms involved in this embodiment are first explained.

[0194] Financial transaction risk: refers to factors that affect the security and stability of financial system transactions.

[0195] Transaction risks: may include fraud risk, credit risk, security and compliance risk, etc.

[0196] System risks: These may include risks of external data source anomalies, system failure (bug) risks, risks of basic hardware service anomalies, and so on.

[0197] Risk detection: It can refer to detecting risk events or risky transactions from daily system transactions.

[0198] KAFKA: refers to a high-throughput distributed publish-subscribe messaging system.

[0199] HIVE (data warehouse platform): used for offline big data processing.

[0200] Time sequence: used to distinguish different times.

[0201] Attribution analysis: refers to automatically locating the cause of risk when a risk occurs.

[0202] KAFKA offset: used to determine the position of the message queue in KAFKA.

[0203] In related technologies, financial transaction analysis includes real-time and offline methods.

[0204] Offline processing typically involves analyzing T-1 cycle data (e.g., yesterday's data) in a large database related to the business. This primarily involves locating the cause of a risk event or anomaly. Based on this, corresponding expert rules are then developed to provide targeted early warnings and resolution plans. A specific approach to detecting risk fluctuations in financial scenarios is based on existing expert rules.

[0205] Specifically, it may include but not be limited to the following steps A1 to A5.

[0206] Step A1: A risk event occurs during normal transactions, resulting in abnormalities in key indicators such as rejection rate and approval rate (equivalent to the above-mentioned set indicators), indicating possible risks.

[0207] Step A2: Business personnel process risk indicators of different types and dimensions based on risk performance and past experience, such as fraud risk, credit risk, and compliance risk.

[0208] Step A3: Perform weighted summation of risk indicators of different dimensions according to previous expert rules to obtain the current risk score.

[0209] Step A4: Compare the current risk score with the preset risk threshold. If the risk threshold is exceeded, the current risk is prompted.

[0210] Step A5: The dimension corresponding to the first-level risk dimension indicator that exceeds the threshold is used as the first-level risk dimension. Then, within this first-level risk dimension, various indicators are processed for the preset second-level dimensions, such as the corresponding bad customer rate, to obtain the second-level dimension with the highest risk, which serves as the final risk factor. Some solutions may consider multiple dimensions, but because they all manually process indicators for lower-level dimensions, the number of dimensions is always limited.

[0211] In simple terms, the process is as follows Figure 8 As shown, the process includes risk event 801, first-level risk dimension detection 802, and second-level risk dimension detection 803. The data flow for first-level risk dimension detection 802 includes: trading system, offline indicator processing, indicator warning threshold determination (below the threshold), and warning termination. The data flow for second-level risk dimension detection 803 includes: indicator warning threshold determination (above or equal to the threshold), offline second-level down-level indicator processing, indicator warning threshold determination, second-level risk dimension, and final result.

[0212] For real-time processing, it is not a simple matter of specifying a threshold as a risk monitoring standard based on past expert rules. Instead, the indicators of the current cycle are compared with those of the past cycles, and the coefficient of difference between the indicators of the current cycle and the indicators of the past cycles is obtained as the risk coefficient. Then, a threshold is set for this risk coefficient. The essence is to judge the current cycle and the past cycles.

[0213] Specifically, it may include but not be limited to the following steps B1 to B5.

[0214] Step B1: Based on past rules, preset key indicators are used to process the indicator results of each dimension in real time.

[0215] Step B2: Fit and compare the indicator results of the current cycle of each indicator with the indicator results of the previous cycles to determine the risk coefficient of the difference in the current cycle.

[0216] Step B3: Determine the size of the risk coefficient, conduct risk assessment based on the size of the difference, derive a risk score for risk warning, and obtain a first-level risk dimension.

[0217] Step B4: perform time fitting extraction on the lower sticky dimension.

[0218] Step B5: Find the secondary sub-dimension with the greatest difference and obtain the secondary risk factor.

[0219] In simple terms, Figure 9As shown, the data flow of this process may include but is not limited to: trading system (transaction reporting), preset indicator processing, comparison of fit judgment with past cycles, end of warning when below the threshold; when above the threshold, generate warning, extract the first-level risk dimension, perform time fitting extraction on the lower sticky dimension, find the second-level sub-dimension with the largest difference, and obtain the risk factor.

[0220] An analysis of the above two processing methods shows the following main defects.

[0221] The defects of offline risk attribution analysis may include but are not limited to the first to third points below.

[0222] First, the timeliness issue. Offline anomaly analysis requires a T+1 cycle to understand the cause of a problem in the T cycle (for example, the cause of today's problem can only be analyzed tomorrow). In financial scenarios, actual losses have already occurred in the T+1 cycle, and fraud gangs usually change their tactics immediately after making a profit, making it difficult to truly mitigate risks.

[0223] Second, flexibility: Offline data is processed in batches (for example, batch analysis of all yesterday's data). This processing time is long due to the large amount of data. If there are many dimensions, it is difficult to quickly identify risk dimensions.

[0224] Third, the accuracy of risk dimension analysis is poor: risk dimension analysis based on experience and human experience analysis is limited by energy and experience, and it is difficult to make the relevant dimension analysis objective and accurate.

[0225] The shortcomings of real-time risk attribution analysis may include but are not limited to:

[0226] First, the timeliness of verification: transactions are one-time, and after a risk occurs, the effectiveness of the new disposal plan cannot be verified in real time;

[0227] Second, there is no complete early warning and disposal plan: In financial scenarios, conventional risk anomaly detection and attribution solutions currently focus primarily on risk detection, lacking a comprehensive risk detection + rule verification + risk disposal solution.

[0228] Third, the accuracy of risk dimensions is poor: Current real-time risk anomaly attribution analysis solutions are all based on expert rules and fixed maintenance indicator processing. They are unable to meet the needs of increasingly small-scale fraud incidents with small rules. For example, risk events occur in specific channels, specific customer groups, or specific regions of a certain product. Traditional attribution analysis usually only locates channels or customer groups, and it is difficult to find anomalies in this combined dimension. Direct processing of channels or customer groups has poor accuracy and is prone to "accidental injuries."

[0229] Offline risk detection is increasingly unable to meet the timeliness requirements of various fraud, credit and other risk events faced by current financial scenarios. Taking financial fraud as an example, current gang fraud is characterized by gang-based crimes, diverse methods, flexibility and cunningness. Fraud gangs usually use the same method for a period of time, and even the relevant methods change their form immediately after using them once, which puts great pressure on anti-fraud. Offline data risk detection is obviously difficult to meet the requirements of real-time confrontation.

[0230] Real-time risk detection solutions are limited by performance. Consider the following situation: when a financial scenario is rolled out and the number of customers increases, there are many customer segmentation dimensions. Common dimensions include: age group, gender, customer group identification, education, occupation, etc. These dimensions are presented in the form of permutations and combinations, resulting in a large number of segmentation dimensions. Real-time indicator processing is limited by performance, and it is impossible to process all segmentation dimension indicators. Simply processing, analyzing, and issuing warnings for indicators in a certain dimension will result in poor accuracy and the problem of "inadvertently injuring" non-risk customers.

[0231] In response to the demand for real-time risk confrontation in financial scenarios, the fox of this application adopts a real-time indicator processing method. In order to solve the problem that real-time risk detection and attribution cannot verify risk information in real time, this embodiment of the application uses KAFKA to realize real-time playback of transactions, solves the problem of rule verification, and then uses a customized early warning dimension mining algorithm to realize real-time mining of segmented risk dimensions.

[0232] Specifically, this embodiment of the present application has the following features:

[0233] First, by recording the KAFKA offset of statistical messages, transaction playback is achieved, facilitating real-time verification of business rules. This allows for rapid rule verification and tuning, making it suitable for real-time countermeasures against financial fraud transactions.

[0234] Secondly, by comparing the indicator proportions of the regions (equivalent to the upper and lower regions) in the current period (equivalent to the first proportion mentioned above) with the indicator proportions of the regions in the comparison period (equivalent to the second proportion mentioned above), and combining them with set thresholds to calculate the indicator's difference coefficient, we can determine the risk anomalies of the indicators based on parameters in the time and space dimensions, taking into account the temporal and spatial differences in transactions, and accurately detect transaction anomalies.

[0235] Secondly, we have designed an automatic attribution algorithm that, through intelligent combined dimension indicator processing solutions, can quickly and accurately provide prompts for possible risk dimensions, facilitating quick business decision-making.

[0236] Among them, the risk warning and abnormal movement analysis solutions in the financial industry are mainly limited by the following problems, resulting in the above-mentioned defects.

[0237] Traditional offline data processing is difficult to meet the requirements of real-time risk mitigation. Traditional risk discovery and rule verification in the financial industry are all carried out in offline large databases after risk events. Real-time risk warning solutions are difficult to verify in real time. The main reason is that transaction data is one-time, and it is impossible to replay transaction data and verify rules based on transaction data. The impact of time and space on transaction characteristics is not fully considered, and simple practical thresholds are used for risk detection, which is not flexible enough. For segmented warning dimension analysis, expert rules are currently used for manual conjecture verification, which makes it difficult to achieve precise warning dimension analysis.

[0238] The implementation process of this embodiment of the present application is described below.

[0239] Specifically, it may include but not be limited to the following first to fourth parts.

[0240] Part 1: Transaction data collection.

[0241] Transaction data is collected from the original transaction system. Common transactions include account opening, loan application, anti-fraud, etc. The upstream system reports the relevant transactions through the message middleware KAFKA for downstream systems to process the data.

[0242] Part 2: Real-time indicator warning processing.

[0243] Data indicator warnings in related technologies include: first processing key indicators, setting alarm thresholds in advance, and judging whether to generate an alert by determining whether the thresholds are met. The problem faced by traditional methods is that they require frequent adjustments. Financial transactions are usually subject to irregular fluctuations due to factors such as holidays and marketing activities. Simply issuing warnings through thresholds is prone to false alarms.

[0244] For the indicator warning of this embodiment of the present application, full consideration is given to the characteristics of financial transactions being affected by time and space. Space refers to the region. Specifically, in financial transactions, transactions usually include Internet Protocol (IP) addresses, mobile phone numbers, ID numbers, landline telephones and other information. This information can be used to derive the address information of the transaction. For example, the first 6 digits of a mobile phone number represent the location of the mobile phone number, and the first 6 digits of the ID number also represent the province and city (such as the first 6 digits of the ID number 420116 represent Wuhan City, Hubei Province). The first few digits of the communication IP can also obtain the address, and the landline area code can also obtain the address (for example, 0755 represents Shenzhen). In specific scenarios, a priority rule is usually followed, giving priority to the address corresponding to the IP. If there is no IP information, the mobile phone number is taken, and so on.

[0245] Address information can be parsed and derived. Mobile phone numbers can be used to derive relevant locations, and ID card numbers can also be used to obtain relevant address information. Once address information is obtained, relevant indicator data for each region can be processed. Unlike traditional indicator threshold solutions, this system uses a custom set of algorithms to implement indicator warnings. This algorithm determines the relevant stability, and the degree of stability is used to determine whether system transactions are normal.

[0246] The specific indicator processing is to use FLINK to call the messages in KAFKA, and periodically aggregate and calculate the relevant transactions according to the time window, and process indicators such as the 5-minute anti-fraud application volume and the 1-hour loan transaction volume. During the FLINK processing process, it is necessary to aggregate and calculate the relevant transactions in FLINK according to the time window, and record the message offset of the transaction corresponding to the start and end of the window in KAFKA, and record it to facilitate transaction playback and rule verification that may be needed later.

[0247] The early warning process of the second part may include but is not limited to the following S1 to S6.

[0248] S1. Read the KAFKA transaction and record the consumer offset of the start and end positions of the KAFKA window in the current cycle.

[0249] S2. The address code corresponding to the transaction is derived based on the IP, mobile phone number, ID card and other information contained in the transaction.

[0250] S3. After obtaining the address information, use the address as a statistical dimension to process the transaction indicators of each region in the current cycle.

[0251] S4. Query the indicator data of different regions in the past X periods according to the configuration.

[0252] S5. Use intelligent early warning algorithms to conduct risk warnings and determine the indicator risk coefficient W.

[0253] S5 may specifically include but is not limited to the following steps a to e.

[0254] Step a: Encode all addresses in sequence. Assuming there are n areas in total, number all areas starting from 1 to n.

[0255] Step b: Define the region variable i. The value of i starts from 1 and gradually increases to n. i represents the i-th region. A i represents the ratio of the key indicator value of the i-th region in the current period to the key indicator value of all regions, B i Represents the ratio of the key indicator value of the i-th region to the key indicator values ​​of all regions in the comparison period.

[0256] Step c: Define a periodic variable j. When the value of j increases step by step from 1 to the current period x, for each past period, calculate the difference coefficient W of its key indicators from the current period. j .

[0257] For each past region i, if i % < m, and i % < m, then exclude this region in the current risk statistics. m is a set threshold. Dimensions where the proportion of key indicators (such as household registration quantity) is less than this value are considered to have negligible influence, but it is required to be satisfied simultaneously. In the actual scenario, when the proportion of the same region in the current period and the past period is small enough simultaneously, it is removed when calculating the difference coefficient. However, if only the current period is less than this value, or the comparison period is less than this value, it cannot be ignored because such risks are greater in the actual scenario. For example, if a region with no transactions suddenly has a large transaction volume, it indicates a greater risk.

[0258] Step d: The overall difference coefficient W between the current period and the j-th past period. j . The specific calculation can refer to the following formula (1).

[0259] Among them, take the logarithm of the ratio of the proportion in the current period to the proportion in the comparison period, and then multiply the difference between the proportion in the current period and the proportion in the comparison period. The purpose of taking the logarithm is to prevent the excessive influence of the difference in a single region. Then sum over J periods to obtain the overall difference coefficient.

[0260] [[ID=二十一]]W j The larger W is, the greater the difference, indicating that the overall transaction distribution in the current period is different from the past, and risks need to be promptly noticed and excluded. j The smaller W is, the better the distribution fit between the current period and the past period, and there is no obvious risk fluctuation. Through W, j Overall risk anomalies can be quickly and sensitively detected, facilitating business decision-making analysis.

[0261] Therefore, in this solution, in this embodiment of the present application, by reasonably setting the value of m, the following effects can be achieved:

[0262] First, prevent the denominator from being 0 when the proportion of a region is 0, which would make the calculation formula meaningless. That is, even in the case of no transactions, this factor is still considered as a factor for judging transaction risks instead of being directly ignored, ensuring the accuracy of risk anomaly analysis.

[0263] Second, because there is a squared value, the impact on the overall ratio result is extremely small and has little impact on the actual analysis result, but it is not directly ignored, achieving an accurate analysis result.

[0264] Step e: Calculate the average difference coefficient between the current period and the past period as the stability coefficient W.

[0265] The calculation of the stability coefficient W can refer to the following formula (2).

[0266]

[0267] In formula (2), W represents the stability coefficient, Indicates the average difference coefficient between the current period and the past periods.

[0268] 6. According to the calculated stability coefficient W, and the reasonable stability coefficient W t Compare, if W>W t , an early warning is generated.

[0269] The value of the threshold m (equivalent to the first proportion threshold mentioned above) is:

[0270] The threshold m cannot be set simply based on expert experience, otherwise there will be problems such as false alarms. The value of m needs to consider the following two issues.

[0271] The first question is that the regional distribution is even, and the proportion in most regions is at a low level.

[0272] Assuming that 90% of the regions are at a relatively low and uniform level, setting a threshold arbitrarily may result in most regions being excluded, resulting in inaccurate results.

[0273] The second problem is that as time changes, the overall distribution changes and the threshold becomes invalid.

[0274] First, sort all regions from small to large. Suppose that previously it was set that there were 10% of cities with a proportion below one ten-thousandth, but as time changes, the proportion of all cities is above one ten-thousandth, then the threshold will become invalid.

[0275] This embodiment of the present invention designs a threshold calculation method, which may specifically include but is not limited to the following steps C1 to C3.

[0276] Step C1: compose an array of the percentages of the key indicator values ​​of all regions in the current period and the comparison period. Assume that the percentage array of the key indicator values ​​of the current period is a i , the ratio array of the key indicator values ​​of the comparison period is b i .

[0277] Step C2: i , b iSort them from smallest to largest and remove all 0 values.

[0278] Step C3, calculate a respectively i , b i The first sudden increase point a x , b y .

[0279] With a i For example, construct (1, a1), (2, a2), (3, a3)..(i, a i ) Construct a curve and calculate the slope k between two adjacent points i , get the first sudden increase point, satisfy k i >2(k i-1 +k i-2 ), that is, the point is considered to be the first sudden increase point. Using the same method, we can obtain a i , b i The first sudden increase point in the , assuming that they are a x , b y , the final threshold m = min(a x , b y ), which is the smaller of the two.

[0280] In addition, in order to prevent too many areas from being filtered, resulting in too few calculated samples, a safety threshold z can be set.

[0281] Assume that the safety threshold z is 10%, with a i For example, calculate the first point that satisfies the sum of all previous points greater than 10%, assuming it is a j , j is satisfied The same method is used to obtain the first point of b. i Minimum point b k , the calculation of the final threshold m can refer to the following formula (3).

[0282] m=min(a x , b y , a j , b k ) formula (3);

[0283] In formula (3), a x Indicates a i The first sudden increase point, b y Indicates b i The first sudden increase point, a j Express satisfaction The first point, b k Express satisfaction The first point.

[0284] In this case, in this embodiment of the present application, the regional proportions in the current cycle and the comparison cycle are determined respectively, and then the regional proportions in different cycles are determined, the sudden increase point is determined, and then the minimum value is found among the two sudden increase points to determine the threshold value, thereby realizing the setting of the threshold value. It is set according to the actual situation of the current cycle and the comparison cycle, which improves the accuracy of the dynamic setting of the threshold value, and the minimum value of the two is used as the threshold value, which ensures that the subsequent impact on the calculation of the difference coefficient is reduced to a minimum, thereby improving the accuracy of the calculation result.

[0285] The following is an example to illustrate the process of determining the value of m.

[0286] Assume that the proportion of customers in each region in the current cycle (equivalent to the first proportion mentioned above) is 0.06, 0.003, 0.007, 0.012, 0, 0, 0, 0, 0, 0.001, 0.003, 0.12, 0.0026, 0.002, 0.0074, 0.018, 0.25, 0.16, 0.19, 0.0035, 0.14, 0.0205, and assume that the proportion of customers in each region in the previous cycle is 0.06, 0.003, 0.007, 0.012, 0.000, 0.000, 0.001, 0.003, 0.12, 0.0026, 0.002, 0.0074, 0.018, 0.25, 0.16, 0.19, 0.0035, 0.14, 0.0205, respectively. The proportion of customer volume in the district (equivalent to the second proportion mentioned above) is 0.05, 0, 0.008, 0.0025, 0.15, 0, 0.0032, 0.0001, 0.0006, 0.0044, 0.25, 0.141, 0.035, 0.0015, 0.045, 0.0058, 0.0026, 0.16, 0.0006, 0.0032, 0.0015, 0.015, and 0.12 respectively.

[0287] The m value calculation process of step may include but is not limited to the following steps D1 to D4.

[0288] Step D1, sort the current cycle into 0, 0, 0, 0, 0, 0.001, 0.002, 0.0026, 0.003, 0.003, 0.0035, 0.007, 0.0074, 0.012, 0.018, 0.0205, 0.06, 0.12, 0.14, 0.16, 0.19, 0.25, and compare the cycles 0.0001, 0.0006, 0.0006, 0.0015, 0.0015, 0.0025, 0.0026, 0.0032, 0.0032, 0.0044, 0.0058, 0.008, 0.015, 0.035, 0.045, 0.05, 0.12, 0.141, 0.15, 0.16, 0.25.

[0289] Step D2: Remove all zero values ​​in the current period to obtain 0.001, 0.002, 0.0026, 0.003, 0.003, 0.0035, 0.007, 0.0074, 0.012, 0.018, 0.0205, 0.06, 0.12, 0.14, 0.16, 0.19, and 0.25.

[0290] The coordinates are constructed according to the proportion of the current period: (1, 0.001), (2, 0.002), (3, 0.0026), (4, 0.003), (5, 0.003), (6, 0.0035), (7, 0.007), (8, 0.0074), (9, 0.012), (10, 0.018), (11, 0.06), (12, 0.12), (13 , 0.14), (14, 0.16), (15, 0.19), (16, 0.25), the slope between two adjacent coordinates is 0.001, 0.0006, 0.0004, 0, 0.0005, 0.0035, 0.0004, 0.0046, 0.006, 0.0025, 0.0395, 0.06, 0.02, 0.02, 0.03, and the first one that satisfies k i >2(k i-1 +k i-2 ) is 0.02, then the corresponding coordinates are (1, 0.007), and we get a x It is 0.007.

[0291] Use the same method to obtain b y is 0.035.

[0292] Assuming the safety threshold is 0.1, then a j is 0.06, b k is 0.045.

[0293] Step D4: The final m is 0.007.

[0294] m=min(0.007, 0.035, 0.06, 0.045)

[0295] The second part of the real-time indicator warning processing process can be as follows Figure 10 As shown, it includes: derived address information 1001, processed address related results 1002, query of indicator data of past cycles 1003, intelligent stability algorithm 1004, judgment based on stability coefficient whether it exceeds the threshold 1005, generating warning if exceeded 1006, and completion if not exceeded 1007.

[0296] Part III: Auxiliary analysis of risk dimensions.

[0297] Once a warning has been detected, the next step is to conduct a rapid risk attribution analysis, quickly locate the detailed dimensions of the risk, and then take targeted action.

[0298] Traditional attribution analysis is typically conducted offline in a big data environment, using expert experience to process metrics across specific dimensions. For example, an abnormal volume of incoming applications from a channel doesn't typically indicate abnormalities for all customers within that channel; rather, it typically refers to a specific customer group within that channel, requiring rapid identification and intervention. As system volume increases, the number of entry points increases, and the customer base becomes more complex, traditional manual attribution analysis becomes labor-intensive and inaccurate.

[0299] To address the problems existing in stock attribution analysis, this embodiment of the present application designs a new attribution analysis algorithm that does not restrict sub-dimensions and can intelligently prompt possible risk dimension combinations.

[0300] Because the transaction has already passed, the first problem is how to replay the transaction. The solution here is to use the KAFKA message offsets of the window start and end positions recorded in the second step of the solution to read the transaction cached in KAFKA for secondary consumption.

[0301] Taking the anti-fraud application volume indicator as an example, the risk dimension auxiliary analysis process may include:

[0302] The second part records the KAFKA message offsets of the window start and end positions. Messages of the corresponding period are read from KAFKA for secondary consumption. Assume that there are n dimensions and m sub-dimension types, and the top x dimension combinations need to be known.

[0303] Specifically, it may include but not be limited to the following steps E1 to E3.

[0304] Step E1: First, divide the n dimensions into m groups. Note that each group represents a sub-dimension.

[0305] In real scenarios, there are four dimensions, such as channels, risky customer groups, age groups, and regions. First, they are sorted according to risk weights. For example, if the risky customer group priority contributes the most to distinguishing transactions from a business perspective, it will be ranked first. The remaining dimensions are sorted according to weights.

[0306] Step E2: Define a variable i, which is incremented from i=1 to m, and replay the transaction to obtain the indicator values ​​of the top mx dimension in the first group of the period.

[0307] Suppose there are m = 10 dimensions in total, and we need to know the top 5 combined dimensions. The first group has 200 dimensions, so the dimensions that need to be taken out of the first group are the first mx = 50 sub-dimensions.

[0308] Step E3: When i>1, it is necessary to obtain the top(m-i+1)x dimensions of the previous group, combine them with all the dimensions in the current group, and then replay the transaction to obtain the top(mi)x combined dimensions of the i-th group.

[0309] In order to ensure that the dimensions with higher weights have higher accuracy in the result analysis when calculating indicators, during the first aggregation and sorting, all dimensions of the first group are obtained for aggregation calculation to obtain the top mx dimensions in the first group of dimensions. For the second aggregation, the selected mx dimensions of the first group are combined with all dimensions of the second group to obtain the top (m-1)x second-level combined dimensions. When it comes to the last group, the 2x (m-1)-level combined dimensions of the second group are combined with all dimensions of the last group, and then aggregation processing is performed to obtain the final topx m-level combined dimensions as the final dimensions.

[0310] like Figure 11 As shown, there are m groups of dimensions: the first group consists of a1, a2, a3, ..., ak; the second group consists of b1, b2, b3, ..., bk; and the mth group consists of m1, m2, m3, ..., mk. All dimensions in the first group are aggregated to produce the top mx sub-dimensions, ai1, ai2, ..., amx. The mx first-level dimensions selected from the first group are combined with all dimensions in the second group to obtain the top (m-1)x second-level dimensions ab1, ab2, ..., ab(m-1)x, respectively. The resulting x m-level dimensions are ab...m1, ab...m2, ..., ab...mx.

[0311] The process of auxiliary analysis of risk dimensions in the third part is explained through an example.

[0312] Example: Suppose a product has three levels of dimensions (product, channel, and risky customer group) and 11 dimensions. The product category is P1, P2, P3, P4, P5, and P6, for a total of six products. The channels are U1, U2, U3, U4, and U5, for a total of five channels. The risk customer groups are R1, R2, R3, R4, R5, R6, and R7, for a total of seven risky customer groups. Assuming the risk weight ranking of business tasks is risky customer group > product > channel, the business needs to know the top two combined dimensions.

[0313] Specific steps may include but are not limited to the following steps F1 to F3.

[0314] Step F1: Obtain all risk customer groups, replay the transaction data of the corresponding period, and obtain the top 3×2=6 risk customer groups based on the risk customer groups.

[0315] Assume that the risk customer groups are R1, R2, R3, R4, R5, and R7.

[0316] Step F2: Combine the six risk customer groups obtained in step F1 with all products to obtain a total of 6×6=36 second-level combination dimensions. Then, obtain the top2×2=4 second-level combination dimensions, assuming they are R2P4, R3P1, R3P5, and R5P3.

[0317] Step F3: Combine the four secondary combination dimensions and the five channels obtained in the second step to obtain a total of 20 three-level combination dimensions. Then, perform indicator processing to obtain the top two three-level combination dimensions among the 20 combination dimensions, assuming they are R3P1U4 and R5P3U2. Finally, R3P1U4 and R5P3U2 are the target two three-level combination dimensions.

[0318] In this way, it can be suggested that the indicator data under the combined dimension of R3P1U4 and R5P3U2 may be abnormal.

[0319] In the processing of multi-level combined dimension indicators, group by is traditionally used. However, a problem with simple group by is that when there are three levels of combined dimensions, the number of indicators obtained is the product of the number of each dimension. In the example above, the final indicator dimension is 6×5×7=210 dimensions.

[0320] But if there are 5 dimension types, and each dimension has 20 sub-dimensions, there will be 20 5 The real-time calculation memory will overflow and the result data cannot be persisted because the data volume is too large. Assuming that only the top 3 dimensions are needed, the data volume of this embodiment of the present application is 15×12×9×6×3=29160, then the data volume will be 3200000 / 29120=100, and the data volume will be reduced to one percent of the original. Assuming that there are more dimension types and more sub-dimensions, the data volume optimization effect will be more obvious.

[0321] Part 4: Business rule verification (equivalent to the above verification).

[0322] After obtaining the system's auxiliary attribution analysis, the business can formulate relevant rules for specific risk dimensions. After formulating the relevant rules, real-time rule verification is then performed. The business rules are then verified in real time online using a playback solution until the final transaction results meet business expectations.

[0323] Specifically, it may include but not be limited to the following steps G1 to G4.

[0324] Step G1: Based on the results of the attribution analysis and the business's own expert experience, refine the business rules;

[0325] Step G2: Replay the transactions of the corresponding period and apply business rules to transaction approval;

[0326] Step G3: Re-process the key indicators and verify the indicator results;

[0327] Step G4: If the result does not meet expectations, adjust the rules and repeat steps G1 to G3 until the final result meets expectations.

[0328] In simple terms, Figure 12 As shown, the fourth part of business rule verification may include: refining business rules 1201, replaying transactions 1202, indicator processing 1203, verifying results 1204, adjusting business rules 1205, and generating final rules 1206.

[0329] Below, a real business scenario is used as an example to illustrate the data processing process of this application.

[0330] Go to step 5.

[0331] The first step is to identify risks.

[0332] After the business discovers an anomaly through the risk monitoring solution provided by this embodiment of the present application, it then obtains possible combined risk dimensions through the risk attribution solution and makes a decision.

[0333] For example, the number of anti-fraud applications for product A has surged, and the total amount of account opening expenditure S has increased significantly.

[0334] Part 2: Real-time attribution analysis to obtain conclusions.

[0335] The risk attribution module of this solution suggests possible risk dimensions: customer group R3 under channel U1, customer group R5 under channel U3, and customer group R2 under channel U4, which are represented by the combined dimensions U1R3, U3R5, and U4R2 respectively.

[0336] Step 3: Real-time rule verification.

[0337] In actual production, real-time rule verification may have the risk of affecting customer experience or causing mishandling, so real-time rule verification is usually performed using a bypass.

[0338] First, limit the flow of U1R3. If the total amount of account opening and expenditure S does not change significantly, then restore U1R3, and then limit the flow of U3R5. If the overall expenditure is restored, then the final risk dimension of U3R5.

[0339] Step 4: Final rules take effect.

[0340] The response measures for U3R5 will be released to real production and the disposal measures will take effect.

[0341] Step 5: Real-time early warning and disposal.

[0342] In the fourth step, the business rules that meet business expectations have been verified and applied to the latest transactions through generation and publication, thus achieving an overall closed loop.

[0343] This embodiment of the present application has the following technical effects:

[0344] First, a customized algorithm solves the problem of simple and inflexible early warning and discovery methods in related technologies. It also fully considers the impact of time and space on transactions. Financial transactions are significantly affected by holidays and regions. Usually, pre-holiday loans and transactions in hot cities are significantly higher than those on weekdays and in third- and fourth-tier cities. It can globally monitor the stability of compatible transactions and discover related risks.

[0345] Secondly, intelligent risk dimension analysis can accurately indicate the subdivided risk dimensions of abnormal changes and provide real-time reference for the business. This can greatly reduce the analysis time of the business and provide timeliness for risk prevention.

[0346] Thirdly, the platform for real-time business rule verification can verify the limitations of rules in real time and perform real-time online optimization of the rules to improve the timeliness and accuracy of rule application.

[0347] In order to implement the above data processing method, a data processing device according to an embodiment of the present application is provided below in combination with Figure 13 The structural diagram of the data processing device is shown in FIG.

[0348] like Figure 13 As shown, the data processing device 130 includes: a collection unit 1301, a processing unit 1302, a query unit 1303, a determination unit 1304 (for the sake of distinction, it can also be called the first determination unit 1304) and an output unit 1305.

[0349] The collection unit 1301 is configured to collect transaction information of the first business in real time during the transaction process; the transaction information at least includes address information representing the area where the transaction occurs;

[0350] The processing unit 1302 is configured to process the transaction information of the first business based on the address information to obtain a value of a set indicator of the first business in each of the at least two areas in a current period;

[0351] A query unit 1303 is configured to query a value of a set indicator of the first service in each of at least two areas in at least one comparison period;

[0352] a determining unit 1304 configured to determine a risk coefficient of the first service based on a value of a set indicator of the first service in each of the areas in a current period and a value of a set indicator of the first service in each of the areas in at least one comparison period;

[0353] The output unit 1305 is configured to output a risk warning when the risk coefficient is greater than or equal to a risk threshold.

[0354] In some embodiments, the processing unit 1302 is specifically configured to:

[0355] Parsing the address information in the transaction information to determine the region to which the address information belongs;

[0356] Based on the transaction time in the transaction information, determining that the transaction time belongs to the current period;

[0357] The transaction information of the first business is statistically analyzed with region, cycle and set index as statistical dimensions to obtain the value of the set index of the first business in each of the regions in the current cycle.

[0358] In some embodiments, the determining unit 1304 is configured to:

[0359] For each of the at least two areas, determining a first proportion of each area in the current period based on the value of the set indicator of the first service in each area in the current period; the first proportion is used to represent a ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas;

[0360] For each of the at least two areas, determining a second proportion for each area in each comparison period included in the at least one comparison period based on the value of the set indicator of the first service in each area in the at least one comparison period; the second proportion is used to represent a ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas;

[0361] For each comparison period included in the at least one comparison period, determining a difference coefficient between the current period and the comparison period based on the first proportion of each of the regions in the current period, the second proportion of each of the regions in the comparison period, and a first proportion threshold, to obtain at least one difference coefficient;

[0362] A risk coefficient of the first business is determined based on the at least one variance coefficient.

[0363] In some embodiments, the data processing device 130 may further include a second determining unit, which is configured to:

[0364] Determine a first array for the current cycle and a second array for the comparison cycle; the first array includes a first proportion of each of the at least two regions in the current cycle, and the second array includes a second proportion of each of the at least two regions in the comparison cycle; the data in the first array and the second array satisfy a first order;

[0365] constructing a first curve based on the first array, and constructing a second curve based on the second array;

[0366] determining a first reference value in the first array that satisfies a first condition based on the first curve, and determining a second reference value in the second array that satisfies the first condition based on the second curve;

[0367] The first proportion threshold is determined based on at least the first reference value and the second reference value.

[0368] In some embodiments, the second determining unit is further configured to:

[0369] determining a third reference value satisfying a second condition in the first array;

[0370] determining a fourth reference value satisfying the second condition in the second array;

[0371] The first proportion threshold is determined to be a minimum value among the first reference value, the second reference value, the third reference value, and the fourth reference value.

[0372] In some embodiments, the data processing device 130 may further include a risk dimension analysis unit, the risk dimension analysis unit being configured to: determine n risk dimensions of the first business, and divide the n risk dimensions into m groups in descending order of risk weight; n is an integer greater than 1; n is greater than m, and a group includes at least one risk dimension;

[0373] Replaying the transactions of the first business, and selecting m×x first-level risk dimensions from the first group of risk dimensions based on the values ​​of the set indicators, where x is an integer greater than 1;

[0374] Combine the (m-i+1)×x risk dimensions in the i-th group with all dimensions in the i+1-th group to obtain h combinations of risk dimensions; replay the transactions of the first business, and based on the values ​​of the set indicators, select (mi)×x combinations of risk dimensions from the h combinations of risk dimensions; i and h are integers greater than 1;

[0375] Increment i by 1 and re-execute the process of combining the (m-i+1)×x risk dimensions of the i-th group with all dimensions in the i+1-th group to obtain h risk dimension combinations; replay the transaction of the first business and, based on the value of the set indicator, select (mi)×x risk dimension combinations from the h risk dimension combinations; until x risk dimension combinations consisting of m risk dimensions are obtained;

[0376] Based on the values ​​of the set indicators, at least one warning risk dimension combination is determined among the m×x first-level risk dimensions, the combination of the (mi)×x risk dimensions, and the x risk dimension combinations consisting of m risk dimensions.

[0377] In some embodiments, the data processing device 130 may further include a verification unit configured to:

[0378] Determining a warning risk dimension combination to be verified in the at least one warning risk dimension combination;

[0379] limiting the early warning risk dimension combinations other than the early warning risk dimension combination to be verified in the at least one risk dimension combination, so as to perform transaction verification by using the bypass of the first business based on the early warning risk dimension combination to be verified;

[0380] Collecting the values ​​of the indicators set during the transaction verification process;

[0381] When the value of the set indicator satisfies the third condition, the warning risk dimension combination to be verified is determined to be the target risk dimension combination.

[0382] It should be noted that the data processing device provided in the embodiment of the present application includes the various units included, which can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0383] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0384] It should be noted that, in the embodiment of the present application, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0385] To implement the above-mentioned data processing method, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the data processing method provided in the above-mentioned embodiment are implemented.

[0386] The following combination Figure 14 The electronic device 140 shown is a block diagram of the electronic device.

[0387] In one example, the electronic device 140 may be the electronic device described above. Figure 14As shown, the electronic device 140 includes: a processor 1401, at least one communication bus 1402, a user interface 1403, at least one external communication interface 1404, and a memory 1405. The communication bus 1402 is configured to enable communication between these components. The user interface 1403 may include a display screen, and the external communication interface 1404 may include a standard wired interface and a wireless interface.

[0388] The memory 1405 is configured to store instructions and applications executable by the processor 1401, and can also cache data to be processed or processed by the processor 1401 and various modules in the electronic device (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (RAM).

[0389] In a fourth aspect, an embodiment of the present application provides a storage medium, that is, a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the data processing method provided in the above embodiment are implemented.

[0390] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0391] It should be understood that “one embodiment” or “an embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, “in one embodiment” or “in some embodiments” appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0392] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0393] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0394] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0395] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0396] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0397] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0398] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: The method comprises: collecting transaction information of the first business in a transaction process in real time; the transaction information at least includes address information representing the area where the transaction occurs; Based on the address information, the transaction information of the first business is processed to obtain a value of a set indicator of the first business in each of the at least two areas in a current period; querying a value of a set indicator of the first service in each of at least two areas in at least one comparison period; Determining a risk coefficient of the first business based on a value of a set indicator of the first business in each of the areas in a current period and a value of a set indicator of the first business in each of the areas in at least one comparison period; When the risk coefficient is greater than or equal to a risk threshold, outputting a risk warning; processing the transaction information of the first business based on the address information to obtain a value of a set indicator of the first business in each of the at least two areas in the current period, including: Parsing the address information in the transaction information to determine the region to which the address information belongs; Based on the transaction time in the transaction information, determining that the transaction time belongs to the current period; The transaction information of the first business is statistically analyzed with region, cycle and set index as statistical dimensions to obtain the value of the set index of the first business in each of the regions in the current cycle.

2. The method according to claim 1, characterized in that The determining the risk coefficient of the first business based on the value of the set indicator of the first business in each area in the current period and the value of the set indicator of the first business in each area in at least one comparison period includes: For each of the at least two areas, determining a first proportion of each area in the current period based on the value of the set indicator of the first service in each area in the current period; the first proportion is used to represent a ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas; For each of the at least two areas, determining a second proportion for each area in each comparison period included in the at least one comparison period based on the value of the set indicator of the first service in each area in the at least one comparison period; the second proportion is used to represent a ratio of the value of the set indicator of each area to the sum of the values ​​of the set indicators of the at least two areas; For each comparison period included in the at least one comparison period, determining a difference coefficient between the current period and the comparison period based on the first proportion of each of the regions in the current period, the second proportion of each of the regions in the comparison period, and a first proportion threshold, to obtain at least one difference coefficient; A risk coefficient of the first business is determined based on the at least one variance coefficient.

3. The method according to claim 2, characterized in that The method further comprises: Determine a first array for the current cycle and a second array for the comparison cycle; the first array includes a first proportion of each of the at least two regions in the current cycle, and the second array includes a second proportion of each of the at least two regions in the comparison cycle; the data in the first array and the second array satisfy a first order; constructing a first curve based on the first array, and constructing a second curve based on the second array; determining a first reference value in the first array that satisfies a first condition based on the first curve, and determining a second reference value in the second array that satisfies the first condition based on the second curve; The first proportion threshold is determined based on at least the first reference value and the second reference value.

4. The method according to claim 3, characterized in that The determining the first proportion threshold based at least on the first reference value and the second reference value includes: determining a third reference value satisfying a second condition in the first array; determining a fourth reference value satisfying the second condition in the second array; The first proportion threshold is determined to be a minimum value among the first reference value, the second reference value, the third reference value, and the fourth reference value.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determine n risk dimensions of the first business, and divide the n risk dimensions into m groups in descending order of risk weight; n is an integer greater than 1; n is greater than m, and one group includes at least one risk dimension; Replaying the transactions of the first business, and selecting m×x first-level risk dimensions from the first group of risk dimensions based on the values ​​of the set indicators, where x is an integer greater than 1; Combine the (m-i+1)×x risk dimensions in the i-th group with all dimensions in the i+1-th group to obtain h combinations of risk dimensions; replay the transactions of the first business, and based on the values ​​of the set indicators, select (mi)×x combinations of risk dimensions from the h combinations of risk dimensions; i and h are integers greater than 1; Increment i by 1 and re-execute the process of combining the (m-i+1)×x risk dimensions of the i-th group with all dimensions in the i+1-th group to obtain h risk dimension combinations; replay the transaction of the first business and, based on the value of the set indicator, select (mi)×x risk dimension combinations from the h risk dimension combinations; until x risk dimension combinations consisting of m risk dimensions are obtained; Based on the values ​​of the set indicators, at least one warning risk dimension combination is determined among the m×x first-level risk dimensions, the combination of the (mi)×x risk dimensions, and the x risk dimension combinations consisting of m risk dimensions.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determining a warning risk dimension combination to be verified in the at least one warning risk dimension combination; limiting the early warning risk dimension combinations other than the early warning risk dimension combination to be verified in the at least one risk dimension combination, so as to perform transaction verification by using the bypass of the first business based on the early warning risk dimension combination to be verified; Collecting the values ​​of the indicators set during the transaction verification process; When the value of the set indicator satisfies the third condition, the warning risk dimension combination to be verified is determined to be the target risk dimension combination.

7. A data processing device, characterized in that: The device comprises: a collection unit, configured to collect transaction information of the first business in real time during the transaction process; the transaction information at least including address information representing the area where the transaction occurs; a processing unit, configured to process the transaction information of the first business based on the address information, and obtain a value of a set indicator of the first business in each of the at least two areas in a current period; a query unit, configured to query a value of a set indicator of the first service in each of at least two areas in at least one comparison period; a determining unit, configured to determine a risk coefficient of the first business based on a value of a set indicator of the first business in each of the areas in a current period and a value of a set indicator of the first business in each of the areas in at least one comparison period; An output unit, configured to output a risk warning when the risk coefficient is greater than or equal to a risk threshold; The processing unit is specifically used to: Parsing the address information in the transaction information to determine the region to which the address information belongs; Based on the transaction time in the transaction information, determining that the transaction time belongs to the current period; The transaction information of the first business is statistically analyzed with region, cycle and set index as statistical dimensions to obtain the value of the set index of the first business in each of the regions in the current cycle.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the data processing method according to any one of claims 1 to 6 is implemented.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and apparatus for evaluating fraud risk in an electronic commerce transaction

    US20020194119A1