Risk assessment method, device and system based on business data, equipment and medium
By acquiring business data sets, determining the business service type, detecting abnormal data using local outlier algorithms, and evaluating risk levels using third-party reference data, the problem of low risk assessment accuracy for small businesses or individual customers is solved, and a higher accuracy risk assessment is achieved.
Patent Information
- Application Number
- CN202510918707.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Small businesses or individual customers lack sufficient credit history and information transparency, which makes the business data they collect missing, resulting in low risk assessment accuracy and a single business service rule that is difficult to adapt to different business service types, further reducing the accuracy of assessment.
By obtaining the user's business data set, determining the business service type, using local outlier algorithm to detect abnormal data, obtaining reference data sets of third-party institutions, and evaluating the risk level based on the business service type, and using the differentiated data for comprehensive evaluation.
It improves the accuracy of risk assessment, reduces the deviation of assessment, conforms to the actual situation of the business, and improves the accuracy of assessment.
Smart Images

Figure CN120494528A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of business risk assessment and processing, and in particular to a risk assessment method, apparatus, system, equipment and medium based on business data. Background Art
[0002] With the development of technology and the transformation towards informatization and digitalization, businesses are rapidly expanding their online businesses, enriching their online products and services, and increasing user demand for products. As business scenarios become increasingly complex, businesses and institutions need to conduct risk assessments to mitigate the potential for risk events when handling various financial services, thereby safeguarding the interests of both businesses and users.
[0003] One of the commonly used methods is to collect the customer's business data and determine the business service rules, match the acquired business data with the business service rules, and then determine whether there is a risk and the risk level based on the matching results.
[0004] However, the currently commonly used methods have the following technical problems: small businesses or individual customers usually lack sufficient credit history and information transparency, which leads to incomplete business data collected, resulting in large deviations from actual data in subsequent assessments and low accuracy; and as the types of business services continue to increase, a single business service rule is difficult to adapt to different business services, further reducing the accuracy of the assessment. Summary of the Invention
[0005] In view of the above problems, the present application is proposed to provide a business data-based risk assessment method, apparatus, system, device, and medium that overcomes or at least partially solves the above problems, including: A risk assessment method based on business data, the method comprising: After obtaining the user's business data set, determining the corresponding business service type according to the business data set, wherein the business data set includes multiple business data; Determining whether the business data set contains abnormal data based on a local outlier algorithm; If it is determined that the business data set contains abnormal data, obtaining a reference data set, wherein the reference data is a data set obtained from a third-party organization where users complete services corresponding to the business service type; Differentiating data between the business dataset and the reference dataset is determined based on the business service type, and the risk level is evaluated using the differentiating data.
[0006] In a possible implementation, determining whether the service data set contains abnormal data based on a local outlier algorithm includes: Converting each item of business data in the business data set into a pre-examination format to obtain data in multiple formats; After determining the corresponding category of each format data, adding each format data to a preset data set of the corresponding category; Calculate the density value of each data of the format data and the preset data set based on the local outlier algorithm; If any of the density values is greater than a preset value, it is determined that the business data set contains abnormal data; If each of the density values is less than a preset value, it is determined that the business data set does not contain abnormal data.
[0007] In a possible implementation, the calculating the density value of the format data and each data of the preset data set based on the local outlier algorithm includes: Adding the format data and each data of the preset data set to a preset coordinate system to obtain a data coordinate system, wherein the preset coordinate system is a coordinate system constructed according to the values of the preset data set; The distance value between the format data and adjacent data in the data coordinate system is calculated, and an average value of a plurality of the distance values is calculated to obtain a density value.
[0008] In a possible implementation, determining the difference data between the business dataset and the reference dataset based on the business service type includes: determining a plurality of data items based on the item text of the business service type; Extracting corresponding data of each data item from the business data set and the reference data set respectively; Determine the outlier value of the two data corresponding to each data item to obtain distinguished data.
[0009] In a possible implementation, determining the corresponding business service type according to the business data set includes: Mapping each business data item in the business data set to a preset business tree diagram to obtain a mapping tree diagram; Determine the names corresponding to the mapped nodes in the mapping tree diagram to obtain multiple business process names; The business service type is obtained by matching the texts of the multiple business process names with corresponding types in multiple historical business types recorded in a preset database.
[0010] In a possible implementation, the using the distinguishing data to assess the risk level includes: Determining a risk weight value according to the data item corresponding to the distinguishing data; Calculating a risk assessment value using the numerical value of the distinguishing data and the corresponding risk weight value; The risk level is determined according to the magnitude of the risk assessment value.
[0011] A risk assessment device based on business data, comprising: A type determination module is used to determine a corresponding business service type according to a business data set after obtaining a business data set of a user, wherein the business data set includes multiple business data; An abnormal data module, used to determine whether the business data set contains abnormal data based on a local outlier algorithm; A data acquisition module is configured to acquire a reference data set if it is determined that the business data set contains abnormal data, wherein the reference data set is a data set of users completing services corresponding to the business service type obtained from a third-party organization; The risk assessment module is configured to determine difference data between the business data set and the reference data set based on the business service type, and assess a risk level using the difference data.
[0012] In a possible implementation, determining whether the service data set contains abnormal data based on a local outlier algorithm includes: Converting each item of business data in the business data set into a pre-examination format to obtain data in multiple formats; After determining the corresponding category of each format data, adding each format data to a preset data set of the corresponding category; Calculate the density value of each data of the format data and the preset data set based on the local outlier algorithm; If any of the density values is greater than a preset value, it is determined that the business data set contains abnormal data; If each of the density values is less than a preset value, it is determined that the business data set does not contain abnormal data.
[0013] In a possible implementation, the calculating the density value of the format data and each data of the preset data set based on the local outlier algorithm includes: Adding the format data and each data of the preset data set to a preset coordinate system to obtain a data coordinate system, wherein the preset coordinate system is a coordinate system constructed according to the values of the preset data set; The distance value between the format data and adjacent data in the data coordinate system is calculated, and an average value of a plurality of the distance values is calculated to obtain a density value.
[0014] In a possible implementation, determining the difference data between the business dataset and the reference dataset based on the business service type includes: determining a plurality of data items based on the item text of the business service type; Extracting corresponding data of each data item from the business data set and the reference data set respectively; Determine the outlier value of the two data corresponding to each data item to obtain distinguished data.
[0015] In a possible implementation, determining the corresponding business service type according to the business data set includes: Mapping each business data item in the business data set to a preset business tree diagram to obtain a mapping tree diagram; Determine the names corresponding to the mapped nodes in the mapping tree diagram to obtain multiple business process names; The business service type is obtained by matching the texts of the multiple business process names with corresponding types in multiple historical business types recorded in a preset database.
[0016] In a possible implementation, the using the distinguishing data to assess the risk level includes: Determining a risk weight value according to the data item corresponding to the distinguishing data; Calculating a risk assessment value using the numerical value of the distinguishing data and the corresponding risk weight value; The risk level is determined according to the magnitude of the risk assessment value.
[0017] A risk assessment system based on business data, the system comprising: Data processing platform, multiple business terminals and multiple third-party terminals; The data processing platform is connected to the multiple business terminals and the multiple third-party terminals respectively; The data processing platform is suitable for the risk assessment method based on business data as described above.
[0018] A device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of risk assessment based on business data as described above.
[0019] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of risk assessment based on business data as described above.
[0020] This application has the following advantages: In an embodiment of the present application, after obtaining a user's business data set, the present application can determine the corresponding business service type based on the business data set; determine whether the business data set contains abnormal data based on a local outlier algorithm; if it is determined that the business data set contains abnormal data, obtain a reference data set; determine the difference between the business data set and the reference data set based on the business service type, and use the difference data to assess the risk level. After obtaining the business data set and determining that the data set contains abnormal data, third-party data can be obtained and used to conduct a comprehensive assessment with the data of the current business to determine the deviation of the business data, and then conduct a risk assessment based on the deviation data, thereby adapting to the actual situation of the business, improving the accuracy of the assessment, and reducing the deviation of the assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 This is a flowchart of a risk assessment method based on business data provided by an embodiment of the present application; Figure 2 This is a structural block diagram of a risk assessment device based on business data provided by an embodiment of the present application; Figure 3 This is a structural block diagram of a risk assessment system based on business data provided by an embodiment of the present application; Figure 4 It is a structural diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0023] To make the objectives, features, and advantages of this application more readily apparent, the present application is further described below in conjunction with the accompanying drawings and specific embodiments. It is apparent that the embodiments described are only a portion of the embodiments of this application, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments in this application without inventive effort are also within the scope of protection of this application.
[0024] With the development of technology and the transformation towards informatization and digitalization, businesses are rapidly expanding their online businesses, enriching their online products and services, and increasing user demand for products. As business scenarios become increasingly complex, businesses and institutions need to conduct risk assessments to mitigate the potential for risk events when handling various financial services, thereby safeguarding the interests of both businesses and users.
[0025] One of the commonly used methods is to collect the customer's business data and determine the business service rules, match the acquired business data with the business service rules, and then determine whether there is a risk and the risk level based on the matching results.
[0026] However, the currently commonly used methods have the following technical problems: small businesses or individual customers usually lack sufficient credit history and information transparency, which leads to incomplete business data collected, resulting in large deviations from actual data in subsequent assessments and low accuracy; and as the types of business services continue to increase, a single business service rule is difficult to adapt to different business services, further reducing the accuracy of the assessment.
[0027] In order to solve the above technical problems, refer to Figure 1 , showing a flowchart of the steps of a risk assessment method based on business data provided by an embodiment of the present application; In one embodiment, the business data-based risk assessment method is applicable to an enterprise's online management system or business processing system, which may be an online data processing platform.
[0028] In one application scenario, the business data-based risk assessment method is applicable to businesses that engage in online transactions, including banks, insurance companies, or online shopping malls.
[0029] As an example, the risk assessment method based on business data may include: S11. After obtaining a business data set of a user, determining a corresponding business service type according to the business data set, wherein the business data set includes multiple business data.
[0030] In one embodiment, a user performs business processing on an online platform. A business processing involves multiple service contents. Each service content is a step and each service content can correspond to a data. A data corresponding to each service content can be obtained, thereby obtaining multiple business data of the user; finally, the multiple business data can be combined into a data set to obtain a business data set.
[0031] After obtaining the user's business data set, we can determine the service type corresponding to this business processing and perform risk assessment based on the service type. Different service types have different risk assessment requirements or conditions. Subsequently, we can conduct assessment and processing based on the risk assessment requirements or conditions corresponding to the service type, so as to fit the user's actual situation and improve the assessment accuracy.
[0032] In one embodiment, different enterprises have multiple business service types, and each business service type corresponds to a business process with multiple steps. In order to determine the corresponding business service type according to the steps of the business process, as an example, step S11 may include the following sub-steps: S111 , mapping each business data item in the business data set to a preset business tree diagram to obtain a mapping tree diagram.
[0033] S112: Determine the names corresponding to the mapped nodes in the mapping tree diagram to obtain multiple business process names.
[0034] S113 . Match the texts of the multiple business process names with corresponding types in multiple historical business types recorded in a preset database to obtain a business service type.
[0035] In one embodiment, a business tree diagram may be prepared in advance, where each branch of the business tree diagram corresponds to a business process, and each node of each branch is a step of the business process. The data and related content required for the step may be recorded at each node.
[0036] After obtaining the business data set, you can obtain the items and corresponding steps of each business data in the business data set, and then map each business data to a preset business tree diagram according to the steps and items corresponding to the data, thereby obtaining a mapping tree diagram.
[0037] Next, you can identify the names of the mapped nodes in the mapping tree to obtain multiple business process names. Each business process name represents a business step. For example, a user's insurance refund involves steps such as request, identity verification, refund item confirmation, refund amount review, and result notification. Each step corresponds to a business process name.
[0038] Then, text matching can be performed based on the text of each business process name and the text of multiple historical business types recorded in the preset database. When the text of each business process name is the same as each text of a historical business type, it can be determined that the type corresponding to the historical business type is the business service type.
[0039] Assume that there are 5 business process names, and historical business type a also has 5 step texts. The names of the 5 business process names correspond one-to-one with the 5 step texts of historical business type a. Then the type corresponding to historical business type a is the business service type.
[0040] In one embodiment, assuming that there are 5 business process names, historical business type a may have texts of 6 steps, and the names of the 5 business process names correspond one-to-one to the texts of 5 of the steps of historical business type a; historical business type b may have texts of 7 steps, and the names of the 5 business process names correspond one-to-one to the texts of 5 of the steps of historical business type b.
[0041] At this time, you can filter historical business type a and historical business type b, and display historical business type a and historical business type b to the business personnel, so that the business personnel can select one of the two historical business types as the business service type.
[0042] S12. Determine whether the business data set contains abnormal data based on a local outlier algorithm.
[0043] In one embodiment, after determining the business service type, it may be determined whether the business data set contains abnormal data. If it contains abnormal data, there may be risks in this business.
[0044] Optionally, you can use different algorithms to detect whether the data in the dataset is abnormal. For example, the 3sigma algorithm (based on the normal distribution, considers data exceeding 3sigma as abnormal data), the Z-score algorithm (calculate the distance between the data and the mean, and a Z-score of 3 is considered abnormal data), the Boxplot algorithm (finds abnormal data based on the interquartile range (IQR)), and the Isolation Forest algorithm ("isolates" abnormal data by randomly selecting features and split points).
[0045] In order to accurately determine whether a business data set contains abnormal data, the local outlier algorithm can be used to determine whether each business data item in the business data set is an outlier. If the data is an outlier, then the data is abnormal data; conversely, if the data is not an outlier, then the data is normal data.
[0046] As an example, the determining whether the business data set contains abnormal data based on the local outlier algorithm may include the following sub-steps: S121 : Convert each business data item of the business data set into a pre-examination format to obtain data in multiple formats.
[0047] S122: After determining the corresponding category of each format data, add each format data to a preset data set of the corresponding category.
[0048] S123: Calculate density values of the format data and each data of the preset data set based on a local outlier algorithm.
[0049] S124: If any one of the density values is greater than a preset value, it is determined that the business data set contains abnormal data.
[0050] S125: If each density value is less than a preset value, determine that the business data set does not contain abnormal data.
[0051] In one operation mode, since the formats of business data corresponding to different steps may be different, in order to uniformly process multiple business data, each business data may be converted into a pre-examination format to obtain multiple format data of the same format.
[0052] Next, determine the corresponding category of each format data. Specifically, determine the category of the format data in the preset business tree node, and then obtain the data set of the node, which contains data of other different users of the node who handle the same business, to obtain the preset data set.
[0053] Then, each format data may be added to the preset data set of its corresponding category, and the density value (specifically, the LOF value) of the format data and each data in the preset data set may be calculated using the local outlier algorithm.
[0054] If the density value is greater than the preset value, it means that the format data is an outlier in the preset data set of its corresponding category. Conversely, if the density value is less than the preset value, it means that the format data is not an outlier in the preset data set of its corresponding category.
[0055] If any density value is greater than a preset value, it is determined that the business data set contains abnormal data, and the abnormal data is format data with a density value greater than the preset value.
[0056] If each density value is less than a preset value, it is determined that the business data set does not contain abnormal data.
[0057] In one embodiment, in order to fully utilize the information in the data, as an example, the calculation of the density value of the format data and each data of the preset data set based on the local outlier algorithm may include the following sub-steps: S1231. Add the format data and each data of the preset data set to a preset coordinate system to obtain a data coordinate system, where the preset coordinate system is a coordinate system constructed according to the values of the preset data set.
[0058] S1232: Calculate the distance value between the format data and adjacent data in the data coordinate system, and calculate the average value of multiple distance values to obtain a density value.
[0059] In one embodiment, the preset data set may be preprocessed before adding the formatted data to the preset data set. In one operation mode, the preprocessing may be data cleaning, in which outlier data or abnormal data in the preset data set is treated as noise and removed. By removing abnormal data from the preset data set, the abnormal data in the preset data set is prevented from affecting the formatted data. Moreover, from the perspective of data information mining, certain outlier data in the preset data set may contain valuable information that can help discover interesting phenomena and promote progress in practical applications. Therefore, during the data processing process, it is necessary to balance the retention or removal of outliers based on the specific situation to fully utilize the information in the data.
[0060] After the preprocessing is completed, the format data and each data of the preset data set can be added to the preset coordinate system. The preset coordinate system is a coordinate system constructed based on the values of the preset data set that has completed the preprocessing.
[0061] Specifically, the method for constructing the preset coordinate system can be to obtain the first difference and the second difference of each node from the preset data set, wherein the first difference is the difference between the data collected at time t and the data collected at time t-1, and the second difference is the difference between the data collected at time t-1 and the data collected at time t-2; a two-dimensional coordinate system is generated with the first difference as the horizontal coordinate and the second difference as the vertical coordinate to obtain the preset coordinate system.
[0062] The format data and each data of the preset dataset can be added to the preset coordinate system.
[0063] Finally, the distance value between the format data and the adjacent data can be calculated in the data coordinate system, and then the average of multiple distance values can be calculated to obtain the density value.
[0064] In an optional embodiment, abnormal data may also be detected using a pre-trained anomaly detection model.
[0065] In one embodiment, the business data set may be input into a preset anomaly detection model to determine whether the business data set contains abnormal data.
[0066] It should be noted that the preset anomaly detection model can be an anomaly detection model generated by inputting a training data set into an RNN recurrent neural network for training.
[0067] The anomaly detection model training operation may include the following steps: The first step is to obtain the training data set. The input training data set can be a feature sequence within a time window, which can be shown in the following formula: .
[0068] In the above formula, t represents the time step. At time step t=0, the hidden state h(0) is initialized to a zero vector or a random vector.
[0069] The second step is to calculate the hidden state, where h(t) is the hidden state at time step t, X(t) is the input at time step t, W and U are weight matrices, b is the bias term, and tanh is the hyperbolic tangent activation function. The specific formula is as follows: ; The third step is to calculate the output result. The output result is o(t), which is the model output at time step t, W out is the weight matrix of the output layer, b out is the bias term of the output layer. It is shown in the following formula: .
[0070] S13. If it is determined that the business data set contains abnormal data, a reference data set is obtained, wherein the reference data set is a data set obtained from a third-party organization in which users complete services corresponding to the business service type.
[0071] After it is determined that the business dataset contains one or more abnormal data, a reference dataset can be obtained.
[0072] The reference dataset can be a collection of data obtained from a third-party organization, representing users completing services corresponding to the business service type. By obtaining the reference dataset from a third-party organization, the reference dataset can be used as a reference to determine whether the business dataset contains one or more abnormal data with significant deviations, thereby determining a specific risk level.
[0073] In an optional embodiment, a reference data set of services corresponding to the same business service type completed by the same user at a third-party organization may be obtained.
[0074] It is also possible to obtain a reference data set of services corresponding to the same business service type completed by different users at a third-party organization.
[0075] For example, a business dataset might be a dataset of user A purchasing Section XX compulsory auto insurance. A reference dataset might be a dataset of user A purchasing the same Section XX compulsory auto insurance from a third-party institution. A reference dataset might also be a dataset of user B purchasing the same Section XX compulsory auto insurance from a third-party institution.
[0076] S14: Determine difference data between the business dataset and the reference dataset based on the business service type, and use the difference data to evaluate the risk level.
[0077] In one embodiment, since the reference dataset can determine what specific data the user needs to complete the business service type, the numerical deviation of the reference dataset and the business dataset can be compared, and the numerical deviation of the two datasets can be used to determine any abnormalities in the local business service type, and then the risk of this business can be determined based on the data deviation.
[0078] For example, the business service this time is a refund service. The refund the user completes at a third-party agency is 1xx yuan, and the refund applied for this time is 1xxxx yuan, a difference of 100 times. The difference between the two is large, and the risk is high.
[0079] For another example, the business service this time is to purchase annuity insurance services. The user pays 3xxx yuan for the annuity insurance at a third-party institution, and the payment for this application to purchase annuity insurance is 2xxx yuan, which is about 1.5 times the difference. The difference between the two is small, and the risk is small.
[0080] In one embodiment, in order to accurately determine the difference value between the two data sets, as an example, determining the difference data between the business data set and the reference data set based on the business service type may include the following sub-steps: S141. Determine multiple data items based on the project text of the business service type.
[0081] S142: Extract corresponding data of each data item from the business data set and the reference data set respectively.
[0082] S143: Determine the outlier value of the two data corresponding to each data item to obtain distinguished data.
[0083] In a specific operation, the text of each item of the business service type can be determined to obtain the item text. Each item is the service content of the business service. The corresponding data item column can be determined based on the item text.
[0084] Next, the corresponding data of each data item can be extracted from the business data set according to the data item column. Similarly, the corresponding data of each data item can also be extracted from the reference data set according to the data item column.
[0085] Finally, the difference between the two corresponding data of each data item can be calculated to obtain the abnormal value of each data item, thereby obtaining the distinguishing data.
[0086] In one embodiment, there are multiple data items, and different data items have different levels of importance. For example, if the data item is identity information, its importance is higher. If the data item is time, its importance is lower. As an example, the risk level assessment using the distinguishing data may include the following sub-steps: S144. Determine a risk weight value according to the data item corresponding to the distinguishing data.
[0087] S145. Calculate a risk assessment value using the numerical value of the distinguishing data and the corresponding risk weight value.
[0088] S146. Determine the risk level according to the risk assessment value.
[0089] In one operation mode, a weight value of each data item included in the distinguishing data can be determined to obtain a risk weight value. It should be noted that a technician can set a weight value for each data item in advance.
[0090] Next, the risk assessment value may be calculated using the numerical value corresponding to each data item of the distinguishing data and its corresponding risk weight value.
[0091] In one operation mode, the risk assessment value can be calculated as follows: ; In the above formula, c is the risk assessment value, A i is the value of the i-th data item, A k is the value of the kth data item, A l is the value of the lth data item, P i is the risk weight value of the i-th data item, P k is the risk weight value of the kth data item, P l is the risk weight value of the lth data item.
[0092] Finally, the risk level can be determined based on the numerical value of the risk assessment value.
[0093] For example, when the risk assessment value is greater than 10, the risk level may be level one; when the risk assessment value is greater than 10, the risk level may be level two; and when the risk assessment value is greater than 30, the risk level may be level three.
[0094] By determining the risk level, multi-level alarm prompts can be processed according to the risk level.
[0095] For example, an alarm at the first-level risk level can issue an early warning to alert relevant personnel; Level 2 risk level alarm: issues a high-risk warning, prompts emergency measures, and suspends business services; Level 3 risk alarm: Immediately issue an emergency alarm and notify relevant personnel.
[0096] In this embodiment, the embodiment of the present application provides a risk assessment method based on business data, which has the following beneficial effects: after obtaining the user's business data set, the present application can determine the corresponding business service type based on the business data set; determine whether the business data set contains abnormal data based on the local outlier algorithm; if it is determined that the business data set contains abnormal data, obtain a reference data set; determine the difference data between the business data set and the reference data set based on the business service type, and use the difference data to assess the risk level. After obtaining the business data set and determining that the data set contains abnormal data, third-party data can be obtained, and a comprehensive assessment can be performed using the third-party data and the data of this business to determine the deviation of the business data, and then a risk assessment can be performed based on the deviation data, thereby adapting to the actual situation of the business, improving the accuracy of the assessment, and reducing the deviation of the assessment. As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0097] Reference Figure 2 , shows a structural block diagram of a risk assessment device based on business data provided by an embodiment of the present application; Specifically include: A type determination module 201 is configured to determine a corresponding business service type based on a business data set after obtaining a business data set of a user, wherein the business data set includes multiple business data; An abnormal data module 202 is used to determine whether the business data set contains abnormal data based on a local outlier algorithm; The data acquisition module 203 is configured to acquire a reference data set if it is determined that the business data set contains abnormal data, wherein the reference data set is a data set of users completing services corresponding to the business service type acquired from a third-party organization; The risk assessment module 204 is configured to determine difference data between the business dataset and the reference dataset based on the business service type, and assess a risk level using the difference data.
[0098] Optionally, determining whether the business data set contains abnormal data based on a local outlier algorithm includes: Converting each item of business data in the business data set into a pre-examination format to obtain data in multiple formats; After determining the corresponding category of each format data, adding each format data to a preset data set of the corresponding category; Calculate the density value of each data of the format data and the preset data set based on the local outlier algorithm; If any of the density values is greater than a preset value, it is determined that the business data set contains abnormal data; If each of the density values is less than a preset value, it is determined that the business data set does not contain abnormal data.
[0099] Optionally, the calculating the density value of the format data and each data of the preset data set based on the local outlier algorithm includes: Adding the format data and each data of the preset data set to a preset coordinate system to obtain a data coordinate system, wherein the preset coordinate system is a coordinate system constructed according to the values of the preset data set; The distance value between the format data and adjacent data in the data coordinate system is calculated, and an average value of a plurality of the distance values is calculated to obtain a density value.
[0100] Optionally, the determining, based on the business service type, distinguishing data between the business dataset and the reference dataset includes: determining a plurality of data items based on the item text of the business service type; Extracting corresponding data of each data item from the business data set and the reference data set respectively; Determine the outlier value of the two data corresponding to each data item to obtain distinguished data.
[0101] Optionally, determining a corresponding business service type according to the business data set includes: Mapping each business data item in the business data set to a preset business tree diagram to obtain a mapping tree diagram; Determine the names corresponding to the mapped nodes in the mapping tree diagram to obtain multiple business process names; The business service type is obtained by matching the texts of the multiple business process names with corresponding types in multiple historical business types recorded in a preset database.
[0102] Optionally, the using the distinguishing data to assess the risk level includes: Determining a risk weight value according to the data item corresponding to the distinguishing data; Calculating a risk assessment value using the numerical value of the distinguishing data and the corresponding risk weight value; The risk level is determined according to the magnitude of the risk assessment value.
[0103] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0104] Reference Figure 3 , shows a structural block diagram of a risk assessment system based on business data provided by an embodiment of the present application; Specifically including: data processing platform, multiple business terminals and multiple third-party terminals; The data processing platform is connected to the multiple business terminals and the multiple third-party terminals respectively; The data processing platform is applicable to the risk assessment method based on business data as described in the above embodiment.
[0105] Reference Figure 4 , shows a computer device of a business data-based risk assessment method of the present application, which may specifically include the following: The computer device 12 is a general-purpose computing device. The components of the computer device 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0106] The bus 18 represents one or more of several types of bus 18 structures, including a memory bus 18 or memory controller, a peripheral bus 18, an accelerated graphics port, a processor, or a local bus 18 that utilizes any of a variety of bus 18 architectures. Examples of such architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus 18, a Micro Channel Architecture (MAC) bus 18, an Enhanced ISA bus 18, an Audio Video Electronics Standards Association (VESA) local bus 18, and a Peripheral Component Interconnect (PCI) bus 18.
[0107] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0108] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be configured to read and write to non-removable, non-volatile magnetic media (commonly referred to as a "hard drive"). Although Figure 4 Although not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), as well as an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 18 via one or more data media interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42 configured to perform the functions of various embodiments of the present application.
[0109] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in a memory. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules 42, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0110] The computer device 12 may also communicate with one or more external devices 14 (e.g., a keyboard, a pointing device, a display 24, a camera, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network (e.g., the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via the bus 18. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to microcode, device drivers, redundant processing units 16, external disk drive arrays, RAID systems, tape drives, and data backup storage systems 34.
[0111] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the risk assessment method based on business data provided in the embodiment of the present application.
[0112] That is, when the processing unit 16 executes the above program, the following is achieved: After obtaining the user's business data set, determining the corresponding business service type according to the business data set, wherein the business data set includes multiple business data; Determining whether the business data set contains abnormal data based on a local outlier algorithm; If it is determined that the business data set contains abnormal data, obtaining a reference data set, wherein the reference data is a data set obtained from a third-party organization where users complete services corresponding to the business service type; Differentiating data between the business dataset and the reference dataset is determined based on the business service type, and the risk level is evaluated using the differentiating data.
[0113] In an embodiment of the present application, the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the risk assessment method based on business data as provided in all embodiments of the present application is implemented.
[0114] That is, when the program is executed by the processor: After obtaining the user's business data set, determining the corresponding business service type according to the business data set, wherein the business data set includes multiple business data; Determining whether the business data set contains abnormal data based on a local outlier algorithm; If it is determined that the business data set contains abnormal data, obtaining a reference data set, wherein the reference data is a data set obtained from a third-party organization where users complete services corresponding to the business service type; Differentiating data between the business dataset and the reference dataset is determined based on the business service type, and the risk level is evaluated using the differentiating data.
[0115] Any combination of one or more computer-readable media may be employed. A computer-readable medium may be a computer signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program for use by or in connection with an instruction execution system, apparatus, or device.
[0116] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0117] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet using an Internet service provider). The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced.
[0118] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0119] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0120] The above is a detailed introduction to the risk assessment method, system, device and medium based on business data provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A risk assessment method based on business data, characterized in that: The method comprises: After obtaining the user's business data set, determining the corresponding business service type according to the business data set, wherein the business data set includes multiple business data; Determining whether the business data set contains abnormal data based on a local outlier algorithm; If it is determined that the business data set contains abnormal data, obtaining a reference data set, wherein the reference data is a data set obtained from a third-party organization where users complete services corresponding to the business service type; Differentiating data between the business dataset and the reference dataset is determined based on the business service type, and the risk level is evaluated using the differentiating data.
2. The risk assessment method based on business data according to claim 1, characterized in that: The determining whether the business data set contains abnormal data based on a local outlier algorithm includes: Converting each item of business data in the business data set into a pre-examination format to obtain data in multiple formats; After determining the corresponding category of each format data, adding each format data to a preset data set of the corresponding category; Calculate the density value of each data of the format data and the preset data set based on the local outlier algorithm; If any of the density values is greater than a preset value, it is determined that the business data set contains abnormal data; If each of the density values is less than a preset value, it is determined that the business data set does not contain abnormal data.
3. The risk assessment method based on business data according to claim 2, characterized in that: The calculating the density value of each data of the format data and the preset data set based on the local outlier algorithm includes: Adding the format data and each data of the preset data set to a preset coordinate system to obtain a data coordinate system, wherein the preset coordinate system is a coordinate system constructed according to the values of the preset data set; The distance value between the format data and adjacent data in the data coordinate system is calculated, and an average value of a plurality of the distance values is calculated to obtain a density value.
4. The risk assessment method based on business data according to claim 1, characterized in that: The determining, based on the business service type, the distinguishing data between the business data set and the reference data set includes: determining a plurality of data items based on the item text of the business service type; Extracting corresponding data of each data item from the business data set and the reference data set respectively; Determine the outlier value of the two data corresponding to each data item to obtain distinguished data.
5. The risk assessment method based on business data according to any one of claims 1 to 4, characterized in that: The determining the corresponding business service type according to the business data set includes: Mapping each business data item in the business data set to a preset business tree diagram to obtain a mapping tree diagram; Determine the names corresponding to the mapped nodes in the mapping tree diagram to obtain multiple business process names; The business service type is obtained by matching the texts of the multiple business process names with corresponding types in multiple historical business types recorded in a preset database.
6. The risk assessment method based on business data according to any one of claims 1 to 4, characterized in that: The using the distinguishing data to assess the risk level includes: Determining a risk weight value according to the data item corresponding to the distinguishing data; Calculating a risk assessment value using the numerical value of the distinguishing data and the corresponding risk weight value; The risk level is determined according to the magnitude of the risk assessment value.
7. A risk assessment device based on business data, characterized in that: The device comprises: A type determination module is used to determine a corresponding business service type according to a business data set after obtaining a business data set of a user, wherein the business data set includes multiple business data; An abnormal data module, used to determine whether the business data set contains abnormal data based on a local outlier algorithm; A data acquisition module is configured to acquire a reference data set if it is determined that the business data set contains abnormal data, wherein the reference data set is a data set of users completing services corresponding to the business service type obtained from a third-party organization; The risk assessment module is configured to determine difference data between the business data set and the reference data set based on the business service type, and assess a risk level using the difference data.
8. A risk assessment system based on business data, characterized in that: The system includes: a data processing platform, multiple business terminals and multiple third-party terminals; The data processing platform is connected to the multiple business terminals and the multiple third-party terminals respectively; The data processing platform is applicable to the business data-based risk assessment method as described in any one of claims 1 to 6.
9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the risk assessment method based on business data as described in any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer-executable program, and the computer-executable program is used to enable a computer to execute the risk assessment method based on business data according to any one of claims 1 to 6.
Citation Information
Patent Citations
Abnormity detection and attribution method and device, equipment and computer readable storage medium
CN111901171A
Quality treatment and improvement method for industrial Internet platform data
CN113515512A
Data security assessment method and device, equipment and storage medium
CN113656808A