Risk judgment method and device

By identifying risk types and associated characteristics on an internet platform, automatically calculating risk value and frequency scores, and generating a comprehensive score, the system solves the problems of low efficiency and poor reliability in risk assessment in existing technologies, and achieves efficient and accurate risk assessment across all platforms and business scenarios.

CN121504119APending Publication Date: 2026-02-10BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411052676.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies are unable to efficiently and accurately prevent and assess risks throughout the entire lifecycle on internet platforms. In particular, when faced with massive amounts of user data, it is difficult to achieve risk discovery and assessment across multiple subject domains across all platforms and business scenarios, and existing methods have poor reliability.

Method used

By identifying the risk types within the subject set, obtaining the associated characteristic data of the subjects to be judged, and using risk value association characteristics and risk frequency association characteristics for automated judgment, the risk value and frequency scores are calculated, and a comprehensive risk score is generated by combining fixed coefficients, thereby achieving risk discovery and judgment across the entire platform and all business scenarios.

Benefits of technology

It enables efficient and accurate risk discovery and assessment, and can perform automated risk assessments on multiple subject areas on an internet platform, improving the efficiency and reliability of risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504119A_ABST
    Figure CN121504119A_ABST
Patent Text Reader

Abstract

The invention discloses a risk judgment method and device, and relates to the technical field of big data. A specific embodiment of the method comprises the steps of determining risk types existing in a preset main body set, and selecting a plurality of to-be-judged main bodies related to any risk type from the main body set; determining a plurality of associated features of the risk type, and obtaining feature data of the to-be-judged subject in the associated features; determining a value-at-risk judgment result of the to-be-judged subject according to the feature data of the to-be-judged subject in the value-at-risk associated features, and determining a risk frequency judgment result of the to-be-judged subject according to the feature data of the to-be-judged subject in the risk frequency associated features, and determining a risk comprehensive determination result of the to-be-determined subject based on the risk value determination result and the risk frequency determination result of the to-be-determined subject. According to the embodiment, efficient and accurate risk discovery and judgment can be realized through an automatic judgment mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and in particular to a risk assessment method and apparatus. Background Technology

[0002] Currently, internet platforms are plagued by malicious users exploiting platform vulnerabilities and circumventing rules to commit fraud, engage in professional claims, and conduct favoritism, disrupting public order and business operations. Therefore, it is necessary to conduct full lifecycle risk prevention and assessment for all transactions on internet platforms. Existing technologies primarily rely on professional auditors or investigators to conduct in-depth investigations and risk assessments of specific business operations. However, due to variations in investigation cycles, experience cannot be replicated, and it is impossible to comprehensively assess multiple types of risks simultaneously, resulting in low efficiency in risk discovery and assessment. Furthermore, existing technologies suffer from poor reliability; faced with massive amounts of user data, expert experience alone is insufficient for reliable and comprehensive risk discovery and assessment. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a risk assessment method and apparatus that can achieve efficient and accurate risk discovery and assessment across all platforms, business scenarios, and multiple subject domains through automated assessment methods.

[0004] To achieve the above objectives, according to one aspect of the present invention, a risk assessment method is provided.

[0005] The risk assessment method of this invention includes: determining risk types existing in a preset set of subjects; selecting multiple subjects to be assessed related to any risk type from the set of subjects; determining multiple associated features of the risk type; and obtaining feature data of the subjects to be assessed in the associated features; wherein the associated features include: risk value associated features and risk frequency associated features; determining the risk value assessment result of the subjects to be assessed based on the feature data of the subjects to be assessed in the risk value associated features; determining the risk frequency assessment result of the subjects to be assessed based on the feature data of the subjects to be assessed in the risk frequency associated features; and determining the comprehensive risk assessment result of the subjects to be assessed based on the risk value assessment result and the risk frequency assessment result.

[0006] Optionally, the method further includes: determining a first risk threshold corresponding to the risk value association feature based on the feature data of the subject to be determined in the risk value association feature; determining a second risk threshold corresponding to the risk frequency association feature based on the feature data of the subject to be determined in the risk frequency association feature; and identifying the subject to be determined whose feature data of the risk value association feature is greater than the first risk threshold or whose feature data of the risk frequency association feature is greater than the second risk threshold as a risk subject.

[0007] Optionally, determining the risk value determination result of the subject to be determined based on the feature data of the risk value correlation feature includes: obtaining quantile data of the feature data of each risk subject in the risk value correlation feature at a preset first quantile value, and the risk value score of the quantile data within a preset first numerical range; determining the growth multiple of the risk value correlation feature using the quantile data, the risk value score of the quantile data, and a first risk threshold; and determining the risk value score of any risk subject within the first numerical range based on the first risk threshold, the growth multiple, and the feature data of any risk subject in the risk value correlation feature.

[0008] Optionally, determining the risk frequency determination result of the subject to be determined based on the feature data of the risk frequency correlation feature includes: obtaining quantile data of the feature data of each risk subject in the risk frequency correlation feature at a preset second quantile value, and a risk frequency score of the quantile data within a first numerical range; determining the growth multiple of the risk frequency correlation feature using the quantile data, the risk frequency score of the quantile data, and a second risk threshold; and determining the risk frequency score of the risk subject within the first numerical range based on the second risk threshold, the growth multiple, and the feature data of any risk subject in the risk frequency correlation feature.

[0009] Optionally, determining the comprehensive risk assessment result of the subject to be assessed based on the risk value assessment result and the risk frequency assessment result of the subject to be assessed includes: multiplying the risk value score, risk frequency score, risk severity score within a first numerical range for the corresponding risk type, and a fixed coefficient of any risk subject to obtain the comprehensive risk score of the subject to be assessed.

[0010] Optionally, the associated features further include: multiple other associated features; and the method further includes: determining the score of each other associated feature on a preset classification index based on the feature data of the subject to be judged in the other associated features; wherein the classification index includes: information gain or Gini coefficient; determining the preset number of other associated features with the highest scores as display features, and generating a visualization chart of the comprehensive risk score of each risk subject based on the display features.

[0011] Optionally, the method further includes: after acquiring multiple feature data of any subject to be determined based on the associated features, aggregating the multiple feature data based on the identifier of the subject to be determined.

[0012] Optionally, the method further includes: storing the identifiers and comprehensive risk scores of each risk subject in the same risk type in an ordered set of a pre-set cache unit; wherein the elements in the ordered set are arranged in descending order of comprehensive risk scores; storing the identifiers of each risk subject in the risk type, the risk value determination results of each risk subject, the risk frequency determination results, and the feature data of the associated features in a pre-set database table; and transmitting the identifier of the ordered set to a risk processing terminal for processing the risk type.

[0013] To achieve the above objectives, according to another aspect of the present invention, a risk assessment device is provided.

[0014] The risk assessment device of this invention includes: a risk type determination unit, configured to determine risk types existing in a preset set of subjects, and select multiple subjects to be assessed related to any risk type from the set of subjects; a data collection unit, configured to determine multiple associated features of the risk type, and acquire feature data of the subject to be assessed in the associated features; wherein, the associated features include: risk value associated features and risk frequency associated features; and a assessment unit, configured to determine the risk value assessment result of the subject to be assessed based on the feature data of the subject to be assessed in the risk value associated features, determine the risk frequency assessment result of the subject to be assessed based on the feature data of the subject to be assessed in the risk frequency associated features, and determine the comprehensive risk assessment result of the subject to be assessed based on the risk value assessment result and the risk frequency assessment result.

[0015] To achieve the above objectives, according to another aspect of the present invention, an electronic device is provided.

[0016] An electronic device according to the present invention includes: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the risk assessment method provided by the present invention.

[0017] To achieve the above objectives, according to another aspect of the present invention, a computer-readable storage medium is provided.

[0018] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the risk assessment method provided by the present invention.

[0019] According to the technical solution of the present invention, the embodiments described above have the following advantages or beneficial effects: First, the risk types existing in the pre-defined subject set are identified. Then, multiple subjects related to any given risk type are selected from this set. Next, multiple associated characteristics of that risk type are determined, and feature data for each associated characteristic of the subject is obtained. These associated characteristics may include risk value association characteristics and risk frequency association characteristics. Subsequently, the risk value assessment result of the subject is determined based on the feature data of the risk value association characteristics, and the risk frequency assessment result is determined based on the feature data of the risk frequency association characteristics. Finally, a comprehensive risk assessment result is determined based on both the risk value assessment result and the risk frequency assessment result. In this way, an automated assessment method that automatically collects data and performs calculations achieves efficient and accurate risk discovery and assessment across the entire platform, all business scenarios, and multiple subject domains.

[0020] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description

[0021] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein: Figure 1 This is a schematic diagram of the main steps of the risk determination method in this embodiment of the invention; Figure 2 This is a schematic diagram of the overall process of the risk determination method in an embodiment of the present invention; Figure 3 This is a schematic diagram of the components of the risk determination device in an embodiment of the present invention; Figure 4 This is an exemplary system architecture diagram that can be applied thereto according to embodiments of the present invention; Figure 5 This is a schematic diagram of the electronic device structure used to implement the risk determination method in the embodiments of the present invention. Detailed Implementation

[0022] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] It should be noted that, unless otherwise specified, the embodiments of the present invention and the technical features thereof can be combined with each other.

[0024] Figure 1This is a schematic diagram of the main steps of the risk determination method according to an embodiment of the present invention.

[0025] like Figure 1 As shown, the risk assessment method of this invention can be specifically executed according to the following steps: Step S101: Determine the risk types existing in the preset subject set, and select multiple subjects to be judged related to any risk type from the subject set.

[0026] In this embodiment of the invention, the subject set is a collection of multiple subjects of the same subject type. These subject types can include user accounts, service provider employees, stores under the service provider, and partner companies of the service provider. The risk types can be fine-grained risk classification results corresponding to a specific subject set. For example, for user accounts, this could include abnormal claims (e.g., professional claims for profit), frequent returns, frequent exchanges, excessive order IPs, etc.; for employees, it could include exchange of benefits (e.g., receiving money from partners to provide them with benefits), serious breach of trust, etc. The risk types can also belong to coarse-grained risk categories. For example, frequent returns and frequent exchanges belong to the fraud arbitrage risk category. In this step, the risk categories can be sorted out first, and then the risk types in each business scenario can be determined based on the specific business scenario.

[0027] After determining the risk type, multiple entities related to any risk type can be selected from the corresponding set of entities to be assessed. For example, taking the abnormal claim risk type as an example, the relevant entities to be assessed could be accounts that have recently made payments.

[0028] Step S102: Determine multiple associated features of the risk type and obtain the feature data of the subject to be determined in the associated features.

[0029] In this step, multiple associated features related to the risk type can be identified, and feature data for each subject to be determined in each associated feature can be obtained. In practical applications, these associated features may include: risk value associated features and risk frequency associated features. The former is a feature related to the risk value, such as the amount involved, while the latter is related to the frequency of risky behavior. For example, for abnormal claim risk types, the risk value associated feature could be the claim amount (i.e., the total claim amount over a period of time), and the risk frequency associated feature could be the number of claims. In practical applications, after obtaining multiple feature data for any subject to be determined in the associated features, the multiple feature data can be aggregated based on the identifier of the subject to be determined, for example, summing the claim amount and the number of claims for multiple feature data.

[0030] Step S103: Determine the risk value judgment result of the subject to be judged based on the characteristic data of the risk value correlation feature of the subject to be judged, determine the risk frequency judgment result of the subject to be judged based on the characteristic data of the risk frequency correlation feature of the subject to be judged, and determine the comprehensive risk judgment result of the subject to be judged based on the risk value judgment result and the risk frequency judgment result of the subject to be judged.

[0031] Before proceeding with this step, a first risk threshold corresponding to the risk value correlation feature can be determined based on the characteristic data of each subject to be assessed in the risk value correlation feature. For example, the first risk threshold can be determined based on the statistical values ​​of the characteristic data of each subject to be assessed in the risk value correlation feature, such as the 10th quartile, 20th quartile, 30th quartile, mean, median, maximum, or minimum value. A second risk threshold can be determined based on the statistical values ​​of the characteristic data of each subject to be assessed in the risk frequency correlation feature, such as the 10th quartile, 20th quartile, 30th quartile, mean, median, maximum, or minimum value. Subsequently, subjects whose characteristic data of the risk value correlation feature is greater than the first risk threshold, or whose characteristic data of the risk frequency correlation feature is greater than the second risk threshold, can be identified as risk subjects and placed into a risk pool for risk assessment. Other subjects to be assessed, due to their lower risk level, will not be assessed. The risk pool refers to a container used to store the relevant data of risk subjects, and can adopt any applicable structure such as a queue or database table. The first and second risk thresholds can also be directly given by human experience.

[0032] Subsequently, the risk value assessment result of each subject to be assessed can be calculated based on the characteristic data of each subject to be assessed in terms of risk value correlation characteristics, and the risk frequency assessment result of each subject to be assessed can be calculated based on the characteristic data of each subject to be assessed in terms of risk frequency correlation characteristics. Finally, the comprehensive risk assessment result of the subject to be assessed can be determined based on the calculated risk value assessment result and risk frequency assessment result of the subject to be assessed.

[0033] Specifically, for the risk value assessment result, the quantile data of the characteristic data of each risk subject in the risk value correlation feature at a preset first quantile value, and the risk value score of the quantile data within a preset first numerical range can be obtained. For example, the first numerical range is between 1 and 10 points, the first quantile value is 90 (representing the 90th quantile), and the risk value score of the quantile data is 9 points.

[0034] Subsequently, the growth multiple of the risk value correlation feature is determined using the above quantile data, the risk value score of the above quantile data, and the first risk threshold. In practical applications, the relevant formula for the growth multiple (Formula 1) is as follows:

[0035]

[0036] in, Indicates the risk value score. This represents characteristic data indicating the risk subject's correlation with risk value. This indicates the first risk threshold. This indicates the growth multiple of the risk value-related characteristics.

[0037] The above quantile data (quantile data of the 90th quantile) Substituting its risk value score (9 points) into the above formula yields:

[0038]

[0039] Subsequently, based on the first risk threshold, the calculated growth multiple, and the characteristic data of any risk subject in the risk value correlation feature, Formula 1 can be used to determine the risk value score of the risk subject within the first numerical range. In order to limit the score of abnormally large values ​​to the first numerical range, preferably, when the calculated risk value score of a risk subject is greater than 9 points, its final risk value score can be determined as 10 points to prevent the score from exceeding 10 points.

[0040] The calculation method for determining the frequency of risk is similar. First, the quantile data of each risk subject's risk frequency correlation characteristics at a preset second quantile value is obtained, along with the risk frequency score of that quantile data within a first numerical range. The second quantile value may or may not be equal to the first quantile value, as shown in Formula 2.

[0041]

[0042] in, Indicates the frequency of risk. Feature data representing the correlation between the frequency of risks and the characteristics of risk subjects. This indicates the second risk threshold. This indicates the growth rate of the risk frequency-related characteristics.

[0043] Using the 90th percentile as an example again, this quantile data will then be used... The risk frequency score (9 points) for this quantile data and the growth multiple of the risk frequency correlation feature determined by the second risk threshold are as follows:

[0044]

[0045] Subsequently, based on the second risk threshold, the growth rate, and the characteristic data of any risk subject's risk frequency correlation features, the risk frequency score of the risk subject within the first numerical range can be determined using Formula 2. To limit abnormally large scores to the first numerical range, preferably, when a risk subject's calculated risk frequency score exceeds 9 points, its final risk frequency score can be set to 10 points to prevent the score from exceeding 10 points.

[0046] Finally, the comprehensive risk assessment result for the subject to be assessed can be determined based on the risk value assessment result and the risk frequency assessment result, and thus become the final risk assessment result. Preferably, the comprehensive risk score for any risk subject can be obtained by multiplying its risk value score, risk frequency score, risk severity score (pre-set for the corresponding risk type within a first numerical range), and a fixed coefficient, as shown in the following formula:

[0047]

[0048] Where R represents the overall risk score, S represents the risk severity score, and the fixed coefficient in the above formula is 0.1. In practical applications, if S, F, and V are all between 1 and 10, then R ranges from 1 to 100.

[0049] As a preferred embodiment, in addition to the risk value correlation feature and the risk frequency correlation feature, the correlation features for any risk type may further include multiple other correlation features. In one embodiment, the scores of each other correlation feature in a preset classification index can be determined based on the feature data of the subject to be judged in the other correlation features. For example, the above classification index includes information gain or the Gini coefficient, which are known techniques. When using information gain, the known ID3 algorithm or C4.5 algorithm can be used. It can be understood that information gain and the Gini coefficient can measure the degree of influence of a feature on the classification result. The larger the information gain, the greater the influence of the feature on the classification result; the smaller the Gini coefficient, the greater the influence of the feature on the classification result.

[0050] Subsequently, the other related features with the highest scores are identified as display features. Based on these display features, a visual chart (such as a radar chart) of the comprehensive risk score of each risk subject is generated, thereby realizing the visualization of risk assessment results in multiple dimensions.

[0051] As a preferred solution, after obtaining the comprehensive risk score of each risk subject, the identifiers and comprehensive risk scores of each risk subject in the same risk type can be stored in an ordered set of a pre-set cache unit. The elements in the ordered set can be arranged in descending order (from largest to smallest) of the comprehensive risk score. The identifiers of each risk subject in this risk type, the risk value judgment results of each risk subject, the risk frequency judgment results, and the feature data of the associated characteristics are stored in a pre-set database table. Finally, the identifier of the ordered set is transmitted to the risk processing terminal used to process this risk type.

[0052] For example, the above caching unit can be Redis, the ordered set can be a Zset in Redis, and the database table can be a table in a relational database (such as MySQL). Through the above processing, relatively important data—the risk subject's identifier and risk comprehensive score—can be stored in a caching unit with high storage media costs but fast read / write speeds. Risk subjects are then sorted in descending order of their risk comprehensive scores, allowing the risk processing end to efficiently process the risk status of each risk subject in descending order of risk severity. Simultaneously, the risk details data of each risk subject (including risk value assessment results, risk frequency assessment results, and feature data related to associated characteristics) are stored in a large database table. When needed, the risk processing end can query relevant risk details data based on the risk subject's identifier. This query operation can be performed asynchronously, and the read / write performance of the database table can meet the requirements, thus achieving a balance between risk processing efficiency and data storage costs, and avoiding pressure on the storage system from massive amounts of data.

[0053] The following describes a specific embodiment of the present invention; see [link to specific embodiment]. Figure 2 .

[0054] Internet platforms have entered a phase of stringent regulation, making operational security, regulatory security, information security, and legal security in the transaction process indispensable. Throughout the entire operation of an internet platform, a sound and compliant operational order not only meets regulatory requirements but also purifies the operating environment and enhances user experience. However, some malicious users exploit loopholes, circumventing platform rules to engage in fraud, arbitrage, professional claims, and exchanges of favors. These illegal activities, to varying degrees, not only expose businesses and users to various risks but also disrupt public order. Therefore, it is necessary to prevent, intercept, and manage risks throughout the entire lifecycle of internet platform marketing—before, during, and after the event. Identifying risks across various business segments of different entities within massive amounts of data requires in-depth analysis of different entities' business processes to uncover risks or judgment based on historical experience. However, as businesses become increasingly complex and data volumes grow massive, a flexible and universal risk scoring technology solution needs to be developed for different business scenarios of different entities to predict potential risks and construct an integrated compliance risk scoring standard for internet platforms encompassing risk identification, assessment, and governance. Currently, traditional auditing and case investigation scenarios lack a risk scoring technology that covers the entire subject domain (i.e., various entities) to comprehensively identify risks at each business node. Most investigations are based on a single business point, which is not only limited and has a high threshold for investigation, but also cannot conduct joint investigations across multiple business scenarios to comprehensively assess the magnitude of risks and achieve comprehensive identification.

[0055] To address the above issues, this embodiment of the invention begins by analyzing the risk types of various business operations. Starting from four major business themes (including accounts, employees, stores, and companies), it specifically analyzes the risk types of each theme domain, such as exchange of benefits, information security, regulation, violation of values, transaction risk, fraud, corruption, business order, professional claims, and serious breach of trust. Then, risk features are selected based on the specific risk type, and the features that ultimately measure the risk type are selected based on the feature importance of the decision tree. Feature values ​​are calculated according to feature rules, and indicator scores are calculated according to the SFV three-indicator scoring algorithm. Finally, a comprehensive risk score is determined, with a higher score indicating a greater risk.

[0056] During the data cleaning phase, the risk categories of various entities are first sorted out. Then, the risk types for each entity are determined, and the associated fields for each risk type are obtained. Taking abnormal claims as an example, the following associated characteristics of accounts with claims within a certain period can be obtained: account, claim amount, number of claims, total consumption amount, number of items purchased, claim type, product code, primary category code, primary category name, secondary category code, secondary category name, tertiary category code, tertiary category name, brand name, operator account, operator name, etc. Afterwards, data processing and filtering / impact can be performed, such as handling outliers and filtering and imputing null values.

[0057] During the data analysis phase, the first step is to conduct a risk type distribution analysis and determine relevant thresholds. For example, claim records of users who placed orders during the survey period are selected, and the claim amount, number of claims, total consumption amount, number of items purchased, and compensation amount are statistically analyzed. The minimum, 20th, 40th, 50th, 80th, and 90th percentiles, maximum, average, and mode of each business characteristic are analyzed to assess the user's data distribution. Combined with historical professional claim cases, a professional claim threshold is determined. When a user's claim amount exceeds the threshold, that user is placed in a professional claim risk pool, and the risk level is then determined using risk-related algorithms. Subsequently, specific data can be excluded, such as claim data from low-risk corporate accounts, to avoid risk assessment bias.

[0058] In the feature calculation stage, feature selection based on random forests can be performed on relevant features other than claim amount and number of claims. This involves using information gain, Gini coefficient, and other algorithms to filter features. Afterward, feature dimensions are aggregated, and the model is iteratively updated and the results are output. Using information gain and Gini coefficient for feature selection is a known technique. For example, when using information gain, the known ID3 or C4.5 algorithms can be used to calculate the information gain of each feature, and a predetermined number of features with the highest information gain are selected as the results. When using the Gini coefficient, known methods for calculating the Gini coefficient of each feature can be used, and a predetermined number of features with the lowest Gini coefficient are selected as the results. This selection process removes features that have little impact on risk assessment, enhancing the targeting of risk points and reducing subsequent computational load through data dimensionality reduction based on the random forest algorithm.

[0059] In the indicator scoring and risk scoring stages, the aforementioned formulas can be used to calculate the risk value score V and the risk frequency score F, respectively, and the comprehensive risk score R can be obtained based on the preset risk severity score S.

[0060] For example, combine the following formulas:

[0061]

[0062]

[0063] The formula for calculating the Value at Risk (V) score is as follows:

[0064]

[0065] Combine the following formulas:

[0066]

[0067]

[0068] The formula for calculating the risk frequency score F is as follows:

[0069]

[0070] Based on the above steps, a risk scoring scheme based on feature rules and indicator algorithms is provided. This scheme can output risk scores across all platforms, business scenarios, and multiple thematic domains, measuring the risk level of each theme. Subsequent investigations can be launched based on the risk targets identified by the risk score, or the risk can be proactively applied to various stages of the business (including pre-event risk scanning, risk interception, and post-event governance). This breaks with conventional investigation methods and helps internet platforms conduct comprehensive risk scanning. Embodiments of this invention can be used for special audits, case investigations, identifying existing risks, assisting in the execution of audits and investigations, and operational audits, safeguarding the various businesses of service providers.

[0071] It should be noted that the technical solutions of this invention, including the collection, updating, analysis, processing, use, transmission, and storage of user personal information, all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0072] For the foregoing method embodiments, they are described as a series of actions for ease of description. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, and some steps may actually be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential for implementing the present invention.

[0073] To facilitate better implementation of the above-described solutions of the embodiments of the present invention, related apparatus for implementing the above-described solutions is also provided below.

[0074] Please see Figure 3 As shown, the risk assessment device 300 provided in this embodiment of the invention may include: a risk type determination unit 301, a data collection unit 302, and a assessment unit 303.

[0075] The risk type determination unit 301 is used to determine the risk types existing in a preset set of subjects, and select multiple subjects related to any risk type from the set of subjects; the data collection unit 302 is used to determine multiple associated features of the risk type, and obtain the feature data of the subject to be determined in the associated features; wherein, the associated features include: risk value associated features and risk frequency associated features; the determination unit 303 is used to determine the risk value determination result of the subject to be determined based on the feature data of the subject to be determined in the risk value associated features, determine the risk frequency determination result of the subject to be determined based on the feature data of the subject to be determined in the risk frequency associated features, and determine the comprehensive risk determination result of the subject to be determined based on the risk value determination result and the risk frequency determination result.

[0076] In this embodiment of the invention, the determination unit 303 may be further configured to: determine a first risk threshold corresponding to the risk value association feature based on the feature data of the subject to be determined in the risk value association feature; determine a second risk threshold corresponding to the risk frequency association feature based on the feature data of the subject to be determined in the risk frequency association feature; and determine the subject to be determined whose feature data of the risk value association feature is greater than the first risk threshold or whose feature data of the risk frequency association feature is greater than the second risk threshold as a risk subject.

[0077] Preferably, the determination unit 303 may be further configured to: acquire quantile data of the feature data of each risk subject in the risk value correlation feature at a preset first quantile value, and the risk value score of the quantile data in a preset first numerical range; determine the growth multiple of the risk value correlation feature using the quantile data, the risk value score of the quantile data, and a first risk threshold; and determine the risk value score of any risk subject in the first numerical range based on the first risk threshold, the growth multiple, and the feature data of any risk subject in the risk value correlation feature.

[0078] As a preferred embodiment, the determination unit 303 may be further configured to: acquire quantile data of the feature data of each risk subject in the risk frequency correlation feature at a preset second quantile value, and a risk frequency score of the quantile data in a first numerical range; determine the growth multiple of the risk frequency correlation feature using the quantile data, the risk frequency score of the quantile data, and a second risk threshold; and determine the risk frequency score of any risk subject in the first numerical range based on the second risk threshold, the growth multiple, and the feature data of any risk subject in the risk frequency correlation feature.

[0079] In specific applications, the determination unit 303 can be further used to: multiply the risk value score, risk frequency score, risk severity score within the first numerical range preset for the corresponding risk type, and a fixed coefficient of any risk subject to obtain the comprehensive risk score of the risk subject.

[0080] In practical applications, the associated features further include: multiple other associated features; and the data collection unit 302 can be further used to: determine the score of each other associated feature in a preset classification index based on the feature data of the subject to be judged in the other associated features; wherein, the classification index includes: information gain or Gini coefficient; the device 300 can further include: a visualization unit, used to: determine the preset number of other associated features with the highest scores as display features, and generate a visualization chart of the comprehensive risk score of each risk subject based on the display features.

[0081] Preferably, the data collection unit 302 can be further used to: after acquiring multiple feature data of any subject to be determined based on the identifier of the subject to be determined, aggregate the multiple feature data based on the identifier of the subject to be determined.

[0082] Furthermore, in this embodiment of the invention, the device 300 may further include a data storage unit for: storing the identifiers and comprehensive risk scores of each risk subject in the same risk type in an ordered set of a preset cache unit; wherein the elements in the ordered set are arranged in descending order of comprehensive risk scores; storing the identifiers of each risk subject in the risk type, the risk value determination results of each risk subject, the risk frequency determination results, and the feature data of the associated features in a preset database table; and transmitting the identifier of the ordered set to a risk processing terminal for processing the risk type.

[0083] According to the technical solution of this invention, firstly, the risk types existing in a preset set of subjects are determined, and multiple subjects related to any risk type are selected from the subject set; then, multiple associated features of the risk type are determined, and feature data of the subjects to be judged in each associated feature are obtained. These associated features may include risk value associated features and risk frequency associated features; subsequently, the risk value judgment result of the subjects to be judged is determined based on the feature data of the subjects to be judged in the risk value associated features, and the risk frequency judgment result of the subjects to be judged is determined based on the feature data of the subjects to be judged in the risk frequency associated features; finally, the comprehensive risk judgment result of the subjects to be judged is determined based on the risk value judgment result and the risk frequency judgment result. Thus, through an automated judgment method that automatically collects data and performs automatic calculations, efficient and accurate risk discovery and judgment are achieved across all platforms, all business scenarios, and multiple subject domains.

[0084] Figure 4 An exemplary system architecture 400 is shown that can be applied to the risk assessment method or risk assessment device of the present invention.

[0085] like Figure 4 As shown, system architecture 400 may include terminal devices 401, 402, and 403, network 404, and server 405 (this architecture is merely an example; the components included in a specific architecture may be adjusted according to the specific application). Network 404 serves as the medium for providing a communication link between terminal devices 401, 402, and 403 and server 405. Network 404 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0086] Users can use terminal devices 401, 402, and 403 to interact with server 405 via network 404 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 401, 402, and 403, such as risk assessment applications (for example only).

[0087] Terminal devices 401, 402, and 403 can be various electronic devices with displays that support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0088] Server 405 can be a server that provides various services, such as a backend server that supports risk assessment applications operated by users using terminal devices 401, 402, and 403 (for example only). The backend server can process received risk assessment requests and feed back the processing results (such as calculated risk scores - for example only) to terminal devices 401, 402, and 403.

[0089] It should be noted that the risk assessment method provided in the embodiments of the present invention is generally executed by server 405, and correspondingly, the risk assessment device is generally set in server 405.

[0090] It should be understood that Figure 4 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0091] The present invention also provides an electronic device. The electronic device according to an embodiment of the present invention includes: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the risk assessment method provided by the present invention.

[0092] The following is for reference. Figure 5It shows a schematic diagram of the structure of a computer system 500 suitable for implementing an electronic device according to embodiments of the present invention. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0093] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage section 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the computer system 500. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0094] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.

[0095] In particular, according to the embodiments disclosed in this invention, the processes described in the above main step diagrams can be implemented as computer software programs. For example, embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the main step diagrams. In the above embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit 501, it performs the functions defined in the system of this invention.

[0096] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0098] The units described in the embodiments of the present invention can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor can be described as including a risk type determination unit, a data collection unit, and a decision unit. The names of these units do not necessarily limit the specific unit; for example, the risk type determination unit can also be described as "a unit that provides risk types to the data collection unit."

[0099] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, and when the device executes the one or more programs, the steps performed by the device include: determining risk types existing in a preset set of subjects; selecting multiple subjects to be judged related to any risk type from the set of subjects; determining multiple associated features of the risk type; and obtaining feature data of the subject to be judged in the associated features; wherein the associated features include: risk value associated features and risk frequency associated features; determining a risk value judgment result of the subject to be judged based on the feature data of the subject to be judged in the risk value associated features; determining a risk frequency judgment result of the subject to be judged based on the feature data of the subject to be judged in the risk frequency associated features; and determining a comprehensive risk judgment result of the subject to be judged based on the risk value judgment result and the risk frequency judgment result.

[0100] In the technical solution of this invention embodiment, firstly, the risk types existing in a preset set of subjects are determined, and multiple subjects related to any risk type are selected from the subject set; then, multiple associated features of the risk type are determined, and feature data of the subject to be judged in each associated feature are obtained. These associated features may include risk value associated features and risk frequency associated features; subsequently, the risk value judgment result of the subject to be judged is determined based on the feature data of the subject to be judged in the risk value associated features, and the risk frequency judgment result of the subject to be judged is determined based on the feature data of the subject to be judged in the risk frequency associated features; finally, the comprehensive risk judgment result of the subject to be judged is determined based on the risk value judgment result and the risk frequency judgment result. Thus, through an automated judgment method that automatically collects data and performs automatic calculations, efficient and accurate risk discovery and judgment are achieved across all platforms, all business scenarios, and multiple subject domains.

[0101] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A risk assessment method, characterized in that, include: Determine the risk types present in the pre-set set of subjects, and select multiple subjects related to any risk type from the set of subjects to be judged; Multiple associated features of this risk type are determined, and feature data of the subject to be judged based on these associated features are obtained; wherein, the associated features include: risk value associated features and risk frequency associated features; The risk value determination result of the subject to be determined is determined based on the feature data of the risk value correlation feature. The risk frequency determination result of the subject to be determined is determined based on the feature data of the risk frequency correlation feature. The comprehensive risk determination result of the subject to be determined is determined based on the risk value determination result and the risk frequency determination result.

2. The method according to claim 1, characterized in that, The method further includes: Based on the feature data of the subject to be determined in the risk value correlation feature, a first risk threshold corresponding to the risk value correlation feature is determined; based on the feature data of the subject to be determined in the risk frequency correlation feature, a second risk threshold corresponding to the risk frequency correlation feature is determined. The subject to be judged is identified as a risk subject if the feature data of the risk value association feature is greater than the first risk threshold, or if the feature data of the risk frequency association feature is greater than the second risk threshold.

3. The method according to claim 2, characterized in that, The step of determining the risk value assessment result of the subject to be assessed based on the feature data of the risk value correlation characteristics includes: Obtain the quantile data of the feature data of each risk subject in the risk value correlation feature at a preset first quantile value, and the risk value score of the quantile data in a preset first numerical range; The growth multiple of the risk value association feature is determined using the quantile data, the risk value score of the quantile data, and a first risk threshold. The risk value score of a risk subject within a first numerical range is determined based on the first risk threshold, the growth rate, and the characteristic data of any risk subject in the risk value correlation feature.

4. The method according to claim 3, characterized in that, The step of determining the risk frequency determination result of the subject to be determined based on the feature data of the risk frequency correlation characteristics includes: Obtain the quantile data of the feature data of each risk subject in the risk frequency correlation feature at a preset second quantile value, and the risk frequency score of the quantile data in a first numerical range; The growth rate of the risk frequency correlation feature is determined using the quantile data, the risk frequency score of the quantile data, and a second risk threshold. The risk frequency score of the risk subject within a first numerical range is determined based on the second risk threshold, the growth rate, and the characteristic data of any risk subject related to the risk frequency.

5. The method according to claim 4, characterized in that, The determination of the comprehensive risk assessment result of the subject to be assessed based on the risk value assessment result and the risk frequency assessment result includes: The risk comprehensive score of any risk subject is obtained by multiplying its risk value score, risk frequency score, risk severity score within the first numerical range for the corresponding risk type, and a fixed coefficient.

6. The method according to claim 5, characterized in that, The association features further include: multiple other association features; and the method further includes: The scores of each of the other related features are determined based on the feature data of the subject to be judged in the other related features; wherein, the classification indicators include: information gain or Gini coefficient; The other related features with the highest scores are identified as display features, and a visual chart of the comprehensive risk score of each risk subject is generated based on the display features.

7. The method according to claim 1, characterized in that, The method further includes: After acquiring multiple feature data of any subject to be determined in the associated features, the multiple feature data of the subject to be determined in the risk value associated features are summed based on the identifier of the subject to be determined, and the multiple feature data of the subject to be determined in the risk frequency associated features are summed based on the identifier of the subject to be determined.

8. The method according to claim 5, characterized in that, The method further includes: The identifiers and comprehensive risk scores of each risk subject within the same risk type are stored in an ordered set of a pre-set cache unit; wherein the elements in the ordered set are arranged in descending order of comprehensive risk score. The identifiers of each risk subject in this risk type, the risk value determination results of each risk subject, the risk frequency determination results, and the characteristic data of the associated features are stored in a preset database table; The identifier of the ordered set is transmitted to the risk processing terminal used to handle this type of risk.

9. A risk assessment device, characterized in that, include: The risk type determination unit is used to determine the risk types existing in a preset set of subjects, and to select multiple subjects to be determined related to any risk type from the set of subjects; A data collection unit is used to determine multiple associated features of the risk type and acquire feature data of the subject to be judged based on the associated features; wherein, the associated features include: risk value associated features and risk frequency associated features; The determination unit is used to determine the risk value determination result of the subject to be determined based on the feature data of the risk value correlation feature, determine the risk frequency determination result of the subject to be determined based on the feature data of the risk frequency correlation feature, and determine the comprehensive risk determination result of the subject to be determined based on the risk value determination result and the risk frequency determination result.

10. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.