Risk data characteristic value analysis method and device, equipment and medium

By incorporating residual feature values ​​into the risk analysis model, the initial risk analysis model is optimized, which solves the problem of low accuracy in risk scoring in existing technologies and achieves more accurate risk assessment.

CN121958862APending Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing risk analysis models are not optimized for the residuals of the initial assessment, resulting in reduced accuracy of risk scoring.

Method used

By extracting residual feature values ​​to correct initial risk feature values, and quantifying key features related to residuals, the residual correction logic is integrated into the initial risk analysis model to optimize the initial risk analysis model.

Benefits of technology

It significantly improves the accuracy of risk scoring, eliminates initial model bias, and provides an assessment that is closer to the true level of risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958862A_ABST
    Figure CN121958862A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a risk data feature value analysis method, device and equipment and a medium, and the method comprises the steps: obtaining multi-dimensional data of a preset risk type, carrying out the basic feature extraction of the multi-dimensional data, and obtaining data basic features; performing basic risk analysis on the data basic features to obtain initial risk feature values; obtaining a residual feature value according to a preset residual, and combining the initial risk feature value with the residual feature value to generate a risk feature value; extracting key features of the multi-dimensional data by adopting a preset residual analysis model, and performing interval score assignment on the key features to obtain feature interval scores; and constructing an initial risk analysis model according to the feature interval scores and the key features. By integrating the residual error correction logic into the initial risk analysis model, the initial risk analysis model is optimized in a targeted manner to eliminate the deviation of the initial model, and finally the risk scoring accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to a method, apparatus, device, and medium for feature value analysis of risk data. Background Technology

[0002] The feature value of risk data refers to a specific numerical value extracted or calculated from risk-related data (such as user credit, vehicle driving, health status, etc.) that quantifies the level and attributes of risk. It is usually calculated by a risk analysis model, which can transform complex risk analysis logic into a corresponding rule of "feature-interval-score". With its clear rules and simple calculation, its application scenarios have been extended to multiple fields. For example, in the field of healthcare, disease risk levels can be quickly quantified by assigning interval scores based on characteristics such as patient age, basic medical history, and examination indicators, thus assisting in clinical diagnosis and treatment decisions. In the field of fintech, intervals are usually divided and scored based on characteristics such as user credit history, financial status, and behavioral data to quickly quantify the user's credit risk or transaction risk level.

[0003] Most existing risk analysis model construction methods are based on the features of the initial risk model. However, this approach does not optimize for the residuals of the initial assessment, causing the subsequent models used for risk analysis to inherit the biases of the initial model, thereby reducing the accuracy of risk scoring. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and medium for eigenvalue analysis of risk data. By extracting residual eigenvalues ​​to correct initial risk eigenvalues, and quantifying key features related to the residuals, the residual correction logic is integrated into the initial risk analysis model. By comparing the risk eigenvalues ​​with the analysis results of the initial risk analysis model, the initial risk analysis model is optimized in a targeted manner to eliminate initial model bias, thereby ultimately improving the accuracy of risk scoring.

[0005] Firstly, a method for eigenvalue analysis of risk data is provided, including: Acquire multi-dimensional data of preset risk types, and extract basic features from the multi-dimensional data to obtain basic data features; A basic risk analysis is performed on the basic characteristics of the data to obtain initial risk characteristic values; Obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The key features of the multi-dimensional data are extracted using a preset residual analysis model, and the key features are assigned interval scores to obtain feature interval scores. An initial risk analysis model is constructed based on the feature interval scores and the key features, and the key features are scored and analyzed based on the initial risk analysis model to obtain feature scores; The risk feature value is compared with the feature score, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain the target risk analysis model; The target risk characteristic value of the risk data is calculated using the target risk analysis model.

[0006] Secondly, a feature value analysis device for risk data is provided, comprising: The acquisition and extraction module is used to acquire multi-dimensional data of a preset risk type and extract basic features from the multi-dimensional data to obtain basic data features; The analysis module is used to perform basic risk analysis on the basic characteristics of the data to obtain initial risk characteristic values; The acquisition and combination module is used to acquire residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The extraction and scoring module is used to extract key features of the multi-dimensional data using a preset residual analysis model, and to assign interval scores to the key features to obtain feature interval scores. An analysis module is constructed to build an initial risk analysis model based on the feature interval scores and the key features, and to perform scoring analysis on the key features based on the initial risk analysis model to obtain feature scores; The comparison and optimization module is used to compare the risk feature value with the feature score, and adjust and optimize the initial risk analysis model according to the comparison result to obtain the target risk analysis model; The calculation module is used to calculate the target risk feature value of the risk data through the target risk analysis model.

[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned risk data feature value analysis method.

[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned risk data feature value analysis method.

[0009] The aforementioned scheme, implemented by a risk data feature value analysis method, apparatus, computer equipment, and storage medium, acquires multi-dimensional data and extracts basic features, breaking through the limitations of traditional scoring models that rely solely on initial risk model features. It extracts basic features from a wider range of multi-dimensional data, providing more comprehensive risk input for subsequent analysis. Through the analysis of these basic features, a preliminary risk assessment result is quickly obtained, providing an initial reference for more accurate risk analysis. Residual feature values ​​can compensate for the deficiencies of initial risk feature values; combining the two makes the risk feature values ​​closer to the true risk level. Key features related to the residuals are quantified, transforming the driving factors behind the residuals into interval scores that can be used in the scoring model, providing a concrete basis for correcting initial biases. An initial risk analysis model is constructed and feature scores are obtained. Residual-related features are integrated into the scoring model framework, allowing the scoring model to incorporate elements for correcting initial biases for the first time, changing the traditional risk analysis model's reliance solely on initial model features. By comparing risk feature values ​​and feature scores, the biases inherited from the initial risk analysis model are accurately identified, and interval divisions or score settings are adjusted accordingly, ultimately eliminating the biases caused by unoptimized residuals in traditional initial risk analysis models and significantly improving the accuracy of risk scoring. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of an application environment for a feature value analysis method for risk data according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a feature value analysis method for risk data according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a feature value analysis device for risk data according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to one embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] The present invention provides a method for eigenvalue analysis of risk data, which can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can acquire multi-dimensional data of preset risk types, extract basic features from the multi-dimensional data to obtain basic data features; perform basic risk analysis on the basic data features to obtain initial risk feature values; obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; extract key features of the multi-dimensional data using a preset residual analysis model, and assign interval scores to the key features to obtain feature interval scores; construct an initial risk analysis model based on the feature interval scores and the key features, and perform scoring analysis on the key features based on the initial risk analysis model to obtain feature scores; and then assign risk features to the server. The risk feature value is compared with the feature score, and the initial risk analysis model is adjusted and optimized based on the comparison result to obtain the target risk analysis model. The target risk feature value of the risk data is calculated through the target risk analysis model, and the target risk feature value is fed back to the client. This invention provides a feature value analysis device for risk data. For target risk feature value business, the initial risk feature value is corrected by extracting residual feature values. Based on the key features related to the residuals and quantifying them, the residual correction logic is integrated into the initial risk analysis model. The analysis results of the risk feature value and the initial risk analysis model are compared, and the initial risk analysis model is optimized in a targeted manner to eliminate the bias of the initial model, ultimately improving the accuracy of risk scoring. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a separate server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a feature value analysis method for risk data provided in an embodiment of the present invention includes the following steps: S1. Obtain multi-dimensional data of preset risk types, and extract basic features from the multi-dimensional data to obtain basic data features.

[0015] In this embodiment of the invention, the acquisition refers to identifying the "preset risk type" to be analyzed, and then collecting raw data from multiple sources that are directly related to the risk. The basic feature extraction refers to filtering and transforming the collected multi-dimensional raw data to extract "basic information" that can directly reflect the basic attributes of the risk.

[0016] Specifically, the specific type of risk to be addressed (e.g., in the field of auto insurance risk assessment) should be clearly identified. Then, raw data from multiple sources and dimensions should be collected around this risk (e.g., driver's driving history, age, gender, vehicle type, driving area, etc.). Next, key information (basic characteristics) that can initially reflect the basic situation of the risk should be extracted from this raw data. This provides a "clean, usable, and risk-focused" data foundation for subsequent risk analysis (e.g., calculating the initial risk value and identifying key risk factors), avoiding deviations in subsequent analysis due to data clutter or missing information.

[0017] In the healthcare setting, data such as the patient's age, weight, cancer stage, type and dosage of chemotherapy drugs, pre-chemotherapy blood routine indicators, presence of abnormal liver and kidney function, and previous chemotherapy history are collected. From the above data, basic information directly related to "bone marrow suppression risk" is extracted, such as "patient age (≥65 years / <65 years)". Irrelevant data such as the patient's marital status and workplace are removed, and finally, the basic data characteristics required for "bone marrow suppression risk" analysis are obtained.

[0018] In fintech scenarios, user identity information, income data, credit data, consumption data, and other data are collected. From these data, basic information that can initially reflect "overdue risk" is extracted, such as "user age (<22 years old / 22-55 years old / >55 years old)" and "average monthly income (<5000 yuan / 5000-15000 yuan / >15000 yuan)". Non-core data such as user interest tags and browsing history are removed to obtain the basic data characteristics for "consumer loan overdue risk" analysis.

[0019] In this embodiment of the invention, the basic feature extraction of the multi-dimensional data to obtain basic data features includes: The multi-dimensional data is cleaned to obtain standardized data; Basic native features are extracted from the standardized data to obtain basic native features, and basic derived features are constructed based on the basic native features; Redundancy filtering is performed on the basic original features and the basic derived features to obtain the basic data features.

[0020] In this embodiment of the invention, the basic original feature extraction refers to the feature obtained by directly filtering or slightly transforming standardized multi-dimensional data, which can reflect the original attributes of the data and is directly related to the preset risk. The construction refers to the new feature generated based on the basic original features through logical operations, statistical analysis and other methods.

[0021] Specifically, data cleaning of multi-dimensional data includes handling missing values, removing outliers, and standardizing the format of multi-dimensional data to obtain standardized data. Based on the preset risk type, the original fields directly related to the risk are selected from the standardized data. For example, in the field of auto insurance, the driver's driving history, age, gender, vehicle type, driving area, etc., are directly extracted as the basic original features.

[0022] Furthermore, based on the obtained basic original features, new features that can supplement the "risk of accidents" related information are generated through logical operations, statistical integration and other methods. For example, by combining "driver age" and "number of accidents in the past 3 years", "average number of accidents per year = number of accidents in the past 3 years / 3" is calculated to reflect the driver's long-term accident frequency; by combining "driving area" and "longest continuous driving time in a single trip", the basic derived feature of "proportion of driving in high-risk scenarios" is constructed.

[0023] Furthermore, the correlation coefficients (such as Pearson correlation coefficients) between all features were calculated to identify highly correlated feature pairs (e.g., correlation coefficient > 0.8). For example, the correlation coefficient between "number of accidents in the past 3 years" and "average number of accidents per year" reached 0.92 (highly correlated), and "average number of accidents per year" already reflected the long-term accident frequency, so "number of accidents in the past 3 years" was removed. Combining this with the logic of auto insurance business, it was determined that "driver gender" had a weak impact on the accident risk of some vehicle types (such as small cars) (historical data showed that the difference in accident rate caused by gender differences was < 5%), and could not effectively supplement the risk information of other features, so "driver gender" was removed. Finally, the impact of each feature on risk was evaluated through decision trees, analysis of variance, and other methods, and features with extremely low importance (e.g., ranked in the bottom 20%) were removed. It was found that "percentage of driving in high-risk scenarios", "violation rate per vehicle type", "driver age", "vehicle type", and "driving area" ranked in the top 5 in terms of the impact weight on accident risk, and these features were retained.

[0024] In this embodiment of the invention, multi-dimensional data is cleaned to remove outliers and fill in missing values, thus avoiding erroneous data from interfering with subsequent analysis. Basic original features directly reflect the essence of the data, ensuring that risk analysis does not deviate from the original business scenario. Basic derived features can capture the hidden relationships between original features and supplement the risk assessment dimensions. Highly correlated or low-value features (such as "annual revenue") are removed to reduce the computational load of subsequent models and improve the analysis speed.

[0025] In this embodiment of the invention, the scope of data collection is defined by "presetting risk types" and irrelevant data is removed by "basic feature extraction", so that subsequent risk analysis only revolves around "information directly related to the risk", reducing the interference of redundant data on analysis efficiency and accuracy of results.

[0026] S2. Perform basic risk analysis on the basic characteristics of the data to obtain initial risk characteristic values.

[0027] In this embodiment of the invention, the basic risk analysis refers to the process of conducting a preliminary assessment of the target risk based on the selected "basic data characteristics" and using a preset, simple, and universal analysis method.

[0028] Specifically, after extracting the basic data features, the risk level reflected by these basic features is quantitatively assessed using a pre-defined model analysis method, ultimately yielding an "initial risk feature value" that can preliminarily reflect the target risk level.

[0029] In healthcare settings, for example, basic data features include patient age, chemotherapy drug dosage (high / medium / low), preoperative white blood cell count (normal / low), and presence of liver and kidney dysfunction. Risk weights are preset for each feature, such as age ≥65 years (25 points), <65 years (10 points); high-dose chemotherapy (30 points), medium-dose (20 points), and low-dose (10 points). For example, "age 70 years (25 points) + high-dose chemotherapy (30 points)" has a total score of 55 points. Risk levels are assigned based on the total score, with 55 points corresponding to "high risk". Therefore, the initial risk feature value is "high risk of bone marrow suppression (55 points)".

[0030] In fintech scenarios, based on user monthly income, number of overdue payments in the past year, current debt ratio (debt / income), and whether they have stable employment (yes / no), a preset interval scoring rule is established: monthly income ≥ 15,000 yuan (10 points), 5,000-15,000 yuan (20 points), < 5,000 yuan (35 points); 0 overdue payments (10 points), 1-2 overdue payments (30 points), ≥ 3 overdue payments (50 points). The scores corresponding to user characteristics are added together. For example, "monthly income 4,500 yuan (35 points) + 2 overdue payments (30 points)" results in a total score of 65 points. A threshold is set based on the total score. 65 points > 60 points corresponds to "high overdue risk". Therefore, the initial risk characteristic value is "high overdue risk (65 points)".

[0031] In this embodiment of the invention, the step of performing basic risk analysis on the basic data characteristics to obtain initial risk characteristic values ​​includes: The data basic features are subjected to format standardization processing to obtain standardized basic features; Based on preset risk association conditions, the standardized basic features are matched and weighted to obtain a basic risk comprehensive score. The basic risk comprehensive score is calibrated according to the preset business risk quantification standard to obtain the initial risk characteristic value.

[0032] Specifically, clarify the specific content included in the basic data characteristics, such as age, gender, historical risk level, etc. (corresponding to basic covariate factors). Next, corresponding standardization methods are adopted for different types of basic features: For numerical features, such as age, if there are different measurement formats, they are converted into a unified numerical format, and the influence of units is eliminated through normalization or standardization methods; For categorical features, such as gender (male, female), one-hot encoding is used to encode "male" as [1,0] and "female" as [0,1], and the categorical information is converted into a numerical vector form.

[0033] Furthermore, we employ established, mature covariate models that have already been deployed. The standardized basic features are matched with the pre-set risk association rules in the model; for example, when an age falls within a certain range and the historical risk level reaches a certain level, the expected accident frequency and loss amount are determined according to the model rules, and corresponding weights are assigned to each standardized basic feature (these weights are based on the covariate model). (And those determined by business experience), then, the value of each standardized basic feature is multiplied by its corresponding weight, and all the product results are summed to obtain the basic risk comprehensive score. This score integrates the impact of multiple basic features on risk. The specific formula is as follows:

[0034]

[0035] in, Indicates covariate factors. This represents the overall score for basic risk.

[0036] Furthermore, the pre-set business risk quantification standard is formulated in conjunction with specific business scenarios (such as premium pricing and risk level classification in insurance business). First, the mapping relationship between the basic risk comprehensive score and the actual business risk (such as loss amount, accident frequency, etc.) is determined. For example, in the insurance scenario, the higher the basic risk comprehensive score, the higher the expected loss amount and accident frequency. Then, according to business requirements (such as the loss ratio and premium coefficient corresponding to different risk levels), the basic risk comprehensive score is calibrated and adjusted. For example, when the basic risk comprehensive score is in a certain range, it is converted into the corresponding initial risk characteristic value according to business rules.

[0037] In this embodiment of the invention, a unified data format and unit are used to eliminate differences in data from different sources and types, facilitating stable and efficient data processing by subsequent analysis tools or models and avoiding analysis errors caused by format issues. By associating features with risks through preset rules, the weighted representation of the degree of influence of different features on risks can accurately quantify basic risks by comprehensively considering multi-dimensional features, providing more realistic preliminary risk quantification results for subsequent analysis. The risk scores are adjusted according to specific business scenarios, making the risk analysis results more in line with actual business needs and directly serving business decision-making.

[0038] S3. Obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values.

[0039] In this embodiment of the invention, the acquisition refers to the process of obtaining residual feature values ​​from relevant data through a specific method or model.

[0040] Specifically, residual characteristic values ​​are obtained by using residual models and historical residual patterns. These residual characteristic values ​​reflect the degree of deviation between the initial risk characteristic value and the actual risk situation. Then, the initial risk characteristic value and the residual characteristic value are combined to obtain a risk characteristic value that is closer to the true risk level.

[0041] In medical settings, a large amount of historical patient data is collected to calculate the residual between each patient's "actual infection status" and "initial predicted infection probability." Then, models such as gradient boosting decision trees are used to learn the relationship between these residuals and other supplementary patient characteristics (such as operation duration and intraoperative blood loss) to obtain a residual analysis model. For new patients, this model is used to obtain residual feature values. Finally, the initial risk feature values ​​are added to the residual feature values ​​or combined according to weights to obtain more accurate postoperative infection risk feature values.

[0042] In the fintech field, the residuals between the "actual default status" and the "initial default probability prediction" of historical credit users are statistically analyzed. Machine learning models are then used to mine the correlation between the residuals and supplementary features such as user consumption habits and social data to generate a residual analysis model. For new credit applicants, the residual feature values ​​obtained from this model are combined with the initial risk feature values ​​to obtain more accurate default risk feature values, which assist in credit decision-making.

[0043] In this embodiment of the invention, obtaining residual feature values ​​based on preset residuals includes: Based on the basic characteristics of the data, real-time supplementary features are extracted from the multi-dimensional data to obtain real-time supplementary features; Obtain historical supplementary features, and construct a feature residual association table based on the historical supplementary features and preset residuals; The residuals of the real-time supplementary features are matched using the feature residual association table to obtain residual feature values.

[0044] In this embodiment of the invention, the supplementary feature extraction refers to further filtering and extracting features from multi-dimensional data that can compensate for the deficiencies of the basic features and more accurately reflect the risks, based on the existing basic data features. The construction refers to performing correlation analysis between the acquired historical supplementary features and the preset residuals, sorting out the correspondence between different supplementary feature values ​​and residuals, and organizing these patterns into a structured feature residual correlation table.

[0045] Specifically, based on fundamental data characteristics (such as age, gender, and historical risk levels related to basic risks), and guided by these fundamental data characteristics, real-time features that can supplement risk information not covered by the fundamental features are filtered and extracted from multi-dimensional data (covering various risk-related details, such as ADAS (related data) in intelligent driving scenarios). For example, in intelligent driving risk assessment scenarios, ADAS factors (such as response time and recognition accuracy of ADAS systems) are extracted. These extracted features are real-time supplementary features, where ADAS refers to advanced driver assistance systems.

[0046] In detail, historical supplementary features (historical ADAS factors, etc., are obtained, which are consistent with the steps for extracting real-time supplementary features as described above, and will not be repeated here). At the same time, a preset residual is determined (the preset residual is the difference between the initial risk feature value and the actual risk feature value), and it is input into the Gradient Boosting Decision Tree (GBDT) model for training, so that the model learns the association pattern between the supplementary features and the residual. The current real-time supplementary features are then input into the trained GBDT model, and the model predicts the preset residual based on the learned association pattern to obtain the residual feature value.

[0047] GBDT learns the relationship between "features and residuals" by using "historical supplementary features" (features extracted in the past that are related to residuals) and "preset residuals" (the difference between the initial risk feature value and the actual risk feature value). Through this historical data, the model learns the pattern and builds a feature-residual relationship table to reflect the structure of the pattern, providing a data and logical foundation for the subsequent "multiple stump combination learning relationship" in GBDT.

[0048] In detail, GBDT uses a combination of multiple tree stumps, and the specific formula is as follows:

[0049] in, This represents the output of the k-th decision stump, where K is the total number of trees. This represents the residuals predicted using ADAS-related models; This represents a residual prediction model based on the ADAS factor. This represents ADAS factors, i.e., feature variables related to the ADAS system; based on the "feature → residual" relationship (i.e., the pattern carried by the relationship table), the corresponding residual results are matched for the "current supplementary feature," and the final residual feature value output by GBDT is... .

[0050] In this embodiment of the invention, real-time supplementary feature extraction can uncover risk information in multi-dimensional data that is not covered by basic features, enrich the dimensions of risk assessment, and make subsequent risk analysis more comprehensive; by establishing the correlation between supplementary features and residuals through historical data, it provides a reference basis for residual analysis and makes residual prediction more reliable.

[0051] S4. Use a preset residual analysis model to extract the key features of the multi-dimensional data, and assign interval scores to the key features to obtain feature interval scores.

[0052] In this embodiment of the invention, the extraction refers to selecting key features that have a significant impact on the residuals from multi-dimensional data based on a pre-set residual analysis model, and the interval scoring refers to dividing the value range of the extracted key features into different intervals and setting a corresponding score for each interval in advance.

[0053] Specifically, according to the predetermined residual analysis model, features that play a key role in residuals (risk prediction bias) are selected from multi-dimensional data. Then, the values ​​of these key features are divided into different intervals, and each interval is assigned a corresponding score to quantify the degree of residual influence under different values ​​of key features. Finally, feature interval scores are obtained for subsequent more accurate risk analysis and correction.

[0054] In this embodiment of the invention, the step of assigning interval scores to the key features to obtain feature interval scores includes: The key features are extracted using a pre-defined residual analysis model to obtain feature split points; The feature split points are sorted, and the key features are divided into feature value intervals based on the sorted split points to obtain the feature value intervals. The feature value interval is scored and calculated based on the feature residual association table to obtain the feature interval score.

[0055] In this embodiment of the invention, the split point extraction refers to determining, based on a preset residual analysis model, that the value range of key features can be divided into different sub-intervals. The feature value interval division refers to dividing the value range of key features into multiple continuous sub-intervals after sorting according to the extracted split points. The score calculation refers to calculating the score corresponding to each interval for the divided feature value intervals by combining the feature residual association table.

[0056] Specifically, based on a pre-defined residual analysis model (e.g., a residual analysis model based on decision tree stumps, which can determine the split point by analyzing the relationship between residuals and key features), the model analyzes the key features. The model calculates the degree of difference in residuals (the deviation between actual risk and predicted risk) in each sub-interval after the segmentation when different numerical points of the key features are used as split points. Taking the key feature "vehicle mileage" as an example, the model will try different mileage values ​​(e.g., 10,000 km, 20,000 km, etc.) and calculate the variance, mean difference, and other indicators of the residuals in the sub-intervals before and after the split point when the value is used as the split point. Finally, the numerical point that makes the residual difference most significant is selected as the feature split point.

[0057] Furthermore, the extracted feature split points are arranged in ascending order. Then, using these sorted split points as boundaries, the entire value range of the key feature is divided into multiple continuous and non-overlapping sub-intervals. For example, if the split points for "vehicle mileage" are 10,000 km and 30,000 km, and after sorting they are 10,000 and 30,000 km, then the mileage is divided into feature value intervals of (-∞, 10,000], (10,000, 30,000], and (30,000, +∞).

[0058] Furthermore, referring to the feature residual association table, which records the correspondence between different feature value intervals and residuals, as well as the score to be assigned to each interval (the score reflects the degree of influence of the key features on the residuals within the interval), for each defined feature value interval, find the matching interval in the feature residual association table, and then calculate the score of the interval according to the scoring rules in the table.

[0059] In this embodiment of the invention, split point extraction can accurately identify the numerical points in key features that have a significant distinguishing effect on residuals, providing a basis for subsequent interval division and ensuring that the divided intervals can effectively reflect the differences in residuals; discretizing the continuous values ​​of key features into multiple intervals simplifies the representation of key features and facilitates subsequent scoring calculations based on intervals; by combining with the feature residual association table, the residual influence of different intervals of key features is quantified into specific scores, making the degree of residual influence more intuitive and easier to understand, and providing a quantifiable basis for the subsequent correction of risk feature values.

[0060] S5. Construct an initial risk analysis model based on the feature interval scores and the key features, and perform scoring analysis on the key features based on the initial risk analysis model to obtain feature scores.

[0061] In this embodiment of the invention, the construction refers to integrating key features and their corresponding feature interval scores, and organizing them into an initial risk analysis model according to certain logic (such as the structural specifications of the scoring model, the business rules of risk assessment, etc.).

[0062] Specifically, the key features and their corresponding feature interval scores are integrated and an initial risk analysis model is created according to established rules and structure. Then, based on this initial risk analysis model, the actual key feature values ​​are matched to determine the score corresponding to each key feature value. Finally, by processing these scores (such as summation, weighted summation, etc.), a feature score that reflects the risk situation is obtained.

[0063] In detail, the existing key features (such as ADAS emergency braking response time, lane keeping deviation frequency, etc.) and the corresponding feature interval scores for each feature value interval under each key feature are identified (e.g., "emergency braking response time ≤ 0.3 seconds" corresponds to a score of 10, "0.3 seconds - 0.5 seconds" corresponds to a score of 20, etc.). Then, each key feature is integrated with the scores of all its feature value intervals to construct a mapping set of "key feature → feature value interval → feature interval score". Using the "feature-score mapping set" as input, the format and presentation of the initial risk analysis model are planned. In the initial risk analysis model, all key features are listed first. Then, for each key feature, its divided feature value intervals are displayed in sequence, and the feature interval score of each interval is labeled accordingly. At the same time, the rules for using the scores of each key feature in risk analysis are clarified (e.g., whether weighted summation is required, only the rule framework needs to be clarified). Finally, an initial risk analysis model containing key features, feature value intervals, feature interval scores, and usage rules is formed.

[0064] In this embodiment of the invention, the step of scoring and analyzing the key features based on the initial risk analysis model to obtain feature scores includes: The key features are extracted to obtain their actual numerical values. The initial risk analysis model is used to perform interval matching on the actual values ​​of the key features to obtain the matching interval; The initial risk analysis model is used to find the corresponding score for the matching interval to obtain the feature score.

[0065] In this embodiment of the invention, the actual numerical extraction refers to obtaining the true values ​​of key features from specific sample data.

[0066] Specifically, from multi-dimensional data of preset risk types, the true values ​​of key characteristics are obtained, using covariates (such as age, gender, historical risk level, etc.) and ADAS factors as in the example above. For example, when assessing the risk of a vehicle, the actual value of the "response time of the ADAS system" extracted from the vehicle's usage records is 0.5 seconds, and the actual value of the "age" extracted from the owner's information is 35 years old. These are the actual values ​​of key features.

[0067] Furthermore, the actual values ​​of the extracted key features are compared with the corresponding intervals of the key features in the initial risk analysis model to find the interval to which the actual value belongs. For example, in the initial risk analysis model, the intervals for "ADAS system response time" are divided into (-∞, 0.3] seconds, (0.3, 0.6] seconds, and (0.6, +∞) seconds. If the actual value is 0.5 seconds, then the matching interval is (0.3, 0.6] seconds.

[0068] Furthermore, each key feature in the initial risk analysis model corresponds to a specific score in different intervals. After determining the matching interval, the score corresponding to the matching interval is found in the initial risk analysis model. Combining the additive form of the initial risk analysis model, the final risk prediction value is the sum of the scores corresponding to each key feature (covariate and ADAS factor), which not only ensures the predictive ability but also explains the contribution of each feature.

[0069] S6. Compare the risk feature value with the feature score, and adjust and optimize the initial risk analysis model according to the comparison result to obtain the target risk analysis model.

[0070] In this embodiment of the invention, the comparison refers to comparing the risk feature values ​​obtained by the initial risk analysis model plus residual analysis with the feature scores obtained by scoring key features only by the initial risk analysis model. The adjustment and optimization refers to modifying and improving the key elements of the initial risk analysis model based on the comparison difference between the risk feature values ​​and the feature scores.

[0071] Specifically, the risk characteristic values ​​obtained by comprehensively considering factors such as basic risk and residual correction are compared with the characteristic scores obtained by scoring key characteristics solely through the initial risk analysis model. Based on the differences between the two, the key characteristic interval division and score setting of the initial risk analysis model are adjusted and improved, ultimately resulting in a target risk analysis model that more accurately reflects risk, is consistent with the prediction results of the original GBDT model, and has strong interpretability.

[0072] In this embodiment of the invention, adjusting and optimizing the initial risk analysis model based on the comparison results to obtain the target risk analysis model includes: Identify the differences between the risk feature value and the feature score based on the comparison results; Based on the differences, construct a table of reasons for the differences, and determine the adjustment plan for the initial risk analysis model based on the table of reasons for the differences. The initial risk analysis model is initially updated according to the adjustment plan to obtain the risk analysis model to be verified. The target risk analysis model is obtained by effectively validating the risk analysis model to be verified based on the key features.

[0073] In this embodiment of the invention, the identification refers to finding specific samples or feature dimensions that have significant differences in numerical value, trend, or risk level by comparing risk feature values. The construction refers to sorting out the possible causes of the differences for the identified differences and organizing the correspondence between "difference points - potential causes - related features" into a structured table. The initial update refers to modifying the key elements of the initial risk analysis model according to the adjustment plan.

[0074] Specifically, risk characteristic values ​​(such as a vehicle's comprehensive risk score of 85 points, which incorporates basic characteristics) are used to... Initial predictions and ADAS factors The residual correction) and the feature score (the same vehicle score obtained based on the initial risk analysis model is 60 points) are compared on a sample-by-sample or feature-dimension basis. By calculating the difference (85-60=25 points) and the statistical difference frequency (e.g., 30% of the samples have a difference of more than 20 points), samples with significant differences (e.g., vehicles with short ADAS system response time but low initial score) and feature dimensions are identified.

[0075] Furthermore, for the identified discrepancies, the reasons were analyzed: for example, the initial interval division of the "ADAS response time" feature was too coarse (only <0.5 seconds and ≥0.5 seconds), failing to reflect the low residual characteristics of the 0.3-0.5 second interval, resulting in a low feature score. These issues were compiled into a discrepancy reason analysis table; based on the reasons in the table, adjustment plans were determined, such as refining the split point of "ADAS response time" from 0.5 seconds to 0.3 seconds and 0.6 seconds, adding a (0.3, 0.6] interval and increasing the corresponding score.

[0076] Furthermore, the initial risk analysis model is modified according to the adjustment plan. The new risk analysis model to be verified is tested using historical data of key features (such as ADAS response time and age of a large number of vehicles). The difference between the feature scores and risk feature values ​​based on the initial risk analysis model is calculated. If the difference rate drops from 30% to below 5% (meeting the preset threshold standard), and each feature score can accurately reflect the residual pattern (such as the score in the (0.3-0.6) second interval being consistent with the residual correction result of the ADAS factor), then the verification is passed, and the risk analysis model to be verified is upgraded to the target risk analysis model. If it fails, the reasons are re-analyzed and adjusted to finally obtain the target risk analysis model.

[0077] S7. Calculate the target risk characteristic value of the risk data using the target risk analysis model.

[0078] Specifically, pre-defined risk data (multi-dimensional data adapted to the target risk analysis model, such as user credit data, vehicle driving data, etc.) is collected and pre-processed according to model requirements (including data cleaning, format conversion, missing value imputation, outlier handling, etc.) to ensure that the data conforms to the model input specifications. From the pre-processed risk data, key features required by the target risk analysis model are extracted (referring to the core risk factors determined by the model during training and optimization), and the feature values ​​are mapped to the intervals or formats specified by the model. The matched key features are input into the target risk analysis model. The model performs calculations on the input features according to its internally fixed algorithm logic (such as the score weights of each feature interval, combination calculation rules, etc.). For example, the model may use a weighted summation method to accumulate the interval scores of each key feature according to preset weights, and finally obtain the target risk feature value. In this invention, the risk feature value specifically refers to the risk feature value in the field of auto insurance risk assessment.

[0079] As can be seen, in the above scheme, for the target risk characteristic value business, multi-dimensional data of a preset risk type is obtained, and basic features are extracted from the multi-dimensional data to obtain basic data features; basic risk analysis is performed on the basic data features to obtain initial risk characteristic values; residual characteristic values ​​are obtained according to preset residuals, and the initial risk characteristic values ​​and residual characteristic values ​​are combined to generate risk characteristic values; key features of the multi-dimensional data are extracted using a preset residual analysis model, and interval scoring is performed on the key features to obtain feature interval scores; an initial risk analysis model is constructed based on the feature interval scores and the key features, and the initial risk characteristic value is determined based on the initial risk characteristic value. The initial risk analysis model scores the key features to obtain feature scores. The risk feature values ​​are compared with the feature scores, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain the target risk analysis model. The target risk feature values ​​of the risk data are calculated using the target risk analysis model. The initial risk feature values ​​are corrected by extracting residual feature values. Based on the key features related to the residuals and quantifying them, the residual correction logic is integrated into the initial risk analysis model. The analysis results of the risk feature values ​​and the initial risk analysis model are compared, and the initial risk analysis model is optimized in a targeted manner to eliminate the bias of the initial model, ultimately improving the accuracy of risk scoring.

[0080] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0081] In one embodiment, a feature value analysis device for risk data is provided, which corresponds one-to-one with the feature value analysis method for risk data in the above embodiments. For example... Figure 3 As shown, this feature value analysis device for risk data includes an acquisition and extraction module 101, an analysis module 102, an acquisition and combination module 103, an extraction and scoring module 104, a construction and analysis module 105, a comparison and optimization module 106, and a calculation module 107. Detailed descriptions of each functional module are as follows: The acquisition and extraction module 101 is used to acquire multi-dimensional data of a preset risk type and extract basic features from the multi-dimensional data to obtain basic data features. Analysis module 102 is used to perform basic risk analysis on the basic characteristics of the data to obtain initial risk characteristic values; The acquisition and combination module 103 is used to acquire residual feature values ​​according to preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The extraction and scoring module 104 is used to extract the key features of the multi-dimensional data using a preset residual analysis model, and to assign interval scores to the key features to obtain feature interval scores. The analysis module 105 is used to construct an initial risk analysis model based on the feature interval scores and the key features, and to perform scoring analysis on the key features based on the initial risk analysis model to obtain feature scores; The comparison and optimization module 106 is used to compare the risk feature value with the feature score, and adjust and optimize the initial risk analysis model according to the comparison result to obtain the target risk analysis model; The calculation module 107 is used to calculate the target risk feature value of the risk data through the target risk analysis model.

[0082] In one embodiment, the extraction module 101, when performing basic feature extraction on the multi-dimensional data to obtain basic data features, is used for: The multi-dimensional data is cleaned to obtain standardized data; Basic native features are extracted from the standardized data to obtain basic native features, and basic derived features are constructed based on the basic native features; Redundancy filtering is performed on the basic original features and the basic derived features to obtain the basic data features.

[0083] In one embodiment, when the analysis module 102 performs basic risk analysis on the basic characteristics of the data to obtain initial risk characteristic values, it is used to: The data basic features are subjected to format standardization processing to obtain standardized basic features; Based on preset risk association conditions, the standardized basic features are matched and weighted to obtain a basic risk comprehensive score. The basic risk comprehensive score is calibrated according to the preset business risk quantification standard to obtain the initial risk characteristic value.

[0084] In one embodiment, when the acquisition module 103 acquires residual feature values ​​based on preset residuals, it is used to: Based on the basic characteristics of the data, real-time supplementary features are extracted from the multi-dimensional data to obtain real-time supplementary features; Obtain historical supplementary features, and construct a feature residual association table based on the historical supplementary features and preset residuals; The residuals of the real-time supplementary features are matched using the feature residual association table to obtain residual feature values.

[0085] In one embodiment, the scoring module 104, when performing interval scoring on the key features to obtain feature interval scores, is used for: The key features are extracted using a pre-defined residual analysis model to obtain feature split points; The feature split points are sorted, and the key features are divided into feature value intervals based on the sorted split points to obtain the feature value intervals. The feature value interval is scored and calculated based on the feature residual association table to obtain the feature interval score.

[0086] In one embodiment, when the analysis module 105 performs scoring analysis on the key features based on the initial risk analysis model to obtain feature scores, it is used to: The key features are extracted to obtain their actual numerical values. The initial risk analysis model is used to perform interval matching on the actual values ​​of the key features to obtain the matching interval; The initial risk analysis model is used to find the corresponding score for the matching interval to obtain the feature score.

[0087] In one embodiment, when the comparison optimization module 106 adjusts and optimizes the initial risk analysis model based on the comparison results to obtain the target risk analysis model, it is used to: Identify the differences between the risk feature value and the feature score based on the comparison results; Based on the differences, construct a table of reasons for the differences, and determine the adjustment plan for the initial risk analysis model based on the table of reasons for the differences. The initial risk analysis model is initially updated according to the adjustment plan to obtain the risk analysis model to be verified. The target risk analysis model is obtained by effectively validating the risk analysis model to be verified based on the key features.

[0088] This invention provides a feature value analysis device for risk data. For a target risk feature value business, it acquires multi-dimensional data of a preset risk type and extracts basic features from the multi-dimensional data to obtain basic data features; performs basic risk analysis on the basic data features to obtain initial risk feature values; obtains residual feature values ​​based on preset residuals; combines the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; extracts key features from the multi-dimensional data using a preset residual analysis model and assigns interval scores to the key features to obtain feature interval scores; and constructs an initial risk analysis model based on the feature interval scores and the key features. The key features are scored and analyzed according to the initial risk analysis model to obtain feature scores. The risk feature values ​​are compared with the feature scores, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain a target risk analysis model. The target risk feature values ​​of the risk data are calculated through the target risk analysis model. The initial risk feature values ​​are corrected by extracting residual feature values. Based on the key features related to the residuals and quantifying them, the residual correction logic is integrated into the initial risk analysis model. The analysis results of the risk feature values ​​and the initial risk analysis model are compared, and the initial risk analysis model is optimized in a targeted manner to eliminate the bias of the initial model, ultimately improving the accuracy of risk scoring.

[0089] For specific limitations regarding the feature value analysis device for risk data, please refer to the limitations of the feature value analysis method for risk data described above, which will not be repeated here. Each module in the aforementioned feature value analysis device for risk data can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0090] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the server-side functions or steps of a risk data feature analysis method.

[0091] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a risk data feature analysis method.

[0092] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire multi-dimensional data of preset risk types, and extract basic features from the multi-dimensional data to obtain basic data features; A basic risk analysis is performed on the basic characteristics of the data to obtain initial risk characteristic values; Obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The key features of the multi-dimensional data are extracted using a preset residual analysis model, and the key features are assigned interval scores to obtain feature interval scores. An initial risk analysis model is constructed based on the feature interval scores and the key features, and the key features are scored and analyzed based on the initial risk analysis model to obtain feature scores; The risk feature value is compared with the feature score, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain the target risk analysis model; The target risk characteristic value of the risk data is calculated using the target risk analysis model.

[0093] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire multi-dimensional data of preset risk types, and extract basic features from the multi-dimensional data to obtain basic data features; A basic risk analysis is performed on the basic characteristics of the data to obtain initial risk characteristic values; Obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The key features of the multi-dimensional data are extracted using a preset residual analysis model, and the key features are assigned interval scores to obtain feature interval scores. An initial risk analysis model is constructed based on the feature interval scores and the key features, and the key features are scored and analyzed based on the initial risk analysis model to obtain feature scores; The risk feature value is compared with the feature score, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain the target risk analysis model; The target risk characteristic value of the risk data is calculated using the target risk analysis model.

[0094] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0095] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0096] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0097] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. If any software tools or components other than those of our company appear in the embodiments, they are merely illustrative examples and do not represent actual use. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for eigenvalue analysis of risk data, characterized in that, include: Acquire multi-dimensional data of preset risk types, and extract basic features from the multi-dimensional data to obtain basic data features; A basic risk analysis is performed on the basic characteristics of the data to obtain initial risk characteristic values; Obtain residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The key features of the multi-dimensional data are extracted using a preset residual analysis model, and the key features are assigned interval scores to obtain feature interval scores. An initial risk analysis model is constructed based on the feature interval scores and the key features, and the key features are scored and analyzed based on the initial risk analysis model to obtain feature scores; The risk feature value is compared with the feature score, and the initial risk analysis model is adjusted and optimized based on the comparison results to obtain the target risk analysis model; The target risk characteristic value of the risk data is calculated using the target risk analysis model.

2. The feature value analysis method for risk data as described in claim 1, characterized in that, The basic feature extraction of the multi-dimensional data to obtain basic data features includes: The multi-dimensional data is cleaned to obtain standardized data; Basic native features are extracted from the standardized data to obtain basic native features, and basic derived features are constructed based on the basic native features; Redundancy filtering is performed on the basic original features and the basic derived features to obtain the basic data features.

3. The eigenvalue analysis method for risk data as described in claim 1, characterized in that, The basic risk analysis of the data's fundamental characteristics to obtain initial risk characteristic values ​​includes: The data basic features are subjected to format standardization processing to obtain standardized basic features; Based on preset risk association conditions, the standardized basic features are matched and weighted to obtain a basic risk comprehensive score. The basic risk comprehensive score is calibrated according to the preset business risk quantification standard to obtain the initial risk characteristic value.

4. The eigenvalue analysis method for risk data as described in claim 1, characterized in that, The step of obtaining residual feature values ​​based on preset residuals includes: Based on the basic characteristics of the data, real-time supplementary features are extracted from the multi-dimensional data to obtain real-time supplementary features; Obtain historical supplementary features, and construct a feature residual association table based on the historical supplementary features and preset residuals; The residuals of the real-time supplementary features are matched using the feature residual association table to obtain residual feature values.

5. The eigenvalue analysis method for risk data as described in claim 4, characterized in that, The step of assigning interval scores to the key features to obtain feature interval scores includes: The key features are extracted using a pre-defined residual analysis model to obtain feature split points; The feature split points are sorted, and the key features are divided into feature value intervals based on the sorted split points to obtain the feature value intervals. The feature value interval is scored and calculated based on the feature residual association table to obtain the feature interval score.

6. The eigenvalue analysis method for risk data as described in claim 1, characterized in that, The step of scoring and analyzing the key features based on the initial risk analysis model to obtain feature scores includes: The key features are extracted to obtain their actual numerical values. The initial risk analysis model is used to perform interval matching on the actual values ​​of the key features to obtain the matching interval; The initial risk analysis model is used to find the corresponding score for the matching interval to obtain the feature score.

7. The eigenvalue analysis method for risk data as described in claim 1, characterized in that, The step of adjusting and optimizing the initial risk analysis model based on the comparison results to obtain the target risk analysis model includes: Identify the differences between the risk feature value and the feature score based on the comparison results; Based on the differences, construct a table of reasons for the differences, and determine the adjustment plan for the initial risk analysis model based on the table of reasons for the differences. The initial risk analysis model is initially updated according to the adjustment plan to obtain the risk analysis model to be verified. The target risk analysis model is obtained by effectively validating the risk analysis model to be verified based on the key features.

8. A feature value analysis device for risk data, characterized in that, include: The acquisition and extraction module is used to acquire multi-dimensional data of a preset risk type and extract basic features from the multi-dimensional data to obtain basic data features; The analysis module is used to perform basic risk analysis on the basic characteristics of the data to obtain initial risk characteristic values; The acquisition and combination module is used to acquire residual feature values ​​based on preset residuals, and combine the initial risk feature values ​​with the residual feature values ​​to generate risk feature values; The extraction and scoring module is used to extract key features of the multi-dimensional data using a preset residual analysis model, and to assign interval scores to the key features to obtain feature interval scores. An analysis module is constructed to build an initial risk analysis model based on the feature interval scores and the key features, and to perform scoring analysis on the key features based on the initial risk analysis model to obtain feature scores; The comparison and optimization module is used to compare the risk feature value with the feature score, and adjust and optimize the initial risk analysis model according to the comparison result to obtain the target risk analysis model; The calculation module is used to calculate the target risk feature value of the risk data through the target risk analysis model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the eigenvalue analysis method for risk data as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the eigenvalue analysis method for risk data as described in any one of claims 1 to 7.