User processing method, device, equipment and medium based on neural network model

Through the user processing method based on neural network model, the problem of lack of reliable information when manually evaluating the risks of insured users is solved, and more accurate risk assessment and classification are achieved.

CN119006181BActive Publication Date: 2025-05-23ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411472075.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-05-23
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In the prior art, the lack of reliable risk information when manually evaluating the risks of the insured users is likely to lead to evaluation errors without knowing them.

Method used

The user processing method based on the neural network model is adopted to obtain and standardize the data of the insured users, input it into the trained neural network model for inference, obtain the user's risk value, and improve the accuracy of the model by screening and removing historical user data.

Benefits of technology

It improves the accuracy of risk assessment for insured users, reduces the possibility of manual assessment errors, and enhances the reliability of evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006181B_ABST
    Figure CN119006181B_ABST
Patent Text Reader

Abstract

The present application relates to the field of electronic digital data processing technology, and in particular to a user processing method, device, equipment and medium based on a neural network model. The method comprises: obtaining a user data set to be insured, which includes initial data of several users to be insured; performing standardization processing on the initial data of each user to be insured in the user data set to be insured, and obtaining standardized data of each user to be insured in the user data set to be insured; inputting the standardized data of each user to be insured in the user data set to be insured into a trained neural network model for inference, and obtaining the risk value of each user to be insured in the user data set to be insured; and classifying the users to be insured in the user data set to be insured according to the risk value of each user to be insured in the user data set to be insured and several preset risk value thresholds. The present invention can obtain the high and low risk categories of users to be insured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a user processing method, device, equipment and medium based on a neural network model. Background Art

[0002] In the prior art, when determining the premium of the user to be insured, the risk of the user to be insured is often evaluated first, and the premium is divided into several fixed levels, and the premium levels corresponding to different risks of the users to be insured are different; the existing method for evaluating the risk of the user to be insured is usually: manually evaluate the risk of the user to be insured based on the personal information of the user to be insured, but because there is no risk information of the user to be insured for reference when manually evaluating the risk of the user to be insured, it may happen that errors are made in the manual risk assessment of the user to be insured (that is, the assessed risk value does not match the personal information) without the user knowing it; how to obtain the high and low risk categories of the users to be insured so that when manually evaluating the risk of the users to be insured, the evaluation results can be verified according to the high and low risk categories of the users to be insured, so as to avoid the situation where errors are made in the manual risk assessment of the users to be insured without the user knowing it, is a problem that needs to be solved urgently. Summary of the invention

[0003] The purpose of the present invention is to provide a user processing method, device, equipment and medium based on a neural network model to obtain the risk level category of the user to be insured, so that when the risk of the user to be insured is manually assessed, the assessment result can be verified according to the risk level category of the user to be insured, thereby avoiding the situation where errors occur without knowing it when the risk of the user to be insured is manually assessed.

[0004] According to a first aspect of the present invention, a user processing method based on a neural network model is provided, the processing method comprising the following steps:

[0005] A user data set to be insured is obtained, wherein the user data set to be insured includes initial data of a plurality of users to be insured, and the initial data of each user to be insured includes a plurality of attribute information of the corresponding user to be insured.

[0006] The initial data of each user to be insured in the user data set to be insured is standardized to obtain the standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured.

[0007] The standardized data of each user to be insured in the user data set to be insured is input into the trained neural network model for inference to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0008] The users to be insured in the user data set to be insured are classified according to the risk value of each user to be insured in the user data set to be insured and a plurality of preset risk value thresholds.

[0009] Furthermore, the process of obtaining the training user set includes:

[0010] A historical user set is obtained, wherein the historical user set includes a plurality of historical users, and the standardized data of each historical user includes a plurality of standardized attribute information of the corresponding historical user.

[0011] The historical users in the historical user set are grouped according to the corresponding standardized data, so that the standardized attribute information of each historical user included in each group is the same or belongs to the same preset range.

[0012] The historical users in each group are divided into several clusters according to the distance of the standardized data of the historical users in the group.

[0013] According to the risk values ​​of historical users in each cluster, historical users with abnormal risk values ​​in the cluster are screened out.

[0014] The historical user set consisting of the remaining historical users in all clusters is determined as the training user set.

[0015] Furthermore, screening out historical users with abnormal risk values ​​in the cluster according to the risk values ​​of historical users in each cluster includes:

[0016] The number of historical users in the first cluster is obtained; the first cluster is any cluster obtained by dividing the historical users in any group.

[0017] Get the variance of the risk values ​​of historical users in the first cluster.

[0018] The number of historical users in the first cluster, the average number of historical users in all clusters, the variance of risk values ​​of historical users in the first cluster, and the average variance of risk values ​​of historical users in all clusters are matched in a preset screening ratio table; the preset screening ratio table includes several entries, each entry includes a quantity range, an average quantity range, a variance range, an average variance range and a screening ratio.

[0019] If the match is successful, the screening ratio of the successfully matched entries is determined as the screening ratio of the first cluster.

[0020] The historical users in the first cluster are sorted in ascending order according to the corresponding risk values, and the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster are screened out from the first cluster; the first percentage is half of the screening ratio of the first cluster.

[0021] Furthermore, the historical users in each group are divided into several clusters according to the distance of the standardized data of the historical users in the group, including:

[0022] The standardized data of a first historical user is obtained; the first historical user is any user in any group.

[0023] The standardized data of a second historical user is obtained; the second historical user belongs to the same group as the first historical user, and the second historical user is different from the first historical user.

[0024] The distance between the standardized data of the first historical user and the second historical user is obtained according to the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained according to the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

[0025] Furthermore, the corresponding weight acquisition process includes:

[0026] According to the standardized data of historical users in the historical user set and the corresponding risk values, the Pearson correlation coefficient of each standardized attribute information and the risk value is obtained.

[0027] The absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value is normalized.

[0028] The absolute value of each normalized Pearson correlation coefficient is determined as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

[0029] According to a second aspect of the present invention, there is provided a user processing device based on a neural network model, the processing device comprising:

[0030] The first acquisition module is used to acquire a user data set to be insured, wherein the user data set to be insured includes initial data of a number of users to be insured, and the initial data of each user to be insured includes a number of attribute information of the corresponding user to be insured.

[0031] The standardization processing module is used to standardize the initial data of each user to be insured in the user data set to be insured, and obtain the standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured.

[0032] The inference module is used to input the standardized data of each user to be insured in the user data set to be insured into the trained neural network model for inference, so as to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user according to the input standardized data of the user; the trained neural network model is obtained according to the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, and the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0033] The classification module is used to classify the users to be insured in the user data set according to the risk value of each user to be insured in the user data set and a plurality of preset risk value thresholds.

[0034] Furthermore, the processing device further includes a training user set acquisition module, and the training user set acquisition module includes:

[0035] The second acquisition module is used to acquire a historical user set, wherein the historical user set includes a number of historical users, and the standardized data of each historical user includes a number of standardized attribute information of the corresponding historical user.

[0036] The grouping module is used to group the historical users in the historical user set according to the corresponding standardized data, so that the standardized attribute information of each historical user included in each group is the same or belongs to the same preset range.

[0037] The clustering module is used to divide the historical users in each group into several clusters according to the distance of the standardized data of the historical users in the group.

[0038] The screening module is used to screen out historical users with abnormal risk values ​​in a cluster according to the risk values ​​of historical users in each cluster.

[0039] The first determination module is used to determine a historical user set consisting of the remaining historical users in all clusters as a training user set.

[0040] Furthermore, the screening module includes:

[0041] The third acquisition module is used to acquire the number of historical users in a first cluster; the first cluster is any cluster obtained by dividing the historical users in any group.

[0042] The fourth acquisition module is used to obtain the variance of the risk values ​​of historical users in the first cluster.

[0043] A matching module is used to match the number of historical users in the first cluster, the average number of historical users in all clusters, the variance of the risk values ​​of historical users in the first cluster, and the average variance of the risk values ​​of historical users in all clusters in a preset screening ratio table; the preset screening ratio table includes a plurality of entries, each entry includes a quantity range, an average quantity range, a variance range, an average variance range and a screening ratio.

[0044] The second determination module is configured to determine the screening ratio of the successfully matched entries as the screening ratio of the first cluster if the match is successful.

[0045] A sorting module is used to sort the historical users in the first cluster in ascending order of corresponding risk values, and to screen out the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster from the first cluster; the first percentage is half of the screening ratio of the first cluster.

[0046] Furthermore, the clustering module includes:

[0047] The fifth acquisition module is used to acquire the standardized data of the first historical user; the first historical user is any user in any group.

[0048] The sixth acquisition module is used to acquire standardized data of a second historical user; the second historical user and the first historical user belong to the same group, and the second historical user is different from the first historical user.

[0049] A distance acquisition module is used to obtain the distance between the standardized data of the first historical user and the second historical user based on the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained based on the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

[0050] Furthermore, the clustering module further includes a weight acquisition module, and the weight acquisition module includes:

[0051] The seventh acquisition module is used to obtain the Pearson correlation coefficient between each standardized attribute information and the risk value according to the standardized data of the historical users in the historical user set and the corresponding risk value.

[0052] The normalization processing module is used to normalize the absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value.

[0053] The third determination module is used to determine the absolute value of each normalized Pearson correlation coefficient as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

[0054] According to a third aspect of the present invention, there is provided an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned user processing method based on a neural network model when executing the computer program.

[0055] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which implements the above-mentioned user processing method based on the neural network model when executed by a processor.

[0056] Compared with the prior art, the present invention has at least the following beneficial effects:

[0057] The present invention uses a trained neural network model to infer the standardized data of each user to be insured in the user data set to be insured, and obtains the risk value of each user to be insured in the user data set to be insured; because the training user set of the trained neural network model of the present invention is obtained by screening the historical user set, the reasoning accuracy of the trained neural network model is relatively high; and the number of abnormal users screened out in each cluster of the present invention is related to the number of historical users in the cluster and the variance of the risk value in the cluster, and is also related to the average number of historical users in all clusters and the average variance of the risk value in all clusters, which is conducive to reducing the difference in the number of historical users in each cluster after screening and retaining historical users with relatively accurate risk values ​​in each cluster, which is conducive to improving the accuracy of the risk values ​​of the training users in the training user set, and then it is conducive to improving the accuracy of the trained neural network model and improving the accuracy of the final classification of the users to be insured. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0059] Figure 1 A flowchart of a user processing method based on a neural network model provided in Embodiment 1 of the present invention;

[0060] Figure 2 A flowchart of the steps of the process of acquiring a training user set provided in the first embodiment of the present invention;

[0061] Figure 3 A flowchart of the steps of dividing historical users in each group into a plurality of clusters according to the distance of the standardized data of the historical users in the group provided in the first embodiment of the present invention;

[0062] Figure 4 A flowchart of the steps of obtaining the distance between the standardized data of the first historical user and the second historical user according to the standardized data of the first historical user and the standardized data of the second historical user provided in the first embodiment of the present invention;

[0063] Figure 5 A flowchart of the steps of screening out historical users with abnormal risk values ​​in a cluster according to the risk values ​​of historical users in each cluster provided in the first embodiment of the present invention;

[0064] Figure 6 A schematic diagram of a user processing device based on a neural network model provided in Embodiment 2 of the present invention;

[0065] Figure 7 A schematic diagram of a training user set acquisition module provided in Embodiment 2 of the present invention;

[0066] Figure 8 A schematic diagram of a clustering module provided in Embodiment 2 of the present invention;

[0067] Fig. 9 A schematic diagram of a weight acquisition module provided in Embodiment 2 of the present invention;

[0068] Fig.10 This is a schematic diagram of a screening module provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION

[0069] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0070] Embodiment 1:

[0071] According to this embodiment, Figure 1 As shown, a user processing method based on a neural network model is provided, and the processing method comprises the following steps:

[0072] S100, obtaining a user data set to be insured, wherein the user data set to be insured includes initial data of a number of users to be insured, and the initial data of each user to be insured includes a number of attribute information of the corresponding user to be insured.

[0073] As a preferred specific implementation manner, the plurality of attribute information includes age, gender, body mass index, chronic disease history, smoking habits, drinking habits, exercise frequency, occupation type and income.

[0074] S200, standardize the initial data of each user to be insured in the user data set to obtain standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured.

[0075] As a preferred specific implementation, the plurality of standardized attribute information includes standardized age, a first label for characterizing gender, a standardized body mass index, a second label for characterizing the presence or absence of chronic diseases, a standardized smoking frequency, a standardized drinking frequency, a standardized exercise frequency, a third label for characterizing whether an occupation is a high-risk occupation, and standardized income.

[0076] Optionally, the values ​​of the first label, the second label and the third label are 0 or 1; when the first label is 0, it indicates that the gender is male; when the first label is 1, it indicates that the gender is female; when the second label is 0, it indicates that there is no chronic disease; when the second label is 1, it indicates that there is a chronic disease; when the third label is 0, it indicates that the occupation is a low-risk occupation; when the third label is 1, it indicates that the occupation is a high-risk occupation.

[0077] Optionally, the value range of the standardized age, standardized body mass index, standardized smoking frequency, standardized drinking frequency, standardized exercise frequency and standardized income are all 0-1; the standardized age is the ratio of the user's age to a preset maximum age, the standardized body mass index is the ratio of the user's body mass index to a preset maximum body mass index, the standardized smoking frequency is the ratio of the number of days the user smokes per week to 7, the standardized drinking frequency is the ratio of the number of days the user drinks per week to 7, the standardized exercise frequency is the ratio of the number of days the user exercises per week to 7, and the standardized income is the ratio of the user's income to the preset maximum income.

[0078] S300, input the standardized data of each user to be insured in the user data set to be insured into the trained neural network model for inference, and obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, and the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0079] As a preferred specific implementation, the risk value of any historical user is the sum of the first ratio and the second ratio of the corresponding historical user, the first ratio being the ratio of the actual number of claims in the historical time period to the preset number of claims, and the second ratio being the ratio of the claim amount in the historical time period to the preset claim amount. Optionally, the historical time period is the most recent 1 year.

[0080] As a preferred embodiment, Figure 2 As shown, the process of obtaining the training user set includes:

[0081] S310, obtaining a historical user set, wherein the historical user set includes a number of historical users, and the standardized data of each historical user includes a number of standardized attribute information of the corresponding historical user.

[0082] S320, grouping the historical users in the historical user set according to the corresponding standardized data, so that the standardized attribute information of each historical user included in each group is the same or belongs to the same preset range.

[0083] As a preferred specific implementation, the value range of each standardized attribute information is divided into several sub-ranges. If the sub-ranges to which all the standardized attribute information of two historical users belong are the same, then the two historical users belong to the same group. For example, if the value range of each standardized attribute information is divided into two sub-ranges, then the historical users in the historical user set can be divided into two sub-ranges. n Based on this preferred embodiment, the historical users can be initially divided only by judging whether each standardized attribute information of each historical user conforms to the corresponding sub-range, with relatively fast division speed and high division efficiency.

[0084] As a specific implementation, the historical users included in the same group have their standardized ages belonging to the same preset age range, the first label used to characterize gender is the same, the standardized body mass index belongs to the same preset body mass index range, the second label used to characterize the presence or absence of chronic diseases is the same, the standardized smoking frequency belongs to the same preset smoking frequency range, the standardized drinking frequency belongs to the same preset drinking frequency range, the standardized exercise frequency belongs to the same preset exercise frequency range, the third label used to characterize whether the occupation is a high-risk occupation is the same, and the standardized income belongs to the same preset income range.

[0085] S330 , dividing the historical users in each group into a plurality of clusters according to the distances of the standardized data of the historical users in the group.

[0086] In this embodiment, each historical user in the same group is taken as an object to be clustered, and the k-means method is used to realize the division of historical users in the same group, where k is an empirical value; in the process of using the k-means method to divide historical users in the same group, it is necessary to obtain the distance between any two historical users. As a preferred specific implementation method, Figure 3 As shown, S330 includes:

[0087] S331, obtaining standardized data of a first historical user; the first historical user is any user in any group.

[0088] S332, obtaining standardized data of a second historical user; the second historical user and the first historical user belong to the same group, and the second historical user is different from the first historical user.

[0089] S333, obtaining the distance between the standardized data of the first historical user and the second historical user based on the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained based on the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

[0090] In this embodiment, the weighted sum formula is used to obtain the distance between the standardized data of the first historical user and the second historical user, and the difference value of any standardized attribute information of the first historical user and the second historical user is the absolute value of the difference between the standardized attribute information of the first historical user and the second historical user. As a preferred specific implementation method, Figure 4 As shown, S333 includes:

[0091] S3331, obtaining the Pearson correlation coefficient between each standardized attribute information and the risk value according to the standardized data of the historical users in the historical user set and the corresponding risk value.

[0092] Those skilled in the art know that any method for obtaining the Pearson correlation coefficient in the prior art falls within the protection scope of the present invention.

[0093] S3332, normalize the absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value.

[0094] S3333, the absolute value of each normalized Pearson correlation coefficient is determined as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

[0095] Based on this preferred specific implementation, the weight corresponding to the standardized attribute information that is more relevant to the risk value is larger. Therefore, the risk values ​​of historical users in the same cluster are closer, which is conducive to improving the accuracy of subsequent screening of historical users with abnormal risk values ​​based on the risk values ​​of historical users in the cluster.

[0096] S340: Screen out historical users with abnormal risk values ​​in the cluster according to the risk values ​​of the historical users in each cluster.

[0097] In this embodiment, the historical user with abnormal risk value in the cluster refers to the historical user with a large difference in risk value from other historical users in the cluster; as a preferred specific implementation method, Figure 5 As shown, S340 includes:

[0098] S341, obtaining the number of historical users in a first cluster; the first cluster is any cluster obtained by dividing historical users in any group.

[0099] S342, obtaining the variance ratio coefficient of the risk value of the historical users in the first cluster.

[0100] Those skilled in the art know that any method for obtaining variance in the prior art falls within the protection scope of the present invention.

[0101] S343, matching the number of historical users in the first cluster, the average number of historical users in all clusters, the variance of the risk values ​​of historical users in the first cluster, and the average variance of the risk values ​​of historical users in all clusters in a preset screening ratio table; the preset screening ratio table includes a plurality of entries, each entry includes a quantity range, an average quantity range, a variance range, an average variance range, and a screening ratio.

[0102] In this embodiment, the average number of historical users of all clusters is the ratio of the number of historical users included in the historical user set to the number of clusters, and the average variance of risk values ​​of historical users of all clusters is the ratio of the sum of variances of risk values ​​of historical users of all clusters to the number of clusters.

[0103] In this embodiment, the preset screening ratio table is a pre-constructed table. As a preferred specific implementation, the table satisfies the following conditions: for two entries with the same quantity range, the same average quantity range and the same variance range, the screening ratio of the entry with the larger average variance corresponding to the average variance range is smaller than the screening ratio of the entry with the smaller average variance corresponding to the average variance range; for two entries with the same quantity range, the same average quantity range and the same average variance range, the screening ratio of the entry with the larger variance corresponding to the variance range is greater than the screening ratio of the entry with the smaller variance corresponding to the variance range; for two entries with the same quantity range, the same variance range and the same average variance range, the screening ratio of the entry with the larger average quantity corresponding to the average quantity range is smaller than the screening ratio of the entry with the smaller average quantity corresponding to the average quantity range; for two entries with the same average quantity range, the same variance range and the same average variance range, the screening ratio of the entry with the larger quantity corresponding to the quantity range is greater than the screening ratio of the entry with the smaller quantity corresponding to the quantity range. Based on this preferred specific implementation, while taking into account the amount of data for training users, the difference in the number of historical users in each cluster after subsequent screening can be made smaller, thereby avoiding overfitting of the trained neural network model, and the risk values ​​of the historical users in each retained cluster are more accurate, which is conducive to improving the accuracy of reasoning of the trained neural network model.

[0104] S344: If the match is successful, determine the screening ratio of the successfully matched entries as the screening ratio of the first cluster.

[0105] In this embodiment, if the matching is unsuccessful, historical users with abnormal risk values ​​in the cluster will not be screened out.

[0106] In this embodiment, if the quantity range included in a certain entry in the preset screening ratio table includes the number of historical users in the first cluster, the average quantity range included in the entry includes the average number of historical users of all clusters, the variance range included in the entry includes the variance of the risk values ​​of the historical users in the first cluster, and the average variance range included in the entry includes the average variance of the risk values ​​of the historical users of all clusters, then the match is determined to be successful, and the entry is determined to be a successfully matched entry.

[0107] S345, sort the historical users in the first cluster in ascending order of corresponding risk values, and screen out the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster from the first cluster; the first percentage is half of the screening ratio of the first cluster.

[0108] In this embodiment, the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster are determined as historical users with abnormal risk values ​​in the first cluster.

[0109] S350: Determine a historical user set consisting of the remaining historical users in all clusters as a training user set.

[0110] The training user set obtained based on S310-S350 can improve the accuracy of the trained neural network model in reasoning about users to be insured with different attributes.

[0111] S400, classifying the users to be insured in the user data set according to the risk value of each user to be insured in the user data set and a plurality of preset risk value thresholds.

[0112] As a specific implementation, several preset risk value thresholds are arranged in order from small to large, and the users to be insured whose corresponding risk values ​​are less than the minimum risk value threshold are regarded as users of the category with the smallest risk value, and the users to be insured whose corresponding risk values ​​are greater than or equal to the minimum risk value threshold and less than the second smallest risk value threshold are regarded as users of the category with the second smallest risk value, and so on, the users to be insured whose corresponding risk values ​​are greater than or equal to the second largest risk value threshold and less than the largest risk value threshold are regarded as users of the category with the second largest risk value, and the users to be insured whose corresponding risk values ​​are greater than or equal to the largest risk value threshold are regarded as users of the category with the largest risk value.

[0113] As a preferred specific implementation manner, when the users to be insured in the user data set to be insured are classified and stored, the risk value of each user to be insured is also stored.

[0114] This embodiment uses a trained neural network model to infer the standardized data of each user to be insured in the user data set to be insured, and obtains the risk value of each user to be insured in the user data set to be insured; since the training user set of the trained neural network model of this embodiment is obtained by screening the historical user set, the reasoning accuracy of the trained neural network model is relatively high; and the number of abnormal users screened out in each cluster of this embodiment is related to the number of historical users in the cluster and the variance of the risk value in the cluster, and is also related to the average number of historical users in all clusters and the average variance of the risk value in all clusters, which is conducive to reducing the difference in the number of historical users in each cluster after screening and retaining historical users with relatively accurate risk values ​​in each cluster, which is conducive to improving the accuracy of the risk values ​​of training users in the training user set, and then it is conducive to improving the accuracy of the trained neural network model and improving the accuracy of the final classification of users to be insured.

[0115] Embodiment 2:

[0116] According to this embodiment, Figure 6As shown, a user processing device based on a neural network model is provided, and the processing device includes:

[0117] The first acquisition module 100 is used to acquire a user data set to be insured, wherein the user data set to be insured includes initial data of a number of users to be insured, and the initial data of each user to be insured includes a number of attribute information of the corresponding user to be insured.

[0118] As a preferred specific implementation manner, the plurality of attribute information includes age, gender, body mass index, chronic disease history, smoking habits, drinking habits, exercise frequency, occupation type and income.

[0119] The standardization processing module 200 is used to standardize the initial data of each user to be insured in the user data set to be insured, and obtain the standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of attribute information of the corresponding user to be insured after standardized processing.

[0120] As a preferred specific implementation, the plurality of standardized attribute information includes standardized age, a first label for characterizing gender, a standardized body mass index, a second label for characterizing the presence or absence of chronic diseases, a standardized smoking frequency, a standardized drinking frequency, a standardized exercise frequency, a third label for characterizing whether an occupation is a high-risk occupation, and standardized income.

[0121] Optionally, the values ​​of the first label, the second label and the third label are 0 or 1; when the first label is 0, it indicates that the gender is male; when the first label is 1, it indicates that the gender is female; when the second label is 0, it indicates that there is no chronic disease; when the second label is 1, it indicates that there is a chronic disease; when the third label is 0, it indicates that the occupation is a low-risk occupation; when the third label is 1, it indicates that the occupation is a high-risk occupation.

[0122] Optionally, the value range of the standardized age, standardized body mass index, standardized smoking frequency, standardized drinking frequency, standardized exercise frequency and standardized income are all 0-1; the standardized age is the ratio of the user's age to a preset maximum age, the standardized body mass index is the ratio of the user's body mass index to a preset maximum body mass index, the standardized smoking frequency is the ratio of the number of days the user smokes per week to 7, the standardized drinking frequency is the ratio of the number of days the user drinks per week to 7, the standardized exercise frequency is the ratio of the number of days the user exercises per week to 7, and the standardized income is the ratio of the user's income to the preset maximum income.

[0123] The inference module 300 is used to input the standardized data of each user to be insured in the user data set to be insured into the trained neural network model for inference, so as to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user according to the input standardized data of the user; the trained neural network model is obtained according to the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, and the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0124] As a preferred specific implementation, the risk value of any historical user is the sum of the first ratio and the second ratio of the corresponding historical user, the first ratio being the ratio of the actual number of claims in the historical time period to the preset number of claims, and the second ratio being the ratio of the claim amount in the historical time period to the preset claim amount. Optionally, the historical time period is the most recent 1 year.

[0125] As a preferred specific implementation, the processing device further includes a training user set acquisition module, such as Figure 7 As shown, the training user set acquisition module includes:

[0126] The second acquisition module 310 is used to acquire a historical user set, where the historical user set includes a number of historical users, and the standardized data of each historical user includes a number of standardized attribute information of the corresponding historical user.

[0127] The grouping module 320 is used to group the historical users in the historical user set according to the corresponding standardized data, so that each standardized attribute information of each historical user included in each group is the same or belongs to the same preset range.

[0128] As a preferred specific implementation, the value range of each standardized attribute information is divided into several sub-ranges. If the sub-ranges to which all the standardized attribute information of two historical users belong are the same, then the two historical users belong to the same group. For example, if the value range of each standardized attribute information is divided into two sub-ranges, then the historical users in the historical user set can be divided into two sub-ranges. nBased on this preferred embodiment, the preliminary division of historical users can be achieved by simply judging whether each standardized attribute information of each historical user conforms to the corresponding sub-range, with relatively fast division speed and high division efficiency.

[0129] As a specific implementation, the historical users included in the same group have their standardized ages belonging to the same preset age range, the first label used to characterize gender is the same, the standardized body mass index belongs to the same preset body mass index range, the second label used to characterize the presence or absence of chronic diseases is the same, the standardized smoking frequency belongs to the same preset smoking frequency range, the standardized drinking frequency belongs to the same preset drinking frequency range, the standardized exercise frequency belongs to the same preset exercise frequency range, the third label used to characterize whether the occupation is a high-risk occupation is the same, and the standardized income belongs to the same preset income range.

[0130] The clustering module 330 is used to divide the historical users in each group into a number of clusters according to the distance of the standardized data of the historical users in the group.

[0131] In this embodiment, each historical user in the same group is taken as an object to be clustered, and the k-means method is used to realize the division of historical users in the same group, where k is an empirical value; in the process of using the k-means method to divide historical users in the same group, it is necessary to obtain the distance between any two historical users. As a preferred specific implementation method, Figure 8 As shown, the clustering module 330 includes:

[0132] The fifth acquisition module 331 is used to acquire standardized data of a first historical user; the first historical user is any user in any group.

[0133] The sixth acquisition module 332 is used to acquire standardized data of a second historical user; the second historical user and the first historical user belong to the same group, and the second historical user is different from the first historical user.

[0134] The distance acquisition module 333 is used to obtain the distance between the standardized data of the first historical user and the second historical user based on the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained based on the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

[0135] In this embodiment, the weighted sum formula is used to obtain the distance between the standardized data of the first historical user and the second historical user, and the difference value of any standardized attribute information of the first historical user and the second historical user is the absolute value of the difference between the standardized attribute information of the first historical user and the second historical user. As a preferred specific implementation, the clustering module also includes a weight acquisition module, such as Fig. 9 As shown, the weight acquisition module includes:

[0136] The seventh acquisition module 3331 is used to acquire the Pearson correlation coefficient between each standardized attribute information and the risk value according to the standardized data of the historical users in the historical user set and the corresponding risk value.

[0137] Those skilled in the art know that any method for obtaining the Pearson correlation coefficient in the prior art falls within the protection scope of the present invention.

[0138] The normalization processing module 3332 is used to normalize the absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value.

[0139] The third determination module 3333 is used to determine the absolute value of each normalized Pearson correlation coefficient as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

[0140] Based on this preferred specific implementation, the weight corresponding to the standardized attribute information that is more relevant to the risk value is larger. Therefore, the risk values ​​of historical users in the same cluster are closer, which is conducive to improving the accuracy of subsequent screening of historical users with abnormal risk values ​​based on the risk values ​​of historical users in the cluster.

[0141] The screening module 340 is used to screen out historical users with abnormal risk values ​​in a cluster according to the risk values ​​of the historical users in each cluster.

[0142] In this embodiment, the historical user with abnormal risk value in the cluster refers to the historical user with a large difference in risk value from other historical users in the cluster; as a preferred specific implementation method, Fig.10 As shown, the screening module 340 includes:

[0143] The third acquisition module 341 is used to acquire the number of historical users in a first cluster; the first cluster is any cluster obtained by dividing the historical users in any group.

[0144] The fourth acquisition module 342 is used to obtain the variance ratio coefficient of the risk value of the historical users in the first cluster.

[0145] Those skilled in the art know that any method for obtaining variance in the prior art falls within the protection scope of the present invention.

[0146] The matching module 343 is used to match the number of historical users in the first cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the first cluster, and the average variance of the risk values ​​of the historical users in all clusters in a preset screening ratio table; the preset screening ratio table includes a plurality of entries, each entry includes a quantity range, an average quantity range, a variance range, an average variance range and a screening ratio.

[0147] In this embodiment, the average number of historical users of all clusters is the ratio of the number of historical users included in the historical user set to the number of clusters, and the average variance of risk values ​​of historical users of all clusters is the ratio of the sum of variances of risk values ​​of historical users of all clusters to the number of clusters.

[0148] In this embodiment, the preset screening ratio table is a pre-constructed table. As a preferred specific implementation, the table satisfies the following conditions: for two entries with the same quantity range, the same average quantity range and the same variance range, the screening ratio of the entry with the larger average variance corresponding to the average variance range is smaller than the screening ratio of the entry with the smaller average variance corresponding to the average variance range; for two entries with the same quantity range, the same average quantity range and the same average variance range, the screening ratio of the entry with the larger variance corresponding to the variance range is greater than the screening ratio of the entry with the smaller variance corresponding to the variance range; for two entries with the same quantity range, the same variance range and the same average variance range, the screening ratio of the entry with the larger average quantity corresponding to the average quantity range is smaller than the screening ratio of the entry with the smaller average quantity corresponding to the average quantity range; for two entries with the same average quantity range, the same variance range and the same average variance range, the screening ratio of the entry with the larger quantity corresponding to the quantity range is greater than the screening ratio of the entry with the smaller quantity corresponding to the quantity range. Based on this preferred specific implementation, while taking into account the amount of data for training users, the difference in the number of historical users in each cluster after subsequent screening can be made smaller, thereby avoiding overfitting of the trained neural network model, and the risk values ​​of the historical users in each retained cluster are more accurate, which is conducive to improving the accuracy of reasoning of the trained neural network model.

[0149] The second determining module 344 is configured to determine the screening ratio of the successfully matched entries as the screening ratio of the first cluster if the match is successful.

[0150] In this embodiment, if the matching is unsuccessful, historical users with abnormal risk values ​​in the cluster will not be screened out.

[0151] In this embodiment, if the quantity range included in a certain entry in the preset screening ratio table includes the number of historical users in the first cluster, the average quantity range included in the entry includes the average number of historical users of all clusters, the variance range included in the entry includes the variance of the risk values ​​of the historical users in the first cluster, and the average variance range included in the entry includes the average variance of the risk values ​​of the historical users of all clusters, then the match is determined to be successful, and the entry is determined to be a successfully matched entry.

[0152] The sorting module 345 is used to sort the historical users in the first cluster in ascending order of corresponding risk values, and to screen out the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster from the first cluster; the first percentage is half of the screening ratio of the first cluster.

[0153] In this embodiment, the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster are determined as historical users with abnormal risk values ​​in the first cluster.

[0154] The first determination module 350 is used to determine a historical user set consisting of the remaining historical users in all clusters as a training user set.

[0155] The training user set obtained based on modules 310 - 350 can improve the accuracy of the trained neural network model in reasoning about users to be insured with different attributes.

[0156] The classification module 400 is used to classify the users to be insured in the user data set according to the risk value of each user to be insured in the user data set and a plurality of preset risk value thresholds.

[0157] As a specific implementation, several preset risk value thresholds are arranged in order from small to large, and the users to be insured whose corresponding risk values ​​are less than the minimum risk value threshold are regarded as users of the category with the smallest risk value, and the users to be insured whose corresponding risk values ​​are greater than or equal to the minimum risk value threshold and less than the second smallest risk value threshold are regarded as users of the category with the second smallest risk value, and so on, the users to be insured whose corresponding risk values ​​are greater than or equal to the second largest risk value threshold and less than the largest risk value threshold are regarded as users of the category with the second largest risk value, and the users to be insured whose corresponding risk values ​​are greater than or equal to the largest risk value threshold are regarded as users of the category with the largest risk value.

[0158] As a preferred specific implementation manner, when the users to be insured in the user data set to be insured are classified and stored, the risk value of each user to be insured is also stored.

[0159] This embodiment uses a trained neural network model to infer the standardized data of each user to be insured in the user data set to be insured, and obtains the risk value of each user to be insured in the user data set to be insured; since the training user set of the trained neural network model of this embodiment is obtained by screening the historical user set, the reasoning accuracy of the trained neural network model is relatively high; and the number of abnormal users screened out in each cluster of this embodiment is related to the number of historical users in the cluster and the variance of the risk value in the cluster, and is also related to the average number of historical users in all clusters and the average variance of the risk value in all clusters, which is conducive to reducing the difference in the number of historical users in each cluster after screening and retaining historical users with relatively accurate risk values ​​in each cluster, which is conducive to improving the accuracy of the risk values ​​of training users in the training user set, and then it is conducive to improving the accuracy of the trained neural network model and improving the accuracy of the final classification of users to be insured.

[0160] Embodiment three:

[0161] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0162] A user data set to be insured is obtained, wherein the user data set to be insured includes initial data of a plurality of users to be insured, and the initial data of each user to be insured includes a plurality of attribute information of the corresponding user to be insured.

[0163] The initial data of each user to be insured in the user data set to be insured is standardized to obtain the standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured.

[0164] The standardized data of each user to be insured in the user data set to be insured is input into the trained neural network model for inference to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0165] The users to be insured in the user data set to be insured are classified according to the risk value of each user to be insured in the user data set to be insured and a plurality of preset risk value thresholds.

[0166] Embodiment 4:

[0167] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0168] A user data set to be insured is obtained, wherein the user data set to be insured includes initial data of a plurality of users to be insured, and the initial data of each user to be insured includes a plurality of attribute information of the corresponding user to be insured.

[0169] The initial data of each user to be insured in the user data set to be insured is standardized to obtain the standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured.

[0170] The standardized data of each user to be insured in the user data set to be insured is input into the trained neural network model for inference to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters.

[0171] The users to be insured in the user data set to be insured are classified according to the risk value of each user to be insured in the user data set to be insured and a plurality of preset risk value thresholds.

[0172] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0173] Although some specific embodiments of the present invention have been described in detail by way of example, it will be appreciated by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It will also be appreciated by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A user processing method based on a neural network model, characterized in that: The processing method comprises the following steps: Acquire a user data set to be insured, wherein the user data set to be insured includes initial data of a plurality of users to be insured, and each initial data of the user to be insured includes a plurality of attribute information of the corresponding user to be insured; Standardizing the initial data of each user to be insured in the user data set to be insured to obtain standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured; The standardized data of each user to be insured in the user data set to be insured is input into the trained neural network model for inference to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, and the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, and the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters; Classifying the users to be insured in the user data set according to the risk value of each user to be insured in the user data set and a plurality of preset risk value thresholds; The process of obtaining the training user set includes: Acquire a historical user set, wherein the historical user set includes a plurality of historical users, and the standardized data of each historical user includes a plurality of standardized attribute information of the corresponding historical user; Grouping the historical users in the historical user set according to the corresponding standardized data, so that each standardized attribute information of each historical user included in each group is the same or belongs to the same preset range; Divide the historical users in each group into several clusters according to the distance of the standardized data of the historical users in the group; According to the risk values ​​of historical users in each cluster, historical users with abnormal risk values ​​in the cluster are screened out; The historical user set consisting of the remaining historical users in all clusters is determined as the training user set; Screening out historical users with abnormal risk values ​​in a cluster based on the risk values ​​of historical users in each cluster includes: Obtain the number of historical users in the first cluster; the first cluster is any cluster obtained by dividing the historical users in any group; Get the variance of risk values ​​of historical users in the first cluster; The number of historical users in the first cluster, the average number of historical users in all clusters, the variance of risk values ​​of historical users in the first cluster, and the average variance of risk values ​​of historical users in all clusters are matched in a preset screening ratio table; the preset screening ratio table includes a plurality of entries, each of which includes a quantity range, an average quantity range, a variance range, an average variance range, and a screening ratio; the screening ratio table satisfies the following conditions: for two entries with the same quantity range, the same average quantity range, and the same variance range, the screening ratio of the entry with the larger average variance corresponding to the average variance range is smaller than the screening ratio of the entry with the smaller average variance corresponding to the average variance range. Elimination ratio; among two items with the same quantity range, the same average quantity range and the same average variance range, the screening ratio of the items with larger variance corresponding to the variance range is greater than the screening ratio of the items with smaller variance corresponding to the variance range; among two items with the same quantity range, the same variance range and the same average variance range, the screening ratio of the items with larger average quantity corresponding to the average quantity range is less than the screening ratio of the items with smaller average quantity corresponding to the average quantity range; among two items with the same average quantity range, the same variance range and the same average variance range, the screening ratio of the items with larger quantity corresponding to the quantity range is greater than the screening ratio of the items with smaller quantity corresponding to the quantity range; If the match is successful, the screening ratio of the successfully matched entries is determined as the screening ratio of the first cluster; The historical users in the first cluster are sorted in ascending order according to the corresponding risk values, and the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster are screened out from the first cluster; the first percentage is half of the screening ratio of the first cluster.

2. The user processing method based on the neural network model according to claim 1 is characterized in that: The historical users in each group are divided into several clusters according to the distance of the standardized data of the historical users in the group, including: Obtaining standardized data of a first historical user; the first historical user is any user in any group; Acquire standardized data of a second historical user; the second historical user and the first historical user belong to the same group, and the second historical user is different from the first historical user; The distance between the standardized data of the first historical user and the second historical user is obtained according to the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained according to the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

3. The user processing method based on the neural network model according to claim 2 is characterized in that: The corresponding weight acquisition process includes: According to the standardized data of historical users in the historical user set and the corresponding risk values, the Pearson correlation coefficient between each standardized attribute information and the risk value is obtained; The absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value is normalized; The absolute value of each normalized Pearson correlation coefficient is determined as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

4. A user processing device based on a neural network model, characterized in that: The processing device comprises: A first acquisition module is used to acquire a user data set to be insured, wherein the user data set to be insured includes initial data of a plurality of users to be insured, and each initial data of the user to be insured includes a plurality of attribute information of the corresponding user to be insured; A standardization processing module is used to perform standardization processing on the initial data of each user to be insured in the user data set to be insured, so as to obtain standardized data of each user to be insured in the user data set to be insured; the standardized data of each user to be insured includes a number of standardized attribute information of the corresponding user to be insured; An inference module is used to input the standardized data of each user to be insured in the user data set to be insured into a trained neural network model for inference, so as to obtain the risk value of each user to be insured in the user data set to be insured; the trained neural network model is used to infer the risk value of the corresponding user based on the input standardized data of the user; the trained neural network model is obtained based on the standardized data of each training user in the training user set and the corresponding risk value, and the risk value of any training user is positively correlated with the actual number of claims of the training user, and the risk value of any training user is positively correlated with the claim amount of the training user; the training user set is obtained by screening the historical user set, and the screening process includes screening out historical users with abnormal risk values ​​belonging to the same cluster, and the number of historical users with abnormal risk values ​​in different clusters is related to the number of historical users in the corresponding cluster, the average number of historical users in all clusters, the variance of the risk values ​​of the historical users in the corresponding cluster, and the average variance of the risk values ​​of the historical users in all clusters; A classification module, used for classifying the users to be insured in the user data set to be insured according to the risk value of each user to be insured in the user data set to be insured and a plurality of preset risk value thresholds; The processing device further includes a training user set acquisition module, and the training user set acquisition module includes: A second acquisition module is used to acquire a historical user set, wherein the historical user set includes a number of historical users, and the standardized data of each historical user includes a number of standardized attribute information of the corresponding historical user; A grouping module, used for grouping the historical users in the historical user set according to the corresponding standardized data, so that each of the standardized attribute information of the historical users included in each group is the same or belongs to the same preset range; A clustering module, used to divide the historical users in each group into several clusters according to the distance of the standardized data of the historical users in the group; A screening module is used to screen out historical users with abnormal risk values ​​in a cluster according to the risk values ​​of historical users in each cluster; A first determination module is used to determine a historical user set consisting of all remaining historical users in the cluster as a training user set; The screening modules include: The third acquisition module is used to acquire the number of historical users in the first cluster; the first cluster is any cluster obtained by dividing the historical users in any group; A fourth acquisition module, used to obtain the variance of the risk values ​​of historical users in the first cluster; The matching module is used to match the number of historical users in the first cluster, the average number of historical users in all clusters, the variance of risk values ​​of historical users in the first cluster, and the average variance of risk values ​​of historical users in all clusters in a preset screening ratio table; the preset screening ratio table includes a plurality of entries, each of which includes a quantity range, an average quantity range, a variance range, an average variance range, and a screening ratio; the screening ratio table satisfies the following conditions: for two entries with the same quantity range, the same average quantity range, and the same variance range, the screening ratio of the entry with a larger average variance corresponding to the average variance range is smaller than the screening ratio of the entry with a smaller average variance corresponding to the average variance range; Purpose screening ratio; among two items with the same quantity range, the same average quantity range and the same average variance range, the screening ratio of the item with the larger variance corresponding to the variance range is greater than the screening ratio of the item with the smaller variance corresponding to the variance range; among two items with the same quantity range, the same variance range and the same average variance range, the screening ratio of the item with the larger average quantity corresponding to the average quantity range is less than the screening ratio of the item with the smaller average quantity corresponding to the average quantity range; among two items with the same average quantity range, the same variance range and the same average variance range, the screening ratio of the item with the larger quantity corresponding to the quantity range is greater than the screening ratio of the item with the smaller quantity corresponding to the quantity range; A second determination module, configured to determine the screening ratio of the successfully matched entries as the screening ratio of the first cluster if the match is successful; A sorting module is used to sort the historical users in the first cluster in ascending order of corresponding risk values, and to screen out the first percentage of historical users with the smallest risk values ​​and the first percentage of historical users with the largest risk values ​​in the first cluster from the first cluster; the first percentage is half of the screening ratio of the first cluster.

5. The user processing device based on the neural network model according to claim 4, characterized in that: The clustering modules include: A fifth acquisition module is used to acquire standardized data of a first historical user; the first historical user is any user in any group; a sixth acquisition module, configured to acquire standardized data of a second historical user; the second historical user and the first historical user belong to the same group, and the second historical user is different from the first historical user; A distance acquisition module is used to obtain the distance between the standardized data of the first historical user and the second historical user based on the standardized data of the first historical user and the standardized data of the second historical user; the distance between the standardized data of the first historical user and the second historical user is obtained based on the difference value of each standardized attribute information of the first historical user and the second historical user and the corresponding weight.

6. The user processing device based on the neural network model according to claim 5, characterized in that: The clustering module further includes a weight acquisition module, and the weight acquisition module includes: A seventh acquisition module, used for acquiring the Pearson correlation coefficient between each standardized attribute information and the risk value according to the standardized data of the historical users in the historical user set and the corresponding risk value; A normalization processing module, used for normalizing the absolute value of the Pearson correlation coefficient between each standardized attribute information and the risk value; The third determination module is used to determine the absolute value of each normalized Pearson correlation coefficient as the weight of the corresponding attribute information after normalization; the absolute value of any normalized Pearson correlation coefficient is the ratio of the absolute value of the Pearson correlation coefficient to the sum of the absolute values ​​of all Pearson correlation coefficients.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the user processing method based on the neural network model as described in any one of claims 1 to 3 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the user processing method based on a neural network model as described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Risk assessment model training method and device and server

    CN115293336A

  • Building sampling method based on hierarchical clustering and feature selection

    CN118468068A