Risk assessment method and device, computer equipment and storage medium
By comprehensively analyzing users' historical credit, voice, and text data, a risk assessment model is constructed, which solves the problem of difficulty in capturing users' real-time emotions and behavioral anomalies in traditional financial risk assessment, and achieves more accurate risk judgment and cost optimization.
Patent Information
- Application Number
- CN202510749302.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional financial risk assessment models are based on users’ historical credit data and static data, which makes it difficult to capture users’ real-time emotional changes and behavioral abnormalities, resulting in low accuracy of risk judgment and easy misjudgment or omission.
By acquiring and comprehensively analyzing the target user's historical credit data, voice data, and text data, voice feature parameters and text feature parameters are extracted, voice feature vectors and text feature vectors are constructed, and a comprehensive risk result is calculated by combining the Shapley additive interpretation value. When the risk exceeds the threshold, an early warning prompt is output.
It improves the accuracy and reliability of risk assessment, enabling more accurate determination of users' risk levels, reducing risk management costs, and minimizing the bias and limitations of a single data dimension.
Smart Images

Figure CN120852029A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial data processing technology, and in particular to a risk assessment method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the rapid development of computer technology, users may make false statements (lying) during financial transactions due to their own reasons. Financial institutions need to assess the user's risk profile to ensure their risk management capabilities. Traditional risk control models generally construct risk scores based on users' historical credit data and static data. Static data is difficult to capture users' real-time emotional changes and behavioral anomalies during the application process. Furthermore, considering only historical credit data and static data without comprehensively considering other dimensions of data and dynamic behavioral data leads to a one-sided perspective and is prone to misjudgment or omission.
[0003] Therefore, the accuracy of user risk assessment is relatively low, and it is difficult to accurately identify users lying. Summary of the Invention
[0004] Therefore, it is necessary to provide a risk assessment method, apparatus, computer equipment, and storage medium that can accurately assess risks in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a risk assessment method, including:
[0006] Obtain preprocessed data of the target user to be evaluated; the data to be evaluated includes historical credit data, voice data, and text data;
[0007] Based on the historical credit data, a first risk assessment result is determined; the first risk assessment result characterizes the credit anomaly risk of the target user.
[0008] Speech feature parameters are extracted based on the speech data, and speech feature vectors are determined based on the speech feature parameters.
[0009] Text feature parameters are extracted based on the text data, and multiple types of text feature vectors are determined based on the text feature parameters;
[0010] Based on the speech feature vector and the multiple types of text feature vectors, a second risk assessment result is determined; the second risk assessment result characterizes the real-time abnormal risk of the target user.
[0011] Based on the first risk assessment result and the second risk assessment result, a comprehensive risk result is determined.
[0012] In one embodiment, the step of extracting speech feature parameters based on the speech data and determining a speech feature vector based on the speech feature parameters includes:
[0013] Based on the speech data, the fundamental frequency mean, pause time, and standard deviation are extracted;
[0014] The speech feature vector is determined based on the fundamental frequency mean, the pause time, and the standard deviation; the speech feature vector represents the speech stress of the target user.
[0015] In one embodiment, the step of extracting text feature parameters based on the text data and determining multiple types of text feature vectors based on the text feature parameters includes at least two of the following:
[0016] Based on the text data, determine the total number of texts and the number of texts containing the target words, and based on the total number of texts and the number of texts containing the target words, determine the text feature vector of the word type;
[0017] The total number of statements and the parse tree depth of each statement are determined based on the text data, and the text feature vector of the syntax type is determined based on the total number of statements and the parse tree depth.
[0018] Based on the text data, the total number of words and the vector of each word are determined, and based on the total number of words and the vector of each word, the text feature vector of the semantic type is determined;
[0019] Based on the text data, the sentiment score of each word and the total number of words are determined, and the text feature vector of the sentiment type is determined based on the total number of words and the sentiment score.
[0020] In one embodiment, determining the second risk assessment result based on the speech feature vector and multiple types of text feature vectors includes:
[0021] At least one of the text feature vectors of the lexical type, the syntactic type, the semantic type, and the sentiment type is fused to determine a comprehensive text feature vector.
[0022] The second risk assessment result is obtained based on the first weight vector corresponding to the speech feature vector and the second weight vector corresponding to the comprehensive text feature vector.
[0023] In one embodiment, the method further comprises:
[0024] Calculate the Shapley additive interpretation value for each feature vector contained in the historical credit data, the voice data, and the text data;
[0025] The contribution of each eigenvector to the overall risk outcome is determined based on the Shapley additive interpretation value; the contribution indicates whether each eigenvector has a positive or negative impact on the overall risk outcome.
[0026] In one embodiment, the method further comprises:
[0027] Output at least one of the following: the comprehensive risk result, the contribution distribution corresponding to the Shapley additive explanation value, the change curve corresponding to the speech feature vector, and the change curve corresponding to the text feature vector;
[0028] If the overall risk result is greater than or equal to a preset risk threshold, an early warning prompt will be output; the early warning prompt is used to remind the target user to undergo risk review.
[0029] In one embodiment, the method further comprises:
[0030] Obtain the original data to be evaluated from the target user;
[0031] The original data to be evaluated is anonymized to obtain candidate data;
[0032] The candidate data is categorized and labeled according to fraudulent behavior and overdue behavior to obtain candidate data with labeled types.
[0033] The labeled candidate data is cleaned and balanced to obtain the preprocessed data to be evaluated.
[0034] Secondly, this application also provides a risk assessment device, the device comprising:
[0035] The acquisition module is used to acquire preprocessed data to be evaluated from the target user; the data to be evaluated includes historical credit data, voice data, and text data.
[0036] The first risk assessment result module is used to determine a first risk assessment result based on the historical credit data; the first risk assessment result characterizes the credit anomaly risk of the target user.
[0037] The first feature extraction module is used to extract speech feature parameters based on the speech data and determine a speech feature vector based on the speech feature parameters.
[0038] The second feature extraction module is used to extract text feature parameters based on the text data, and to determine multiple types of text feature vectors based on the text feature parameters;
[0039] The second risk assessment result module is used to determine a second risk assessment result based on the voice feature vector and multiple types of text feature vectors; the second risk assessment result characterizes the real-time abnormal risk of the target user.
[0040] The determination module is used to determine the comprehensive risk result based on the first risk assessment result and the second risk assessment result.
[0041] Thirdly, this application also provides a computer device, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in any embodiment of this application.
[0042] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any embodiment of this application.
[0043] The aforementioned risk assessment method comprehensively considers historical credit data, voice data, and text data to evaluate the risk of target users from different perspectives. Historical credit data reflects a user's past credit status, while voice and text data capture the user's current behavioral characteristics and real-time status, reducing the bias and limitations that may arise from a single data dimension. This makes the risk assessment more comprehensive and multi-dimensional, thereby improving its accuracy and reliability. Furthermore, based on the comprehensive risk results, the risk level of target users can be determined more accurately, improving resource utilization efficiency and reducing risk management costs. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is an application environment diagram illustrating a risk assessment method according to an exemplary embodiment;
[0046] Figure 2 This is a flowchart illustrating a risk assessment method according to an exemplary embodiment;
[0047] Figure 3 This is a flowchart illustrating a risk assessment method according to an exemplary embodiment;
[0048] Figure 4 This is a flowchart illustrating a risk assessment method according to an exemplary embodiment;
[0049] Figure 5 This is a structural block diagram of a risk assessment apparatus according to an exemplary embodiment;
[0050] Figure 6 This is an internal structural diagram of a computer device according to an exemplary embodiment. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] The terms "first," "second," and "third" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "at least one" is used to indicate one or more; "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0053] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0054] The risk assessment method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0055] In one exemplary embodiment, such as Figure 2 As shown, a risk assessment method is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps S201 to S206. Wherein:
[0056] Step S201: Obtain the preprocessed target user's data to be evaluated; the data to be evaluated includes historical credit data, voice data, and text data.
[0057] In this embodiment, preprocessing refers to the process of collecting, cleaning, and transforming the raw data before analyzing or modeling it. Preprocessing may include, but is not limited to, at least one of data integration, data cleaning, data standardization, and data reduction.
[0058] In this embodiment of the application, historical credit data may include, but is not limited to, user identity information, user income information, user credit transaction information, and public record information. Public record information may include administrative penalty records and legal penalty records, etc.
[0059] Step S202: Based on the historical credit data, determine the first risk assessment result; the first risk assessment result characterizes the credit anomaly risk of the target user.
[0060] In some embodiments, determining the first risk assessment result based on the historical credit data includes:
[0061] Based on the target user's credit score, income level, and historical repayment history, the first risk assessment result is determined.
[0062] For example, a method for determining the result of a first risk assessment The method is as follows: ;in, Indicates the target user's credit score; Indicates the income level of the target users; Indicates the target user's historical repayment history; , , The reference coefficients, which respectively indicate credit score, income level, and historical repayment history, can be determined using statistical regression methods.
[0063] Step S203: Extract speech feature parameters based on the speech data, and determine the speech feature vector based on the speech feature parameters.
[0064] In this application embodiment, the speech feature parameters may include, but are not limited to, at least one of formant frequency, bandwidth, spectral flux, mean fundamental frequency, fundamental frequency fluctuation, pause time, and speech energy.
[0065] In one embodiment, the terminal extracts speech feature parameters such as emotion, stress, and speech rate changes from the speech data, determines the speech feature vector, and can perform normalization processing.
[0066] Step S204: Extract text feature parameters based on the text data, and determine multiple types of text feature vectors based on the text feature parameters.
[0067] In this embodiment of the application, the text feature parameters may include, but are not limited to, at least one of the following: word frequency, total number of words, sentence length, sentiment tendency, sentence complexity, and total number of sentences.
[0068] In this embodiment of the application, the type of text feature vector may include, but is not limited to, lexical type, syntactic type, semantic type and sentiment type.
[0069] In one embodiment, the terminal extracts multi-level text feature parameters from text data and determines multiple types of text feature vectors based on bag-of-words model, word embedding method, or deep learning method.
[0070] Step S205: Based on the speech feature vector and the multiple types of text feature vectors, determine the second risk assessment result; the second risk assessment result characterizes the real-time abnormal risk of the target user.
[0071] In some embodiments, the terminal can determine the weight values corresponding to each feature vector (voice feature vector, text feature vector) according to the integrated algorithm, and determine the second risk assessment result based on each feature vector and its corresponding weight value.
[0072] For example, a method for determining the result of a second risk assessment The method is as follows: ;in, Indicates the i-th real-time feature vector (real-time feature vectors include speech feature vectors and multiple types of text feature vectors); Indicates the weight value corresponding to each real-time feature vector.
[0073] Step S206: Based on the first risk assessment result and the second risk assessment result, determine the comprehensive risk result.
[0074] In some embodiments, the terminal determines the comprehensive risk result based on the product of the first risk assessment result and the corresponding first weight, and the product of the second risk assessment result and the corresponding second weight.
[0075] For example, a method for determining comprehensive risk outcomes The method is as follows: ;in, Indicates the results of the first risk assessment; Instructions for a second risk assessment; , Indicates the weight adjustment parameter. ; Indicates the i-th real-time feature vector (real-time feature vectors include speech feature vectors and multiple types of text feature vectors); Indicates the weight value corresponding to each real-time feature vector.
[0076] The aforementioned risk assessment method comprehensively considers historical credit data, voice data, and text data to evaluate the risk of target users from different perspectives. Historical credit data reflects a user's past credit status, while voice and text data capture the user's current behavioral characteristics and real-time status, reducing the bias and limitations that may arise from a single data dimension. This makes the risk assessment more comprehensive and multi-dimensional, thereby improving its accuracy and reliability. Furthermore, based on the comprehensive risk results, the risk level of target users can be determined more accurately, improving resource utilization efficiency and reducing risk management costs.
[0077] In one exemplary embodiment, such as Figure 3 As shown, step S203 includes steps S302 to S304. Wherein:
[0078] Step S302: Extract the fundamental frequency mean, pause time and standard deviation based on the speech data.
[0079] In this embodiment, the fundamental frequency mean can indicate the average value of the fundamental frequency in the speech data / signal. The fundamental frequency is the frequency of vocal cord vibration, which determines the pitch of the speech. During phonation, the frequency of vocal cord vibration changes over time, and the fundamental frequency mean can be obtained by averaging the fundamental frequency values that change over time.
[0080] In this embodiment, the pause time can indicate the period during which the target user intentionally or unintentionally pauses pronunciation. The terminal can process the speech data using speech analysis software to detect the presence or absence of speech signals and determine the start and end times of the pause.
[0081] For example, the terminal can extract speech feature parameters, such as the fundamental frequency mean, using audio feature extraction tools (such as OpenSMILE). Standard deviation pause time .
[0082] Step S304: Determine the speech feature vector based on the fundamental frequency mean parameter, the pause time parameter, and the standard deviation parameter; the speech feature vector represents the speech stress of the target user.
[0083] For example, a method for determining speech feature vectors The (Voice Stress Index) is calculated as follows: ;in, Indicative standard deviation; Indicates the fundamental frequency mean; Indicate the pause time; Indicates the pause time function to enhance the weighting of abnormal pauses.
[0084] In this embodiment, the mean fundamental frequency reflects the basic pitch level of speech. Changes in the tension of the target user's vocal cords under stress will cause changes in the fundamental frequency, and this subtle change can be captured by the mean fundamental frequency. Pause duration reflects the rhythm and fluency of speech; stress may cause the target user to pause more or have abnormally long pauses, thus providing clues for stress assessment. Standard deviation reflects the fluctuation of features such as the fundamental frequency. Stress usually reduces the stability of speech features, increasing the standard deviation, which helps to characterize dynamic changes. By using features such as the mean fundamental frequency, pause duration, and standard deviation, quantitative indicators are provided for assessing speech stress. Compared to subjective judgment, quantitative feature parameters can more accurately and objectively reflect the stress information contained in the target user's speech, reducing interference from human factors and making speech stress assessment more scientific and reliable.
[0085] In one embodiment, the step of extracting text feature parameters based on the text data and determining multiple types of text feature vectors based on the text feature parameters includes at least two of the following:
[0086] Based on the text data, determine the total number of texts and the number of texts containing the target words, and based on the total number of texts and the number of texts containing the target words, determine the text feature vector of the word type;
[0087] The total number of statements and the parse tree depth of each statement are determined based on the text data, and the text feature vector of the syntax type is determined based on the total number of statements and the parse tree depth.
[0088] Based on the text data, the total number of words and the vector of each word are determined, and based on the total number of words and the vector of each word, the text feature vector of the semantic type is determined;
[0089] Based on the text data, the sentiment score of each word and the total number of words are determined, and the text feature vector of the sentiment type is determined based on the total number of words and the sentiment score.
[0090] For example, a terminal can calculate the frequency of words appearing in text data using Term Frequency-Inverse Document Frequency (TF-IDF). This is a text feature vector for determining word type. The method is as follows: ;in, Indicates the frequency of occurrence of target words in text data; Total number of texts indicated; Indicator words; Indicates the number of texts containing the target word.
[0091] In some embodiments, the terminal can extract syntactic structure information based on text data, such as sentence length (number of characters / words), depth of the syntactic tree / number of nodes, etc.; the terminal can construct a syntactic parse tree based on each sentence (i), thus generating a text feature vector that determines the syntactic type. The method for calculating the average syntax tree height is as follows: ;in, Indicates the parse tree depth of the i-th statement; Indicates the total number of statements in the text.
[0092] In some embodiments, the terminal encodes text data using the BERT (Bidirectional Encoder Representations from Transformers) model to obtain contextual semantic parameters; and embeds these parameters as sentences using specific tags (such as [cls]). Average pooling is performed on each word vector to obtain the text feature vector of the speech type. : ;in, Total number of indicator words; Indicates the word vector corresponding to the i-th word output by the BERT model.
[0093] For example, a text feature vector for determining sentiment type The method is as follows: ;in, Total number of indicator words; Indicator words The corresponding sentiment score.
[0094] In this embodiment, different types of text feature vectors, such as lexical type, syntactic type, semantic type, and sentiment type, reflect the characteristics of the text from different perspectives. For example, lexical type features can reflect the frequency and distribution of specific words in the text, helping to identify the text's theme and focus; syntactic type features can demonstrate the structural complexity and language organization of the text; semantic type features can reveal the meaning and internal logic of the text; and sentiment type features reflect the emotional tendency contained in the text. By combining these multi-dimensional features, potential risks can be identified more promptly and accurately, reducing misjudgments and omissions caused by incomplete or inaccurate information, and improving the reliability of risk assessment.
[0095] In one embodiment, determining the second risk assessment result based on the speech feature vector and multiple types of text feature vectors includes:
[0096] At least one of the text feature vectors of the lexical type, the syntactic type, the semantic type, and the sentiment type is fused to determine a comprehensive text feature vector.
[0097] The second risk assessment result is obtained based on the first weight vector corresponding to the speech feature vector and the second weight vector corresponding to the comprehensive text feature vector.
[0098] For example, the terminal extracts features from text data using feature functions to obtain multiple types of text feature vectors; weights are assigned to each type of text feature vector for feature fusion to obtain a comprehensive text feature vector. A method for determining the comprehensive text feature vector... The method is as follows: ;in, Indicates the weights corresponding to each text feature vector; Indicates text feature functions.
[0099] For example, the second risk assessment result can be calculated using a sigmoid function. The terminal maps the linear combination of the second risk assessment results to probability values between (0,1). The closer the probability value is to 1, the higher the likelihood of short-term risk in the target user's current business behavior; conversely, the closer the probability value is to 0, the lower the likelihood of short-term risk in the target user's current business behavior. A method for determining the second risk assessment result... The method is as follows: ;in, Indicates the normalized speech feature vector; Indicates the comprehensive text feature vector; Indicator bias item; Indicates the first weight vector; Indicates the second weight vector.
[0100] In this embodiment, different types of text feature vectors describe the text from aspects such as vocabulary, syntax, semantics, and sentiment. Fusing these feature vectors integrates information from various aspects, avoiding the limitations of a single feature vector. For example, vocabulary features reflect the text's theme and key information, syntactic features demonstrate the organization and structure of language, semantic features reveal the text's deeper meaning, and sentiment features reflect the text's emotional tendency. Through feature fusion, the assessment model can capture more comprehensive and detailed textual information, thereby more accurately assessing risk. By combining speech feature vectors and comprehensive text feature vectors, the degree of negative emotion and the level of risk posed by the target user can be further accurately determined.
[0101] In one embodiment, the method further includes:
[0102] Calculate the Shapley additive interpretation value for each feature vector contained in the historical credit data, the voice data, and the text data;
[0103] The contribution of each eigenvector to the overall risk outcome is determined based on the Shapley additive interpretation value; the contribution indicates whether each eigenvector has a positive or negative impact on the overall risk outcome.
[0104] In this embodiment of the application, Shapley Additive exPlanations (SHAP) is an explanatory analysis tool based on Shapley values in cooperative game theory, which can be used to explain the specific contribution of each feature vector in the prediction results of machine learning models.
[0105] In some embodiments, the SHAP value assigns a contribution value to each feature vector, characterizing the positive or negative impact of that feature vector on the current judgment of "lying" or "not lying." For example, when the model outputs a comprehensive risk assessment result that is biased towards "lying," the SHAP value can help identify which feature vectors contribute the most to that judgment: some feature vectors have positive SHAP values, indicating that these feature vectors tend to support the conclusion of "lying"; some feature vectors have negative SHAP values, meaning that these feature vectors tend to support the conclusion of "not lying."
[0106] In traditional risk assessment scenarios, whether using complex machine learning or deep learning models, the decision-making process is often difficult to understand intuitively. In contrast, in this embodiment, by calculating the SHAP value, the role of each feature vector in the model's decision-making can be clearly seen, clarifying its contribution direction (positive or negative) and degree to the overall risk outcome. For example, in credit risk assessment, the impact of feature vectors such as the target user's income level and historical delinquency frequency on the final risk score can be clearly understood, making the model's decision-making process transparent and interpretable.
[0107] In one embodiment, the method further includes:
[0108] Output at least one of the following: the comprehensive risk result, the contribution distribution corresponding to the Shapley additive explanation value, the change curve corresponding to the speech feature vector, and the change curve corresponding to the text feature vector;
[0109] If the overall risk result is greater than or equal to a preset risk threshold, an early warning prompt will be output; the early warning prompt is used to remind the target user to undergo risk review.
[0110] In some embodiments, the terminal displays the comprehensive risk results, the contribution distribution corresponding to the Shapley additive explanation value, the change curves corresponding to the speech feature vector, and the change curves corresponding to the text feature vector on a visual dashboard in real time, facilitating operators to quickly locate the source of risk; and, it acquires pre-set risk thresholds. ,exist In such cases, an early warning message will be output; among which, Indicates the overall risk outcome.
[0111] In some embodiments, the step of outputting an early warning when the overall risk result is greater than or equal to a preset risk threshold includes at least one of the following:
[0112] The first method involves using a terminal-based voice device to output the warning message.
[0113] The second method: output the warning prompt on the display screen of the terminal;
[0114] The third method: The warning prompt is output based on the vibration of the motor in the terminal;
[0115] The fourth method involves sending the warning notification to a third-party platform; the warning notification is then displayed by the third-party platform.
[0116] In some embodiments, the warning message can be displayed on the terminal's screen in the form of text, list, etc.; or it can be sent to a third-party platform via email / SMS and displayed on the third-party platform.
[0117] In this embodiment, by displaying the comprehensive risk results in real time, the overall risk status of the target user can be intuitively understood. By showing the contribution distribution corresponding to the Shapley additive explanatory value, the influence of each feature vector on the comprehensive risk result can be clearly presented. Operators can understand which factors are the main reasons for the increase or decrease in risk, which helps to analyze the source of risk in depth and provide a basis for risk prevention and control. Furthermore, the change curves corresponding to the voice feature vector and text feature vector can intuitively show the changing trend of these features over time or other factors, which is conducive to timely capture of potential risk signals. When the comprehensive risk result is greater than or equal to the preset risk threshold, an early warning prompt is output, which can promptly remind relevant personnel to conduct risk review of the target user, thereby effectively preventing the further expansion of risk and reducing potential losses.
[0118] In one embodiment, the method further includes:
[0119] Obtain the original data to be evaluated from the target user;
[0120] The original data to be evaluated is anonymized to obtain candidate data;
[0121] The candidate data is categorized and labeled according to fraudulent behavior and overdue behavior to obtain candidate data with labeled types.
[0122] The labeled candidate data is cleaned and balanced to obtain the preprocessed data to be evaluated.
[0123] In one embodiment, the terminal can use the SHA-256 hash algorithm to desensitize the collected raw data to be evaluated (raw voice / text data, etc.), remove or mask sensitive information, and obtain alternative data; thereby reducing the interference of sensitive information on subsequent feature extraction and model evaluation processing and ensuring the privacy and security of the target user.
[0124] In one embodiment, the terminal categorizes and labels candidate data based on whether the fraudulent behavior is "lying" (L, value 0 or 1) and whether the overdue behavior is "overdue" (D, value 0 or 1), resulting in labeled candidate data. The set of labeled candidate data is as follows: ;in, Indicates the amount of data.
[0125] In one embodiment, data with anomalies (such as blurred audio or missing text) in the labeled candidate data is cleaned by setting a noise threshold. When the quality of the data sample is below the noise threshold If this happens, the data sample is filtered to obtain cleaned candidate data; an exemplary filtering method is as follows: ;in, Indicator function; Indicates the i-th data sample.
[0126] In one embodiment, the terminal performs data balancing on the cleaned candidate data: oversampling is performed using the SMOTE (Synthetic Minority Over-sampling Technique) algorithm to ensure that each type of data sample satisfies the following formula: ;in, Indicates the number of data samples of the "lying" type; Indicates the number of data samples that are "not lying".
[0127] In this embodiment, data anonymization can hide or replace sensitive information in the data, protecting the privacy of target users and reducing the risk of data leakage. Classifying and labeling the data clearly defines different types of data, making the data characteristics more explicit; this helps in subsequent analysis and modeling, better understanding the relationship between the data and fraudulent / overdue behaviors, providing a foundation for accurate user risk assessment. Data cleaning and balancing reduce errors during model training, improving model stability and reliability, making evaluation results more accurate and credible, and enhancing the model's generalization ability and prediction accuracy.
[0128] This application also provides an application scenario in which the above-described risk assessment method is applied. Specifically, the risk assessment method is applied in this scenario as follows:
[0129] The terminal acquires the target user's original data to be evaluated in the data collection and preprocessing module, and performs data labeling and classification on the original data to obtain labeled candidate data. The labeled candidate data is then cleaned and balanced to obtain preprocessed data to be evaluated. In the multimodal feature extraction module, speech and text features are extracted from the data to be evaluated, resulting in speech feature vectors and text feature vectors. In the dynamic behavior analysis model construction module, a first risk assessment result is calculated based on historical credit data in the data to be evaluated; a second risk assessment result is calculated based on the extracted speech and text feature vectors; and the contribution of each feature vector to the overall risk result is calculated using the SHAP method. In the real-time early warning and decision support module, an early warning trigger mechanism is set up. When the overall risk result is greater than or equal to a preset risk threshold, an early warning prompt is output; and a decision support interface is displayed. The decision support interface outputs at least one of the following: the overall risk result, the contribution distribution corresponding to the Shapley additive explanation value, the change curve corresponding to the speech feature vector, and the change curve corresponding to the text feature vector.
[0130] The aforementioned risk assessment method comprehensively considers historical credit data, voice data, and text data to evaluate the risk of target users from different perspectives. Historical credit data reflects a user's past credit status, while voice and text data capture the user's current behavioral characteristics and real-time status, reducing the bias and limitations that may arise from a single data dimension. This makes the risk assessment more comprehensive and multi-dimensional, thereby improving its accuracy and reliability. Furthermore, based on the comprehensive risk results, the risk level of target users can be determined more accurately, improving resource utilization efficiency and reducing risk management costs.
[0131] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0132] Based on the same inventive concept, this application also provides a risk assessment apparatus for implementing the risk assessment method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more risk assessment apparatus embodiments provided below can be found in the limitations of the risk assessment method described above, and will not be repeated here.
[0133] In one exemplary embodiment, such as Figure 5 As shown, a risk assessment device is provided, comprising:
[0134] The acquisition module 10 is used to acquire preprocessed data to be evaluated of the target user; the data to be evaluated includes historical credit data, voice data and text data.
[0135] The first risk assessment result module 20 is used to determine a first risk assessment result based on the historical credit data; the first risk assessment result characterizes the credit anomaly risk of the target user.
[0136] The first feature extraction module 30 is used to extract speech feature parameters based on the speech data and determine a speech feature vector based on the speech feature parameters.
[0137] The second feature extraction module 40 is used to extract text feature parameters based on the text data, and to determine multiple types of text feature vectors based on the text feature parameters;
[0138] The second risk assessment result module 50 is used to determine a second risk assessment result based on the voice feature vector and multiple types of text feature vectors; the second risk assessment result characterizes the real-time abnormal risk of the target user;
[0139] The determination module 60 is used to determine the comprehensive risk result based on the first risk assessment result and the second risk assessment result.
[0140] In one embodiment, the first feature extraction module 30 is configured to perform the following steps:
[0141] Based on the speech data, the fundamental frequency mean, pause time, and standard deviation are extracted;
[0142] The speech feature vector is determined based on the fundamental frequency mean, the pause time, and the standard deviation; the speech feature vector represents the speech stress of the target user.
[0143] In one embodiment, the second feature extraction module 40 is configured to perform at least two of the following steps:
[0144] Based on the text data, determine the total number of texts and the number of texts containing the target words, and based on the total number of texts and the number of texts containing the target words, determine the text feature vector of the word type;
[0145] The total number of statements and the parse tree depth of each statement are determined based on the text data, and the text feature vector of the syntax type is determined based on the total number of statements and the parse tree depth.
[0146] Based on the text data, the total number of words and the vector of each word are determined, and based on the total number of words and the vector of each word, the text feature vector of the semantic type is determined;
[0147] Based on the text data, the sentiment score of each word and the total number of words are determined, and the text feature vector of the sentiment type is determined based on the total number of words and the sentiment score.
[0148] In one embodiment, the second risk assessment result module 50 is used to perform the following steps:
[0149] At least one of the text feature vectors of the lexical type, the syntactic type, the semantic type, and the sentiment type is fused to determine a comprehensive text feature vector.
[0150] The second risk assessment result is obtained based on the first weight vector corresponding to the speech feature vector and the second weight vector corresponding to the comprehensive text feature vector.
[0151] In one embodiment, the apparatus further includes:
[0152] The first calculation module is used to calculate the Shapley additive interpretation value corresponding to each feature vector contained in the historical credit data, the voice data, and the text data;
[0153] The second calculation module is used to determine the contribution of each eigenvector to the overall risk outcome based on the Shapley additive interpretation value; the contribution indicates whether each eigenvector has a positive or negative impact on the overall risk outcome.
[0154] In one embodiment, the apparatus further includes:
[0155] The output module is used to output at least one of the following: the comprehensive risk result, the contribution distribution corresponding to the Shapley additive explanation value, the change curve corresponding to the speech feature vector, and the change curve corresponding to the text feature vector;
[0156] The early warning module is used to output an early warning prompt when the overall risk result is greater than or equal to a preset risk threshold; the early warning prompt is used to remind the target user to undergo risk review.
[0157] In one embodiment, the acquisition module 10 is configured to perform the following steps:
[0158] Obtain the original data to be evaluated from the target user;
[0159] The original data to be evaluated is anonymized to obtain candidate data;
[0160] The candidate data is categorized and labeled according to fraudulent behavior and overdue behavior to obtain candidate data with labeled types.
[0161] The labeled candidate data is cleaned and balanced to obtain the preprocessed data to be evaluated.
[0162] Each module in the aforementioned risk assessment device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0163] In one embodiment, a computer device is provided, which can be any terminal, and its internal structure diagram can be as follows: Figure 6 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a sample analysis method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0164] Those skilled in the art will understand that Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0165] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0166] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0167] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0168] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0169] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A risk assessment method, characterized in that, The method includes: Obtain preprocessed data of the target user to be evaluated; the data to be evaluated includes historical credit data, voice data, and text data; Based on the historical credit data, a first risk assessment result is determined; the first risk assessment result characterizes the credit anomaly risk of the target user. Speech feature parameters are extracted based on the speech data, and speech feature vectors are determined based on the speech feature parameters. Text feature parameters are extracted based on the text data, and multiple types of text feature vectors are determined based on the text feature parameters; Based on the speech feature vector and the multiple types of text feature vectors, a second risk assessment result is determined; the second risk assessment result characterizes the real-time abnormal risk of the target user. Based on the first risk assessment result and the second risk assessment result, a comprehensive risk result is determined.
2. The method according to claim 1, characterized in that, The step of extracting speech feature parameters based on the speech data and determining speech feature vectors based on the speech feature parameters includes: Based on the speech data, the fundamental frequency mean, pause time, and standard deviation are extracted; The speech feature vector is determined based on the fundamental frequency mean, the pause time, and the standard deviation; the speech feature vector represents the speech stress of the target user.
3. The method according to claim 1, characterized in that, The step of extracting text feature parameters based on the text data and determining multiple types of text feature vectors based on the text feature parameters includes at least two of the following: Based on the text data, determine the total number of texts and the number of texts containing the target words, and based on the total number of texts and the number of texts containing the target words, determine the text feature vector of the word type; The total number of statements and the parse tree depth of each statement are determined based on the text data, and the text feature vector of the syntax type is determined based on the total number of statements and the parse tree depth. Based on the text data, the total number of words and the vector of each word are determined, and based on the total number of words and the vector of each word, the text feature vector of the semantic type is determined; Based on the text data, the sentiment score of each word and the total number of words are determined, and the text feature vector of the sentiment type is determined based on the total number of words and the sentiment score.
4. The method according to claim 3, characterized in that, The determination of the second risk assessment result based on the speech feature vector and multiple types of text feature vectors includes: At least one of the text feature vectors of the lexical type, the syntactic type, the semantic type, and the sentiment type is fused to determine a comprehensive text feature vector. The second risk assessment result is obtained based on the first weight vector corresponding to the speech feature vector and the second weight vector corresponding to the comprehensive text feature vector.
5. The method according to claim 1, characterized in that, The method further includes: Calculate the Shapley additive interpretation value for each feature vector contained in the historical credit data, the voice data, and the text data; The contribution of each eigenvector to the overall risk outcome is determined based on the Shapley additive interpretation value; the contribution indicates whether each eigenvector has a positive or negative impact on the overall risk outcome.
6. The method according to claim 5, characterized in that, The method further includes: Output at least one of the following: the comprehensive risk result, the contribution distribution corresponding to the Shapley additive explanation value, the change curve corresponding to the speech feature vector, and the change curve corresponding to the text feature vector; If the overall risk result is greater than or equal to a preset risk threshold, an early warning prompt will be output; the early warning prompt is used to remind the target user to undergo risk review.
7. The method according to claim 1, characterized in that, The method further includes: Obtain the original data to be evaluated from the target user; The original data to be evaluated is anonymized to obtain candidate data; The candidate data is categorized and labeled according to fraudulent behavior and overdue behavior to obtain candidate data with labeled types. The labeled candidate data is cleaned and balanced to obtain the preprocessed data to be evaluated.
8. A risk assessment device, characterized in that, The device includes: The acquisition module is used to acquire preprocessed data to be evaluated from the target user; the data to be evaluated includes historical credit data, voice data, and text data. The first risk assessment result module is used to determine a first risk assessment result based on the historical credit data; the first risk assessment result characterizes the credit anomaly risk of the target user. The first feature extraction module is used to extract speech feature parameters based on the speech data and determine a speech feature vector based on the speech feature parameters. The second feature extraction module is used to extract text feature parameters based on the text data, and to determine multiple types of text feature vectors based on the text feature parameters; The second risk assessment result module is used to determine a second risk assessment result based on the voice feature vector and multiple types of text feature vectors; the second risk assessment result characterizes the real-time abnormal risk of the target user. The determination module is used to determine the comprehensive risk result based on the first risk assessment result and the second risk assessment result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.