Credit risk assessment method and device based on multi-modal, equipment and medium

By extracting voiceprint feature vectors from real-time voice data from customer terminal devices and combining them with credit data to calculate the voiceprint stress index, the problem of insufficient utilization of real-time financial stress and unstructured data in existing credit risk assessment methods is solved, thus achieving more accurate credit risk assessment.

CN122155829APending Publication Date: 2026-06-05PING AN TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing credit risk assessment methods rely on static historical data, making it difficult to capture borrowers' current real financial stress or potential fraudulent intentions in real time. Furthermore, they do not make full use of unstructured data from voice interactions, resulting in low assessment accuracy.

Method used

By extracting the voiceprint feature vector from real-time voice data on the customer's terminal device, calculating the voiceprint stress index, and combining it with credit data and a pre-trained risk scoring model, a comprehensive risk score is generated by taking into account the customer's emotional state and credit history.

Benefits of technology

It improves the accuracy of credit risk assessment, enabling it to reflect customers' potential credit risks in real time and provide comprehensive risk assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155829A_ABST
    Figure CN122155829A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence and is suitable for the financial field, and discloses a credit risk assessment method, device, equipment and medium based on multi-modal, the method comprises the following steps: acquiring real-time voice data of a client to be evaluated, extracting a voiceprint feature vector of the real-time voice data on the terminal side of a client terminal device, and calculating a voiceprint stress index; acquiring credit investigation data of the client to be evaluated and calculating a credit investigation score of the client to be evaluated; calculating the variance of the voiceprint stress index and the credit score, and calculating a credit data fusion weight of the client to be evaluated based on the calculated variance; inputting the voiceprint stress index, the credit investigation score and the credit data fusion weight into a pre-trained risk score model, calculating a comprehensive risk score of the client to be evaluated; comparing the comprehensive risk score of the client to be evaluated with a preset risk threshold, and outputting a risk assessment result of the client to be evaluated according to the comparison result. The accuracy of the credit risk assessment of the client is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and is applicable to the financial field, particularly relating to a multimodal credit risk assessment method, apparatus, equipment, and medium. Background Technology

[0002] Credit risk assessment is a core component of modern financial risk control systems, and its accuracy and timeliness directly affect the sound operation of lending businesses. Currently, mainstream assessment methods mainly rely on the analysis of static historical data. However, with the diversification of financial scenarios and the increasing complexity of risk patterns, the limitations of existing technological solutions are becoming increasingly apparent.

[0003] First, traditional credit scoring models heavily rely on static, structured data such as repayment history and debt ratios. This type of data is significantly lagging, making it difficult to capture borrowers' current financial pressure or potential fraudulent intentions in real time and with sensitivity. Furthermore, the model completely ignores unstructured data generated during interactions such as voice communication, including emotional fluctuations and contradictory statements, resulting in lower accuracy in assessing groups lacking stable credit records, such as micro-enterprise owners and freelancers. Second, although voiceprint recognition technology has been applied to identity verification, current voiceprint recognition technology primarily confirms borrower identity and does not deeply couple the rich biometric information contained in voice, such as abnormal speech rate due to stress, fundamental frequency jitter, or voice tremor, into the credit risk model. Therefore, these solutions cannot effectively infer changes in a borrower's willingness or ability to repay through voice features. Third, while some current AI-assisted risk control systems attempt to introduce text semantic analysis, their data dimensions remain limited, mainly processing text information such as loan application forms, and lacking sufficient support and fusion capabilities for multimodal data such as voice and video. This makes it difficult for the system to identify actual high-risk fraudulent activities that appear compliant on the surface in written materials but reveal abnormal signs such as nervousness or concealment in voice communication.

[0004] Therefore, improving the accuracy of customer credit risk assessment is a technical problem that urgently needs to be solved. Summary of the Invention

[0005] This invention provides a multimodal credit risk assessment method, apparatus, equipment, and medium to address the technical problem of improving the accuracy of customer credit risk assessment.

[0006] In a first aspect, the present invention provides a multimodal credit risk assessment method, comprising: Acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. Obtain the credit data of the customer to be evaluated, and calculate the credit score of the customer to be evaluated based on the credit data; Calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance; The voiceprint pressure index, credit score, and credit data fusion weights are input into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comprehensive risk score of the customer to be evaluated is compared with a preset risk threshold, and the risk assessment result of the customer to be evaluated is output based on the comparison result.

[0007] Secondly, the present invention provides a multimodal credit risk assessment device, comprising: The first acquisition module is used to acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. The second acquisition module is used to acquire the credit data of the customer to be evaluated and calculate the credit score of the customer to be evaluated based on the credit data. The first calculation module is used to calculate the variance of the voiceprint stress index and the credit score, and to calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance. The second calculation module is used to input the voiceprint pressure index, credit score and credit data fusion weight into the pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comparison module compares the comprehensive risk score of the customer to be evaluated with a preset risk threshold, and outputs the risk assessment result of the customer to be evaluated based on the comparison result.

[0008] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described multimodal credit risk assessment method.

[0009] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described multimodal credit risk assessment method.

[0010] The aforementioned multimodal credit risk assessment method, apparatus, equipment, and medium, in which real-time voice data of the customer to be assessed is acquired through a client, and voiceprint feature vectors of the real-time voice data are extracted at the client's terminal device. A voiceprint stress index is calculated based on these voiceprint feature vectors. Credit data of the customer to be assessed is acquired, and a credit score is calculated based on this data. The variance between the voiceprint stress index and the credit score is calculated, and a credit data fusion weight is calculated based on the calculated variance. The voiceprint stress index, credit score, and credit data fusion weight are input into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be assessed. The comprehensive risk score of the customer to be assessed is compared with a preset risk threshold, and the risk assessment result is output based on the comparison result. In this invention, by extracting the voiceprint feature vectors of the real-time voice data at the client's terminal device and calculating the voiceprint stress index based on these vectors, the psychological stress level of the customer during communication can be reflected, indicating their potential credit risk. Furthermore, by inputting the voiceprint stress index, credit score, and credit data fusion weights into a pre-trained risk scoring model, a comprehensive risk score for the customer to be assessed is calculated. This comprehensive approach takes into account the customer's emotional state, credit history, and other relevant factors, providing a complete risk assessment result and effectively improving the accuracy of the customer's credit risk assessment. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an application environment for a multimodal credit risk assessment method according to an embodiment of the present invention.

[0013] Figure 2 This is a flowchart illustrating a multimodal credit risk assessment method according to an embodiment of the present invention.

[0014] Figure 3 yes Figure 2 A schematic diagram of a specific implementation method for step S10.

[0015] Figure 4 yes Figure 2 A flowchart illustrating a specific implementation of step S20.

[0016] Figure 5This is a schematic diagram of a multimodal credit risk assessment device according to an embodiment of the present invention.

[0017] Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.

[0018] Figure 7 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The multimodal credit risk assessment method provided in this invention can be applied to, for example... Figure 1 In the application environment, Figure 1 This is a schematic diagram of an application environment for a multimodal credit risk assessment method according to an embodiment of the present invention; wherein, the client communicates with the server via a network. The server can obtain real-time voice data of the customer to be assessed through the client, extract the voiceprint feature vector of the real-time voice data at the client's terminal device, and calculate the voiceprint stress index based on the voiceprint feature vector; obtain the credit data of the customer to be assessed, and calculate the credit score of the customer to be assessed based on the credit data; calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight of the customer to be assessed based on the calculated variance; input the voiceprint stress index, credit score and credit data fusion weight into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be assessed; compare the comprehensive risk score of the customer to be assessed with a preset risk threshold, and output the risk assessment result of the customer to be assessed based on the comparison result. In the present invention, by extracting the voiceprint feature vector of the real-time voice data at the client's terminal device and calculating the voiceprint stress index based on the voiceprint feature vector, the psychological stress level of the customer during the communication process can be reflected, reflecting their potential credit risk. Furthermore, by inputting the voiceprint stress index, credit score, and credit data fusion weights into a pre-trained risk scoring model, a comprehensive risk score for the customer to be assessed is calculated. This model comprehensively considers the customer's emotional state, credit history, and other relevant factors, providing a comprehensive risk assessment result and effectively improving the accuracy of customer credit risk assessment. The invention will now be described in detail through specific embodiments.

[0021] Please see Figure 2 As shown, Figure 2This is a flowchart illustrating a multimodal credit risk assessment method provided in an embodiment of the present invention. The multimodal credit risk assessment method specifically includes the following steps: S10: Acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector from the customer's terminal device, and calculate the voiceprint stress index based on the voiceprint feature vector. Specifically, in this embodiment of the invention, by acquiring the customer's voice data in real time and extracting the voiceprint feature vector, the customer's emotions and psychological state can be accurately grasped, thereby providing important emotional data support for credit risk assessment. The calculation of the voiceprint stress index can reflect the customer's psychological stress level during communication, and thus reflect their potential credit risk. Specifically, as follows... Figure 3 The above, Figure 3 yes Figure 2 A flowchart illustrating a specific implementation of step S10. Specifically, it includes the following steps S11-S12: S11: Extract the Mel-frequency cepstral coefficient feature sequence of the real-time speech data, and calculate the acoustic indicators of the real-time speech data based on the Mel-frequency cepstral coefficient feature sequence. The acoustic indicators include at least the harmonic-to-noise ratio (HNR) enhancement index and the fundamental frequency jitter index. Specifically, in this embodiment of the invention, the Mel-frequency cepstral coefficient is an effective feature used to characterize speech signals, capable of capturing the frequency characteristics of human speech. This feature extraction method can convert speech signals into a digital form that can be used for subsequent analysis, thereby providing basic data for voiceprint analysis. By calculating acoustic indicators such as the HNR enhancement index and the fundamental frequency jitter index, the customer's speech characteristics can be analyzed in depth. The HNR enhancement index reflects the clarity of speech, while the fundamental frequency jitter index reflects the stability and emotional state of speech. These indicators provide a basis for subsequent stress index calculations, helping to more comprehensively understand the customer's psychological state. For example, in the credit application process, banks can assess a customer's emotions and psychological state by analyzing the customer's real-time speech data. If the customer has a low HNR when answering questions, it may indicate that they are nervous or anxious, which may affect their credit risk assessment. Abnormal fluctuations in the base frequency jitter index may also indicate customer dissatisfaction with loan terms or feelings of pressure. This information helps banks assess a customer's willingness and ability to repay, thereby making more reasonable credit decisions. Specifically, this includes the following steps S111: S111: Based on the Mel-frequency cepstral coefficient feature sequence, fundamental frequency detection is performed on the real-time speech data to obtain the fundamental frequency of the real-time speech data, and the harmonic-to-noise ratio (HNR) enhancement index is calculated based on the fundamental frequency of the real-time speech data. Specifically, in this embodiment of the invention, fundamental frequency detection using the Mel-frequency cepstral coefficient feature sequence can effectively capture the main pitch information in the speech signal. The fundamental frequency is the most basic sound frequency in the speech signal, representing the speaker's emotion and psychological state. Accurate fundamental frequency detection provides a basis for subsequent acoustic index calculations, ensuring the reliability of the overall evaluation.

[0022] The calculation formula for the harmonic-to-noise ratio enhancement index is shown in formula (1): (1) in, Indicates the harmonic-to-noise ratio enhancement index. The first term representing the fundamental frequency Second harmonic This represents the spectral energy within the frequency band. Indicates the non-harmonic frequency band. It represents the spectral energy within the frequency band.

[0023] Specifically, in this embodiment of the invention, the harmonic-to-noise ratio (HNR) enhancement index is one of the important parameters for evaluating speech quality, reflecting the clarity and fluency of the speech signal. By calculating the HNR enhancement index, the ratio of harmonic to non-harmonic components in speech can be quantified, thereby reflecting the customer's emotional state. For example, a higher HNR is generally associated with a good psychological state, while a lower HNR may indicate tension, anxiety, or other negative emotions. In financial scenarios, such as when a customer applies for a loan or credit card, the bank can collect the customer's real-time voice data through voice calls. By combining the analysis results of the base frequency and the HNR enhancement index, the bank can dynamically adjust its credit decisions.

[0024] S12: Input the acoustic indicators into a preset stress mapping model to calculate the voiceprint stress index of the customer to be evaluated. Specifically, in this embodiment of the invention, by inputting the acoustic indicators into a preset stress mapping model, the psychological stress level of the customer can be quantified. This model can transform complex acoustic features into a voiceprint stress index, making the assessment results more intuitive and easier to operate. The calculation of the voiceprint stress index can provide important decision support for financial institutions. For example, when conducting customer assessments, banks can use the voiceprint stress index to quickly determine the customer's psychological state and adjust the credit review process according to the stress level. For example, if the voiceprint stress index shows that the customer's stress is high, the bank may choose to conduct a more in-depth risk assessment or provide more financial consulting services to reduce potential default risks. Specifically, this includes the following steps S121-S122: S121: Normalize the harmonic-to-noise ratio (HNR) enhancement index and the fundamental frequency jitter index to obtain normalized HNR enhancement index and fundamental frequency jitter index. Specifically, in this embodiment of the invention, by normalizing the acoustic indices, the dimensional differences between different indices can be eliminated, allowing each index to be compared on the same scale. The normalized indices can more effectively reflect the customer's true psychological state, reduce noise and bias, thereby improving the prediction accuracy of the subsequent stress mapping model.

[0025] S122: The normalized harmonic-to-noise ratio enhancement index and fundamental frequency jitter index are input into a preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated. Specifically, in this embodiment of the invention, the normalized acoustic indicators are input using a preset pressure mapping model to calculate the customer's voiceprint pressure index in real time. This dynamic assessment can promptly reflect the psychological changes of customers during communication, providing real-time evidence for credit risk assessment.

[0026] The formula for calculating the voiceprint stress index of the customer to be evaluated is shown in formula (2): (2) in, Indicates the voiceprint pressure index. This represents the normalized harmonic-to-noise ratio enhancement index. This represents the normalized fundamental frequency jitter index. and These are preset positive weighting coefficients, and .

[0027] S20: Obtain the credit data of the customer to be evaluated, and calculate the credit score of the customer based on the credit data. Specifically, in this embodiment of the invention, the acquisition of credit data and the calculation of the score can provide objective quantitative indicators for the customer's credit history. This ensures that the evaluation is based on reliable historical data, making the credit score results highly authoritative and comparable. As an important component of traditional credit assessment, credit scoring provides basic data support for overall risk assessment, helping to identify the customer's credit status and repayment ability, thereby improving the comprehensiveness of the assessment. Specifically, as follows... Figure 4 The above, Figure 4 yes Figure 2 A flowchart illustrating a specific implementation of step S20. Specifically, it includes the following steps S21-S22: S21: Obtain multi-dimensional structured credit data of the customer to be evaluated, and perform data preprocessing on the obtained structured credit data. Specifically, in this embodiment of the invention, by obtaining multi-dimensional structured credit data, such as income, debt, repayment history, etc., a comprehensive understanding of the customer's credit status can be achieved, avoiding information bias caused by relying on only a single indicator. This provides a solid data foundation for subsequent credit scoring. For example, in the financial field, when issuing loans or credit cards, banks first obtain multi-dimensional credit data of the customer, including income, existing debt, credit card usage, and repayment history. By preprocessing this data, the integrity and accuracy of the data are ensured.

[0028] S22: Based on preset scoring rules, the various credit indicators in the preprocessed structured credit data are weighted to generate an initial credit score. Specifically, in this embodiment of the invention, different credit indicators are comprehensively evaluated to generate an initial credit score, enabling financial institutions to quickly obtain an intuitive credit risk assessment result. For example, when issuing loans or credit cards, banks use preset scoring rules, such as the weight of certain indicators in the overall score, to weight these preprocessed data to generate an initial credit score.

[0029] S23: Normalize the initial credit score to obtain a standardized credit score for the customer to be evaluated. Specifically, in this embodiment of the invention, normalizing the initial credit score can eliminate differences between different scoring systems, making credit scores comparable.

[0030] S30: Calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight for the customer to be evaluated based on the calculated variance. Specifically, in this embodiment of the invention, by calculating the variance of the voiceprint stress index and the credit score, the stability of these two indicators can be assessed, thereby providing a basis for calculating the fusion weight. This effectively identifies the relative importance between the two, making the final credit risk assessment more objective and accurate. The fusion weight calculated based on the variance can flexibly adjust the influence of different data sources, ensuring that the assessment results better reflect the customer's true credit risk status.

[0031] S40: The voiceprint stress index, credit score, and credit data fusion weights are input into the pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be assessed. Specifically, in this embodiment of the invention, information from multiple data sources is integrated into the pre-trained risk scoring model, which can fully utilize the advantages of multimodal data and improve the accuracy and reliability of credit risk assessment. The comprehensive risk score can comprehensively consider the customer's emotional state, credit history, and other relevant factors, thereby providing a comprehensive risk assessment result.

[0032] S41: The voiceprint stress index, credit score, and credit data fusion weights are input into an adaptive weighted neural network to fine-tune the credit data fusion weights, generating contextualized fusion weights. Specifically, in this embodiment of the invention, the adaptive weighted neural network can dynamically adjust the weights of various indicators according to the specific circumstances and context of different customers. This flexibility makes risk assessment more realistic and better reflects the customer's credit risk. The adaptive weighted neural network allows for dynamic adjustment of the weights of various indicators based on the specific circumstances and context of different customers, thus better reflecting the customer's credit risk.

[0033] S42: Based on the contextualized fusion weights, the voiceprint stress index and credit score are non-linearly weighted and fused to generate a comprehensive risk score for the customer to be assessed. Specifically, by fusing the voiceprint stress index with the credit score, a comprehensive risk score can be generated. This comprehensive risk score not only reflects the customer's credit history but also considers their emotional state and psychological stress. This provides financial institutions with a more comprehensive perspective on customer risk, helping to make more accurate credit decisions.

[0034] S50: Compare the comprehensive risk score of the customer to be assessed with a preset risk threshold, and output the risk assessment result of the customer to be assessed based on the comparison result. Specifically, in this embodiment of the invention, by comparing the comprehensive risk score with a preset risk threshold, the credit risk level of a customer can be quickly determined. This helps to make decisions in a short time and supports financial institutions or relevant parties in taking necessary risk management measures. Outputting the risk assessment result based on the comparison result can provide clear guidance for credit decisions, reduce potential credit risks, and improve overall risk management efficiency.

[0035] In this embodiment of the invention, after calculating the voiceprint pressure index based on the voiceprint feature vector and before calculating the variance of the voiceprint pressure index and the credit score, the method further includes: injecting Laplace noise into the voiceprint feature vector of the extracted real-time voice data based on preset privacy protection rules and sensitivity rules; encrypting the voiceprint pressure index after injecting Laplace noise on the client terminal device side, and uploading the encrypted voiceprint pressure index to the cloud server.

[0036] The formula for injecting Laplacian noise into the extracted real-time speech data's speaker feature vector is shown in formula (3): (3) in, This represents the speaker signature feature vector after injecting Laplacian noise. Indicates sensitivity.

[0037] Specifically, in this embodiment of the invention, injecting Laplace noise into the voiceprint feature vector can effectively protect customer privacy. Laplace noise is a commonly used differential privacy technique that can reduce the risk of data leakage while preserving data usability. This is especially important for handling sensitive personal information, ensuring that customer voiceprint data is not maliciously used. Encrypting the voiceprint pressure index after noise injection further enhances data security. This ensures that even if data is intercepted during transmission, attackers cannot decipher this information, thereby effectively preventing data leakage and misuse.

[0038] As can be seen, in the above scheme, by extracting the voiceprint feature vector of the real-time voice data at the client's terminal device and calculating the voiceprint stress index based on the voiceprint feature vector, the psychological stress level of the customer during communication can be reflected, indicating their potential credit risk. Furthermore, by inputting the voiceprint stress index, credit score, and credit data fusion weights into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be assessed, a comprehensive risk assessment result can be provided, taking into account the customer's emotional state, credit history, and other relevant factors, effectively improving the accuracy of customer credit risk assessment.

[0039] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0040] In one embodiment, a multimodal credit risk assessment device is provided, which corresponds one-to-one with the multimodal credit risk assessment method described in the above embodiments. For example... Figure 5 As shown, Figure 5 This is a schematic diagram of a multimodal credit risk assessment device according to an embodiment of the present invention. The multimodal credit risk assessment device includes a first acquisition module 51, a second acquisition module 52, a first calculation module 53, a second calculation module 54, and a comparison module 55. Detailed descriptions of each functional module are as follows: The first acquisition module 51 is used to acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. The second acquisition module 52 is used to acquire the credit data of the customer to be evaluated and calculate the credit score of the customer to be evaluated based on the credit data. The first calculation module 53 is used to calculate the variance of the voiceprint pressure index and the credit score, and to calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance. The second calculation module 54 is used to input the voiceprint pressure index, credit score and credit data fusion weight into the pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comparison module 55 compares the comprehensive risk score of the customer to be evaluated with a preset risk threshold, and outputs the risk assessment result of the customer to be evaluated based on the comparison result.

[0041] In one embodiment, the first acquisition module 51 is specifically used for: Extract the Mel frequency cepstral coefficient feature sequence of the real-time speech data, and calculate the acoustic index of the real-time speech data based on the Mel frequency cepstral coefficient feature sequence. The acoustic index includes at least the harmonic-to-noise ratio enhancement index and the fundamental frequency jitter index. The acoustic indicators are input into a preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated.

[0042] In one embodiment, the first acquisition module 51 is further configured to: The fundamental frequency of the real-time speech data is detected based on the Mel frequency cepstral coefficient feature sequence to obtain the fundamental frequency of the real-time speech data, and the harmonic-to-noise ratio enhancement index is calculated based on the fundamental frequency of the real-time speech data. The formula for calculating the harmonic-to-noise ratio enhancement index is as follows: in, Indicates the harmonic-to-noise ratio enhancement index. The first term representing the fundamental frequency Second harmonic This represents the spectral energy within the frequency band. Indicates the non-harmonic frequency band. It represents the spectral energy within the frequency band.

[0043] In one embodiment, the first acquisition module 51 is further configured to: The harmonic noise ratio enhancement index and the fundamental frequency jitter index are normalized to obtain the normalized harmonic noise ratio enhancement index and fundamental frequency jitter index. The normalized harmonic noise ratio enhancement index and the fundamental frequency jitter index are input into the preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated. The formula for calculating the voiceprint stress index of the customer to be evaluated is as follows: in, Indicates the voiceprint pressure index. This represents the normalized harmonic-to-noise ratio enhancement index. This represents the normalized fundamental frequency jitter index. and These are preset positive weighting coefficients, and .

[0044] In one embodiment, the second acquisition module 52 is specifically used for: Obtain multi-dimensional structured credit data of the customer to be evaluated, and perform data preprocessing on the obtained structured credit data; Based on preset scoring rules, the various credit indicators in the preprocessed structured credit data are weighted to generate an initial credit score. The initial credit score is normalized to obtain a standardized credit score for the customer to be evaluated.

[0045] In one embodiment, the second calculation module 54 is specifically used for: The voiceprint pressure index, credit score, and credit data fusion weights are input into an adaptive weighted neural network to fine-tune the credit data fusion weights and generate contextualized fusion weights. Based on the contextualized fusion weights, the voiceprint stress index and credit score are nonlinearly weighted and fused to generate a comprehensive risk score for the customer to be evaluated.

[0046] In one embodiment, the multimodal-based credit risk assessment device is further used for: Based on preset privacy protection rules and sensitivity rules, Laplace noise is injected into the voiceprint feature vector of the extracted real-time speech data; On the client terminal device side, the voiceprint pressure index after injecting Laplace noise is encrypted, and the encrypted voiceprint pressure index is uploaded to the cloud server.

[0047] This invention provides a multimodal credit risk assessment device. By extracting the voiceprint feature vector from real-time voice data at the customer's terminal device and calculating a voiceprint stress index based on the voiceprint feature vector, it can reflect the customer's psychological stress level during communication and thus their potential credit risk. Furthermore, by inputting the voiceprint stress index, credit score, and credit data fusion weights into a pre-trained risk scoring model, a comprehensive risk score for the customer to be assessed is calculated. This comprehensive approach considers the customer's emotional state, credit history, and other relevant factors, providing a comprehensive risk assessment result and effectively improving the accuracy of customer credit risk assessment.

[0048] Specific limitations regarding the multimodal credit risk assessment device can be found in the limitations of the multimodal credit risk assessment method described above, and will not be repeated here. Each module in the aforementioned multimodal credit risk assessment device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0049] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side multimodal credit risk assessment method.

[0050] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 7 As shown, Figure 7 This is another schematic diagram of a computer device according to an embodiment of the present invention. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of a multimodal credit risk assessment method.

[0051] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. Obtain the credit data of the customer to be evaluated, and calculate the credit score of the customer to be evaluated based on the credit data; Calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance; The voiceprint pressure index, credit score, and credit data fusion weights are input into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comprehensive risk score of the customer to be evaluated is compared with a preset risk threshold, and the risk assessment result of the customer to be evaluated is output based on the comparison result.

[0052] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. Obtain the credit data of the customer to be evaluated, and calculate the credit score of the customer to be evaluated based on the credit data; Calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance; The voiceprint pressure index, credit score, and credit data fusion weights are input into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comprehensive risk score of the customer to be evaluated is compared with a preset risk threshold, and the risk assessment result of the customer to be evaluated is output based on the comparison result.

[0053] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0054] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0055] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0056] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multimodal credit risk assessment method, characterized in that, include: Acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. Obtain the credit data of the customer to be evaluated, and calculate the credit score of the customer to be evaluated based on the credit data; Calculate the variance of the voiceprint stress index and the credit score, and calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance; The voiceprint pressure index, credit score, and credit data fusion weights are input into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comprehensive risk score of the customer to be evaluated is compared with a preset risk threshold, and the risk assessment result of the customer to be evaluated is output based on the comparison result.

2. The multimodal credit risk assessment method according to claim 1, characterized in that, The process of acquiring real-time voice data of the customer to be evaluated, extracting the voiceprint feature vector of the real-time voice data on the end-user side of the customer's terminal device, and calculating the voiceprint pressure index based on the voiceprint feature vector includes: Extract the Mel frequency cepstral coefficient feature sequence of the real-time speech data, and calculate the acoustic index of the real-time speech data based on the Mel frequency cepstral coefficient feature sequence. The acoustic index includes at least the harmonic-to-noise ratio enhancement index and the fundamental frequency jitter index. The acoustic indicators are input into a preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated.

3. The multimodal credit risk assessment method according to claim 2, characterized in that, The process involves extracting the Mel-frequency cepstral coefficient feature sequence from the real-time speech data and calculating acoustic metrics based on the Mel-frequency cepstral coefficient feature sequence. These acoustic metrics include at least a harmonic-to-noise ratio (HNR) enhancement metric and a fundamental frequency jitter metric, including: The fundamental frequency of the real-time speech data is detected based on the Mel frequency cepstral coefficient feature sequence to obtain the fundamental frequency of the real-time speech data, and the harmonic-to-noise ratio enhancement index is calculated based on the fundamental frequency of the real-time speech data. The formula for calculating the harmonic-to-noise ratio enhancement index is as follows: in, Indicates the harmonic-to-noise ratio enhancement index. The first term representing the fundamental frequency Second harmonic This represents the spectral energy within the frequency band. Indicates the non-harmonic frequency band. It represents the spectral energy within the frequency band.

4. The multimodal credit risk assessment method according to claim 2, characterized in that, The step of inputting the acoustic indicators into a preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated includes: The harmonic noise ratio enhancement index and the fundamental frequency jitter index are normalized to obtain the normalized harmonic noise ratio enhancement index and fundamental frequency jitter index. The normalized harmonic noise ratio enhancement index and the fundamental frequency jitter index are input into the preset pressure mapping model to calculate the voiceprint pressure index of the customer to be evaluated. The formula for calculating the voiceprint stress index of the customer to be evaluated is as follows: in, Indicates the voiceprint pressure index. This represents the normalized harmonic-to-noise ratio enhancement index. This represents the normalized fundamental frequency jitter index. and These are preset positive weighting coefficients, and .

5. The multimodal credit risk assessment method according to claim 1, characterized in that, The process of obtaining the credit data of the customer to be evaluated and calculating the credit score of the customer to be evaluated based on the credit data includes: Obtain multi-dimensional structured credit data of the customer to be evaluated, and perform data preprocessing on the obtained structured credit data; Based on preset scoring rules, the various credit indicators in the preprocessed structured credit data are weighted to generate an initial credit score. The initial credit score is normalized to obtain a standardized credit score for the customer to be evaluated.

6. The multimodal credit risk assessment method according to claim 1, characterized in that, The step of inputting the voiceprint stress index, credit score, and credit data fusion weights into a pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be assessed includes: The voiceprint pressure index, credit score, and credit data fusion weights are input into an adaptive weighted neural network to fine-tune the credit data fusion weights and generate contextualized fusion weights. Based on the contextualized fusion weights, the voiceprint stress index and credit score are nonlinearly weighted and fused to generate a comprehensive risk score for the customer to be evaluated.

7. The multimodal credit risk assessment method according to claim 1, characterized in that, After calculating the voiceprint stress index based on the voiceprint feature vector, and before calculating the variance of the voiceprint stress index and the credit score, the method further includes: Based on preset privacy protection rules and sensitivity rules, Laplace noise is injected into the voiceprint feature vector of the extracted real-time speech data; On the client terminal device side, the voiceprint pressure index after injecting Laplace noise is encrypted, and the encrypted voiceprint pressure index is uploaded to the cloud server.

8. A multimodal credit risk assessment device, characterized in that, include: The first acquisition module is used to acquire real-time voice data of the customer to be evaluated, extract the voiceprint feature vector of the real-time voice data on the end side of the customer's terminal device, and calculate the voiceprint pressure index based on the voiceprint feature vector. The second acquisition module is used to acquire the credit data of the customer to be evaluated and calculate the credit score of the customer to be evaluated based on the credit data. The first calculation module is used to calculate the variance of the voiceprint stress index and the credit score, and to calculate the credit data fusion weight of the customer to be evaluated based on the calculated variance. The second calculation module is used to input the voiceprint pressure index, credit score and credit data fusion weight into the pre-trained risk scoring model to calculate the comprehensive risk score of the customer to be evaluated. The comparison module compares the comprehensive risk score of the customer to be evaluated with a preset risk threshold, and outputs the risk assessment result of the customer to be evaluated based on the comparison result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multimodal credit risk assessment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the multimodal credit risk assessment method as described in any one of claims 1 to 7.