VSA-based approval method, system and device, and storage medium
Through a VSA-based method, user audio data is obtained and a hierarchical deep learning model is established, which solves the problem of being unable to evaluate customers with no credit history in traditional financial risk approval, and realizes real-time dynamic risk assessment and coverage improvement.
Patent Information
- Application Number
- CN202510792814.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional financial risk approval technology is unable to effectively assess customers without credit records, resulting in approximately 30% of potential financial customers being denied loans, hindering the advancement of inclusive finance.
Through a VSA-based method, user audio data is obtained, 15-dimensional sound feature information is extracted, and a hierarchical deep learning model is established to perform real-time dynamic risk assessment.
It has achieved real-time dynamic risk assessment of customers with no credit records, increased the coverage rate by 30%, improved the reliability and accuracy of approval, and avoided misjudgments caused by data lag in traditional methods.
Smart Images

Figure CN120634710A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technology, and more specifically, to an approval method, system, device and storage medium based on VSA (Voice Stress Analysis). Background Art
[0002] Traditional financial risk approval technology mainly implements risk assessment based on traditional methods such as credit reports, income certificates, and historical repayment records. It relies on structured data (such as salary flows, credit card records, etc.) and cannot cover "long-tail customers" with no credit records, such as freelancers and small and micro business owners. As a result, about 30% of potential financial customers are rejected for loans due to lack of traditional credit data, making it difficult to promote inclusive finance. Summary of the Invention
[0003] The embodiments of the present invention provide a VSA-based approval method, system, storage medium, device and computer program product, which obtain acoustic features through unstructured user audio data and establish an unstructured data credit model, provide an evaluation basis for customers with no credit history, and can realize real-time dynamic risk assessment.
[0004] According to the first aspect of the present invention, an embodiment of the present invention provides an approval method based on VSA, which includes: obtaining the user's initial audio data; preprocessing the initial audio data to obtain preprocessed standard audio data; hierarchically extracting 15-dimensional sound feature information based on the standard audio data, the 15-dimensional sound feature information including: vowel prolongation rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, resonance peak slope, speech-breathing coordination, kurtosis value, Lyapunov exponent, fractal dimension, intonation index, and emotional deviation; training the initial model based on the 15-dimensional sound feature information to obtain a hierarchical deep learning model; obtaining audio data to be approved according to the user's approval request and inputting the audio data to be approved into the hierarchical deep learning model, and the hierarchical deep learning model outputs the approval result.
[0005] According to the above-mentioned embodiment of the present invention, acoustic features are obtained through unstructured user audio data and an unstructured data credit model is established to provide an evaluation basis for customers with no credit records. In addition, real-time dynamic risk assessment can be achieved by analyzing real-time audio data awaiting approval.
[0006] In some embodiments of the present invention, preprocessing the audio data includes: eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering silence data and / or background sound interference data in the first audio data to obtain second audio data; removing conversation segments irrelevant to approval in the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; converting the fifth audio data into sixth audio data with a preset sampling rate and a preset format; generating seventh audio data that eliminates volume differences among recording devices based on the sixth audio data; and compressing the seventh audio data to obtain the standard audio data.
[0007] In some embodiments of the present invention, the hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model.
[0008] In some embodiments of the present invention, the risk scoring model is obtained by the following steps: obtaining target features in the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; constructing a DeepFM model by combining an FM model and a DNN model; training the DeepFM model based on the target features to obtain the risk scoring model; and dynamically adjusting the decision boundary of the risk scoring model according to the business scenario.
[0009] In some embodiments of the present invention, the multimodal feature fusion network includes: a physical layer, a timing layer, and an emotional layer; wherein, the physical layer performs a first feature processing on the 15-dimensional sound feature information to obtain spectral feature information; the timing layer performs a second feature processing on the 15-dimensional sound feature information to obtain timing feature information; the emotional layer performs a third feature processing on the 15-dimensional sound feature information to obtain emotional feature information and dynamically weights the emotional feature information.
[0010] According to the second aspect of the present invention, an embodiment of the present invention provides an approval system based on VSA, which includes: an initial data acquisition module for acquiring the user's initial audio data; a preprocessing module for preprocessing the initial audio data to obtain preprocessed standard audio data; a 15-dimensional sound feature information extraction module for hierarchically extracting 15-dimensional sound feature information based on the standard audio data, wherein the 15-dimensional sound feature information includes: vowel prolongation rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, resonance peak slope, speech-breathing coordination, kurtosis value, Lyapunov index, fractal dimension, intonation index, and emotional deviation; a model training module for training the initial model based on the 15-dimensional sound feature information to obtain a hierarchical deep learning model; and an approval request module for acquiring audio data to be approved according to the user's approval request and inputting the audio data to be approved into the hierarchical deep learning model, and the hierarchical deep learning model outputs an approval result.
[0011] According to the above-mentioned embodiment of the present invention, acoustic features are obtained through unstructured user audio data and an unstructured data credit model is established to provide an evaluation basis for customers with no credit records. In addition, real-time dynamic risk assessment can be achieved by analyzing real-time audio data awaiting approval.
[0012] In some embodiments of the present invention, the preprocessing module preprocesses the audio data, including: eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering silence data and / or background sound interference data in the first audio data to obtain second audio data; removing conversation segments irrelevant to approval in the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; converting the fifth audio data into sixth audio data with a preset sampling rate and a preset format; generating seventh audio data that eliminates volume differences between recording devices based on the sixth audio data; and compressing the seventh audio data to obtain the standard audio data.
[0013] In some embodiments of the present invention, the hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model.
[0014] In some embodiments of the present invention, the risk scoring model is obtained by the following steps: obtaining target features in the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; constructing a DeepFM model by combining an FM model and a DNN model; training the DeepFM model based on the target features to obtain the risk scoring model; and dynamically adjusting the decision boundary of the risk scoring model according to the business scenario.
[0015] In some embodiments of the present invention, the multimodal feature fusion network includes: a physical layer, a timing layer, and an emotional layer; wherein, the physical layer performs a first feature processing on the 15-dimensional sound feature information to obtain spectral feature information; the timing layer performs a second feature processing on the 15-dimensional sound feature information to obtain timing feature information; the emotional layer performs a third feature processing on the 15-dimensional sound feature information to obtain emotional feature information and dynamically weights the emotional feature information.
[0016] According to the third aspect of the present invention, an embodiment of the present invention provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the computer performs the following operations: the operations include the steps included in the VSA-based approval method described in any of the above embodiments.
[0017] According to a fourth aspect of the present invention, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory is used to store one or more computer-readable instructions, wherein the one or more computer-readable instructions, when executed by the processor, can implement the VSA-based approval method as described in any one of the above embodiments.
[0018] According to a fifth aspect of the present invention, an embodiment of the present invention provides a computer program product including a computer program, which, when executed by a processor, implements the VSA-based approval method as described in any one of the above embodiments.
[0019] From the above, it can be seen that the VSA-based approval method, system, storage medium, device and computer program product provided by the embodiments of the present invention obtain acoustic features through unstructured user audio data and establish an unstructured data credit model, providing an evaluation basis for customers with no credit record, and by analyzing real-time audio data to be approved, real-time dynamic risk assessment can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 1 is a flowchart of a VSA-based approval method according to Example 1 of the present invention;
[0021] Figure 22 is a flowchart of a VSA-based approval method according to Example 2 of the present invention;
[0022] Figure 3 is a schematic diagram of a process for collecting and preprocessing sound data according to an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of a process for constructing a hierarchical deep learning model according to one embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of a process for constructing a multimodal feature fusion network according to an embodiment of the present invention;
[0025] Figure 6 2 is a schematic diagram of the architecture of a VSA-based approval system according to Example 3 of the present invention. DETAILED DESCRIPTION
[0026] The various aspects of the present invention are described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, well-known modules, units and their connections, links, communications or operations are not shown or described in detail. In addition, the described features, architectures or functions can be combined in any manner in one or more embodiments. It should be understood by those skilled in the art that the various embodiments described below are only for illustration and are not intended to limit the scope of protection of the present invention. It can also be easily understood that the modules or units or processing methods in the various embodiments described herein and shown in the drawings can be combined and designed in various different configurations.
[0027] The following is a brief explanation of the terms used below.
[0028] App:Application, application.
[0029] U-Net structure: a convolutional neural network architecture.
[0030] WAV format: the most basic lossless audio format.
[0031] WAV format: Waveform Audio File Format, a lossless audio format stored based on PCM (Pulse Code Modulation) technology.
[0032] PCM format: Pulse Code Modulation, is a technology that converts analog audio signals into digital signals. PCM format is the "mother tongue" of digital audio and is also the underlying data format of formats such as WAV.
[0033] Hilbert Envelope Analysis: Hilbert Envelope Analysis, a signal processing technique based on the Hilbert Transform, is mainly used to analyze the amplitude modulation characteristics of non-stationary signals (such as vibration signals and speech signals), extract the envelope information of the signal, and thus reveal the instantaneous amplitude and frequency changes of the signal.
[0034] Glottal source signal: The excitation signal generated by the glottis (the vibrating area of the vocal cords) during speech production is the initial energy source of the speech waveform. It is essentially a periodic or quasi-periodic sequence of airflow pulses that directly determines the fundamental characteristics of speech pitch (fundamental frequency) and timbre.
[0035] Vocal tract response: Vocal tract response refers to the frequency response produced by the filtering effect of the shape and size of the vocal tract when the glottal source signal passes through the vocal tract (the resonance cavity from the glottis to the lips, including the laryngeal cavity, pharyngeal cavity, oral cavity and nasal cavity). It determines the resonance characteristics (such as resonance peaks) and timbre details of the speech.
[0036] Pure glottal waveform: A pure glottal waveform refers to the raw time-domain signal generated solely by glottal vibration, without the influence of subsequent processes such as vocal tract filtering and nasal coupling. This idealized representation of the glottal source signal is used to isolate and study the essential characteristics of vocal cord vibration.
[0037] Core Vowel Segment: The stable part of vowel pronunciation.
[0038] DTW algorithm: Dynamic Time Warping, dynamic time warping algorithm, an algorithm used to measure the similarity between two time series (such as speech signals).
[0039] MFCC: Mel-Frequency Cepstral Coefficients, a feature extraction method used in speech processing and audio analysis, used to convert speech signals into a set of feature parameters that can characterize the auditory characteristics of the human ear.
[0040] VAD detection: Voice Activity Detection (VAD) can accurately segment valid speech regions from continuous signals and eliminate invalid silence or noise segments, thereby improving the efficiency and robustness of speech processing systems.
[0041] LPC technology: Linear Predictive Coding, linear predictive coding technology.
[0042] Pearson Correlation Coefficient: Pearson Correlation Coefficient, an indicator in statistics that measures the degree of linear correlation between two variables.
[0043] Lyapunov index: Lyapunov characteristic index, used to represent the numerical characteristics of the average exponential divergence rate of adjacent trajectories in phase space.
[0044] LSTM: Long Short-Term Memory, long short-term memory network.
[0045] GMM model: Gaussian Mixture Model, a model based on probability statistics.
[0046] HMM model: Hidden Markov Model, a statistical time series model.
[0047] KL divergence: Kullback-Leibler Divergence, also known as relative entropy, is an indicator in information theory that measures the difference between two probability distributions.
[0048] GDPR: General Data Protection Regulation.
[0049] CASIA Public Datasets: A series of datasets created and maintained by the Institute of Automation, Chinese Academy of Sciences.
[0050] XGBoost: eXtreme Gradient Boosting, a machine learning algorithm based on the gradient boosting framework (GradientBoosting).
[0051] Focal Loss: Focal loss, a loss function designed to solve the problem of sample class imbalance in target detection.
[0052] SHAP: SHapley Additive exPlanations, an interpretability tool for explaining the predictions of machine learning models.
[0053] TensorRT: A high-performance deep learning inference optimizer and runtime library designed to accelerate the deployment of trained deep learning models on NVIDIA GPUs while reducing memory usage and power consumption.
[0054] BIM: Building Information Modeling.
[0055] [Example 1]
[0056] Figure 1 4 is a flowchart of a VSA-based approval method according to Example 1 of the present invention.
[0057] like Figure 1 As shown, in embodiment 1 of the present invention, the VSA-based approval method may include at least the following steps S11, S12, S13, S14 and S15, which are described in detail below.
[0058] In step S11, the user's initial audio data is obtained.
[0059] In some embodiments, the initial audio data includes but is not limited to customer audio obtained and recorded from one or more of the following scenarios: customer service hotline, video face-to-face signing, smart question and answer, mobile app, and other scenarios.
[0060] In step S12, the initial audio data is preprocessed to obtain preprocessed standard audio data.
[0061] In some embodiments, preprocessing the audio data includes: eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering silence data and / or background sound interference data in the first audio data to obtain second audio data; removing conversation segments unrelated to approval in the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; converting the fifth audio data into sixth audio data with a preset sampling rate and a preset format; generating seventh audio data that eliminates volume differences among recording devices based on the sixth audio data; and compressing the seventh audio data to obtain the standard audio data.
[0062] In an optional implementation manner, preprocessing the audio data specifically includes the following steps:
[0063] (1) Spectral subtraction is used to process steady-state noise (such as air conditioning noise), and a deep noise reduction network with a U-Net structure is used to eliminate non-stationary noise (such as sudden noise and keyboard tapping). The input layer of the U-Net structure receives the time-frequency spectrum of the noisy speech, and the output layer reconstructs the clean speech. The fusion of spectral subtraction and a deep network (deep noise reduction network, U-Net) can eliminate steady-state and non-stationary noise in audio data, and the signal-to-noise ratio (SNR) can be improved to ≥20dB1.
[0064] (2) A dual threshold algorithm based on short-time energy (STE) and zero-crossing rate (ZCR) is used to segment valid speech segments and filter out silence or background human voice interference in audio data.
[0065] (3) Use ASR (automatic speech recognition) technology to verify whether the voice content is relevant to the business (such as loan purpose, loan amount, etc.) and eliminate irrelevant dialogue fragments.
[0066] (4) Use frequency domain perturbation technology (such as Random Phase Shifting) to achieve voiceprint anonymization. Using frequency domain perturbation technology to desensitize voiceprints and retain only emotional / psychological characteristics for scoring can avoid identity information leakage. In this way, while retaining emotional characteristics (such as tension), privacy regulations such as GDPR can be met. Taking the CASIA public dataset and data from financial institutions as an example, compared with traditional frequency domain masking technology, the use of frequency domain perturbation technology to process speech reduces speech continuity loss by 60%.
[0067] (5) The speech stream is cut into segments of a preset fixed length (e.g., 10 seconds per segment). The portion that is less than the preset length is padded with zeros to avoid the problem of window mismatch during subsequent feature extraction.
[0068] (6) Convert the audio data into a structure with a preset sampling rate and a preset format. For example, the unified preset sampling rate is 44.1kHz to cover the human ear's perception range of 20Hz-20kHz; the conversion format is WAV / PCM format to be compatible with different analysis tools.
[0069] (7) By calculating the root mean square energy (RMS) of each speech segment and linearly scaling it to a standard value of -20dBFS (decibel full scale), the average volume of different speech segments is unified to the same energy level, eliminating volume differences caused by recording devices. This can solve the spectrum distortion problem caused by hardware differences in traditional technologies and support data consistency across platforms (such as phones, apps, and face-to-face signing devices).
[0070] (8) Use the μ-law compression and expansion algorithm (e.g., μ = 255) to map the 16-bit PCM signal to 8 bits, preserving speech details while reducing high-frequency noise sensitivity and data size.
[0071] In step S13, 15-dimensional sound feature information is extracted hierarchically based on the standard audio data. The 15-dimensional sound feature information includes: vowel lengthening rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, formant slope, speech-respiration coordination, kurtosis value, Lyapunov exponent, fractal dimension, intonation index, and emotional deviation. Among them, vowel lengthening rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, and fundamental frequency standard deviation are sound feature information used for physiological basis layer detection; speech rate gradient change information, unnatural pause information, formant slope, and speech-respiration coordination are sound feature information used for dynamic behavior layer detection; kurtosis value, Lyapunov exponent, and fractal dimension are sound feature information used for high-order system layer detection; and intonation index and emotional deviation are sound feature information used for emotional stress layer detection.
[0072] In some embodiments, 15-dimensional sound feature information is extracted by the following method:
[0073] (1) Vowel lengthening rate: Linear predictive coding (LPC) technology is used to process the speech signal, accurately extract the starting and ending points of the target vowel segment, and determine the duration of the core vowel segment.
[0074] Calculate the core vowel duration ratio, T_vowel, as T_vowel = (t_end - t_start) / total_duration, where t_end is the starting point of the target vowel segment, t_start is the ending point of the target vowel segment, and total_duration is the total duration of the target vowel segment. If the core vowel duration ratio exceeds a preset recommended threshold (e.g., 200ms), it is determined to be a deceptive feature, indicating possible financial risk, such as intentional misleading.
[0075] (2) Amplitude Shimmer: Based on Hilbert envelope analysis, the amplitude fluctuation characteristics of the sound wave are extracted. By calculating the difference ratio of the peak values of adjacent cycles, the degree of short-term amplitude fluctuation in the steady state of speech is quantitatively described. Amplitude shimmer information can reflect the vibration health of vocal organs such as the vocal cords and the tension of the autonomic nervous system. For example, in financial scenarios, abnormal amplitude shimmer may mean that the participants are in a state of high tension or anxiety, thereby affecting the objectivity and stability of financial decision-making. Therefore, abnormal amplitude shimmer can indicate financial risks.
[0076] (3) Fundamental frequency jitter rate (Jitter): The fundamental frequency trajectory is extracted through Morlet wavelet transform, and the percentage fluctuation of the frequency change in adjacent cycles is calculated. When the fluctuation rate exceeds the volatility threshold (for example, 3%), it is considered abnormal and indicates risk. Abnormal fluctuations in the fundamental frequency jitter rate directly reflect the stability of vocal cord vibration and the regulatory function of the autonomic nervous system. In scenarios such as financial transactions, the fundamental frequency jitter rate of those who spread false information or individuals with potential risks often changes significantly. Therefore, abnormal behavior can be identified by the fundamental frequency jitter rate.
[0077] (4) Glottal closure rate: The IAIF (Iterative Adaptive Inverse Filtering) algorithm is used to separate the glottal source signal and the vocal tract response from the original speech signal to obtain a pure glottal waveform. Furthermore, the turbulent noise caused by incomplete closure is quantified. When the glottal closure rate is greater than the glottal closure rate threshold (for example, 15%), it is determined that the user has engaged in fraudulent behavior. In financial risk assessment, by analyzing the changes in the glottal closure rate of the customer or the counterparty, it can be used to assist in identifying whether the customer has deliberately concealed information.
[0078] (5) Micro-tremor energy ratio: Wavelet decomposition (such as Morlet wavelet) is used to separate the energy of the preset frequency band (6-12Hz frequency band energy), and the ratio of the preset frequency band energy to the total sound wave energy is calculated. When the ratio exceeds the preset threshold, it is determined to be abnormal tremor and a risk is prompted. For example, the micro-tremor energy ratio of normal speech is less than 0.5%, while it can reach 1.2-2.3% in a deceptive state. Combined with baseline calibration (micro-tremor level during neutral conversation), when the micro-tremor energy ratio exceeds the individual baseline value by a preset multiple (for example, 1.8 times), it is determined to be abnormal tremor and a financial risk is prompted (possibly suspected fraud).
[0079] (6) Fundamental frequency standard deviation (F0-SD): The fundamental frequency sequence is extracted through short-time Fourier transform and its standard deviation is calculated. The fundamental frequency standard deviation reflects the degree of fluctuation of the overall fundamental frequency of the speech. For example, when F0-SD exceeds a preset threshold (e.g., 35Hz), the speaker may be deliberately controlling the speech, such as attempting to conceal negative information in financial activities or misleading risk assessments.
[0080] (7) Speech rate gradient change: Detecting sudden changes in speech rate using the DTW algorithm, specifically including the following steps: a. DTW constructs a cumulative distance matrix through speech framing (for example, the time range of the framing is 20-40ms) and feature extraction (such as MFCC). Then, through optimal path search, dynamic programming is used to find the path with the minimum cumulative distance, reflecting the nonlinear expansion and contraction of the time axis. b. Determining sudden changes in speech rate: The path slope exceeds the normal range (for example, the slope is greater than 1.2 syllables / second2 during deception), and the cumulative distance peak increases significantly (for example, the cumulative distance peak of a deceptive person can reach 2.3 times that of normal speech). The rationality of the speech rate change is verified by combining the endpoint information of voice activity detection (VAD). In financial risk prediction, by detecting speech rate gradient changes during the conversation process, potential abnormal information and behavior patterns can be determined.
[0081] (8) Unnatural pauses: VAD is used to detect unnatural pauses. The specific steps include the following: VAD technology identifies speech segments and silence segments based on the energy, spectrum, and other characteristics of the audio signal, and counts the frequency of non-semantic pauses. The normal silence segment frequency is 20%-35%, while the frequency of silence segments for fraudsters can exceed 5 times / minute due to increased cognitive load. In the financial field, by monitoring unnatural pauses, we can gain insight into the cognitive state and psychological stress changes of both parties in the communication, and provide early warning of risks.
[0082] (9) Formant slope: Using LPC technology, the vocal tract is modeled as a full-pole model. The LPC coefficients are calculated frame by frame, and the time-varying slope of F1 / F2 (F1 / F2 are the first two formant peaks) is obtained through polynomial fitting. In the deceptive state, abnormal vocal tract muscle tension causes the formant slope to exceed a preset threshold (e.g., 35 Hz / ms). In financial risk prediction, changes in the formant slope can be used as an indicator to assess the psychological state of traders or financial participants, helping to determine whether they are engaging in risky behavior.
[0083] (10) Speech-respiration coordination: The VAD algorithm is used to identify speech segments and silent breathing segments. Coordination is quantified by the proportion of inhalation time (normal is 20%-35%, and fraudsters generally have less than 15%) and the slope of respiratory recovery (which decreases by more than 40% under stress). The Pearson correlation coefficient between speech rate and inhalation depth is also calculated. Normal speech is greater than 0.7, while fraudsters have a value less than 0.4. In financial scenarios, coordination indicators can help assess potential risks in financial transactions.
[0084] (11) Kurtosis value after dynamic range compression: Calculate the average of the fourth power deviations of all sample points from the mean and the fourth power of the standard deviation. Then, divide the average of the fourth power deviations of all sample points from the mean as the numerator and the fourth power of the standard deviation as the denominator and subtract 3 to obtain the kurtosis value.
[0085] Normal speech has a kurtosis close to zero, exhibiting a natural decay characteristic; whereas deceptive speech has concentrated peaks or a flattened distribution, resulting in a kurtosis greater than or less than 0. In financial risk prediction, analyzing the kurtosis value after dynamic range compression can identify abnormal speech signals and, in turn, uncover potential risk points in financial activities.
[0086] (12) Lyapunov exponent: Based on Takens' embedding theorem, the one-dimensional speech signal is reconstructed into a high-dimensional phase space. The Wolf algorithm is used to track the evolution of adjacent trajectories in the phase space and calculate the maximum Lyapunov exponent. When the exponent exceeds a preset threshold (e.g., 0.8), it indicates that the speech signal has a high degree of chaos, that is, the speaker may be under physiological stress or cognitive load. This can be used to identify abnormal behavior in financial risk prediction, for example.
[0087] (13) Fractal dimension: First, the speech signal is converted into a time-domain waveform matrix and normalized. A short-time Fourier transform is used to generate a time-frequency graph as a two-dimensional fractal analysis object. Second, a box size sequence is set, and the time-frequency graph is covered with different square grids. The minimum number of boxes covering non-zero energy points is counted, and then a linear regression is performed on log(N(s)) and log(1 / s) to obtain the fractal dimension D, where log(N(s)) is the logarithm of the number of boxes required to cover the fractal object, and log(1 / s) is the logarithm of the reciprocal of the box size. When the fractal dimension is less than a preset threshold (e.g., 1.2), it is considered to be intentionally controlled speech, which can be used to assist behavioral analysis in financial risk prediction.
[0088] (14) Intonation Rising Index: The LSTM model establishes long-term dependencies between speech fundamental frequency sequences through cell states and gated units. In deception scenarios, abnormal temporal fluctuations in fundamental frequency (e.g., a rise of more than 40Hz within 0.5 seconds) will trigger unnatural rising tone detection, indicating possible risky behavior.
[0089] (15) Emotional Deviation: A GMM model is used to model the probability density of neutral emotional speech features and construct a baseline acoustic feature distribution. The speech features to be tested are extracted, decoded through the HMM model to obtain the observed probability distribution, and the KL divergence with the neutral speech model is calculated. When the KL divergence exceeds a preset threshold (e.g., 1.2), it indicates that there is significant emotional stress or cognitive load in the speech, which can be used for financial risk prediction.
[0090] Using voice sentiment analysis to identify a customer's willingness to repay (e.g., the correlation between tension and deception) can replace some paper materials or structured data. Furthermore, by building unstructured data credit models based on acoustic features (such as speech speed, intonation fluctuations, and emotional consistency), this can provide an assessment basis for customers without a credit history (e.g., freelancers), increasing potential user coverage by 30% and thus breaking through traditional reliance on credit data. Furthermore, using voice sentiment analysis to identify whether a customer is engaging in fraudulent behavior not only avoids the misjudgment caused by forged documents (such as false bank statements) during manual review of documents such as proof of income, but also enables objective, quantitative analysis of a customer's true emotions and psychological state, improving approval reliability.
[0091] In step S14, the initial model is trained according to the 15-dimensional sound feature information to obtain a hierarchical deep learning model.
[0092] In some embodiments, the hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model.
[0093] In a further embodiment, the multimodal feature fusion network includes: a physical layer, a temporal layer, and an emotional layer. The physical layer performs a first feature processing on the 15-dimensional sound feature information to obtain spectral feature information; the temporal layer performs a second feature processing on the 15-dimensional sound feature information to obtain temporal feature information; and the emotional layer performs a third feature processing on the 15-dimensional sound feature information to obtain emotional feature information, and the emotional feature information is dynamically weighted.
[0094] In some embodiments, the risk scoring model is obtained by the following steps: obtaining target features in the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; constructing a deep factorization machine DeepFM model by combining a factorization machine FM model and a deep neural network DNN model; training the DeepFM model according to the target features to obtain the risk scoring model; and dynamically adjusting the decision boundary of the risk scoring model according to the business scenario.
[0095] In step S15, based on the user's approval request, the audio data to be approved is obtained and input into the hierarchical deep learning model, which then outputs the approval result. Real-time analysis of voice interaction (e.g., sudden changes in acoustic parameters during the Q&A session) can promptly capture changes in a customer's financial situation (e.g., vocal tremors caused by sudden financial stress), enabling millisecond-level risk warnings.
[0096] The above-mentioned VSA-based approval method of Example 1 of the present invention is adopted to obtain acoustic features through unstructured user audio data and establish an unstructured data credit model, so as to provide an evaluation basis for customers with no credit records. In addition, by analyzing real-time audio data to be approved, real-time dynamic risk assessment can be achieved, which solves the data lag and static problems caused by the current traditional financial risk approval method based on credit reports due to the long update cycle of credit reports (usually 1-3 months) and the inability to reflect the real-time economic status of customers (such as sudden unemployment, short-term debt, etc.).
[0097] [Example 2]
[0098] Figure 2 4 is a flowchart of a VSA-based approval method according to embodiment 2 of the present invention.
[0099] like Figure 2 As shown, in embodiment 2 of the present invention, the VSA-based approval method includes at least the following steps S21, S22, S23 and S24, which are described in detail below.
[0100] Step S21: sound data collection and preprocessing. In this embodiment, Figure 3 As shown, step S21 at least includes the following steps S211, S212, S213, S214, S215, S216, S217, S218 and S219, which are described in detail below.
[0101] In step S211, customer audio is collected in scenarios such as customer service hotlines, video face-to-face interviews, intelligent Q&A, and mobile apps. In this embodiment, customer audio collection supports ordinary microphones (0.5m diameter antennas, mobile phone microphones, etc.), making the equipment cost only 1 / 33 of that of traditional voiceprint recognition systems, thereby reducing hardware costs. In addition, it can adapt to scenarios such as telephone customer service (focusing on intonation analysis), video face-to-face interviews (synchronizing micro-expression data), and mobile terminals (wireless transmission). Furthermore, feature weights are automatically assigned through the attention mechanism, achieving multi-scenario coverage and improving scenario compatibility.
[0102] In step S212, spectral subtraction is used to process steady-state noise (such as air conditioning sounds), and a deep noise reduction network with a U-Net structure is used to eliminate non-stationary noise (such as sudden noise and keyboard tapping). The input layer of the U-Net structure receives the time-frequency spectrum of the noisy speech, and the output layer reconstructs the clean speech. By combining spectral subtraction with a deep network (deep noise reduction network, U-Net) to eliminate steady-state and non-stationary noise in audio data, the signal-to-noise ratio (SNR) can be increased to ≥20dB1.
[0103] In step S213, a dual threshold algorithm based on short-time energy and zero-crossing rate is used to segment the effective speech segments, and to filter out silence or background human voice interference in the audio data.
[0104] In step S214, the ASR technology is used to verify whether the voice content is relevant to the business (such as the purpose of the loan, the loan amount, etc.), and irrelevant dialogue segments are eliminated.
[0105] In step S215, frequency domain perturbation technology is used to anonymize the voiceprint. Using frequency domain perturbation technology to desensitize the voiceprint, retaining only emotional / psychological characteristics for scoring, can prevent identity leakage. This allows emotional characteristics (such as stress) to be preserved while meeting privacy regulations such as GDPR.
[0106] In step S216, the speech stream is cut into segments of a preset fixed length (eg, 10 seconds per segment), wherein the portion that is less than the preset length is padded with zeros to avoid a window mismatch problem during subsequent feature extraction.
[0107] In step S217, the audio data is converted to a structure with a preset sampling rate and a preset format. For example, the unified preset sampling rate is 44.1kHz to cover the human ear's perception range of 20Hz-20kHz; the format is converted to WAV / PCM format to be compatible with different analysis tools.
[0108] In step S218, the RMS energy of each speech segment is calculated and linearly scaled to a standard value of -20dBFS (decibel full scale), thereby unifying the average volume of different speech segments to the same energy level and eliminating volume differences caused by recording equipment. This solves the spectral distortion problem caused by hardware differences in traditional technologies and supports data consistency across platforms (such as phones, apps, face-to-face signing devices, etc.).
[0109] In step S219, detail-preserving data compression is performed by using a μ-law companding algorithm (eg, μ=255) to map the 16-bit PCM signal to 8 bits, thereby preserving speech details while reducing high-frequency noise sensitivity and data size.
[0110] Step S22: 15-dimensional sound feature hierarchical extraction. Step S22 is the same as step S13 in Example 1 and will not be described again here.
[0111] Step S23: Layered deep learning model construction. In this embodiment, Figure 4 As shown, step S23 at least includes the following steps S231, S232, S233, S234 and S235, which are described in detail below.
[0112] In step S231, a multimodal feature fusion network is designed. In this embodiment, Figure 5 As shown, the multimodal feature fusion network is derived through physical layer (1D-CNN) feature processing, temporal dynamic layer (Bi-LSTM) modeling, and emotional layer (Attention) attention weighting. In further implementations, similar to the concept of BIM technology integration cloud platform, a multidimensional database is established, integrating acoustic features (spectral, temporal, etc.) with external data (such as repayment records), thereby achieving full lifecycle management of risk decisions.
[0113] Among them, physical layer feature processing includes: using 1D-CNN (one-dimensional convolutional neural network) to extract spectral features (such as fundamental frequency, resonance peak, etc.), and capturing local frequency domain patterns through convolution kernels, such as using a 3×1 convolution kernel to extract the harmonic structure of the short-time spectrum.
[0114] Temporal dynamic layer modeling includes: using Bi-LSTM (bidirectional long short-term memory network) to process temporal features such as speaking rate and pause duration, and capturing long-term dependencies, such as calculating the time series changes of intonation fluctuations through hidden layer units.
[0115] Emotional layer attention weighting includes: introducing a multi-head attention mechanism (Multi-head Attention) to dynamically weight emotional features (such as jitter and glottal closure rate). For example, the correlation weights of different psychological features are calculated using the query-key matrix.
[0116] Separate modeling of the physical layer, timing layer, and emotional layer can achieve hierarchical feature decoupling analysis and avoid misjudgments caused by the coupling of dialects, environmental noise, and psychological characteristics in traditional technologies.
[0117] like Figure 5 As shown in Figure 1, the multimodal feature fusion network design also includes: feature splicing, fully connected layer, and fusion feature output.
[0118] Among them, through feature splicing, the spectral feature information obtained by physical layer processing, the time series feature information obtained by time series layer processing, and the emotional feature information obtained by emotional layer processing and dynamically weighted are spliced in a certain order. For example, the spectral feature information is spliced first, then the time series feature information, and finally the emotional feature information. Furthermore, the spectral feature information, time series feature information, and emotional feature information are merged in the feature dimension to form a longer feature vector. Therefore, through feature splicing, the features obtained by independent processing of the physical layer, time series layer, and emotional layer are integrated together, providing a more comprehensive and rich feature representation for the subsequent fully connected layer processing.
[0119] The fully connected layer receives the concatenated feature vector obtained from the feature concatenation step and performs further feature extraction and transformation on the concatenated features to explore the complex relationships between the features. Each input unit in the fully connected layer establishes a one-to-one connection with all output units through a weight matrix, meaning that each element in the input feature vector participates in the calculation of all output units. Specifically, the fully connected layer performs a linear transformation on the concatenated feature vector, calculating it using a weight matrix W and a bias vector b to obtain the output vector y, calculated as y = Wx + b, where x is the concatenated feature vector. An activation function, such as the ReLU (Rectified Linear Unit) function, is then applied to the result of the linear transformation to introduce nonlinear factors and enhance the model's expressiveness. The output after the activation function serves as the final output of the fully connected layer. Based on the features processed by the physical layer, temporal layer, and emotional layer, the fully connected layer maps these features into a new feature space through linear transformation and nonlinear activation, preparing for the final fused feature output.
[0120] The fused feature output is a direct result of the fully connected layer's output. After processing by the fully connected layer, the resulting feature vector is the final fused feature. This fused feature incorporates the various feature information extracted by the physical layer, temporal layer, and sentiment layer. Further conversion, mining, and fusion by the fully connected layer avoids misjudgments caused by the coupling of dialects, environmental noise, and psychological characteristics used in traditional techniques, providing a more accurate and discriminative feature representation for subsequent applications.
[0121] In step S232, the risk scoring model is constructed and optimized. Specifically, the following steps are included:
[0122] (1) Feature importance screening. Feature selection is performed based on XGBoost. Key dimensions (such as formant shift and fundamental frequency standard deviation) are screened through gain analysis (Gain Score), and redundant features (such as dialect-related fundamental frequency differences) are eliminated.
[0123] (2) Hybrid model training. DeepFM (deep factorization machine) is used to combine shallow feature crosstalk with deep nonlinear modeling. For example, the shallow part uses FM (factorization machine) to capture the second-order interaction between the physical layer and the emotional layer; the deep part uses DNN (deep neural network) to learn high-order feature combinations.
[0124] (3) Dynamic threshold adjustment. By introducing a federated learning framework, the scoring threshold is optimized in real time and the approval threshold (i.e., decision boundary) is dynamically adjusted based on business scenarios and business risk preferences (e.g., economic cycle fluctuations, loan usage, etc.). Compared with traditional static models (e.g., fixed voiceprint similarity thresholds), the flexibility of the approval strategy is increased by more than 30%. For example, during low-risk periods, the approval threshold can be lowered to increase the approval rate.
[0125] In step S233, the layered loss function and joint training are performed. Specifically, the following steps are included:
[0126] (1) Hierarchical loss design. The physical layer loss uses mean square error (MSE) to constrain the spectrum reconstruction accuracy (such as FFT reconstruction error); the high-level semantic loss uses cross-entropy to optimize the risk classification results and combines it with focal loss to solve the sample imbalance problem.
[0127] (2) Joint training strategy: Alternating training is used to freeze high-level network parameters to train the physical layer, and then unfreeze all parameters for end-to-end fine-tuning to avoid gradient conflicts.
[0128] In step S234, the model is verified and continuously optimized. In this embodiment, on the one hand, by integrating SHAP (SHapley Additive exPlanations) value analysis, the impact of 15-dimensional features on the score is visualized, such as the high-risk contribution corresponding to each acoustic dimension such as the gradient change of speech speed, thereby enhancing interpretability and being able to support regulatory audit requirements. On the other hand, by deploying the online learning module, misjudgment samples (such as customer voices that have been approved but overdue) are received in real time, and the model parameters are updated through small-batch gradient descent to achieve incremental learning. Thus, the incremental learning mechanism is used to support online updates of the model, and the iteration efficiency of misjudgment samples is increased by 50%.
[0129] In step S235, lightweight deployment and real-time decision-making are performed. In this embodiment, lightweight deployment includes using TensorRT to quantize the model (INT8 precision) and prune the model to achieve model compression, thereby increasing the inference speed to within 5ms to meet the real-time requirements of phone approval.
[0130] In some embodiments, a dual-threshold mechanism is used to implement rule engine linkage. For example, when the model output score is greater than 90, the application is automatically approved; when the model output score is less than 60, the user's approval request is directly rejected; when the model output score is between 60 and 90, a manual review is triggered, and the weight of the associated voice keyword detection (such as "urgent need for funds") is adjusted (for example, the weight of the "faster speaking speed" feature corresponding to "urgent need for funds" is increased by 30%).
[0131] Step S24: Risk Decision and Feedback: The model is optimized through feedback such as the results of automated approvals, corrections from manual reviews, and downstream business feedback (such as actual customer repayment performance after loan issuance).
[0132] The above-mentioned VSA-based approval method of Example 2 of the present invention is adopted to achieve the following effects through 15-dimensional analysis of sound signals and deep learning dynamic modeling: (1) An unstructured data credit model is established through acoustic features (such as speaking speed, intonation fluctuation, and emotional consistency) to provide an evaluation basis for customers with no credit records, thus achieving coverage of the group with no credit records. (2) Real-time dynamic risk assessment is achieved based on real-time analysis of voice interaction. (3) Voiceprint biometric recognition is achieved by detecting identity fraud through comparison with the voiceprint library, and psychological layer feature analysis is performed through parameters such as fundamental frequency mutation and energy anomaly to identify lying behavior (accuracy is increased by about 25%), achieving anti-fraud and objective quantitative analysis of the customer's true emotions and psychological state. (4) Frequency domain perturbation technology is used to desensitize the voiceprint, retaining only emotional / psychological features for scoring and approval, which can avoid identity information leakage and enhance privacy protection.
[0133] [Example 3]
[0134] Figure 6 2 is a schematic diagram of the architecture of a VSA-based approval system according to Example 3 of the present invention.
[0135] like Figure 6 As shown, the VSA-based approval system includes: an initial data acquisition module 310 , a pre-processing module 320 , a 15-dimensional sound feature information extraction module 330 , a model training module 340 , and a request approval module 350 .
[0136] The initial data acquisition module 310 is used to acquire the user's initial audio data. In some embodiments, the initial audio data includes but is not limited to customer audio acquired and recorded from one or more of the following scenarios: customer service hotline, video face-to-face interview, intelligent question-and-answer session, mobile app, etc.
[0137] The preprocessing module 320 is configured to preprocess the initial audio data to obtain preprocessed standard audio data.
[0138] In some embodiments, the preprocessing module 320 preprocesses the audio data, including: eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering silence data and / or background sound interference data in the first audio data to obtain second audio data; removing conversation segments irrelevant to approval in the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; converting the fifth audio data into sixth audio data with a preset sampling rate and a preset format; generating seventh audio data that eliminates volume differences among recording devices based on the sixth audio data; and compressing the seventh audio data to obtain the standard audio data.
[0139] The 15-dimensional sound feature information extraction module 330 is used to hierarchically extract 15-dimensional sound feature information based on the standard audio data. The 15-dimensional sound feature information includes: vowel lengthening rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, formant slope, speech-respiration coordination, kurtosis value, Lyapunov index, fractal dimension, intonation index, and emotional deviation.
[0140] The model training module 340 is used to train the initial model based on the 15-dimensional sound feature information to obtain a hierarchical deep learning model. In this embodiment, the model training module is constructed with reference to the idea of integrating BIM technology into the cloud platform. Thus, based on the concepts of BIM and cloud platform, seamless integration of acoustic data and financial business systems can be achieved. Furthermore, localized acoustic analysis can reduce the cloud computing load and avoid the high energy consumption problem compared to traditional 5G communication (sensor data is transmitted to the cloud server for processing in real time via the 5G network).
[0141] In some embodiments, the hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model.
[0142] In a further embodiment, the multimodal feature fusion network includes: a physical layer, a temporal layer, and an emotional layer. The physical layer performs a first feature processing on the 15-dimensional sound feature information to obtain spectral feature information; the temporal layer performs a second feature processing on the 15-dimensional sound feature information to obtain temporal feature information; and the emotional layer performs a third feature processing on the 15-dimensional sound feature information to obtain emotional feature information, and the emotional feature information is dynamically weighted.
[0143] In some embodiments, the risk scoring model is obtained by the following steps: obtaining target features in the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; constructing a deep factorization machine DeepFM model by combining a factorization machine FM model and a deep neural network DNN model; training the DeepFM model according to the target features to obtain the risk scoring model; and dynamically adjusting the decision boundary of the risk scoring model according to the business scenario.
[0144] The request approval module 350 is used to obtain the audio data to be approved based on the user's approval request and input the audio data to the layered deep learning model, which then outputs the approval result. In this embodiment, the request approval module is a real-time approval module based on the TensorRT engine. It has the characteristics of low data consumption and high response time, thus supporting wireless transmission within a range of 20 meters (referring to 2.4GHz technology) and millisecond-level decision-making, which is suitable for high-concurrency scenarios such as mobile apps.
[0145] In a further implementation manner, the VSA-based approval system of Example 3 of the present invention reserves a modular interface to support future fusion with biometrics (such as heart rate fluctuations, etc.) and environmental sensor data to form an extensible architecture.
[0146] The above-mentioned VSA-based approval system of Example 3 of the present invention obtains acoustic features through unstructured user audio data and establishes an unstructured data credit model to provide an evaluation basis for customers with no credit record. In addition, by analyzing real-time audio data to be approved, real-time dynamic risk assessment can be achieved.
[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by combining software with a hardware platform. Based on this understanding, all or part of the contribution of the technical solution of the present invention to the background art can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.
[0148] Correspondingly, embodiments of the present invention further provide a computer-readable storage medium having computer-readable instructions or a program stored thereon. When executed by a processor, the computer-readable instructions or program causes the computer to perform the following operations, which include the steps included in the VSA-based approval method described in any of the above embodiments, and are not further described here. The storage medium may include, for example, an optical disk, a hard disk, a floppy disk, a flash memory, a magnetic tape, and the like.
[0149] In addition, embodiments of the present invention further provide a computer device comprising a memory and a processor, wherein the memory is configured to store one or more computer-readable instructions or programs, wherein the one or more computer-readable instructions or programs, when executed by the processor, implement the VSA-based approval method described in any of the above embodiments. The computer device may be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or the like.
[0150] Embodiments of the present invention also provide a computer program product comprising a computer program, the computer program including program code for executing the VSA-based approval method shown in the flowchart. When the computer program product is executed in a computer system, the program code causes the computer system to implement the VSA-based approval method provided by the embodiments of the present disclosure.
[0151] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0152] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will appreciate that the technical solutions described in the above embodiments may be modified or some of the technical features may be replaced with equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A VSA-based approval method, characterized in that: The approval methods include: Get the user's initial audio data; Preprocessing the initial audio data to obtain preprocessed standard audio data; Extracting 15-dimensional sound feature information in layers according to the standard audio data, wherein the 15-dimensional sound feature information includes: vowel lengthening rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, formant slope, speech-respiration coordination, kurtosis value, Lyapunov index, fractal dimension, intonation index, and emotional deviation; The initial model is trained according to the 15-dimensional sound feature information to obtain a hierarchical deep learning model; The audio data to be approved is obtained according to the user's approval request and the audio data to be approved is input into the hierarchical deep learning model, and the hierarchical deep learning model outputs an approval result.
2. The approval method according to claim 1, wherein: Preprocessing the audio data includes: Eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering out silence data and / or background sound interference data in the first audio data to obtain second audio data; removing the dialogue segments irrelevant to the approval from the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; Converting the fifth audio data into sixth audio data having a preset sampling rate and a preset format; generating seventh audio data based on the sixth audio data, wherein the volume difference between the recording devices is eliminated; The seventh audio data is compressed to obtain the standard audio data.
3. The approval method according to claim 1, wherein: The hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model. The risk scoring model is obtained through the following steps: Obtaining target features from the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; Combine the factorization machine FM model and the deep neural network DNN model to build a deep factorization machine DeepFM model; Training the DeepFM model according to the target features to obtain the risk scoring model; and, The decision boundary of the risk scoring model is dynamically adjusted according to the business scenario.
4. The approval method according to claim 3, wherein: The multimodal feature fusion network includes: a physical layer, a temporal layer, and an emotional layer; wherein, performing first feature processing on the 15-dimensional sound feature information through the physical layer to obtain spectrum feature information; Performing second feature processing on the 15-dimensional sound feature information through the time sequence layer to obtain time sequence feature information; The emotional layer performs third feature processing on the 15-dimensional sound feature information to obtain emotional feature information and dynamically weights the emotional feature information.
5. A VSA-based approval system, characterized in that: The approval system includes: An initial data acquisition module is used to obtain the user's initial audio data; A preprocessing module, configured to preprocess the initial audio data to obtain preprocessed standard audio data; a 15-dimensional sound feature information extraction module, configured to extract 15-dimensional sound feature information in layers based on the standard audio data, wherein the 15-dimensional sound feature information includes: vowel lengthening rate, amplitude tremor information, fundamental frequency jitter rate, glottal closure rate, micro-vibration energy ratio, fundamental frequency standard deviation, speech rate gradient change information, unnatural pause information, formant slope, speech-respiration coordination, kurtosis value, Lyapunov exponent, fractal dimension, intonation index, and emotional deviation; A model training module, configured to train an initial model based on the 15-dimensional sound feature information to obtain a hierarchical deep learning model; The approval request module is used to obtain the audio data to be approved according to the user's approval request and input the audio data to be approved into the hierarchical deep learning model, and the hierarchical deep learning model outputs the approval result.
6. The approval system according to claim 5, wherein: The preprocessing module preprocesses the audio data, including: Eliminating steady-state noise and non-steady-state noise in the initial audio data to obtain first audio data; filtering out silence data and / or background sound interference data in the first audio data to obtain second audio data; removing the dialogue segments irrelevant to the approval from the second audio data to obtain third audio data; generating fourth audio data with anonymized voiceprint based on the third audio data; cutting the fourth audio data into fifth audio data of a preset length; Converting the fifth audio data into sixth audio data having a preset sampling rate and a preset format; generating seventh audio data based on the sixth audio data, wherein the volume difference between the recording devices is eliminated; The seventh audio data is compressed to obtain the standard audio data.
7. The approval system according to claim 5, wherein: The hierarchical deep learning model includes a multimodal feature fusion network and a risk scoring model. The risk scoring model is obtained through the following steps: Obtaining target features from the 15-dimensional sound feature information using a machine learning algorithm based on a gradient boosting framework; Combine the factorization machine FM model and the deep neural network DNN model to build a deep factorization machine DeepFM model; Training the DeepFM model according to the target features to obtain the risk scoring model; and, The decision boundary of the risk scoring model is dynamically adjusted according to the business scenario.
8. A computer-readable storage medium storing computer-readable instructions, characterized in that: The computer-readable instructions are executed by a processor to implement the approval method according to any one of claims 1 to 4.
9. A computer device comprising a memory and a processor, The memory stores computer-readable instructions, characterized in that: The processor executes the computer-readable instructions to implement the approval method according to any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the approval method according to any one of claims 1 to 4 is implemented.
Citation Information
Cited By
Intelligent dialogue method and system for accompanying old people based on user portraits
CN120998202A