An intelligent risk control voice monitoring method and system based on a multimodal algorithm

Through the intelligent risk control voice monitoring method based on multimodal algorithm, the problems of low risk control voice monitoring efficiency, large subjective deviation and monitoring lag in the existing technology are solved, real-time risk judgment and early warning push of voice are realized, and risk control effect and efficiency are improved.

CN119626260BActive Publication Date: 2025-05-30WESHARE TECH SERVICES (SHENZHEN) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163105.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-30
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing risk control voice monitoring technology has efficiency bottlenecks, subjective deviations and repayment monitoring lag, making it difficult to achieve real-time and effective monitoring, resulting in poor risk control and high cost.

Method used

An intelligent risk control speech monitoring method based on multimodal algorithm is adopted. By performing feature extraction and text recognition model training on the trained speech, combining word vector model and semantic correlation analysis, a risk discrimination model is generated, and speech risk status is judged in real time and early warning information is pushed.

Benefits of technology

It improves the intelligence level of risk control voice monitoring, enhances the accuracy and efficiency of voice recognition, reduces the misjudgment rate, realizes real-time monitoring of voice repayment without the platform, and improves the recall rate and risk control effect of risk judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119626260B_ABST
    Figure CN119626260B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent decision-making technology, and discloses an intelligent risk control voice monitoring method and system based on a multimodal algorithm, including: extracting features from training voices, constructing a pre-trained text recognition model for voice features, and training the pre-trained text recognition model using the voice features and training texts; dividing the digital information and text information in the training texts, and training the pre-trained word vector model using the preprocessed texts and training word vectors; performing vector fusion on the training word vectors, digital information, and rule text data, and performing semantic association analysis on the training texts and rule text data; generating a feature vector corresponding to the context vector and semantic association graph, constructing a pre-trained risk discrimination model for the feature vector, and training the pre-trained risk discrimination model using the feature vector and training categories. The present invention can improve the intelligence of risk control voice monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent risk control voice monitoring method and system based on a multimodal algorithm, and belongs to the field of intelligent decision-making technology. Background Art

[0002] Currently, in the current risk control business scenario, for voice monitoring, traditional methods mainly rely on manual sampling review, that is, samples are extracted from a large number of telephone recordings at a certain ratio, and professional quality inspection personnel rely on experience and preset rules to listen and judge whether there are any violations. For the voice monitoring of repayments, if the repayment occurs outside the official platform channel, the existing technology can often only detect it retrospectively through financial reconciliation or abnormal feedback, lacking real-time and effective monitoring means. The repayment voices within the official platform are mainly reviewed routinely according to established processes and data records. Therefore, the existing technology has the following defects: 1. Efficiency bottleneck: The efficiency of manual sampling of voices is extremely low, and it is difficult to achieve full coverage in the face of a large amount of data. Limited sampling samples are likely to cause omissions of violations, making it impossible to control risks in a timely manner. 2. Subjective deviation: Different quality inspection personnel have different evaluation criteria for the same voice due to differences in experience and understanding, resulting in poor stability and accuracy of evaluation results. Moreover, long-term manual operation is prone to fatigue and being influenced by emotions, further weakening the credibility of quality inspection. 3. Repayment monitoring lag: It is difficult to monitor the repayment voices outside the official platform in a timely manner, and can only passively wait for subsequent financial links to discover abnormalities. During this period, the risk of capital flow increases, and financial institutions are prone to potential losses, and the cost of retrospective investigation afterwards is high.

[0003] Therefore, the intelligentization of risk control voice monitoring in the existing technology is insufficient. Summary of the Invention

[0004] The present invention provides an intelligent risk control voice monitoring method and system based on a multimodal algorithm, and its main purpose is to improve the intelligentization of risk control voice monitoring.

[0005] To achieve the above purpose, an intelligent risk control voice monitoring method based on a multimodal algorithm provided by the present invention includes:

[0006] Obtain training voices and training texts corresponding to the training voices, extract features from the training voices to obtain voice features, construct a pre-trained text recognition model for the voice features, and use the voice features and the training texts to train the pre-trained text recognition model to obtain a trained text recognition model;

[0007] Obtain the training word vectors corresponding to the training text, divide the digital information and text information in the training text, preprocess the text information to obtain the preprocessed text, construct a pre-trained word vector model for the preprocessed text, and use the preprocessed text and the training word vectors to train the pre-trained word vector model to obtain a trained word vector model;

[0008] Obtain the preset rule text data, perform vector fusion on the training word vectors, the digital information, and the rule text data to obtain the context vector, and perform semantic association analysis on the training text and the rule text data to obtain the semantic association graph;

[0009] Generate the feature vectors corresponding to the context vector and the semantic association graph, obtain the training categories corresponding to the feature vectors, construct a pre-trained risk discrimination model for the feature vectors, and use the feature vectors and the training categories to train the pre-trained risk discrimination model to obtain a trained risk discrimination model;

[0010] Obtain the speech to be recognized at the current moment, and output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model, and the trained risk discrimination model, and use the risk category to determine whether the speech to be recognized is in a risk state;

[0011] When the speech to be recognized is in a risk state, generate a risk warning message corresponding to the speech to be recognized according to the risk category, the speech to be recognized, and the text content, and push the risk warning message to the corresponding risk control management personnel to obtain the risk control speech monitoring result of the speech to be recognized.

[0012] Optionally, the feature extraction of the training speech to obtain speech features includes:

[0013] Perform speech preprocessing on the training speech using the following formula to obtain preprocessed speech:

[0014]

[0015] where represents the preprocessed speech, represents the preprocessing function, t represents the time variable, and s(t) represents the training speech;

[0016] Convert the preprocessed speech to the time-frequency domain using the following formula to obtain the spectral feature matrix:

[0017]

[0018] where denotes the spectral feature matrix, f denotes the frequency variable, n denotes the time frame index, w() denotes the Hamming window function, When the Hamming window function is at it represents the center point of the time variable contained within the Hamming window function. denotes the imaginary part;

[0019] Take the spectral feature matrix as the speech feature.

[0020] Optionally, using the speech feature and the training text to train the pre-trained text recognition model to obtain a trained text recognition model, including:

[0021] Perform feature mapping on the speech feature to obtain an input speech sequence;

[0022] Generate an output text sequence corresponding to the training text;

[0023] In the pre-trained text recognition model, use the following formula to calculate the text recognition probability of the input speech sequence with respect to the output text sequence:

[0024]

[0025] where denotes the model output probability, denotes the input speech sequence, denotes the output text sequence, denotes the total number of words in the vocabulary, denotes the total duration, t denotes the time variable, denotes the probability calculated by the pre-trained text recognition model of outputting under the given ;

[0026] Based on the text recognition probability, use the following formula to calculate the model loss value corresponding to the pre-trained text recognition model:

[0027]

[0028] where denotes the model loss value, denotes the vocabulary size, denotes the th time variable and the th word's true probability, denotes the text recognition probability, denotes the total duration;

[0029] Training the pre-trained text recognition model with the model loss value to obtain a trained text recognition model.

[0030] Optionally, the dividing the digital information and text information in the training text includes:

[0031] Extracting the numerical values and text information in the training text;

[0032] Constructing a regular expression for the first numerical value in the numerical values;

[0033] Extracting the first information of the first numerical value through the regular expression;

[0034] Using a preset information extraction model to extract the second information of the second numerical value in the numerical values;

[0035] Taking the first information and the second information as digital information.

[0036] Optionally, the preprocessing the text information to obtain a preprocessed text includes:

[0037] Obtaining a preset stop word list and punctuation symbol library;

[0038] Removing the stop words and punctuation symbols in the text information according to the stop word list and the punctuation symbol library to obtain a preprocessed text.

[0039] Optionally, the training the pre-trained word vector model with the preprocessed text and the training word vectors to obtain a trained word vector model includes:

[0040] In the pre-trained word vector model, calculating the word vector probability of the preprocessed text with respect to the training word vectors using the following formula:

[0041]

[0042] where represents the word vector probability, represents the training word vector of the central word in the preprocessed text, represents the training word vector of the context word in the preprocessed text, represents the vocabulary set in the vocabulary, represents the serial number of the vocabulary in the preprocessed text, represents from the value of jumping to the serial numbers of other vocabularies;

[0043] Based on the word vector probability, calculating the model loss index corresponding to the pre-trained word vector model using the following formula:

[0044]

[0045] Among them, represents the model loss index, represents the total number of words in the vocabulary, represents the context window size, represents that given the center word the probability of the context word appearing, represents the serial number of the word in the preprocessed text, represents starting from the value for jumping to the serial number of other words;

[0046] The pre-trained word vector model is trained through the model loss index to obtain a trained word vector model.

[0047] Optionally, the vector fusion of the training word vectors, the digital information, and the rule text data to obtain a context vector includes:

[0048] Using a preset multi-layer neural network to convert the training word vectors and the digital information into hidden layer text vectors;

[0049] Converting the rule text data into hidden layer rule vectors;

[0050] Using the following formula to calculate the attention score between the hidden layer text vector and the hidden layer rule vector:

[0051]

[0052] Among them, represents the attention score, represents the hidden layer text vector, represents the hidden layer rule vector, represents the th hidden layer text vector in represents the th hidden layer rule vector in represents the weight parameter of represents the weight parameter of represents the bias;

[0053] Using the following formula to calculate the attention weight corresponding to the attention score:

[0054]

[0055] Among them, represents the attention weight, represents the attention score, represents the serial number of the hidden layer rule vector in represents the th hidden layer text vector and the th hidden layer rule vector, and the attention score therebetween, represents the number of hidden layer rule vectors in

[0056] Based on the attention weight, the training word vector, the digital information, and the rule text data are vectorially fused by the following formula to obtain a context vector:

[0057]

[0058] Among them, represents the context vector, represents the attention weight, represents the th hidden layer rule vector in represents the number of hidden layer rule vectors in

[0059] Optionally, the semantic association analysis of the training text and the rule text data to obtain a semantic association graph includes:

[0060] Extracting a first key semantics in the training text and a second key semantics in the rule text data;

[0061] Calculating the association strength between the first key semantics and the second key semantics by the following formula:

[0062]

[0063] Among them, represents the association strength, represents the first key semantics, represents the second key semantics;

[0064] Generating a semantic association graph of the training text and the rule text data through the association strength.

[0065] Optionally, the model training of the pre-trained risk discrimination model using the feature vector and the training category to obtain a trained risk discrimination model includes:

[0066] In the pre-trained risk discrimination model, the risk category corresponding to the feature vector is calculated using the following formula:

[0067]

[0068] where represents the risk category, g represents the classification function, are the parameters of the pre-trained risk discrimination model;

[0069] The risk loss value between the risk category and the training category is calculated using the following formula:

[0070]

[0071]

[0072]

[0073] where represents the risk loss value, represents the number of samples, represents the th true risk category label corresponding to the training category, represents the risk category probability output by the pre-trained risk discrimination model, represents the semantic association consistency loss function, represents the true association label between the first key semantics u and the second key semantics v, represents the association probability of the semantic association graph output by the pre-trained risk discrimination model, and represent the trade-off coefficients, represents the semantic association graph in the edge, represents the risk classification loss function;

[0074] The pre-trained risk discrimination model is trained through the risk loss value to obtain a trained risk discrimination model.

[0075] To solve the above problems, the present invention also provides an intelligent risk control voice monitoring system based on a multi-modal algorithm. The system includes:

[0076] A model training module, configured to obtain training speech and training text corresponding to the training speech, extract features from the training speech to obtain speech features, construct a pre-trained text recognition model for the speech features, and use the speech features and the training text to train the pre-trained text recognition model to obtain a trained text recognition model;

[0077] A word vector training module, which is used to obtain the training word vectors corresponding to the training text, divide the digital information and text information in the training text, perform text preprocessing on the text information to obtain preprocessed text, construct a pre-trained word vector model for the preprocessed text, and use the preprocessed text and the training word vectors to train the pre-trained word vector model to obtain a trained word vector model;

[0078] An association analysis module, which is used to obtain preset rule text data, perform vector fusion on the training word vectors, the digital information, and the rule text data to obtain context vectors, and perform semantic association analysis on the training text and the rule text data to obtain a semantic association graph;

[0079] A risk discrimination module, which is used to generate feature vectors corresponding to the context vectors and the semantic association graph, obtain the training categories corresponding to the feature vectors, construct a pre-trained risk discrimination model for the feature vectors, and use the feature vectors and the training categories to train the pre-trained risk discrimination model to obtain a trained risk discrimination model;

[0080] A status judgment module, which is used to obtain the speech to be recognized at the current moment, output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model, and the trained risk discrimination model, and use the risk category to judge whether the speech to be recognized is in a risk state;

[0081] An information push module, which is used to generate risk warning information corresponding to the speech to be recognized according to the risk category, the speech to be recognized, and the text content when the speech to be recognized is in a risk state, and push the risk warning information to the corresponding risk control management personnel to obtain the risk control speech monitoring result of the speech to be recognized.

[0082] Compared with the problems described in the background art, in the embodiments of the present invention, by extracting features from the training speech and using STFT for time-frequency domain conversion, the characteristic changes of the speech signal at different times and frequencies can be effectively captured. Compared with traditional feature extraction methods, it can better process information such as formants and tones in speech, providing a richer and more accurate feature representation for subsequent speech recognition, thereby improving the speech recognition accuracy, especially in complex environments (such as with background noise, large changes in speech speed, etc.), the speech recognition effect is significantly improved. Further, in the embodiments of the present invention, by using the preprocessed text and the training word vectors to train the pre-trained word vector model, in a multi-corpus joint training manner, in addition to the general corpus, a professional corpus in the financial field is introduced. The generated word vectors can better capture the semantic information of specific words in the financial field, making the text representation more in line with the requirements of the risk control business. In subsequent multi-modal fusion and risk assessment, it can more accurately understand the financial-related semantics in the speech text, reducing the misjudgment rate. Further, in the embodiments of the present invention, by performing semantic association analysis on the training text and the rule text data, by explicitly modeling the semantic relationship between the speech text and the rule text, deep semantic association information can be mined. For example, for some implicitly expressed intentions or non-standard descriptions of repayment information, the semantic association graph can assist the model in better understanding its potential risks and improving the recall rate of risk judgment to avoid risk omission. Therefore, the intelligent risk control speech monitoring method and system based on multi-modal algorithms provided by the embodiments of the present invention can improve the intelligence of risk control speech monitoring through an intelligent neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a schematic flow chart of an intelligent risk control speech monitoring method based on multi-modal algorithms provided by an embodiment of the present invention;

[0084] Figure 2 It is a schematic diagram of a module for implementing the intelligent risk control speech monitoring method based on multi-modal algorithms provided by an embodiment of the present invention.

[0085] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0087] An embodiment of the present application provides an intelligent risk control voice monitoring method based on a multi-modal algorithm. The execution subject of the intelligent risk control voice monitoring method based on the multi-modal algorithm includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the intelligent risk control voice monitoring method based on the multi-modal algorithm can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0088] Embodiment 1:

[0089] Referring to Figure 1 As shown, it is a flowchart of an intelligent risk control voice monitoring method based on a multi-modal algorithm provided by an embodiment of the present invention. In this embodiment, the intelligent risk control voice monitoring method based on the multi-modal algorithm includes:

[0090] S1. Obtain a training voice and a training text corresponding to the training voice, extract features from the training voice to obtain voice features, construct a pre-trained text recognition model for the voice features, and use the voice features and the training text to train the pre-trained text recognition model to obtain a trained text recognition model.

[0091] In an embodiment of the present invention, the training voice refers to a training sample for subsequent training of a neural network model, including voice information obtained from various call lines, communication links that may involve informal repayment voice interactions, etc., and the training text refers to the text data corresponding to the training voice.

[0092] Furthermore, in an embodiment of the present invention, by extracting features from the training voice, using STFT for time-frequency domain conversion can effectively capture the characteristic changes of the voice signal at different times and frequencies. Compared with traditional feature extraction methods, it can better process information such as formants and pitches in the voice, provide a richer and more accurate feature representation for subsequent voice recognition, thereby improving the voice recognition accuracy, especially the voice recognition effect in complex environments (such as having background noise, large speed changes, etc.) is significantly improved.

[0093] In an embodiment of the present invention, the extracting features from the training voice to obtain voice features includes: performing voice preprocessing on the training voice using the following formula to obtain preprocessed voice:

[0094]

[0095] Wherein, represents the preprocessed voice, represents the preprocessing function, t represents the time variable, and s(t) represents the training voice;

[0096] Convert the preprocessed speech into the time-frequency domain using the following formula to obtain a spectral feature matrix:

[0097]

[0098] where, represents the spectral feature matrix, f represents the frequency variable, n represents the time frame index, w() represents the Hann window function, represents the center point of the time variable contained inside the Hann window function when the Hann window function is at , represents the imaginary part;

[0099] Use the spectral feature matrix as the speech feature.

[0100] where, includes a series of operations such as a noise filtering algorithm and a volume normalization algorithm based on spectral analysis. Further, the pre-trained text recognition model refers to a speech-to-text model with an untrained LSTM-CTC architecture.

[0101] In one embodiment of the present invention, training the pre-trained text recognition model using the speech feature and the training text to obtain a trained text recognition model includes: performing feature mapping on the speech feature to obtain an input speech sequence; generating an output text sequence corresponding to the training text; in the pre-trained text recognition model, calculating the text recognition probability of the input speech sequence with respect to the output text sequence using the following formula:

[0102]

[0103] where, represents the model output probability, represents the input speech sequence, represents the output text sequence, represents the total number of words in the vocabulary, represents the total duration, t represents the time variable, represents the probability calculated by the pre-trained text recognition model of outputting given ;

[0104] Based on the text recognition probability, calculate the model loss value corresponding to the pre-trained text recognition model using the following formula:

[0105]

[0106] where, represents the model loss value, represents the vocabulary size, represents the -th time variable for the true probability of the -th word, represents the text recognition probability, represents the total duration;

[0107] The pre-trained text recognition model is trained using the model loss value to obtain a trained text recognition model.

[0108] It should be noted that if the -th word output by the pre-trained text recognition model is correct, then is 1, otherwise is 0.

[0109] Optionally, the process of feature mapping the speech features refers to the process of feature mapping and encoding the speech features, which can be implemented by an encoder. Further, the process of training the pre-trained text recognition model using the model loss value to obtain a trained text recognition model refers to the process of optimizing the weights and biases of the model using a neural network parameter optimization algorithm to reduce the model loss value.

[0110] S2. Obtain the training word vectors corresponding to the training text, divide the digital information and text information in the training text, perform text preprocessing on the text information to obtain a preprocessed text, construct a pre-trained word vector model for the preprocessed text, and use the preprocessed text and the training word vectors to train the pre-trained word vector model to obtain a trained word vector model.

[0111] In the embodiment of the present invention, the training word vectors refer to the word vector training samples corresponding to the training text.

[0112] In an embodiment of the present invention, the dividing the digital information and text information in the training text includes: extracting the digital values and text information in the training text; constructing a regular expression for the first value in the digital values; extracting the first information of the first value through the regular expression; extracting the second information of the second value in the digital values using a preset information extraction model; and using the first information and the second information as digital information.

[0113] Among them, the first value refers to the repayment amount, the second value refers to the repayment account number, and the information extraction model refers to a model for extracting repayment account number information, such as a natural language extraction NLP model.

[0114] Optionally, the process of extracting the first information of the first value through the regular expression refers to: for repayment amount extraction, using a regular expression Match the digital string as the possible repayment amount.

[0115] In an embodiment of the present invention, the text preprocessing of the text information to obtain the preprocessed text includes: obtaining a preset stop word list and punctuation library; removing the stop words and punctuation from the text information according to the stop word list and the punctuation library to obtain the preprocessed text.

[0116] Among them, the stop word list and the punctuation library refer to the pre-constructed database of stop words and punctuation that need to be cleaned.

[0117] Further, in the embodiment of the present invention, the pre-trained word vector model is trained by using the preprocessed text and the training word vector, and in the multi-corpus joint training method, in addition to the general corpus, a financial domain professional corpus is introduced, so that the generated word vector can better capture the semantic information of specific words in the financial domain, making the text representation more in line with the risk control business requirements. In subsequent multi-modal fusion and risk assessment, the financial-related semantics in the speech text can be understood more accurately, reducing the misjudgment rate.

[0118] Among them, the pre-trained word vector model refers to the untrained word vector model Word2Vec.

[0119] In an embodiment of the present invention, the pre-trained word vector model is trained by using the preprocessed text and the training word vector to obtain the trained word vector model, including: in the pre-trained word vector model, the word vector probability of the preprocessed text with respect to the training word vector is calculated by using the following formula:

[0120]

[0121] Among them, represents the word vector probability, represents the training word vector of the central word in the preprocessed text, represents the training word vector of the context word in the preprocessed text, represents the vocabulary set in the vocabulary, represents the serial number of the vocabulary in the preprocessed text, represents from the value of jumping to the serial number of other vocabulary;

[0122] Based on the word vector probability, the model loss index corresponding to the pre-trained word vector model is calculated by using the following formula:

[0123]

[0124] Among them, represents the model loss index, represents the total number of words in the vocabulary, represents the context window size, represents the probability of the context word appearing given the center word ; represents the serial number of the word in the preprocessed text, represents the value for jumping from to the serial number of other words;

[0125] The pre-trained word vector model is trained through the model loss index to obtain a trained word vector model.

[0126] Optionally, the process of training the pre-trained word vector model through the model loss index to obtain a trained word vector model refers to the process of optimizing the weights and biases of the model using a neural network parameter optimization algorithm, thereby reducing the model loss value.

[0127] S3. Obtain preset rule text data, perform vector fusion on the trained word vectors, the digital information, and the rule text data to obtain context vectors, and perform semantic association analysis on the training text and the rule text data to obtain a semantic association graph.

[0128] In the embodiments of the present invention, the rule text data refers to pre-constructed rule text data, such as data in databases such as a keyword library, an official repayment process text library, etc.

[0129] In an embodiment of the present invention, the performing vector fusion on the trained word vectors, the digital information, and the rule text data to obtain context vectors includes: converting the trained word vectors and the digital information into hidden layer text vectors using a preset multi-layer neural network; converting the rule text data into hidden layer rule vectors; calculating the attention score between the hidden layer text vectors and the hidden layer rule vectors using the following formula:

[0130]

[0131] Among them, represents the attention score, represents the hidden layer text vector, represents the hidden layer rule vector, represents the th hidden layer text vector in represents the The rule vectors of the hidden layer denote the weight parameters of denote the weight parameters of denote the bias;

[0132] Calculate the attention weights corresponding to the attention scores using the following formula:

[0133]

[0134] where denotes the attention weights, denotes the attention scores, denotes the serial number of the rule vector of the hidden layer in denotes the th hidden layer text vector and the th hidden layer rule vector denotes the number of rule vectors of the hidden layer in

[0135] Based on the attention weights, use the following formula to perform vector fusion on the training word vectors, the digital information, and the rule text data to obtain the context vector:

[0136]

[0137] where denotes the context vector, denotes the attention weights, denotes the th rule vector of the hidden layer in denotes the number of rule vectors of the hidden layer in

[0138] wherein, the multi-layer neural network refers to a neural network including an input layer, a hidden layer, and an output layer, such as a BP neural network.

[0139] Optionally, the process of converting the rule text data into rule vectors of the hidden layer can be implemented by a multi-layer neural network.

[0140] Furthermore, the embodiment of the present invention performs semantic association analysis on the training text and the rule text data, so as to explicitly model the semantic relationship between the speech text and the rule text, and can mine deep semantic association information. For example, for some implicitly expressed intentions or non-standard descriptions of repayment information, the semantic association graph can assist the model to better understand its potential risks, improve the recall rate of risk judgment, and avoid risk omission.

[0141] In one embodiment of the present invention, the semantic association analysis of the training text and the rule text data to obtain a semantic association graph includes: extracting the first key semantics in the training text and the second key semantics in the rule text data; calculating the association strength between the first key semantics and the second key semantics using the following formula:

[0142]

[0143] wherein, represents the association strength, represents the first key semantics, represents the second key semantics;

[0144] Generate a semantic association graph of the training text and the rule text data through the association strength.

[0145] Among them, the first key semantics and the second key semantics refer to the key information in the training text and the rule text data, such as words, phrases or concepts, etc.

[0146] Optionally, the process of generating the semantic association graph of the training text and the rule text data through the association strength refers to connecting two key semantics through the magnitude of the association strength, using the key semantics as graph nodes, and the edge weight of the edge between two nodes being the association strength.

[0147] S4. Generate a feature vector corresponding to the context vector and the semantic association graph, obtain the training category corresponding to the feature vector, construct a pre-trained risk discrimination model for the feature vector, and use the feature vector and the training category to train the pre-trained risk discrimination model to obtain a trained risk discrimination model.

[0148] In an embodiment of the present invention, the feature vector refers to the result of splicing the context vector and the semantic association graph. Further, the training category refers to the training sample of the risk category corresponding to the feature vector.

[0149] In one embodiment of the present invention, the use of the feature vector and the training category to train the pre-trained risk discrimination model to obtain a trained risk discrimination model includes: in the pre-trained risk discrimination model, calculating the risk type corresponding to the feature vector using the following formula:

[0150]

[0151] wherein, represents the risk type, g represents the classification function, are the parameters of the pre-trained risk discrimination model;

[0152] Calculate the risk loss value between the risk category and the training category using the following formula:

[0153]

[0154]

[0155]

[0156] where represents the risk loss value, represents the number of samples, represents the th true risk category label corresponding to the training category, represents the risk category probability output by the pre-trained risk discrimination model, represents the semantic association consistency loss function, represents the true association label between the first key semantics u and the second key semantics v, represents the association probability of the semantic association graph output by the pre-trained risk discrimination model, and represents the trade-off coefficient, represents the semantic association graph edge in, represents the risk classification loss function;

[0157] Perform model training on the pre-trained risk discrimination model through the risk loss value to obtain a trained risk discrimination model.

[0158] Optionally, the process of performing model training on the pre-trained risk discrimination model through the risk loss value to obtain a trained risk discrimination model refers to the process of optimizing the weights and biases of the model using a neural network parameter optimization algorithm to reduce the model loss value.

[0159] where and are trade-off coefficients used to balance the importance of the two tasks, represents the true association label between semantic units u and v (1 if the association label exists, otherwise 0).

[0160] S5. Obtain the speech to be recognized at the current moment, output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model, and the trained risk discrimination model, and use the risk category to determine whether the speech to be recognized is in a risk state.

[0161] Optionally, the process of outputting the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model, and the trained risk discrimination model is the process of using the trained models to recognize the text and risk category corresponding to the speech, which is consistent with the process of recognizing the training text and training category corresponding to the foregoing recognition training text, and will not be elaborated further herein.

[0162] Among them, the risk category includes multiple categories, among which there are non-risk categories, such as safety, and risk categories. If there is a risk category, it means that the speech to be recognized is in a risk state. For example, the risk category includes risks, risks of defaulting on repayment off the platform, no risks, etc.

[0163] S6. When the speech to be recognized is in a risk state, generate a risk warning message corresponding to the speech to be recognized according to the risk category, the speech to be recognized, and the text content, and push the risk warning message to the corresponding risk control management personnel to obtain the risk voice monitoring result of the speech to be recognized.

[0164] Optionally, the process of generating a risk warning message corresponding to the speech to be recognized according to the risk category, the speech to be recognized, and the text content means that when the risk intelligent discrimination model determines that there is a risk, a warning message including the risk category, the speech to be recognized, and the text content is generated. Further, the process of pushing the risk warning message to the corresponding risk control management personnel means that the generated warning message will be promptly pushed to the risk control management personnel through various methods (such as text messages, emails, internal message pushing in the risk control system, etc.) so that they can take corresponding measures quickly.

[0165] Compared with the problems described in the background art, in the embodiments of the present invention, by extracting features from the training speech and using STFT for time-frequency domain conversion, the characteristic changes of the speech signal at different times and frequencies can be effectively captured. Compared with traditional feature extraction methods, it can better process information such as formants and pitches in speech, providing a richer and more accurate feature representation for subsequent speech recognition, thereby improving the speech recognition accuracy. Especially in complex environments (such as with background noise and large changes in speech speed), the speech recognition effect is significantly improved. Further, in the embodiments of the present invention, by using the preprocessed text and the training word vectors to train the pre-trained word vector model, in a multi-corpus joint training manner, in addition to the general corpus, a professional corpus in the financial field is introduced. In this way, the generated word vectors can better capture the semantic information of specific financial terms, making the text representation more in line with the requirements of the risk control business. In subsequent multi-modal fusion and risk assessment, the financial-related semantics in the speech text can be understood more accurately, reducing the misjudgment rate. Further, in the embodiments of the present invention, by performing semantic association analysis on the training text and the rule text data, by explicitly modeling the semantic relationship between the speech text and the rule text, deep semantic association information can be mined. For example, for some implicitly expressed intentions or non-standard descriptions of repayment information, the semantic association graph can assist the model in better understanding its potential risks and improving the recall rate of risk judgment to avoid risk omission. Therefore, the intelligent risk control speech monitoring method and system based on the multi-modal algorithm provided by the embodiments of the present invention can improve the intelligence of risk control speech monitoring through an intelligent neural network.

[0166] Embodiment 2:

[0167] As Figure 2 shown, it is a functional module diagram of an intelligent risk control speech monitoring system based on a multi-modal algorithm of the present invention.

[0168] The intelligent risk control speech monitoring system 200 based on a multi-modal algorithm of the present invention can be installed in an electronic device. According to the functions implemented, the intelligent risk control speech monitoring system based on a multi-modal algorithm can include a model training module 201, a word vector training module 202, an association analysis module 203, a risk discrimination module 204, a status judgment module 205, and an information push module 206. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0169] In the embodiments of the present invention, the functions of each module / unit are as follows:

[0170] The model training module 201 is configured to obtain training speech and the training text corresponding to the training speech, extract features from the training speech to obtain speech features, construct a pre-trained text recognition model for the speech features, and use the speech features and the training text to train the pre-trained text recognition model to obtain a trained text recognition model;

[0171] The word vector training module 202 is configured to obtain the training word vectors corresponding to the training text, divide the digital information and text information in the training text, preprocess the text information to obtain preprocessed text, construct a pre-trained word vector model for the preprocessed text, and use the preprocessed text and the training word vectors to train the pre-trained word vector model to obtain a trained word vector model;

[0172] The association analysis module 203 is configured to obtain preset rule text data, perform vector fusion on the training word vectors, the digital information, and the rule text data to obtain context vectors, and perform semantic association analysis on the training text and the rule text data to obtain a semantic association graph;

[0173] The risk discrimination module 204 is configured to generate feature vectors corresponding to the context vectors and the semantic association graph, obtain the training categories corresponding to the feature vectors, construct a pre-trained risk discrimination model for the feature vectors, and use the feature vectors and the training categories to train the pre-trained risk discrimination model to obtain a trained risk discrimination model;

[0174] The status judgment module 205 is configured to obtain the speech to be recognized at the current moment, output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model, and the trained risk discrimination model, and use the risk category to judge whether the speech to be recognized is in a risk state;

[0175] The information push module 206 is configured to generate risk warning information corresponding to the speech to be recognized according to the risk category, the speech to be recognized, and the text content when the speech to be recognized is in a risk state, and push the risk warning information to the corresponding risk control management personnel to obtain the risk control voice monitoring result of the speech to be recognized.

[0176] Specifically, each module in the intelligent risk control voice monitoring system 200 based on the multi-modal algorithm in the embodiments of the present invention adopts the same technical means as those in the Figure 1 intelligent risk control voice monitoring method based on the multi-modal algorithm described above, and can produce the same technical effects, which will not be elaborated here.

[0177] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An intelligent risk control voice monitoring method based on a multimodal algorithm, characterized in that: The method comprises: Acquire a training speech and a training text corresponding to the training speech, perform feature extraction on the training speech to obtain speech features, construct a pre-trained text recognition model of the speech features, and perform model training on the pre-trained text recognition model using the speech features and the training text to obtain a trained text recognition model; Obtaining a training word vector corresponding to the training text, dividing the digital information and the text information in the training text, performing text preprocessing on the text information to obtain a preprocessed text, constructing a pretrained word vector model for the preprocessed text, and performing model training on the pretrained word vector model using the preprocessed text and the training word vector to obtain a trained word vector model; Obtaining preset regular text data, performing vector fusion on the training word vector, the digital information and the regular text data to obtain a context vector, performing semantic association analysis on the training text and the regular text data to obtain a semantic association graph; Generate a feature vector corresponding to the context vector and the semantic association graph, obtain a training category corresponding to the feature vector, construct a pre-trained risk discrimination model of the feature vector, and perform model training on the pre-trained risk discrimination model using the feature vector and the training category to obtain a trained risk discrimination model; Obtain the speech to be recognized at the current moment, output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model and the trained risk discrimination model, and use the risk category to determine whether the speech to be recognized is in a risky state; When the voice to be recognized is in a risky state, risk warning information corresponding to the voice to be recognized is generated according to the risk category, the voice to be recognized and the text content, and the risk warning information is pushed to the corresponding risk control management personnel to obtain the risk control voice monitoring result of the voice to be recognized.

2. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The step of extracting features from the training speech to obtain speech features includes: The training speech is preprocessed using the following formula to obtain the preprocessed speech: in, represents the preprocessed speech, represents the preprocessing function, t represents the time variable, and s(t) represents the training speech; The pre-processed speech is converted into the time-frequency domain using the following formula to obtain a spectrum feature matrix: in, represents the spectrum feature matrix, f represents the frequency variable, n represents the time frame index, and w() represents the Heining window function. Indicates that the Hening window function is When , the center point of the time variable contained in the Haining window function, represents the imaginary part; The spectrum feature matrix is ​​used as speech feature.

3. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The method of performing model training on the pre-trained text recognition model by using the speech features and the training text to obtain a trained text recognition model includes: Performing feature mapping on the speech features to obtain an input speech sequence; Generate an output text sequence corresponding to the training text; In the pre-trained text recognition model, the text recognition probability of the input speech sequence with respect to the output text sequence is calculated using the following formula: in, represents the model output probability, represents the input speech sequence, Represents the output text sequence, represents the total number of words in the vocabulary, represents the total duration, t represents the time variable, Indicates that the pre-trained text recognition model is calculated at a given Lower Output probability; Based on the text recognition probability, the model loss value corresponding to the pre-trained text recognition model is calculated using the following formula: in, Represents the model loss value, represents the vocabulary size, Indicates The time variable The true probability of a word, represents the text recognition probability, Indicates the total duration; The pre-trained text recognition model is trained using the model loss value to obtain a trained text recognition model.

4. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The dividing of the digital information and the text information in the training text includes: Extracting numerical values ​​and text information from the training text; Constructing a regular expression for a first value among the digital values; Extracting first information of the first value by using the regular expression; Extracting second information of a second value from the digital value using a preset information extraction model; The first information and the second information are used as digital information.

5. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The step of performing text preprocessing on the text information to obtain a preprocessed text includes: Get the preset stop word list and punctuation library; According to the stop word list and the punctuation mark library, the stop words and punctuation marks in the text information are removed to obtain a preprocessed text.

6. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The method of performing model training on the pre-trained word vector model using the pre-processed text and the training word vector to obtain a trained word vector model includes: In the pre-trained word vector model, the word vector probability of the pre-processed text with respect to the training word vector is calculated using the following formula: in, represents the word vector probability, Represents the central word in the preprocessed text The training word vectors of Represents the context words in the preprocessed text The training word vectors of represents the set of words in the vocabulary, Represents the ordinal number of the vocabulary in the preprocessed text, Indicates from Jump to the value of other vocabulary numbers; Based on the word vector probability, the model loss index corresponding to the pre-trained word vector model is calculated using the following formula: in, represents the model loss index, represents the total number of words in the vocabulary, represents the context window size, Indicates that given the central word The context word appears in the case of The probability of Represents the ordinal number of the vocabulary in the preprocessed text, Indicates from Jump to the value of other vocabulary numbers; The pre-trained word vector model is trained using the model loss index to obtain a trained word vector model.

7. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The step of performing vector fusion on the training word vector, the digital information and the regular text data to obtain a context vector includes: Using a preset multi-layer neural network to convert the training word vector and the digital information into a hidden layer text vector; Converting the rule text data into a hidden layer rule vector; The attention score between the hidden layer text vector and the hidden layer rule vector is calculated using the following formula: in, represents the attention score, represents the hidden layer text vector, represents the hidden layer rule vector, express Middle hidden layer text vectors, express Middle hidden layer rule vector, express The weight parameter, express The weight parameter, Indicates bias; The attention weight corresponding to the attention score is calculated using the following formula: in, represents the attention weight, represents the attention score, express The serial number of the hidden layer rule vector in , Indicates The hidden layer text vector is The attention score between the hidden layer rule vectors, express The number of hidden layer rule vectors in ; Based on the attention weight, the training word vector, the digital information and the regular text data are vector-fused using the following formula to obtain a context vector: in, represents the context vector, represents the attention weight, express Middle hidden layer rule vector, express The number of hidden layer rule vectors in .

8. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The performing semantic association analysis on the training text and the rule text data to obtain a semantic association graph includes: Extracting a first key semantics in the training text and a second key semantics in the rule text data; The association strength between the first key semantics and the second key semantics is calculated using the following formula: in, represents the strength of association, Indicates the first key semantics, Indicates the second key semantics; A semantic association graph between the training text and the regular text data is generated according to the association strength.

9. The intelligent risk control voice monitoring method based on multimodal algorithm according to claim 1, characterized in that: The using the feature vector and the training category to perform model training on the pre-trained risk discrimination model to obtain a trained risk discrimination model includes: In the pre-trained risk discrimination model, the risk type corresponding to the feature vector is calculated using the following formula: in, represents the risk type, g represents the classification function, are the parameters of the pre-trained risk discrimination model; The risk loss value between the risk type and the training category is calculated using the following formula: in, represents the risk loss value, represents the number of samples, Indicates The true risk category labels corresponding to the training categories, represents the risk category probability output by the pre-trained risk discrimination model, represents the semantic association consistency loss function, represents the true association label between the first key semantics u and the second key semantics v, represents the association probability of the semantic association graph output by the pre-trained risk discrimination model, and represents the trade-off coefficient, Representing semantic association graph Middle side, represents the risk classification loss function; The pre-trained risk discrimination model is trained using the risk loss value to obtain a trained risk discrimination model.

10. An intelligent risk control voice monitoring system based on a multimodal algorithm, characterized in that: The system comprises: A model training module is used to obtain a training speech and a training text corresponding to the training speech, perform feature extraction on the training speech to obtain speech features, construct a pre-trained text recognition model of the speech features, and perform model training on the pre-trained text recognition model using the speech features and the training text to obtain a trained text recognition model; A word vector training module is used to obtain the training word vector corresponding to the training text, divide the digital information and text information in the training text, perform text preprocessing on the text information to obtain a preprocessed text, construct a pretrained word vector model of the preprocessed text, and perform model training on the pretrained word vector model using the preprocessed text and the training word vector to obtain a trained word vector model; An association analysis module is used to obtain preset regular text data, perform vector fusion on the training word vector, the digital information and the regular text data to obtain a context vector, and perform semantic association analysis on the training text and the regular text data to obtain a semantic association graph; A risk discrimination module is used to generate a feature vector corresponding to the context vector and the semantic association graph, obtain a training category corresponding to the feature vector, construct a pre-trained risk discrimination model of the feature vector, and perform model training on the pre-trained risk discrimination model using the feature vector and the training category to obtain a trained risk discrimination model; A state judgment module is used to obtain the speech to be recognized at the current moment, output the text content and risk category corresponding to the speech to be recognized through the trained text recognition model, the trained word vector model and the trained risk discrimination model, and use the risk category to judge whether the speech to be recognized is in a risky state; The information push module is used to generate risk warning information corresponding to the voice to be recognized according to the risk category, the voice to be recognized and the text content when the voice to be recognized is in a risk state, and push the risk warning information to the corresponding risk control management personnel to obtain the risk control voice monitoring result of the voice to be recognized.

Citation Information

Patent Citations

  • Voice quality inspection method, device, computer equipment and storage medium

    CA3156142A1

  • Text classification method and device and storage medium

    CN114860930A