Fraud countering method based on key business voiceprint

By building a hierarchical analysis framework and AI model for collaborative analysis, combined with voiceprint feature extraction and speech synthesis detection, we have achieved accurate identification and real-time blocking of fraudulent calls, solved the problem of high misjudgment rate in existing technologies, and improved the ability to prevent telecommunications network fraud.

CN120808791APending Publication Date: 2025-10-17EASTERN COMM +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511299462.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and accurately identify fraudulent calls, resulting in a high misjudgment rate, which affects user experience and governance effectiveness.

Method used

Build a fraud prevention method based on key business voiceprints. By constructing a hierarchical analysis framework, using AI discriminant and generative models for collaborative analysis, and combining voiceprint feature extraction and speech synthesis detection, we can achieve accurate identification and real-time blocking of fraudulent behavior.

Benefits of technology

It significantly improves the accuracy of identifying fraudulent calls, reduces the misjudgment rate, enhances the ability to prevent telecommunications network fraud, and ensures user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808791A_ABST
    Figure CN120808791A_ABST
Patent Text Reader

Abstract

The invention discloses a fraud countering method based on key business voiceprints. The method comprises the following steps: constructing a normal call scene and fraud-related call scene classification sample library; the AI discriminant model and the generative model are cooperatively researched and judged; extracting voiceprint features and constructing a key business classification library; overlapping voiceprint comparison to carry out comprehensive study and judgment; carrying out comprehensive study and judgment by superimposing speech synthesis detection; and dealing with fraud risks. According to the invention, processing measures such as short message early warning, number shutdown, real-time interception and the like can be carried out on the call which is confirmed to be involved in fraud through detection in the telephone scene, and otherwise, the call is considered to be normal and is played. According to the method, on the basis of not influencing normal calls of other networks in the current network, the current network does not need to be greatly transformed, and accurate identification and timely and effective treatment measures of the suspected fraud phone can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of communication network security, and particularly relates to a fraud countermeasure method based on key service voiceprints. BACKGROUND

[0002] To do a good job in precise management, it is urgent to solve the problem of how to efficiently and accurately identify fraudulent calls, and to use advanced technologies such as AI large models to expand the intelligent identification application of fraud-related risk numbers, while reducing the number of cases and reducing the perception of bad service caused by governance, and overall improving the ability to prevent fraud in the telecommunications network. SUMMARY

[0003] In view of the defects in the prior art, the purpose of the present application is to provide a technical scheme of a fraud countermeasure method based on key service voiceprints, which realizes accurate identification and real-time blocking of fraudulent behavior in voice calls through the construction of a hierarchical analysis framework.

[0004] The fraud countermeasure method based on key service voiceprints comprises the following steps: Constructing a normal call scenario and a fraud-related call scenario classification sample library; Using sample library data for model training, using AI discriminant model and generative model for collaborative research and judgment; wherein, for calls with high fraud risk, directly implement SMS warning, number closure or real-time interception; for calls with low fraud risk, continue the voiceprint comparison and speech synthesis detection process; Voiceprint feature extraction and construction of key service classification library, covering the feature library of normal users and fraudsters with high-risk behavior; Using the key service voiceprint library to compare the voiceprints of medium and low risk calls, and combining the model collaborative research and judgment results for comprehensive research and judgment; Performing speech synthesis detection on medium and low risk call voice, and combining the model collaborative research and judgment results for comprehensive research and judgment; According to the secondary research and judgment of medium and low risk call voice, calls confirmed by detection as fraud-related in the phone scenario are directly implemented with SMS warning, number closure or real-time interception, and normal calls are passed through.

[0005] The fraud countermeasure method based on key service voiceprints comprises the following steps: Based on anti-fraud APP case data, and combined with the privacy number voice quality inspection platform, extract the fraud communication features of the case audio, and classify through entity labeling; normal communication scenario classification is based on short calls, high dispersion features, covering legal call behaviors such as express delivery, takeout, service follow-up.

[0006] The fraud countermeasures method based on key business voice prints has the characteristics that the step of collaborative judgment of the AI discriminant model and the generative model includes: training the black and white samples by using the discriminant model, constructing a key business classification expert model, and realizing fraud behavior recognition through semantic understanding and intent analysis; combining the generative model, the output of the discriminant model is judged again to improve the classification accuracy and generalization ability.

[0007] The fraud countermeasures method based on key business voice prints has the characteristics that the step of voice print feature extraction and construction of key business classification library includes: based on the high-risk audio data output by the AI inference engine, the voice print recognition technology is used to extract features, and a double feature library containing the voice print feature set of fraudsters and the voice print feature set of normal users of high-risk behavior patterns is established.

[0008] The fraud countermeasures method based on key business voice prints has the characteristics that the step of comprehensive judgment by voice print comparison includes: for the call voice, the voice print comparison is performed by using the key business classification voice print library, the comparison result is high risk, and the short message warning, number closing or real-time interception are directly implemented, the comparison result is medium and low risk, and the confidence score of collaborative judgment is superimposed, and after weighting, the same as the set threshold value enters the next disposal link.

[0009] The fraud countermeasures method based on key business voice prints has the characteristics that the step of comprehensive judgment by voice print comparison includes: for the call voice, the voice print comparison is performed by using the key business classification voice print library, the comparison result is high risk, and the short message warning, number closing or real-time interception are directly implemented, the comparison result is medium and low risk, and the confidence score of collaborative judgment is superimposed, and after weighting, the same as the set threshold value enters the next disposal link.

[0010] The fraud countermeasures method based on key business voice prints has the characteristics that the voice print comparison adopts the PLDA algorithm, and the similarity threshold τ is set as: , wherein, is the average score of impostors, is the standard deviation, is an adjustable coefficient.

[0011] The fraud countermeasures method based on key business voice prints has the characteristics that the voice print comparison adopts the PLDA algorithm, and the similarity threshold τ is set as: Phase discontinuity detection: , wherein, is the inter-frame phase jump amount, Spectral slope anomaly detection: , wherein β is the spectral slope of the current speech segment, is the average spectral slope of normal speech, is the standard deviation of the statistical distribution of the spectral slope.

[0012] The fraud prevention method based on key business voiceprints is characterized in that the comprehensive analysis and judgment adopts a weighted decision-making mechanism, and the calculation formula is: ,in, Indicates the final decision result, usually a binary output. I represents the indicator function, which outputs 1 when the condition is met, otherwise it outputs 0. Indicates the The original score of each judgment module, represents the adaptive weight, Indicates the decision threshold, which is used to determine whether the comprehensive score meets the trigger condition.

[0013] The fraud prevention method based on key business voiceprints is characterized in that the update mechanism of the dual feature library includes: Online update: The feature center is updated using the sliding window averaging method. The formula is: ,in, represents the updated feature center value, For the forgetting factor, represents the historical feature center value before updating, The feature vector representing the new input; Offline update: The feature extraction model is fully retrained every month.

[0014] The technical advantages of the present invention are embodied in: 1) Through the dual-feature design of the voiceprint database, it maintains the recognition accuracy of known fraud patterns while effectively detecting new fraud behaviors; 2) Adopting a generative model-enhanced analysis and judgment mechanism to significantly reduce the misjudgment rate; 3) Dynamic weight allocation algorithm ensures real-time processing performance.

[0015] This invention is fully compatible with existing network architectures and enables rapid deployment through standardized interfaces, providing a new technical approach to combating telecom fraud. Specifically, by collecting and analyzing the voiceprint characteristics of legitimate individuals who exhibit consistent fraudulent behavior, the system can more accurately identify genuine fraudsters and avoid misjudgments due to behavioral similarities. This innovative design allows the system to maintain a high recall rate while keeping the false alarm rate extremely low. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a schematic diagram of the collaborative analysis process of AI discriminant models and generative models; Figure 2is a comprehensive judgment process schematic diagram based on the collaborative judgment result and the voiceprint comparison result; Figure 3 is a comprehensive judgment process schematic diagram based on the collaborative judgment result and the voice synthesis detection result; Figure 4 is a flowchart of fraud risk disposal; Figure 5 is a flowchart of the implementation steps of the present application. DETAILED DESCRIPTION

[0017] The present application will be further described in detail below in combination with the drawings and specific embodiments.

[0018] The application discloses a fraud countermeasure method based on key business voiceprints, which first adopts an adversarial sample enhancement technology and a semi-supervised learning method to construct a multi-dimensional voiceprint feature sample library containing normal call scenarios (such as express delivery, take-out and other short-time high-frequency calls) and fraud-related scenarios; secondly, an innovative discriminative model based on an improved ECAPA-TDNN (Emphasized Channel Attention Propagation and Aggregation in Time Delay Neural Network) architecture is used to realize voiceprint feature extraction, and a generative model based on a DeepSeek (a language model developed by DeepSeek company) large model is used to realize a collaborative analysis mechanism of knowledge-enhanced semantic understanding through an attention weight fusion algorithm to realize dynamic decision-making; then, multi-dimensional verification is performed through x-vector (a voiceprint embedding vector based on a deep neural network, which converts variable-length speech into a fixed-dimensional feature representation through a time pooling layer) voiceprint feature comparison and WaveGAN (a speech synthesis model of a generative adversarial network, composed of a generator and a discriminator)-based voice synthesis detection; finally, intelligent graded disposal is implemented on the confirmed fraud-related calls. The application realizes a technical closed loop of "feature extraction-model collaboration-dynamic disposal", maintains the compatibility of the existing network, simultaneously reduces the false positive rate, and significantly improves the anti-fraud engine capability.

[0019] The fraud countermeasure method of the application comprises the following steps: Constructing a normal call scenario and fraud-related call scenario classification sample library; Using the sample library data to train the model, and using an AI discriminative model and a generative model to collaboratively judge; wherein, for the calls judged as high fraud risk, fraud risk disposal (SMS warning, number suspension or real-time interception) is directly implemented; for the calls judged as low fraud risk, voiceprint comparison and voice synthesis detection processes are continued; The voiceprint feature extraction and the construction of the key business classification library cover the feature library of normal users and fraudsters (for example, high frequency and high dispersion) of high-risk behaviors. The voiceprint comparison of the key business voiceprint library is used for the voiceprint comparison of the medium and low risk call voice, and the comprehensive judgment is made combined with the model cooperative judgment result. The voiceprint synthesis detection is performed on the medium and low risk call voice, and the comprehensive judgment is made combined with the model cooperative judgment result. According to the secondary judgment of the medium and low risk call voice, the call in the telephone scene confirmed as involving fraud through detection is directly implemented fraud risk disposal (SMS warning, number closing or real-time interception), and the normal call is passed.

[0020] The steps of constructing the normal call scene and the call scene involving fraud sample library include: based on the anti-fraud APP case data, and combined with the privacy number voice quality detection platform, the fraud communication features of the case audio are extracted, and classified through entity annotation; the normal communication scene classification is based on short call and high dispersion features, covering express delivery, takeout, service return visit and other legal call behaviors.

[0021] The steps of the above-mentioned AI discriminant model and generative model cooperative judgment include: using discriminant model to train black and white samples, constructing key business classification expert model, realizing fraud behavior recognition through semantic understanding and intent analysis; combined with the generative model, the output of the discriminant model is judged twice to improve the classification accuracy and generalization ability.

[0022] The steps of the above-mentioned voiceprint feature extraction and the construction of the key business classification library include: based on the high-risk audio data output by the AI inference engine, the voiceprint recognition technology is used to extract features, and a double feature library containing the voiceprint feature set of fraudsters and the voiceprint feature set of normal users of high-risk behavior mode is established.

[0023] The steps of the above-mentioned voiceprint comparison comprehensive judgment include: for the call voice, the voiceprint comparison is performed using the key business classification voiceprint library, the comparison result is high risk, and the SMS warning, number closing or real-time interception is directly implemented, the comparison result is medium and low risk, and the confidence score of the cooperative judgment is superimposed, and after weighting, the same as the set threshold value enters the next disposal link.

[0024] The steps of the above-mentioned voiceprint synthesis detection comprehensive judgment include: for the call voice, the voice synthesis detection is used to identify whether it is a synthetic or false generated voice, and the confidence score of the cooperative judgment is superimposed if it is a synthetic or false generated voice, and after weighting, the same as the set threshold value is directly implemented SMS warning, number closing or real-time interception, and the one that does not meet the threshold value is passed.

[0025] The voiceprint comparison uses PLDA algorithm, and the similarity threshold τ is set as: ,in, is the mean score of the impostors, is the standard deviation, is an adjustable coefficient.

[0026] The above-mentioned speech synthesis detection includes: Phase discontinuity detection: ,in, is the inter-frame phase jump variable, Spectral slope anomaly detection: ,in, is the spectrum slope of the current speech segment, is the average spectral slope of normal speech, is the standard deviation of the statistical distribution of the spectral slope.

[0027] The above comprehensive assessment adopts a weighted decision-making mechanism, and the calculation formula is: ,in, Indicates the final decision result, usually a binary output. I represents the indicator function, which outputs 1 when the condition is met, otherwise it outputs 0. Indicates the The original score of each judgment module, represents the adaptive weight, Indicates the decision threshold, which is used to determine whether the comprehensive score meets the trigger condition.

[0028] The update mechanism of the dual signature library includes: Online update: The feature center is updated using the sliding window averaging method. The formula is: ,in, represents the updated feature center value, For the forgetting factor, represents the historical feature center value before updating, The feature vector representing the new input; Offline update: The feature extraction model is fully retrained every month.

[0029] The innovative features of the present invention also lie in: 1. Voiceprint feature extraction process The improved ECAPA-TDNN architecture is used for voiceprint feature extraction. The network structure includes: Input layer: processes speech signals with a sampling rate of 16kHz, a frame length of 25ms, and a frame shift of 10ms; Feature extraction layer: SE-Res2Block module stacking, the formula is expressed as: , where x is the input, For output, sigmoid function, ReLU activation function, and fully connected layer weights, denotes a pooling layer.

[0030] 2. Dual feature library design The application innovatively constructs a key business voiceprint feature library containing dual feature categories , which is defined as .

[0031] : Fraudster voiceprint feature set. The voiceprint samples of this dataset mainly come from verified telecommunication network fraud cases, such as voice recordings of cases provided by regulatory units, voice recordings of fraud calls reported by users and confirmed by artificial or intelligent auditing, etc. This ensures a high confidence level of the "fraud-related" label for this part of the data.

[0032] : Normal user voiceprint feature set of high-risk behavior patterns. This dataset aims to distinguish between real fraudsters and legitimate users (such as delivery personnel, customer service personnel, etc.) who only have similar behavior patterns (such as short high-frequency calls, high dispersion calls, etc.) to fraudsters. Its voiceprint samples come from: audio recordings that have been screened for potential high-risk behavior characteristics such as high-frequency calls, high call dispersion, short call duration, etc., but have been verified as legitimate calls. Verification methods can include but are not limited to: comparing known legitimate business number white lists, combining user business type judgments, or confirming the legitimacy of call content through customer service visits, user feedback, etc.

[0033] Feature extraction: For the above-mentioned audio data (including and source audio), advanced voiceprint feature extraction techniques are used. For example, the x-vector extraction framework based on deep learning can be used. This framework usually contains several layers of time-delay neural network or convolutional neural network layers to capture temporal or local acoustic features, a statistical pooling layer (such as mean and standard deviation pooling) to aggregate frame-level features into segment-level features, and several fully connected layers to generate fixed-dimensional embedding vectors. The extraction process may involve speech activity detection to remove silent segments, and setting appropriate frame length (such as 25ms), frame shift (such as 10ms), and number of mel filter banks (such as 64 or 80). The extracted x-vector is the voiceprint feature representation of the corresponding speaker.

[0034] Library construction and indexing: The extracted x-vector feature vectors and their source category labels ( or ) and associated information (such as speaker identification, case number, etc.) are stored in the database. To improve the efficiency of subsequent comparison, an index structure can be established for the feature vector, for example, using a tree-based index or a hash-based index.

[0035] Library update: The voiceprint library supports a dynamic update mechanism. As new fraud cases are identified and new high-risk behavior normal user samples are accumulated, voiceprint features can be continuously extracted and added to the library, while obsolete or no longer representative voiceprint data can be removed as needed.

[0036] By constructing a dual feature library containing , the present application can better distinguish between normal calls that are similar but essentially different, while accurately identifying known fraud patterns, thereby significantly reducing the false positive rate for legitimate users in the voiceprint comparison process and improving the overall accuracy and user experience of the anti-fraud engine.

[0037] 3. Dual-scene voice feature library construction Based on the adversarial sample generation technique (by adding carefully designed small perturbations to the original normal or fraud-related samples, generating adversarial samples that the model cannot correctly classify, and adding these samples to the training set to enhance the model's robustness to small changes or potential attacks) and semi-supervised contrastive learning [using a small number of labeled fraud / normal samples, while using a large number of unlabeled call recordings. Through a contrastive learning framework, a representation function is learned, such that voice segments from the same call scenario (even if unlabeled) or the same speaker (if known) are closer in feature space, while segments from different scenarios or speakers are farther apart. This allows the model to learn effective voice / scenario representations from unlabeled data, improving performance in cases where labeled data is insufficient], using fraud case intelligence and privacy number voice quality inspection platform data, a multi-dimensional feature sample library is established containing normal call scenarios (such as express delivery, food delivery, etc. high dispersion short-time calls) and fraud-related scenarios, i.e.: in addition to the speech pattern, it can also include call duration distribution features, call time regularity features (such as abnormal calls at night), call graph features associated with the caller number (such as a large number of unrelated users called in a short period of time), acoustic features of the voice signal itself (such as prosody, abnormal speech rate), etc.

[0038] Optimizing feature space distribution using deep metric learning:

[0039] : Contrastive loss function value, used to measure the model's ability to distinguish the similarity of sample pairs, : Binary label, indicating the category of the sample pair, ​​ : neural network for input sample , extracted feature vector, : margin, preset threshold hyperparameter.

[0040] 4. hybrid intelligent analysis engine The discriminative model is constructed based on the improved ECAPA-TDNN architecture, and the generative model DeepSeek knowledge enhancement module cooperates with the judgment mechanism: Discriminative model architecture: 1) Language recognition module: A CNN-BiLSTM hybrid architecture is adopted, and the loss function is a cross-entropy with class weight: wherein, : weighted cross-entropy loss value, used to measure the difference between the model prediction probability distribution and the real label distribution; : summation symbol, cumulative calculation for all classes or samples; : weight coefficient, used to adjust the loss contribution of different classes or samples; : real label, usually in one-hot encoding form; : natural logarithm of the model prediction probability; 2) Speech transcription module: Based on the Conformer-CTC model, the attention mechanism calculation formula is: wherein, denotes the query matrix, denotes the key-value matrix, denotes the value matrix, denotes the matrix transpose, denotes the scaling factor; 3) intent understanding module: Transformer architecture, 12-layer encoder; multi-head attention calculation formula: wherein, Q denotes the query matrix, K denotes the key-value matrix, V denotes the value matrix, denotes the output of the th attention head, denotes the concatenation of the outputs of multiple heads along the feature dimension, denotes the output transformation matrix, used to map the concatenated features back to the original dimension; The output fraud intent probability distribution discriminative model includes: Language recognition module: CNN-BiLSTM hybrid architecture Speech transcription module: Conformer-CTC model Intention understanding module: Transformer architecture with multi-head attention mechanism.

[0041] Generative model coordination mechanism: 1) Semantic completion process: For medium and low-risk text output by the discriminative model, use the DeepSeek model for completion: where represents the joint probability of generating complete output sequence y under the condition of input text x, represents the output sequence, consisting of T discrete units, represents the continuous multiplication symbol, represents the input text, and serves as the context basis for the generative model, represents the anti-fraud knowledge context.

[0042] 2) Dynamic weight fusion algorithm: Model coordination score calculation: where represents the final coordination score of the model, obtained by fusing the output scores of different models, is the sigmoid function, is the learnable parameter matrix, updated through online learning, represents the detection model score, represents the generative model score, represents the bias.

[0043] Weight self-adaptive adjustment strategy: where represents the adaptive weight of the th module, reflecting its importance in the current task, represents the exponential function, used to nonlinearly amplify the accuracy differences of each module, is the recent accuracy of each module, represents the temperature coefficient, , represents the accuracy of the th or th module, usually calculated through a sliding window or exponential decay method, is the summation symbol.

[0044] 5. Multi-factor security verification Double verification for medium and low-risk calls: 1) Voiceprint comparison: PLDA (Probabilistic Linear Discriminant Analysis: a statistical method for voiceprint similarity calculation) algorithm is used to calculate the similarity score. Given two samples We usually calculate the log likelihood ratio (LLR) of the same speaker:

[0045] : Two samples belong to the same speaker, : Two samples belong to different speakers;

[0046] 2) Speech synthesis detection: based on phase spectrum analysis and prosodic feature anomaly detection Comprehensive analysis and judgment adopts a weighted decision mechanism: Where Decision represents the final decision result, usually binary output, I represents the indicator function, output 1 when the condition is met, otherwise output 0, represents the original score of the th analysis module, is the adaptive weight, represents the decision threshold value, used to determine whether the comprehensive score meets the trigger condition.

[0047] Key business voiceprint library construction: Innovatively build a voiceprint library containing double features:

[0048] Among them is the voiceprint feature of the fraudster, is the voiceprint feature of the normal user with high similarity to the behavior pattern of the fraudster. By building a double feature library, the generalization ability of the system to new fraud methods is significantly improved.

[0049] 6. Feature library update mechanism: Online update: use sliding window average method to update feature center: , where represents the updated feature center value, is the forgetting factor, representing the historical feature center value before updating, represents the newly input feature vector.

[0050] Offline update: retrain the feature extraction model every month.

[0051] The purpose of the present application is to accurately identify the fraud behavior existing in the voice call, and to implement timely and effective interception, shutdown, early warning and other disposal measures for the suspected fraud call behavior and the calling and called numbers. Embodiment 1

[0052] Figure 1 is an AI discriminant model and generative model collaborative judgment process schematic diagram, including the following steps: Step 110, obtaining the voice file or real-time voice stream of the call, first performing language recognition, including Chinese, foreign language and Mandarin and dialect recognition classification, input: 16kHz single channel audio, features: 80-dimensional Mel filter bank features, frame length 25ms, frame shift 10ms, model: CNN (3 layers)-BiLSTM (2 layers) hybrid structure, output: probability distribution of 13 languages and dialects; Step 120, converting the voice into text using ASR transcription technology, acoustic model: Conformer (attention dimension 256, FFN dimension 1024), language model: 4-gram hybrid model based on Transformer, decoding: beam search (beam size=10); Step 130, refining the specific fraud scene using the NLP intent understanding model after the transcribed text; Step 140, using the deepseek large model to complete the semantics for the AI discriminant model intent understanding judgment as medium and low risk and the transcribed text that is not smooth, for example: input segment: "you are suspected of... immediately transfer the funds into... account", the completion result is: "you are suspected of... please immediately transfer the funds into the safe account"; Step 150, using the deepseek large model for secondary intent understanding analysis for the AI discriminant model intent understanding judgment as medium and low risk text, identifying whether the voice content is related to fraud. Embodiment 2

[0053] Figure 2 is a comprehensive judgment process schematic diagram based on the collaborative judgment result and the voiceprint comparison result, including the following steps: Step 210, obtaining the voice file or real-time voice stream of the call; Step 220, extracting the feature vector related to the voiceprint in the voice; Step 230, comparing with the key business voiceprint library in the system one by one, and taking the highest confidence score; Step 240, the voiceprint comparison judgment as medium and low risk needs to combine the results of the previous collaborative judgment to make comprehensive judgment, that is, the total sum of the weighted proportion of the confidence scores of the two is used to judge whether it is related to fraud. Embodiment 3

[0054] Figure 3 is a comprehensive judgment process schematic diagram based on the collaborative judgment result and the speech synthesis detection result, comprising the following steps: Step 310, obtaining a speech file or a real-time speech stream of the call; Step 320, extracting a feature vector related to synthetic false generation in the speech; Step 330, using a speech synthesis false generation detection model to detect the speech, and identifying whether the speech has traces of synthesis or false generation; Step 340, the speech synthesis detection judgment is high risk, and the result of the previous collaborative judgment needs to be combined for comprehensive judgment, that is, whether it involves fraud is judged according to the sum of the confidence scores of the two after proportional weighting. Embodiment 4

[0055] Figure 4 is a process schematic diagram of fraud risk disposal, comprising the following judgment logic: The AI discriminant model judges that it is high risk to involve fraud, and directly enters the risk disposal process; The AI discriminant model judges that it is low risk, and uses deepseek for semantic completion and intent understanding secondary judgment, that is, collaborative judgment; The collaborative judgment is high risk to involve fraud, and directly enters the risk disposal process; The collaborative judgment is low risk, and voiceprint comparison and speech synthesis detection can be performed at the same time; wherein: The voiceprint comparison is high risk, and directly enters the risk disposal process; the voiceprint comparison is low risk, and the comprehensive judgment is combined with the result of the collaborative judgment; The speech synthesis detection is low risk, and the call is directly processed; the speech synthesis detection is high risk, and the comprehensive judgment is combined with the result of the collaborative judgment.

[0056] The present application realizes accurate judgment of voice call fraud by multi-modal data fusion, voiceprint comparison, speech synthesis detection and other AI model collaborative analysis, and uses the world knowledge base and Chinese semantic understanding advantage of deepseek to realize accurate judgment of voice call fraud, solves the defects of slow response, low reasoning efficiency and single anti-fraud strategy of traditional call behavior analysis model in the face of the rapid evolution of fraud methods, and the poor user experience after disposal caused by low recognition accuracy.

[0057] The method of the present application can realize accurate identification and timely and effective disposal measures for suspected fraud calls without affecting the normal calls of other networks in the existing network, without the need for large modification of the existing network.

[0058] It should be noted that the above only describes the preferred embodiments of the present application, and does not limit the patent protection scope of the present application.

Claims

1. A fraud prevention method based on key business voiceprints, characterized by The following steps are involved: Build a classification sample library of normal call scenarios and fraudulent call scenarios; Model training is conducted using sample library data, with AI discriminative and generative models working together to analyze and judge calls. Calls identified as high fraud risk are directly subject to SMS alerts, number closure, or real-time interception. Voiceprint comparison and speech synthesis detection processes continue for calls identified as medium to low fraud risk. Extract voiceprint features and build a key business classification database, covering the characteristics of normal users and fraudsters with high-risk behaviors; Use the key business voiceprint database to compare the voiceprints of low- and medium-risk calls, and conduct comprehensive analysis based on the results of the model's collaborative analysis; Perform speech synthesis testing on low- and medium-risk calls, and conduct comprehensive analysis based on the results of collaborative model analysis. Based on the secondary analysis of low- and medium-risk call voices, calls that are detected and confirmed to be fraudulent in telephone scenarios will be directly issued SMS warnings, number shutdowns or real-time interceptions, while normal calls will be allowed to go through.

2. A fraud prevention method based on key business voiceprints as claimed in claim 1, characterized in that The steps of constructing a classification sample library of normal call scenarios and fraud-related call scenarios include: based on anti-fraud APP case data and combined with the privacy number voice quality inspection platform, extracting the fraudulent communication features of the recordings involved in the case, and classifying them through entity labeling; normal communication scenario classification is based on short calls and high discrete features, covering legal call behaviors such as express delivery, takeout, and service return visits.

3. A fraud prevention method based on key business voiceprints as claimed in claim 1, characterized in that The steps of collaborative analysis of the AI ​​discriminant model and the generative model include: using the discriminant model to train black and white samples, building a key business classification expert model, and identifying fraudulent behavior through semantic understanding and intent analysis; combining the generative large model to conduct secondary analysis on the output of the discriminant model to improve classification accuracy and generalization ability.

4. The fraud prevention method based on key business voiceprint according to claim 1 is characterized in that The steps of extracting voiceprint features and building a key business classification library include: extracting features based on high-risk audio data output by the AI ​​inference engine using voiceprint recognition technology, and establishing a dual feature library containing a voiceprint feature set of fraudsters and a voiceprint feature set of normal users with high-risk behavior patterns.

5. The fraud prevention method based on key business voiceprint according to claim 1 is characterized in that The steps of comprehensive assessment through voiceprint comparison include: performing voiceprint comparison on call voice using a voiceprint library classified by key business categories; if the comparison result is high risk, SMS warning, number shutdown or real-time interception will be directly implemented; if the comparison result is medium or low risk, the confidence score of collaborative assessment will be added, and after weighting, if the set threshold is met, the next step of disposal will also be entered.

6. A fraud prevention method based on key business voiceprints as claimed in claim 1, characterized in that The steps of comprehensive analysis and judgment through the speech synthesis detection include: using speech synthesis detection to identify whether the call voice is synthesized or falsely generated speech, and if it is determined to be synthesized or falsely generated speech, the confidence score of the collaborative analysis is superimposed, and after weighting, if it meets the set threshold, SMS warning, number shutdown or real-time interception is directly implemented, and if it does not meet the threshold, it is released.

7. The fraud prevention method based on key business voiceprint according to claim 1 is characterized in that The voiceprint comparison adopts the PLDA algorithm, and the similarity threshold τ is set as: ,in, is the mean score of the impostors, is the standard deviation, is an adjustable coefficient.

8. The fraud prevention method based on key business voiceprint according to claim 1 is characterized in that The speech synthesis detection includes: Phase discontinuity detection: , where Δφ(t) is the inter-frame phase jump variable, Spectral slope anomaly detection: , where β is the spectrum slope of the current speech segment, is the average spectral slope of normal speech, is the standard deviation of the statistical distribution of the spectral slope.

9. The fraud prevention method based on key business voiceprint according to claim 1, characterized in that The comprehensive assessment adopts a weighted decision-making mechanism, and the calculation formula is: , where Decision represents the final decision result, usually a binary output, and I represents the indicator function, which outputs 1 when the condition is met, otherwise it outputs 0. Indicates the The original score of each judgment module, represents the adaptive weight, Indicates the decision threshold, which is used to determine whether the comprehensive score meets the trigger condition.

10. A fraud prevention method based on key business voiceprints as claimed in claim 1, characterized in that The update mechanism of the dual feature library includes: Online update: The feature center is updated using the sliding window averaging method. The formula is: ,in, represents the updated feature center value, For the forgetting factor, represents the historical feature center value before updating, The feature vector representing the new input; Offline update: The feature extraction model is fully retrained every month.

Citation Information

Patent Citations

  • Voiceprint identification method and device

    CN106847292A

  • Voiceprint retrieval method and device and electronic equipment

    CN112447178A

  • Model-based verbal skill recommendation method, device, computer equipment and storage medium

    CN113688221A

  • Anti-fraud system and method

    CN114641003A

  • Context semantic understanding-based customer service method, equipment and program product

    CN119151552A