A Fraud Prevention Communication Method Based on Voice Changer Recognition
By combining multi-dimensional speech feature analysis and deep learning model training with dynamic recognition thresholds and multi-level feedback control, the problem of recognition difficulties in traditional voiceprint recognition under voice-changing technology has been solved, and efficient anti-fraud communication has been achieved in complex environments.
Patent Information
- Application Number
- CN202411246360.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-09-06
AI Technical Summary
Modern voice-changing technology can produce realistic voices, making it difficult for traditional voiceprint recognition technology to accurately identify voice-changing behavior in fraud prevention. Furthermore, fluctuations in communication network quality increase the difficulty of identification.
By acquiring multi-dimensional voice feature data, training and mapping deep learning models, establishing dynamic recognition threshold models, and combining user historical call data and call environment noise data, a multi-level feedback control system is constructed to perform personalized recognition and call quality optimization. Multi-criteria decision fusion analysis is then conducted to improve the accuracy and robustness of recognition.
It significantly improves the accuracy and robustness of voice-changing recognition, reduces false positives and false negatives, enhances the system's adaptability and user experience, and ensures effective fraud prevention protection in complex environments.
Smart Images

Figure CN119132308B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voiceprint recognition technology, and in particular to an anti-fraud communication method based on voice-changing recognition. Background Technology
[0002] Voice alteration is a technology designed to identify and detect spoofing or imitation in voice. This technology is particularly important in the fields of speech recognition and security because advanced voice alteration techniques can mimic the voice characteristics of a specific individual, rendering traditional voiceprint recognition systems unreliable. With the rapid development of communication technologies and artificial intelligence, voice communication has become increasingly prevalent in daily life and business activities. However, the increasing prevalence of advanced voice alteration technologies makes it increasingly difficult to rely solely on traditional voiceprint recognition techniques for identity verification and fraud prevention.
[0003] Traditional anti-fraud communication methods often suffer from the following problems: Modern voice-changing technology is becoming increasingly sophisticated, capable of producing highly realistic voices, making it more difficult to identify fraud based solely on voice characteristics. For example, some advanced voice-changing software can mimic a specific person's voice, including details such as timbre, tone, and accent. Fluctuations in communication network quality can cause voice distortion, increasing the difficulty of identification. For instance, in areas with poor signal, degraded call quality can affect the accuracy of voiceprint analysis. Setting appropriate recognition thresholds to balance false alarms (misidentifying normal voices as voice changes) and false negatives (failing to recognize voice changes) is a challenge; for example, the system might misidentify voice changes caused by a cold as voice changes, or fail to recognize highly realistic voice-changing techniques. Summary of the Invention
[0004] Therefore, it is necessary for the present invention to provide a fraud prevention communication method based on voice changer recognition to solve at least one of the above-mentioned technical problems.
[0005] To achieve the above objectives, a fraud prevention communication method based on voice changer recognition includes the following steps:
[0006] Step S1: Obtain multi-dimensional voice feature data of users, including timbre, frequency features, language rhythm, emotional features and pause patterns; predict the possibility of voice change based on multi-dimensional voice feature data to obtain voiceprint feature prediction data; train a deep learning model based on voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0007] Step S2: Simulate and analyze the voice-changing recognition behavior under different call states based on the voiceprint feature map data to obtain voice-changing recognition simulation data; establish a dynamic recognition threshold model based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data.
[0008] Step S3: Adjust the control parameters in real time according to the preset target recognition accuracy and dynamic recognition threshold model to obtain control parameter adjustment data; construct a multi-level feedback control system based on the control parameter adjustment data to obtain an adaptive recognition model;
[0009] Step S4: Obtain user's historical call data, perform personalized analysis on the user's historical call data, and extract features to obtain user profile data; predict abnormal behavior based on user profile data and adaptive recognition model, and use machine learning technology to optimize the acquisition of recognition thresholds to obtain personalized recognition scheme data.
[0010] Step S5: Obtain call environment noise data and network quality data, and construct a call quality model based on the call environment noise data and network quality data; optimize the call quality model across scenarios, and adjust the communication parameters in real time to obtain call quality optimization data;
[0011] Step S6: Perform multi-criteria decision fusion analysis on the personalized identification scheme data and call quality optimization data to obtain the final anti-fraud identification scheme data.
[0012] This invention, through comprehensive analysis of these features, can more accurately capture the speaker's unique voiceprint, reducing the possibility of misjudgment and omission. These features include timbre, frequency characteristics, speech rhythm, emotional characteristics, and pause patterns. By analyzing multi-dimensional speech feature data, it predicts whether the voice has undergone voice-changing processing, obtaining voiceprint feature prediction data; this step helps to detect potential voice-changing behavior early. Using the voiceprint feature prediction data to train a deep learning model, the model can learn to recognize various complex speech features, improving the accuracy and robustness of recognition. By mapping the voiceprint state under different call conditions, voiceprint feature map data is obtained, helping the system better understand and analyze the performance of speech features in different environments. Simulation analysis of voice-changing recognition behavior under different call conditions yields voice-changing recognition simulation data. This helps the system accurately identify voice-changing behavior under different environments and conditions. Based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data, a dynamic recognition threshold model is established, allowing for flexible adjustment of the recognition threshold, improving the sensitivity and accuracy of recognition. Based on the preset target recognition accuracy and the dynamic recognition threshold model, control parameters are adjusted in real time to ensure the system is always in the optimal recognition state. A multi-level feedback control system is constructed, continuously optimizing the system based on control parameters to improve its adaptability and recognition accuracy. Historical call data is acquired from users, undergoing personalized analysis and feature extraction to create user profiles, enabling a better understanding of users' voice characteristics and behavioral patterns. Based on user profile data and an adaptive recognition model, abnormal behavior is predicted, and machine learning techniques are used to obtain the optimal recognition threshold, resulting in personalized recognition scheme data. Call environment noise data and network quality data are acquired to construct a call quality model, accurately reflecting the impact of call quality on recognition. The call quality model is optimized across scenarios, and communication parameters are adjusted in real time to ensure the system's recognition performance in different environments. Multi-criteria decision fusion analysis is performed on personalized recognition scheme data and call quality optimization data to comprehensively evaluate various factors, improving recognition accuracy and reliability; comprehensive recognition decisions are made considering multiple factors to improve overall system performance; and multi-criteria decision-making reduces false alarms and false negatives, improving system usability and user experience.
[0013] Preferably, step S1 includes the following steps:
[0014] Step S11: Obtain the user's historical voice data and perform a short-time Fourier transform on the historical voice data to obtain the voice time-frequency data;
[0015] Step S12: Extract Mel frequency cepstral coefficients from the speech time-frequency data to obtain timbre feature data; extract frequency features based on fundamental frequency contours from the speech time-frequency data to obtain frequency feature data;
[0016] Step S13: Use a speech activity detection algorithm to identify speech and non-speech segments in the speech time-frequency data to obtain language rhythm feature data;
[0017] Step S14: Analyze the duration of speech segments and the pause patterns of non-speech segments based on the language rhythm feature data to obtain pause pattern data.
[0018] Step S15: Analyze the pitch variation, energy distribution, and speech rate of the speech time-frequency data, and extract emotional features to obtain emotional feature data; perform harmonic-noise ratio analysis on the speech time-frequency data, and evaluate the harmonic structure of the speech to obtain harmonic structure data;
[0019] Step S16: Merge the timbre feature data, frequency feature data, language rhythm feature data, pause pattern data, and emotional feature data into multi-dimensional speech feature data based on the harmonic structure data;
[0020] Step S17: Use support vector machines to classify multi-dimensional speech feature data and score the probability of voice change to obtain voiceprint feature prediction data.
[0021] Step S18: Train a deep learning model based on the voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0022] This invention utilizes technologies such as STFT, MFCC, fundamental frequency profile, and VAD to extract speech features from multiple dimensions, providing high-resolution and high-precision speech analysis. Combining timbre, frequency, speech rhythm, pause patterns, emotional features, and harmonic structure, it provides comprehensive speech feature data, improving recognition accuracy and robustness. By dynamically adjusting the recognition threshold and implementing a multi-level feedback control system, it enhances the system's adaptability under different call conditions, ensuring recognition stability and reliability. Based on users' historical speech data and personalized analysis, it provides personalized recognition schemes, improving the system's ability to recognize individual users and reducing false positives and false negatives. Deep learning models are used for training and mapping to capture complex speech features, improving the system's ability to detect voice-changing behavior and its robustness. A multi-criteria decision model comprehensively evaluates various factors for comprehensive recognition decisions, improving the overall system performance and reducing false positives and false negatives. Overall, this anti-fraud communication method based on voice-changing recognition significantly improves the security and accuracy of voice communication through multi-dimensional and multi-level comprehensive analysis and optimization. It effectively copes with advanced voice-changing technologies and complex communication environments, providing more reliable anti-fraud protection.
[0023] Preferably, step S18 includes the following steps:
[0024] Step S181: Add noise and time scaling to the voiceprint feature prediction data, and perform normalization processing to obtain the voiceprint feature training dataset.
[0025] Step S182: Train the preset deep learning model based on the voiceprint feature training dataset, and adjust the parameters of the model using the cross-validation method to obtain the voiceprint feature model. The deep learning model includes a multi-layer convolutional neural network structure for extracting local features, a long short-term memory network layer for capturing temporal dependencies, an attention mechanism layer for highlighting important features, and a fully connected layer.
[0026] Step S183: Simulate different call conditions to obtain conditional simulation data; apply the voiceprint feature model to the conditional simulation data to obtain conditional response data;
[0027] Step S184: Map the voiceprint status and features under different call conditions based on the conditional response data and conditional simulation data to obtain voiceprint feature map data.
[0028] This invention enhances the diversity of training data by adding noise and time scaling, improving the model's generalization ability and enabling it to better adapt to various real-world call environments. Normalization reduces data bias, improving the stability and efficiency of model training and ensuring features are compared on the same scale. Extracting local speech features improves the model's ability to recognize subtle speech features; capturing temporal dependencies in speech signals improves the model's understanding of dynamic speech changes; highlighting important features enhances the model's focus on key speech features, improving recognition accuracy. Simulating different call conditions comprehensively tests the model's performance in various real-world environments, improving its adaptability; obtaining model response data under different conditions helps evaluate and optimize model performance, ensuring its accuracy and robustness in various environments. Mapping voiceprint states and features under different call conditions generates voiceprint feature map data, providing intuitive visualization information to help understand and analyze speech features; integrating feature data under various conditions improves the model's ability to recognize complex speech environments, ensuring stability and reliability in practical applications. Optimizing model parameters improves the model's robustness and generalization ability, ensuring stable performance under different call conditions. By adding noise and time scaling, the diversity of training data is enhanced, improving the model's adaptability to various real-world call environments. Multi-layer convolutional neural networks and long short-term memory (LSTM) layers are used to extract local features and capture temporal dependencies, improving the model's ability to recognize speech features. An attention mechanism layer highlights important features, enhancing the model's focus on key speech features and improving recognition accuracy. Cross-validation optimizes model parameters, improving robustness and generalization ability, ensuring stable performance under different call conditions. Simulating various call conditions comprehensively tests and optimizes model performance, improving its adaptability and accuracy in various environments. Generating voiceprint feature map data provides intuitive visualization information, aiding in understanding and analyzing speech features, and improving the model's ability to recognize complex speech environments. Overall, these steps, through systematic data processing and model optimization, improve the accuracy and robustness of voice-changing recognition, ensuring effectiveness and reliability in practical applications, and providing a solid technical guarantee for anti-fraud communications.
[0029] Preferably, step S2 includes the following steps:
[0030] Step S21: Perform statistical analysis on the feature distribution, clustering, and outliers of the voiceprint feature map data to obtain voiceprint feature statistics.
[0031] Step S22: Based on the voiceprint feature statistics, simulate and analyze the voice change recognition behavior of the voiceprint feature map data to obtain voice change recognition simulation data;
[0032] Step S23: Evaluate the recognition performance of the voice changer recognition simulation data and calculate the recognition accuracy, false alarm rate and false negative rate under different conditions to obtain the recognition performance evaluation data;
[0033] Step S24: Based on the identification performance evaluation data, conduct performance impact analysis on different environmental parameters and establish a relationship model between environmental parameters and identification performance to obtain the environmental impact model;
[0034] Step S25: Set initial recognition thresholds for different call states based on the environmental impact model and recognition performance evaluation data, adjust the thresholds according to real-time environmental parameters, train the machine learning model, and thus obtain a dynamic recognition threshold model.
[0035] This invention provides a comprehensive overview of the distribution of feature data, helping to understand the performance of voice features under different call conditions and improving the quality of the model's foundational data. Identifying and grouping similar voiceprint features helps discover potential voice-changing patterns, improving the accuracy of voice-changing recognition. Detecting and analyzing abnormal feature points enhances the system's ability to recognize abnormal voice behavior, reducing false alarms and missed alarms. Voice-changing recognition simulations based on detailed statistical data provide a comprehensive understanding of voice-changing behavior, enhancing the system's accuracy and reliability in voice-changing recognition. A systematic evaluation of the accuracy, false alarm rate, and missed alarm rate of voice-changing recognition provides a detailed understanding of model performance, guiding model optimization and adjustment. Model optimization is performed based on performance evaluation data, improving the overall performance and stability of recognition. The impact of different environmental parameters on recognition performance is revealed, providing a deeper understanding of environmental factors and ensuring the model's adaptability and robustness in various environments. A relationship model between environmental parameters and recognition performance is established, providing a basis for dynamically adjusting recognition parameters and ensuring continuous optimization of recognition performance. By setting reasonable initial recognition thresholds for different call states, the initial accuracy of the recognition system is improved. The recognition thresholds are dynamically adjusted based on real-time environmental parameters to ensure accuracy and robustness under various call conditions. Through continuous machine learning model training, the recognition threshold model is continuously optimized, enhancing the system's adaptability and recognition performance. Overall, these steps, through systematic data processing, analysis, and model optimization, significantly improve the accuracy and robustness of voice-changing recognition, ensuring its effectiveness and reliability in practical applications and providing a solid technical guarantee for fraud prevention communications.
[0036] Preferably, step S3 includes the following steps:
[0037] Step S31: Extract the key control parameters and their influence range from the dynamic identification threshold model to obtain the control parameter influence matrix;
[0038] Step S32: Set the initial control parameter values according to the preset target recognition accuracy and control parameter influence matrix to obtain the initial control parameter set;
[0039] Step S33: Perform parameter sensitivity testing on the initial control parameter set based on the degree of influence of each control parameter on the recognition accuracy, thereby obtaining parameter sensitivity data; use the gradient descent method to adaptively adjust the parameters based on the parameter sensitivity data, thereby obtaining control parameter adjustment data;
[0040] Step S34: Based on the control parameters, adjust the data and monitor the accuracy of recognition, processing delay and resource consumption in real time to design a multi-level feedback control loop, thereby obtaining a multi-level feedback control system. The multi-level feedback control loop includes two levels: fast response and long-term optimization.
[0041] Step S35: Run the multi-level feedback control system on the simulation test platform, collect system response data, and perform control effect analysis to obtain control system performance data;
[0042] Step S36: Based on the performance data of the control system, the multi-level feedback control system is encapsulated into a deployable adaptive recognition model, thereby obtaining the adaptive recognition model.
[0043] This invention extracts key control parameters and their influence ranges, clarifying the key variables in the model and improving its understandability and controllability. The control parameter influence matrix provides the specific impact of each parameter on recognition performance, offering a scientific basis for subsequent parameter optimization and adjustment. Based on the target recognition accuracy and the control parameter influence matrix, initial control parameter values are rationally set to improve the accuracy and performance of the system's initial operation. The optimized initial control parameter set reduces model debugging and adjustment time, improving development efficiency. Sensitivity testing clarifies the specific impact of each control parameter on recognition accuracy, providing optimization directions. Gradient descent is used for parameter adjustment, improving the accuracy and efficiency of control parameter adjustment and ensuring the optimized performance of the recognition system. By designing a multi-level feedback control loop with fast response and long-term optimization, the system can quickly respond to short-term fluctuations while achieving long-term optimization. Real-time monitoring of recognition accuracy, processing latency, and resource consumption improves the overall system performance and stability. Simulation testing platforms are used to run the multi-level feedback control system, identifying and resolving potential problems in advance to ensure successful system deployment. System response data is collected for control effect analysis, providing comprehensive performance data and a scientific basis for system optimization. The optimized multi-level feedback control system is encapsulated into a deployable adaptive recognition model, facilitating practical applications and improving the model's usability and deployment efficiency. Based on comprehensive performance data, the optimized adaptive recognition model exhibits higher stability and reliability, ensuring excellent performance in real-world applications. Overall, these steps, through systematic parameter extraction, optimization, and multi-level feedback control design, significantly improve the accuracy, robustness, and adaptability of the voice-changing recognition system, providing an efficient and reliable technical solution for fraud prevention communications.
[0044] Preferably, step S4 includes the following steps:
[0045] Step S41: Obtain user historical call data and perform de-identification processing on the user historical call data to obtain de-identified historical call data;
[0046] Step S42: Perform time series analysis on the de-identified historical call data and extract user call pattern features to obtain user call pattern feature data;
[0047] Step S43: Construct a social network graph of users based on call subjects and frequencies using anonymized historical call data to obtain social network feature data;
[0048] Step S44: Based on user call pattern feature data and social network feature data, perform user classification using a clustering algorithm to obtain user profile data;
[0049] Step S45: Based on user profile data and adaptive recognition model, perform abnormal behavior prediction to obtain abnormal behavior prediction data;
[0050] Step S46: Optimize the identification threshold based on the abnormal behavior prediction data using machine learning technology to obtain personalized identification scheme data.
[0051] This invention effectively protects users' personal privacy information through anonymization, avoiding the risks of sensitive data leakage and privacy infringement; it ensures the security of historical call data during processing and storage, complying with data protection regulations and standards, and enhancing the credibility and transparency of data management. Through time series analysis, it extracts users' call pattern characteristics, including call frequency and time-of-day preferences, to gain a deeper understanding of users' behavioral habits and communication patterns; based on call pattern characteristic data, it personalizes services, such as optimizing communication plans and providing customized recommendations, enhancing user experience and satisfaction. It constructs users' social network graphs, analyzes call partners and their frequency, revealing the density and activity of users' relationships within their social circles; through social network characteristic data, it identifies and understands users' influence and key roles in social networks, providing data support for social interaction. Combining user profile data and adaptive recognition models, it effectively predicts abnormal patterns and events in user communication behavior, promptly identifying and responding to potential risks. It improves anti-fraud capabilities, reduces potential damage to user and system security caused by abnormal behavior, and protects communication security and user interests. By leveraging machine learning techniques to dynamically adjust identification thresholds based on abnormal behavior prediction data, personalized anti-fraud identification solutions are achieved. Optimizing these thresholds effectively improves the accuracy and real-time response capabilities of the anti-fraud system, while reducing false alarm and false negative rates. In summary, these steps combine data processing, analysis, and prediction technologies, providing multi-dimensional data support and analytical capabilities for anti-fraud communication methods, effectively enhancing system security, user experience, and operational efficiency.
[0052] Preferably, step S5 includes the following steps:
[0053] Step S51: Obtain ambient noise data and network quality data for the call;
[0054] Step S52: Evaluate the call quality based on the call environment noise data and network quality data to obtain a call quality model. The call quality evaluation includes audio clarity evaluation, continuity evaluation, and echo cancellation effect evaluation.
[0055] Step S53: Collect scenario call data in various real-world scenarios, and use transfer learning techniques to optimize the call quality model's generalization ability based on the scenario call data, thereby obtaining optimized call quality data.
[0056] This invention, by acquiring ambient noise and network quality data, enables the system to perceive the noise level and network conditions of the real-time communication environment, facilitating the adjustment of communication parameters to optimize call quality. It provides fundamental data support, offering necessary input information for subsequent call quality assessment and optimization. Based on the acquired environmental data, the system can accurately assess call quality, including key indicators such as audio clarity, call continuity, and echo cancellation effectiveness. The assessment results can be used to adjust communication parameters in real time, optimizing the call experience and ensuring good call quality for users under various environmental conditions. By collecting call data in multiple real-world scenarios and applying transfer learning techniques, the generalization ability of the call quality model is optimized, enabling it to adapt to different communication scenarios and environmental conditions. This ensures the effectiveness and stability of call quality optimization strategies in various real-world scenarios, improving the system's reliability and practicality in real-world applications. Continuous collection and analysis of scenario call data allows for continuous optimization of the call quality model, maintaining the system's adaptability to changes in the communication environment and providing stable call service quality. In summary, these steps integrate data acquisition, analysis, and model optimization technologies, providing crucial call quality assessment and optimization capabilities for anti-fraud communication methods, thereby significantly improving the system's user experience and operational efficiency.
[0057] Preferably, step S6 includes the following steps:
[0058] Step S61: Perform data cleaning and standardization on the personalized identification scheme data and call quality optimization data to obtain personalized identification standard data and call quality standard data;
[0059] Step S62: Extract timbre features, emotional features, and environmental noise features from the personalized identification standard data and the call quality standard data, and use the analytic hierarchy process (AHP) to assign feature weights, thereby obtaining feature weight assignment data.
[0060] Step S63: Input the feature weight allocation data, personalized identification standard data, and call quality standard data into the preset multi-criteria decision model to obtain multi-criteria decision data;
[0061] Step S64: Identify potential abnormal behaviors and fraudulent activities in the multi-criteria decision data, and conduct risk assessment and threshold adjustment to obtain the final anti-fraud identification scheme data.
[0062] This invention ensures the consistency and accuracy of personalized identification scheme data and call quality optimization data through data cleaning and standardization, improving the reliability of subsequent analysis and decision-making. Cleaning and standardization remove noise and errors from the data, making subsequent data analysis and decision-making more accurate and reliable. Processing data according to unified standards facilitates comparison and integration between different data sources, enhancing the operability and application value of the data. By extracting timbre features, emotional features, and environmental noise features, the key features of call quality and personalized identification schemes are analyzed in depth, providing strong support for subsequent decision-making. Using the analytic hierarchy process (AHP) for feature weight allocation allows for an objective assessment of the contribution of each feature to the final decision, improving the scientific nature and accuracy of the decision. Comprehensive consideration and weighing of different types of data helps establish a comprehensive decision-making model, improving the system's decision-making efficiency and accuracy. Through a multi-criteria decision-making model that comprehensively considers multiple factors such as feature weights, personalized identification standards, and call quality standards, the anti-fraud identification scheme can be more comprehensively evaluated and optimized. Based on a scientific decision-making model, subjective bias can be effectively reduced, improving the objectivity and credibility of the decision. Through multi-criteria decision-making, the system can more accurately identify potential anomalies and fraudulent behaviors, taking timely preventative and responsive measures to protect communication security and user rights. Identifying abnormal and fraudulent behaviors based on multi-criteria decision data effectively enhances the system's ability to perceive and respond to potential risks. In-depth assessment and analysis of identified abnormal behaviors quantifies risk levels, providing a basis for developing targeted preventative measures. The thresholds of the fraud prevention scheme are dynamically adjusted based on real-time risk assessment results, ensuring the system's stable and efficient operation under different scenarios. In summary, these steps effectively integrate data analysis, decision models, and risk management technologies, providing comprehensive data support and a scientific basis for decision-making in fraud prevention communication methods. Attached Figure Description
[0063] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0064] Figure 1 This is a flowchart illustrating the steps of the anti-fraud communication method based on voice changer recognition according to the present invention.
[0065] Figure 2 for Figure 1 A detailed flowchart of step S1;
[0066] Figure 3 for Figure 1 A detailed flowchart of step S2. Detailed Implementation
[0067] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0068] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0069] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0070] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a fraud prevention communication method based on voice-changing recognition, the method comprising the following steps:
[0071] Step S1: Obtain multi-dimensional voice feature data of users, including timbre, frequency features, language rhythm, emotional features and pause patterns; predict the possibility of voice change based on multi-dimensional voice feature data to obtain voiceprint feature prediction data; train a deep learning model based on voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0072] Step S2: Simulate and analyze the voice-changing recognition behavior under different call states based on the voiceprint feature map data to obtain voice-changing recognition simulation data; establish a dynamic recognition threshold model based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data.
[0073] Step S3: Adjust the control parameters in real time according to the preset target recognition accuracy and dynamic recognition threshold model to obtain control parameter adjustment data; construct a multi-level feedback control system based on the control parameter adjustment data to obtain an adaptive recognition model;
[0074] Step S4: Obtain user's historical call data, perform personalized analysis on the user's historical call data, and extract features to obtain user profile data; predict abnormal behavior based on user profile data and adaptive recognition model, and use machine learning technology to optimize the acquisition of recognition thresholds to obtain personalized recognition scheme data.
[0075] Step S5: Obtain call environment noise data and network quality data, and construct a call quality model based on the call environment noise data and network quality data; optimize the call quality model across scenarios, and adjust the communication parameters in real time to obtain call quality optimization data;
[0076] Step S6: Perform multi-criteria decision fusion analysis on the personalized identification scheme data and call quality optimization data to obtain the final anti-fraud identification scheme data.
[0077] In this embodiment of the invention, reference Figure 1 The above is a flowchart illustrating the steps of an anti-fraud communication method based on voice changer recognition according to the present invention. In this example, the anti-fraud communication method based on voice changer recognition includes the following steps:
[0078] Step S1: Obtain multi-dimensional voice feature data of users, including timbre, frequency features, language rhythm, emotional features and pause patterns; predict the possibility of voice change based on multi-dimensional voice feature data to obtain voiceprint feature prediction data; train a deep learning model based on voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0079] This invention employs audio processing libraries such as Librosa or PyAudio for audio loading and spectral analysis, extracting duct length, formants, and acoustic waveforms. Fourier transform or short-time Fourier transform (STFT) is applied to extract spectral information. Voice activity detection algorithms, such as VAD (Voice Activity Detection), are used to identify speech and non-speech segments. Emotional information is extracted by analyzing speech rate and pitch changes. Natural pauses in speech are detected by analyzing the duration of speech segments and the pause patterns of non-speech segments. Machine learning algorithms such as Support Vector Machines (SVM) are used to classify multi-dimensional speech feature data and predict the voice change probability score for each speech sample. Based on the predicted voiceprint feature data, deep learning models, such as Convolutional Neural Networks (CNN) and Long Short-Term Memory Networks (LSTM), are designed to train the voiceprint recognition model. This model maps voiceprint states and features under different call conditions, generating voiceprint feature map data for subsequent analysis and simulation.
[0080] Step S2: Simulate and analyze the voice-changing recognition behavior under different call states based on the voiceprint feature map data to obtain voice-changing recognition simulation data; establish a dynamic recognition threshold model based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data.
[0081] This invention uses voiceprint feature map data to simulate and analyze voice-changing recognition behavior under different call states, and evaluates the recognition accuracy, false alarm rate, and false negative rate; combined with call environment parameters and voice-changing recognition simulation data, a dynamic recognition threshold model is established using statistical methods or machine learning algorithms; the initial recognition threshold is determined for different call states and dynamically adjusted as real-time environmental parameters change.
[0082] Step S3: Adjust the control parameters in real time according to the preset target recognition accuracy and dynamic recognition threshold model to obtain control parameter adjustment data; construct a multi-level feedback control system based on the control parameter adjustment data to obtain an adaptive recognition model;
[0083] This invention, in its embodiments, adjusts control parameters in real time based on a preset target recognition accuracy and dynamic recognition threshold model, such as adjusting the model's learning rate and decision boundary. A multi-level feedback control system is designed based on the control parameter adjustment data, including two levels: rapid response and long-term optimization. The rapid response level quickly adjusts real-time call quality and voice-changing recognition performance, while the long-term optimization level performs more in-depth systemic optimization based on historical data.
[0084] Step S4: Obtain user's historical call data, perform personalized analysis on the user's historical call data, and extract features to obtain user profile data; predict abnormal behavior based on user profile data and adaptive recognition model, and use machine learning technology to optimize the acquisition of recognition thresholds to obtain personalized recognition scheme data.
[0085] This invention de-identifies user historical call data to ensure data privacy and security. It employs time series analysis and clustering algorithms to extract user call pattern features, such as call frequency and time-of-day preferences. Based on user profile data and an adaptive recognition model, it predicts abnormal user call behavior and utilizes machine learning techniques to optimize personalized recognition thresholds, ensuring timely identification and response to abnormal behavior.
[0086] Step S5: Obtain call environment noise data and network quality data, and construct a call quality model based on the call environment noise data and network quality data; optimize the call quality model across scenarios, and adjust the communication parameters in real time to obtain call quality optimization data;
[0087] This invention acquires call environment noise data and network quality data, such as audio signal-to-noise ratio and network latency, and uses the collected call environment data to construct a call quality evaluation model to evaluate audio clarity, continuity, and echo effects. Transfer learning technology is used to optimize the model's generalization ability to ensure call quality optimization effects in various real-world scenarios.
[0088] Step S6: Perform multi-criteria decision fusion analysis on the personalized identification scheme data and call quality optimization data to obtain the final anti-fraud identification scheme data.
[0089] This invention cleans and standardizes personalized identification scheme data and call quality optimization data to ensure data quality and consistency. The cleaned data is then input into a preset multi-criteria decision model, such as the analytic hierarchy process or decision tree, for comprehensive decision analysis. Considering factors such as feature weights, user profiles, and call quality, the optimal anti-fraud identification scheme is formulated.
[0090] This invention, through comprehensive analysis of these features, can more accurately capture the speaker's unique voiceprint, reducing the possibility of misjudgment and omission. These features include timbre, frequency characteristics, speech rhythm, emotional characteristics, and pause patterns. By analyzing multi-dimensional speech feature data, it predicts whether the voice has undergone voice-changing processing, obtaining voiceprint feature prediction data; this step helps to detect potential voice-changing behavior early. Using the voiceprint feature prediction data to train a deep learning model, the model can learn to recognize various complex speech features, improving the accuracy and robustness of recognition. By mapping the voiceprint state under different call conditions, voiceprint feature map data is obtained, helping the system better understand and analyze the performance of speech features in different environments. Simulation analysis of voice-changing recognition behavior under different call conditions yields voice-changing recognition simulation data. This helps the system accurately identify voice-changing behavior under different environments and conditions. Based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data, a dynamic recognition threshold model is established, allowing for flexible adjustment of the recognition threshold, improving the sensitivity and accuracy of recognition. Based on the preset target recognition accuracy and the dynamic recognition threshold model, control parameters are adjusted in real time to ensure the system is always in the optimal recognition state. A multi-level feedback control system is constructed, continuously optimizing the system based on control parameters to improve its adaptability and recognition accuracy. Historical call data is acquired from users, undergoing personalized analysis and feature extraction to create user profiles, enabling a better understanding of users' voice characteristics and behavioral patterns. Based on user profile data and an adaptive recognition model, abnormal behavior is predicted, and machine learning techniques are used to obtain the optimal recognition threshold, resulting in personalized recognition scheme data. Call environment noise data and network quality data are acquired to construct a call quality model, accurately reflecting the impact of call quality on recognition. The call quality model is optimized across scenarios, and communication parameters are adjusted in real time to ensure the system's recognition performance in different environments. Multi-criteria decision fusion analysis is performed on personalized recognition scheme data and call quality optimization data to comprehensively evaluate various factors, improving recognition accuracy and reliability; comprehensive recognition decisions are made considering multiple factors to improve overall system performance; and multi-criteria decision-making reduces false alarms and false negatives, improving system usability and user experience.
[0091] Preferably, step S1 includes the following steps:
[0092] Step S11: Obtain the user's historical voice data and perform a short-time Fourier transform on the historical voice data to obtain the voice time-frequency data;
[0093] Step S12: Extract Mel frequency cepstral coefficients from the speech time-frequency data to obtain timbre feature data; extract frequency features based on fundamental frequency contours from the speech time-frequency data to obtain frequency feature data;
[0094] Step S13: Use a speech activity detection algorithm to identify speech and non-speech segments in the speech time-frequency data to obtain language rhythm feature data;
[0095] Step S14: Analyze the duration of speech segments and the pause patterns of non-speech segments based on the language rhythm feature data to obtain pause pattern data.
[0096] Step S15: Analyze the pitch variation, energy distribution, and speech rate of the speech time-frequency data, and extract emotional features to obtain emotional feature data; perform harmonic-noise ratio analysis on the speech time-frequency data, and evaluate the harmonic structure of the speech to obtain harmonic structure data;
[0097] Step S16: Merge the timbre feature data, frequency feature data, language rhythm feature data, pause pattern data, and emotional feature data into multi-dimensional speech feature data based on the harmonic structure data;
[0098] Step S17: Use support vector machines to classify multi-dimensional speech feature data and score the probability of voice change to obtain voiceprint feature prediction data.
[0099] Step S18: Train a deep learning model based on the voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0100] As an embodiment of the present invention, reference is made to... Figure 2 As shown, Figure 1 A detailed flowchart of step S1 is shown below. In this embodiment of the invention, step S1 includes the following steps:
[0101] Step S11: Obtain the user's historical voice data and perform a short-time Fourier transform on the historical voice data to obtain the voice time-frequency data;
[0102] This invention retrieves historical voice data from a voice database or user communication records, ensuring data security and privacy protection. The voice signal is processed in frames, each typically approximately 20-40 milliseconds long. Overlapping portions can be set to half the frame length. A window function (such as the Hanning window) is applied to each frame to reduce spectral leakage. An FFT (Fast Fourier Transform) is then performed on each windowed frame to obtain spectral information. The spectral information of each frame is represented as time-frequency domain data, i.e., the spectral distribution on the time axis.
[0103] Step S12: Extract Mel frequency cepstral coefficients from the speech time-frequency data to obtain timbre feature data; extract frequency features based on fundamental frequency contours from the speech time-frequency data to obtain frequency feature data;
[0104] In this embodiment of the invention, a Mel filter bank is applied to the time-frequency data of each frame to convert the linear spectrum into a Mel frequency spectrum. The logarithm is taken to obtain the power spectrum, and then a Discrete Cosine Transform (DCT) is performed to obtain the MFCC coefficients. The MFCC coefficients are extracted as timbre feature data. The fundamental frequency information of the speech signal is analyzed to extract the fundamental frequency profile; frequency features, such as fundamental frequency variation patterns and frequency distribution, are extracted based on the fundamental frequency profile.
[0105] Step S13: Use a speech activity detection algorithm to identify speech and non-speech segments in the speech time-frequency data to obtain language rhythm feature data;
[0106] In this embodiment of the invention, a speech activity detection algorithm is applied to the time-frequency data of each frame to identify speech segments and non-speech segments, and to analyze rhythmic features such as the duration of speech segments and the pause patterns of non-speech segments.
[0107] Step S14: Analyze the duration of speech segments and the pause patterns of non-speech segments based on the language rhythm feature data to obtain pause pattern data.
[0108] The embodiments of the present invention analyze the duration distribution of speech segments and the pause duration of non-speech segments based on language rhythm feature data.
[0109] Step S15: Analyze the pitch variation, energy distribution, and speech rate of the speech time-frequency data, and extract emotional features to obtain emotional feature data; perform harmonic-noise ratio analysis on the speech time-frequency data, and evaluate the harmonic structure of the speech to obtain harmonic structure data;
[0110] This invention analyzes pitch variations, energy distribution, and speech rate in speech time-frequency data to extract emotional features; calculates the ratio of harmonic components to noise components in the speech signal, and evaluates the harmonic structure of the speech.
[0111] Step S16: Merge the timbre feature data, frequency feature data, language rhythm feature data, pause pattern data, and emotional feature data into multi-dimensional speech feature data based on the harmonic structure data;
[0112] This invention integrates timbre features, frequency features, language rhythm features, pause pattern data, and emotional feature data into a multi-dimensional speech feature dataset.
[0113] Step S17: Use support vector machines to classify multi-dimensional speech feature data and score the probability of voice change to obtain voiceprint feature prediction data.
[0114] In this embodiment of the invention, multi-dimensional speech feature data is input into an SVM classifier to predict and classify the voiceprint features.
[0115] Step S18: Train a deep learning model based on the voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
[0116] This invention designs a deep learning model based on voiceprint feature prediction data, such as a convolutional neural network (CNN) or a long short-term memory network (LSTM); under different call conditions, it maps voiceprint states and features to generate voiceprint feature map data.
[0117] This invention utilizes technologies such as STFT, MFCC, fundamental frequency profile, and VAD to extract speech features from multiple dimensions, providing high-resolution and high-precision speech analysis. Combining timbre, frequency, speech rhythm, pause patterns, emotional features, and harmonic structure, it provides comprehensive speech feature data, improving recognition accuracy and robustness. By dynamically adjusting the recognition threshold and implementing a multi-level feedback control system, it enhances the system's adaptability under different call conditions, ensuring recognition stability and reliability. Based on users' historical speech data and personalized analysis, it provides personalized recognition schemes, improving the system's ability to recognize individual users and reducing false positives and false negatives. Deep learning models are used for training and mapping to capture complex speech features, improving the system's ability to detect voice-changing behavior and its robustness. A multi-criteria decision model comprehensively evaluates various factors for comprehensive recognition decisions, improving the overall system performance and reducing false positives and false negatives. Overall, this anti-fraud communication method based on voice-changing recognition significantly improves the security and accuracy of voice communication through multi-dimensional and multi-level comprehensive analysis and optimization. It effectively copes with advanced voice-changing technologies and complex communication environments, providing more reliable anti-fraud protection.
[0118] Preferably, step S18 includes the following steps:
[0119] Step S181: Add noise and time scaling to the voiceprint feature prediction data, and perform normalization processing to obtain the voiceprint feature training dataset.
[0120] This invention selects appropriate noise types and intensities from the original voiceprint feature prediction data and adds them to the speech signal; it then performs time scaling on the speech signal to adjust the playback speed and generate speech segments of different lengths; ensuring that adding noise and time scaling do not significantly change the recognizability of the speech signal, i.e., the speech features can still maintain their recognizability. After adding noise and time scaling, normalization is applied to the processed voiceprint feature data; this ensures that the normalized data has similar numerical ranges across different feature dimensions, thereby improving the stability and convergence speed of model training.
[0121] Step S182: Train the preset deep learning model based on the voiceprint feature training dataset, and adjust the parameters of the model using the cross-validation method to obtain the voiceprint feature model. The deep learning model includes a multi-layer convolutional neural network structure for extracting local features, a long short-term memory network layer for capturing temporal dependencies, an attention mechanism layer for highlighting important features, and a fully connected layer.
[0122] This invention constructs the network structure of a deep learning model, determines the number of neurons, activation function, and connection method for each layer; defines an appropriate loss function (such as cross-entropy loss) and optimizer (such as the Adam optimizer) to set the correct objective for the model training process; and uses cross-validation to adjust the model's hyperparameters (such as learning rate, batch size, etc.) to obtain the best model performance.
[0123] Step S183: Simulate different call conditions to obtain conditional simulation data; apply the voiceprint feature model to the conditional simulation data to obtain conditional response data;
[0124] Based on the voiceprint feature training dataset generated in the previous steps, this invention creates multiple simulated scenarios; adjusts parameters in the simulated scenarios, such as noise level, network quality, and voice quality, to generate voiceprint feature data under different call conditions; and ensures that the simulated data can reflect the diversity and complexity that may occur in real calls.
[0125] Step S184: Map the voiceprint status and features under different call conditions based on the conditional response data and conditional simulation data to obtain voiceprint feature map data.
[0126] In this embodiment of the invention, conditional simulation data is input into a pre-trained deep learning voiceprint feature model; based on the model output, the voiceprint state and features under different call conditions are analyzed and mapped; voiceprint feature map data is generated, which reflects the changes and responses of voiceprint features under different call conditions.
[0127] This invention enhances the diversity of training data by adding noise and time scaling, improving the model's generalization ability and enabling it to better adapt to various real-world call environments. Normalization reduces data bias, improving the stability and efficiency of model training and ensuring features are compared on the same scale. Extracting local speech features improves the model's ability to recognize subtle speech features; capturing temporal dependencies in speech signals improves the model's understanding of dynamic speech changes; highlighting important features enhances the model's focus on key speech features, improving recognition accuracy. Simulating different call conditions comprehensively tests the model's performance in various real-world environments, improving its adaptability; obtaining model response data under different conditions helps evaluate and optimize model performance, ensuring its accuracy and robustness in various environments. Mapping voiceprint states and features under different call conditions generates voiceprint feature map data, providing intuitive visualization information to help understand and analyze speech features; integrating feature data under various conditions improves the model's ability to recognize complex speech environments, ensuring stability and reliability in practical applications. Optimizing model parameters improves the model's robustness and generalization ability, ensuring stable performance under different call conditions. By adding noise and time scaling, the diversity of training data is enhanced, improving the model's adaptability to various real-world call environments. Multi-layer convolutional neural networks and long short-term memory (LSTM) layers are used to extract local features and capture temporal dependencies, improving the model's ability to recognize speech features. An attention mechanism layer highlights important features, enhancing the model's focus on key speech features and improving recognition accuracy. Cross-validation optimizes model parameters, improving robustness and generalization ability, ensuring stable performance under different call conditions. Simulating various call conditions comprehensively tests and optimizes model performance, improving its adaptability and accuracy in various environments. Generating voiceprint feature map data provides intuitive visualization information, aiding in understanding and analyzing speech features, and improving the model's ability to recognize complex speech environments. Overall, these steps, through systematic data processing and model optimization, improve the accuracy and robustness of voice-changing recognition, ensuring effectiveness and reliability in practical applications, and providing a solid technical guarantee for anti-fraud communications.
[0128] Preferably, step S2 includes the following steps:
[0129] Step S21: Perform statistical analysis on the feature distribution, clustering, and outliers of the voiceprint feature map data to obtain voiceprint feature statistics.
[0130] Step S22: Based on the voiceprint feature statistics, simulate and analyze the voice change recognition behavior of the voiceprint feature map data to obtain voice change recognition simulation data;
[0131] Step S23: Evaluate the recognition performance of the voice changer recognition simulation data and calculate the recognition accuracy, false alarm rate and false negative rate under different conditions to obtain the recognition performance evaluation data;
[0132] Step S24: Based on the identification performance evaluation data, conduct performance impact analysis on different environmental parameters and establish a relationship model between environmental parameters and identification performance to obtain the environmental impact model;
[0133] Step S25: Set initial recognition thresholds for different call states based on the environmental impact model and recognition performance evaluation data, adjust the thresholds according to real-time environmental parameters, train the machine learning model, and thus obtain a dynamic recognition threshold model.
[0134] As an embodiment of the present invention, reference is made to... Figure 3 As shown, Figure 1 A detailed flowchart of step S2 is shown below. In this embodiment of the invention, step S2 includes the following steps:
[0135] Step S21: Perform statistical analysis on the feature distribution, clustering, and outliers of the voiceprint feature map data to obtain voiceprint feature statistics.
[0136] This invention, through calculating various statistical indicators (such as mean, variance, skewness, kurtosis, etc.) of voiceprint feature data, can reveal the distribution of voiceprint features. Statistical methods and functions, such as descriptive statistics, histograms, and kernel density estimation, can be used to analyze the distribution characteristics of voiceprint features. Clustering algorithms (such as K-means clustering, hierarchical clustering, etc.) are used to perform cluster analysis on voiceprint features to discover clusters and patterns in the data. Samples can be clustered into different groups based on the similarity of voiceprint features, and the feature differences and similarities between each group can be analyzed. Outlier detection algorithms (such as box plots, Z-scores, isolated forests, etc.) are used to identify outliers in the voiceprint feature data. Outliers may indicate abnormal situations during data acquisition or processing and require special attention and handling.
[0137] Step S22: Based on the voiceprint feature statistics, simulate and analyze the voice change recognition behavior of the voiceprint feature map data to obtain voice change recognition simulation data;
[0138] This invention employs voiceprint feature extraction tools to extract voiceprint features from sound data, such as MFCC (Melbourne Frequency Cepstral Coefficients) based feature extraction or other voiceprint feature extraction methods. These features may include spectral characteristics, temporal characteristics, and energy characteristics of the sound. Simulation analysis tools are used to simulate voice changes in the voiceprint features by altering parameters such as frequency, amplitude, and resonance. Digital signal processing techniques and sound synthesis algorithms can be used to implement voice change simulation. The simulated voiceprint features are compared and analyzed with the original voiceprint features to obtain simulated data for voice change recognition. Machine learning algorithms and classifiers can be used to classify and recognize the voiceprint features, evaluating the performance and accuracy of voice change recognition.
[0139] Step S23: Evaluate the recognition performance of the voice changer recognition simulation data and calculate the recognition accuracy, false alarm rate and false negative rate under different conditions to obtain the recognition performance evaluation data;
[0140] This invention selects appropriate evaluation metrics to measure the performance of voice changer recognition, such as accuracy, false positive rate, false negative rate, precision, and recall. These metrics help evaluate the overall performance and recognition capability of the system. The simulated voice changer recognition data is divided into training and testing sets. The model is trained using the training set, and then its performance is evaluated using the testing set. Techniques such as cross-validation can be used to reduce the bias of the evaluation results and improve their reliability. Evaluation metrics, such as accuracy, false positive rate, and false negative rate, are calculated under different conditions based on the test results. These metrics can be calculated using a confusion matrix, based on the number of true positives, true negatives, false positives, and false negatives.
[0141] Step S24: Based on the identification performance evaluation data, conduct performance impact analysis on different environmental parameters and establish a relationship model between environmental parameters and identification performance to obtain the environmental impact model;
[0142] This invention collects different environmental parameters (such as noise level, echo conditions, and speech quality) and corresponding recognition performance evaluation data. The data is then organized for analysis and modeling. Statistical analysis and data visualization methods are used to analyze the impact of different environmental parameters on recognition performance. Correlation analysis and analysis of variance can be used to determine the relationship between environmental parameters and performance indicators. Based on the analysis results, a model of the relationship between environmental parameters and recognition performance is established. Machine learning algorithms such as linear regression, logistic regression, and support vector machines can be used to build the model, and the model is then trained and validated. The performance and accuracy of the established environmental impact model are evaluated, and optimization and adjustments are made as needed. Cross-validation and model evaluation metrics (such as R-squared value and mean squared error) can be used to evaluate the quality of the model.
[0143] Step S25: Set initial recognition thresholds for different call states based on the environmental impact model and recognition performance evaluation data, adjust the thresholds according to real-time environmental parameters, train the machine learning model, and thus obtain a dynamic recognition threshold model.
[0144] This invention organizes voiceprint features and simulated voice changer recognition data into the format required for a machine learning model. Typically, feature data is used as input variables, and the voice changer recognition result is used as the target variable. The dataset is ensured to contain sufficient samples and labels for effective model training. Voiceprint features are further processed and transformed as needed. This may include feature scaling, dimensionality reduction, and feature selection to improve model performance and generalization ability. A suitable machine learning model is selected based on task requirements and data characteristics. Common voiceprint recognition models include Support Vector Machines (SVM), Random Forest, and Deep Neural Networks. The selected model is trained using the prepared dataset. This involves dividing the dataset into training and validation sets and iteratively optimizing the model's parameters and hyperparameters to best adapt the model to the training data. The performance of the trained model is evaluated using a test set. Common evaluation metrics include accuracy, precision, recall, and F1 score. Based on the evaluation results, the model can be adjusted and optimized to improve its performance. Once the requirements are met, the trained model can be saved for use in subsequent voiceprint recognition tasks.
[0145] This invention provides a comprehensive overview of the distribution of feature data, helping to understand the performance of voice features under different call conditions and improving the quality of the model's foundational data. Identifying and grouping similar voiceprint features helps discover potential voice-changing patterns, improving the accuracy of voice-changing recognition. Detecting and analyzing abnormal feature points enhances the system's ability to recognize abnormal voice behavior, reducing false alarms and missed alarms. Voice-changing recognition simulations based on detailed statistical data provide a comprehensive understanding of voice-changing behavior, enhancing the system's accuracy and reliability in voice-changing recognition. A systematic evaluation of the accuracy, false alarm rate, and missed alarm rate of voice-changing recognition provides a detailed understanding of model performance, guiding model optimization and adjustment. Model optimization is performed based on performance evaluation data, improving the overall performance and stability of recognition. The impact of different environmental parameters on recognition performance is revealed, providing a deeper understanding of environmental factors and ensuring the model's adaptability and robustness in various environments. A relationship model between environmental parameters and recognition performance is established, providing a basis for dynamically adjusting recognition parameters and ensuring continuous optimization of recognition performance. By setting reasonable initial recognition thresholds for different call states, the initial accuracy of the recognition system is improved. The recognition thresholds are dynamically adjusted based on real-time environmental parameters to ensure accuracy and robustness under various call conditions. Through continuous machine learning model training, the recognition threshold model is continuously optimized, enhancing the system's adaptability and recognition performance. Overall, these steps, through systematic data processing, analysis, and model optimization, significantly improve the accuracy and robustness of voice-changing recognition, ensuring its effectiveness and reliability in practical applications and providing a solid technical guarantee for fraud prevention communications.
[0146] Preferably, step S3 includes the following steps:
[0147] Step S31: Extract the key control parameters and their influence range from the dynamic identification threshold model to obtain the control parameter influence matrix;
[0148] This invention selects a suitable dynamic recognition threshold model as the research object, such as a probability distribution-based model or a machine learning-based classifier model; it determines possible key control parameters, such as threshold, weights, and learning rate. A set of experiments is designed, changing the value of one control parameter in each experiment while keeping other parameters constant. The recognition accuracy and control parameter values are recorded for each experiment. The impact of each control parameter on the recognition accuracy is analyzed based on the experimental data, forming a control parameter influence matrix. Statistical methods, sensitivity analysis, and other techniques can be used to evaluate the range of influence of the control parameters.
[0149] Step S32: Set the initial control parameter values according to the preset target recognition accuracy and control parameter influence matrix to obtain the initial control parameter set;
[0150] This invention provides an embodiment that determines a preset target recognition accuracy rate. Based on specific application requirements and performance specifications, a target accuracy rate is set as a reference. Combining the control parameter influence matrix and the target recognition accuracy rate, initial control parameter values are derived from the preset target accuracy rate. Considering the model's adjustable range and practical application scenarios, the initial control parameters are appropriately adjusted and optimized to ensure that the expected performance is achieved in actual operation.
[0151] Step S33: Perform parameter sensitivity testing on the initial control parameter set based on the degree of influence of each control parameter on the recognition accuracy, thereby obtaining parameter sensitivity data; use the gradient descent method to adaptively adjust the parameters based on the parameter sensitivity data, thereby obtaining control parameter adjustment data;
[0152] This invention performs sensitivity tests on each parameter in the initial control parameter set to evaluate the impact of each parameter on the recognition accuracy. This can be done by fixing other parameters one by one, changing the value of one parameter, and then observing the change in recognition accuracy. Based on the parameter sensitivity data, methods such as gradient descent are used to adaptively adjust the control parameters to achieve the expected recognition accuracy. Optimization algorithms such as gradient descent can calculate the gradient of the parameters based on the parameter sensitivity data and update the parameter values.
[0153] Step S34: Based on the control parameters, adjust the data and monitor the accuracy of recognition, processing delay and resource consumption in real time to design a multi-level feedback control loop, thereby obtaining a multi-level feedback control system. The multi-level feedback control loop includes two levels: fast response and long-term optimization.
[0154] In the fast response level of this invention, a real-time monitoring system is set up to collect data on indicators such as recognition accuracy, processing latency, and resource consumption. Based on the real-time monitoring data, a control loop is designed to quickly respond to changes in system performance. Using control parameter adjustment data and real-time monitoring data, a feedback control algorithm (such as a proportional-integral-derivative controller) is used to adjust the control parameters to keep recognition accuracy, processing latency, and resource consumption within acceptable ranges. In the long-term optimization level, historical data is collected, including control parameter adjustment data, real-time monitoring data, and system performance evaluation indicators. Based on historical data, performance evaluation and analysis are performed to understand the long-term performance trend and existing problems of the system. Optimization algorithms (such as genetic algorithms, particle swarm optimization algorithms, etc.) are used in conjunction with historical data to optimize the control parameters to improve system performance and stability. Based on the optimization results, the control parameters are updated to gradually approach the optimal value, and the system performance is continuously monitored, with necessary adjustments and improvements made.
[0155] Step S35: Run the multi-level feedback control system on the simulation test platform, collect system response data, and perform control effect analysis to obtain control system performance data;
[0156] This invention designs experiments to verify the performance of a multi-level feedback control system. A typical set of input samples can be selected, and the recognition accuracy, processing delay, and resource consumption can be evaluated through actual system operation. Experimental data, including indicators such as recognition accuracy, processing delay, and resource consumption, are collected, analyzed, and evaluated to verify the performance and effectiveness of the multi-level feedback control system. Based on the experimental results, necessary adjustments and improvements are made to the system to further enhance performance and optimize the control strategy.
[0157] Step S36: Based on the performance data of the control system, the multi-level feedback control system is encapsulated into a deployable adaptive recognition model, thereby obtaining the adaptive recognition model.
[0158] Based on experimental verification results and performance evaluation, this invention analyzes the advantages and disadvantages of the system and makes improvements and optimizations. According to experimental data and performance evaluation, the control parameters in the multi-level feedback control system are adjusted and optimized to further improve the system's performance and stability. After optimization and adjustment, the system's performance is re-evaluated, and necessary iterations and improvements are made based on the experimental verification results. Finally, an optimized dynamic recognition threshold model is obtained. This model has good recognition performance and stability, and can adaptively adjust the threshold according to real-time changes to adapt to different environments and application requirements.
[0159] This invention extracts key control parameters and their influence ranges, clarifying the key variables in the model and improving its understandability and controllability. The control parameter influence matrix provides the specific impact of each parameter on recognition performance, offering a scientific basis for subsequent parameter optimization and adjustment. Based on the target recognition accuracy and the control parameter influence matrix, initial control parameter values are rationally set to improve the accuracy and performance of the system's initial operation. The optimized initial control parameter set reduces model debugging and adjustment time, improving development efficiency. Sensitivity testing clarifies the specific impact of each control parameter on recognition accuracy, providing optimization directions. Gradient descent is used for parameter adjustment, improving the accuracy and efficiency of control parameter adjustment and ensuring the optimized performance of the recognition system. By designing a multi-level feedback control loop with fast response and long-term optimization, the system can quickly respond to short-term fluctuations while achieving long-term optimization. Real-time monitoring of recognition accuracy, processing latency, and resource consumption improves the overall system performance and stability. Simulation testing platforms are used to run the multi-level feedback control system, identifying and resolving potential problems in advance to ensure successful system deployment. System response data is collected for control effect analysis, providing comprehensive performance data and a scientific basis for system optimization. The optimized multi-level feedback control system is encapsulated into a deployable adaptive recognition model, facilitating practical applications and improving the model's usability and deployment efficiency. Based on comprehensive performance data, the optimized adaptive recognition model exhibits higher stability and reliability, ensuring excellent performance in real-world applications. Overall, these steps, through systematic parameter extraction, optimization, and multi-level feedback control design, significantly improve the accuracy, robustness, and adaptability of the voice-changing recognition system, providing an efficient and reliable technical solution for fraud prevention communications.
[0160] Preferably, step S4 includes the following steps:
[0161] Step S41: Obtain user historical call data and perform de-identification processing on the user historical call data to obtain de-identified historical call data;
[0162] This invention provides an embodiment of obtaining users' historical call data from communication service providers or system logs, ensuring privacy protection for users' call data. Sensitive information is removed using desensitization techniques, such as using hash functions or encryption algorithms to desensitize phone numbers. After desensitization, the desensitized historical call data is saved for subsequent analysis and processing.
[0163] Step S42: Perform time series analysis on the de-identified historical call data and extract user call pattern features to obtain user call pattern feature data;
[0164] This invention performs time series analysis on anonymized historical call data, including statistics and analysis of call time, call duration, call frequency, etc. Based on the results of the time series analysis, it extracts user call pattern characteristics, such as average call duration, distribution of call frequency, call activity, etc., and saves the extracted user call pattern characteristic data for subsequent social network construction and user classification.
[0165] Step S43: Construct a social network graph of users based on call subjects and frequencies using anonymized historical call data to obtain social network feature data;
[0166] This invention constructs a call relationship graph between users based on anonymized historical call data, with each user as a node and call relationships as edges connecting nodes. The weights of the edges are determined based on call frequency and the information of the callers, representing the degree of intimacy or frequency of communication between users. Based on the constructed call relationship graph, social network feature data, such as node degree centrality and betweenness centrality, are extracted for subsequent user classification and personalized identification schemes.
[0167] Step S44: Based on user call pattern feature data and social network feature data, perform user classification using a clustering algorithm to obtain user profile data;
[0168] This invention combines user call pattern feature data and social network feature data to form a user feature vector. Clustering algorithms (such as K-means, hierarchical clustering, etc.) are used to cluster the user feature vectors, grouping similar users into the same category. Based on the clustering results, a user classification label is assigned to each user for subsequent abnormal behavior prediction and personalized identification schemes.
[0169] Step S45: Based on user profile data and adaptive recognition model, perform abnormal behavior prediction to obtain abnormal behavior prediction data;
[0170] This invention establishes an abnormal behavior prediction model based on user profile data. Machine learning algorithms (such as support vector machines, decision trees, etc.) can be used to train the model to predict normal user behavior patterns. An adaptive recognition model is used, combined with the user's personalized features and historical call data, to predict and judge the user's current call behavior, determine whether abnormal behavior exists, and obtain abnormal behavior prediction data based on the prediction results to identify whether the user has abnormal behavior.
[0171] Step S46: Optimize the identification threshold based on the abnormal behavior prediction data using machine learning technology to obtain personalized identification scheme data.
[0172] This invention employs machine learning techniques, such as classification algorithms (logistic regression, support vector machines, etc.) or clustering algorithms (K-means, DBSCAN, etc.), to process abnormal behavior prediction data. Training and testing datasets are used for model training and evaluation, and appropriate evaluation metrics (such as accuracy, recall, F1 score, etc.) are selected to assess model performance. The parameters of the classifier or clusterer are adjusted, and the model is optimized through methods such as cross-validation to achieve better performance. Based on the model's evaluation results, the optimal identification threshold or clustering threshold is selected to classify users into normal and abnormal behaviors. Personalized identification scheme data, including the selected optimal threshold and corresponding model parameters, is saved for subsequent abnormal behavior identification and application.
[0173] This invention effectively protects users' personal privacy information through anonymization, avoiding the risks of sensitive data leakage and privacy infringement; it ensures the security of historical call data during processing and storage, complying with data protection regulations and standards, and enhancing the credibility and transparency of data management. Through time series analysis, it extracts users' call pattern characteristics, including call frequency and time-of-day preferences, to gain a deeper understanding of users' behavioral habits and communication patterns; based on call pattern characteristic data, it personalizes services, such as optimizing communication plans and providing customized recommendations, enhancing user experience and satisfaction. It constructs users' social network graphs, analyzes call partners and their frequency, revealing the density and activity of users' relationships within their social circles; through social network characteristic data, it identifies and understands users' influence and key roles in social networks, providing data support for social interaction. Combining user profile data and adaptive recognition models, it effectively predicts abnormal patterns and events in user communication behavior, promptly identifying and responding to potential risks. It improves anti-fraud capabilities, reduces potential damage to user and system security caused by abnormal behavior, and protects communication security and user interests. By leveraging machine learning techniques to dynamically adjust identification thresholds based on abnormal behavior prediction data, personalized anti-fraud identification solutions are achieved. Optimizing these thresholds effectively improves the accuracy and real-time response capabilities of the anti-fraud system, while reducing false alarm and false negative rates. In summary, these steps combine data processing, analysis, and prediction technologies, providing multi-dimensional data support and analytical capabilities for anti-fraud communication methods, effectively enhancing system security, user experience, and operational efficiency.
[0174] Preferably, step S5 includes the following steps:
[0175] Step S51: Obtain ambient noise data and network quality data for the call;
[0176] This invention allows the use of specialized noise sensors or microphone equipment to measure noise levels in a call environment. Noise data can be expressed as decibels (dB) or other noise metrics. Network performance monitoring tools or network quality measurement equipment can be used to measure network latency, packet loss rate, bandwidth, and other metrics during a call. This data can be used to evaluate network stability and communication quality.
[0177] Step S52: Evaluate the call quality based on the call environment noise data and network quality data to obtain a call quality model. The call quality evaluation includes audio clarity evaluation, continuity evaluation, and echo cancellation effect evaluation.
[0178] This invention utilizes ambient noise data from calls to analyze the signal-to-noise ratio (SNR) of audio signals. A lower SNR value typically indicates poor audio clarity. Signal processing techniques (such as filtering and noise reduction algorithms) can be used to enhance and denoise the audio signal to improve clarity. Call continuity is assessed based on network quality data, including latency and packet loss rate. Higher latency or packet loss rates can lead to call interruptions or audio discontinuity; network optimization techniques (such as congestion control and traffic scheduling) can be employed to improve call continuity. The effectiveness of echo cancellation algorithms is evaluated by analyzing echo signals during calls and audio signals captured by the microphone. Echo cancellation technology can be used to eliminate echoes during calls, improving call quality.
[0179] Step S53: Collect scenario call data in various real-world scenarios, and use transfer learning techniques to optimize the call quality model's generalization ability based on the scenario call data, thereby obtaining optimized call quality data.
[0180] This invention demonstrates calls in various real-world scenarios, covering different environmental noise levels and network quality conditions. Call data is collected, including ambient noise data, network quality data, and call quality assessment results.
[0181] By leveraging transfer learning techniques, a previously trained call quality model is combined with newly collected scenario call data to optimize and improve its generalization ability. Methods such as knowledge distillation and domain adaptation in transfer learning can be used to transfer knowledge from the previous model to new scenarios, improving the model's performance in these new environments. The optimized call quality model is then used to predict and evaluate call quality optimization data, providing more accurate guidance and suggestions for call quality optimization.
[0182] This invention, by acquiring ambient noise and network quality data, enables the system to perceive the noise level and network conditions of the real-time communication environment, facilitating the adjustment of communication parameters to optimize call quality. It provides fundamental data support, offering necessary input information for subsequent call quality assessment and optimization. Based on the acquired environmental data, the system can accurately assess call quality, including key indicators such as audio clarity, call continuity, and echo cancellation effectiveness. The assessment results can be used to adjust communication parameters in real time, optimizing the call experience and ensuring good call quality for users under various environmental conditions. By collecting call data in multiple real-world scenarios and applying transfer learning techniques, the generalization ability of the call quality model is optimized, enabling it to adapt to different communication scenarios and environmental conditions. This ensures the effectiveness and stability of call quality optimization strategies in various real-world scenarios, improving the system's reliability and practicality in real-world applications. Continuous collection and analysis of scenario call data allows for continuous optimization of the call quality model, maintaining the system's adaptability to changes in the communication environment and providing stable call service quality. In summary, these steps integrate data acquisition, analysis, and model optimization technologies, providing crucial call quality assessment and optimization capabilities for anti-fraud communication methods, thereby significantly improving the system's user experience and operational efficiency.
[0183] Preferably, step S6 includes the following steps:
[0184] Step S61: Perform data cleaning and standardization on the personalized identification scheme data and call quality optimization data to obtain personalized identification standard data and call quality standard data;
[0185] This invention cleans personalized identification scheme data and call quality optimization data, including handling missing values, outliers, and duplicate values. Data cleaning tools or data processing libraries in programming languages (such as Python) (such as pandas) can be used for data cleaning. The cleaned data is then standardized to convert it to a uniform scale and range to eliminate dimensional differences between different data points. Common standardization methods include mean-variance standardization and maximum-minimum standardization.
[0186] Step S62: Extract timbre features, emotional features, and environmental noise features from the personalized identification standard data and the call quality standard data, and use the analytic hierarchy process (AHP) to assign feature weights, thereby obtaining feature weight assignment data.
[0187] This invention employs audio processing technology to extract timbre features from personalized recognition standard data and call quality standard data. Common timbre features include pitch, audio spectrum features (such as frequency, energy, spectral centroid, etc.), and time-domain and frequency-domain features of sound. Sentiment analysis technology is used to extract emotional features from the personalized recognition standard data and call quality standard data. These emotional features can include characteristics such as the emotional tone, speech rate, and intonation of the voice. Environmental noise-related features are also extracted from the personalized recognition standard data and call quality standard data. Noise analysis algorithms or noise recognition techniques can be used to extract features such as noise level and noise spectrum.
[0188] Step S63: Input the feature weight allocation data, personalized identification standard data, and call quality standard data into the preset multi-criteria decision model to obtain multi-criteria decision data;
[0189] This invention, based on specific needs and problems, selects appropriate multi-criteria decision-making methods, such as weighted summation, TOPSIS (Technique for Order Preference by Similarity to Ideal Solution), and fuzzy comprehensive evaluation. A multi-criteria decision-making model is constructed by assigning weights to feature weighted data, personalized identification standard data, and call quality standard data. The feature weighted data, personalized identification standard data, and call quality standard data are input into the multi-criteria decision-making model for calculation and evaluation. Based on the model's calculation rules and weights, multi-criteria decision data is obtained for subsequent analysis and judgment.
[0190] Step S64: Identify potential abnormal behaviors and fraudulent activities in the multi-criteria decision data, and conduct risk assessment and threshold adjustment to obtain the final anti-fraud identification scheme data.
[0191] This invention employs machine learning techniques or rule engines to identify potential abnormal and fraudulent behaviors in multi-criteria decision data. Supervised learning can be performed based on known abnormal behavior and fraud patterns, or unsupervised learning techniques can be used to detect abnormal patterns in the data. A risk assessment is then conducted based on the identified abnormal and fraudulent behaviors. The assessment method may include calculating a risk score or probability, evaluating the risk level based on the severity of the abnormal behavior and the likelihood of fraud. Based on the risk assessment results, the thresholds used to judge abnormal and fraudulent behaviors are adjusted. Depending on the actual situation and needs, the thresholds are changed to balance the risks of false positives and false negatives, achieving the best anti-fraud effect. Based on the results after risk assessment and threshold adjustment, the final anti-fraud identification scheme data is obtained, including identifiers and related information of those identified as abnormal or fraudulent.
[0192] This invention ensures the consistency and accuracy of personalized identification scheme data and call quality optimization data through data cleaning and standardization, improving the reliability of subsequent analysis and decision-making. Cleaning and standardization remove noise and errors from the data, making subsequent data analysis and decision-making more accurate and reliable. Processing data according to unified standards facilitates comparison and integration between different data sources, enhancing the operability and application value of the data. By extracting timbre features, emotional features, and environmental noise features, the key features of call quality and personalized identification schemes are analyzed in depth, providing strong support for subsequent decision-making. Using the analytic hierarchy process (AHP) for feature weight allocation allows for an objective assessment of the contribution of each feature to the final decision, improving the scientific nature and accuracy of the decision. Comprehensive consideration and weighing of different types of data helps establish a comprehensive decision-making model, improving the system's decision-making efficiency and accuracy. Through a multi-criteria decision-making model that comprehensively considers multiple factors such as feature weights, personalized identification standards, and call quality standards, the anti-fraud identification scheme can be more comprehensively evaluated and optimized. Based on a scientific decision-making model, subjective bias can be effectively reduced, improving the objectivity and credibility of the decision. Through multi-criteria decision-making, the system can more accurately identify potential anomalies and fraudulent behaviors, taking timely preventative and responsive measures to protect communication security and user rights. Identifying abnormal and fraudulent behaviors based on multi-criteria decision data effectively enhances the system's ability to perceive and respond to potential risks. In-depth assessment and analysis of identified abnormal behaviors quantifies risk levels, providing a basis for developing targeted preventative measures. The thresholds of the fraud prevention scheme are dynamically adjusted based on real-time risk assessment results, ensuring the system's stable and efficient operation under different scenarios. In summary, these steps effectively integrate data analysis, decision models, and risk management technologies, providing comprehensive data support and a scientific basis for decision-making in fraud prevention communication methods.
[0193] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0194] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A fraud prevention communication method based on voice changer recognition, characterized in that, Includes the following steps: Step S1: Obtain multi-dimensional voice feature data of users, including timbre, frequency features, language rhythm, emotional features and pause patterns; predict the possibility of voice change based on multi-dimensional voice feature data to obtain voiceprint feature prediction data; train a deep learning model based on voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data. Step S2: Simulate and analyze the voice-changing recognition behavior under different call states based on the voiceprint feature map data to obtain voice-changing recognition simulation data; establish a dynamic recognition threshold model based on the environmental parameters corresponding to different call states and the voice-changing recognition simulation data. Step S3: Adjust the control parameters in real time according to the preset target recognition accuracy and dynamic recognition threshold model to obtain control parameter adjustment data; construct a multi-level feedback control system based on the control parameter adjustment data to obtain an adaptive recognition model; Step S4: Obtain user historical call data, perform personalized analysis on the user historical call data, and extract features to obtain user profile data; Abnormal behavior is predicted based on user profile data and adaptive recognition models, and the recognition threshold is optimized through machine learning technology to obtain personalized recognition solution data. Step S5: Obtain call environment noise data and network quality data, and construct a call quality model based on the call environment noise data and network quality data; The call quality model is optimized across scenarios, and communication parameters are adjusted in real time to obtain call quality optimization data. Step S6: Perform multi-criteria decision fusion analysis on the personalized identification scheme data and call quality optimization data to obtain the final anti-fraud identification scheme data.
2. The anti-fraud communication method based on voice changer recognition according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain the user's historical voice data and perform a short-time Fourier transform on the historical voice data to obtain the voice time-frequency data; Step S12: Extract Mel frequency cepstral coefficients from the speech time-frequency data to obtain timbre feature data; extract frequency features based on fundamental frequency contours from the speech time-frequency data to obtain frequency feature data; Step S13: Use a speech activity detection algorithm to identify speech and non-speech segments in the speech time-frequency data to obtain language rhythm feature data; Step S14: Analyze the duration of speech segments and the pause patterns of non-speech segments based on the language rhythm feature data to obtain pause pattern data. Step S15: Analyze the pitch variation, energy distribution, and speech rate of the speech time-frequency data, and extract emotional features to obtain emotional feature data; perform harmonic-noise ratio analysis on the speech time-frequency data, and evaluate the harmonic structure of the speech to obtain harmonic structure data; Step S16: Merge the timbre feature data, frequency feature data, language rhythm feature data, pause pattern data, and emotional feature data into multi-dimensional speech feature data based on the harmonic structure data; Step S17: Use support vector machines to classify multi-dimensional speech feature data and score the probability of voice change to obtain voiceprint feature prediction data. Step S18: Train a deep learning model based on the voiceprint feature prediction data, and map the voiceprint state and features under different call conditions to obtain voiceprint feature map data.
3. The anti-fraud communication method based on voice-changing recognition according to claim 2, characterized in that, Step S18 includes the following steps: Step S181: Add noise and time scaling to the voiceprint feature prediction data, and perform normalization processing to obtain the voiceprint feature training dataset. Step S182: Train the preset deep learning model based on the voiceprint feature training dataset, and adjust the parameters of the model using the cross-validation method to obtain the voiceprint feature model. The deep learning model includes a multi-layer convolutional neural network structure for extracting local features, a long short-term memory network layer for capturing temporal dependencies, an attention mechanism layer for highlighting important features, and a fully connected layer. Step S183: Simulate different call conditions to obtain conditional simulation data; apply the voiceprint feature model to the conditional simulation data to obtain conditional response data; Step S184: Map the voiceprint status and features under different call conditions based on the conditional response data and conditional simulation data to obtain voiceprint feature map data.
4. The anti-fraud communication method based on voice changer recognition according to claim 3, characterized in that, Step S2 includes the following steps: Step S21: Perform statistical analysis on the feature distribution, clustering, and outliers of the voiceprint feature map data to obtain voiceprint feature statistics. Step S22: Based on the voiceprint feature statistics, simulate and analyze the voice change recognition behavior of the voiceprint feature map data to obtain voice change recognition simulation data; Step S23: Evaluate the recognition performance of the voice changer recognition simulation data and calculate the recognition accuracy, false alarm rate and false negative rate under different conditions to obtain the recognition performance evaluation data; Step S24: Based on the identification performance evaluation data, conduct performance impact analysis on different environmental parameters and establish a relationship model between environmental parameters and identification performance to obtain the environmental impact model; Step S25: Set initial recognition thresholds for different call states based on the environmental impact model and recognition performance evaluation data, adjust the thresholds according to real-time environmental parameters, train the machine learning model, and thus obtain a dynamic recognition threshold model.
5. The anti-fraud communication method based on voice changer recognition according to claim 4, characterized in that, Step S3 includes the following steps: Step S31: Extract the key control parameters and their influence range from the dynamic identification threshold model to obtain the control parameter influence matrix; Step S32: Set the initial control parameter values according to the preset target recognition accuracy and control parameter influence matrix to obtain the initial control parameter set; Step S33: Perform parameter sensitivity testing on the initial control parameter set based on the degree of influence of each control parameter on the recognition accuracy, thereby obtaining parameter sensitivity data; use the gradient descent method to adaptively adjust the parameters based on the parameter sensitivity data, thereby obtaining control parameter adjustment data; Step S34: Based on the control parameters, adjust the data and monitor the accuracy of recognition, processing delay and resource consumption in real time to design a multi-level feedback control loop, thereby obtaining a multi-level feedback control system. The multi-level feedback control loop includes two levels: fast response and long-term optimization. Step S35: Run the multi-level feedback control system on the simulation test platform, collect system response data, and perform control effect analysis to obtain control system performance data; Step S36: Based on the performance data of the control system, the multi-level feedback control system is encapsulated into a deployable adaptive recognition model, thereby obtaining the adaptive recognition model.
6. The anti-fraud communication method based on voice changer recognition according to claim 5, characterized in that, Step S4 includes the following steps: Step S41: Obtain user historical call data and perform anonymization processing on the user historical call data to obtain anonymized historical call data; Step S42: Perform time series analysis on the de-identified historical call data and extract user call pattern features to obtain user call pattern feature data; Step S43: Construct a social network graph of users based on call subjects and frequencies using anonymized historical call data to obtain social network feature data; Step S44: Based on user call pattern feature data and social network feature data, perform user classification using a clustering algorithm to obtain user profile data; Step S45: Based on user profile data and adaptive recognition model, perform abnormal behavior prediction to obtain abnormal behavior prediction data; Step S46: Optimize the identification threshold based on the abnormal behavior prediction data using machine learning technology to obtain personalized identification scheme data.
7. The anti-fraud communication method based on voice changer recognition according to claim 6, characterized in that, Step S5 includes the following steps: Step S51: Obtain ambient noise data and network quality data for the call; Step S52: Evaluate the call quality based on the call environment noise data and network quality data to obtain a call quality model. The call quality evaluation includes audio clarity evaluation, continuity evaluation, and echo cancellation effect evaluation. Step S53: Collect scenario call data in various real-world scenarios, and use transfer learning techniques to optimize the call quality model's generalization ability based on the scenario call data, thereby obtaining optimized call quality data.
8. The anti-fraud communication method based on voice changer recognition according to claim 7, characterized in that, Step S6 includes the following steps: Step S61: Perform data cleaning and standardization on the personalized identification scheme data and call quality optimization data to obtain personalized identification standard data and call quality standard data; Step S62: Extract timbre features, emotional features, and environmental noise features from the personalized identification standard data and the call quality standard data, and use the analytic hierarchy process (AHP) to assign feature weights, thereby obtaining feature weight assignment data. Step S63: Input the feature weight allocation data, personalized identification standard data, and call quality standard data into the preset multi-criteria decision model to obtain multi-criteria decision data; Step S64: Identify potential abnormal behaviors and fraudulent activities in the multi-criteria decision data, and conduct risk assessment and threshold adjustment to obtain the final anti-fraud identification scheme data.
Citation Information
Patent Citations
Anti-fraud system based on robust voiceprint
CN118116391A
Data update method, client, and electronic device
US20180366128A1