Real-time AI-based spurious voice detection and response system and method in telecommunication networks.

TR202614412A2Pending Publication Date: 2026-09-21TURK TELEKOMUNIKASYON A S
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TR202614412
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-21

Smart Images

  • Figure 00000014_0000
    Figure 00000014_0000
Patent Text Reader

Abstract

It is a system that enables the real-time detection and prevention of AI-assisted voice spoofing, identity impersonation and similar fraud attempts that may occur during voice calls made on telecommunication networks, and its features include a communication network (2), processor (3), deepfake detection module (4), telecommunication risk scoring engine (5), dynamic verification module (6) and decision and action module (7).
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF Real-time AI-based spurious voice detection in telecommunication networks and intervention system and method Technical Area The invention addresses potential issues that may arise during voice calls made over telecommunication networks. AI-powered voice spoofing, identity impersonation, and similar fraudulent attempts are real. An integrated security system that enables timely detection and prevention. 10 It is related to the system and method that presents the mechanism. State of the Art Today, there are 15 different measures aimed at preventing fraud in telecommunication networks. These methods generally include number verification systems, blacklisting, etc. blacklisting mechanisms, filtering systems based on user complaints, spam It is based on call blocking applications and call signaling analysis. However, voice cloning, which has emerged in recent years with the development of artificial intelligence technologies, is 20 (voice cloning), artificial speech generation (text-to-speech), real-time voice conversion (Voice conversion) and deepfake-based fraud methods pose significant risks to existing systems. This leads to it being largely inadequate. A significant portion of existing systems only show which number the call came from or 25 assessing which network the call is being transmitted over, the actual voice transmitted during the call. It does not analyze the content. Therefore, AI-generated or imitated voices... It is not possible to detect it. In addition, existing solutions mostly perform analysis after the call or user 30 They take action after the complaint is filed. Therefore, they act before the attempted fraud takes place. Intervention is not possible before the damage occurs, and detection is only possible after the damage has happened. It is possible. Another shortcoming is that existing systems largely focus on the signaling layer. The focus is on SIP messages, call setup logs, or signaling data. 2 Despite being examined, the audio data of the call was transmitted through the media. This is not being evaluated. Therefore, it is not being determined whether the sound is natural or artificially produced. It cannot be determined. Current solutions also make decisions based on only a single data source. For example, 5 only call numbers, only user complaints, or only blacklists are considered This is taken into account. In return, call behavior, subscriber history, device reliability, SIM changes that evaluate different parameters together, such as network mobility and voice analysis There is no integrated decision-making mechanism. The dynamic risk scoring structure is also used in existing fraud systems at the operator level. They are not sufficiently developed. Systems mostly rely on predefined rules and static thresholds. It operates based on its values ​​and adapts to newly emerging fraud methods. It is unable to provide this. In addition, existing solutions aim to verify the user when a suspicious situation is detected. It does not run a dynamic and automated verification mechanism during the call. This Therefore, making the right decision in high-risk but unconfirmed scenarios is crucial. It is becoming more difficult. In conclusion, current systems are weak in detecting AI-assisted voice spoofing; Generating risk by evaluating data sources together, real-time intervention. It is inadequate in both implementation and dynamic verification during the call. This situation involves a multi-layered risk assessment operating within the telecommunications operator's infrastructure. A new generation of fraud prevention that implements and takes real-time action 25 This reveals the need for the system. The application with registration number US9699660B1 was revealed as a result of technical investigations. Summary: “Techniques for detecting telecommunications fraud; real-time data 30 an analysis of risk models typically used in authentication applications in between, call metadata that is continuously streamed to a database server the implementation and as the database server receives phone usage data, phone usage This involves deriving patterns. The database server then transmits them in a stream. In order to detect SIM box or SIM cloning fraud within the data, 35 derived comparing phone usage patterns with patterns of phone use for fraudulent purposes. 3 It compares. Within a large set of numerous calls, this type of comparison... Comparison results showing the likelihood of fraud in authentication applications in the form of a risk score derived using typically available risk models It is possible." As can be seen, the invention relates to the detection of telecommunications fraud, and in addition to this... a structure that can provide a solution to the disadvantages mentioned above It does not mention it. In conclusion, due to the negative aspects described above and the current solutions, topic 10 Due to its shortcomings, it has become necessary to make improvements in the relevant technical field. Purpose of the Invention The invention represents a new breakthrough in this field, unlike the structures used in the current technology. The aim is to create a structure with different technical specifications that bring these elements together. The primary purpose of the invention is to improve communication during voice calls made over telecommunication networks. potential AI-powered voice spoofing, identity impersonation, and similar scams an integrated system that enables the real-time detection and prevention of such attempts. The goal is to develop systems and methods that provide a security mechanism. Another aim of the invention is to transmit not only call signaling but also the sound transmitted during a call. It is also able to make assessments by analyzing the data. Thus, the existing systems can be identified. AI-generated voices that it struggles to produce, voice cloning attempts, and deepfake 25 Fraud scenarios based on this tactic can be identified with a higher accuracy rate. Unlike existing techniques, this invention does not rely on a single data source; voice analysis results, call behavior, subscriber history, device reliability information, SIM By evaluating the changes and network security parameters together, a multi-layered 30 Telekom creates a Risk Score. This enables more accurate and reliable decisions. It can be obtained. Current solutions mostly involve analyzing the situation after the call or responding to user complaints. It operates in conjunction with the invention, providing real-time analysis while the call is ongoing. 35 4 They are able to carry out the operation and intervene before the fraudulent attempt is completed. In this way, users can be protected from suffering financial and emotional harm. Another advantage of the invention is its ability to take dynamic action depending on the level of risk. Based on the generated risk score, the system can decide whether to allow the call to continue normally or not. can send security alerts to the user, initiate additional verification processes, or In high-risk situations, it can limit or terminate the call. This structure This provides a flexible and adaptive safety mechanism that is not bound by static rules. is being done. The invention also operates centrally within the operator network, therefore on the user side. It does not require any additional applications or hardware. Thus, it is available to all subscribers. Standard and widespread protection can be provided, security solutions by the user Furthermore, it does not require establishment or management. Thanks to the dynamic verification mechanism, the system provides additional verification in suspicious cases. By initiating these processes, it can reduce false alarm rates. This feature benefits both the user. It both improves the user experience and increases the level of security. Thanks to the invention's AI-based and learnable nature, it can detect newly emerging frauds. The methods can be adapted, the system can improve itself over time, and It can maintain its effectiveness in the face of the changing threat environment. In conclusion, this invention enables sound analysis in the media landscape and addresses multi-layered risk management. using a scoring mechanism, performing real-time intervention, dynamic 25 thanks to its ability to support verification processes and integrate with operator infrastructure. It offers significant advantages over existing technologies and introduces new technologies to telecommunication networks. This offers more effective protection against next-generation fraud attempts. The invention relates to communication in the telecommunications sector, specifically IP-based voice communication. Fraud that may occur during voice calls made on their networks It is used to detect and prevent unauthorized access to 4G and 5G technologies. Mobile communication systems, including IMS (IP Multimedia Subsystem) based voice. services, VoLTE (Voice over LTE) and VoNR (Voice over New Radio) calls, enterprise VoIP infrastructures, international voice interconnect networks, and operator fraud management 35 It can be implemented in systems. The invention also enables AI-powered voice imitation (deepfake), identity theft, and similar audio-related activities. real-time detection and prevention of fraudulent activities It is used in the field of telecommunications security systems. In this context, the invention; mobile operators, fixed-line providers, virtual mobile network operators, and cloud-based communications 5 It is suitable for use by service providers. To fulfill the purposes described above, the invention is used in telecommunication networks. AI-assisted voice spoofing that may occur during voice calls, Real-time detection of identity theft and similar fraud attempts and 10 It is a system that enables prevention, and its feature is;  The voice call initiated by the scammer was transmitted via mobile voice service. the communication network representing the operator infrastructure,  Voice flow and call-related signaling passing through the communication network The information is routed, and the incoming voice call is processed through specific control steps. 15 A processor that detects anomalies, fraud, or risky content.  It runs on the processor, analyzing the audio stream using artificial intelligence algorithms. Deepfake, voice cloning, artificial speech generation, and voice anomaly detection; extracts acoustic and behavioral characteristics from voice data transmitted during a call, Sound 20 using pre-trained machine learning and deep learning models a deepfake that classifies and analyzes a sample and generates an Audio Reliability Score. capture module,  The processor processes voice analysis results, call behavior, and subscriber history. SIM changes are combined with device reliability and network security parameters. It brings in scores that combine with specific weighting coefficients and awards 25 for each call. It calculates a dynamic Telecom Risk Score, and this score is discussed during the negotiations. The telecommunications risk scoring engine, which is constantly updated with new data as it develops,  Results generated by the telecommunications risk scoring engine running on the processor Telekom Risk is transferred and if the determined threshold values ​​are exceeded. Voice verification, 30, initiates additional verification processes based on the risk level in its score. Dynamic systems that perform behavioral analysis and real-time security checks. verification module,  The Telecommunications Risk Score and verification results are running on the processor. Evaluation and continuation of the call, warning given to the user, additional verification. 35 decisions regarding the implementation or limitation / termination of the call are made automatically. decision and action module that enables its implementation 6 It includes. The structural and characteristic features and all the advantages of the invention are given in the figures below. Thanks to the detailed explanation written with references to the figures, it becomes clearer. This will be understood, and therefore the evaluation should also take these figures and detailed explanations into account. 5 It must be done by taking it. Figures that will help understand the invention. Figure 1 shows a general representation of the system that is the subject of the invention. 10 The drawings do not necessarily need to be scaled and are necessary for understanding the invention. Details that are not present may have been overlooked. Furthermore, at least to a large extent... Elements that are identical or at least have substantially identical functions are numbered the same. It is shown. 15 Explanation of Part References 1. Scammer 2. Communication network 20 3. Processor 4. Deepfake detection module 5. Telecommunications risk scoring engine 6. Dynamic verification module 7. Decision and action module 25 8. Customer Detailed Description of the Invention In this detailed explanation, the preferred configurations of the invention are only 30% better understood in the subject. in order to facilitate understanding and without imposing any limiting effects It is explained. The invention addresses potential issues that may arise during voice calls made over telecommunication networks. AI-powered voice spoofing, identity impersonation, and similar fraudulent attempts are real 35 7 an integrated security system that enables timely detection and prevention It is related to the system and method that presents the mechanism. The elements and functions used in the system and method that are the subject of the invention are as follows: The scammer (1) tries to convince the customer (8) to benefit from various issues. is a person. Communication network (2) represents the operator infrastructure where mobile voice service is provided. It is an element. 10 The processor (3) passes the incoming voice call through certain control steps to detect anomalies, fraud or the unit that identifies risky content. Deepfake capture module (4) runs on processor (3) and converts the audio stream to artificial intelligence 15 By analyzing with algorithms, deepfake, voice cloning, artificial speech generation and voice It is an AI-powered module that detects anomalies. The telecom risk scoring engine (5) runs on processor (3) and the voice analysis results, call behavior, subscriber history, SIM changes, device reliability, and network 20 It is the unit that creates a dynamic Telecommunications Risk Score by combining security parameters. Dynamic verification module (6) runs on processor (3) and the specified Telecom Risk Based on the risk level in your score, you can initiate additional verification processes and voice verification. By performing behavioral analysis and real-time security checks, we reduced fraudulent call rates by 25%. It is the reducing module. The decision and action module (7) runs on the processor (3) and the Telecom Risk Score Evaluation and continuation of the call, warning given to the user, additional verification. Decisions regarding the implementation or limitation / termination of the call are automatically made on the 30th. It is the module that enables its implementation. Customer (8) is a mobile operator subscriber targeted by scammers (1). The invention works on the following principle: 35 8 Real-time voice calls made within the telecommunications operator's infrastructure. the analysis involves combining the call's voice data with network security parameters. based on evaluation and taking automated action according to the results obtained It is based on. The system was initially the communication network (2) of the voice call initiated by the fraudster (1) It starts working by transmitting the information through the system. Communication begins after call setup. voice stream and call signaling information passing through the network (2), processor (3) It is directed to the analysis components contained within. In the first stage, the audio stream is captured by the AI-powered Deepfake capture module (4). This module processes various acoustic and behavioral data from the voice data transmitted during a call. It extracts features such as the frequency distribution, harmonic structure, and speech of the sound during analysis. rhythm, pauses, energy levels, spectral coherence, formant transitions, and human Natural variations specific to the voice are being evaluated. 15 Deepfake capture module (4), pre-trained machine learning and deep learning It classifies the sound sample using models. In this context, the sound is classified as the natural human voice. Is it a voice produced using voice cloning methods, or an artificial text-to-speech conversion? whether it is a speech or an audio file created using real-time voice conversion methods 20 The aim is to determine this. As a result of the analysis, the system will assign a Sound Reliability Score. It is manufactured by SGS. The voice analysis results obtained later were used by the telecom risk scoring engine (5) It is evaluated together with other security parameters. Telecom risk scoring engine 25 (5) does not only use voice analysis results; it also uses call behavior information, subscriber history, SIM card changes, device reliability information, call frequency, short-term the number of searches conducted within the network, records of previously received complaints, and network security. It also takes these events into account. Within this scope, the system;  Sound Reliability Score,  Call Behavior Score,  Subscriber Trust Score,  Device Safety Score, 35  Network Trust Score 9 These scores are generated and combined using specific weighting coefficients. Thus A dynamic Telecommunications Risk Score is calculated for each call. The Telecommunications Risk Score does not remain constant during the call; new data is generated as the conversation progresses. It is constantly updated using [method]. Thus, 5 that appears low risk at the beginning of the call However, as the interview progressed, scenarios exhibiting suspicious behavior were also identified. It is possible. The results produced by the Risk Scoring Engine (5) are sent to the dynamic validation module (6) is being transmitted. If the system exceeds the defined threshold values, additional verification will be performed. 10 This initiates the processes. At this stage, automatic verification is requested from the user or the caller. Questions can be posed, the phonetic characteristics of the answers given can be re-analyzed, and the sound... The consistency of this example with previous examples is being evaluated. This reduces the likelihood of false alarms. While reducing the use of fake voices, it also ensures accurate detection. All analysis results obtained in the final stage are processed by the decision and action module (7). This module is being evaluated based on the Telecommunications Risk Score and verification results. allowing the call to continue normally, the customer is given a security warning (8). can send, apply additional verification, or block the call in high-risk situations. It can limit or terminate it. 20 This invention thus prevents existing systems from making decisions based solely on call numbers or blacklists. Unlike other systems, it analyzes voice content, call behavior, subscriber security data, and By evaluating network parameters together, we perform a multi-layered analysis and It can intervene in fraud attempts in real time. 25 The steps involved in the process carried out with the system that is the subject of the invention are listed below:  The voice call initiated by the scammer (1) over the communication network (2) transmission,  After call setup, the voice stream passing through the communication network (2) and 30 signaling information regarding the call is entered into the analysis components within the processor (3). guidance,  The routed audio stream runs on the deepfake capture module (4) which is running on the processor (3) processed by,  Acoustic data is extracted from the audio data carried during the call by the Deepfake capture module (4). and the extraction of behavioral characteristics,  Deepfake capture module (4) by pre-trained machine learning and deepfake Classification of audio samples using learning models,  Generating the Sound Reliability Score (SGS) as a result of the analysis,  The obtained voice analysis results are compared with other telecommunications risk scoring engines (5) Evaluation and scoring together with safety parameters, 5  Combining the generated scores using specific weighting coefficients and for each call Calculating a dynamic Telecommunications Risk Score,  As the discussion continues, the Telecommunications Risk Score will be continuously updated using new data. updating,  Dynamic validation of the results produced by the telecom risk scoring engine (5) 10 transfer to module (6),  If the specified threshold values ​​are exceeded, the dynamic verification module (6) will perform the verification. initiating additional verification processes,  All the analysis results obtained are processed by the decision and action module (7) evaluation, 15  Decision and action module (7) by the Telecom Risk Score and verification depending on the results, the call may be allowed to continue as normal. sending a security alert to the customer (8), applying additional verification or the call Automatic implementation of decisions to restrict or terminate [these measures].

Claims

11 REQUESTS 1. Issues that may arise during voice calls made on telecommunication networks. AI-powered voice spoofing, identity impersonation, and similar fraudulent attempts It is a system that enables real-time detection and prevention, and its feature is; 5  Mobile voice service through which the voice call initiated by the scammer (1) was transmitted communication network (2) which represents the operator infrastructure to which it is given  Voice flow and call signaling passing through the communication network (2) the information is routed, and the incoming voice call is passed through certain control steps. Processor that detects anomalies, fraud or risky content (3), 10  By analyzing the sound stream with artificial intelligence algorithms, which run on the processor (3) Deepfake, voice cloning, artificial speech generation, and voice anomaly detection; extracts acoustic and behavioral characteristics from voice data transmitted during a call, voice using pre-trained machine learning and deep learning models deepfake 15 classifies the sample and produces an Audio Reliability Score as a result of the analysis. capture module (4),  The processor (3) runs on the voice analysis results, call behaviors, subscriber history, SIM changes are combined with device reliability and network security parameters. It retrieves the generated scores, combines them with specific weighting coefficients, and for each call... A dynamic Telecom Risk Score is calculated, and the score will be discussed further in the 20... The telecom risk scoring engine (5), which is constantly updated with new data as it develops,  Produced by the telecommunications risk scoring engine (5) running on the processor (3). The results are reported and if the determined threshold values ​​are exceeded, Telekom Risk Voice verification, which initiates additional verification processes based on the risk level in its score, Dynamic 25 that performs behavioral analysis and real-time security checks verification module (6),  The Telecom Risk Score and verification results running on the processor (3). Evaluation and continuation of the call, warning given to the user, additional verification. automatic decisions on implementation or limitation / termination of the call Decision and action module (7) 30 that enables its implementation It includes.

2. The system complies with Claim 1 and its features include: Voice Reliability Score, Call Behavior Score, The Subscriber Trust Score, Device Trust Score, and Network Trust Score are all determined by telecommunications companies. It includes a risk scoring engine (5). 35 12 3. The system complies with Claim 1 and its feature is; during the analysis, it analyzes the frequency distribution and harmonics of the sound. its structure, speech rhythm, pause durations, energy levels, spectral coherence, to evaluate formant transitions and natural variations specific to the human voice It includes a structured deepfake capture module (4).

4. The system complies with Claim 1 and its feature is automatic verification for the user or caller. It includes a dynamic validation module (6) structured to ask questions.

5. The system compliant with Claim 1, and its feature is that it is the mobile operator targeted by fraudsters (1). The decision is configured to send a security alert to the customer (8) who is a subscriber and 10 It includes action module (7).

6. Issues that may arise during voice calls made over telecommunication networks. AI-powered voice spoofing, identity impersonation, and similar fraudulent attempts It is a method that enables real-time detection and prevention, and its feature is; 15  The voice call initiated by the scammer (1) over the communication network (2) transmission,  After call setup, the voice stream passing through the communication network (2) and The signaling information regarding the call is analyzed within the processor (3). redirection to its components, 20  The routed audio stream runs on the deepfake capture module (3) (4) by processing,  Deepfake capture module (4) from the audio data carried during the call Extraction of acoustic and behavioral characteristics,  Deepfake capture module (4) by pre-trained machine learning and 25 Classification of audio samples using deep learning models,  Generating the Sound Reliability Score (SGS) as a result of the analysis,  The obtained voice analysis results are evaluated by the telecommunications risk scoring engine (5) Evaluation and scoring together with other security parameters,  The generated scores are combined using specific weighting coefficients, and each call is 30 Calculating a dynamic Telecommunications Risk Score,  As the discussion continues, the Telecommunications Risk Score will be continuously updated using new data. updating,  Dynamic validation of the results produced by the telecom risk scoring engine (5) Transfer to module (6), 35 13  If the specified threshold values ​​are exceeded, the dynamic verification module (6) initiation of additional verification processes by,  All the analysis results obtained are processed by the decision and action module (7) evaluation,  Decision and action module (7) by, Telecom Risk Score and verification 5 depending on the results, the call may be allowed to continue as normal. sending a security alert to the customer (8), applying additional verification or the call automatic implementation of decisions to restrict or terminate It includes the steps of the process.

7. The method is in accordance with Claim 6 and its feature is; voice by telecommunications risk scoring engine (5). Reliability Score, Call Behavior Score, Subscriber Trust Score, Device Trust Score and This involves establishing a Network Trust Score.

8. The method is compliant with Request 6, and its feature is that the sound is captured by the Deepfake capture module (4) 15 frequency distribution, harmonic structure, speech rhythm, pause durations, energy levels, spectral coherence, formant transitions, and natural characteristics specific to the human voice It is the evaluation of variations.

9. The method is compliant with Request 6, and its feature is that the sound is captured by the Deepfake capture module (4) 20 Is it a natural human voice, a voice produced using voice cloning, or text-to-speech? is it converted artificial speech or real-time voice conversion method? The goal is to determine whether it is a fabricated sound or not.

10. The method is compliant with claim 6 and its feature is; by the dynamic validation module (6) 25 This involves automatically presenting verification questions to the user or the caller.