Device for recognizing context of voice phishing in real time during call and operating method thereof
The deep learning-based voice phishing detection system addresses inefficiencies in existing methods by processing voice signals locally, extracting features, and using continual learning to detect voice phishing in real-time, enhancing computational efficiency and fraud detection.
Patent Information
- Application Number
- PCT/KR2025/004854
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
Existing voice phishing detection methods are inefficient due to the need for data communication with databases and servers, leading to reduced computational efficiency and difficulty in real-time detection of electronic financial fraud.
A real-time voice phishing detection system utilizing a deep learning-based model that collects voice signals, extracts text and voice features, calculates risk levels, and detects phishing without database communication, employing a multi-modal approach and continual learning.
Enables efficient, real-time detection of voice phishing with improved computational efficiency and robust performance, reducing the risk of financial fraud by identifying suspicious calls and patterns.
Smart Images

Figure KR2025004854_16102025_PF_FP_ABST
Abstract
Description
Device and operating method for recognizing the context of voice phishing in real time during a call
[0001] The present invention relates to a real-time voice phishing detection technology utilizing an artificial intelligence-based model, and more specifically, to a technical idea that enables real-time voice phishing detection by applying voice features and text features to a deep learning model.
[0002] Electronic financial fraud cases leveraging AI and victims' personal information have been on the rise recently, and financial authorities are struggling to prevent electronic financial fraud due to increasingly sophisticated techniques compared to the past.
[0003] This makes it difficult to judge voice phishing, such as electronic financial fraud, and thus, technology is needed to assist users by providing accurate judgment criteria.
[0004] For this reason, technologies are being developed to detect electronic financial fraud in advance and prevent damage.
[0005] Recently, with the advancement of artificial intelligence technology, methods for detecting electronic financial fraud using artificial intelligence technology are being actively studied.
[0006] Representative examples of related prior art include Korean Patent No. 10-2392950.
[0007] Existing voice phishing detection methods have clear limitations in computational efficiency, as they require communication with databases and servers to analyze call content and convert it into text to detect voice phishing.
[0008] Conventional techniques determine the intent of electronic financial fraud by comparing data on electronic financial fraud cases stored in a database and server with data obtained from actual terminals.
[0009] The data extracted in this way is compared with the data in the DB and used for detection, but in this process, computational efficiency is significantly reduced.
[0010] The purpose of the present invention is to detect voice phishing more efficiently and to determine whether a call is a voice phishing call more quickly and in real time.
[0011] The purpose of the present invention is to enable real-time detection of electronic financial fraud by utilizing an AI model that applies a methodology that enables real-time detection and omits data communication processes with a DB and server to improve computational efficiency.
[0012] The present invention aims to provide robust performance in voice phishing detection through a multi-modal method and a deep learning-based algorithm (model).
[0013] A real-time voice phishing recognition device according to one embodiment may include a voice signal collection unit that collects in real time a voice signal transmitted to a communication terminal after a call is established, a text information extraction unit that converts the collected voice signal into text to extract text information, a voice feature information extraction unit that extracts voice feature information for the voice signal in parallel with the extraction of the text information, a risk calculation unit that calculates a risk level based on the extracted text information and the extracted voice feature information using a voice phishing detection model, and a voice phishing detection unit that detects voice phishing based on the calculated risk level.
[0014] The system may further include a detection model learning unit that learns the voice phishing detection model by inputting a public data set generated from voice files of actual damage cases related to voice phishing crimes according to an embodiment.
[0015] According to one embodiment, the detection model learning unit can crawl a script on a web page for the voice file and process performance measurement and model learning through an STT (speech to text) processor.
[0016] The detection model learning unit according to one embodiment can process personal information de-identification and stopword removal processes as the STT (speech to text) process.
[0017] According to an embodiment, the voice phishing detection unit can detect voice phishing by recognizing the context of voice phishing in real time based on the calculated risk level.
[0018] The risk calculation unit according to one embodiment can generate the text information using a pre-learned transfer learning model, generate multimodal data by combining the generated text information and the extracted voice feature information, classify the generated multimodal data using a model applying a Continual Learning learning methodology, and calculate a risk for the classified multimodal data.
[0019] A method of operating a real-time voice phishing recognition device according to an embodiment may include a step of collecting a voice signal transmitted to a communication terminal in real time after a call is established, a step of converting the collected voice signal into text to extract text information, a step of extracting voice feature information for the voice signal in parallel with the extraction of the text information, a step of calculating a risk level based on the extracted text information and the extracted voice feature information using a voice phishing detection model, and a step of detecting voice phishing based on the calculated risk level.
[0020] The method of operating a real-time voice phishing recognition device according to one embodiment may further include a step of learning a model by inputting a public data set generated from a voice file of an actual damage case related to a voice phishing crime, or learning a model through a speech to text (STT) processor by crawling a script on a web page for the voice file.
[0021] The step of calculating a risk level based on the extracted text information and the extracted voice feature information according to one embodiment may include a step of generating the text information using a pre-learned transfer learning model, a step of generating multimodal data by combining the generated text information and the extracted voice feature information, a step of classifying the generated multimodal data using a model applying a Continual Learning learning methodology, and a step of calculating a risk level for the classified multimodal data.
[0022] In one embodiment, voice phishing can be detected more efficiently and whether a call is a voice phishing call can be determined more quickly and in real time.
[0023] According to one embodiment, electronic financial fraud can be detected in real time by utilizing an AI model that applies a methodology that enables real-time detection and omits data communication with the DB and server to improve computational efficiency.
[0024] According to one embodiment, robust performance in voice phishing detection can be provided through a multi-modal approach and a deep learning-based algorithm (model).
[0025] FIG. 1 is a drawing illustrating a real-time voice phishing recognition device according to an embodiment.
[0026] FIG. 2 is a drawing illustrating the entire operation of analyzing text and classifying voice using a real-time voice phishing recognition device according to one embodiment.
[0027] Figure 3 is a diagram explaining the learning process of a model using a pre-learned transfer learning model.
[0028] Figure 4 is a diagram illustrating a training process for detecting voice phishing using learning data and test data.
[0029] Figure 5 is a diagram illustrating an embodiment of detecting voice phishing in an audio data set.
[0030] Figure 6 is a drawing explaining an embodiment of notifying a user terminal of voice phishing.
[0031] FIG. 7 is a diagram illustrating an operation method of a real-time voice phishing recognition device according to an embodiment.
[0032] Specific structural or functional descriptions of embodiments according to the concept of the present invention disclosed in this specification are merely illustrative for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms and are not limited to the embodiments described in this specification.
[0033] Embodiments according to the concept of the present invention may have various modifications and take various forms, and thus, embodiments are illustrated in the drawings and described in detail in this specification. However, this is not intended to limit embodiments according to the concept of the present invention to specific disclosed forms, but rather includes modifications, equivalents, or alternatives that fall within the spirit and technical scope of the present invention.
[0034] While terms such as "first" or "second" may be used to describe various components, these components should not be limited by these terms. These terms are intended solely to distinguish one component from another. For example, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component," without departing from the scope of the invention.
[0035] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Expressions that describe relationships between components, such as "between," "immediately between," or "directly adjacent to," should be interpreted similarly.
[0036] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the present invention. The singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" are intended to specify the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0037] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0038]
[0039] Hereinafter, embodiments will be described in detail with reference to the attached drawings. However, the scope of the patent application is not limited or restricted by these embodiments. The same reference numerals in each drawing represent the same components.
[0040] FIG. 1 is a drawing illustrating a real-time voice phishing recognition device (100) according to one embodiment.
[0041] A real-time voice phishing recognition device (100) according to one embodiment can more efficiently detect voice phishing and more quickly determine whether a call is a voice phishing call in real time. Furthermore, to improve computational efficiency, the device omits data communication with a database and server, utilizes an AI model that applies a methodology that enables real-time detection, and can detect electronic financial fraud in real time. Furthermore, the device can provide robust performance in voice phishing detection through a multi-modal approach and a deep learning-based algorithm (model).
[0042] To this end, a real-time voice phishing recognition device (100) according to one embodiment may include a voice signal collection unit (110), a text information extraction unit (120), a voice feature information extraction unit (130), a risk calculation unit (140), and a voice phishing detection unit (150).
[0043] The voice signal collection unit (110) can collect voice signals transmitted to a communication terminal in real time after a call is established.
[0044] The voice signal collection unit (110) can continuously collect voice signals generated from the communication terminal in real time after a call is connected. The voice signal collection unit (110) can continuously monitor and analyze voice signals generated during a call. The collected voice signals can capture all voice activity occurring during a call and be used for subsequent processing or analysis.
[0045] The text information extraction unit (120) can extract text information by converting the collected voice signal into text.
[0046] The text information extraction unit (120) converts collected voice signals into text. This function extracts the conversation content contained in the voice signals into text format. This conversion from voice signals to text information is accomplished using automatic speech recognition technology, enabling rapid processing of information conveyed during a call. This function helps effectively identify and manage important information or meanings arising from a call.
[0047] The text information extraction unit (120) can convert a voice signal into text using the STT (Speech-to-Text) function.
[0048] Specifically, the text information extraction unit (120) receives the collected voice signal and preprocesses the transmitted voice data to perform processes such as noise removal and signal strengthening.
[0049] Additionally, preprocessed voice data can be input into a deep learning model or an acoustic model to extract each sound unit (phoneme) or feature. Furthermore, the text information extraction unit (120) performs a process of converting the extracted sound units into text, which corresponds to the process of mapping a voice signal into text using a voice recognition engine.
[0050] Next, the text information extraction unit (120) outputs the recognized text.
[0051] The voice feature information extraction unit (130) can extract voice feature information for the voice signal in parallel with the extraction of text information.
[0052] The voice feature information extraction unit (130) analyzes and extracts various features occurring in a voice signal, allowing for mathematical representations of voice characteristics such as frequency, energy, and frequency variation. This voice feature information can be combined with text information to more accurately understand and utilize information derived from the voice signal, such as context. For example, it can be used to determine the emotional content of a voice signal or to verify the identity of a speaker.
[0053] For example, voice feature information can be used to analyze the emotional content of a voice. For example, by analyzing voice characteristics such as speed, pitch, and volume, the speaker's emotions can be identified. This allows for the detection of emotional states occurring during a call.
[0054] Additionally, unique characteristics of a speaker can be extracted from a voice signal to identify that individual. Voice characteristics of a speaker are related to their voice waveform, pronunciation patterns, intonation, and other factors. This can be used to identify the speaker during a call for security purposes or customer service purposes.
[0055] Voiceprints (voice identification) can be generated by identifying individual speakers' voice characteristics. These can be used to identify and authenticate users.
[0056] These technologies can be used to provide enhanced security and personalization capabilities in speech recognition and speech processing systems.
[0057] The risk calculation unit (140) can calculate the risk based on the extracted text information and the extracted voice feature information using a voice phishing detection model.
[0058] The risk calculation unit (140) analyzes information extracted during voice calls to assess the likelihood of fraudulent activity, such as voice phishing, and determine the risk level. The voice phishing detection model analyzes voice and text data to detect predefined patterns or characteristics to identify suspicious calls. This reduces the user's exposure to voice phishing and enhances security.
[0059] For example, voice feature analysis, language analysis, conversation pattern analysis, etc. can be performed using predefined patterns or characteristics.
[0060] Specifically, voice phishing detection models can analyze specific patterns in voice signals. For example, in voice phishing calls, callers often speed up the conversation to create a sense of urgency. In such cases, detecting changes in the speed of the voice signal can identify suspicious calls.
[0061] Additionally, voice phishing detection models can identify suspicious patterns by analyzing the language and phrases used in voice signals. For example, they can detect frequently used voice phishing-related terms or phrases during calls and flag them as suspicious.
[0062] The risk calculation unit (140) can analyze conversation patterns to identify suspicious calls. For example, voice phishing calls tend to have a specific question-and-answer pattern that repeats, and these patterns can be detected to identify suspicious calls.
[0063] The voice phishing detection unit (150) can detect voice phishing based on the calculated risk level. In particular, the voice phishing detection unit (150) can detect voice phishing by recognizing the context of voice phishing in real time based on the calculated risk level.
[0064] For example, thresholds or threshold ranges can be defined in stages, and alarms such as warning, alert, and danger of voice phishing can be provided based on these.
[0065] In addition, the real-time voice phishing recognition device (100) according to one embodiment may further include a detection model learning unit (160).
[0066] According to an embodiment, a detection model learning unit (160) can learn the voice phishing detection model by inputting a public data set generated from voice files of actual damage cases related to voice phishing crimes.
[0067] For example, you can utilize the public data set of actual cases of damage related to voice phishing crimes provided by the Financial Supervisory Service.
[0068] In particular, voices related to habitual crime, loan fraud, and investigative agency impersonation can be utilized as public data sets.
[0069] To this end, the detection model learning unit (160) can crawl the script on the web page for the voice file and utilize it for performance measurement and model learning of the STT Process.
[0070] As another example, AI hub-civil complaint (call center) inquiry / response data can be utilized as a public data set.
[0071] This open data set of phone conversations provided by AI Hub can be leveraged from the financial and insurance domain, a common area of voice phishing. Voices can be categorized into product subscription and cancellation, transfer / withdrawal / loan service-related voices, and more. Labeled JSON data can be used to measure the performance of speech-to-text (STT) processes and train models.
[0072] During the training process, the model can be trained using voice data extracted from actual voice phishing cases. This allows the model to learn various voice phishing patterns that can occur in real-world situations and improve reliable detection capabilities. The voice data can be in Korean. Furthermore, voice data in various languages, such as Chinese and English, can be used, and model training can improve detection capabilities for those languages.
[0073] According to one embodiment, the detection model learning unit (160) can crawl a script on a web page for the voice file and process performance measurement and model learning through an STT (speech to text) processor.
[0074] Web pages may include news articles or forum posts containing audio files of voice phishing incidents.
[0075] The detection model learning unit (160) can retrieve the HTML code of a selected webpage and operate a crawler to extract necessary information. The crawler accesses the webpage via its URL and retrieves the HTML document.
[0076] The detection model learning unit (160) can extract text portions containing script information from crawled HTML documents. This process identifies specific tags or classes in the HTML document and extracts the text contained therein.
[0077] Next, the detection model learning unit (160) can input the extracted text into a STT (Speech to Text) processor to convert speech into text. The STT processor can apply technology to analyze speech signals and convert them into text.
[0078] The detection model training unit (160) uses the converted text to measure the performance of the voice phishing detection model and can update or retrain the model as needed. This allows the model to recognize various patterns occurring in actual scripts and more accurately detect voice phishing.
[0079] The detection model learning unit (160) according to one embodiment is a STT (speech to text) process, and can process personal information de-identification and stopword removal processes.
[0080] The detection model learning unit (160) according to one embodiment can de-identify personally identifiable information from extracted text data for security and privacy protection, and remove meaningless words or phrases to improve the accuracy of the data.
[0081] The text data extracted through the STT process can be identified and de-identified to identify portions containing personal information. For example, personally identifiable information such as resident registration numbers, phone numbers, and addresses can be identified and masked or deleted to enhance security.
[0082] Additionally, data accuracy can be improved by removing meaningless words and phrases from the extracted text. This refines the text data, allowing voice phishing detection models to focus on meaningful information.
[0083] This process can improve the learning process of voice phishing detection models and contribute to personal information protection by making the STT process more safe and effective.
[0084] A detection model learning unit (160) according to one embodiment can generate text information by utilizing a pre-trained transfer learning model. This model is specialized for natural language processing tasks and is used to process text data to generate meaningful information.
[0085] The detection model learning unit (160) according to one embodiment generates multimodal data by combining generated text information and extracted voice feature information, which provides richer information by simultaneously considering voice and text data and detects voice phishing.
[0086] According to one embodiment, the detection model learning unit (160) can learn the model by applying the Continual Learning learning methodology. This methodology allows the model to continuously learn and adapt to new data, thereby continuously improving the model's performance.
[0087] The control unit (170) can be interpreted as a central processing unit (CPU) and can perform various operations and process data within the system.
[0088] In particular, the control unit (170) can read commands from memory, interpret and execute the commands, and can also perform arithmetic operations such as addition, subtraction, multiplication, and division.
[0089] In addition, the control unit (170) can handle data storage and retrieval, and can also perform the function of reading data from memory and storing the results of performing operations back into memory.
[0090] In addition, the control unit (170) can manage the execution flow of the program, and in particular, can control the flow of the program using commands such as conditional statements (if-else) or iterative statements (for, while). In addition, the control unit (170) can have a small and fast memory device called a register placed inside, and this register can be used to temporarily store data or perform operations.
[0091] The control unit (170) can process and take appropriate action when an external event or exceptional situation occurs, and can quickly access data and instructions by using cache memory that is faster than the main memory.
[0092] In addition, the control unit (170) can use a system bus to communicate with memory or input / output devices, and can provide various power management functions to minimize power consumption.
[0093] FIG. 2 is a drawing illustrating the entire operation of analyzing text and classifying voice using a real-time voice phishing recognition device according to one embodiment.
[0094] The Contrastive task used in the Whisper architecture is a self-supervised learning technique that enables the model to effectively extract and learn information from speech signals.
[0095] Whisper segments and transforms speech signals into continuous time frames and generates random noise masks. It also learns latent representations for the encoder and minimizes contrastive loss through context network training.
[0096] For example, the text information extraction unit (120) utilizes the Whisper model to convert speech to text (STT) and generate voice signal feature values. In addition, it utilizes the KoELECTRA model to generate text feature values from text generated from the STT model.
[0097] The voice feature information extraction unit (130) extracts MFCC (Mel-Frequency Cepstral Coefficient) and Mel spectrogram from the audio data set from the voice signal collection unit (110), and shows an embodiment of using the corresponding feature as an input for a deep neural network model to make a final prediction.
[0098] To achieve this, multimodal data can be generated by combining text information (text features) and speech features. Continual Learning can then be applied to train a model to classify the generated multimodal data. Through the "knowledge distillation process" of the trained model, the multimodal data classification model can be verified, and the final data classification and risk assessment can be calculated.
[0099] The present invention is a voice phishing recognition and detection technology in the field of communication security, which simplifies the data communication process in an existing voice phishing detection system to improve computational efficiency and detection performance.
[0100] Unlike existing systems, the present invention simplifies the communication process by eliminating the need for communication with a database (DB) or a server located remotely.
[0101] All operations used in the present invention are processed solely by an artificial intelligence-based model within the software, and the model enables real-time recognition and detection of voice phishing context from voice information during a call.
[0102] Furthermore, the present invention can more efficiently detect voice phishing and more quickly determine whether a call is a phishing call in real time. Existing voice phishing detection methods require communication with a database and server to analyze call content, convert it to text, and detect voice phishing. This requires significant computational efficiency, and the present invention addresses these limitations.
[0103] Figure 3 is a diagram explaining the learning process (300) of a model using a pre-learned transfer learning model.
[0104] The training process (300) of a pre-trained transfer learning model first receives a token sequence from input data for training. Each token is converted into a dense vector through an embedding layer, and the embedding is used to represent the meaning of a word as a low-dimensional vector in space.
[0105] Before inputting the token sequence to the generator, some tokens can be masked. Masked tokens are not input to the generator; instead, the generator can generate tokens for these positions. This process can be used for data augmentation.
[0106] The masked positions in the original input can be replaced with tokens sampled from the generator and input to the discriminator. The masked positions in the original input data are replaced with tokens sampled from the generator, and this modified input can be input to the discriminator. The discriminator can determine whether the input is actually masked or generated by the generator.
[0107] The discriminator can classify each token based on the differences between the original token sequence and the modified input. This classification can determine whether each token actually belongs to the original data or to data generated by the generator. Therefore, the discriminator performs a binary classification task, classifying each token as either a legitimate call or a voice phishing script.
[0108] In order for the deep learning model inherent in the present invention to not stop at temporary detection but to adapt well to future call situations, a Continual Learning methodology can be used during the model training process.
[0109] Through this methodology, the weights within the model are updated to suit the changing data during the learning process for newly entered data, and data that has completed learning and detection is deleted.
[0110] In addition, model lightweighting techniques such as knowledge distillation and model quantization can be used to ensure real-time performance and reduce computational weight. Since all computations are processed within the device of the invention for voice phishing detection, these can be utilized to reduce the weight of the internal model. In particular, model lightweighting can be performed through knowledge distillation from a model with a complex computational volume to a lightweight model.
[0111] Finally, voice phishing context recognition is performed by utilizing the aforementioned feature information and model, and the predicted probability value within the model generated during the prediction operation process is calculated and quantified to present the risk level to the user in order to provide more detailed judgment criteria for voice phishing detection.
[0112] Figure 4 is a diagram explaining a training process (400) for detecting voice phishing using learning data and test data.
[0113] To train the KoELECTRA model using the Train data, the Generator and Discriminator are used to distinguish and classify normal script data and phishing script data.
[0114] To achieve this, the KoELECTRA model loads train data. This data contains both normal and phishing scripts, and each script must be encoded in an appropriate format. For example, text data can be converted into numbers through tokenization and encoding.
[0115] The KoELECTRA model consists of a Generator and a Discriminator. The Generator generates phishing scripts from training data. The Discriminator distinguishes between scripts generated by the Generator and actual phishing scripts.
[0116] The generator generates phishing scripts from training data, and can distinguish between normal and phishing scripts.
[0117] The Discriminator compares the generated phishing script with the actual phishing script to determine whether it is phishing or not, and the Generator and Discriminator can handle binary classification problems.
[0118] The Generator can be trained to classify the generated phishing scripts as actual phishing scripts, and the Discriminator can be trained to accurately classify the generated scripts as phishing scripts.
[0119] The Generator and Discriminator can be trained alternately. The Generator can improve the phishing scripts it generates to fool the Discriminator, while the Discriminator can be trained to identify the phishing scripts generated by the Generator.
[0120] This alternate training process is based on the core principles of Generative Adversarial Networks (GANs). After training, the model's performance is evaluated using test data. The Generator generates phishing scripts from the test data, and the Discriminator evaluates the ability to distinguish the generated scripts from actual phishing scripts.
[0121] If the model's performance is unsatisfactory, additional fine-tuning can be performed. Fine-tuning is the process of improving model performance by adjusting the training data. Through this process, the KoELECTRA model can improve its ability to distinguish and classify legitimate and phishing scripts using the Generator and Discriminator.
[0122] Meanwhile, the present invention can collect train data and test data. This data contains a mix of normal and voice phishing calls.
[0123] The collected data must be preprocessed into an appropriate format, which includes converting the voice signal into a feature vector, labeling it, and processing it into a form that can be used for model training.
[0124] In the process of training a model using preprocessed Train data, the model can learn features that can distinguish between normal calls and voice phishing calls.
[0125] Typically, a deep learning algorithm is used to train a model, and the train data is used to adjust the weights so that the model can correctly classify the input data.
[0126] The learned model is applied to test data to evaluate its performance. This process allows us to evaluate whether the model accurately classifies the given speech data.
[0127] Evaluation is typically based on accuracy, which measures how accurately the model can distinguish between legitimate calls and voice phishing.
[0128] Through this process, the Voice Phishing Classification Model can be trained based on the Train data and then use the Test data to identify and classify normal calls and voice phishing calls.
[0129] As a result of the training process (400) for detecting voice phishing using learning data and test data, it can be confirmed that the accuracy standard shows a performance of approximately 96%, as shown in [Table 1] below.
[0130]
[0131] [Table 1]
[0132]
[0133]
[0134] FIG. 5 is a diagram illustrating an embodiment (500) of detecting voice phishing in an audio data set.
[0135] In the embodiment (500) of Fig. 5, an embodiment is shown in which MFCC (Mel-Frequency Cepstral Coefficient) and Mel spectrogram are extracted from an audio data set, and the corresponding features are used as inputs of a deep neural network model to make a final prediction.
[0136] Drawing symbol 510 is MFCC, which represents a numerical value that represents the unique characteristics of a feature sound that can be extracted from an audio signal.
[0137] Drawing symbol 520 is a Mel spectrogram, which represents a spectrogram in which the frequency is converted to a Mel scale.
[0138] As a result of the embodiment (500) of detecting voice phishing in an audio data set, it can be confirmed that the accuracy standard shows a performance of approximately 94%, as shown in [Table 2] below.
[0139]
[0140] [Table 2]
[0141]
[0142]
[0143] Ultimately, the present invention can provide a Deep Learning-based model that detects voice phishing using voice phishing voice data.
[0144] By performing the STT Process, text is extracted from voice data to build voice and text data, and two detection models can be built using the KoELECTRA model and text data, and the DNN model and voice data.
[0145] Each of them shows robust performance with accuracy standards of 98.34% and 94%, respectively, and can be confirmed to have innovative insights in preventing voice phishing crimes.
[0146] Figure 6 is a drawing (600) explaining an embodiment of notifying a user terminal of voice phishing.
[0147] The target users of the present invention are not limited to groups vulnerable to voice phishing, such as the elderly and the disabled, but can also be ordinary people who are not vulnerable to voice phishing.
[0148] To this end, software installed on the user terminal can collect voice signals with the target user's consent, and the collected voice signals are used to detect voice phishing.
[0149] In the present invention, a voice signal transmitted to a user's communication terminal is extracted, and in a real-time call situation, the present invention converts the extracted voice signal into text (STT) and simultaneously generates unique voice feature information indicated in the voice signal itself.
[0150] Additionally, the context of voice phishing can be recognized using a deep learning-based model based on the generated voice feature information and text information.
[0151] The deep learning-based models mentioned in the present invention include both pre-trained models for text and speech recognition and models capable of continuous additional learning, such as new types of deep learning models.
[0152] The present invention can improve computing efficiency by providing a process for effectively detecting and quickly determining voice-like fraud.
[0153] Additionally, in addition to the vulnerable group of voice-like fraud, it targets general users and can detect voice-like fraud in real time by collecting user voices with consent.
[0154] By converting voice signals into text and extracting unique voice feature information, a deep learning-based model can be used to identify voice-like fraud contexts. Continuous learning methods can be used to allow the model to adapt to future call situations, and model lightweighting techniques can be introduced to improve real-time performance.
[0155] By using techniques such as knowledge distillation and model quantization, complex models can be reduced to lightweight models, allowing all computations for voice-like fraud detection to be handled on-device, reducing the computational load. The predicted probability values generated by the model can be presented as risk levels to provide users with more detailed judgment criteria for voice-like fraud detection.
[0156] In particular, the present invention can detect voice phishing in real time, thereby preventing financial damage to individuals and organizations due to fraudulent activities, thereby creating a more stable financial environment, which can be beneficial to both consumers and financial institutions.
[0157] Users can participate more actively in economic activities by gaining confidence in the security of online financial transactions.
[0158] The present invention can improve social safety through voice-like fraud detection technology, and can help ordinary users make calls more safely, in addition to vulnerable groups protected from voice-like fraud.
[0159] Additionally, since personal voices are collected with consent, a high level of ethics is maintained in terms of personal information protection, which can contribute to strengthening social trust.
[0160] The present invention effectively detects and quickly assesses voice-like fraud, thereby minimizing economic losses and preventing fraud, thereby enhancing the financial security of businesses and individuals. This innovative technology for detecting voice-like fraud can increase trust among financial institutions, businesses, and general users, thereby enhancing social stability.
[0161] FIG. 7 is a diagram illustrating an operation method of a real-time voice phishing recognition device according to an embodiment.
[0162] The method of operating a real-time voice phishing recognition device according to one embodiment collects in real time a voice signal transmitted to a communication terminal after a call is established (step 701).
[0163] Next, the collected voice signal is converted into text to extract text information (step 702), and in parallel with the extraction of text information, voice feature information for the voice signal is extracted (step 703).
[0164] The method of operating a real-time voice phishing recognition device according to an embodiment calculates a risk level based on the extracted text information and the extracted voice feature information using a voice phishing detection model (step 704).
[0165] To this end, the operating method of the real-time voice phishing recognition device may include a process of learning a model by inputting a public data set generated from a voice file of an actual damage case related to a voice phishing crime, or learning a model through a speech to text (STT) processor by crawling a script on a web page for the voice file.
[0166] To calculate risk, a pre-trained transfer learning model is used to generate the text information, and multimodal data is created by combining the generated text information with the extracted speech feature information. Furthermore, a model employing the Continual Learning methodology is used to classify the generated multimodal data, and a risk index for the classified multimodal data is calculated.
[0167] Next, voice phishing can be detected based on the calculated risk level (step 705).
[0168] Ultimately, the present invention can be used to build an AI model that monitors electronic financial transactions in real time. This model can be trained to detect suspicious patterns or behaviors in electronic financial transactions.
[0169] Using the present invention, unlike existing methods, this model can omit the data communication process between the DB and the server, and instead, can directly receive and analyze data of electronic financial transactions in real time.
[0170] The present invention allows for real-time analysis of electronic financial transactions to assess the potential for fraud. This allows for real-time detection and prevention of fraud, even during the transaction process.
[0171] Using the present invention, if suspicious patterns or behaviors are detected during electronic financial transactions, the transaction can be stopped, necessary measures can be taken to prevent fraud, and the user's financial information and assets can be protected.
[0172]
[0173] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0174] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0175] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0176] Although the embodiments described above have been described with limited drawings, those skilled in the art will recognize that various modifications and variations can be made based on the above description. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0177] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. A voice signal collection unit that collects voice signals transmitted to a communication terminal in real time after the call is set; A text information extraction unit that converts the collected voice signal into text and extracts text information; In parallel with the extraction of the above text information, a voice feature information extraction unit that extracts voice feature information for the voice signal; A risk calculation unit that calculates a risk based on the extracted text information and the extracted voice feature information using a voice phishing detection model; and A voice phishing detection unit that detects voice phishing based on the risk level calculated above. A real-time voice phishing recognition device including:
2. In paragraph 1, A detection model learning unit that learns the voice phishing detection model by inputting a public data set generated from voice files of actual damage cases related to voice phishing crimes. A real-time voice phishing recognition device including:
3. In paragraph 2, The above detection model learning unit, A real-time voice phishing recognition device characterized in that it crawls a script on a web page for the above voice file and processes performance measurement and model learning through an STT (speech to text) processor.
4. In paragraph 3, The above detection model learning unit, A real-time voice phishing recognition device that processes personal information de-identification and stopword removal processes as the above STT (speech to text) process.
5. In paragraph 1, The above voice phishing detection unit, A real-time voice phishing recognition device that detects voice phishing by recognizing the context of voice phishing in real time based on the risk level calculated above.
6. In paragraph 1, The above risk calculation section is, Generate the above text information by utilizing a pre-trained transfer learning model, A real-time voice phishing recognition device characterized in that it generates multimodal data by combining the generated text information and the extracted voice feature information, classifies the generated multimodal data using a model that applies a Continual Learning learning methodology, and calculates a risk level for the classified multimodal data.
7. A step of collecting voice signals transmitted to a communication terminal in real time after setting up the phone; A step of converting the collected voice signal into text and extracting text information; In parallel with the extraction of the above text information, a step of extracting voice feature information for the voice signal; A step of calculating a risk level based on the extracted text information and the extracted voice feature information using a voice phishing detection model; and Steps for detecting voice phishing based on the risk level calculated above A method of operating a real-time voice phishing recognition device including:
8. In paragraph 7, A step of training a model by inputting a public data set generated from voice files of actual damage cases related to voice phishing crimes, or training a model through a STT (speech to text) processor by crawling the script on a web page for the voice files. A method of operating a real-time voice phishing recognition device including:
9. In paragraph 7, The step of calculating the risk level based on the extracted text information and the extracted voice feature information is as follows: A step of generating the text information by utilizing a pre-trained transfer learning model; A step of generating multimodal data by combining the generated text information and the extracted voice feature information; A step of classifying the generated multimodal data using a model that applies the Continual Learning learning methodology; and Step for calculating risk for the above classified multimodal data A method for operating a real-time voice phishing recognition device, characterized in that it includes:
Citation Information
Patent Citations
Communication network fraud identification method and apparatus, and electronic device
CN111601000A
System for monitoring phishing attack and preventing method
KR1020150057338A
Apparatus and method for detecting smishing message
KR1020170024777A
Method and system for preventing bank fraud
KR1020170060958A
Method of continuously producing cannabidiol from Cannabis sp. and uses thereof
KR1020210040710A