Method and device for determining the malicious nature of a telephone communication.
The method and device for determining the malicious nature of telephone communications utilize advanced data analysis techniques to assess the likelihood of a call being fraudulent, addressing the limitations of existing solutions and enhancing user safety by providing a reliable means to identify and respond to suspicious calls.
Patent Information
- Application Number
- FR2023014756
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
AI Technical Summary
Existing solutions for detecting malicious telephone calls are not reliable, as fraudsters can easily impersonate legitimate callers by manipulating telephone number headers, making it difficult for users to distinguish genuine from fraudulent communications.
A method and device that determine the malicious nature of a telephone communication by obtaining data from the communication and analyzing it using techniques such as voice recognition, speech analysis, and text transcription to assign a score indicating the likelihood of the call being fraudulent, allowing users to take appropriate action based on the score.
This solution provides a more reliable means of identifying fraudulent calls by leveraging advanced data analysis techniques, enabling users to make informed decisions about engaging with suspicious communications and potentially preventing identity theft and fraud.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for determining the malicious nature of a telephone communication.
[0001] 1. Field of the invention
[0002] The invention belongs to the field of telecommunications and relates in particular to a method for detecting / characterizing a telephone call as malicious / fraudulent.
[0003] 2. Prior Art
[0004] Users of telecommunications services (telephony, video telephony, SMS, RCS, Instant messaging, email, etc.) are receiving more and more unsolicited electronic communications, which can cause a lot of inconvenience. This is the case, for example, when they receive calls made during phishing campaigns. The hackers behind these campaigns aim to obtain personal information from their contacts in order to subsequently commit identity theft / fraud.
[0005] We are thus observing an increase in telephone calls from fraudsters encouraging their interlocutors to validate, to their detriment, (for example on their mobile terminal) banking operations and / or authorizations to access their account (context of the Payment Services Directive No. 2). During these malicious calls, the fraudster most often claims to be from the bank of his interlocutor and develops a discourse which takes up a pre-established discussion framework with appropriate banking vocabulary, words belonging to the lexical field of urgency and a very directive tone. For an expert, these communications are quite characteristic / detectable but for the customer / interlocutor called they are not always easy to detect.
[0006] Solutions exist to combat such actions. They are mainly based on so-called "anti-spam" mechanisms which allow, on the basis of the caller's telephone number, to block the call. However, it has become very easy for malicious people to usurp the identity of the customer's / caller's bank by, for example, declaring the bank's telephone number in the SIP FROM header of the SIP (Session Initiation Protocol) protocol allowing the establishment of the communication.
[0007] There is therefore a need for a more reliable solution that can detect / characterize a malicious telephone call.
[0008] 3. Statement of the invention
[0009] The invention improves the state of the art and proposes a method for deter-
[0010]
[0011]
[0012]
[0013]
[0014]
[0015] determination of the malicious nature of a first communication established between a transmitting terminal and a receiving terminal, said method being implemented by a determination device and characterized in that the method comprises the following steps: - obtaining, based on at least one first piece of data received by said receiving terminal through said first communication, at least one second piece of data; - determination, based on the value of said at least one second piece of data, of the malicious nature of said first communication. Advantageously, according to the invention, a communication established between a transmitting terminal and a receiving terminal can be determined as fraudulent, i.e. sent by a malicious person, based on the value of a second data item itself obtained based on a first data item received by the receiving terminal. The second data can be a Boolean data type or correspond to a score, for example between 0 and 1. This score is then compared to a reference threshold, for example 0.6. If the assigned score is higher (or higher or equal) than this reference threshold, then the communication can be considered a fraudulent communication. Conversely, if the assigned score is lower (or lower or equal) than this reference threshold, then the communication can be considered a non-fraudulent communication. Also, when the determination device is understood by the receiving terminal, the user of the receiving terminal will be able, for example, in the case of a telephone call, to take the call or not depending on the value of the second data. Note that the process can be located on a server of an operator participating in the communication. According to a particular embodiment of the invention, a method as described above is characterized in that said at least one first data item corresponds at least to one data item of audio and / or video and / or text type. This embodiment of the invention makes it possible, for example, to determine / characterize the communication as fraudulent based on the result of processing the audio stream (first data within the meaning of the invention) received by the receiving terminal (case of a telephone communication). The processing may, for example, comprise: - a step of voice recognition of the caller (identification of the speaker / caller); - a step of analyzing all or part of the verbatim(s) obtained following the transcription of the vocal content of the audio stream in the form of an exploitable text (in English speech to text or automatic speech recognition); - a stage of analysis of rhythm, accent, intonation, flow, the grammatical agreements, verb tenses and their conjugations, syntax, pauses in the caller's speech, etc.; - etc.
[0016] When the first data item comprises, in addition to the audio component, a video component, the processing may further comprise a step of analyzing the posture of the caller and / or facial expressions or even an eye tracking step which makes it possible, for example, to determine that the caller is reading a pre-established script, etc.
[0017] According to a particular embodiment of the invention, a method as described above is characterized in that said at least one second data item is obtained from a server following the sending, by said receiving terminal, of said first data item via a second communication established between said receiving terminal and server.
[0018] This embodiment of the invention makes it possible to obtain, from a server specialized in the detection of fraudulent communications, data / information (second data within the meaning of the invention) indicating whether the communication established between the sending terminal and the receiving terminal is fraudulent or not. Thus, thanks to the use of the specialized server (for example a voice server) the determination method according to the invention is made particularly reliable and efficient.
[0019] According to a particular embodiment of the invention, the specialized server can integrate a conversational agent capable of dialoguing with the interlocutor of the receiving terminal. A conversational agent is understood to mean a computerized dialog automaton capable of dialoguing with a user, for example, via a mobile terminal, a connected object, a computer, a dashboard of a vehicle, etc. The conversational agent generates messages (for example vocal in the case of a voice assistant) intended for the client / user based on data sent by the client. Note that the dialog can be done vocally and / or through a human-machine interface (screen, keyboard, etc.).
[0020] Note that the second communication can be established automatically between the sending terminal and the receiving terminal, for example, as soon as the first communication is received / accepted by the receiving terminal or when an event is detected during the first communication (for example when detecting a keyword spoken by the caller).
[0021] Alternatively, the second communication is established / triggered, for example, by the user of the receiving terminal via a human-machine interface.
[0022] Also note that the second communication may be an encrypted communication in order to guarantee a certain level of security.
[0023] According to a particular embodiment of the invention, a method as described above is characterized in that said first and second communications are joined together so as to create a teleconference.
[0024] A teleconference is understood to mean a conference conducted remotely (for example, a telephone conference, a videoconference, a virtual reality conference, etc.). Note that in the case of a telephone conference, it can be configured to allow a called party to participate in a communication already established, or configured so that the called party simply listens to the communication already established without participating in it.
[0025] This embodiment of the invention allows the specialized server to obtain the first data in real time. Thus, its processing can also be carried out in real time and the sending of the second data can be carried out as quickly as possible, facilitating, for example, the decision-making process regarding an action to be carried out at the level of the first communication.
[0026] (According to a particular embodiment of the invention, a method as described above is characterized in that said at least one first data item corresponds to at least one first audio stream and in that said at least one second data item is obtained as a function of the result of an analysis of said first audio stream.
[0027] This embodiment of the invention makes it possible to determine that the first communication is fraudulent on the basis of processing / analysis of an audio stream received, by the receiving terminal, from the sending terminal (for example based on the result of the analysis of the text transcription of the vocal content of the incoming audio stream, i.e. the speech spoken by the caller).
[0028] According to a particular embodiment of the invention, a method as described above is characterized in that said at least one second data item is further obtained as a function of the result of an analysis of a second audio stream transmitted by said receiving terminal to said transmitting terminal via said first communication.
[0029] This embodiment of the invention makes it possible to determine that the first communication is fraudulent, also based on the result of processing / analysis of an audio stream sent by the receiving terminal to the sending terminal (for example based on the result of the analysis of the text transcription of the vocal content of the outgoing audio stream, i.e. the speech spoken by the called party).
[0030] According to a particular embodiment of the invention, a method as described above is characterized in that said first communication is interrupted as a function of the value of said at least one second data item.
[0031] This mode of implementation of the invention makes it possible to hang up / interrupt the first communication when the value of the second data indicates that it is fraudulent.
[0032] According to a particular embodiment of the invention, a method as described above is characterized in that a second audio stream transmitted by said receiving terminal to said transmitting terminal via said first communication is modified as a function of the value of said at least one second data item.
[0033] This embodiment of the invention makes it possible, for example, to alter the audio stream emitted by the receiving terminal as a function of the value of the second data item, i.e. when the first communication is determined to be fraudulent. In other words, when the user responds to a fraudster, the method can, for example, alter all or part of the voice response of the called party so that the latter cannot transmit personal / sensitive information to the fraudster.
[0034] The invention also relates to a device for determining the malicious nature of a first communication established between a transmitting terminal and a receiving terminal, characterized in that the device comprises: - a module for obtaining, as a function of at least one first item of data received by said receiving terminal through said first communication, at least one second item of data; - a module for determining, based on said at least one second piece of data, the malicious nature of said first communication.
[0035] The term module can correspond to a software component as well as to a hardware component or a set of hardware and software components, a software component itself corresponding to one or more computer programs or subroutines or more generally to any element of a program capable of implementing a function or a set of functions as described for the modules concerned. In the same way, a hardware component corresponds to any element of a hardware assembly capable of implementing a function or a set of functions for the module concerned (integrated circuit, smart card, memory card, etc.).
[0036] Terminal characterized in that it comprises a determination device as described above.
[0037] The invention also relates to a computer program comprising instructions for implementing the above method according to any of the particular embodiments described above, when said program is executed by a processor. The method can be implemented in various ways, in particular in hard-wired form or in software form. This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0038] The invention also relates to a recording medium or readable information medium by a computer, and comprising instructions of a computer program as mentioned above. The recording media mentioned above may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a hard disk. On the other hand, the recording media may correspond to a transmissible medium such as an electrical or optical signal, which can be conveyed via an electrical or optical cable, by radio or by other means. The programs according to the invention may in particular be downloaded from a network such as the Internet.
[0039] Alternatively, the recording media may correspond to an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the method in question.
[0040] This determination device and this computer program have characteristics and advantages similar to those described previously in relation to the determination method.
[0041] 4. List of figures
[0042] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:
[0043] [Fig-1] [Fig.l] illustrates an example of an implementation environment according to a particular embodiment of the invention,
[0044] [Fig.2] [Fig.2] illustrates the architecture of a device implementing the invention according to a particular embodiment,
[0045] [Fig.3] [Fig.3] illustrates steps of the determination method according to a particular embodiment of the invention,
[0046] 5. Description of an embodiment of the invention
[0047] [Fig. 1] illustrates an example of an environment for implementing the invention according to a particular embodiment. The environment represented in [Fig. 1] comprises a communication network 100 to which the terminals 101 and 102 belonging respectively to the users UT1 and UT2 are connected.
[0048] The communication network 100 may be a mobile telecommunications network with an access network of the GSM, EDGE, 3G, 3G+, 4G, 5G, WiFi, etc. type or a fixed telecommunications network with an access network of the ADSL, Fiber, VDSL, WiFi, etc. type. It may implement an RCS type architecture. The telecommunications network 100 may also be an IP telecommunications network implemented by a VoIP (Voice over IP) network architecture of the IMS (IP Multimedia Subsystem) type or any other network architecture. The telecommunications network 100 may also correspond to a group of telecommunications networks of different operators, interconnected with each other via interconnection servers.
[0049] The terminals 101 and 102 can be any type of terminal allowing communication sessions of the telephone, videophone, instant messaging, SMSoIP, etc. type to be established. The terminals 101 and 102 correspond for example to a mobile phone, a landline phone, a smartphone (smart phone in English), a tablet, a television connected to a communication network, a connected object, an autonomous car, or a personal computer. The terminals 101 and 102 can transmit and receive all types of communications, via the communication network 100. In our example, the terminal 101 belongs to a hacker UT1 and the terminal 102 to one of his targets.
[0050] The terminal 102 is capable of implementing the determination method according to a particular embodiment.
[0051] According to a particular embodiment of the invention, the determination method can be distributed partially and / or totally between the terminal 102 and a determination device associated with it.
[0052] The implementation environment of the invention also comprises a specialized server 103 for fraud detection. The server 103 is, for example, connected to one or more databases (not shown) updated regularly and listing elements making it possible to determine whether a communication is fraudulent or not. For example, the databases can store the voice prints of the fraudsters, the scripts spoken by the fraudsters, telephone numbers or identifiers associated with the fraudsters, etc. The server 103 is capable of analyzing all or part of the content of an audio and / or video and / or textual communication via dedicated modules.These modules can for example perform a speech recognition process, voice recognition, optical character recognition, computer vision, transcription of the vocal content of an audio stream, analysis of the rhythm, accent, intonation, rate, grammatical agreements, verb tenses and their conjugations, syntax, speech pauses of the vocal content of an audio stream, etc.
[0053] The server 103 may further comprise a conversational agent capable of establishing communications with other terminals. For example, the conversational agent may establish a voice communication with the terminals 101 and 102 and more particularly with the users UT1 and / or UT2 through the telecommunications network 100. More specifically, the conversational agent may for example vocally transmit to the user UT2 the result of an analysis of all or part of a communication (second data within the meaning of the invention). Note that, in this case, the conversational agent may correspond to a voice assistant.
[0054] The conversational agent, which in our example is hosted / executed at the server 103, is connected to the communication network 100. The terminal 102 can further integrate a software client of the conversational agent allowing it, for example, to establish communication with the conversational agent via the network 100 (for example IP and / or circuit) and RCS (Rich Communication Services) technology.
[0055] According to a particular embodiment of the invention, the conversational agent can be implemented by any type of connected terminal having the architecture of a computer such as, for example, and in a non-limiting manner, a gateway, a games console, a television, an ATM, a router, a tablet, a personal computer, a smartphone, etc. Note that the conversational agent can be distributed partially or totally between the terminal 102 and the server 103.
[0056] [Fig.2] illustrates a device 200 configured to implement the determination method according to a particular embodiment. According to this particular embodiment of the invention, the device 200 has the conventional architecture of a computer, and notably comprises a memory MEM, a processing unit UT, equipped for example with a processor PROC, and controlled by the computer program PG stored in memory MEM. The computer program PG comprises instructions for implementing the steps of the determination method as described previously, when the program is executed by the processor PROC.
[0057] At initialization, the code instructions of the computer program PG are for example loaded into a memory before being executed by the processor PROC. The processor PROC of the processing unit UT notably implements the steps of the determination method according to any one of the particular embodiments described in relation to FIGS. 1 and 3, according to the instructions of the computer program PG.
[0058] The device 200 comprises an OBT obtaining module capable of obtaining at least one second piece of data received, for example, from a conversational agent or a server located in the network 100. The second piece of data is determined as a function of a first piece of data received by the terminal 102 through a communication established between the terminal 101 and the terminal 102 via the telecommunications network 100. The first piece of data may for example comprise an audio stream and / or text and / or a video stream.
[0059] According to a particular embodiment, the OBT module is capable of communicating (transmitting and / or receiving messages) via an IP network and / or circuit and RCS technology.
[0060] According to a particular embodiment of the invention, the device 200 may comprise an analysis module ANA capable of analyzing the first data received through the communication established between the terminal 101 and the terminal 102. The determination device and more particularly the ANA analysis module can obtain the first data from a third-party device or directly from an audio and / or video rendering device (speaker, screen, etc.) of the terminal 102. The analysis module can execute all types of processes allowing the analysis of the first data. For example, the ANA module can execute a process of speech recognition, voice recognition, optical character recognition, computer vision, transcription of the vocal content of an audio stream, analysis of the rhythm, accent, intonation, flow, grammatical agreements, verb tenses and their conjugations, syntax, speech pauses of the vocal content of an audio stream, etc.
[0061] According to a particular embodiment, the device 200 may comprise a digital storage module (for example the memory MEM) capable of storing all or part of the first data and / or the second data.
[0062] The device 200 further comprises a DETER module capable of determining, as a function of the value of the second data item, the fraudulent nature of the communication established between the terminal 101 and the terminal 102.
[0063] According to a particular embodiment, the device 200 comprises a human-machine interface module IHM capable of restoring vocally, visually or via vibrations (screen, speaker, vibrating motor, etc.) the second data to a user.
[0064] [Fig. 3] illustrates steps of the determination method according to one of the particular embodiments of the invention presented previously in support of [Fig. 1] and [Fig. 2], the method being executed on the terminal 102.
[0065] During the first step 300, the terminal 102 receives a request to establish a telephone communication from the terminal 101. The user UT2 accepts the communication and the user UT1 engages in conversation with the user UT2. During the discussion, the user UT2 has a doubt about the fraudulent nature of the call. He then triggers the execution of the determination method. The triggering can for example be done via a human-machine interface of the terminal 102 (keyword spoken, mouse click / button, etc.).
[0066] According to a particular embodiment, the execution of the determination method is triggered automatically when the user UT2 accepts the communication.
[0067] According to a particular embodiment, a notification (sound, message, etc.) is sent to the terminal 101 and its user UT1 when the determination method is triggered.
[0068] All or part of the incoming audio stream (first data within the meaning of the invention), i.e. received by the terminal 102 from the terminal 101, is then duplicated by the terminal 102 then transmitted (step 301) in real time to the server 103. The terminal 102 communicates, for example, with a conversational agent / voice assistant of the server 103 via RCS (Rich Communication Services) technology.
[0069] Alternatively, the terminal 102 establishes a conference call including the terminals 101, 102 and the voice assistant of the server 103. Thus, the voice assistant being integrated into the conversation, it has real-time knowledge of the entirety of the voice exchanges (i.e. the conversation) carried out between the two interlocutors UT1 and UT2.
[0070] Note that all or part of the incoming audio stream can be recorded by the method in order, for example, to carry out its analysis (in real time or a posteriori) or to feed a database with a set of data relating to fraud management.
[0071] The server 103 then obtains (step 400) all or part of the incoming audio stream and then analyzes it (step 401). To do this, the server 103 can rely on different mechanisms / methods and knowledge bases. The server can, for example, execute a speech recognition method, a voice recognition method, a transcription of the vocal content of the audio stream, an analysis of the rhythm, the accent, the intonation, the flow rate, the grammatical agreements, the verb tenses and their conjugations, the syntax, the speech pauses of the vocal content of the audio stream, in order to determine the nature of the communication. These methods are already known per se and will not be described in more detail here.
[0072] More concretely, the server 103 can for example analyze the voice of the interlocutor UT1, determine a voice print and compare it with voice prints contained within a database listing the voice prints of known fraudsters.
[0073] Alternatively or cumulatively, the server 103 can analyze the rhythm, accent, intonation, flow, grammatical agreements, verb tenses and their conjugations, syntax, pauses of the speech of the interlocutor UT1. The server can thus determine whether the speech uttered by the user UT1 is a directive, anxiety-provoking and / or threatening speech or whether the speech takes up known characteristics of fraudster practices.
[0074] Alternatively or cumulatively, the server 103 can analyze the text transcription of the voice content of the audio stream, i.e. the speech uttered by the user UT1, in order to determine whether the speech is close to a pre-established script of fraud.
[0075] Alternatively or cumulatively, the server 103 can obtain and analyze the audio stream transmitted, by the receiving terminal 102, to the transmitting terminal 101 in order to determine whether the speech of the user UT2 intended for the user UT1 is likely to be prejudicial to the user UT2 (for example via the analysis of the entire of the conversation established between users UT1 and UT2 or via only the analysis of the responses of user UT2).
[0076] The server 103 obtains for each analysis (for example from each module specialized in a particular analysis) a score indicating the degree of confidence in the communication. The server 103 can then consolidate an overall score by taking into account or not weightings associated with each score obtained.
[0077] According to a particular embodiment, the incoming flow is divided into a plurality of segments and the analysis is carried out on all or part of the segments.
[0078] During step 402, the server 103 sends to the terminal 102 the overall score or all of the scores obtained as a result of each analysis (second data within the meaning of the invention).
[0079] The score(s) are then received by the terminal 102 during step 302. The terminal 102 analyzes the score(s) received and when the score(s) exceed an alert threshold, different actions can be carried out by the terminal 102 (cumulatively or not): - the restitution via a human-machine interface of the terminal 102 of at least one piece of information reminding the user of the risk associated with the conversation and the actions not to be carried out. This information can be rendered vocally and / or visually, for example in the form of text displayed on the screen of the terminal 102. - sending a notification to the caller indicating that the call is about to be recorded; - the end of the communication established with terminal 101 or the conference call; - etc.
[0080] According to a particular embodiment, the terminal 102 can, depending on the value of the score(s) received, for example when the communication is determined to be fraudulent, alter all or part of the audio stream emitted by the terminal 102 to the terminal 101 so that the called party (UT2) cannot transmit personal / sensitive information to the fraudster.
[0081] It goes without saying that the embodiment described above has been given for purely indicative purposes and is in no way limiting, and that numerous modifications can easily be made by those skilled in the art without departing from the scope of the invention.
Claims
Claims
1. Method for determining the malicious nature of a first communication established between a transmitting terminal (101) and a receiving terminal (102), said method being implemented by a determination device (200) and characterized in that the method comprises the following steps: - obtaining (302), as a function of at least one first data item received by said receiving terminal through said first communication, at least one second data item; - determining (302), as a function of the value of said at least one second data item, the malicious nature of said first communication.
2. Method according to claim 1 wherein said at least one first data item corresponds at least to one data item of audio and / or video and / or text type.
3. Method according to claim 1 wherein said at least one second data item is obtained from a server (103) following the sending, by said receiving terminal (102), of said first data item via a second communication established between said receiving terminal (102) and said server (103).
4. The method of claim 3 wherein said first and second communications are joined together to create a teleconference.
5. Method according to claim 1 wherein said at least one first data item corresponds to at least one first audio stream and in that said at least one second data item is obtained as a function of the result of an analysis (401) of said first audio stream.
6. Method according to claim 1 wherein said at least one second data item is further obtained as a function of the result of an analysis of a second audio stream transmitted by said receiving terminal (102) to said transmitting terminal (101) via said first communication.
7. Method according to claim 1 wherein said first communication is interrupted depending on the value of said at least one second data item.
8. The method of claim 1 wherein a second audio stream transmitted by said receiving terminal (102) to said terminal transmitter (101) through said first communication is modified according to the value of said at least one second data.
9. Device for determining the malicious nature of a first communication established between a transmitting terminal (101) and a receiving terminal (102) characterized in that the device comprises: - a module for obtaining (OBT), as a function of at least one first data item received by said receiving terminal through said first communication, at least one second data item; - a module for determining (DETER), as a function of said at least one second data item, the malicious nature of said first communication.
10. Terminal characterized in that it comprises a device according to claim 9.
11. A computer program comprising instructions for implementing the method according to any one of claims 1 to 8, when the program is executed by a processor.
Citation Information
Patent Citations
Multilayer set of neural networks
GB2584827A
Determining scam risk during a voice call
US20170142252A1
Fraud Detection System and Method
US20210383410A1
Prohibiting voice attacks
US20220407886A1