Fraud risk intervention method, terminal equipment and storage medium
By integrating multimodal call data locally on the terminal device for fraud risk assessment, the problem of failure of protection mechanisms and data leakage of number tagging and cloud voice recognition in existing technologies is solved, and more efficient fraud risk identification and privacy protection are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing fraud call prevention solutions rely on number tagging and cloud-based voice recognition, which are prone to failure of protection mechanisms and data leakage. They are also inadequate to deal with new types of fraud and cannot effectively protect user privacy.
Fraud risk identification is performed locally on the terminal device using multimodal call data, including call audio stream, semantic text stream, and interactive event stream. Fraud risk assessment and intervention are carried out by fusing multimodal data, avoiding data upload to the cloud.
It improves the accuracy and robustness of fraud risk identification, protects user privacy data, prevents data leakage and abuse, and enhances security protection effectiveness.
Smart Images

Figure CN122002289A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of terminal technology, and in particular to a method for intervening in fraud risk, a terminal device, and a storage medium. Background Technology
[0002] With the rapid development of communication technology, fraudulent calls are becoming increasingly prevalent, and their methods are constantly evolving, exhibiting highly sophisticated and intelligent features. Currently, mainstream fraudulent call prevention solutions mainly rely on two methods: one is number tagging, where terminal devices extract the caller's number and upload it to a remote server when a call is detected. The server then matches the number against a database of fraudulent numbers. If the caller's number matches a number in the database, interception or alerts are triggered. The other method is keyword recognition based on cloud-based speech recognition, where the call audio is uploaded to the cloud in real time, automatically converted into text using speech recognition technology, and compared with a pre-set database of sensitive words to determine the presence of fraud risk.
[0003] However, both methods have significant limitations. Firstly, the logic for number tagging relies on historically tagged numbers; fraudsters can easily disable the protection mechanism by simply changing the number or using an unlisted one. Secondly, cloud-based voice recognition requires uploading call content to a cloud platform, posing risks of data leakage, misuse, or illegal storage.
[0004] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the date of this application, nor does it imply that it can be considered prior art in this disclosure. Summary of the Invention
[0005] This specification provides a method, terminal device, and storage medium for intervening in fraud risks. It utilizes multimodal call data locally on the terminal device to identify and intervene in fraud risks, thereby improving the accuracy and robustness of fraud risk identification, while preventing the leakage of user privacy data during transmission or storage.
[0006] Firstly, this specification provides a method for intervening in fraud risk, applied to a terminal device. The method includes: in response to detecting a call event, obtaining the caller ID of the current call; performing a trusted verification on the caller ID; if the caller ID fails the trusted verification, collecting multimodal call data during the call, and performing a fraud risk determination locally on the terminal device based on the multimodal call data to obtain a fraud risk level, wherein the multimodal call data includes at least two data streams among call audio stream, semantic text stream, or interactive event stream; and performing a corresponding risk intervention operation locally on the terminal device based on the fraud risk level.
[0007] In some embodiments, when the multimodal call data includes three modalities of data streams: call audio stream, semantic text stream, and interactive event stream, fraud risk determination is performed based on the multimodal call data to obtain a fraud risk level, including: performing fraud risk assessment based on the call audio stream to obtain a first fraud risk score; performing fraud risk assessment based on the semantic text stream to obtain a second fraud risk score; performing fraud risk assessment based on the interactive event stream to obtain a third fraud risk score; and performing a comprehensive determination based on the first fraud risk score, the second fraud risk score, and the third fraud risk score to obtain the fraud risk level.
[0008] In some embodiments, a fraud risk assessment is performed based on the call audio stream to obtain a first fraud risk score, including: extracting the target voiceprint feature of the caller from the call audio stream; and comparing the target voiceprint feature with each voiceprint feature stored in a first database for similarity, and determining the first fraud risk score based on the comparison result, wherein the first database stores the voiceprint feature of at least one fraudulent user.
[0009] In some embodiments, a fraud risk assessment is performed based on the semantic text stream to obtain a second fraud risk score, including: inputting the semantic text stream into a fraud recognition model deployed locally on the terminal device, so as to perform semantic understanding and fraud intent recognition on the semantic text stream through the fraud recognition model to obtain a second fraud risk score, wherein the fraud recognition model is a model obtained by performing knowledge distillation on a basic large language model to obtain a lightweight model, and fine-tuning the lightweight model using labeled fraud sample data.
[0010] In some embodiments, a fraud risk assessment is performed based on the interaction event stream to obtain a third fraud risk score, including: extracting the target interactive behavior features of the current call from the interaction event stream; and comparing the target interactive behavior features with the interactive behavior features stored in a second database to obtain the third fraud risk score, wherein the second database stores at least one interactive behavior feature of a historical call that has engaged in fraudulent behavior.
[0011] In some embodiments, the fraud risk level is obtained by comprehensively judging the first fraud risk score, the second fraud risk score, and the third fraud risk score, including: determining the weights of the three modalities; weighting and fusing the first fraud risk score, the second fraud risk score, and the third fraud risk score based on the weights of the three modalities to obtain a total fraud risk score; and determining the fraud risk level based on the total fraud risk score.
[0012] In some embodiments, collecting multimodal call data during a call includes: during the call, by calling a system interface provided by the operating system, collecting a first audio signal output from the speaker of the terminal device and a second audio signal input from the microphone, and mixing the first audio signal and the second audio signal to obtain the call audio stream; converting the call audio stream into text data using speech recognition technology to obtain the semantic text stream; and during the call, by calling a system interface provided by the operating system, collecting operation events in the user interface of the terminal device, and arranging them in chronological order to obtain the interaction event stream.
[0013] In some embodiments, the interaction event stream includes at least one of the following operation events: application switching event, preset control operation event, paste event, share screenshot event, screen recording event, remote control event, enable accessibility service permission, or application installation event.
[0014] In some embodiments, performing trusted verification on the calling number includes: obtaining a set of numbers from at least two different data sources in the terminal device to dynamically generate a set of trusted numbers; comparing the calling number with the numbers in the set of trusted numbers to identify whether the calling number belongs to the set of trusted numbers; and determining whether the calling number passes trusted verification based on the identification result.
[0015] In some embodiments, obtaining a set of numbers from at least two different data sources in the terminal device to dynamically generate a trusted set of numbers includes: obtaining a first set of numbers from the terminal device's address book; obtaining order data related to logistics and delivery from a lifestyle application installed on the terminal device, and extracting the delivery person's number from the order data to obtain a second set of numbers; obtaining logistics notification SMS messages from the terminal device's local SMS database, and extracting the delivery person's number from the logistics notification SMS messages to obtain a third set of numbers; obtaining a fourth set of numbers from the terminal device's historical call records; and generating the trusted set of numbers based on the first set of numbers, the second set of numbers, the third set of numbers, and the fourth set of numbers.
[0016] In some embodiments, based on the fraud risk level, a corresponding risk intervention operation is performed locally on the terminal device, including: when the fraud risk level is higher than or equal to a preset level, monitoring the interaction between the user and the financial application locally on the terminal device, and performing a corresponding risk intervention operation when a specific operation is detected.
[0017] In some embodiments, the specific operation includes at least one of the following: a transfer operation, a payment operation, an operation to modify account security settings, or an operation to activate an automatic deduction agreement.
[0018] In some embodiments, the risk intervention operation includes at least one of the following: pop-up prompts, voice prompts, light prompts, or floating text prompts.
[0019] In some embodiments, the call event is an incoming call event; The method of collecting multimodal call data during a call includes: collecting multimodal call data during a call in response to detecting a call connection establishment event.
[0020] Secondly, this specification provides a terminal device, comprising: at least one storage medium storing at least one instruction set; and at least one processor communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and executes a fraud risk intervention method as described in any of the first aspects above, according to the instructions of the at least one instruction set.
[0021] Thirdly, this specification provides a computer-readable non-transitory storage medium, wherein the computer-readable non-transitory storage medium stores at least one instruction set, which, when executed by at least one processor, implements the method for intervening in fraud risk as described in any one of the first aspects above.
[0022] The methods for preventing fraud risks, as well as other functions of the terminal devices and storage media provided in this specification, are partially listed in the following description. The inventive aspects of the methods for preventing fraud risks, as well as the terminal devices and storage media provided in this specification, can be fully explained through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A schematic diagram illustrating an intervention scenario for fraud risk provided by an embodiment of this specification is shown; Figure 2 A hardware schematic diagram of a computing device provided according to an embodiment of this specification is shown; Figure 3 A flowchart is shown of an intervention method for fraud risk provided according to an embodiment of this specification; Figure 4A schematic diagram of another fraud risk intervention process provided according to an embodiment of this specification is shown. Detailed Implementation
[0025] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.
[0026] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.
[0027] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0028] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0029] The embodiments described in this specification are applicable to scenarios requiring intervention to mitigate fraud risks. Intervention to mitigate fraud risks refers to automatically identifying the presence and level of fraud risk during a user's call using a terminal device, and then performing corresponding risk intervention operations based on the fraud risk level to avoid fraud risks for the user.
[0030] As mentioned earlier, number-based solutions heavily rely on historical data and known blacklists, making them ill-equipped to combat new types of fraud using newly registered numbers, virtual operator number segments, or overseas number spoofing. Solutions based on cloud-based voice recognition and keyword matching require uploading user call recordings to the cloud, posing risks of data leakage, misuse, or illegal storage.
[0031] Unlike the solutions described above, the embodiments in this specification provide a method for identifying and intervening in fraud risks using multimodal call data locally on a terminal device. The multimodal call data includes data from at least two modalities: audio, semantics, or interactive behavior. By fusing multimodal call data, the terminal device can more comprehensively capture the implicit patterns and behavioral anomalies of fraudulent rhetoric, effectively identifying variant fraudulent content that has been disguised, has replaced keywords, or uses non-standard expressions, significantly improving the accuracy and robustness of fraud risk identification. Simultaneously, all multimodal data collection, analysis, and decision-making processes are completed locally on the terminal device, eliminating the need to upload call audio, semantic text, or interactive behavior to a cloud server. This fundamentally avoids the risk of user privacy data being leaked, misused, or illegally retained during transmission or storage, balancing security effectiveness with compliance requirements for personal information protection.
[0032] Figure 1 A schematic diagram illustrating an intervention scenario for fraud risk provided according to embodiments of this specification is shown. Figure 1 As shown, the fraud risk intervention scenario 001 may include terminal device 110 and anti-fraud server 120.
[0033] Terminal device 110 refers to a hardware entity that directly faces the user, possesses data input / output capabilities, and provides interactive capabilities to the user. Its core characteristic is that it provides functional services through local or network communication, serving as a carrier connecting the digital and physical worlds. Terminal device 110 may include smartphones, tablets, laptops, desktop computers, in-vehicle devices, virtual reality devices, augmented reality devices, etc.
[0034] Terminal device 110 can have multiple applications installed. These applications include, but are not limited to: calling applications, financial applications, web browser applications, search applications, chat applications, shopping applications, lifestyle applications, video applications, social networking applications, etc. Users can make voice calls through calling applications. It should be noted that the application types listed above are categorized according to the functions they provide. In actual application, an application can belong to one or more of the types listed above. For example, when an application supports both voice calls and shopping functions, it can be classified as both a calling application and a lifestyle application. When an application supports both financial payment and shopping functions, it is classified as both a financial application and a lifestyle application.
[0035] See also Figure 1 The terminal device 110 is equipped with a local anti-fraud engine. This local anti-fraud engine can implement the fraud risk intervention methods described in the embodiments of this specification. Specifically, the local anti-fraud engine can have the ability to sense events or data in one or more applications installed on the terminal device. The local anti-fraud engine can also make fraud risk decisions based on the sensed data / events, thereby intervening in fraud risks.
[0036] The local anti-fraud engine refers to a software functional unit integrated into the terminal device that has the ability to identify and intervene in fraud risks. In some embodiments, the local anti-fraud engine can be deployed in the operating system layer or security service framework of the terminal device. Deployment methods include pre-installation by the device manufacturer, push notifications via operating system updates, or download and installation by the user from the official app store. In some embodiments, the local anti-fraud engine can also be embedded in various applications installed on the terminal device in the form of a Software Development Kit (SDK). After the local anti-fraud engine is deployed locally on the terminal device 110, it can execute fraud risk intervention schemes locally on the terminal device 110 during user calls, completing fraud risk identification and intervention without relying on the cloud-based anti-fraud server 120, thus balancing response efficiency and user privacy protection.
[0037] As an example, when a terminal device receives a call event through a calling application, the local anti-fraud engine can detect (perceive) the call event. In response to detecting the call event, it can obtain the caller's number and perform a trusted verification of that number. During the trusted verification process, the local anti-fraud engine can use a set of trusted numbers obtained from one or more applications installed on the terminal device as the verification basis. If the caller's number fails the trusted verification, the local anti-fraud engine collects multimodal call data during the call and performs a fraud risk assessment locally on the terminal device based on the multimodal data to obtain a fraud risk level. The multimodal call data can include at least two of the following data streams: call audio stream, semantic text stream generated from the call audio stream, or user interaction event streams during the call. Furthermore, the local anti-fraud engine can perform corresponding risk intervention operations locally on the terminal device based on the fraud risk level. For example, during a call, when the local anti-fraud engine detects (perceives) that the user is performing certain operations (such as a money transfer) in a financial application, it can trigger a risk intervention operation to alert the user of the potential fraud risk.
[0038] See also Figure 1A local anti-fraud engine can include a scheduler, one or more databases, and one or more artificial intelligence models. The scheduler is responsible for controlling and scheduling the execution logic of intervention methods for fraud risk. During the scheduling process, it can access one or more databases to obtain the required data and can also invoke one or more fraud detection models to support risk decisions.
[0039] In some embodiments, the aforementioned database and fraud identification model can be pre-deployed to the terminal device by the anti-fraud server 120. The anti-fraud server 120 can maintain a database (knowledge base) related to fraud risk identification, including but not limited to a blacklist of fraudulent numbers, a database for storing the voiceprint characteristics of fraudsters, and a database for storing the interactive behavior characteristics of fraudulent calls. The anti-fraud server 120 can also maintain one or more pre-trained fraud identification models. For example, the anti-fraud server 120 uses labeled fraud sample data to perform knowledge distillation and fine-tuning on a basic large language model to generate a lightweight fraud identification model. The anti-fraud server 120 can deploy the aforementioned database or fraud identification model to the terminal device 110 according to a preset deployment strategy (such as a periodic strategy, a request-based strategy, etc.).
[0040] The fraud risk intervention method described in the embodiments of this specification can be executed by the terminal device 110, specifically by the local anti-fraud engine deployed on the terminal device 110. In this case, the terminal device 110 may store data or instructions for implementing the fraud risk intervention method, and can execute or be used to execute the data or instructions. In some embodiments, the terminal device 110 may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device. The specific implementation of the fraud risk intervention method will be described later and will not be elaborated here.
[0041] Figure 2 A hardware schematic diagram of a computing device 200 provided according to an embodiment of this specification is shown. The computing device 200 can be used as... Figure 1 The terminal device 110 and computing device 200 can execute the fraud risk intervention methods described in this specification.
[0042] like Figure 2 As shown, the computing device 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing device 200 may also include a communication port 250 and an internal communication bus 210. The computing device 200 may also include I / O components 240.
[0043] The internal communication bus 210 can connect different system components, including storage medium 230, processor 220 and communication port 250.
[0044] I / O component 240 supports input / output between computing device 200 and other components.
[0045] Communication port 250 is used for data communication between computing device 200 and the outside world. For example, communication port 250 can be used for data communication between computing device 200 and a network. Communication port 250 can be a wired communication port or a wireless communication port.
[0046] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may be computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc.
[0047] Processor 220 is communicatively connected to storage medium 230. Processor 220 is used to execute at least one instruction set described above. When computing device 200 is running, processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the fraud risk intervention method provided in this specification. Processor 220 can execute all the steps included in the fraud risk intervention method. Processor 220 can be in the form of one or more processors. In some embodiments, processor 220 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, or any combination thereof.
[0048] Just to illustrate the point, Figure 2Only the case where the computing device 200 includes one processor 220 is shown. However, it should be noted that the computing device 200 may also include multiple processors 220 in this specification. Therefore, the operation and / or method steps disclosed in this specification may be executed by one processor or by multiple processors in combination. For example, if the processor 220 of the computing device 200 in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0049] Figure 3 A flowchart of a fraud risk intervention method P300 provided according to an embodiment of this specification is shown. This method can be performed by a terminal device acting as the called party. Figure 3 As shown, the intervention method P300 for fraud risk includes steps S310-S340.
[0050] S310: In response to detecting a call event, obtain the caller's number for this call.
[0051] In some embodiments, a call event can be an incoming call event, that is, an event indicating that the terminal device has received an incoming call notification. In some embodiments, a call event can also be a call connection establishment event, that is, an event indicating that the called party answers the call.
[0052] The caller ID refers to the phone number used by the caller. When a terminal device receives an incoming call notification, the caller ID is usually displayed in the interface of the calling application.
[0053] In some embodiments, the local anti-fraud engine can passively monitor call events and legitimately obtain the caller's number through standardized interfaces provided by the operating system. For example, the local anti-fraud engine runs locally on the terminal device as a background system service. The local anti-fraud engine registers a corresponding listener with the operating system's telephony management framework to subscribe to changes in call status. When events such as incoming calls, call connection establishment, or call termination occur, the operating system actively calls back the listening interface registered by the local anti-fraud engine, passing callback parameters containing the event type. Thus, the local anti-fraud engine can detect call events in real-time and with low power consumption without active polling. After receiving the callback of a call event (such as an incoming call or call connection establishment event), the local anti-fraud engine can extract the caller's number from the callback parameters.
[0054] S320: Perform a trusted verification on the calling number.
[0055] The trusted verification process is used to determine whether the caller ID belongs to a known trusted contact. The local anti-fraud engine compares the obtained caller ID with a set of trusted numbers dynamically generated locally on the terminal device to determine whether the number belongs to a known contact in the address book, a recent service provider, or a frequently contacted non-address book number, and outputs a judgment conclusion based on the comparison result. The judgment conclusion can be either "trusted verification passed" or "trusted verification failed".
[0056] In some embodiments, the terminal device may obtain a set of numbers from at least two different data sources within the terminal device to dynamically generate a set of trusted numbers. The terminal device compares the calling number with the numbers in the set of trusted numbers to identify whether the calling number belongs to the set of trusted numbers. Based on the identification result, the terminal device determines whether the calling number has passed the trusted verification.
[0057] In the embodiments of this specification, S320 can be considered as a preliminary fraud risk screening performed by the terminal device on the caller ID itself. By verifying the caller ID, it is possible to quickly and effectively identify whether the caller belongs to a known trusted contact. The terminal device can then decide on subsequent execution procedures based on the preliminary screening results. For example, if the preliminary screening result indicates a pass, there is no need to continue implementing stricter risk decision-making strategies during the call, thereby avoiding eavesdropping. If the preliminary screening result indicates a fail, then more stringent analysis and decision-making strategies are adopted during subsequent calls, thereby reducing the losses caused to the user by fraud risks.
[0058] In some embodiments, the aforementioned data sources may include address books, lifestyle applications, SMS databases, and historical call records. Address books are databases of contact numbers actively saved by the called party. Numbers in the address book represent people the called party explicitly trusts. Lifestyle applications refer to various service applications, such as food delivery platforms, express delivery platforms, ride-hailing platforms, and financial platforms. These applications may have official customer service numbers and temporary delivery driver numbers; although these numbers are not in the address book, they have genuine interactions with the called party and are therefore highly credible. SMS databases may also store delivery driver numbers from logistics services. Numbers in historical call records are numbers that have actually communicated with the called party and are also highly credible.
[0059] The terminal device can obtain the first set of numbers from its address book. Specifically, the terminal device can read the contact numbers in its locally stored address book by calling the contact reading interface provided by the operating system, thus obtaining the first set of numbers.
[0060] The terminal device can obtain order data related to logistics and delivery from the lifestyle applications installed on the terminal device, and extract the delivery person's number from the order data to obtain a second set of numbers. Specifically, the terminal device obtains logistics and delivery order data by monitoring various lifestyle applications installed on the terminal device, and extracts relevant numbers such as "delivery person's phone number" and "rider's contact information" from the order data to obtain a second set of numbers.
[0061] The terminal device can also obtain logistics notification SMS messages from its local SMS database and extract the delivery person's number from these messages to obtain a third set of numbers. Specifically, the terminal device can read logistics notification SMS messages through an SMS database interface. Subsequently, the terminal device can extract the delivery person's number from these messages to obtain a third set of numbers.
[0062] The terminal device can obtain the fourth number set from its historical call records. Specifically, the terminal device can call the operating system's call log interface to read incoming and outgoing call records. The terminal device can then add all numbers from the historical call records to the fourth number set.
[0063] Furthermore, the terminal device can generate a trusted number set based on the first number set, the second number set, the third number set, and the fourth number set. Specifically, the terminal device can merge and deduplicate the numbers in the first, second, third, and fourth number sets to form a trusted number set. It should be noted that the trusted number set can be generated either in advance before the call event begins or dynamically after the call event is triggered; this is not limited here.
[0064] In the above scheme, the terminal device generates a set of trusted numbers based on the first set of numbers, the second set of numbers, the third set of numbers, and the fourth set of numbers. This can effectively integrate trusted numbers from different data sources, thereby helping to improve the accuracy of trusted verification results.
[0065] After generating the trusted number set, the terminal device compares the calling number with each number in the trusted number set to determine whether the calling number belongs to the trusted number set. If the calling number belongs to the trusted number set, it means the calling number belongs to a trusted contact, thus confirming that the calling number has passed trusted verification. In this case, the terminal device does not need to continue executing S330-S340. If the calling number does not belong to the trusted number set, it means the calling number does not belong to a trusted contact, thus confirming that the calling number has failed trusted verification. In this case, the terminal device continues executing S330-S340.
[0066] S330: If the calling number fails the trusted verification, multimodal call data is collected during the call, and fraud risk determination is performed locally on the terminal device based on the multimodal call data to obtain the fraud risk level. The multimodal call data includes at least two data streams, such as call audio stream, semantic text stream, or interactive event stream.
[0067] In some embodiments, when the call event in S310 is an incoming call event, in S330, the terminal device may, in response to the caller ID failing trusted verification and detecting a call connection establishment event, collect multimodal call data during the call. In some embodiments, when the call event in S310 is a call connection establishment event, in S330, the terminal device may, in response to the caller ID failing trusted verification, collect multimodal call data during the call.
[0068] In this specification, a call audio stream refers to the raw voice data transmitted in real time during telephone communication, typically existing as a continuous audio signal. A semantic text stream refers to text content with semantic information generated after automatic speech recognition (ASR) and natural language processing of the call audio stream. An interactive event stream refers to a sequence of actions or key events performed by the caller and / or the called party during the call.
[0069] Different data streams require different acquisition methods. The acquisition methods for each type of data stream are explained below.
[0070] In some embodiments, the call audio stream can be acquired as follows: during a call, by calling a system interface provided by the operating system, a first audio signal output from the speaker of the terminal device and a second audio signal input from the microphone are acquired, and the first and second audio signals are mixed to obtain the call audio stream. The mixing process may include aligning the first and second audio signals, noise reduction, echo cancellation, and volume equalization. The call audio stream obtained after mixing is a complete, clear, and time-consistent stereo or mono call audio stream.
[0071] Semantic text streams can be acquired as follows: The terminal device converts the call audio stream into text data using speech recognition technology to obtain the semantic text stream. Specifically, the terminal device can input the call audio stream into a speech recognition engine. The speech recognition engine converts the speech signal frame by frame into a corresponding text sequence, distinguishes and annotates the spoken content of the called party and the calling party, and generates a structured semantic text stream. The semantic text stream obtained in this way not only retains the core semantic information of the call audio stream but also possesses timestamp alignment, role identification, and searchability.
[0072] Interactive event streams can be collected as follows: During a call, the terminal device calls the system interface provided by the operating system to collect operation events in the user interface of the terminal device, and arranges them in chronological order to obtain the interactive event stream. The interactive event stream completely records the called party's operation trajectory during the call, or in other words, records the operation trajectory guided by the calling party to the called party.
[0073] In some embodiments, the interaction event stream may include at least one of the following operation events: application switching event, operation event of preset control, paste event, share screenshot event, screen recording event, remote control event, enable accessibility service permission, and application installation event.
[0074] Among these, the application switching event refers to the called party switching to other applications during a call, such as switching to a banking app or a browser. The preset control operation event captures the called party's actions on specific interface elements, such as clicking high-risk buttons like "transfer," "confirm payment," or "copy verification code." The paste event detects the reading or pasting of clipboard content, such as pasting verification codes or bank card numbers. The screenshot sharing event detects the called party taking a screenshot of the current screen and then sharing, saving, or uploading it. The screen recording event detects the start or stop of third-party screen recording. The remote control event monitors whether remote desktop, screen mirroring, or third-party remote assistance tools are enabled. Granting accessibility service permissions refers to the called party granting accessibility permissions to specific applications. The application installation event refers to the installation of a new application on the terminal device during a call.
[0075] The events listed above all represent situations where there is a high risk of fraud during the call, potentially leading to financial losses for the user. Collecting these events and creating an interaction event stream helps analyze the potential for fraud in the call from the perspective of interactive behavior.
[0076] When the collected multimodal call data includes call audio streams, semantic text streams, and interactive event streams, the multimodal call data can comprehensively and accurately reconstruct the voice content, semantic information, and called party's operational behavior in the call scenario, providing a high-quality, interpretable, and interconnected multidimensional data foundation for fraud risk assessment.
[0077] In some embodiments, multimodal call data includes two data streams: a call audio stream and a semantic text stream. In some embodiments, multimodal call data includes two data streams: a call audio stream and an interaction event stream. In some embodiments, multimodal call data includes two data streams: a semantic text stream and an interaction event stream. In some embodiments, multimodal call data includes three data streams: a call audio stream, a semantic text stream, and an interaction event stream.
[0078] When multimodal call data includes three data streams: call audio stream, semantic text stream, and interactive event stream, the terminal device performs a fraud risk assessment based on the call audio stream to obtain a first fraud risk score. The terminal device then performs a fraud risk assessment based on the semantic text stream to obtain a second fraud risk score. Finally, the terminal device performs a fraud risk assessment based on the interactive event stream to obtain a third fraud risk score. After obtaining the first, second, and third fraud risk scores, the terminal device can make a comprehensive judgment based on these scores to obtain the fraud risk level.
[0079] When terminal devices use the first fraud risk score, the second fraud risk score, and the third fraud risk score for comprehensive judgment, they can fully perceive potential fraud signals from three dimensions: voice features, semantic content, and user interaction behavior. This significantly improves the accuracy and robustness of fraud risk identification, thereby providing users with smarter and more proactive security protection.
[0080] The process of determining the first fraud risk score, the second fraud risk score, and the third fraud risk score is described below.
[0081] In some embodiments, the first fraud risk score can be obtained through the following process: The terminal device extracts the target voiceprint feature of the caller from the call audio stream. The terminal device compares the similarity of the target voiceprint feature with each voiceprint feature stored in a first database, and determines the first fraud risk score based on the comparison result. The first database stores the voiceprint features of at least one fraudulent user. That is, the first database stores the voiceprint features corresponding to at least one known fraudulent number. The terminal device compares the extracted target voiceprint feature of the caller with each voiceprint feature stored in the first database. Based on the similarity between the target voiceprint feature and the voiceprint features stored in the first database, the terminal device generates a corresponding first fraud risk score. The higher the first fraud risk score, the more similar the target voiceprint feature is to the voiceprint features stored in the first database, and the greater the suspicion of fraud by the caller.
[0082] Because voiceprints are difficult to forge and uniquely linked to the physiological and habitual characteristics of the speaker, the first fraud risk score identified by voiceprint feature comparison has higher accuracy and anti-deception capabilities compared to traditional methods that rely solely on numbers, keywords, etc.
[0083] In some embodiments, the second fraud risk score can be obtained through the following process: the terminal device inputs a semantic text stream into a fraud recognition model deployed locally on the terminal device, so as to perform semantic understanding and fraud intent recognition on the semantic text stream through the fraud recognition model to obtain the second fraud risk score.
[0084] The aforementioned fraud detection model is obtained by performing knowledge distillation on a basic large language model to obtain a lightweight model, and then fine-tuning and training the lightweight model using labeled fraud sample data. The terminal device inputs a semantic text stream into the fraud detection model deployed locally on the terminal device. The fraud detection model can identify whether the semantic text stream contains typical fraudulent phrases or suspicious intentions such as inducing transfers, impersonation, or fake prize winnings, and outputs a second fraud risk score accordingly. The higher the second fraud risk score, the greater the suspicion of fraud by the caller.
[0085] By inputting the real-time translated semantic text stream into a fraud detection model deployed locally on the terminal, a deep contextual understanding of the call content and accurate insight into fraudulent intent are achieved. Because this model is a specialized model that performs knowledge distillation on a large language model and fine-tunes its training using fraudulent script data, it can effectively identify novel and complex leading phrases and logical traps from a semantic dimension, without being limited by a fixed list of keywords. This results in a highly accurate secondary fraud risk score.
[0086] In some embodiments, the third fraud risk score can be obtained through the following process: The terminal device extracts the target interactive behavior features of the current call from the interaction event stream. The terminal device compares the target interactive behavior features with the interactive behavior features stored in a second database to obtain the third fraud risk score. The second database stores the interactive behavior features of at least one historical call with fraudulent behavior. The terminal device compares the extracted target interactive behavior features with the interactive behavior features stored in the second database. Based on the similarity between the target interactive behavior features and the interactive behavior features stored in the second database, the terminal device generates a corresponding third fraud risk score. The higher the third fraud risk score, the more similar the target interactive behavior features are to the interactive behavior features stored in the second database, and the greater the suspicion of fraud by the caller.
[0087] By analyzing the caller's operational sequence during the call (such as switching to financial applications, clicking the transfer button, etc.), the system can directly capture the substantial high-risk behaviors triggered by fraudulent inducement on the user's side. Since these interactive behavioral characteristics are key, objective, and difficult-to-fake evidence of the user's being scammed, the third fraud risk score derived from this comparison can accurately reflect that the fraud risk has moved from the verbal inducement stage to the substantive operational stage, providing reliable evidence to support the system's immediate intervention. After determining the first fraud risk score, the second fraud risk score, and the third fraud risk score, the terminal device can perform a weighted fusion of the first fraud risk score, the second fraud risk score, and the third fraud risk score to obtain a fraud risk level.
[0088] In some embodiments, the terminal device can determine the weights of the three modalities, and based on the weights of the three modalities, perform a weighted fusion of the first fraud risk score, the second fraud risk score, and the third fraud risk score to obtain a total fraud risk score. The terminal device determines the fraud risk level based on the total fraud risk score.
[0089] For example, the first fraud risk score is S1, the second fraud risk score is S2, and the third fraud risk score is S3. Among them, the weight corresponding to the call audio stream is W1, the weight corresponding to the semantic text stream is W2, and the weight corresponding to the interactive event stream is W3. Then the total score S = (S1×W1) + (S2×W2) + (S3×W3).
[0090] By weighted and fused the first, second, and third fraud risk scores, the terminal device constructs a multimodal comprehensive risk assessment system that integrates voiceprint features, deep semantic intent, and user behavior evidence. This system can cross-verify fraud risks from three orthogonal dimensions: physiological identity, semantic logic, and operational actions, increasing the difficulty for fraudsters to evade detection and thus systematically improving the accuracy and robustness of fraud risk determination.
[0091] In some embodiments, the weights of the three modalities can be preset. For example, the weight W1 corresponding to the call audio stream is greater than the weight W2 corresponding to the semantic text stream, and the weight W2 corresponding to the semantic text stream is greater than the weight W3 corresponding to the interactive event stream.
[0092] In some embodiments, the weights of the three modalities can be dynamically adjusted based on the context information of the current call.
[0093] For example, when the calling number makes a high number of calls within a recent preset time period and the call duration is short, the weight W1 corresponding to the call audio stream can be appropriately increased, while the weights corresponding to the other two modalities can be decreased accordingly. As another example, if the current call content contains keywords such as "verification code," "transfer," "screen sharing," "don't tell family," or "act now," the weight W2 corresponding to the semantic text stream can be appropriately increased, while the weights corresponding to the other two modalities can be decreased accordingly. Furthermore, if the current call is a late-night call, the weights W2 corresponding to the semantic text stream and W3 corresponding to the interactive event stream can be appropriately increased, while the weight W1 corresponding to the call audio stream can be decreased.
[0094] The dynamic adjustment of the three modal weights improves the accuracy and adaptability of terminal devices in fraud detection, enabling them to intelligently focus on the most reliable judgment criteria according to different call scenarios.
[0095] After obtaining the overall fraud risk score, the terminal device can determine the fraud risk level based on this score. Specifically, after obtaining the overall fraud risk score, the terminal device can determine the fraud risk level based on the numerical range of the overall fraud risk score. Assume that fraud risk levels are divided into low risk, medium risk, and high risk. When the overall fraud risk score is below a first threshold, the fraud risk level is "low risk." When the overall fraud risk score is between the first and second thresholds, the fraud risk level is "medium risk." When the overall fraud risk score exceeds the second threshold, the fraud risk level is "high risk." The first threshold is less than the second threshold.
[0096] In some embodiments, the terminal device may include a hardware-based trusted execution environment (TEA). All sensitive data related to the call content (including but not limited to the call audio stream, the semantic text stream translated from it, and voiceprint features extracted from the call audio stream) is encrypted and isolated and stored in a secure storage area within the TEA. All processing steps, including feature extraction, model inference, and risk scoring calculation, are completed only within the protected memory space of the TEA. This approach ensures that sensitive data is hardware-level isolated from the terminal device's ordinary operating system and upper-layer application environment throughout its entire lifecycle of storage and computation, thereby preventing malicious theft or tampering from the system software layer and further enhancing the security of privacy data.
[0097] S340: Based on the fraud risk level, perform corresponding risk intervention operations locally on the terminal device.
[0098] In some embodiments, risk intervention operations include at least one of the following: pop-up prompts, voice prompts, light prompts, and floating text prompts.
[0099] In this context, a pop-up notification refers to a full-screen or half-screen confirmation window that is forcibly displayed on the terminal device. The window displays a message similar to "The current call may be involved in fraud. Do you want to continue?" and provides options such as "Confirm to continue," "View details," or "Contact the anti-fraud center." In this situation, the user can only continue subsequent operations after manually selecting "Confirm to continue" in the window.
[0100] Voice broadcast prompts refer to the terminal device automatically playing preset warning voice messages, such as: "Please note that this call may involve fraud. Please be vigilant and do not perform any operations such as transferring money."
[0101] A visual cue refers to a terminal device emitting a visual signal with a specific color (such as flashing red) or frequency to intuitively convey the risk status.
[0102] Floating text prompts refer to the display of a prompt bar or bubble at the top of the screen, with content such as "This call is at risk of fraud, please be vigilant," which continues throughout the call without interrupting interface operations.
[0103] In some embodiments, the terminal device can employ different risk intervention operations based on the fraud risk level. For example, when the fraud risk level is high (e.g., medium and high risk), the terminal device employs a stronger risk intervention operation, such as a pop-up notification or a voice prompt, or both. When the fraud risk level is low (e.g., low risk), the terminal device employs a weaker risk intervention operation, such as a light notification or a floating text notification.
[0104] In some embodiments, when the fraud risk level is higher than or equal to a preset level, the terminal device can trigger monitoring of the financial application. That is, the terminal device locally monitors the user's interaction with the financial application and executes corresponding risk intervention operations when a specific operation is detected.
[0105] In some embodiments, the specific operations monitored by the terminal device for financial applications may include at least one of the following: transfer operations, payment operations, operations to modify account security settings, and operations to enable automatic deduction protocols.
[0106] For example, assuming the preset risk level is medium, if the S330 identifies a fraud risk level as medium or high, the terminal device locally monitors the user's interactions with the financial application. If a transfer or payment operation is detected, a pop-up notification or voice prompt is executed; if an automatic deduction agreement activation is detected, a pop-up notification is executed; if an account security settings modification is detected, a voice prompt is executed. It should be noted that the above examples illustrating the correspondence between specified operations and risk intervention operations are merely possible examples. Other correspondences may be used in practical applications, and this manual does not limit this.
[0107] The terminal device monitors the interaction between the user and the financial application locally. It has the characteristics of low latency and can accurately block the fraud at the critical moment (such as before the user performs a sensitive operation), thereby achieving effective risk prevention and control.
[0108] Figure 4 A schematic diagram illustrates another fraud risk intervention process provided according to embodiments of this specification. The following is in conjunction with... Figure 4 An example is provided to illustrate a relatively complete fraud risk intervention process.
[0109] like Figure 4 As shown, the local anti-fraud engine in the terminal device registers a corresponding listener with the operating system's telephony management framework to subscribe to changes in call status. When a call event (such as an incoming call or a call connection establishment event) occurs, the operating system calls back the listener interface registered by the local anti-fraud engine, passing a callback parameter containing the event type. Thus, the local anti-fraud engine detects the occurrence of the call event. After receiving the callback for the call event, the local anti-fraud engine can extract the caller ID of this call from the callback parameter.
[0110] Furthermore, the local anti-fraud engine reads contacts, order / logistics notification messages from lifestyle applications, logistics notification SMS messages from the SMS database, and historical call records through the operating system's interface to dynamically generate a set of trusted numbers. The anti-fraud engine then compares the caller ID of the current call with the numbers in this set of trusted numbers to obtain a verifiable result for the caller ID's trustworthiness.
[0111] If the trusted verification result is successful, the local anti-fraud engine can end the intervention process without performing any further operations.
[0112] If the trusted verification result fails, the local anti-fraud engine can collect multimodal call data during the call. This multimodal call data can include call audio streams, semantic text streams, or interactive event streams. The local anti-fraud engine performs a fraud risk assessment based on the call audio stream to obtain a first fraud risk score; it performs a fraud risk assessment based on the semantic text stream to obtain a second fraud risk score; and it performs a fraud risk assessment based on the interactive event stream to obtain a third fraud risk score. The terminal device can then make a comprehensive judgment based on the first, second, and third fraud risk scores to determine the fraud risk level of the call.
[0113] In cases of high fraud risk (e.g., medium or high risk), the local anti-fraud engine can monitor the user's interactions with financial applications locally. When it detects specific operations (e.g., transfers, payments, changes to account security settings, or activation of automatic deduction agreements), it can perform high-level risk intervention operations, such as pop-up prompts or voice prompts, or even both.
[0114] When the fraud risk level is low (e.g., low risk), the local anti-fraud engine can directly use low-intervention risk intervention operations to provide risk warnings to users, such as using light prompts or floating text prompts.
[0115] In summary, the embodiments of this specification provide a method, terminal device, and storage medium for intervening in fraud risks. Fraud risks are determined locally on the terminal device based on multimodal call data. This multimodal call data includes data from at least two modalities: audio, semantics, or interactive behavior. By fusing multimodal call data, the terminal device can more comprehensively capture implicit patterns and behavioral anomalies in fraudulent statements, effectively identifying variant fraudulent content that has been disguised, has replaced keywords, or uses non-standard expressions, significantly improving the accuracy and robustness of fraud risk identification. Simultaneously, all multimodal data collection, analysis, and decision-making processes are completed locally on the terminal device, eliminating the need to upload call audio, text, or behavior logs to a cloud server. This fundamentally avoids the risk of user privacy data being leaked, misused, or illegally retained during transmission or storage, balancing security protection effectiveness with personal information protection compliance requirements.
[0116] It should be noted that all user data collected and processed during the implementation of the technical solutions described in this manual (including but not limited to call content, interactive behavior, and device information) strictly adheres to the core principles of "legality, legitimacy, necessity" and "informed consent." Related data activities are conducted only within the scope of the user's explicit and voluntary authorization; all data processing flows are completed locally on the user's terminal device or in a legally established secure computing environment, aiming to achieve the anti-fraud security purpose stated in this solution. The original data will not be retained, shared, or used for any other purpose beyond the aforementioned objectives, thus fully protecting user privacy rights and meeting applicable data protection laws and regulations.
[0117] This specification, in another aspect, provides a non-transitory storage medium storing at least one set of executable instructions for fraud risk intervention. When the executable instructions are executed by a processor, they instruct the processor to implement the steps of the fraud risk intervention method P300 described in this specification. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on a computing device 200, the program code causes the computing device 200 to perform the steps of the fraud risk intervention method P300 described in this specification. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on the computing device 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on computing device 200, partially on computing device 200, as a standalone software package, partially on computing device 200 and partially on a remote computing device, or entirely on a remote computing device.
[0118] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0119] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.
[0120] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.
[0121] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.
[0122] Every patent, patent application, publication of a patent application, and other material, such as articles, books, specifications, publications, documents, and literature (excluding any related historical examination documents), cited in this disclosure is incorporated herein for all purposes, including, for example, in the specification and claims of this disclosure. However, in the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the foregoing and those used in this disclosure, the descriptions, definitions, and / or terms used in this disclosure shall prevail.
[0123] Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.
Claims
1. A method for intervening in fraud risk, applied to a terminal device, the method comprising: In response to the detection of a call event, obtain the caller ID of this call; The calling number is verified for reliability. If the calling number fails the trusted verification, multimodal call data is collected during the call, and fraud risk determination is performed locally on the terminal device based on the multimodal call data to obtain the fraud risk level. The multimodal call data includes at least two data streams, such as call audio stream, semantic text stream, or interactive event stream. as well as Based on the fraud risk level, corresponding risk intervention operations are performed locally on the terminal device.
2. The intervention method according to claim 1, wherein, When the multimodal call data includes data streams of three modes: call audio stream, semantic text stream, and interactive event stream, fraud risk determination is performed based on the multimodal call data to obtain a fraud risk level, including: A fraud risk assessment is performed based on the call audio stream to obtain a first fraud risk score; A second fraud risk score is obtained by performing a fraud risk assessment based on the semantic text stream. A third fraud risk score is obtained by performing a fraud risk assessment based on the aforementioned interactive event stream; and The fraud risk level is obtained by comprehensively judging the first fraud risk score, the second fraud risk score and the third fraud risk score.
3. The intervention method according to claim 2, wherein, A fraud risk assessment is performed based on the call audio stream to obtain a first fraud risk score, including: The target voiceprint features of the caller are extracted from the call audio stream; and The target voiceprint feature is compared with the voiceprint features stored in the first database for similarity, and the first fraud risk score is determined based on the comparison results. The first database stores the voiceprint features of at least one fraudulent user.
4. The intervention method according to claim 2, wherein, Based on the semantic text stream, a fraud risk assessment is performed to obtain a second fraud risk score, including: The semantic text stream is input into a fraud detection model deployed locally on the terminal device, so as to perform semantic understanding and fraud intent identification on the semantic text stream through the fraud detection model to obtain a second fraud risk score. The fraud detection model is obtained by performing knowledge distillation on a basic large language model to obtain a lightweight model, and then fine-tuning the lightweight model using labeled fraud sample data.
5. The intervention method according to claim 2, wherein, Based on the aforementioned interactive event stream, a fraud risk assessment is performed to obtain a third fraud risk score, including: Extract the target interactive behavior features of this call from the interaction event stream; and The target interactive behavior features are compared with the interactive behavior features stored in the second database to obtain the third fraud risk score. The second database stores at least one interactive behavior feature of a historical call that has engaged in fraudulent behavior.
6. The intervention method according to claim 2, wherein, The fraud risk level is determined by comprehensively assessing the first fraud risk score, the second fraud risk score, and the third fraud risk score, including: Determine the weights for the three modes; Based on the weights of the three modalities, the first fraud risk score, the second fraud risk score, and the third fraud risk score are weighted and fused to obtain a total fraud risk score; and The fraud risk level is determined based on the overall fraud risk score.
7. The intervention method according to claim 2, wherein, Collect multimodal call data during the call, including: During a call, the system interface provided by the operating system is called to collect the first audio signal output by the speaker of the terminal device and the second audio signal input by the microphone, and the first audio signal and the second audio signal are mixed to obtain the call audio stream; The call audio stream is converted into text data using speech recognition technology to obtain the semantic text stream; and During the call, the system interface provided by the operating system is called to collect operation events in the user interface of the terminal device and arrange them in chronological order to obtain the interactive event stream.
8. The intervention method according to claim 7, wherein, The interactive event stream includes at least one of the following operation events: application switching event, preset control operation event, paste event, share screenshot event, screen recording event, remote control event, enable accessibility service permission, or application installation event.
9. The intervention method according to claim 1, wherein, The trusted verification of the calling number includes: Obtain a set of numbers from at least two different data sources in the terminal device to dynamically generate a set of trusted numbers; The calling number is compared with numbers in the trusted number set to identify whether the calling number belongs to the trusted number set; and Based on the identification results, it is determined whether the calling number has passed the trusted verification.
10. The intervention method according to claim 9, wherein, Obtaining a set of numbers from at least two different data sources in the terminal device to dynamically generate a set of trusted numbers includes: Obtain a first set of numbers from the address book of the terminal device; Order data related to logistics and delivery is obtained from the lifestyle applications installed on the terminal device, and the delivery person's number is extracted from the order data to obtain a second set of numbers; The system obtains logistics notification SMS messages from the local SMS database of the terminal device, and extracts the delivery person's number from the logistics notification SMS messages to obtain a third set of numbers; Obtain the fourth set of numbers from the historical call records of the terminal device; and The trusted number set is generated based on the first number set, the second number set, the third number set, and the fourth number set.
11. The intervention method according to claim 1, wherein, Based on the fraud risk level, corresponding risk intervention operations are performed locally on the terminal device, including: When the fraud risk level is higher than or equal to a preset level, the terminal device locally monitors the user's interaction with the financial application and performs corresponding risk intervention operations when a specific operation is detected.
12. The intervention method according to claim 11, wherein, The specific operation includes at least one of the following: transfer operation, payment operation, modification of account security settings operation, or activation of automatic deduction agreement operation.
13. The intervention method according to claim 1, wherein, The risk intervention operation includes at least one of the following: pop-up prompts, voice broadcast prompts, light prompts, or floating text prompts.
14. The intervention method according to claim 1, wherein, The call event is an incoming call event; The method of collecting multimodal call data during a call includes: collecting multimodal call data during a call in response to detecting a call connection establishment event.
15. A terminal device, comprising: At least one storage medium storing at least one instruction set; as well as At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and executes the fraud risk intervention method as described in any one of claims 1-14 according to the instructions of the at least one instruction set.
16. A computer-readable non-transitory storage medium, wherein, The computer-readable non-transitory storage medium stores at least one set of instructions, which, when executed by at least one processor, implement the method for intervening in fraud risk as described in any one of claims 1-14.