Aging-adaptive anti-fraud intelligent protection system based on multi-modal detection and LSTM model

By constructing an anti-fraud intelligent protection system based on multimodal detection and LSTM models, the problems of insufficient multimodal detection depth and real-time protection capabilities in existing technologies have been solved. It achieves full coverage detection and visual interpretation of multimodal content, adapts to the needs of the elderly, and enhances their risk awareness and protection capabilities.

CN121387239AInactive Publication Date: 2026-01-23白霄腾
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511444365.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing intelligent anti-fraud systems lack depth in multimodal content detection, lack real-time protection capabilities, and have poor interpretability of detection results, making it difficult to understand the basis for risk assessment.

Method used

We construct an age-friendly anti-fraud intelligent protection system based on multimodal detection and LSTM models, including a front-end interface subsystem, a back-end service subsystem, an intelligent protection and anti-fraud core module and a model library. It integrates multimodal active detection, passive real-time detection, deepfake experience and voice interaction submodules, and introduces visualization interpretation technology for detection results.

Benefits of technology

It achieves full coverage detection of multimodal content, adapts to the physiological characteristics of the elderly, provides visualized interpretation of detection results, supports the improvement of risk awareness among the elderly, and enables collaborative protection between users and their families.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121387239A_ABST
    Figure CN121387239A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an aging-adaptive anti-fraud intelligent protection system based on multi-modal detection and an LSTM (Long Short Term Memory) model, which comprises an anti-fraud detection system architecture consisting of a front-end interface subsystem, a rear-end service subsystem, an intelligent protection anti-fraud core module and a model library, the intelligent protection anti-cheating core module comprises a multi-mode active detection uploading sub-module, a passive real-time detection sub-module, a deep forging experience sub-module, an anti-cheating learning community sub-module and a voice interaction sub-module. According to the aging-adaptive anti-fraud intelligent protection system based on the multi-modal detection and the LSTM model, by constructing a hybrid model of CNN, LSTM and frequency domain analysis, image and video Deepfake abnormal data are accurately detected, logic dimensions depend on LLMs and are combined with RAG and CoT technologies, text and audio modal special detection are synchronously adapted, screen video streams are captured based on H.264 coding, real-time detection is performed on each frame through a lightweight detector, and the anti-fraud intelligent protection system based on the multi-modal detection and the LSTM model is obtained. And triggering a pop-up window and a voice alarm and synchronizing a family member end to form double-end collaborative protection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of artificial intelligence, and particularly relates to an anti-fraud intelligent protection system for the elderly based on multi-modal detection and an LSTM model. BACKGROUND

[0002] In today's high-speed development of information technology, artificial intelligence has penetrated into all aspects of our life. Texts, images, audios and videos generated by intelligent AI are increasingly common, and their quality and fidelity have also significantly improved, gradually reaching the level of false information output. For example, AI face changing technology (Deepfake) can generate videos that are almost indistinguishable from real people, and fake audios can also imitate the voices of specific persons through AI technology. These technologies are easily exploited by unscrupulous people to create false information or fraudulent content. Therefore, in the face of the dilemma of imperfect AI laws and regulations and supervision, only effective identification of false and counterfeit content can protect our safety and rights and interests. Advanced technical tools need to be actively used to ensure information security, especially when sensitive or suspicious information is encountered. Verification through anti-fraud tools can reduce the risk of being cheated.

[0003] According to the invention patent with the Chinese patent publication number CN117910452A, an intelligent anti-fraud technology and system based on large language model adaptive iteration are disclosed. The intelligent anti-fraud system based on large language model adaptive iteration can effectively identify and prevent various different text, picture, voice and video fraud content, improve user safety and trust; at the same time, it has the ability to adaptively update the iteration model, which can continuously identify new fraud methods. However, the intelligent anti-fraud system relies on single analysis of LLM after converting multi-modal content into text, and the multi-modal detection depth is insufficient, lacking technical processing of non-text content, and lacking passive real-time protection capability. The detection result has poor interpretability, which makes it difficult to understand the risk basis. Based on the above technical problems, an anti-fraud intelligent protection system for the elderly based on multi-modal detection and an LSTM model is proposed. SUMMARY

[0004] (I) Technical problems solved In view of the deficiencies of the prior art, the application provides an anti-fraud intelligent protection system for the elderly based on multi-modal detection and an LSTM model, which has the advantages of constructing a multi-modal detection system and active and passive defense mechanisms, and introducing a detection result visualization interpretation technology. The problems of the existing intelligent anti-fraud system in the above background technology, such as multi-modal detection depth, real-time protection capability and detection result interpretation and understanding, need to be further improved.

[0005] (II) Technical solutions In order to achieve the above-mentioned construction of multi-modal detection system and active and passive defense mechanism, and introduce the visual explanation technology of detection result, the technical scheme is provided as follows: an anti-fraud intelligent protection system based on multi-modal detection and LSTM model, comprising an anti-fraud detection system architecture composed of a front-end interface subsystem, a back-end service subsystem, an intelligent protection fraud prevention core module and a model library; The front-end interface subsystem comprises a main content block, a navigation route and a footer block. The intelligent protection fraud prevention core module comprises a multi-modal active detection upload sub-module, a passive real-time detection sub-module, a deep fake experience sub-module, an anti-fraud learning community sub-module and a voice interaction sub-module. The back-end service subsystem comprises a user management module, a data storage management module and a security protection module, an API interface component and a server group. The model library comprises a text detection and fraud risk analysis model, an image and video Deepfake detection model, an audio Deepfake detection model, a voice interaction auxiliary model and an auxiliary support model.

[0006] Preferably, the main content block comprises a detection upload page, a result analysis page, an education learning page and a personal information page, and the footer block is provided with a page link entrance corresponding to the main content block, and the link jump of different pages is realized through the navigation route. The detection upload page comprises a text detection block, a picture detection block, a video detection block and an audio detection block, which are used for submitting the upload detection of the corresponding block content. The result analysis page outputs the visual result corresponding to the detection analysis in the main content block, including abnormal probability evaluation, risk type judgment view, text interpretation display, process risk display, image analysis display and related case analysis display. The education learning page is provided with a test learning block, a community exchange block and a knowledge popularization block, the test learning block comprises fraud prevention test and face changing revelation, the fraud prevention test is used for learning anti-fraud experience, the face changing revelation is used for simulating fraud voice and fraud video recording experience, and the community exchange block and the knowledge popularization block are used for anti-fraud event exchange and anti-fraud knowledge propaganda.

[0007] Preferably, the user management module records user basic attributes and login record information, the user basic attributes comprise user ID, user name and user role, and the login record information comprises login ID, login password and IP address. The data storage management module includes a vector database storing vectorized text segments and enabling similarity segment retrieval, a cloud log database storing user and system operation error data, and an associative database storing interrelated data of a fraud type table, an anti-fraud education community record table, a user information table, a passive real-time detection table, a deepfake experience module table, and a detection task record table; The security protection module includes a network security layer for protecting system boundaries and internal network communication, an application security layer for protecting system applications, a data security layer for protecting static storage and dynamic transmission data, an authentication management layer for authorizing user device access, and an endpoint security layer for protecting user devices accessing the system; The network security layer uses a firewall and a DDoS protection network to control access authorization and identify and mitigate distributed denial-of-service attacks; the application security layer uses secure coding specifications and audit vulnerability scanning technology to check code and security vulnerabilities through automated tools; the data security layer uses TLS / SSL protocols to encrypt network transmission data, encrypt sensitive data in the database, and display desensitized data when displayed to unauthorized personnel; The authentication management layer implements multi-factor authentication login by combining passwords, mobile phone verification codes, and biometric features, and authorizes user operations based on role and attribute access control, API key, and SSL certificate security management; the endpoint security layer has virus intrusion monitoring software deployed in the system to monitor host device and service port security.

[0008] Preferably, the security protection module further includes a monitoring and auditing layer that collects all log files for correlation analysis to determine potential threat data, records key operations in audit logs for post-traceback, and monitors and handles security events in real time based on a security operations center; The API interface component includes an external interface for terminal connection and an internal interface for connecting the smart care fraud prevention core module, and a collaborative work subsystem is accessed through the external interface and the internal interface; The server group includes an Nginx domain name configuration server for managing website domain name configuration, a load balancing server for implementing server resource balanced allocation, a data backup server for periodic data backup, an API database server for API interface data, and an API interface server for processing user requests and providing data services; User requests are uploaded to the Nginx domain name configuration server and the load balancing server through the collaborative work system to achieve load balancing of the request, the request is distributed to idle API database servers and API interface servers through the load balancing server, and the response result is sent to the user's mobile terminal device through the collaborative work subsystem.

[0009] Preferably, the multi-modal active detection upload sub-module includes a text detection module, a picture detection module, a video detection module, and a sound detection module that access a detection upload page to perform deepfake detection and fraud risk detection. The text detection module inputs a TXT format text file and outputs a confidence, a judgment, and an analysis result in Json format. The picture detection module inputs a png or jpg format picture file and outputs a confidence, a judgment result, and an explanatory picture in Json format. The video detection module inputs an MP4 or avi format video file and outputs a frame risk array, a video overall risk, and an explanatory picture in Json format. The sound detection module inputs an MP3 or WAW format audio file and outputs a confidence, a judgment result, and a transcription text in Json format. The passive real-time detection sub-module detects content with deepfake based on screen picture capture. It monitors the client screen in real time and captures video content frames. The passive real-time detection sub-module inputs video stream data encoded in H.264 format. If the detected content is normal, it outputs video stream data encoded in H.264 format. If the detected content has abnormalities, it outputs a pop-up window on the screen and a voice broadcast warning. The voice interaction sub-module includes a large language model and a visual language model, as well as a sound transcription module, an operation module, and a memory module. The large language model inputs a Json format text and outputs an operation plan step. The visual language model inputs a picture file and outputs an operation method and pixel coordinates. The sound transcription module inputs an audio file and outputs a transcription text. The operation module inputs an operation method and pixel coordinates and outputs a Json format operation result. The memory module inputs a Json format operation result and outputs a null value.

[0010] Preferably, the deepfake experience sub-module includes an AI voice simulation function block and an AI face changing function block deployed in a test learning block. The AI voice simulation function block records and uploads voice files autonomously, and the AI face changing function block uploads picture files autonomously to generate AI face changing videos. The fraud prevention learning community sub-module includes a fraud case list deployed in a community exchange block and a case type analysis deployed in a knowledge popularization block.

[0011] Preferably, the text detection and fraud risk analysis model includes a Fast-DetectGPT model, a language large model, a CoT enhanced thinking chain model, and a RAG model. The Fast-DetectGPT model is based on conditional probability curvature calculation, which quickly determines whether the text is AI-generated by comparing the probability difference between human-selected words and machine model-selected words; The language large model distinguishes between ambiguous statements based on contextual understanding to identify risk types in annotated text, generates a detection report from complex detection results, and displays key risk points with points; The thought chain CoT enhanced model splits complex text into single sentence reasoning, evaluates risks sentence by sentence and scores risks; The RAG model generates a retrieval enhanced model with a DPR dual encoder sub-model to convert knowledge base text and query text into high-dimensional vectors by calculating similarity through inner product; A GLM text embedding sub-model splits the knowledge base text into fragments and converts them into semantic vectors; A FAISS vector database stores the embedded case vectors; The image and video Deepfake detection model is a multi-modal hybrid detection model combining CNN-LSTM-frequency domain analysis, which includes a spatial feature detection model based on convolutional neural network CNN, a time series feature detection model based on long short-term memory network LSTM, and a frequency feature detection model based on Fourier transform and wavelet analysis; The spatial feature detection model extracts local spatial details of images and video frames based on CNN, identifies subtle traces of AI forgery, and uses an integrated detection framework optimization model that combines MesoNet, Xception, and EfficientNet; The time series feature detection model processes the inter-frame temporal correlation of videos based on LSTM to capture the temporal anomaly features of AI-forged videos; The frequency feature detection model captures the abnormal features of AI-forged images in the frequency domain based on frequency transform technology.

[0012] Preferably, the audio Deepfake detection model includes an audio AI generation detection model trained based on the DECRO dataset, which extracts prosody, tone, and spectral features of audio to automatically identify AI speech defects; It also includes an automatic speech recognition ASR model that converts audio signals into text through an interface, uses noise reduction preprocessing and semantic error correction optimization, and adapts to elderly speech scenarios; The voice interaction assistance model includes a speech-to-text STT model and a natural language processing NLP model. The STT model uses an end-to-end speech recognition scheme to optimize elderly speech features, and the NLP model is based on intent classification and entity extraction technology to analyze user voice command requirements; The auxiliary support model is a Deepfake generation class model for a deepfake experience submodule, a generator-discriminator double network structure is constructed based on a generative adversarial network (GAN), an output AI fake video is generated by a generator migrating facial expressions and replacing a background according to a picture and voice text uploaded by a user, and a generation effect is optimized by a discriminator.

[0013] (Three) beneficial effects Compared with the prior art, the application provides an anti-fraud intelligent protection system suitable for the aging based on multi-modal detection and an LSTM model, which has the following beneficial effects: 1. The anti-fraud intelligent protection system suitable for the aging based on multi-modal detection and the LSTM model, by constructing a hybrid model of CNN, LSTM and frequency domain analysis, accurately detects image and video Deepfake abnormal data, and its logical dimension relies on LLMs combined with RAG and CoT technology to analyze text fraud tactics and correlate with legal cases, and synchronously adapts to text and audio modal special detection, realizing multi-modal content full coverage.

[0014] 2. The anti-fraud intelligent protection system suitable for the aging based on multi-modal detection and the LSTM model, by integrating STT and NLP models to build a voice interaction system, supporting a voice instruction uploading-voice broadcasting result operation process, its interface uses large font, high contrast rendering and click vibration feedback to adapt to the physiological characteristics of the elderly, and its detection result is converted into popular point text by LLMs and output by TTS voice to avoid complex reading problems of the elderly.

[0015] 3. The anti-fraud intelligent protection system suitable for the aging based on multi-modal detection and the LSTM model, based on H.264 encoding to capture screen video streams, realizes real-time detection of each frame through a lightweight detector, triggers a pop-up window and voice alarm immediately after detecting a risk, and synchronously forms a user-family member double-end collaborative protection on the family member side, aiming to solve the problem of insufficient risk awareness of the elderly.

[0016] 4. The anti-fraud intelligent protection system suitable for the aging based on multi-modal detection and the LSTM model, by image detection, Grad-CAM heat maps and frequency spectrum maps are output and fake areas are marked, by text detection, LLMs are used for sentence-by-sentence semantic analysis, by video detection, frame risk arrays are generated and the highest risk frame is marked and time sequence defects are displayed, realizing visual result explanation of detection conclusion-technical basis-case correlation. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 It is a detailed design schematic diagram of the system architecture of the application; Figure 2 It is a general structure design schematic diagram of the intelligent anti-fraud detection of the application; Figure 3The intelligent anti-scam core module architecture of the application is shown in the figure; Figure 4 The multi-modal upload detection module flowchart of the application is shown in the figure; Figure 5 The real-time passive protection detection module flowchart of the application is shown in the figure; Figure 6 The voice interaction module flowchart of the application is shown in the figure; Figure 7 The system top-level data flow conversion schematic diagram of the application is shown in the figure; Figure 8 The system 0th-level data flow conversion schematic diagram of the application is shown in the figure; Figure 9 The system 1st-level data flow conversion schematic diagram of the application is shown in the figure; Figure 10 The cross-end application architecture schematic diagram of the application is shown in the figure; Figure 11 The database design schematic diagram of the application is shown in the figure; Figure 12 The detection upload page real machine schematic diagram of the application is shown in the figure; Figure 13 The result analysis page real machine first part schematic diagram of the application is shown in the figure; Figure 14 The result analysis page real machine second part schematic diagram of the application is shown in the figure; Figure 15 The education learning page real machine first part schematic diagram of the application is shown in the figure; Figure 16 The education learning page real machine second part schematic diagram of the application is shown in the figure. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the embodiments of the application and the drawings. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0019] Embodiment one In this embodiment, the multi-modal active detection is tested, and the result analysis is as follows: 1) Audio detection result analysis: the system performs well in audio analysis, and can accurately identify synthesized audio and real audio. Especially in the comparison of spectral characteristics, the system can quickly identify the special noise and inconsistency of the fake audio, and the recognition rate is 95%. When facing audio with large background noise, the performance of the system decreases, but it can still identify obvious fake features; Improvement suggestions: Consider introducing more noise processing algorithms to improve detection capabilities in low SNR situations.

[0020] 2) Image detection result analysis: The image Deepfake detection effect is very ideal. The system successfully marks the images synthesized using face swapping technology, can accurately locate the forged area through light and shadow inconsistency, edge blur and other features, and the recognition accuracy is 98%; when the image quality is low (such as low resolution, excessive compression), the recognition accuracy decreases, and some forged details may be missed.

[0021] Improvement suggestions: Low-quality image processing algorithms can be enhanced, such as introducing super-resolution reconstruction technology.

[0022] 3) Video detection result analysis: The video detection module can effectively identify videos made using deepfake technology (such as GANs), with an accuracy of 96%. Based on time sequence consistency checking and face tracking technology, the system identifies unnatural changes between video frames and marks the forged parts. In long-time, high-frame-rate videos, the processing time is longer, and occasionally there may be a delay in fake detection.

[0023] Improvement suggestions: Video processing algorithms can be optimized to shorten detection time and improve processing speed.

[0024] Example Two In this example, the passive real-time detection test case is analyzed as follows: 1) Webpage or social platform monitoring result analysis: The system can timely identify and mark potential fake or fraud information, especially in analyzing text and link content. Tests show that the detection accuracy of fake information reaches 94%. During the detection process, the system can accurately identify common fraud links, phishing websites, etc. For some more concealed fraud methods (such as disguised malicious ads), the system's detection effect is slightly decreased, and it is recommended to optimize the detection mechanism for dynamic content (such as dynamically loaded ads).

[0025] 2) Video monitoring result analysis: The system can identify fake videos (such as Deepfake technology processed videos) in real time in video monitoring, with an accuracy of 95%. Whether it is image synthesis or sound synthesis, the system can timely feedback through inter-frame consistency and audio spectrum analysis. When the video quality is low and the frame rate is too low, the system may have some delay in analysis, and the accuracy decreases slightly. It is recommended to further enhance the robustness of the algorithm in low-quality video recognition.

[0026] 3) Voice message monitoring result analysis: For voice message monitoring, the system successfully identified fake audio with an accuracy rate of 96%. Especially when analyzing voice spectrum and tone, the system can effectively identify the characteristics of fake audio. In a noisy environment, the system's recognition ability decreases, and the processing of audio environmental noise needs to be strengthened to improve the accuracy of recognition.

[0027] Example Three In this example, the text-based fraud information recognition function is tested, and the result analysis includes: 1) Common fraud advertisement text: The system can accurately identify the "winning" and "free trial" fraud keywords contained in the text, with an accuracy rate of 92%.

[0028] 2) Phishing link recognition: The system can identify phishing websites in the text with an accuracy rate of 88%. However, there are still occasional missed detections for some shortened or hidden links.

[0029] 3) Indirect fraud language: For fraud information expressed indirectly or ambiguously, the system's accuracy rate is 75%. This type of fraud information often uses complex language patterns, and the system needs to optimize its context semantic analysis capabilities.

[0030] 4) False offer information recognition: For texts containing false offers or emergency promotions, the system can successfully identify and timely warn, with an accuracy rate of 90%.

[0031] Comprehensive analysis: The system performs well in identifying common fraud advertisement texts and phishing links, and can accurately mark fraud information and issue timely warnings. For complex fraud language, especially for ambiguous fraud information, the recognition ability is lacking. The existing sentiment analysis and context understanding model has certain limitations in this regard. The system also performs well in detecting false offer information, can timely identify and warn users, and has a low false positive rate.

[0032] Example Four In this example, the voice-based fraud detection function is tested, and the result analysis includes: 1) Common fraud voice recognition: The system performs well in identifying voices with "urgent", "offer" or "limited time" words, with an accuracy rate of 90%. These keywords often appear in fraud calls, and the system can accurately detect and provide warnings.

[0033] ​​2) Tone Analysis: The system accurately identified 85% of the speech containing threatening or urgent tones. Most fraudulent speech contains strong tone variations, especially urgent and threatening tones, but in complex situations (e.g., when emotional transitions are more natural), the system's recognition may not be satisfactory.

[0034] 3) Speech Rate Detection: The system effectively detected fast speech rates in 87% of the cases. Fraudulent calls often use fast speech rates to exert pressure, and the system performs well in such cases.

[0035] 4) Keyword Detection: The system accurately identified common fraudulent keywords such as "offer" and "winning" in 89% of the cases. This keyword detection is crucial for detecting fraudulent information.

[0036] Overall Analysis: The system performs well in identifying common fraudulent speech patterns, tone variations, and speech rates, especially for urgent and threatening tones and fast speech rates. In emotionally intense speech detection, the system's performance is slightly less satisfactory, especially when emotional expressions are more natural. The system needs to better analyze emotional changes and speech fluency to improve its detection capabilities for complex emotional speech.

[0037] Example Five In this example, the image-based fraudulent information recognition function is tested, and the results analysis includes: 1) Fake QR Code Identification: The system accurately identifies fake QR codes in images and warns users that the QR code may be a phishing link, with an accuracy rate of 90%. 2) Fake Advertisement Image Identification: The system can identify fake advertisements in images, especially those with misleading offers or prizes, with an accuracy rate of 88%.

[0038] 2) Web Screenshot Image Identification: The system performs well in detecting fraudulent content in web screenshots, successfully identifying elements of phishing websites with an accuracy rate of 85%.

[0039] 3) Real-Time Image Stream Detection: The system can effectively identify and warn of fraudulent information in real-time image streams with an accuracy rate of 87%, demonstrating strong real-time detection capabilities. Overall Analysis: The system performs well in handling fake QR codes, fake advertisements, and web screenshots, accurately identifying fraudulent elements in images and promptly issuing warnings. In real-time image stream detection, the system's performance still has room for improvement, as it may miss some detections in high-speed, high-volume input scenarios.

[0040] Example Six In this embodiment, the user behavior-based fraud information recognition function is tested, and the result analysis includes: 1) User browsing fraudulent website behavior: The system can timely detect the user's browsing behavior on fraudulent websites and issue real-time warnings with an accuracy rate of 90%. Clicking on fraudulent advertisements: The system's detection accuracy rate for user clicks on fraudulent advertisements is 88%. The system can quickly identify the fraudulent nature of the advertisements and provide timely feedback.

[0041] 2) Receiving suspicious emails: The system monitors the user's receipt of suspicious emails and successfully identifies fraudulent emails with an accuracy rate of 85%.

[0042] 3) Real-time monitoring of user operations: The system can dynamically update based on the user's real-time behavior, providing timely warnings with an accuracy rate of 87%.

[0043] Overall analysis: The system performs well in monitoring user behavior and identifying fraud risks, especially when browsing fraudulent websites and clicking on fraudulent advertisements, providing timely and effective warnings. When monitoring the user's email receipt behavior, the system shows some accuracy, but the identification accuracy rate for some less obvious email fraud is slightly lower. The real-time monitoring function can respond quickly in most cases, but there is still room for optimization in high-frequency user operations and complex scenarios.

[0044] In summary, the old-age adaptation anti-fraud intelligent care system based on multi-modal detection and LSTM model, by constructing a hybrid model of CNN, LSTM and frequency domain analysis, accurately detects image and video Deepfake abnormal data, its logical dimension relies on LLMs combined with RAG and CoT technology, analyzes text fraud tactics and correlates with legal cases, and synchronously adapts to text and audio modal special detection, achieving full coverage of multi-modal content; By integrating STT and NLP models to build a voice interaction system, it supports the operation process of voice instruction upload-voice broadcast results, its interface uses large font, high contrast rendering and click vibration feedback to adapt to the physiological characteristics of the elderly, and its detection results are converted into popular point text by LLMs, combined with TTS voice output, avoiding complex reading problems for the elderly; Based on H.264 encoding to capture screen video streams, a lightweight detector is used to realize real-time detection of each frame, and a pop-up window and voice alarm are triggered immediately after detecting risks, and a user-family double-end collaborative protection is formed on the family side, aiming to solve the problem of insufficient risk awareness of the elderly; Grad-CAM heat maps and frequency spectrum graphs are output through image detection, and fake areas are marked, LLMs are used for sentence-by-sentence semantic analysis through text detection, and frame risk arrays are generated through video detection, the highest risk frame is marked and time sequence defects are displayed, realizing the visualization of the detection conclusion-technical basis-case association result explanation.

[0045] The related modules involved in the system are hardware system modules or functional modules combined with computer software programs or protocols and hardware in the prior art. The computer software programs or protocols involved in the functional modules are known to those skilled in the art and are not improvements of the system. The improvement of the system is the interaction or connection relationship between the modules, that is, the improvement of the overall structure of the system to solve the corresponding technical problems of the system.

[0046] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. An age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model, characterized in that, The system architecture for fraud detection includes a front-end interface subsystem, a back-end service subsystem, a smart fraud prevention core module, and a model library. The front-end interface subsystem includes a main content section, navigation routes, and a footer section. The core module for intelligent fraud prevention includes a multimodal active detection and uploading submodule, a passive real-time detection submodule, a deepfake experience submodule, a fraud prevention learning community submodule, and a voice interaction submodule. The backend service subsystem includes a user management module, a data storage management module and a security protection module, an API interface component and a server group; The model library includes text detection and fraud risk analysis models, image and video deepfake detection models, audio deepfake detection models, voice interaction assistance models, and auxiliary support models.

2. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The main content section includes a detection upload page, a result analysis page, an education and learning page, and a personal information page. The footer section deploys page link entries corresponding to the main content section, and links to different pages are implemented through navigation routing. The detection upload page includes text detection blocks, image detection blocks, video detection blocks, and audio detection blocks, which are used to submit the content of the corresponding blocks for upload detection; The results analysis page outputs visual results corresponding to the detection and analysis in the main content section, including anomaly probability assessment, risk type judgment view, text interpretation display, process risk display, image analysis display, and related case analysis display; The educational learning page includes a test and learning block, a community exchange block, and a knowledge popularization block. The test and learning block includes anti-fraud tests and face-swapping debunking. Users can learn anti-fraud experience based on the anti-fraud test and experience recording fraudulent voice and videos based on the face-swapping debunking. The community exchange block and the knowledge popularization block are used for exchanging information on anti-fraud incidents and promoting anti-fraud knowledge.

3. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The user management module records user basic attributes and login record information. The user basic attributes include user ID, user name and user role. The login record information includes login ID, login password and IP address. The data storage management module includes a vector database storing vectorized text fragments and a vector database for similarity fragment retrieval, a cloud log database storing user and system operation error data, and a relational database. The relational database stores a fraud type table, an anti-fraud education community record table, a user information table, a passive real-time detection table, a deepfake experience module table, and a detection task record table for interrelated data. The security protection module includes a network security layer for protecting system boundaries and internal network communication, an application security layer for protecting system applications, a data security layer for protecting static storage and dynamic data transmission, an authentication management layer for authorizing access to user equipment, and an endpoint security layer for protecting user equipment accessing the system. The network security layer employs firewalls and DDoS protection to control network access authorization, as well as to identify and mitigate distributed denial-of-service attacks; the application security layer uses secure coding standards and audit vulnerability scanning technology, and uses automated tools to check code for security vulnerabilities; the data security layer uses TLS / SSL protocols to encrypt data transmitted over the network, encrypt sensitive data in the database, and display de-identified data when unauthorized personnel are shown the data. The authentication management layer combines passwords, mobile verification codes, and biometrics to achieve multi-factor authentication login, and authorizes user operations based on role and attribute access control, and manages API keys and SSL certificates securely; the endpoint security layer deploys virus intrusion monitoring software in the system to monitor the security of host devices and service ports.

4. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The security protection module also includes a monitoring and auditing layer, which collects all log files for correlation analysis to determine potential threat data, records key operations in the audit logs for post-event traceability, and handles security incidents based on real-time monitoring by the security operations center. The API interface component includes an external interface for terminal connection and an internal interface for connecting to the Smart Protection Anti-Fraud Core Module. A collaborative working subsystem is connected based on the external and internal interfaces. The server group includes an Nginx domain name configuration server for managing website domain name configuration, a load balancing server for achieving balanced allocation of server resources, a data backup server for periodically backing up data, an API database server for API interface data, and an API interface server for processing user requests and providing data services. User requests are uploaded to the Nginx domain configuration server and load balancer server through the collaborative work system to achieve load balancing of requests. The load balancer server distributes the requests to idle API database servers and API interface servers, and sends the response results to the user's mobile device through the collaborative work subsystem.

5. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The multimodal active detection upload submodule includes a text detection module, an image detection module, a video detection module, and an audio detection module that are connected to the detection upload page to perform deepfake detection and fraud risk detection. The text detection module takes a TXT format text file as input and outputs a JSON format confidence score, judgment, and analysis result; the image detection module takes a PNG or JPG format image file as input and outputs a JSON format confidence score, judgment result, and explanatory image; the video detection module takes an MP4 or AVI format video file as input and outputs a JSON format frame risk array, overall video risk, and explanatory image. The sound detection module takes MP3 and WAW audio files as input and outputs JSON format confidence scores, judgment results, and transcribed text. The passive real-time detection submodule detects deepfake content based on screen capture. It monitors the client screen in real time and captures video content frames. The passive real-time detection submodule is input with H.264 format encoded video stream data. If the detected content is normal, it outputs H.264 format encoded video stream data. If the detected content is abnormal, it outputs a pop-up window and voice warning on the screen. The voice interaction submodule includes a deployed large language model and a visual language model, as well as a voice transcription module, an operation module, and a memory module. The large language model takes JSON format text as input and outputs operation plan steps. The visual language model takes an image file as input and outputs the operation method and pixel coordinates. The audio file is input into the sound transcription module, and the transcribed text is output. The operation module inputs the operation method and pixel coordinates, and outputs the operation result in JSON format. The memory module takes a JSON-formatted operation result as input and outputs an empty value.

6. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The deepfake experience submodule includes an AI voice-mimicking function module and an AI face-swapping function module deployed in the test and learning block. The AI ​​voice-mimicking function module autonomously records and uploads voice files, and the AI ​​face-swapping function module autonomously uploads image files, which are then merged to generate an AI face-swapping video. The anti-fraud learning community submodule includes a list of fraud cases deployed in the community exchange block and case type analysis deployed in the knowledge popularization block.

7. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The text detection and fraud risk analysis model includes the Fast-DetectGPT model, the language big model, the CoT enhanced thinking chain model, and the retrieval enhanced generation RAG model. The Fast-DetectGPT model is based on conditional probability curvature calculation. By comparing the probability differences between words selected by humans and words selected by machine models, it can quickly determine whether the text is generated by AI. The language big data model understands the context to distinguish the fraudulent intent of ambiguous statements, automatically identifies the risk type in the labeled text, generates a detection report from complex detection results, and displays key risk points in points. The CoT enhanced thinking chain model breaks down complex text into single-sentence reasoning, assesses risk sentence by sentence, and scores risk. The retrieval enhancement generation RAG model uses the DPR dual encoder sub-model to convert the knowledge base text and query text into high-dimensional vectors and calculates similarity through inner product; it uses the GLM text embedding sub-model to segment the knowledge base text and convert it into semantic vectors; and it uses the FAISS vector database to store the embedded case vectors. The image and video Deepfake detection model is a multimodal hybrid detection model that combines CNN-LSTM-frequency domain analysis. The multimodal hybrid detection model includes a spatial feature detection model based on convolutional neural network (CNN), a temporal feature detection model based on long short-term memory network (LSTM), and a frequency feature detection model based on Fourier transform and wavelet analysis. The spatial feature detection model is based on CNN to extract local spatial details of images and video frames, identify subtle traces of AI forgery, and adopts an integrated detection framework that combines MesoNet, Xception and EfficientNet to optimize the model. The temporal feature detection model is based on the inter-frame temporal correlation of LSTM processing of video to capture the temporal anomaly features of AI-forged videos; The frequency feature detection model captures abnormal features in the frequency domain of AI-forged images based on frequency transformation technology.

8. The age-adaptive anti-fraud intelligent protection system based on multimodal detection and LSTM model according to claim 1, characterized in that, The audio deepfake detection model includes an audio AI generation detection model, which is trained on the DECRO dataset and extracts the prosody, tone and spectral features of the audio to automatically identify AI speech defects; it also includes an automatic speech recognition (ASR) model, which calls an interface to convert audio signals into text, and uses noise reduction preprocessing and semantic error correction optimization to adapt to the speech scenarios of the elderly. The voice interaction assistance model includes a speech-to-text (STT) model and a natural language processing (NLP) model. The STT model adopts an end-to-end speech recognition scheme to optimize the speech characteristics of the elderly. The NLP model is based on intent classification and entity extraction technology to parse the user's voice command requirements. The auxiliary support model is a Deepfake generation model used in the Deepfake Experience submodule. The Deepfake generation model is based on a Generative Adversarial Network (GAN) to construct a dual network structure of generator and discriminator. Based on the user-uploaded images and voice text, the generator transfers facial expressions and replaces the background to generate and output AI-forged videos. The discriminator optimizes the generation effect.

Citation Information

Patent Citations

  • Intelligent anti-fraud technology and system based on large language model adaptive iteration

    CN117910452A