Anti-social worker attack information cleaning and security detection system and method of online customer service system

By using an information cleaning and security detection system for online customer service systems, combined with deep learning models and threat intelligence APIs, the problem of customer service systems being unable to defend against social engineering attacks has been solved. This system enables full-dimensional security analysis of text, URLs, and attachments, thereby improving network security protection capabilities.

CN121009541APending Publication Date: 2025-11-25国家电网有限公司客户服务中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511032556.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing customer service systems struggle to effectively identify and defend against social engineering attacks, leading to compromised terminal devices that become springboards for internal network threats. Existing protection devices also fall short of providing comprehensive security analysis of text content, URL links, and attachments.

Method used

An anti-social engineering attack information cleaning and security detection system for online customer service was designed, including an information receiving and cleaning module, an information classification module, a security analysis module, and an information alarm module. It utilizes Kafka for real-time streaming data processing, combines the DeepSeek model and third-party threat intelligence APIs to perform security detection on text, URLs, and attachments, dynamically manages blacklists and whitelists, and aggregates security analysis results through Kafka to take corresponding measures.

Benefits of technology

It enables real-time identification and defense against social engineering attacks, reduces the risk of attacks on online customer service systems, protects the security of terminal devices, improves the full-dimensional coverage of network security protection, and reduces the impact on the emotions and equipment of customer service personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009541A_ABST
    Figure CN121009541A_ABST
Patent Text Reader

Abstract

The invention relates to a social worker attack prevention information cleaning and security detection system and method of an online customer service system. The system comprises an information receiving and cleaning module; an information classification module; the security analysis module comprises three independent units, namely a text risk analysis unit, a link security detection unit and an attachment security detection unit; and an information alarm module. Compared with the prior art, the method has the following advantages: uncivilized languages are filtered through flow cleaning, so that the emotion and mentality of customer service personnel can be prevented from being influenced, and subsequent work can be better carried out; social worker languages can be recognized through text analysis for timely early warning, so that sensitive information is prevented from being leaked by entering a snare; safety analysis is carried out on the URL and the attachment, and the situation that a computer is poisoned, lost and the like due to the fact that online customer service staff click a fishing link and receive malicious attachments is avoided; real-time warning is performed through security detection of texts, URLs and attachments, so that the network security protection range is expanded, and threat risks are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital information transmission technology, specifically to an anti-social engineering attack information cleaning and security detection system and method for online customer service systems. Background Technology

[0002] With the rapid development of information technology, customer service systems have evolved from traditional telephone channels to multimedia channels such as intelligent voice and online customer service. They have also evolved from transmitting traditional voice messages to transmitting audio, video, URLs, compressed files, and other formats. As intelligent customer service continues to develop, attacks targeting customer service systems are becoming increasingly serious.

[0003] In addition to the necessary system functional components, customer service systems also pose a significant threat to customer service personnel, as their terminals are vulnerable to social engineering attacks, which can compromise them and turn them into springboards for hackers to move laterally within the intranet office area.

[0004] Therefore, current customer service systems only deploy traditional firewalls, intrusion detection systems (IDS), and web application firewalls (WAF) at the network boundary, targeting the network and application layers. These devices primarily detect and block traffic based on known traffic characteristics, making it difficult to effectively identify social engineering attacks, spam, or public opinion risks. Faced with increasingly complex security threats, the existing protection architecture is no longer sufficient to meet actual needs. It needs to be upgraded to a comprehensive security analysis system covering text content, URL links, and attachments. This system should detect potential threats in real time and automatically trigger alerts, forming a more robust risk defense system.

[0005] Explanation of relevant terms:

[0006] Social engineering attacks, short for social engineering attacks, are cyberattacks that utilize "social engineering" techniques. In computer science, social engineering refers to influencing someone's psychology through legitimate communication, causing them to take certain actions or reveal confidential information. This is often considered a form of deception to gather information, commit fraud, and infiltrate computer systems.

[0007] IDS: Intrusion Detection System, is a security technology system used to monitor abnormal activity in a network or system, identify potential security threats, and issue timely alerts.

[0008] WAF: Web Application Firewall, is a security device or software specifically designed to protect web applications. Summary of the Invention

[0009] This invention addresses the shortcomings of existing technologies by providing real-time alerts for security detection of text, URLs, and attachments, thereby increasing the scope of network security protection and reducing threat risks.

[0010] On the one hand, this invention proposes an anti-social engineering attack information cleaning and security detection system for online customer service systems, comprising:

[0011] The information receiving and cleaning module is used to receive raw information data from customers and hackers and clean the information, filtering out spam and social engineering language from the text. The cleaned data is stored in the database.

[0012] The information classification module uses Kafka's real-time streaming data processing to classify the information stream and distribute different types of data to the corresponding units in the security analysis module for security detection.

[0013] The security analysis module comprises three independent units:

[0014] The text risk analysis unit is used to identify potential malicious content or risks in text.

[0015] The link security detection unit is used to perform security checks on URLs, query blacklists and whitelists, analyze URLs, and detect malicious behavior;

[0016] The attachment security detection unit is used to detect and block attachments that carry malicious code or risky content;

[0017] The information alerting module aggregates all security analysis results via Kafka, including text security, link security, and attachment security. Analysis results are stored in a Kafka topology, using keys or message identifiers to distinguish different result types. Based on the risk of text, links, and attachments, the module comprehensively assesses the overall threat level and takes security response measures based on the assessment results. For example, malicious links or attachments are blocked or alerted; real attack data, such as text, links, and attachments, are labeled.

[0018] Preferably, it also includes: a management and control module; serving as a centralized management console for displaying security incident and threat detection reports; and providing a user feedback interface where users can report potential false alarms or missed alarms to further optimize related modules.

[0019] Secondly, the present invention provides a method for cleaning and security testing of social engineering attack information in an online customer service system implemented on the aforementioned cleaning and security testing system, comprising the following processes:

[0020] S100: After the raw information data of customers and hackers is sent, it is received by the Kafka of the information receiving and cleaning module and cleaned. The information is filtered to remove spam and social engineering language from the text, and the cleaned data is stored in the database.

[0021] S200: The information classification module uses Kafka's real-time streaming data processing to perform streaming processing, classify the information stream, and distribute different types of data to the corresponding units in the security analysis module for security detection.

[0022] S300: The text risk analysis unit identifies potential malicious content or risks in text;

[0023] S400: The link security detection unit performs security checks on URLs, queries blacklists and whitelists, analyzes URLs, and detects malicious behavior;

[0024] S500: The attachment security detection unit detects and blocks attachments that carry malicious code or risky content;

[0025] S600: The information alert module aggregates all security analysis results through Kafka, including text security, link security, and attachment security. The analysis results are stored in a Kafka toptic, using a key or message identifier to distinguish different types of results. Based on the risks of text, links, and attachments, the overall threat level is comprehensively judged, and security response measures are taken based on the assessment results.

[0026] Preferably, step S300 includes the following process:

[0027] It uses a locally deployed DeepSeek large model for intelligent text processing, covering the entire process of text preprocessing, sentiment analysis, text classification, keyword extraction, named entity recognition, and malicious content detection.

[0028] S310: Utilizes DeepSeek's built-in high-efficiency word segmenter to replace traditional word segmentation tools, and achieves sentiment analysis through customized prompting engineering and few-sample fine-tuning;

[0029] S320: Model-based multi-task learning capability enables text classification and named entity recognition;

[0030] S330: For malicious content detection, it combines DeepSeek's semantic understanding capabilities with a custom malicious feature library to identify phishing messages, malicious rhetoric, and social engineering attack patterns through model inference. At the same time, it uses the model's keyword extraction function to replace traditional algorithms, achieving end-to-end intelligent identification of public opinion information and risky texts.

[0031] Preferably, S400 includes the following process:

[0032] S410: Blacklist / whitelist detection uses blacklist / whitelist data based on DNS or a database for inspection;

[0033] S420: For URLs not in the blacklist or whitelist, additional security checks are performed by calling a third-party threat intelligence API to detect whether URLs not in the blacklist or whitelist point to phishing, malware distribution sites, or ad fraud.

[0034] S430: Analyze the URL redirection chain to uncover the hidden final destination address.

[0035] Preferably, step S410 includes the following process:

[0036] S411: Basic URL parsing;

[0037] S412: Domain Reputation Test;

[0038] S413: DNS and IP Detection;

[0039] S414: HTTP Response Analysis;

[0040] S415: SSL / TLS verification.

[0041] Preferably, the malicious behavior detection in S420 includes the following process:

[0042] S421: URL fingerprint recognition;

[0043] S422: URL content scanning; Text related to URL content scanning can be submitted to the text analysis unit for text analysis and public opinion identification.

[0044] Preferably, the blacklist and whitelist are dynamically managed. URLs that pass security checks are updated to the database, and caching technologies such as Redis and Memcached are used to accelerate the querying of blacklists and whitelists and the caching of URL detection results. URLs that pass security checks are synchronized to the whitelist and a trusted identifier is added, while URLs that fail security checks are synchronized to the blacklist and a risk identifier is added.

[0045] Preferably, step S500 includes the following specific process:

[0046] S510: Uses the open-source content analysis tool Apache Tika to parse file types, verify the consistency between file content and extension, prevent malicious files disguised as normal files, configures a whitelist of file types, only allows attachments of specific formats to be uploaded, and uses Nginx as a reverse proxy server to limit the request body and avoid uploading large files.

[0047] S520: Submit the content of the attached document to the text risk analysis unit for text analysis and public opinion identification;

[0048] S530: Utilizes the ClamAV open-source virus scanning engine to detect whether attachments contain viruses, Trojans, or malicious macro code in real time during attachment upload;

[0049] S540: Simulates the execution of attachments in a sandbox environment to detect whether their behavior has malicious characteristics;

[0050] S550: Attachments are uploaded to a third-party threat intelligence center for detection. Links contained within are automatically submitted to the link security detection unit for link security analysis. Attachment detection results are identified and synchronized to a sample library. The sample library uses hash matching technology to record the hash value of attachments. Configuring the sample library can improve the system's ability to detect malicious and suspicious attachments. Attachments that match the sample library are not uploaded to the sandbox, which avoids duplicate detection of the same attachments, improving attachment detection speed and user interaction experience.

[0051] S560: Attachments that require security testing will be notified by Kafka that the attachment is under testing; attachments that fail security testing or are marked as threatening in the sample library will be blocked or flagged as having security risks, while attachments without risks will be marked as safe.

[0052] The advantages of this invention over the prior art are as follows:

[0053] Firstly, online customer service representatives may exhibit uncivilized behavior due to customer mood swings. Traffic filtering removes such language, preventing it from affecting the emotions and mindset of customer service staff and allowing them to better perform their duties. Secondly, hackers can use online customer service for social engineering attacks. Text analysis can identify social engineering language and provide timely warnings to prevent users from falling into traps and leaking sensitive information. Security analysis of URLs and attachments prevents online customer service staff from clicking phishing links or receiving malicious attachments, which could lead to computer infections or compromised systems. Real-time alerts based on the security detection of text, URLs, and attachments increase the scope of network security protection and reduce threat risks. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the information cleaning and security detection system and method for preventing social engineering attacks in an online customer service system according to the present invention. Detailed Implementation

[0055] A method for preventing social engineering attacks and performing information cleaning and security detection in an online customer service system includes the following processes:

[0056] S100: After the raw information data from customers and hackers is sent, it is received by the Kafka information receiving and cleaning module and cleaned to filter out spam and social engineering language from the text. The cleaned data is stored in a database, such as PostgreSQL, or in a NoSQL database, such as MongoDB.

[0057] S200: The information classification module uses Kafka's real-time streaming data processing Streams to perform streaming processing, classify the information stream, and distribute different types of data to the corresponding units in the security analysis module for security detection.

[0058] S300: The text risk analysis unit identifies potential malicious content or risks in the text; S300 includes the following process:

[0059] It uses a locally deployed DeepSeek large model for intelligent text processing, covering the entire process of text preprocessing, sentiment analysis, text classification, keyword extraction, named entity recognition (NER), and malicious content detection.

[0060] S310: Utilizes DeepSeek's built-in high-efficiency word segmenter to replace traditional word segmentation tools, and achieves sentiment analysis through customized prompting engineering and few-sample fine-tuning;

[0061] S320: Model-based multi-task learning capabilities enable text classification and named entity recognition without relying on independent tools such as FastText and HanLP;

[0062] S330: For malicious content detection, it combines DeepSeek's semantic understanding capabilities with a custom malicious feature library to identify phishing information, malicious rhetoric, and social engineering attack patterns through model inference. At the same time, it uses the model's keyword extraction function to replace traditional algorithms, achieving end-to-end intelligent identification of public opinion information and risky texts.

[0063] S400: The link security detection unit performs security checks on the URL, queries blacklists and whitelists, analyzes the URL, and detects malicious behavior; S400 includes the following processes:

[0064] S410: Blacklist / whitelist detection uses DNS-based or database-based blacklist / whitelist data, such as Redis or MySQL; the blacklist consists of known malicious, phishing, or insecure URLs or domains, while the whitelist consists of known, trusted, and legitimate URLs or domains; S410 includes the following process:

[0065] S411: Basic URL parsing;

[0066] S412: Domain Reputation Test;

[0067] S413: DNS and IP Detection;

[0068] S414: HTTP Response Analysis;

[0069] S415: SSL / TLS verification;

[0070] S420: For URLs not in the blacklist or whitelist, additional security checks are performed by calling a third-party threat intelligence API to detect whether URLs not in the blacklist or whitelist point to phishing, malware distribution sites, or ad fraud; the malicious behavior detection in S420 includes the following process:

[0071] S421: URL fingerprint recognition;

[0072] S422: URL content scanning; Text related to URL content scanning can be submitted to the text risk analysis unit for text analysis and public opinion identification.

[0073] S430: Analyze the URL redirection chain to uncover the hidden final destination address;

[0074] S500: The attachment security detection unit detects and blocks attachments carrying malicious code or risky content; S500 includes the following specific processes:

[0075] S510: Uses the open-source content analysis tool Apache Tika to parse file types, verify the consistency between file content and extension, prevent malicious files disguised as normal files, configures a whitelist of file types, only allows attachments of specific formats to be uploaded, and uses Nginx as a reverse proxy server to limit the request body and avoid uploading large files.

[0076] S520: The content of the attached document is submitted to the text analysis unit for text analysis and public opinion identification;

[0077] S530: Utilizes the ClamAV open-source virus scanning engine to detect whether attachments contain viruses, Trojans, or malicious macro code in real time during attachment upload;

[0078] S540: Simulates the execution of attachments in the Cuckoo Sandbox environment to detect whether their behavior has malicious characteristics;

[0079] S550: Attachments are uploaded to a third-party threat intelligence center for detection. Included links are automatically submitted to the link security detection unit for link security analysis. Attachment detection results are identified and synchronized to a sample database. The sample database uses hash matching technology to record the hash value of attachments. Configuring the sample database improves the system's ability to detect malicious and suspicious attachments. Attachments matching the sample database are not uploaded to the sandbox, avoiding duplicate detection of identical attachments and improving attachment detection speed and user experience. The sample database can use a NoSQL database, such as Elasearch or MongoDB.

[0080] S560: Attachments that require security testing will be notified by Kafka that the attachment is under testing; attachments that fail security testing or are marked as threatening in the sample library will be blocked or flagged as having security risks; attachments without risks will be marked as safe.

[0081] S600: The information alert module aggregates all security analysis results through Kafka, including text security, link security, and attachment security. The analysis results are stored in a Kafka toptic topic, using a key or message identifier to distinguish different types of results. Based on the risks of text, links, and attachments, the overall threat level is comprehensively judged, and security response measures are taken based on the assessment results.

[0082] The blacklist and whitelist are dynamically managed. URLs that pass security checks are updated to the database, and caching technologies such as Redis and Memcached are used to accelerate the querying of blacklists and whitelists and the caching of URL detection results. URLs that pass security checks are synchronized to the whitelist and a trusted identifier is added, while URLs that fail security checks are synchronized to the blacklist and a risk identifier is added.

Claims

1. A social engineering attack prevention and security detection system for online customer service systems, characterized in that, include: The information receiving and cleaning module is used to receive raw information data from customers and hackers and clean the information, filtering out spam and social engineering language from the text. The cleaned data is stored in the database. The information classification module uses Kafka's real-time streaming data processing to classify the information stream and distribute different types of data to the corresponding units in the security analysis module for security detection. The security analysis module comprises three independent units: The text risk analysis unit is used to identify potential malicious content or risks in text. The link security detection unit is used to perform security checks on URLs, query blacklists and whitelists, analyze URLs, and detect malicious behavior; The attachment security detection unit is used to detect and block attachments that carry malicious code or risky content; The information alert module aggregates all security analysis results through Kafka, including text security, link security, and attachment security. The analysis results are stored in a Kafka toptic, and different types of results are distinguished by key or message identifier. Based on the risks of text, links, and attachments, the overall threat level is comprehensively judged, and security response measures are taken based on the assessment results.

2. The anti-social engineering attack information cleaning and security detection system for an online customer service system according to claim 1, characterized in that, Also includes: Management and control module; It serves as a centralized management console for displaying security incident and threat detection reports; Provide a user feedback interface, where users can report potential false alarms or missed alarms to further optimize related modules.

3. A method for cleaning and security testing information against social engineering attacks in an online customer service system implemented on the cleaning and security testing system described in claim 1, characterized in that, The process includes the following: S100: After the raw information data of customers and hackers is sent, it is received by the Kafka of the information receiving and cleaning module and cleaned. The information is filtered to remove spam and social engineering language from the text, and the cleaned data is stored in the database. S2 00: The information classification module uses Kafka's real-time streaming data processing to perform streaming processing, classify the information stream, and distribute different types of data to the corresponding units in the security analysis module for security detection; S300: The text risk analysis unit identifies potential malicious content or risks in text; S400: The link security detection unit performs security checks on URLs, queries blacklists and whitelists, analyzes URLs, and detects malicious behavior; S500: The attachment security detection unit detects and blocks attachments that carry malicious code or risky content; S6 00: The information alert module aggregates all security analysis results through Kafka, including text security, link security, and attachment security. The analysis results are stored in a Kafka toptic, using a key or message identifier to distinguish different types of results. Based on the risks of text, links, and attachments, the module comprehensively judges the overall threat level and takes security response measures based on the assessment results.

4. The method for preventing social engineering attacks and performing information cleaning and security detection in an online customer service system according to claim 3, characterized in that, S300 includes the following process: It uses a locally deployed DeepSeek large model for intelligent text processing, covering the entire process of text preprocessing, sentiment analysis, text classification, keyword extraction, named entity recognition, and malicious content detection. S310: Utilizes DeepSeek's built-in high-efficiency word segmenter to replace traditional word segmentation tools, and achieves sentiment analysis through customized prompting engineering and few-sample fine-tuning; S320: Model-based multi-task learning capability enables text classification and named entity recognition; S330: For malicious content detection, it combines DeepSeek's semantic understanding capabilities with a custom malicious feature library to identify phishing messages, malicious rhetoric, and social engineering attack patterns through model inference. At the same time, it uses the model's keyword extraction function to replace traditional algorithms, achieving end-to-end intelligent identification of public opinion information and risky texts.

5. The method for preventing social engineering attacks and performing security detection in an online customer service system according to claim 3, characterized in that, S400 includes the following process: S410: Blacklist / whitelist detection uses blacklist / whitelist data based on DNS or a database for inspection; S420: For URLs not in the blacklist or whitelist, additional security checks are performed by calling a third-party threat intelligence API to detect whether URLs not in the blacklist or whitelist point to phishing, malware distribution sites, or ad fraud. S430: Analyze the URL redirection chain to uncover the hidden final destination address.

6. The method for preventing social engineering attacks and performing security detection in an online customer service system according to claim 5, characterized in that, S410 includes the following process: S411: Basic URL parsing; S412: Domain Reputation Detection; S413: DNS and IP Detection; S414: HTTP Response Analysis; S415: SSL / TLS Authentication.

7. The method for preventing social engineering attacks and performing security detection in an online customer service system according to claim 5, characterized in that, The malicious behavior detection in S420 includes the following process: S421: URL fingerprint recognition; S422: URL content scanning; Text related to URL content scanning can be submitted to the text risk analysis unit for text analysis and public opinion identification.

8. The method for preventing social engineering attacks and performing security detection in an online customer service system according to claim 5, characterized in that, The blacklist and whitelist are dynamically managed and updated to the database through the URLs detected by security checks. Redis and Memcached caching technologies are used to accelerate the query of the blacklist and whitelist and the caching of URL detection results. URLs that pass the security check are synchronized to the whitelist and marked as trusted; URLs that fail the security check are synchronized to the blacklist and marked as risky.

9. The method for preventing social engineering attacks and performing security detection in an online customer service system according to claim 3, characterized in that, The S500 includes the following specific processes: S510: Uses the open-source content analysis tool Apache Tika to parse file types, verify the consistency between file content and extension, prevent malicious files disguised as normal files, configures a whitelist of file types, only allows attachments of specific formats to be uploaded, and uses Nginx as a reverse proxy server to limit the request body and avoid uploading large files. S520: The content of the attached document is submitted to the text analysis unit for text analysis and public opinion identification; S530: Utilizes the ClamAV open-source virus scanning engine to detect whether attachments contain viruses, Trojans, or malicious macro code in real time during attachment upload; S540: Simulates the execution of attachments in a sandbox environment to detect whether their behavior has malicious characteristics; S550: Attachments are uploaded to a third-party threat intelligence center for detection. Links contained within are automatically submitted to the link security detection module for link security analysis. Attachment detection results are identified and synchronized to a sample database. The sample database uses hash matching technology to record the hash value of attachments. Configuring the sample database can improve the system's ability to detect malicious and suspicious attachments. Attachments that match the sample database are not uploaded to the sandbox, which avoids duplicate detection of the same attachments, improving attachment detection speed and user interaction experience. S560: Attachments that require security testing are notified by Kafka that the attachments are being tested. Attachments that fail security testing or are identified as threatening in the sample library will be blocked or flagged for security risks. Attachments that do not pose a risk will be marked with a security symbol.