Chat content risk detection method, device and equipment and storage medium thereof

By crawling and processing chat content in instant messaging tools, using sensitive thesaurus and semantic risk detection models for risk detection, the problem of information transmission security in financial insurance business is solved, and effective protection of personal privacy and identity information is achieved.

CN120163167APending Publication Date: 2025-06-17CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315033.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

During the process of financial and insurance business, how to ensure the security of information transmission while using instant messaging tools and prevent personal privacy or identity information from being exposed in the chat interface.

Method used

By crawling the live chat content in the instant messaging tool, textual processing and word segmentation processing, risk detection is performed using preset sensitive thesaurus and semantic risk detection models, and chat content is processed in combination with risk processing strategies.

Benefits of technology

It effectively avoids customers exposing their personal privacy or identity information to the chat interface when conducting online financial business, improving information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163167A_ABST
    Figure CN120163167A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, is applied to a scene of online instant communication between a user and a customer service in a financial insurance business service process, and relates to a chat content risk detection method, device and equipment and a storage medium thereof.The method comprises the steps that real-time chat content to be displayed in a target instant communication tool is captured; carrying out textualization processing on the real-time chat content; performing word segmentation processing on the target text content, identifying all sensitive words in a word segmentation result, and obtaining a first risk detection result; inputting the target text content into a semantic risk detection model to obtain a second risk detection result; and performing risk processing on the real-time chat content in combination with the first risk detection result, the second risk detection result and a preset risk processing strategy. Before the chat content is displayed on the chat interface, personal privacy or identity information is prevented from being exposed in the chat interface when a customer performs financial service online handling by adopting an instant messaging tool, and the information security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applied to the scenario where users have online instant communication with customer service during the process of financial insurance business services. It involves a method, device, equipment and its storage medium for detecting risks in chat content. Background Art

[0002] With the rapid development of financial insurance business, insurance institutions have launched more and more insurance types of business. The information exchange between users and customer service, as well as between the upstream and downstream of business processing, has become increasingly frequent, often involving users sending personal information for handling insurance business to customer service, and the upstream and downstream of insurance business handling forwarding and verifying customer handling information, etc.

[0003] Currently, instant messaging tools, due to their instantaneity and speed, are increasingly used in daily communication and data transmission processes. For regular daily chats or ordinary data transmission, they do not require too high security performance. However, with more and more financial business handling gradually combined with instant messaging tools, how to fully ensure the security of information transmission while making full use of the advantages of instant messaging tools has become an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a method, device, equipment and its storage medium for detecting risks in chat content to solve the problem of how to fully ensure the security of information transmission when existing financial business handling is gradually combined with instant messaging tools.

[0005] In a first aspect, the embodiments of this application provide a method for detecting risks in chat content, which adopts the following technical solutions:

[0006] A method for detecting risks in chat content includes the following steps:

[0007] Based on the message sending instruction of the target user, capture the real-time chat content to be displayed in the target instant messaging tool;

[0008] Perform text processing on the real-time chat content to obtain the target text content;

[0009] Perform word segmentation on the target text content, and use a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain the first risk detection result, where the first risk detection result represents the detection result obtained by performing risk detection only based on key word segmentation comparison;

[0010] Input the target text content into the pre-trained semantic risk detection model to obtain the second risk detection result predicted by the semantic risk detection model, where the second risk detection result represents the detection result obtained by performing risk detection based on deep semantic analysis;

[0011] Combine the first risk detection result, the second risk detection result, and a preset risk handling strategy to perform risk handling on the real-time chat content.

[0012] In a second aspect, an embodiment of the present application further provides a chat content risk detection device, which adopts the following technical solution:

[0013] A chat content risk detection device includes:

[0014] A real-time chat content capture module, configured to capture the real-time chat content to be displayed in the target instant messaging tool based on the message sending instruction of the target user;

[0015] A target text content acquisition module, configured to perform text processing on the real-time chat content to obtain the target text content;

[0016] A first risk detection result acquisition module, configured to perform word segmentation on the target text content, and use a preset sensitive word library to identify all sensitive words in the corresponding word segmentation result, and obtain the first risk detection result, where the first risk detection result represents the detection result obtained by performing risk detection only based on key word segmentation comparison;

[0017] A second risk detection result acquisition module, configured to input the target text content into the pre-trained semantic risk detection model to obtain the second risk detection result predicted by the semantic risk detection model, where the second risk detection result represents the detection result obtained by performing risk detection based on deep semantic analysis;

[0018] A risk handling module, configured to combine the first risk detection result, the second risk detection result, and a preset risk handling strategy to perform risk handling on the real-time chat content.

[0019] In a third aspect, an embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0020] A computer device includes a memory and a processor. A computer-readable instruction is stored in the memory, and when the processor executes the computer-readable instruction, the steps of the above-mentioned chat content risk detection method are implemented.

[0021] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0022] A computer-readable storage medium stores computer-readable instructions thereon, and when the computer-readable instructions are executed by a processor, the steps of the chat content risk detection method as described above are implemented.

[0023] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0024] In the chat content risk detection method according to the embodiments of the present application, real-time chat content to be displayed in a target instant messaging tool is captured; the real-time chat content is texturized to obtain target text content; the target text content is segmented, and all sensitive words in the corresponding segmentation results are identified by comparing with a preset sensitive word library to obtain a first risk detection result; the target text content is input into a pre-trained semantic risk detection model to obtain a second risk detection result; and the real-time chat content is risk-processed by combining the first risk detection result, the second risk detection result, and a preset risk processing strategy. The chat content risk detection method is applied before the chat content is displayed on the target chat interface, so as to prevent customers from exposing personal privacy or identity information in the chat interface when handling financial business online using an instant messaging tool, and improve information security. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0027] Figure 2 is a flowchart of an embodiment of the chat content risk detection method according to the present application;

[0028] Figure 3 is Figure 2 a flowchart of a specific embodiment of step 201;

[0029] Figure 4 is Figure 2 a flowchart of a specific embodiment of step 202;

[0030] Figure 5 is Figure 2 a flowchart of a specific embodiment of step 203;

[0031] Figure 6 is Figure 2 a flowchart of a specific embodiment of step 204;

[0032] Figure 7 is a flowchart of a specific embodiment of the risk detection function plug-in processing in the chat content risk detection method described in this application;

[0033] Figure 8 is Figure 2 a flowchart of a specific embodiment of step 205;

[0034] Figure 9 is a schematic structural diagram of an embodiment of the chat content risk detection device according to this application;

[0035] Figure 10 is a schematic structural diagram of an embodiment of the computer device according to this application. Detailed implementation manners

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0037] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0038] In order to enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0039] Such as Figure 1As shown in the figure, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0040] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0041] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.

[0042] The server 103 may be a server that provides various services, such as a background server that provides support for the pages displayed on the terminal device 101.

[0043] It should be noted that the chat content risk detection method provided by the embodiments of the present application is generally executed by the server. Correspondingly, the chat content risk detection device is generally set in the server.

[0044] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0045] Continuing to refer to Figure 2 , a flowchart of an embodiment of the chat content risk detection method according to the present application is shown. The chat content risk detection method includes the following steps:

[0046] Step 201, based on the message sending instruction of the target user, capture the real-time chat content to be displayed in the target instant messaging tool.

[0047] In this embodiment, the target instant messaging tools include enterprise WeChat, WeChat, QQ, etc., and also include instant messaging tools independently developed within the company for users to communicate with customer service, such as the online communication plugin between users and customer service in the mobile APP for insurance business; it also includes internal communication instant tools among employees within the insurance company.

[0048] Specifically, the target users refer to users who handle insurance business or employee users during the transmission of insurance handling materials. Correspondingly, the real-time chat content to be displayed refers to the content of the message to be sent input by the user in the chat box of the target instant messaging tool.

[0049] Step 202: Perform text processing on the real-time chat content to obtain target text content.

[0050] In this embodiment, when performing text processing on the real-time chat content to obtain target text content, it should be understood that the real-time chat content not only includes the text content input by the user, but also includes pictures and voice clips sent by the user. By performing text processing on the real-time chat content to obtain target text content, it is convenient to subsequently uniformly perform risk detection on the target text content, so as to perform security processing on the user chat content and avoid information leakage.

[0051] Step 203: Perform word segmentation on the target text content, and use a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain a first risk detection result, where the first risk detection result represents the detection result obtained by performing risk detection only based on key word segmentation comparison.

[0052] Specifically, perform risk detection at the keyword level in combination with a preset sensitive word library to obtain a risk detection result at the word level.

[0053] Step 204: Input the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result predicted by the semantic risk detection model, where the second risk detection result represents the detection result obtained by performing risk detection based on deep semantic analysis.

[0054] Specifically, perform deep semantic analysis on the target text content in combination with a semantic risk detection model, and obtain a risk detection result at the semantic level according to the deep semantic analysis result.

[0055] Step 205: Combine the first risk detection result, the second risk detection result and a preset risk handling strategy to perform risk handling on the real-time chat content.

[0056] Specifically, based on the word-level risk detection result and the semantic-level risk detection result, the real-time chat content is processed for risks according to a preset risk handling strategy, realizing risk handling of the user's chat input content by combining low feature dimensions and high feature dimensions simultaneously.

[0057] The chat content risk detection method provided by this application is applied before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information in the chat interface when conducting online financial business using instant messaging tools, thereby improving information security. Of course, it can also be applied to the online handling of medical services using instant messaging tools to prevent personal privacy or identity information from being exposed in the chat interface when patient users provide personal information, improving information security.

[0058] In this embodiment, the real-time chat content to be displayed in the target instant messaging tool is captured; the real-time chat content is texturized to obtain the target text content; the target text content is segmented, and all sensitive words in the corresponding segmentation result are identified by comparing with a preset sensitive word library to obtain the first risk detection result; the target text content is input into a pre-trained semantic risk detection model to obtain the second risk detection result; the real-time chat content is processed for risks by combining the first risk detection result, the second risk detection result, and a preset risk handling strategy. The chat content risk detection method is applied before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information in the chat interface when conducting online financial business using instant messaging tools, thereby improving information security.

[0059] Continue to refer to Figure 3 , Figure 3 is Figure 2 The flowchart of a specific embodiment of step 201 includes:

[0060] Step 301: Obtain the API interface according to the real-time chat content corresponding to the target instant messaging tool, and create a corresponding real-time capture task;

[0061] Specifically, taking the car insurance business user APP, Ping An Good Driver as an example, a real-time chat content real-time capture task is created through the API interface for obtaining real-time chat content in Ping An Good Driver.

[0062] Step 302: Start the real-time capture task to capture the real-time chat content to be displayed in the target instant messaging tool.

[0063] Specifically, start the real-time capture task to capture the real-time chat content that the Ping An Good Driver user is about to display and send in the chat box.

[0064] By creating a real-time scraping task, scrape the real-time chat content that the Ping An Good Driver user is about to display and send in the chat box, so as to ensure that before the user sends the instant chat content, risk detection is performed on the content sent by the user, and to prevent the user from directly sending sensitive information to the display interface, resulting in a greater risk of information leakage.

[0065] Continue to refer to Figure 4 , Figure 4 Yes Figure 2 The flowchart of a specific embodiment of step 202 includes:

[0066] Step 401, upload the scraped real-time chat content to the pre-trained content type recognition model to identify the content category of the real-time chat content;

[0067] Specifically, the pre-trained content type recognition model can combine the representation form of the input content in the instant messaging tool chat box to identify the content category of the input content. It should be understood that, combined with existing instant messaging tools, when the user inputs text content, voice content, and picture content, the corresponding display representation forms of different input contents in the chat box are different. The recognition model that can distinguish the display representation forms can be trained in advance using the display representation forms corresponding to different categories of input contents in the chat box, and the trained recognition model is used as the content type recognition model.

[0068] Step 402, if the content category of the real-time chat content is a non-text category, perform text processing on the real-time chat content to obtain the target text content;

[0069] Step 403, if the content category of the real-time chat content is a text category, directly set the real-time chat content as the target text content.

[0070] Specifically, the step of performing text processing on the real-time chat content to obtain the target text content when the content category of the real-time chat content is a non-text category, that is, the implementation method of step 402, includes: if the content category of the real-time chat content is a voice category, perform voice-to-text processing on the real-time chat content to obtain the target text content; if the content category of the real-time chat content is a displayed picture, use OCR text recognition technology to recognize the text information part in the displayed picture to obtain the target text content.

[0071] Continue to refer to Figure 5 , Figure 5 Yes Figure 2 The flowchart of a specific embodiment of step 203 includes:

[0072] Step 501: Input the target text content into a natural language-based text tokenization tool to obtain the tokenization result.

[0073] Specifically, the natural language-based text tokenization tool, for example: the jieba tokenization tool based on NLP (Natural Language Processing), or the BERT text tokenization tool based on NLP, etc.

[0074] Step 502: Use a loop comparison method to compare each token in the tokenization result with all the standard sensitive words included in the standard word library part of the sensitive word library, and identify all the standard sensitive words included in the tokenization result.

[0075] Specifically, the standard word library part in the sensitive word library contains a set of standard sensitive vocabulary pre-set in combination with common business secrets, customer privacy, and relevant laws and regulations involved in the financial insurance business.

[0076] Step 503: Use a loop comparison method to compare each token in the tokenization result with all the custom sensitive words included in the custom word library part of the sensitive word library, and identify all the custom sensitive words included in the tokenization result.

[0077] Specifically, the custom word library part in the sensitive word library contains a set of custom sensitive vocabulary pre-set in combination with the user's input behavior, input habits, and colloquial expression habits.

[0078] Step 504: Organize all the standard sensitive words and all the custom sensitive words included in the tokenization result, and perform a risk detection on the target text content to obtain the first risk detection result.

[0079] Specifically, organizing all the standard sensitive words and all the custom sensitive words included in the tokenization result includes integrating all the standard sensitive words and all the custom sensitive words to obtain the standard sensitive words and custom sensitive words involved in the target text content, ensuring that corresponding sensitive words can be obtained for different users.

[0080] In this embodiment, before performing the steps of tokenizing the target text content, using the preset sensitive word library to compare and identify all the sensitive words in the corresponding tokenization result, and obtaining the first risk detection result, the method further includes: removing the redundant characters in the target text content; performing outlier deletion and correction processing on the target text content. To ensure the accuracy of the target text content and avoid processing the character content in the target text content, the character content generally includes emoji, dynamic emoji, etc.

[0081] Continue to refer to Figure 6 , Figure 6 is Figure 2 A flowchart of a specific embodiment of step 204, including:

[0082] Step 601, input the target text content into a pre-trained semantic risk detection model based on the LSTM algorithm, and extract the keyword dependency relationship, keyword part of speech, word frequency, and phrase structure contained in the target text content;

[0083] Specifically, a pre-trained semantic risk detection model based on the LSTM algorithm is used. When constructing it, the LSTM (Long Short-Term Memory Network) is utilized, which can fully combine the processing characteristics of the long short-term memory network when the user inputs long segments of speech or long segments of text content, and identify the long and short-term dependency relationships between various words in the target text content.

[0084] In this embodiment, several segments of text content containing risk information and having been scored for risk index values are obtained in advance. These several segments of text content are used as a training data set and input into a semantic risk detection model based on the LSTM algorithm to be trained, and the corresponding text features are extracted. According to the correlation correspondence between the text features and the risk information and risk index values respectively, the semantic risk detection model based on the LSTM algorithm is pre-trained to obtain the pre-trained semantic risk detection model based on the LSTM algorithm. It should be understood that the LSTM (Long Short-Term Memory Network) is used to utilize its capture of the long-term information dependency relationship. In actual processing, it is still based on the text feature recognition of the encoder-decoder mode to extract the corresponding text features, such as: keyword part of speech, word frequency, and phrase structure. The pre-trained semantic risk detection model based on the LSTM algorithm can combine the correlation correspondence between the text features and the risk information and risk index values respectively, and perform risk information detection and calculate the risk index value for the subsequent actual input chat content.

[0085] Step 602, use the keyword dependency relationship, keyword part of speech, word frequency, and phrase structure contained in the target text content as risk prediction features, predict the overall risk of the target text content, and output the overall risk index value;

[0086] By using the keyword dependency relationship, keyword part of speech, word frequency, and phrase structure contained in the target text content as risk prediction features, predicting the overall risk of the target text content, and outputting the overall risk index value, the overall risk detection of the target text content is realized at the deep semantic level. Compared with only detecting at the sensitive word level, it provides a more accurate and more comprehensive risk detection result.

[0087] Step 603: Use the overall risk index value as the second risk detection result predicted by the semantic risk detection model.

[0088] Continue to refer to Figure 7 , in some alternative implementation manners, before performing step 205, the method further includes a step of processing the risk detection function in a plug-in manner. Figure 7 FIG. is a flowchart of a specific embodiment of processing the risk detection function in a plug-in manner in the chat content risk detection method described in this application, and includes the following steps:

[0089] Step 701: Package the pre-trained content type recognition model, the preset sensitive word library, the pre-trained semantic risk detection model, the natural language-based text tokenization tool, and the preset risk processing strategy into a portable plug-in.

[0090] Step 702: Pre-install the portable plug-in into the startup function area of the target instant messaging tool.

[0091] Step 703: When the target user sends the message sending instruction in the chat box of the target instant messaging tool, trigger the startup of the portable plug-in in the startup function area.

[0092] Step 704: Use the pre-trained content type recognition model in the portable plug-in to perform content type recognition on the user input content in the chat box of the target instant messaging tool, and

[0093] Step 705: Perform text processing according to the content type recognition result to obtain the target text content.

[0094] Step 706: Use the preset sensitive word library and the natural language-based text tokenization tool in the portable plug-in to perform risk detection on the target text content to obtain the first risk detection result.

[0095] Step 707: Use the pre-trained semantic risk detection model in the portable plug-in to perform risk detection on the target text content to obtain the second risk detection result.

[0096] In a plug-in manner, implant the portable plug-in into the chat box function area of the target instant messaging tool, so that before the user inputs information, the corresponding risk detection plug-in can be started through a one-key startup method, and subsequently, when the user inputs information, the risk detection of the chat input content can be performed automatically and intelligently.

[0097] Continue to refer to Figure 8 , Figure 8 is Figure 2The flowchart of a specific embodiment of step 205 includes:

[0098] Using the preset risk handling strategy in the portable plugin, perform risk handling on the user input content in the chat box of the target instant messaging tool, specifically including:

[0099] Step 801: Calculate the proportion of the text volume of all sensitive words in the target text content according to the first risk detection result;

[0100] Step 802: Obtain the overall risk index value of the target text content according to the second risk detection result;

[0101] Step 803: Combine the risk weights corresponding to the first risk detection result and the second risk detection result respectively, the text proportion, and the overall risk index value, and perform comprehensive calculation to obtain the comprehensive risk value corresponding to the target text content;

[0102] Specifically, for example: the proportion of the text volume of all sensitive words in the target text content is 20%, the overall risk index value of the target text content is 50%, the sum of the risk weights corresponding to the first risk detection result and the second risk detection result is 1, the risk weight corresponding to the first risk detection result is 60% and the risk weight corresponding to the second risk detection result is 40%, then the comprehensive calculation result is: 20%×60% + 50%×40%, which is equal to 0.32, that is, the comprehensive risk value corresponding to the target text content is 0.32.

[0103] Step 804: Determine the size relationship between the comprehensive risk value and the preset risk threshold by comparison;

[0104] Step 805: If the comprehensive risk value exceeds the preset risk threshold, perform risk handling on the user input content in the chat box of the target instant messaging tool according to the preset risk handling strategy.

[0105] Specifically, the step of performing risk processing on the user input content in the chat box of the target instant messaging tool according to the preset risk processing strategy includes: displaying and intercepting the real-time chat content, and sending a risk prompt to the instant sending end of the target user; identifying the position information of the risk information in the user input content, and performing censor display through a pre-set censor processing component; displaying and intercepting the real-time chat content, and enabling a preset secure transmission interface to send the real-time chat content to the target receiving end; for example, transmitting the actual real-time chat content, including ID card information, personal privacy information, etc., to the target receiving end through the secure transmission interface to avoid transmission to the chat interface. Displaying and intercepting the real-time chat content, and performing encryption and compression processing using a preset encryption method to generate an encrypted file to be transmitted, and enabling a preset secure transmission interface to send the encrypted file to be transmitted to the target receiving end; for example, encrypting the actual real-time chat content, including ID card information, personal privacy information, etc., and then transmitting it to the target receiving end through the secure transmission interface to avoid transmission to the chat interface, and performing encryption processing to further avoid information leakage to a greater extent.

[0106] It should be understood that the above steps of performing risk processing on the user input content in the chat box of the target instant messaging tool can perform risk processing separately or jointly. Specifically, for example: only displaying and intercepting the real-time chat content and sending a risk prompt to the instant sending end of the target user. When sending the risk prompt, the prompt methods include supporting real-time pop-up windows, email notifications, and SMS reminders. At the same time, detailed prompt logs are provided, such as: message sender, message input time, etc.; or after displaying and intercepting the real-time chat content and sending a risk prompt to the instant sending end of the target user, enabling a preset secure transmission interface to send the real-time chat content to the target receiving end.

[0107] This application captures the real-time chat content to be displayed in the target instant messaging tool; performs text processing on the real-time chat content to obtain the target text content; performs word segmentation on the target text content, and uses a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain the first risk detection result; inputs the target text content into a pre-trained semantic risk detection model to obtain the second risk detection result; combines the first risk detection result, the second risk detection result and the preset risk processing strategy to perform risk processing on the real-time chat content. Applying the chat content risk detection method before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information on the chat interface when conducting online financial business using the instant messaging tool, and improving information security.

[0108] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0109] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0110] In the embodiments of the present application, the real-time chat content to be displayed in the target instant messaging tool is captured; the real-time chat content is texturized to obtain the target text content; the target text content is segmented, and all sensitive words in the corresponding segmentation results are identified by comparing with a preset sensitive word library to obtain the first risk detection result; the target text content is input into a pre-trained semantic risk detection model to obtain the second risk detection result; the real-time chat content is risk-processed by combining the first risk detection result, the second risk detection result, and a preset risk processing strategy. The chat content risk detection method is applied before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information on the chat interface when conducting online financial business using instant messaging tools, thereby improving information security.

[0111] Further referring to Figure 9 As an implementation of the method shown above Figure 2 In an embodiment of the chat content risk detection device provided by the present application, which corresponds to the method embodiment shown in Figure 2 This device can be specifically applied to various electronic devices.

[0112] As Figure 9 shown, the chat content risk detection device 900 described in this embodiment includes: a real-time chat content capture module 901, a target text content acquisition module 902, a first risk detection result acquisition module 903, a second risk detection result acquisition module 904, and a risk processing module 905. Among them:

[0113] The real-time chat content capture module 901 is used to capture the real-time chat content to be displayed in the target instant messaging tool based on the message sending instruction of the target user;

[0114] A target text content acquisition module 902, configured to perform text processing on the real-time chat content to obtain target text content;

[0115] A first risk detection result acquisition module 903, configured to perform word segmentation on the target text content, and use a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation result, so as to obtain a first risk detection result, where the first risk detection result represents a detection result obtained by performing risk detection only based on key word segmentation comparison;

[0116] A second risk detection result acquisition module 904, configured to input the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result predicted by the semantic risk detection model, where the second risk detection result represents a detection result obtained by performing risk detection based on deep semantic analysis;

[0117] A risk processing module 905, configured to perform risk processing on the real-time chat content by combining the first risk detection result, the second risk detection result, and a preset risk processing strategy.

[0118] This application captures the real-time chat content to be displayed in the target instant messaging tool; performs text processing on the real-time chat content to obtain target text content; performs word segmentation on the target text content, and uses a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation result, so as to obtain a first risk detection result; inputs the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result; combines the first risk detection result, the second risk detection result, and a preset risk processing strategy to perform risk processing on the real-time chat content. The chat content risk detection method is applied before the chat content is displayed on the target chat interface, so as to prevent customers from exposing personal privacy or identity information in the chat interface when handling financial business online using an instant messaging tool, and improve information security.

[0119] In this embodiment, the real-time chat content capture module 901 includes: a real-time capture task creation unit and a real-time chat content capture unit. Wherein:

[0120] The real-time capture task creation unit is configured to obtain an API interface for the real-time chat content corresponding to the target instant messaging tool, and create a corresponding real-time capture task;

[0121] The real-time chat content capture unit is configured to start the real-time capture task and capture the real-time chat content to be displayed in the target instant messaging tool.

[0122] In this embodiment, the target text content acquisition module 902 includes: a content type recognition unit, a text conversion unit, and a direct setting unit. Specifically:

[0123] The content type recognition unit is configured to upload the captured real-time chat content to a pre-trained content type recognition model to identify the content category of the real-time chat content;

[0124] The text conversion unit is configured to, if the content category of the real-time chat content is a non-text category, perform text conversion on the real-time chat content to obtain the target text content;

[0125] The direct setting unit is configured to, if the content category of the real-time chat content is a text category, directly set the real-time chat content as the target text content.

[0126] In this embodiment, the target text content acquisition module 902 includes: a speech-to-text unit and an OCR recognition unit. Specifically:

[0127] The speech-to-text unit is configured to, if the content category of the real-time chat content is a voice category, perform speech-to-text conversion on the real-time chat content to obtain the target text content;

[0128] The OCR recognition unit is configured to, if the content category of the real-time chat content is a displayed picture, use OCR text recognition technology to recognize the text information part in the displayed picture to obtain the target text content.

[0129] In this embodiment, the first risk detection result acquisition module 903 includes: a word segmentation processing unit, a standard sensitive word recognition unit, a custom sensitive word recognition unit, and a first risk detection result acquisition unit. Specifically:

[0130] The word segmentation processing unit is configured to input the target text content into a text word segmentation tool based on natural language to obtain a word segmentation processing result;

[0131] The standard sensitive word recognition unit is configured to use a loop comparison method to compare each word in the word segmentation processing result with all standard sensitive words included in the standard word library part of the sensitive word library, and identify all standard sensitive words included in the word segmentation processing result;

[0132] The custom sensitive word recognition unit is configured to use a loop comparison method to compare each word in the word segmentation processing result with all custom sensitive words included in the custom word library part of the sensitive word library, and identify all custom sensitive words included in the word segmentation processing result;

[0133] The first risk detection result obtaining unit is configured to sort out all standard sensitive words and all custom sensitive words included in the word segmentation processing result, and perform risk detection on the target text content to obtain the first risk detection result.

[0134] In this embodiment, the second risk detection result acquisition module 904 includes: a deep semantic extraction unit, an overall risk prediction unit, and a second risk detection result obtaining unit. Among them:

[0135] The deep semantic extraction unit is configured to input the target text content into a pre-trained semantic risk detection model based on the LSTM algorithm, and extract the keyword dependency relationship, keyword part of speech, word frequency, and phrase structure included in the target text content;

[0136] The overall risk prediction unit is configured to use the keyword dependency relationship, keyword part of speech, word frequency, and phrase structure included in the target text content as risk prediction features, predict the overall risk of the target text content, and output an overall risk index value;

[0137] The second risk detection result obtaining unit is configured to use the overall risk index value as the second risk detection result predicted by the semantic risk detection model.

[0138] In this embodiment, the chat content risk detection device 900 further includes: a plug-in packaging module, a risk detection plug-in installation module, a trigger start module, a content type recognition module, a text processing module, a plug-in first risk detection module, and a plug-in second risk detection module. Among them:

[0139] The plug-in packaging module is configured to package the pre-trained content type recognition model, the preset sensitive word library, the pre-trained semantic risk detection model, the natural language-based text word segmentation tool, and the preset risk processing strategy into a portable plug-in;

[0140] The risk detection plug-in installation module is configured to pre-install the portable plug-in into the startup function area of the target instant messaging tool;

[0141] The trigger start module is configured to trigger the startup of the portable plug-in in the startup function area when the target user sends the message sending instruction in the chat box of the target instant messaging tool;

[0142] The content type recognition module is configured to use the pre-trained content type recognition model in the portable plug-in to perform content type recognition on the user input content in the chat box of the target instant messaging tool;

[0143] A text processing module for performing text processing based on the content type recognition result to obtain the target text content;

[0144] A first risk detection module for the plug-in, which uses the preset sensitive word library and the text segmentation tool based on natural language in the portable plug-in to perform risk detection on the target text content to obtain the first risk detection result;

[0145] A second risk detection module for the plug-in, which uses the pre-trained semantic risk detection model in the portable plug-in to perform risk detection on the target text content to obtain the second risk detection result.

[0146] In this embodiment, the risk processing module 905 includes a text volume ratio calculation unit, an overall risk index value calculation unit, a comprehensive calculation unit, a comparison and judgment unit, and a risk processing unit. Among them:

[0147] The text volume ratio calculation unit is used to calculate the text volume ratio of all sensitive words in the target text content according to the first risk detection result;

[0148] The overall risk index value calculation unit is used to obtain the overall risk index value of the target text content according to the second risk detection result;

[0149] The comprehensive calculation unit is used to perform comprehensive calculation by combining the risk weights corresponding to the first risk detection result and the second risk detection result, the text ratio, and the overall risk index value to obtain the comprehensive risk value corresponding to the target text content;

[0150] The comparison and judgment unit is used to judge the size relationship between the comprehensive risk value and the preset risk threshold through comparison;

[0151] The risk processing unit is used to, if the comprehensive risk value exceeds the preset risk threshold, perform risk processing on the user input content in the chat box of the target instant messaging tool according to the preset risk processing strategy.

[0152] In this embodiment, the risk processing module 905 further includes a first risk processing subunit, a second risk processing subunit, a third risk processing subunit, and a fourth risk processing subunit. Among them:

[0153] The first risk processing subunit is used to display and intercept the real-time chat content and send a risk prompt to the instant sending end of the target user;

[0154] The second risk processing subunit is used to identify the position information of the risk information in the user input content and perform masked display through a pre-set masking processing component;

[0155] The third risk processing subunit is configured to perform display interception on the real-time chat content and enable a preset secure transmission interface to send the real-time chat content to a target receiving end;

[0156] The fourth risk processing subunit is configured to perform display interception on the real-time chat content, perform encryption and compression processing by using a preset encryption method to generate an encrypted file to be transmitted, and enable a preset secure transmission interface to send the encrypted file to be transmitted to a target receiving end.

[0157] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0158] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be completed at the same moment, but can be executed at different moments. Their execution order does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0159] To solve the above technical problems, the embodiments of the present application further provide a computer device. Specifically, please refer to Figure 10 , Figure 10 which is the basic structural block diagram of the computer device in this embodiment.

[0160] The computer device 10 includes a memory 10a, a processor 10b, and a network interface 10c that are communicatively connected to each other through a system bus. It should be noted that Figure 10Only a computer device 10 with a component memory 10a, a processor 10b, and a network interface 10c is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0161] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0162] The memory 10a includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10a can be an internal storage unit of the computer device 10, such as the hard disk or memory of the computer device 10. In other embodiments, the memory 10a can also be an external storage device of the computer device 10, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 10. Of course, the memory 10a can also include both the internal storage unit of the computer device 10 and its external storage device. In this embodiment, the memory 10a is generally used to store the operating system and various application software installed on the computer device 10, such as computer-readable instructions of a chat content risk detection method. In addition, the memory 10a can also be used to temporarily store various data that have been output or will be output.

[0163] In some embodiments, the processor 10b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 10b is generally used to control the overall operation of the computer device 10. In this embodiment, the processor 10b is used to run the computer-readable instructions stored in the memory 10a or process data, such as running the computer-readable instructions of the chat content risk detection method.

[0164] The network interface 10c may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device 10 and other electronic devices.

[0165] The computer device proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to the scenario where users and customer service have online instant communication during the financial insurance business service process. This application captures the real-time chat content to be displayed in the target instant messaging tool; performs text processing on the real-time chat content to obtain the target text content; performs word segmentation on the target text content, and uses a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain the first risk detection result; inputs the target text content into the pre-trained semantic risk detection model to obtain the second risk detection result; combines the first risk detection result, the second risk detection result and the preset risk processing strategy to perform risk processing on the real-time chat content. The chat content risk detection method is applied before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information in the chat interface when handling financial business online, thereby improving information security.

[0166] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to execute the steps of the chat content risk detection method as described above.

[0167] The computer-readable storage medium proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to the scenario where users have online instant communication with customer service during the process of financial insurance business services. This application captures the real-time chat content to be displayed in the target instant messaging tool; performs text processing on the real-time chat content to obtain the target text content; performs word segmentation on the target text content, and uses a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain the first risk detection result; inputs the target text content into a pre-trained semantic risk detection model to obtain the second risk detection result; combines the first risk detection result, the second risk detection result, and a preset risk handling strategy to perform risk handling on the real-time chat content. The chat content risk detection method is applied before the chat content is displayed on the target chat interface to prevent customers from exposing personal privacy or identity information in the chat interface when handling financial business online using instant messaging tools, thereby improving information security.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.

[0169] Obviously, the above-described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The preferred embodiments of this application are given in the drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of this application in other related technical fields is equally within the scope of the patent protection of this application. The non-company software tools or components that appear in the embodiments of this application are only for illustrative purposes and do not represent actual use.

Claims

1. A chat content risk detection method, characterized in that: The steps include: Based on the target user's message sending instruction, the real-time chat content to be displayed in the target instant messaging tool is captured; Performing text processing on the real-time chat content to obtain target text content; Perform word segmentation processing on the target text content, and use a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain a first risk detection result, wherein the first risk detection result represents a detection result obtained by performing risk detection only based on key word segmentation comparison; Inputting the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result predicted by the semantic risk detection model, wherein the second risk detection result represents a detection result obtained by performing risk detection according to deep semantic analysis; The real-time chat content is risk-handled in combination with the first risk detection result, the second risk detection result and a preset risk handling strategy.

2. The chat content risk detection method according to claim 1, characterized in that: The step of capturing the real-time chat content to be displayed in the target instant messaging tool based on the message sending instruction of the target user specifically includes: Acquire an API interface according to the real-time chat content corresponding to the target instant messaging tool, and create a corresponding real-time crawling task; The real-time capture task is started to capture the real-time chat content to be displayed in the target instant messaging tool.

3. The chat content risk detection method according to claim 1, characterized in that: The step of text-processing the real-time chat content to obtain target text content specifically includes: Uploading the captured real-time chat content to a pre-trained content type recognition model to identify the content category of the real-time chat content; If the content category of the real-time chat content is a non-text category, text processing is performed on the real-time chat content to obtain the target text content; If the content category of the real-time chat content is a text category, the real-time chat content is directly set as the target text content.

4. The chat content risk detection method according to claim 3, characterized in that: The step of performing word segmentation processing on the target text content, and using a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain the first risk detection result specifically includes: Inputting the target text content into a text segmentation tool based on natural language to obtain a segmentation processing result; By adopting a circular comparison method, the segmented words in the segmentation processing result are compared with all standard sensitive words contained in the standard word library part in the sensitive word library, and all standard sensitive words contained in the segmentation processing result are identified; By using a circular comparison method, the segmented words in the segmentation processing result are compared with all the custom sensitive words contained in the custom word library part in the sensitive word library, and all the custom sensitive words contained in the segmentation processing result are identified; All standard sensitive words and all custom sensitive words included in the word segmentation processing result are sorted, and risk detection is performed on the target text content to obtain the first risk detection result.

5. The chat content risk detection method according to claim 4, characterized in that: The step of inputting the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result predicted by the semantic risk detection model specifically includes: Input the target text content into a pre-trained semantic risk detection model based on the LSTM algorithm to extract the keyword dependency, keyword part of speech, word frequency and phrase structure contained in the target text content; Using the keyword dependency, keyword part of speech, word frequency and phrase structure contained in the target text content as risk prediction features, predicting the overall risk of the target text content, and outputting an overall risk index value; The overall risk index value is used as the second risk detection result predicted by the semantic risk detection model.

6. The chat content risk detection method according to claim 5, characterized in that: Before executing the step of performing risk processing on the real-time chat content in combination with the first risk detection result, the second risk detection result and a preset risk processing strategy, the method further includes: Encapsulating the pre-trained content type recognition model, the preset sensitive word library, the pre-trained semantic risk detection model, the natural language-based text segmentation tool, and the preset risk handling strategy into a portable plug-in; Preinstalling the portable plug-in into the startup function area of ​​the target instant messaging tool; When the target user sends the message sending instruction in the chat box of the target instant messaging tool, triggering the activation of the portable plug-in in the activation function area; Using the pre-trained content type recognition model in the portable plug-in to identify the content type of the user input content in the chat box of the target instant messaging tool, and, Perform text processing according to the content type recognition result to obtain the target text content; Using the preset sensitive word library and the natural language-based text segmentation tool in the portable plug-in, risk detection is performed on the target text content to obtain the first risk detection result; The pre-trained semantic risk detection model in the portable plug-in is used to perform risk detection on the target text content to obtain the second risk detection result.

7. The chat content risk detection method according to claim 6, characterized in that: The step of performing risk processing on the real-time chat content in combination with the first risk detection result, the second risk detection result and a preset risk processing strategy specifically includes: Using the preset risk handling strategy in the portable plug-in, risk handling is performed on the user input content in the chat box of the target instant messaging tool, specifically including: According to the first risk detection result, calculating the text volume ratio of all sensitive words in the target text content; According to the second risk detection result, obtaining an overall risk index value of the target text content; Combine the risk weights corresponding to the first risk detection result and the second risk detection result, the text proportion, and the overall risk index value, and perform a comprehensive calculation to obtain a comprehensive risk value corresponding to the target text content; By comparison, the magnitude relationship between the comprehensive risk value and the preset risk threshold is determined; If the comprehensive risk value exceeds a preset risk threshold, risk processing is performed on the user input content in the chat box of the target instant messaging tool according to the preset risk processing strategy.

8. A chat content risk detection device, characterized in that: include: A real-time chat content capture module is used to capture the real-time chat content to be displayed in the target instant messaging tool based on the message sending instruction of the target user; A target text content obtaining module is used to perform text processing on the real-time chat content to obtain target text content; A first risk detection result acquisition module is used to perform word segmentation processing on the target text content, and use a preset sensitive word library to compare and identify all sensitive words in the corresponding word segmentation results to obtain a first risk detection result, wherein the first risk detection result represents a detection result obtained by performing risk detection only based on key word comparison; A second risk detection result acquisition module is used to input the target text content into a pre-trained semantic risk detection model to obtain a second risk detection result predicted by the semantic risk detection model, wherein the second risk detection result represents a detection result obtained by performing risk detection according to deep semantic analysis; The risk processing module is used to perform risk processing on the real-time chat content in combination with the first risk detection result, the second risk detection result and a preset risk processing strategy.

9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the chat content risk detection method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the chat content risk detection method according to any one of claims 1 to 7 are implemented.