Message interception method and device, electronic device and storage medium
By identifying keywords and types in the intercepted request text, and combining them with an interception keyword library and tags, risk prediction and semantic analysis are performed. This solves the message blocking problem caused by frequent user request interception, and achieves more efficient message interception and user control.
Patent Information
- Application Number
- CN202511066320.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-07-31
AI Technical Summary
In existing technologies, frequent interception requests from users can lead operators to misjudge them as large-scale fraudulent messages, resulting in global message blocking and reducing the accuracy of message interception.
By acquiring keywords from the interception request text, identifying the request type, and generating a personal interception tendency label, combined with a pre-set interception keyword library and global interception labels, risk prediction and semantic analysis are performed to determine the risk category of the original message and precisely control the interception operation.
It improves the accuracy of message interception, reduces excessive blocking of user numbers by operators, and protects user privacy and security.
Smart Images

Figure CN120812595B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and is applicable to the fields of financial technology and medical technology. In particular, it relates to a message interception method and apparatus, electronic device and storage medium. Background Technology
[0002] Currently, many applications and fraudulent organizations push messages to users via messaging services. In fintech scenarios, these messages may be related to loan services, insurance product promotions, etc. In the medical technology field, they may push messages related to drug or medical service activities. Users can send the fields "TD" or "R" to have the operator unbind their phone number from the company / service in the messaging channel, thus blocking the company / service from sending messages to that phone number again. However, when users frequently send these fields, the operator will determine that the user has received large-scale fraudulent messages, and will then globally block the user's phone number from receiving messages, causing the user to not receive any messages and need to contact the operator to restore related services, ultimately reducing the accuracy of message blocking.
[0003] Therefore, improving the accuracy of message interception has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main objective of this application is to provide a message interception method, apparatus, electronic device, and storage medium, which aims to improve the accuracy of message interception.
[0005] To achieve the above objectives, a first aspect of this application proposes a message interception method, the method comprising:
[0006] Obtain the interception request text and the transmission number used to transmit the interception request text;
[0007] Keyword extraction is performed on the interception request text to obtain the interception keywords;
[0008] Based on the interception keywords, the interception request text is type-identified to obtain the interception request type, and the interception request type is used as a personal interception tendency label for the transmission number;
[0009] Obtain the original message text; wherein, the transmission number is the receiving number of the original message text;
[0010] Based on a preset block keyword library and the individual block tendency tags, the risk of the original message text is predicted to obtain the message risk category;
[0011] The original message text may be intercepted or not, depending on the message risk category.
[0012] In some embodiments, the step of performing risk prediction on the original message text based on a preset interception keyword library and the individual interception tendency label to obtain a message risk category includes:
[0013] Obtain the interception tag library, which stores multiple personal interception tendency tags;
[0014] Based on the tag type of the individual interception tendency tag, determine the number of tags of each tag type in the interception tag library;
[0015] Global blocking tags are selected from multiple individual blocking tendency tags based on the number of tags;
[0016] The risk type of the original message text is identified by the interception keyword library, the personal interception tendency label, and the global interception label.
[0017] In some embodiments, the step of identifying the risk type of the original message text based on the interception keyword library, the personal interception tendency label, and the global interception label to obtain the message risk type of the original message text includes:
[0018] A text query is performed on the interception keyword database based on the original message text to obtain query results; wherein, the query results indicate that there is text in the original message text that matches the interception keyword database, or that there is no text in the original message text that matches the interception keyword database;
[0019] If the query result indicates that there is no text in the original message text that matches the interception keyword database, content matching is performed based on the original message text, the personal interception tendency tag, and the global interception tag to obtain a matching score;
[0020] The message risk type is determined based on the matching score and a preset score threshold.
[0021] In some embodiments, the step of performing content matching based on the original message text, the personal blocking tendency label, and the global blocking label to obtain a matching score includes:
[0022] Semantic analysis is performed on the original message text to obtain text semantic features;
[0023] A matching degree analysis is performed on the semantic features of the text and the global interception tags to obtain a first score;
[0024] A second score is obtained by performing a matching degree analysis on the semantic features of the text and the personal interception tendency tags;
[0025] A matching score is obtained based on the first score and the second score.
[0026] In some embodiments, determining the message risk type based on the matching score and a preset score threshold includes:
[0027] If the first score or the second score is greater than or equal to the preset score threshold, then the message risk type is determined to be a high-risk message;
[0028] If the first score and the second score are both less than the preset score threshold, then the message risk type is determined to be a low-risk message.
[0029] In some embodiments, before performing risk prediction on the original message text based on a preset interception keyword library and the individual interception tendency label to obtain the message risk category, the method further includes:
[0030] The interception request text is subjected to intent recognition to obtain the request type; the request type indicates whether the interception request text requests the addition of the interception keyword or requests the deletion of the interception keyword;
[0031] The interception keyword database is updated according to the request type and the interception keyword;
[0032] The interception tag library is updated based on the individual interception tendency tag of the transmitted number.
[0033] In some embodiments, after updating the intercept keyword database according to the request type and the intercept keyword, the method further includes:
[0034] Once the intercepted keyword database is updated, the preset response text is obtained;
[0035] Send the response text to the transmission number.
[0036] To achieve the above objectives, a second aspect of this application provides a message interception device, the device comprising:
[0037] The first text acquisition module is used to acquire the interception request text and the transmission number used to transmit the interception request text;
[0038] The keyword extraction module is used to extract keywords from the intercepted request text to obtain intercepted keywords;
[0039] The type identification module is used to identify the type of the interception request text based on the interception keyword, obtain the interception request type, and use the interception request type as a personal interception tendency label for the transmission number.
[0040] The second text acquisition module is used to acquire the original message text; wherein, the transmission number is the receiving number of the original message text;
[0041] The risk prediction module is used to predict the risk of the original message text based on a preset block keyword library and the personal block tendency label, and obtain the message risk category.
[0042] The message interception module is used to intercept the original message text or not intercept the original message text according to the message risk category.
[0043] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0044] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0045] The message interception method, apparatus, electronic device, and storage medium proposed in this application acquire interception request text transmitted by a transmission number and extract interception keywords from the interception request text. Then, based on the interception keywords, the interception request type of the interception request text is determined, and this interception request type is used as a personal interception tendency label for the transmission number, ensuring that users can independently control the content intercepted. The original message text is acquired, and then the risk category of the received original message text is determined according to the interception keyword library and the personal interception tendency label. Finally, based on the risk category, it is determined whether to intercept the original message text. The embodiments of this application can reduce the risk of operators excessively blocking user numbers while protecting user privacy and security, ultimately improving the accuracy of message interception. Attached Figure Description
[0046] Figure 1 This is a flowchart of the message interception method provided in the embodiments of this application;
[0047] Figure 2 This is another flowchart of the message interception method provided in the embodiments of this application;
[0048] Figure 3 yes Figure 1 The flowchart of step S105 in the process;
[0049] Figure 4 yes Figure 3 The flowchart of step S304 in the process;
[0050] Figure 5 yes Figure 4 The flowchart of step S402 in the document;
[0051] Figure 6 yes Figure 4 The flowchart of step S403 in the process;
[0052] Figure 7 This is a schematic diagram of the message interception device provided in the embodiments of this application;
[0053] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0055] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0057] First, let's analyze some of the terms used in this application:
[0058] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field combining computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, message intent recognition, message extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0059] Information extraction is a text processing technique that extracts factual messages of a specified type, such as entities, relationships, and events, from natural language text and outputs them as structured data. It's a technique for extracting specific messages from text data. Text data consists of concrete units, such as sentences, paragraphs, and chapters. Text messages are also composed of smaller concrete units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of information extraction. Of course, information extraction techniques can extract messages of various types.
[0060] Currently, many applications and fraudulent organizations push messages to users via messaging services. In fintech scenarios, these messages may be related to loan services, insurance product promotions, etc. In the medical technology field, they may push messages related to drug or medical service activities. Users can send the fields "TD" or "R" to have the operator unbind their phone number from the company / service in the messaging channel, thus blocking the company / service from sending messages to that phone number again. However, when users frequently send these fields, the operator will determine that the user has received large-scale fraudulent messages, and will then globally block the user's phone number from receiving messages, causing the user to not receive any messages and need to contact the operator to restore related services, ultimately reducing the accuracy of message blocking.
[0061] Therefore, improving the accuracy of message interception has become an urgent technical problem to be solved.
[0062] Based on this, embodiments of this application provide a message interception method and apparatus, electronic device and storage medium, aiming to improve the accuracy of message interception.
[0063] The message interception method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the message interception method in this application is described.
[0064] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0065] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0066] The message interception method provided in this application relates to the field of wireless communication technology. The message interception method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the message interception method, but is not limited to the above forms.
[0067] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0068] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user messages, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0069] Figure 1 This is an optional flowchart of the message interception method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.
[0070] Step S101: Obtain the interception request text and the transmission number used to transmit the interception request text.
[0071] Step S102: Extract keywords from the intercepted request text to obtain the intercepted keywords.
[0072] Step S103: Based on the interception keywords, the interception request text is type-identified to obtain the interception request type, and the interception request type is used as a personal interception tendency label for the transmission number.
[0073] Step S104: Obtain the original message text; wherein, the transmission number is the receiving number of the original message text.
[0074] Step S105: Based on the preset interception keyword library and personal interception tendency tags, perform risk prediction on the original message text to obtain the message risk category.
[0075] Step S106: Based on the message risk category, the original message text may be intercepted or not.
[0076] Steps S101 to S106 of this embodiment involve acquiring the interception request text transmitted by the transmitting number and extracting interception keywords from the interception request text. Then, based on the interception keywords, the interception request type of the interception request text is determined, and this interception request type is used as a personal interception tendency label for the transmitting number to ensure that users can independently control the content being intercepted. The original message text is acquired, and then the risk category of the received original message text is determined based on the interception keyword library and the personal interception tendency label. Finally, based on the risk category, it is determined whether to intercept the original message text. This embodiment can reduce the risk of operators excessively blocking user numbers while protecting user privacy and security, ultimately improving the accuracy of message interception.
[0077] In step S101 of some embodiments, the interception request text refers to a request sent by a user through a message, application, or other messaging channel, aimed at controlling and managing the content of message interception. The transmission number can be the phone number of the user sending the interception request text. Taking a financial scenario as an example, the interception request text sent by the user through the transmission number can be "interception keywords: loan, insurance policy, repayment, payment". Taking a medical scenario as another example, the interception request text can be: "interception keywords: health consultation, drug recommendation, drug price reduction", and is not limited thereto.
[0078] In step S102 of some embodiments, the interception keywords are user-defined keywords in the interception request text. For example, if the interception request text is "interception keywords: loan, policy, repayment, payment", then the extracted interception keywords are "loan", "policy", "repayment" and "payment" as interception keywords.
[0079] Keyword extraction can be achieved by setting predefined rules to match specific words or phrases in the intercepted request text. Alternatively, it can be compared with a predefined dictionary to extract keywords that match the message content. Furthermore, keyword extraction can also be performed using support vector machines, BERT-based pre-trained models, and methods based on term frequency-inverse document frequency (TF-IDF).
[0080] In step S103 of some embodiments, the interception request type is the category of interception behavior expressed by the interception request text. The personal interception preference label is an identifier generated individually for each user based on the user's interception request type, used to describe the user's preference for intercepting specific types of messages.
[0081] In some embodiments, semantic recognition can be performed on the intercepted request keywords first. A pre-trained classification model can be used to classify the semantic features and contextual relationships of the intercepted request keywords. For example, when the intercepted keyword is "loan" or "credit card application," the intercepted request type is identified as an advance loan. When the intercepted keyword is "fast loan disbursement," "interest-free loan," "free medicine," "free medical check-up," or "high bonus," the intercepted request type is identified as fraud.
[0082] In step S104 of some embodiments, the original message text refers to the message text that has not been intercepted or judged before being sent by the operator or other service provider. For example, a bank sends a message to a user, where the user's mobile phone number is the transmission number, and the original message text is the content of the message.
[0083] Please see Figure 2 In some embodiments, prior to step S105, the message interception method provided in this application may include, but is not limited to, steps S201 to S203:
[0084] Step S201: Perform intent recognition on the interception request text to obtain the request type; the request type represents whether the interception request text requests to add interception keywords or requests to delete interception keywords.
[0085] Step S202: Update the interception keyword database according to the request type and interception keywords.
[0086] Step S203: Update the interception tag database based on the individual interception tendency tag of the transmission number.
[0087] In step S201 of some embodiments, intent recognition refers to analyzing the interception request text sent by the user to determine the user's specific needs and intent, thereby deriving the corresponding request type. In this embodiment, the request types include: adding interception keywords and deleting interception keywords. Intent recognition can be implemented by distinguishing user needs through rule matching, keyword recognition, or context-based analysis.
[0088] For example, if the interception request text is "Add interception keywords: loan, credit card application", the request will be identified as a request to add interception keywords through intent recognition; while if the user sends "Remove interception keywords: credit card application", the request type will be a request to delete interception keywords.
[0089] In step S202 of some embodiments, the interception keyword library is a database that stores interception keywords, including multiple transmission numbers and the interception keyword corresponding to each transmission number. First, the interception keywords in the request text are parsed. Then, based on the request type, it is determined whether to add new keywords to the interception keyword library or delete existing keywords.
[0090] In some embodiments, once the interception tag library update is complete, a preset response text is retrieved. The response text may be a notification to the user that the interception operation has been successfully completed, such as "Interception keyword added successfully" or "Interception keyword deleted successfully." The response text is then sent to the transmission number. This response text can be sent via SMS, email, or other suitable communication methods to ensure that the user receives timely feedback on the operation result.
[0091] In step S203 of some embodiments, the interception tag library is a database that stores user interception messages, which stores multiple transmission numbers and personal interception tendency tags corresponding to each transmission number.
[0092] Furthermore, based on the request type, it is determined whether to add new blocking keywords or delete existing ones. If it is an addition operation, the new blocking keyword is used for type identification to obtain a new personal blocking tendency tag, which is then added to the tag library. If it is a deletion operation, the corresponding personal blocking tendency tag is updated based on the remaining blocking keywords.
[0093] For example, if a user requests to "add intercepted keywords: loan, credit card application, drug recommendation", the newly added intercepted keywords are type-identified, and the corresponding personal interception tendency tags "loan-related" and "medical-related" are added to the interception tag library. If the user requests to "delete intercepted keyword: drug recommendation", the intercepted keywords corresponding to the user's transmission number in the intercepted keyword library will be "loan, credit card application". After re-type identification, the user's personal interception tendency tag will be updated to "loan-related", thus updating the interception tag library.
[0094] Steps S201 to S203, as illustrated in this embodiment, introduce an intent recognition mechanism to accurately determine the specific intent of the user's request, i.e., whether to add or delete interception keywords, thereby ensuring the flexibility and personalization of the message interception strategy. Subsequently, the interception keyword library and interception tag library are updated in real time according to the user's request, and the interception strategy is further refined by combining personal interception preference tags of the transmission number. Based on the user's past interception behavior and preferences, the interception tag library is automatically updated with interception messages that match the user's needs, thereby providing a more accurate and personalized interception service.
[0095] In step S105 of some embodiments, please refer to Figure 3 Step S105 may include, but is not limited to, steps S301 to S304:
[0096] Step S301: Obtain the interception tag library, which stores multiple personal interception tendency tags.
[0097] Step S302: Determine the number of tags of each tag type in the interception tag library based on the tag type of the personal interception tendency tag.
[0098] Step S303: Filter out global blocking tags from multiple individual blocking tendency tags based on the number of tags.
[0099] Step S304: Based on the interception keyword library, personal interception tendency tags, and global interception tags, identify the risk type of the original message text to obtain the message risk type of the original message text.
[0100] In step S301 of some embodiments, the interception tag library is a database that stores user interception messages, which stores multiple transmission numbers and personal interception tendency tags corresponding to each transmission number.
[0101] In step S302 of some embodiments, the number of personal blocking tendency tags of each type in the blocking tag library is counted. For example, the blocking tag library stores personal blocking tendency tags in categories such as "financial", "medical", and "advertising", and the number of tags under each category is counted. Among them, there are 50 personal blocking tendency tags for "financial", 30 personal blocking tendency tags for "medical", and 10 personal blocking tendency tags for "advertising".
[0102] In step S303 of some embodiments, the tag category with the most tags is determined, and the personal blocking tendency tag corresponding to the tag category is used as the global blocking tag.
[0103] In step S304 of some embodiments, please refer to Figure 4 Step S304 may include, but is not limited to, steps S401 to S403:
[0104] Step S401: Perform a text query on the interception keyword database based on the original message text to obtain the query results; wherein, the query results indicate that there is text in the original message text that matches the interception keyword database, or that there is no text in the original message text that matches the interception keyword database.
[0105] Step S402: If the query result indicates that there is no text in the original message text that matches the interception keyword library, perform content matching based on the original message text, personal interception tendency tags, and global interception tags to obtain a matching score.
[0106] Step S403: Determine the message risk type based on the matching score and the preset score threshold.
[0107] In step S401 of some embodiments, text query refers to comparing the original message text with the intercepted keyword library to determine whether the original message text contains any matching intercepted keywords. For example, if the original message text is "Your loan application has been approved," and the intercepted keyword library contains the keyword "loan," then "loan" will be found and a matching result will be returned. If the original message text is "Your health check report has been completed," and the intercepted keyword library does not contain any keywords related to "health check report," the query result will indicate that there is no text matching the intercepted keyword library.
[0108] In step S402 of some embodiments, if the query result indicates that there is text in the original message text that matches the interception keyword library, then the message risk type of the original message text is directly determined as a high-risk message so that it can be intercepted in the future.
[0109] If the query results indicate that the original message text does not contain any text that matches the intercepted keyword database, please refer to [link / reference]. Figure 5 Step S402 may also include, but is not limited to, steps S501 to S504:
[0110] Step S501: Perform semantic analysis on the original message text to obtain text semantic features.
[0111] Step S502: Perform a matching degree analysis on the text semantic features and global interception tags to obtain the first score.
[0112] Step S503: Perform a matching degree analysis on the text semantic features and personal interception tendency labels to obtain a second score.
[0113] Step S504: Obtain the matching score based on the first score and the second score.
[0114] In step S501 of some embodiments, a semantic recognition model can be trained using natural language processing (NLP) technology, and the semantic recognition model can be used to perform word segmentation, part-of-speech tagging, syntactic analysis and other processing on the original message text in order to extract the semantic features of the original message text, i.e., text semantic features.
[0115] In step S502 of some embodiments, the first score is the similarity between text semantic features and global interception tendency tags. In some embodiments, the user needs to confirm that global interception mode is enabled before proceeding. Global interception tags can effectively reduce missed interceptions because they are defined based on broad user needs and common types of risky messages. Even if some message text does not directly match the user's personal interception tendency tag, this embodiment can still determine whether it belongs to high-risk content through global tags, thereby preventing these messages from being missed.
[0116] The calculation of the first score can include comparing keywords in the original message text with keywords in the global interception tags. If the original message text contains keywords related to the global interception tags, a matching score is calculated based on the number of matching keywords; a higher score indicates a stronger matching. Another commonly used method is based on the Term Frequency-Inverse Document Frequency (TF-IDF) model. By calculating the TF-IDF value of each word in the original message text and the global interception tags, the importance and relevance of keywords in the text are evaluated, and then methods such as cosine similarity are used for matching analysis. This method can more accurately determine the importance of keywords in the text and their relationship with the tags. Alternatively, word vector models (such as Word2Vec or GloVe) can be used to convert words into vectors, and the similarity between the word vectors of the original message text and the global tags can be calculated. In some other embodiments, a text classification method based on a machine learning model can also be used. Using a trained classification model, the probability of the original message text belonging to a certain tag category can be calculated based on its content, and this probability value is used as the matching score.
[0117] In step S503 of some embodiments, the second score is the similarity between text semantic features and personal interception tendency tags. The calculation method of the second score is the same as that of the embodiment of step S502, and will not be repeated here.
[0118] In step S504 of some embodiments, the matching score includes a first score and a second score, which are independent of each other and are not calculated.
[0119] Steps S501 to S504 shown in the embodiments of this application provide a more accurate message interception method by combining semantic analysis, matching degree analysis, and multi-dimensional score evaluation. By performing semantic analysis on the original message text, the message content can be understood in depth, avoiding the problems of false or missed interception caused by traditional simple keyword-based matching.
[0120] In step S403 of some embodiments, please refer to Figure 6 Step S403 includes, but is not limited to, steps S601 to S602:
[0121] Step S601: If the first score or the second score is greater than or equal to the preset score threshold, then the message risk type is determined to be a high-risk message.
[0122] Step S602: If the first score and the second score are less than the preset score threshold, then the message risk type is determined to be a low-risk message.
[0123] In step S601 of some embodiments, the calculated first score or second score is compared according to a preset scoring threshold. If the first score or second score is greater than or equal to the preset scoring threshold, it indicates that the content of the message meets the interception criteria and may be fraudulent, advertising, or other irrelevant messages. Ultimately, the message is classified as a high-risk message and intercepted. For example, if the first score is 0.85, the second score is 0.75, and the preset scoring threshold is 0.8, the message is classified as a high-risk message and interception is performed.
[0124] In step S602 of some embodiments, if both the first score and the second score are less than a preset score threshold, it indicates that the message has a low degree of matching with high-risk content and meets the standard of a normal message. Finally, the risk type of the message is determined to be a low-risk message and no interception is required.
[0125] Steps S601 to S602, as illustrated in this embodiment, enable a more intelligent and accurate assessment of message risk types through a comprehensive judgment of the first and second scores. In this embodiment, the scoring threshold can be adjusted according to actual conditions to improve the accuracy of interception. By classifying message texts into high-risk and low-risk messages, unwanted spam and fraudulent messages are efficiently filtered out, while false interception of normal messages is reduced, thus improving user experience.
[0126] Steps S401 to S403 shown in this embodiment combine a keyword blocking library, personal blocking preference tags, and global blocking tags to intelligently determine the risk level of messages based on the user's actual needs and common risk types, thus avoiding the problems of false blocking and missed blocking in traditional blocking methods.
[0127] Steps S301 to S304, as illustrated in this embodiment, improve the accuracy and flexibility of message interception by combining a comprehensive analysis of the interception keyword library, personal interception preference tags, and global interception tags, avoiding false and missed interceptions while meeting user interception needs. Through intelligent filtering and updating of the interception tag library, it can adapt to changing user needs in real time, making message interception more accurate and efficient.
[0128] In step S106 of some embodiments, the risk category predicted in step S105 will be intercepted if the message is determined to be high-risk (such as fraud, harassment, etc.) to prevent it from being sent to the user. If the message is determined to be low-risk, it will not be intercepted and will be allowed to be sent normally. For example, if the content of the original message text is determined to be fraudulent, the message will be intercepted and the user will not receive it; however, if the message is a normal bank notification, it will not be intercepted and will be delivered to the user normally.
[0129] By following the steps above, we can accurately intercept user messages and protect users from fraudulent and harassing messages.
[0130] Please see Figure 7 This application also provides a message interception device that can implement the above-described message interception method. The device includes:
[0131] The first text acquisition module 701 is used to acquire the interception request text and the transmission number used to transmit the interception request text;
[0132] Keyword extraction module 702 is used to extract keywords from the intercepted request text to obtain the intercepted keywords;
[0133] The type recognition module 703 is used to perform type recognition on the interception request text based on the interception keywords, obtain the interception request type, and use the interception request type as a personal interception tendency label for the transmission number;
[0134] The second text acquisition module 704 is used to acquire the original message text; wherein, the transmission number is the receiving number of the original message text;
[0135] The risk prediction module 705 is used to predict the risk of the original message text based on a preset block keyword library and personal block tendency tags to obtain the message risk category;
[0136] The message interception module 706 is used to intercept the original message text or not intercept the original message text based on the message risk category.
[0137] The specific implementation of this message interception device is basically the same as the specific implementation of the message interception method described above, and will not be repeated here.
[0138] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned message interception method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0139] Please see Figure 8 , Figure 8The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0140] The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0141] The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the message interception method of the embodiments of this application.
[0142] The 803 input / output interface is used to implement message input and output.
[0143] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0144] Bus 805 transmits messages between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);
[0145] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0146] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described message interception method.
[0147] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0148] The message interception method, device, electronic device, and storage medium provided in this application acquire interception request text transmitted by a transmission number and extract interception keywords from the interception request text. Then, based on the interception keywords, the interception request type of the interception request text is determined, and this interception request type is used as a personal interception tendency label for the transmission number, ensuring that users can independently control the content intercepted. The original message text is acquired, and then the risk category of the received original message text is determined according to the interception keyword library and the personal interception tendency label. Finally, based on the risk category, it is determined whether to intercept the original message text. This application embodiment can reduce the risk of operators excessively blocking user numbers while protecting user privacy and security, ultimately improving the accuracy of message interception.
[0149] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0150] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0152] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0153] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0154] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0156] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0159] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A message interception method, characterized in that, The method includes: Obtain the interception request text and the transmission number used to transmit the interception request text; Keyword extraction is performed on the interception request text to obtain the interception keywords; Based on the interception keywords, the interception request text is type-identified to obtain the interception request type, and the interception request type is used as a personal interception tendency label for the transmission number; Obtain the original message text; wherein, the transmission number is the receiving number of the original message text; Based on a preset block keyword library and the individual block tendency tags, the risk of the original message text is predicted to obtain the message risk category; The original message text may be intercepted or not, depending on the message risk category. The step of performing risk prediction on the original message text based on a preset interception keyword library and the individual interception tendency label to obtain the message risk category includes: Obtain the interception tag library, which stores multiple personal interception tendency tags; Based on the tag type of the individual interception tendency tag, determine the number of tags of each tag type in the interception tag library; Global blocking tags are selected from multiple individual blocking tendency tags based on the number of tags; A text query is performed on the interception keyword database based on the original message text to obtain query results; wherein, the query results indicate that there is text in the original message text that matches the interception keyword database, or that there is no text in the original message text that matches the interception keyword database; If the query result indicates that there is no text in the original message text that matches the interception keyword database, content matching is performed based on the original message text, the personal interception tendency tag, and the global interception tag to obtain a matching score; The message risk type is determined based on the matching score and a preset score threshold. The step of performing content matching based on the original message text, the personal interception tendency label, and the global interception label to obtain a matching score includes: Semantic analysis is performed on the original message text to obtain text semantic features; A matching degree analysis is performed on the semantic features of the text and the global interception tags to obtain a first score; A second score is obtained by performing a matching degree analysis on the semantic features of the text and the personal interception tendency tags; The matching score is obtained based on the first score and the second score.
2. The method according to claim 1, characterized in that, The process of determining the message risk type based on the matching score and a preset score threshold includes: If the first score or the second score is greater than or equal to the preset score threshold, then the message risk type is determined to be a high-risk message; If the first score and the second score are both less than the preset score threshold, then the message risk type is determined to be a low-risk message.
3. The method according to claim 1, characterized in that, Before performing risk prediction on the original message text based on a preset interception keyword library and the individual interception tendency label to obtain the message risk category, the method further includes: The interception request text is subjected to intent recognition to obtain the request type; the request type indicates whether the interception request text requests the addition of the interception keyword or requests the deletion of the interception keyword; Update the interception keyword database according to the request type and the interception keyword; The interception tag library is updated based on the individual interception tendency tag of the transmitted number.
4. The method according to claim 3, characterized in that, After updating the interception keyword database according to the request type and the interception keyword, the method further includes: Once the intercepted keyword database is updated, the preset response text is obtained; Send the response text to the transmission number.
5. A message interception device, characterized in that, The device includes: The first text acquisition module is used to acquire the interception request text and the transmission number used to transmit the interception request text; The keyword extraction module is used to extract keywords from the intercepted request text to obtain intercepted keywords; The type identification module is used to identify the type of the interception request text based on the interception keyword, obtain the interception request type, and use the interception request type as a personal interception tendency label for the transmission number. The second text acquisition module is used to acquire the original message text; wherein, the transmission number is the receiving number of the original message text; The risk prediction module is used to predict the risk of the original message text based on a preset block keyword library and the personal block tendency label, and obtain the message risk category. The message interception module is used to intercept the original message text or not intercept the original message text according to the message risk category. The step of performing risk prediction on the original message text based on a preset interception keyword library and the individual interception tendency label to obtain the message risk category includes: Obtain the interception tag library, which stores multiple personal interception tendency tags; Based on the tag type of the individual interception tendency tag, determine the number of tags of each tag type in the interception tag library; Global blocking tags are selected from multiple individual blocking tendency tags based on the number of tags; A text query is performed on the interception keyword database based on the original message text to obtain query results; wherein, the query results indicate that there is text in the original message text that matches the interception keyword database, or that there is no text in the original message text that matches the interception keyword database; If the query result indicates that there is no text in the original message text that matches the interception keyword database, content matching is performed based on the original message text, the personal interception tendency tag, and the global interception tag to obtain a matching score; The message risk type is determined based on the matching score and a preset score threshold. The step of performing content matching based on the original message text, the personal interception tendency label, and the global interception label to obtain a matching score includes: Semantic analysis is performed on the original message text to obtain text semantic features; A matching degree analysis is performed on the semantic features of the text and the global interception tags to obtain a first score; A second score is obtained by performing a matching degree analysis on the semantic features of the text and the personal interception tendency tags; The matching score is obtained based on the first score and the second score.
6. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Short message identification method and device and electronic equipment
CN109684639A