Short message processing method, electronic equipment and computer readable storage medium

By detecting and converting abnormal feature characters in SMS, combining suspicious feature quantization and semantic classification models, the illegal suspiciousness of SMS is automatically identified, which solves the problem of missed blocking in the traditional spam SMS monitoring system and realizes efficient identification of illegal SMS.

CN120456031APending Publication Date: 2025-08-08ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410174374.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional spam message monitoring systems are difficult to effectively identify strategies such as slow-sending and content mutations used by criminals, resulting in a large number of illegal text messages being missed, and the keyword strategy is too loose or too strict to affect user communication needs.

Method used

By detecting abnormal feature characters in the SMS, converting them into normal text, legality evaluation is performed, and combining suspicious feature quantization model, clustering analysis and semantic classification model, the illegal suspiciousness of the SMS is automatically identified without policy configuration.

Benefits of technology

It realizes automatic recognition of the mutated features of SMS content and keywords, reduces the complexity and workload of illegal SMS recognition, and improves the recognition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456031A_ABST
    Figure CN120456031A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a short message processing method, electronic equipment and a computer readable storage medium. The short message processing method comprises the following steps: receiving a short message to be audited; detecting whether each character in the short message belongs to one or more preset abnormal feature characters, and evaluating a detection result to obtain a first evaluation result; under the condition that the characters in the short message contain the abnormal feature characters, converting the abnormal feature characters into normal characters; evaluating the legality of the short message content converted into the normal characters to obtain a second evaluation result; obtaining the illegal suspicion degree of the short message according to the first evaluation result and the second evaluation result; and generating an auditing result of the short message according to the illegal suspiciousness. According to the scheme of the embodiment, the variation characteristics of the short message content and the keyword are automatically identified, strategy configuration is not needed, and the complexity and workload of illegal short message identification are greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communications, and in particular to a text message processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] As a fundamental telecommunications service, SMS offers advantages such as ease of use, wide coverage, and high reach, making it the preferred channel for numerous scenarios, including financial transactions, verification codes, express delivery, and disaster warning notifications. However, criminals exploit these advantages, combining them with internet applications and URL (Uniform Resource Locator) links to engage in spam and even telecommunications-related illegal activities. These activities are diverse, covert, and covert, affecting all aspects of life, causing harm to society. Criminals evade traditional spam monitoring systems by using tactics such as slow delivery to multiple numbers, penetrating traffic flows, and content variation and keyword targeting. Furthermore, traditional spam monitoring systems rely on user complaints for keyword blocking, resulting in a significant number of illicit messages being missed. Consequently, overly lax spam monitoring policies can easily lead to exploitation, while overly strict policies can disrupt user communication needs. Traditional spam monitoring systems face significant challenges. Summary of the Invention

[0003] Embodiments of the present disclosure provide a text message processing method, an electronic device, and a computer-readable storage medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for processing a text message, the method comprising:

[0005] Receive SMS messages awaiting review;

[0006] Detecting whether each character in the text message belongs to one or more preset abnormal characteristic characters, and evaluating the detection results to obtain a first evaluation result;

[0007] If the characters in the text message contain abnormal characteristic characters, convert the abnormal characteristic characters into normal text;

[0008] Evaluate the legitimacy of the short message after conversion into normal text to obtain a second evaluation result;

[0009] Determining the suspicious degree of illegality of the text message according to the first evaluation result and the second evaluation result;

[0010] The review result of the text message is generated according to the illegal suspicion level.

[0011] In a second aspect, an embodiment of the present disclosure provides an electronic device, which may include:

[0012] one or more processors;

[0013] a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the text message processing method;

[0014] One or more input / output (I / O) interfaces are connected between the processor and the memory and configured to implement information interaction between the processor and the memory.

[0015] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the text message processing method is implemented.

[0016] The disclosed embodiment evaluates whether the characters in a text message are abnormal and whether the text is legal, and determines the illegality suspicion of the text message based on the evaluation results, thereby determining whether to intercept the text message based on the illegality suspicion. This realizes automatic recognition of variant features of text message content and keywords without the need for policy configuration, greatly reducing the complexity and workload of illegal text message recognition, and adding the recognition results of variant feature characters to the judgment of the illegality suspicion of the text message, thereby improving the recognition rate of illegal text messages. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In the accompanying drawings of the embodiments of the present disclosure:

[0018] Figure 1 A flowchart of a method for processing short messages provided in an embodiment of the present disclosure;

[0019] Figure 2 A first structural diagram of the text message processing method provided in an embodiment of the present disclosure;

[0020] Figure 3 A second structural diagram of the text message processing method provided in an embodiment of the present disclosure;

[0021] Figure 4 A third structural diagram of the SMS processing method provided in an embodiment of the present disclosure;

[0022] Figure 5 A flowchart of a method for processing SMS messages in which the suspicious feature quantification model, cluster analysis module, and semantic classification model provided in an embodiment of the present disclosure are all involved;

[0023] Figure 6 A schematic diagram of a method for processing short messages provided in an embodiment of the present disclosure;

[0024] Figure 7 This is a block diagram of the electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the technical solution of the present disclosure, the communication perception data processing method and computer-readable storage medium provided by the embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0026] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the illustrated embodiments may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that the present disclosure will be thorough and complete and will fully understand the scope of the present disclosure to those skilled in the art.

[0027] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing the detailed embodiments with reference to the accompanying drawings.

[0028] The present disclosure may be described with reference to plan views and / or cross-sectional views by way of ideal schematic views of the present disclosure. Therefore, the exemplary illustrations may be modified according to manufacturing techniques and / or tolerances.

[0029] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0030] The terms used in this disclosure are only used to describe specific embodiments and are not intended to limit the disclosure. As used in this disclosure, the term "and / or" includes any and all combinations of one or more related enumerated items. As used in this disclosure, the singular forms "a" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. As used in this disclosure, the terms "comprising" and "made of" specify the presence of the features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.

[0031] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meanings as those commonly understood by those skilled in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined in this disclosure.

[0032] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the drawings illustrate the specific shapes of the regions of the elements, but are not intended to be limiting.

[0033] As a fundamental telecommunications service, SMS offers advantages such as ease of use, wide coverage, and high reach, making it the preferred channel for numerous scenarios, including financial account transactions, verification codes, express logistics, and disaster warning notifications. However, criminals exploit these advantages, combining them with internet applications and URL (Uniform Resource Locator) links to send spam messages and engage in telecom violations. These activities are diverse, covert, and covert, affecting all aspects of life, causing harm to society. Criminals evade traditional spam monitoring systems by exploiting traffic-penetrating strategies through slow delivery of multiple numbers and keyword strategies through content variation. Furthermore, traditional spam monitoring systems rely on user complaints for keyword strategies, leading to a significant number of illegal messages being missed. Overly lax policies can easily lead to exploitation, while overly strict policies can impact users' normal communication needs. Traditional solutions face significant challenges.

[0034] The disclosed embodiment evaluates whether the characters in a text message are abnormal and whether the text is legal, and determines the illegality suspicion of the text message based on the evaluation results, thereby determining whether to intercept the text message based on the illegality suspicion. This realizes automatic recognition of variant features of text message content and keywords without the need for policy configuration, greatly reducing the complexity and workload of illegal text message recognition, and adding the recognition results of variant feature characters to the judgment of the illegality suspicion of the text message, thereby improving the recognition rate of illegal text messages.

[0035] The SMS processing method of the embodiments of the present disclosure can be applied to any electronic device that needs to receive SMS messages, such as a terminal device or a server. The terminal device may include, but is not limited to, an in-vehicle device, a user equipment (UE), a mobile device, a computing device, a wearable device, and the like. For example, it may include, but is not limited to, a cellular phone, a cordless phone, a personal digital assistant (PDA), a portable computer, and the like. The SMS processing method can be implemented by a processor calling computer-readable program instructions stored in a memory, or it can be implemented by a server.

[0036] The following is a detailed introduction to the embodiments of the present disclosure.

[0037] The present disclosure provides a method for processing short messages. Figure 1 As shown, the method includes steps S101-S106:

[0038] S101: Receive a text message to be reviewed.

[0039] S102: Detect whether each character in the text message belongs to one or more preset abnormal characteristic characters, and evaluate the detection results to obtain a first evaluation result.

[0040] S103: When the characters in the text message include abnormal characteristic characters, convert the abnormal characteristic characters into normal text.

[0041] S104: Evaluate the legitimacy of the short message after conversion into normal text to obtain a second evaluation result.

[0042] S105: Obtain the illegality suspicion degree of the text message according to the first evaluation result and the second evaluation result.

[0043] S106. Generate a review result of the text message based on the illegality suspicion level.

[0044] In the embodiment of the present disclosure, after the suspicious feature quantification model 202 determines the audit result of releasing or intercepting the text message according to the illegal suspicion level, corresponding response information can be sent according to the audit result.

[0045] In the disclosed embodiment, determining the illegal suspicion of a text message can be achieved using a preset suspicious feature quantification model 202. Before the text message to be reviewed is input into the suspicious feature quantification model 202, the text message to be reviewed can first be pre-identified by the pre-identification module 201, and the text message can be released or blocked based on the identification result.

[0046] In the embodiment of the present disclosure, Figure 2 、 Figure 3 As shown, the pre-identification module 201 includes: a protocol interface module 2011 and a distribution agent module 2012;

[0047] The pre-identification module 201 pre-identifies the received SMS and determines the review result of releasing or blocking the SMS based on the identification result, including:

[0048] The protocol interface module 2011 matches the sender number and the receiver number of the SMS message with the blacklist and whitelist numbers, and determines the review result of releasing or blocking the SMS message based on the first matching result;

[0049] The distribution agent module 2012 matches the content of the text messages that the protocol interface module 2011 cannot determine whether to release or intercept with the black and white list features, determines the review result of releasing or intercepting the text messages based on the second matching result, and distributes the text messages that cannot be determined whether to release or intercept to any one or more of the suspicious feature quantification model 202, the cluster analysis module 203 and the semantic classification model 204.

[0050] In the embodiment of the present disclosure, based on the SMS network security interface, the protocol interface module 2011 receives a real-time SMS submission request from the SMS center or other message / file sources forwarded by the convergence gateway of the existing network. Each of the sender number and the receiver number of the SMS is matched with the stored blacklist and whitelist numbers. If it matches the blacklist number, the SMS is intercepted. If it matches the whitelist number, the SMS is released. If it cannot be determined whether it matches the blacklist or whitelist number, or if it does not match the blacklist or whitelist number, it is sent to the distribution agent module 2012 for further identification. According to the principle of minimizing data security processing, only the necessary fields of the relevant SMS can be extracted, converted into internal messages, and forwarded to the distribution agent module 2012.

[0051] In the embodiment of the present disclosure, the distribution agent module 2012 matches the content of the text message with the pre-stored black and white list features. When it matches the black list features, the content of the illegal text message is added to the black feature sample library, and an audit result of intercepting the text message is generated. When it matches the white list features, an audit result of releasing the text message is generated. When it cannot be determined whether it matches the black and white list features, or when it does not match either the black and white list features, the distribution agent module 2012 distributes the text message to any one or more of the suspicious feature quantification model 202, the cluster analysis module 203 and the semantic classification model 204 for further processing of the text message.

[0052] In the embodiment of the present disclosure, the distribution agent module 2012 hot caches illegal feature strings (for example, including but not limited to illegal URLs (Uniform Resource Locators), QQ numbers, bank accounts, mobile phone numbers, etc.) in the text message content, matches them with the black and white list features after de-interference, and returns the review result of release or interception to the protocol interface module 2011. Based on the review result, the protocol interface module 2011 notifies the aggregation gateway to release or intercept the text message.

[0053] In the disclosed embodiment, the distribution agent module 2012 receives response information returned by the suspicious feature quantification model 202, the cluster analysis module 203, and the semantic classification model 204. The response information includes the audit results (such as release or interception), comprehensively evaluates the results returned by each model and module, and returns the release or interception response information in real time to the protocol interface module 2011. The distribution agent module 2012 can also control the business process scheduling of each model and module call, control the message flow of each model and module call, generate log files and send them to the data analysis module 205, and aggregate the suspicious records of each model / module to generate log files and send them to the statistical storage module 206.

[0054] In the embodiment of the present disclosure, the suspicious feature quantification model 202 calculates the illegal suspicion of the text message and generates an audit result for releasing or blocking the text message based on the calculation result.

[0055] In the embodiment of the present disclosure, calculating the illegal suspicion level of a text message includes:

[0056] Detect whether each character in the text message belongs to one or more preset abnormal characteristic characters;

[0057] Evaluate the detection result of whether the characters in the text message are abnormal characteristic characters to obtain a first evaluation result;

[0058] When the characters in the text message contain abnormal characteristic characters, the abnormal characteristic characters are converted into normal text;

[0059] Evaluate the legitimacy of the text in the text message to obtain a second evaluation result;

[0060] The illegality suspicion degree of the text message is obtained according to the first evaluation result and the second evaluation result.

[0061] In the embodiment of the present disclosure, the evaluation may be a score, and the first evaluation result and the second evaluation result are obtained by scoring.

[0062] In the embodiment of the present disclosure, the suspicious feature quantification model 202 is a model obtained by training a preset neural network based on SMS samples containing abnormal feature characters.

[0063] In the disclosed embodiment, a library of abnormal characteristic characters may be pre-set, and this library may be used to determine which characters are abnormal and which are normal. Based on the various abnormal characteristic characters in this library, a large number of SMS samples containing abnormal characteristic characters may be pre-created, and a preset neural network may be trained to obtain a suspicious feature quantification model 202.

[0064] In the embodiment of the present disclosure, the abnormal characteristic characters may include but are not limited to any one or more of the following: characters within an abnormal Unicode range, characters in an abnormal language distribution (such as a mixture of multiple languages), characters of an abnormal character type, abnormal word segmentation, uncommon characters, variant characters (for example, including but not limited to variant letters, numbers, Chinese characters, emoticons, emojis), Chinese homophonic characters, etc.

[0065] In the disclosed embodiment, the suspicious feature quantification model 102 supports real-time reasoning and identification of abnormal feature characters in text messages. It can focus on identifying suspicious abnormal feature characters such as illegal URLs and code numbers, adapt to the interference of variant feature characters such as variant numbers / letters / Chinese characters / emoji expressions, Chinese homophonic characters, and multilingual mixtures, and determine the illegal suspicion of the text message based on the recognition results.

[0066] In the embodiment of the present disclosure, the suspicious feature quantification model 202 may process SMS messages in three aspects: feature engineering, feature generalization, and suspiciousness vectorization scoring.

[0067] 1. Feature Engineering: Vectorize the text of SMS messages and count abnormal features in SMS messages based on the text vectors.

[0068] Input for feature engineering: a text message.

[0069] Feature engineering algorithm: The suspicious feature quantification model 202 evaluates the detection result of whether the characters in the text message are abnormal feature characters to obtain a first evaluation result.

[0070] The suspicious feature quantification model 202 evaluates the detection result of whether the characters in the text message are abnormal feature characters, and obtains a first evaluation result, which may include:

[0071] In the case that each character is not an abnormal characteristic character, the preset value is used as the first evaluation result.

[0072] The preset value can be set accordingly according to different application scenarios. The detailed data of the preset value is not limited here. For example, the preset value can be 0 or a very small value.

[0073] The suspicious feature quantification model 202 evaluates the detection result of whether the characters in the text message are abnormal feature characters, and obtains a first evaluation result, which may include:

[0074] If any character is an abnormal characteristic character, the proportion of the character in all characters of the text message is counted;

[0075] The statistical proportions are calculated to obtain the first evaluation result.

[0076] The statistical proportions are calculated to obtain the first evaluation results, including:

[0077] Normalize the proportion of each abnormal characteristic character in the text message to obtain a normalized value;

[0078] A first evaluation result is obtained based on the normalized value of each abnormal characteristic character and the corresponding weight value.

[0079] Among them, based on the normalized values of each abnormal feature character and the corresponding weight values, a first evaluation result is obtained, including:

[0080] When there is one abnormal feature character, multiply the normalized value corresponding to the abnormal feature character by the weight value corresponding to the abnormal feature character as the first evaluation result;

[0081] When there are multiple abnormal feature characters, sum the products of the normalized values corresponding to each abnormal feature character and the corresponding weight values as the first evaluation result.

[0082] For example, a text message is "The craftsmanship of this product is more than 200%". It is statistically determined whether each character in this text message is within the common unicode range, whether it belongs to rare characters, variant characters, special characters (such as abnormal character types, abnormal word segmentations, rare characters, etc.). After statistics, it is found that there are variant characters "産" and "手艺". Then calculate the proportion of these characters in the text message. For example, "産" accounts for 1 / 14 and "手艺" accounts for 1 / 7. The statistical values (such as the proportions 1 / 14 and 1 / 7) can be normalized. The normalization method adopted here is not limited. For example, it can include but is not limited to: min-max normalization, standard deviation normalization, decimal ratio normalization, etc. If the normalized values a and b are obtained for 1 / 14 and 1 / 7 respectively, and the weight of "産" in the text message is preset as W1 and the weight of "手艺" in the text message is preset as W2, then a×W1 + b×W2 can be calculated to obtain a value as the first evaluation result.

[0083] The suspicious feature quantization model 202 evaluates the detection result of whether the characters in the text message belong to abnormal feature characters to obtain a first evaluation result, and it can also include:

[0084] Count the number of abnormal feature characters included in all the characters of the text message;

[0085] Calculate the proportion of the number of abnormal feature characters in the total number of all characters in the text message as the first evaluation result. Output of feature engineering: the first evaluation result.

[0086] 2. Feature generalization: Generalize variant features, such as, including but not limited to: place names, names, digital strings, single words in word segmentations, pinyin of rare characters, etc. (that is, normalize the text in the text message so that all the text in the text message is converted into normal text), to achieve coverage processing of variant features such as characters with similar shapes and similar sounds, and reduce the influence of temporal and regional factors.

[0087] Input of feature generalization: A text message.

[0088] Algorithm for feature generalization: Convert abnormal feature characters into normal text.

[0089] Before scoring the legality of the text in a short message, when the characters in the short message contain abnormal feature characters, convert the abnormal feature characters into normal text.

[0090] If there are no abnormal feature characters (such as rare characters, characters with similar shapes, characters with similar pronunciations, etc.) in the short message, then all the characters contained in the short message are normal text. Based on these normal texts, the legality of the short message content can be directly evaluated; if there are abnormal feature characters in the short message, then the abnormal feature characters in the short message need to be converted into normal text first, and then based on these normal texts, the legality of the short message is evaluated. This normal text refers to the text that is commonly used and can be understood by most people (such as more than 95% of the people).

[0091] Converting abnormal feature characters into normal text may include but is not limited to:

[0092] Obtain a preset mapping table, which contains the corresponding relationship between abnormal feature characters and normal text;

[0093] Determine the normal text corresponding to the abnormal feature characters to be converted from the preset mapping table;

[0094] Convert the abnormal feature characters in the short message into the determined normal text.

[0095] This mapping table contains one or more abnormal feature characters, and the normal text corresponding to the one or more abnormal feature characters. There can be multiple such mapping tables. Different types of abnormal feature characters can respectively correspond to the corresponding mapping tables. For example, it can include but is not limited to: traditional Chinese character mapping table, character / word with similar shape mapping table, homophone / word mapping table, variant character mapping table, multi-language mapping table, graphic character mapping table, etc. A common abnormal feature character mapping table can also be established. For different abnormal feature characters, the corresponding normal text can be found in the corresponding mapping table.

[0096] For example, for the traditional Chinese character "産", the corresponding normal text "产" can be obtained from the traditional Chinese character mapping table. For the graphic character the corresponding normal text "phone" can be obtained from the graphic character mapping table. For the character "恋接", the corresponding normal text "link" can be obtained from the homophone word mapping table.

[0097] When an abnormal feature character can correspond to multiple normal texts, each normal text can be respectively replaced into the short message, and the semantics of the replaced short message can be analyzed to find the short message with the most correct or smoothest semantic expression and retain it, so as to complete the conversion of the abnormal feature characters in the short message into normal text.

[0098] Before the suspicious feature quantification model 202 evaluates the legitimacy of the text message content after conversion to normal text, the method may further include:

[0099] Remove the specified type of text in the text message to evaluate the legitimacy of the remaining text content in the text message; or

[0100] Reduce the weight of specified types of text in SMS messages to evaluate the legitimacy of all text content in SMS messages.

[0101] The specified type of text may include time, place name, name, etc., which are interference texts during the legitimacy evaluation and can be processed before the evaluation. For example, after the text of the SMS is segmented, the specified type of text in the SMS can be removed or its weight can be reduced.

[0102] Output of feature generalization: converted normal text.

[0103] 3. Suspicion Quantization Scoring: The text feature vector corresponding to the text containing legitimate text serves as the input to the computational sub-model (e.g., a pre-defined multi-layer neural network model) in the suspicious feature quantification model 202. The computational sub-model calculates the suspiciousness of the text message. The computational sub-model is trained on a multi-layer neural network model (including feature engineering parameters) using training data containing a large amount of illegitimate text message sample data and legitimate annotated text message sample data.

[0104] Input for the suspiciousness vectorization score: SMS features are generalized and converted into normal text.

[0105] Suspicion vectorization scoring algorithm: Evaluate the legitimacy of the content corresponding to normal text to obtain a second evaluation result.

[0106] The computing sub-model itself has the ability to focus on identifying suspicious features such as illegal URLs and code numbers. For example, it can compare the content of a text message (for example, a paragraph or certain words in a text message) with pre-stored illegal content, detect the similarity between the text message content and the illegal content, and score the legitimacy of the text message based on the size of the similarity. All scores of the text message content are normalized, and the normalized values are multiplied by the weights corresponding to the corresponding text message content. The sum of all products is obtained to obtain a value, which can be regarded as a legitimacy score as the second evaluation result.

[0107] Output of the quantized suspicion score: the illegal suspicion degree obtained based on the first evaluation result and the second evaluation result.

[0108] The suspicious feature quantification model 202 may obtain the illegal suspicion degree of the text message according to the first evaluation result and the second evaluation result, which may include: performing calculations on the first evaluation result and the second evaluation result to obtain the illegal suspicion degree of the text message.

[0109] The first evaluation result and the second evaluation result are calculated to obtain the illegal suspicion degree of the text message, including:

[0110] The first evaluation result and the second evaluation result are weightedly calculated to obtain the illegal suspicion degree of the text message.

[0111] For example, multiply the first evaluation result by the corresponding weight value to obtain the first numerical value. The calculation sub-model itself can give a judgment result of whether the current text message is an illegal and suspicious text message, that is, the second evaluation result. Multiply the second evaluation result by the corresponding weight value to obtain the second numerical value. Add the first numerical value and the second numerical value to give the final illegality and suspicion degree of the text message.

[0112] In the embodiment of the present disclosure, it can be pre-specified that the illegal suspicion value of normal text messages is between 0-0.1, and the illegal suspicion value of illegal text messages containing illegal URLs, code numbers, etc. is greater than 50. For text messages with illegal suspicion values between 0.1-50, they can be further input into the clustering analysis module 203 and / or the semantic classification model 204 for judgment.

[0113] In the embodiment of the present disclosure, the suspicious feature quantification model 202 can return a release response message for a normal text message, and can return an interception response message for an illegal text message, and add the detected abnormal features to the black feature sample library.

[0114] In the disclosed embodiment, the suspicious feature quantification model 202 realizes the automatic recognition of the variation features of SMS content and keywords based on the suspicious abnormal features + calculation sub-model. It can draw inferences from one example and identify illegal SMS in real time without the need for policy configuration, which greatly reduces the complexity and workload of on-site policy operation and maintenance.

[0115] In an embodiment of the present disclosure, the method may further include:

[0116] The frequency of occurrence of the text message is judged, and / or the semantics and emotions of the text message are judged, and the review result of releasing or blocking the text message is determined based on the judgment result.

[0117] In the embodiment of the present disclosure, the short message may be sent to the cluster analysis module 203, which determines the frequency of occurrence of the short message and releases or intercepts the short message according to a first determination result; and / or,

[0118] The text message is sent to the semantic classification model 204, which judges the semantics and emotions of the text message and releases or intercepts the text message based on the second judgment result.

[0119] In the embodiment of the present disclosure, the cluster analysis module 203 determines the frequency of occurrence of short messages, which may include:

[0120] Divide the text of the SMS into words and extract keywords;

[0121] Keywords constitute the main content of the text message;

[0122] The frequency of occurrence of the main content is counted as the frequency of occurrence of the text message.

[0123] In the embodiment of the present disclosure, the clustering analysis module 203 can use text segmentation technology and feature extraction technology to construct the main content of the text message (also known as the main feature string), and cluster text messages with similar content on the entire network based on the main content, so as to determine the frequency of occurrence of the current text message and review the text message based on the frequency of occurrence.

[0124] In the embodiment of the present disclosure, the text messages may first be preprocessed, for example, including but not limited to: removing stop words from the text message content, normalizing the variation features, and then performing text segmentation, extracting core keywords from the text segmentation, constructing the main content from the core keywords, counting the frequency of occurrence of the main content, and thus calculating the frequency of occurrence of the text message. Text messages with a higher frequency of occurrence (for example, a frequency of occurrence greater than or equal to a preset frequency threshold) may be treated as suspected illegal text messages.

[0125] In the embodiment of the present disclosure, the main purpose of the cluster analysis performed by the cluster analysis module 203 is not to extract core keywords, because some fraudulent text messages do not contain obvious keywords. The idea and main purpose of the cluster analysis is to extract features from each text message, calculate the distance between the feature and the feature of the text message that has been received (such as calculating the Hamming distance), that is, determine the similarity between the current text message content and the text message content that has been received based on the distance. When the similarity exceeds the preset similarity threshold, the frequency of occurrence of the text message can be increased by 1.

[0126] In the embodiment of the present disclosure, for example, an embodiment of cluster analysis performed by the cluster analysis module 203 is given below:

[0127] 1. Divide the text of the SMS into words and obtain an N-dimensional feature vector (N is a positive integer);

[0128] 2. Set a weight (TF-IDF) for each word. For example, the weight of "invoice" is 5.

[0129] 3. Calculate the hash value for the feature vector. For example, the hash value of "invoice" is 110101.

[0130] 4. Weight and accumulate the feature vectors of all words, that is, W = Hash × weight, where W is the feature value corresponding to any word, Hash is the hash value corresponding to the word, and weight is the weight value corresponding to the word. In the calculation of W = Hash × weight, if a 1 is encountered in the hash value, the hash value and the weight value are multiplied positively, and if a 0 is encountered in the hash value, the hash value and the weight value are multiplied negatively. For example, the hash value "110101" of "invoice" is weighted to obtain: W(invoice) = 110101 × 5 = 5 5 -55 -5 5;

[0131] ⑤ Perform dimensionality reduction on the feature words of all words: merge and accumulate the feature values corresponding to all words in the text message by column, and set the result of each column to 1 if it is greater than 0 and to 0 if it is less than 0;

[0132] ⑥ Use the result obtained after dimensionality reduction as the feature of the text message, for example, 110101.

[0133] In the embodiment of the present disclosure, the cluster analysis module 203 may also determine the social normality of the text message.

[0134] In the embodiment of the present disclosure, the cluster analysis module 203 determines the social normality of the text message, which may include:

[0135] Determining any one or more of the following information about the text message: the incoming and outgoing text message volume of the sender number, the discreteness of the recipient number, and the frequency of interaction between the sender number and the recipient number;

[0136] The social normality of the text message is judged based on the determined information.

[0137] In the embodiment of the present disclosure, the cluster analysis module 203 can analyze the sender number input and output degree (or the SMS input and output volume of the sender number) of SMS based on SNS (social networking site) social analysis. Since the number of SMS messages sent and received by a normal user is fixed within a range, the number is usually not large; if it is identified that a certain number has sent a large number of SMS messages but received a small number of SMS messages, the sender number can be regarded as a suspected illegal user.

[0138] In the embodiment of the present disclosure, the cluster analysis module 203 can perform a dispersion analysis of the recipient numbers of text messages based on SNS (social networking site) social analysis. Since illegal users usually send a large number of text messages to different users, and normal users usually focus on a few friends when sending multiple text messages, and the number is not particularly large, therefore, if a certain number is identified as sending the same text message to many different users, the sending number can be regarded as a suspected illegal user.

[0139] In the embodiment of the present disclosure, the cluster analysis module 203 can perform interaction frequency analysis on the sender and receiver of text messages based on SNS (social networking site) social analysis. This is because normal users generally interact with their friends who receive text messages, while illegal users generally have no connection with the receiving users. If it is identified that there has been no previous interaction between the sender of the text message and the receiver of the text message, the sender's number can be regarded as a suspected illegal user.

[0140] In the embodiment of the present disclosure, the cluster analysis module 203 may also perform a dialing behavior analysis on the sender of the short message.

[0141] In the embodiment of the present disclosure, the cluster analysis module 203 mines and analyzes that the sender number has a large out-degree and a small in-degree, the recipient dispersion is high (a large number of short messages are sent to different recipients), there is no short message interaction between the sender and the recipient, or the sender user has dialing behavior, then it can be determined as an illegal user.

[0142] In the disclosed embodiment, cluster analysis module 203 addresses the problem of criminals exploiting the "unreasonable" approach of circumventing traditional regulatory systems. It can identify criminals sending low-frequency, batch-based messages from a large number of different numbers. Furthermore, this function can be applied to manual review, using the 80 / 20 principle to quickly and efficiently identify illegal short messages with the highest business impact.

[0143] In the embodiment of the present disclosure, the semantics and emotions of a text message may be determined using the semantic classification model 104 . The semantic classification model 204 is obtained by adjusting a pre-trained language model using illegal short message samples.

[0144] In the disclosed embodiment, the semantic classification model 204 is based on a pre-trained large language model (or pre-trained language model) at the bottom layer, and the upper layer is based on SMS domain model training, focusing on identifying negative features such as semantic content and emotional intent.

[0145] In the embodiment of the present disclosure, the semantic classification model 204 is implemented in a B+Z manner, where B represents Base, i.e., the basic pre-trained large language model, and Z represents the illegal SMS recognition adaptation model. The basic pre-trained large language model B has tens of billions of parameters and has strong natural language semantic analysis, emotion recognition, logical reasoning, and induction capabilities. On this basis, the illegal SMS samples of Z are superimposed, and the pre-trained large language model B is trained to achieve fine-tuning of the pre-trained large language model B and obtain the semantic classification model 204. The semantic classification model 204 supports the key identification of illegal SMS with features of bad semantic content (such as: loan collection, pornography, gangs, gambling, fraud, etc.) and bad emotional intentions (such as: intimidation, deception, temptation, inducement, etc.), and returns a release or interception response to the distribution agent module 2012 in real time.

[0146] In an embodiment of the present disclosure, a response message of release can be returned for text messages that do not contain negative semantic content features and negative emotional intention features, and a response message of interception can be returned for text messages that contain negative semantic content features and / or negative emotional intention features, and the detected negative semantic content features and / or negative emotional intention features can be entered into a black feature sample library.

[0147] In an embodiment of the present disclosure, the method may further include:

[0148] Inputting the text message into the multidimensional analysis model 205;

[0149] The multidimensional analysis model 205 performs at least one question-answer interaction with a preset natural language model based on the preset multidimensional prompt words for the text message, and releases or intercepts the text message according to the interaction result.

[0150] In the embodiment of the present disclosure, the short message may be input into the multidimensional analysis model 205 if a preset condition is met. In the embodiment of the present disclosure, the preset condition may include but is not limited to:

[0151] It is impossible to determine whether to release or intercept the SMS based on the suspicion of illegality, or,

[0152] In any multiple review results obtained, the same SMS message may be released or blocked in different ways.

[0153] In the embodiment of the present disclosure, for example, any multiple review results determined by the suspicious feature quantification model 202, the cluster analysis module 203, and the semantic classification model 204 may result in different processing results for releasing or blocking the same text message.

[0154] In the embodiment of the present disclosure, the any plurality may refer to any two or all three.

[0155] In the embodiment of the present disclosure, when any multiple of the review results determined by the suspicious feature quantification model 202, the cluster analysis module 203 and the semantic classification model 204 identify and judge the same text message, the distribution agent module 2012 will receive the review results of the text message by any multiple models or modules. If the distribution agent module 2012 detects that any multiple of the review results of the suspicious feature quantification model 202, the cluster analysis module 203 and the semantic classification model 204 for the same text message are different, the text message can be sent to the multidimensional analysis model 205, which will further identify and judge the text message to achieve automatic review of the text message.

[0156] In the embodiment of the present disclosure, the multidimensional analysis model 205 can pre-customize the prompt words of the multidimensional prompt in the field of SMS anti-fraud. The multidimensional prompt can perform at least one round of interaction with a preset natural language model based on these prompt words (the natural language model pre-customizes the prompt words for the SMS field so that the natural language model eliminates the illusion problem). In some scenarios, multiple rounds of interaction are even required in a step-by-step manner (similar to the question-and-answer interaction of ChatGPT). The multidimensional analysis model 205 comprehensively scores and determines the returned results, which can replace manual review and automatically review the SMS identified by each model and module, greatly reducing operation and maintenance manpower and costs.

[0157] In the embodiment of the present disclosure, the prompt word can be set accordingly according to different SMS content scenarios, and there is no limitation on the detailed prompt word.

[0158] In the embodiment of the present disclosure, the multidimensional analysis model 205 can directly return a release response message to the distribution agent module 2012 for a normal text message, and can return an interception response message to the distribution agent module 2012 for an illegal text message, and add the illegal content features identified in the judgment process to the black feature sample library.

[0159] In the embodiment of the present disclosure, Figure 4 As shown, the multidimensional analysis model 205 can also be directly connected to the cluster analysis module 203. The multidimensional analysis model 205 generates a file. The multidimensional analysis model 205 can judge and process the text messages based on the occurrence of the same content recorded in the file to realize the review of the text messages.

[0160] In the embodiment of the present disclosure, Figure 5 As shown, when the suspicious feature quantification model 202, the cluster analysis module 203 and the semantic classification model 204 are used to process the text message together, the text message processing method may further include steps S501-S503:

[0161] S501, the suspicious feature quantification model 202 calculates the illegal suspicion of the text message and determines the review result of releasing or blocking the text message based on the calculation result;

[0162] S502, the cluster analysis module 203 determines the frequency of occurrence of text messages for which the suspicious feature quantification model 202 cannot determine whether to release or intercept, and determines the audit result of whether to release or intercept the text messages based on the judgment result;

[0163] S503, the semantic classification model 204 judges the semantics and / or emotions of the text messages that the cluster analysis module 203 cannot determine whether to release or intercept, and releases or intercepts the text messages based on the audit result of the judgment result.

[0164] In the embodiment of the present disclosure, the suspicious feature quantification model 202, the cluster analysis module 203, and the semantic classification model 204 can be combined as needed to realize a combined processing method. The combined processing method can be a combination of any two or three, and can be combined with the multidimensional analysis model 205; the execution order of the combined models and modules can be defined according to needs and is not limited here. The following is an embodiment of the combination of the suspicious feature quantification model 202, the cluster analysis module 203, the semantic classification model 204, and the multidimensional analysis model 205.

[0165] In the disclosed embodiment, the suspicious feature quantification model 202 has a high processing performance and can perform the first round of identification and judgment on text messages. For text messages that cannot be identified and judged, they are sent to the cluster analysis module 203 for frequency clustering of similar content and social analysis for further identification and judgment. For text messages that the cluster analysis module 203 cannot identify and judge, the text message (or the similar content of the text message clustered into a text message) can be sent to the semantic classification model 204 (with low processing performance) for further identification and judgment. Finally, the multidimensional analysis model 205 performs multidimensional analysis instead of manual review, and automatically reviews the text messages identified by each model and module. This embodiment fully considers the model processing performance and functional characteristics, and the division of labor and cooperation among each model and module, the coordination of processes, and the optimal combination of mutual complementation to accurately and comprehensively identify and handle illegal text messages.

[0166] In the embodiment of the present disclosure, Figure 3 、 Figure 4 As shown, a statistics storage module 206 can also be provided. The statistics storage module 206 receives and analyzes the records of suspected illegal SMS messages collected by the distribution agent module 2012 from various models / modules, generates log files, and stores the records. Based on the above-mentioned log files, real-time statistics can be generated and multi-dimensional statistical results can be stored in the database.

[0167] In the embodiment of the present disclosure, Figure 3 、 Figure 4 As shown, a management and maintenance module 207 may also be provided, and the management and maintenance module 207 may perform any one or more of the following operations:

[0168] 1. Business log query: Supports conditional query on the log record interface of suspected illegal SMS messages identified and handled by each model and module, which is used to handle user complaints and post-audit verification.

[0169] 2. DashBoard statistical report: Statistics present the business status of each process identification and disposal of each model and module, model effect evaluation, etc.

[0170] 3. Manual review: supports the frequency of occurrence of similar content clustering and drill-down detailed review, which facilitates the auditor to audit and correct the suspected SMS results identified and handled by various models and modules of the system; among them, the review results of illegal / normal SMS content are sent to the distribution agent module 2012 and added to the content feature black / white sample library (i.e. black feature sample library / white feature sample library) 208.

[0171] 4. Sample management: Manual review results, data import from third-party (such as public security, information security, etc.) systems, and addition to the sample library 208 after data cleaning.

[0172] 5. Model training: supports specifying batch sample data, specifying model training, and effect evaluation.

[0173] In the disclosed embodiment, sample library 208 can store black / white text messages that have been manually reviewed and verified, or imported from a third-party system (such as the Public Security Information Security Center), after data cleaning, for real-time feature matching and model training. It also stores samples of illegal URLs, QQ numbers, bank account numbers, mobile phone numbers, etc., which can be used for real-time feature matching.

[0174] In the embodiment of the present disclosure, Figure 6 As shown, the following is an example of SMS processing based on all the above models and modules, including steps S601-S620:

[0175] S601. Number hits whitelist: After the SMS message flows into the system of the embodiment of the present disclosure, the protocol interface module 2011 performs whitelist matching on the sender number and the receiver number of the SMS message, and directly releases the message if a match is found.

[0176] S602, the number hits the blacklist: the protocol interface module 2011 matches the sender number and the receiver number of the SMS message to the blacklist, and directly intercepts the number if it hits the blacklist.

[0177] S603 , the number does not hit the blacklist / whitelist: the SMS message corresponding to the number is sent to the distribution agent module 2012 .

[0178] S604 , content hits whitelist feature: the distribution agent module 2012 matches the SMS content with the whitelist features in its own hot cache in real time. If the similarity exceeds the preset similarity threshold, a release response is returned to the protocol interface module 2011 .

[0179] S605 , content hits blacklist feature: the distribution agent module 2012 matches the SMS content with the blacklist feature in its own hot cache in real time. If the similarity exceeds the preset similarity threshold, an interception response is returned to the protocol interface module 2011 .

[0180] S606 , the content does not hit the black / white list feature: the SMS message corresponding to the content is sent to the suspicious feature quantification model 202 .

[0181] S607 , low suspicion: the suspicious feature quantification model 202 gives an illegal suspicion score. For SMS messages with an illegal suspicion score lower than a preset first score threshold, a release response is returned to the protocol interface module 2011 via the distribution agent module 2012 .

[0182] S608 , high suspicion: the suspicious feature quantification model 202 gives an illegal suspicion score. For SMS messages with an illegal suspicion score higher than a preset second score threshold, an interception response is returned to the protocol interface module 2011 via the distribution agent module 2012 .

[0183] S609 , medium suspicion: the suspicious feature quantification model 202 gives an illegal suspicion score, and sends SMS messages with a medium illegal suspicion score (short messages higher than or equal to the first score threshold, lower than or equal to the second score threshold) to the cluster analysis module 203 .

[0184] S610 , clustering high frequency: the cluster analysis module 203 counts the text messages with similar content and a frequency higher than a preset frequency threshold and abnormal social behaviors, and returns an interception response to the protocol interface module 2011 through the distribution agent module 2012 .

[0185] S611 , clustering medium and low frequencies: the cluster analysis module 203 counts the text messages with similar text message contents and whose occurrence frequencies are lower than a preset frequency threshold, and sends them to the semantic classification model 204 .

[0186] S612. Normal content and emotion: The semantic classification model 204 focuses on identifying the negative features of the semantic content and emotional intention of the text message, identifies the text messages with normal content and emotion, and returns a release response to the protocol interface module 2011 through the distribution agent module 2012.

[0187] S613, illegal content and emotion: The semantic classification model 204 focuses on identifying the negative features of the semantic content and emotional intention of the text message, identifies illegal text messages with abnormal content and / or emotion, and returns an interception response to the protocol interface module 2011 through the distribution agent module 2012.

[0188] S614 , adding to the black and white list feature sample library: Identify illegal text messages with abnormal content and / or emotions, and add the text messages to the black and white list feature sample library of the sample library 108 .

[0189] S615, conditional triggering: According to the triggering conditions configured by the management and maintenance module 207 (for example, the identification results of the suspicious feature quantification model 202 and the semantic classification model 204 are inconsistent, etc.), the SMS is sent to the multi-dimensional analysis module 205, which uses the customized multi-dimensional prompt words in the SMS anti-fraud field and interacts with the natural language model step by step and through multiple rounds of question and answer to comprehensively score the returned results.

[0190] S616. If the multi-dimensional score is lower than the preset score threshold, the sample is added to the whitelist feature library: After multiple rounds of prompt word interaction, if the comprehensive score returned by the natural language model is lower than the score threshold, the sample is added to the whitelist feature library.

[0191] S617. If the multi-dimensional score is higher than the preset score threshold, it will be added to the blacklist feature sample library: After multiple rounds of prompt word interaction, if the comprehensive score returned by the natural language model is higher than the score threshold, it will be added to the blacklist feature sample library.

[0192] S618. Each process processes the storage record: The statistical storage module 206 collects the records of suspected illegal SMS messages of each model / module, generates a log file, and stores the records.

[0193] S619, manual review and addition to the whitelist feature sample library: The operation and maintenance personnel use the management and maintenance module 207 to audit the above-mentioned storage records. If it is found that the system has misjudged normal SMS as illegal SMS, the SMS will be added to the whitelist feature sample library after review and correction.

[0194] S620, manual review and blacklist feature sample library: Operation and maintenance personnel use the management and maintenance module 207 to audit the above-mentioned storage records. If it is found that the system has misjudged illegal SMS as normal SMS, the SMS will be added to the blacklist feature sample library after review and correction.

[0195] The present disclosure also provides an electronic device 700, such as Figure 7 As shown, the electronic device 700 includes:

[0196] One or more processors 701;

[0197] a memory 702 storing one or more programs, which, when executed by the one or more processors 701, enable the one or more processors 701 to implement the method for processing text messages;

[0198] One or more input / output I / O interfaces 703 are connected between the processor 701 and the memory 702 and are configured to implement information exchange between the processor 701 and the memory 702 .

[0199] Among them, the processor 701 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 702 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 703 is connected between the processor 701 and the memory 702, and can realize information exchange between the processor 701 and the memory 702, including but not limited to a data bus (Bus), etc.

[0200] In some embodiments, the processor 701 , the memory 702 , and the I / O interface 703 are connected to each other via a bus 704 , and further connected to other components of the computing device.

[0201] The embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the text message processing method is implemented.

[0202] Those skilled in the art will appreciate that all or some of the functional modules / units disclosed above may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0203] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be performed by several physical components in cooperation.

[0204] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; compact disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cassettes, tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0205] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for processing short messages, characterized in that: The method comprises: Receive SMS messages awaiting review; Detecting whether each character in the text message belongs to one or more preset abnormal characteristic characters, and evaluating the detection results to obtain a first evaluation result; If the characters in the text message contain abnormal characteristic characters, convert the abnormal characteristic characters into normal text; Evaluate the legitimacy of the text message content after conversion into normal text to obtain a second evaluation result; Determining the suspicious degree of illegality of the text message according to the first evaluation result and the second evaluation result; The review result of the text message is generated according to the illegal suspicion level.

2. The method for processing short messages according to claim 1, wherein: The detecting whether each character in the text message belongs to one or more preset abnormal characteristic characters, and evaluating the detection result to obtain a first evaluation result, includes: When any character belongs to the abnormal characteristic character, the proportion of the character in all characters of the text message is counted; The statistical proportions are calculated to obtain the first evaluation result.

3. The method for processing short messages according to claim 2, wherein: The performing operation on the statistical proportion to obtain the first evaluation result includes: Normalize the proportion of each abnormal characteristic character in the text message to obtain a normalized value; The first evaluation result is obtained based on the normalized numerical value of each abnormal characteristic character and the corresponding weight value.

4. The method for processing short messages according to any one of claims 1 to 3, wherein: Before evaluating the legitimacy of the text message content after conversion into normal text, the method further includes: Removing the specified type of text from the text message to evaluate the legitimacy of the remaining text content in the text message; or The weight of the specified type of text in the text message is reduced to evaluate the legitimacy of all text contents in the text message.

5. The method for processing short messages according to any one of claims 1 to 3, wherein: The converting the abnormal characteristic characters into normal text includes: Obtaining a preset mapping table, wherein the mapping table contains a correspondence between abnormal characteristic characters and normal text; Determine the normal text corresponding to the abnormal characteristic character to be converted from a preset mapping table; The abnormal characteristic characters in the text message are converted into determined normal text.

6. The method for processing short messages according to any one of claims 1 to 3, wherein: The method further comprises: The frequency of occurrence of the text message is judged, and / or the semantics and emotions of the text message are judged, and the review result of releasing or intercepting the text message is determined based on the judgment result.

7. The method for processing short messages according to claim 6, wherein: The determining of the frequency of occurrence of the text message includes: Dividing the text of the SMS into words and extracting keywords; The keywords constitute the main content of the text message; The occurrence frequency of the main content is counted as the occurrence frequency of the short message.

8. The method for processing short messages according to claim 6, wherein: The method further comprises: Based on the preset multi-dimensional prompt words, at least one question-answer interaction is performed on the text message to the preset natural language model, and the review result of releasing or blocking the text message is determined according to the interaction result.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A memory having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the SMS processing method according to any one of claims 1 to 8; One or more input / output (I / O) interfaces are connected between the processor and the memory and configured to implement information interaction between the processor and the memory.

10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the text message processing method according to any one of claims 1 to 8.