SMS processing methods and related equipment
By generating a basic character set and evaluating weights, and combining noise word filtering and weight thresholds to identify the SMS messages to be identified, the problem of high false interception rate of black and gray market SMS messages in existing technologies is solved, and illegal SMS messages are accurately intercepted while ensuring timely delivery.
Patent Information
- Application Number
- CN202310466770.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-04-24
AI Technical Summary
Existing technologies have a high false interception rate when identifying and blocking black and gray market text messages, and cannot accurately identify variant text message content such as variant characters, traditional characters, and emoji symbols.
By acquiring historical SMS messages, counting character frequencies to generate a basic character set, and evaluating the probability of characters in allowing and blocking SMS messages based on weights, the SMS messages to be identified are identified using weights, and a target character set is generated for identification by combining noise word filtering and weight thresholds.
It improves the accuracy of SMS blocking, reduces false blocking, and ensures accurate blocking of illegal SMS messages while guaranteeing timely delivery.
Smart Images

Figure CN116614817B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to SMS processing methods and related equipment. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. It should not be construed as an admission that the description herein is prior art.
[0003] SMS messages are text or digital messages sent or received by users through their devices, characterized by high timeliness, strong reach, and non-recallability. Many businesses use SMS to fulfill service requests such as login verification and customer acquisition promotions, ensuring timely and effective processing of these requests. However, malicious actors who threaten internet security also exploit SMS for illegal activities. To counter this, they often use graphic and text messages that deviate from standard SMS formats to circumvent risk detection systems, aiming to send illegal messages to recipients and induce them to click on websites or add phone numbers or other communication accounts. These altered graphic and text messages, such as those using variant characters, traditional characters, internet slang, or emojis, can circumvent keyword detection, allowing the messages to be successfully sent. Therefore, accurately intercepting malicious SMS messages while ensuring timely delivery is a major challenge for those providing commercial SMS services.
[0004] Currently, risk detection providers generally use a basic character set consisting of English characters, common punctuation marks, and common Chinese characters to identify and block black and gray market SMS messages. However, this basic character set does not take into account the actual usage of these characters in SMS services. In other words, the selection of this basic character set is detached from specific SMS services, resulting in a high false blocking rate when using this basic character set to block black and gray market SMS messages. Summary of the Invention
[0005] The SMS processing method and related equipment provided in this invention at least solve the problem of high false interception rate when intercepting SMS messages from black and gray industries in the prior art.
[0006] According to one aspect of the present invention, a method for processing text messages is provided, comprising:
[0007] Retrieve historical SMS messages, including historical allowed SMS messages and historical blocked SMS messages;
[0008] The characters in the historical SMS messages are statistically analyzed to obtain the first total frequency of occurrence of the characters in the historical released SMS messages within a predetermined time period, and the second total frequency of occurrence of the characters in the historical blocked SMS messages within a predetermined time period. Based on the characters in the historical SMS messages, a basic character set is generated.
[0009] Based on the first total occurrence frequency and the second total occurrence frequency, the weight of each character in the basic character set is determined, and the weight is used to reflect the probability of the character appearing in the SMS message that needs to be blocked;
[0010] The basic character set is used to identify the SMS message to be identified based on the weight, in order to determine whether the SMS message to be identified needs to be blocked.
[0011] In some embodiments, the step of determining the weight of each character in the basic character set based on the first total occurrence frequency and the second total occurrence frequency includes:
[0012] Based on the first total occurrence frequency, a first total number of characters belonging to the historical released SMS messages in the basic character set is determined, and based on the second total occurrence frequency, a second total number of characters belonging to the historical blocked SMS messages in the basic character set is determined;
[0013] Calculate the ratio of the first total number of characters to the second total number of characters to obtain the weighting coefficient;
[0014] The weight of the character is determined based on the first total frequency of occurrence, the second total frequency of occurrence, and the weight coefficient.
[0015] In some embodiments, the step of determining the weight of the character based on the first total frequency of occurrence, the second total frequency of occurrence, and the weight coefficient includes:
[0016] The first total frequency of occurrence is subtracted from the product of the second total frequency of occurrence and the weight coefficient to obtain the initial weight value. The initial weight value is then normalized to obtain the weight.
[0017] In some embodiments, before recognizing the SMS message to be identified based on the weights using the basic character set, the method further includes:
[0018] The SMS message to be identified is preprocessed using predetermined noise words. During preprocessing, if the noise words are present in the SMS message to be identified, the noise words are removed from the SMS message to be identified. The noise words are words that would be intercepted in the SMS message to be identified when the SMS message to be identified is identified using the basic character set based on the weight.
[0019] In some embodiments, the step of identifying the SMS message to be identified based on the weight using the basic character set includes:
[0020] Based on the weight, a first weight threshold is determined, wherein the probability of a character with a weight greater than the first weight threshold appearing in a text message that needs to be blocked is less than the probability of a character with a weight less than the first weight threshold appearing in a text message that needs to be blocked.
[0021] A target character set is generated based on the first weight threshold and the basic character set, so as to identify the SMS message to be identified using the target character set.
[0022] In some embodiments, the step of identifying the SMS message to be identified using the target character set includes:
[0023] Based on the number of characters in the target character set that appear in the text message to be identified, it is determined whether the text message to be identified needs to be intercepted.
[0024] In some embodiments, before generating the target character set based on the first weight threshold and the base character set, the method further includes:
[0025] Determine a first candidate value for the first weight threshold, and determine a candidate character set for the target character set based on the first candidate value and the basic character set;
[0026] The candidate character set is used to identify the historical SMS messages in order to determine a first proportion of the historical SMS messages that were blocked using the candidate character set.
[0027] Determine whether the first proportion value reaches the preset first proportion threshold. If so, use the first candidate value as the value of the first weight threshold.
[0028] In some embodiments, the step of identifying the historical SMS messages using the candidate character set includes:
[0029] Based on the number of characters appearing in the candidate character set from the characters in the historically released SMS messages and the historically blocked SMS messages, it is determined whether the historically released SMS messages and / or the historically blocked SMS messages need to be blocked.
[0030] In some embodiments, the step of identifying the SMS message to be identified based on the weight using the basic character set further includes:
[0031] The weight of each character to be identified in the SMS message to be identified is determined. The weight of the character to be identified is obtained by assigning the weight of the character that is the same as the character to be identified in the basic character set to the character to be identified.
[0032] Based on the weight of each of the characters to be identified, the average weight of the characters to be identified is determined, wherein the average weight is obtained by averaging the weights of each of the characters to be identified.
[0033] If the average weight is less than a predetermined second weight threshold, the SMS message to be identified is intercepted.
[0034] In some embodiments, before determining whether the average weight is less than a predetermined second weight threshold, the method further includes:
[0035] A second candidate value for the second weight threshold is determined, and the weights of the characters in each historical released SMS message are added together and averaged to obtain a first historical average weight, and the weights of the characters in each historical blocked SMS message are added together and averaged to obtain a second historical average weight.
[0036] Determine whether the first historical average weight and the second historical average weight are less than the second candidate value. If so, intercept the historical released SMS and the historical blocked SMS, and obtain the total number of the intercepted historical released SMS and the historical blocked SMS, so as to determine the second proportion value of the intercepted historical blocked SMS based on the total number of SMS.
[0037] Determine whether the second proportion value reaches the preset second proportion threshold. If so, use the second candidate value as the value of the second weight threshold.
[0038] In some embodiments, before counting the characters in the historical text messages, the method further includes:
[0039] The historical text messages are classified according to different text message types, and the characters of the historical text messages belonging to the same text message type are counted, and the basic character set is generated based on the characters of the historical text messages.
[0040] According to another aspect of the present invention, an electronic device is also provided, comprising: a processor, and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method steps described above.
[0041] According to another aspect of the invention, a non-transitory machine-readable medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform the above-described method steps.
[0042] Beneficial effects of the embodiments of the present invention:
[0043] In this embodiment of the invention, the basic character set is obtained from historical SMS messages. The selection of this basic character set can be dynamically adjusted according to the actual service. By evaluating the frequency of occurrence of each character in the basic character set in the allowed and blocked SMS messages, the weight of each character is evaluated. When using the basic character set to identify the SMS message to be identified, the probability of the corresponding character in the SMS message to be identified appearing in the SMS message that needs to be blocked can be determined based on the weight of the character. This allows for a high degree of accuracy in blocking, thereby accurately blocking the SMS messages that need to be blocked while ensuring the timely delivery of SMS messages.
[0044] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating a text message processing method according to an embodiment of the present invention.
[0047] Figure 2 This is a schematic diagram illustrating the classification and processing of historical SMS messages according to an embodiment of the present invention;
[0048] Figure 3 This is a schematic diagram of the character weight calculation process provided in an embodiment of the present invention;
[0049] Figure 4 This is a schematic diagram of the process for generating a target character set according to an embodiment of the present invention;
[0050] Figure 5This is a schematic diagram illustrating the process of identifying SMS messages using different identification strategies according to an embodiment of the present invention;
[0051] Figure 6 A schematic diagram illustrating the adjustment process of the first weight threshold provided in an embodiment of the present invention;
[0052] Figure 7 This is a schematic diagram of a process for identifying a text message by calculating its average weight, according to an embodiment of the present invention.
[0053] Figure 8 This is a schematic diagram of the adjustment process of the second weight threshold provided in an embodiment of the present invention. Detailed Implementation
[0054] Embodiments of this embodiment will now be described in more detail with reference to the accompanying drawings. While some embodiments of this embodiment are shown in the drawings, it should be understood that this embodiment can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this embodiment. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this embodiment.
[0055] When using a basic character set comprised of English characters, common punctuation marks, and common Chinese characters to intercept malicious SMS messages, the selection of this basic character set is detached from the specific SMS service, resulting in a high false interception rate. To address this issue, such as... Figure 1 As shown, the first embodiment of the present invention provides a text message processing method, which includes the following steps:
[0056] Step S11: Obtain historical SMS messages, including historical released SMS messages and historical blocked SMS messages. It should be noted that historical released SMS messages in this embodiment are those actually allowed to be sent to the recipient, commonly referred to as legitimate SMS messages by those skilled in the art. Even if the corresponding historical SMS message was mistakenly blocked during historical transmission, this embodiment treats it as a released SMS message allowed to be sent to the recipient. Historical blocked SMS messages in this embodiment are those that are not allowed to be sent to the recipient and are essentially blocked, commonly referred to as illegal SMS messages by those skilled in the art. Even if the corresponding historical SMS message was mistakenly sent to the recipient during historical transmission, this embodiment treats it as a blocked SMS message not allowed to be sent to the recipient. Step S12: Statistically analyze the characters in the historical SMS messages to obtain the first total frequency of occurrence of the characters in historical released SMS messages within a predetermined time period, and the second total frequency of occurrence of the characters in historical blocked SMS messages within a predetermined time period. Based on the characters in the historical SMS messages, generate a basic character set. Therefore, the selection of the basic character set provided in this embodiment of the invention is based on the actual SMS service, and is not a fixed character set that does not conform to the usage rules of relevant characters in the actual SMS service. The characters in the basic character set will change with the actual SMS service, which lays a solid foundation for improving the interception accuracy of SMS messages that need to be blocked.
[0057] Step S13: Determine the weight of each character in the basic character set based on the first total frequency of occurrence and the second total frequency of occurrence. The weight is used to reflect the probability of the character appearing in the SMS message that needs to be intercepted.
[0058] Step S14: Use a basic character set to identify the SMS message to be identified based on weights to determine whether it needs to be blocked. Specifically, if the basic character set is obtained based on historical logistics notification SMS messages, then the basic character set can be used to identify logistics notification SMS messages to be identified, thereby determining whether the SMS message to be identified needs to be blocked, with a high blocking accuracy.
[0059] As can be seen from the above, the basic character set in this embodiment of the invention is obtained based on historical SMS messages of the corresponding type. The selection of this basic character set can be dynamically adjusted according to the actual service. By evaluating the frequency of occurrence of each character in the basic character set in the released and blocked SMS messages, the weight of the characters is evaluated. When using the basic character set to identify the SMS message to be identified, the probability of the corresponding character in the SMS message to be identified appearing in the SMS message that needs to be blocked can be determined based on the weight of the characters. This allows for a determination of whether the SMS message to be identified needs to be blocked, resulting in a high interception accuracy. Thus, while ensuring the timely delivery of SMS messages, the SMS messages that need to be blocked can be accurately blocked.
[0060] To ensure that the obtained basic character set is more closely aligned with specific SMS services, thereby improving the accuracy of intercepting SMS messages that need to be blocked, in step S12, before statistically analyzing the characters present in historical SMS messages, the method provided in this embodiment of the invention further includes: classifying historical SMS messages according to different SMS types, statistically analyzing the characters present in historical SMS messages belonging to the same SMS type, and generating a basic character set based on the characters present in historical SMS messages. For example... Figure 2 As shown, the SMS types in this embodiment of the invention include e-commerce marketing, e-commerce notification, logistics marketing, logistics notification, internet finance marketing, internet finance notification, social marketing, and social notification. Among them, internet finance refers to internet finance, an emerging financial system that relies on internet tools such as payment, cloud computing, social networks, search engines, and applications to provide services such as capital financing, payment, and information intermediation. Because different types of SMS messages have different industry attributes, the characters used in actual service will differ. For example, the characters used most frequently in historical e-commerce marketing SMS messages are quite different from those used in historical logistics notification SMS messages. Therefore, this embodiment of the invention classifies historical SMS messages to obtain a basic character set for identifying SMS messages of the corresponding SMS type, thereby improving the interception accuracy. For example, when using the basic character set obtained from historical e-commerce marketing SMS messages to identify e-commerce marketing SMS messages, the identification accuracy is higher. Therefore, the selection of the basic character set provided in this embodiment of the invention is based on the actual SMS service of the corresponding SMS type. It is a character set adjusted according to the usage rules of relevant characters in the actual SMS service; that is, the characters in this basic character set will change with different SMS messages. For example, e-commerce marketing SMS messages will correspond to a basic character set, internet finance notification SMS messages will correspond to a basic character set, social marketing SMS messages will correspond to a basic character set, and so on. Secondly, since the content of verification code SMS messages is basically uniform, verification code SMS messages will also correspond to a basic character set. When identifying the SMS message to be identified, after determining which SMS message type it belongs to, the basic character set obtained from historical SMS messages of the corresponding SMS type is used to identify the SMS message to be identified, resulting in a low false identification rate and a high interception accuracy rate.
[0061] like Figure 3 As shown, step S13, which involves determining the weight of each character in the basic character set based on the first total frequency of occurrence and the second total frequency of occurrence, includes:
[0062] Step S21: Determine the total number of characters of the first type in the basic character set that belong to historical released text messages according to the first total occurrence frequency, and determine the total number of characters of the second type in the basic character set that belong to historical intercepted text messages according to the second total occurrence frequency. For example, characters such as "您 (you)", "买 (buy)", "的 (of)", "期 (period)"... in the basic character set all belong to historical released text messages. The first total occurrence frequency of "您" is n1, the first total occurrence frequency of "买" is n2, the first total occurrence frequency of "的" is n3, and the first total occurrence frequency of "期" is n4. Then the total number of characters of the first type Nl is the sum of n1, n2, n3, and n4. Similarly, calculate the total number of characters of the second type Ni according to the second total occurrence frequencies of each character in the basic character set that belong to historical intercepted text messages.
[0063] Step S22: Calculate the ratio of the total number of characters of the first type to the total number of characters of the second type to obtain the weight coefficient. That is, the weight coefficient ω = Nl / Ni.
[0064] Step S23: Determine the weight of the character according to the first total occurrence frequency, the second total occurrence frequency, and the weight coefficient.
[0065] In the embodiment of the present invention, the step of determining the weight of the character in Step S23 according to the first total occurrence frequency, the second total occurrence frequency, and the weight coefficient includes: subtract the product of the second total occurrence frequency and the weight coefficient from the first total occurrence frequency to obtain the initial weight value, and perform normalization processing on the initial weight value to obtain the weight. Specifically, if the initial weight value is represented as W c , the first total occurrence frequency is represented as A n , the second total occurrence frequency is represented as B n , then W c = A n - ω * B n . In the embodiment of the present invention, the trigonometric function is used to perform normalization processing on the initial weight value. Then, W c = arctan(A n - ω * B n ), so that the weight is distributed in the interval (-1, 1). When a character is more likely to appear in text messages that need to be intercepted, it is more likely to take a negative value. Thus, when there is a significant difference in the quantity level between normal text messages that do not need to be intercepted and illegal text messages that need to be intercepted, due to the fact that the weight of the character is related to the weight coefficient, even a small number of illegal text messages can increase the weight of illegal text message characters through the weight assignment formula, and there is no situation where the basic character set remains rigid while the text message service changes, adapting to the flexible and changeable real - world text message recognition scenarios, ensuring the recognition accuracy for illegal text messages that need to be intercepted, and thus accurately intercepting text messages that are not allowed to be sent to the recipient.
[0066] If the existing basic character set, consisting of English characters, common punctuation marks, and common Chinese characters, is used to identify SMS messages containing words that circumvent interception using homophones, rare names, or nicknames with symbols, it will misidentify SMS messages that do not need to be intercepted, resulting in a high misidentification rate. To further improve interception accuracy and precisely identify the SMS messages to be identified, in step S14, before identifying the SMS messages using the basic character set based on weights, the method provided in this embodiment further includes: preprocessing the SMS messages to be identified using pre-determined noise words. During preprocessing, if noise words are present in the SMS messages to be identified, they are removed. The noise words are those words that would be intercepted when identifying the SMS messages using the basic character set based on weights, even if they do not need to be intercepted. These noise words can include the aforementioned words that circumvent interception using homophones, rare names, and nicknames with symbols.
[0067] like Figure 4 As shown, in step S14, the step of identifying the SMS message to be identified based on weight using the basic character set includes:
[0068] Step S31: Based on the weights, determine the first weight threshold, wherein the probability of a character with a weight greater than the first weight threshold appearing in a text message that needs to be blocked is the same as the probability of a character with a weight less than the first weight threshold appearing in a text message that needs to be blocked.
[0069] Step S32: Generate a target character set based on the first weight threshold and the basic character set, and use the target character set to identify the SMS message to be identified. That is, based on the first weight threshold, the target character set is selected from the basic character set. The characters in the target character set are characters with a weight greater than the first weight threshold, or characters with a weight less than the first weight threshold.
[0070] Specifically, such as Figure 5As shown, in step S32 of this embodiment, the step of identifying the SMS message to be identified using the target character set includes: determining whether the SMS message to be identified needs to be intercepted based on the number of characters appearing in the target character set among the characters to be identified in the SMS message. In this embodiment, the target character set generated based on the first weight threshold and the basic character set includes a white character set and / or a black character set. The specific thresholds of the first weight threshold for the white character set and the black character set can be the same or different. The specific threshold can be determined based on the frequency of each character appearing in historical SMS messages in historical released SMS messages and historical intercepted SMS messages. Here, for ease of description and distinction, this embodiment pre-sets the value of the first weight threshold for the white character set as the first threshold and pre-sets the value of the first weight threshold for the black character set as the second threshold. After assigning weight scores to all characters in the basic character set in step S23, this embodiment, based on the actual situation of historical SMS service, identifies all characters in the basic character set whose weight scores are greater than the first threshold as the white character set. That is, the white character set consists of characters with a low probability of appearing in SMS messages that need to be intercepted. The number of characters not appearing in the white character set is calculated based on the number of characters in the text message content that appear in the white character set. Therefore, if the number of characters not appearing in the white character set does not exceed a pre-set first threshold, the text message to be identified can be considered a trustworthy text message (i.e., a text message that does not need to be intercepted). Otherwise, there is a risk of it being judged as an illegal text message (i.e., a text message that needs to be intercepted). Thus, the white character set is used to intercept text messages containing characters above the first threshold that do not belong to the white character set. The selection logic for the black character set is similar to that of the white character set. All characters in the basic character set with a weight lower than the second threshold are classified into the black character set. That is, the black character set consists of characters with a high probability of appearing in text messages that need to be intercepted. When using it, the number of characters in the text message to be identified that appear in the black character set determines whether the text message to be identified needs to be intercepted. Specifically, if the number of characters in the text message to be identified that appear in the black character set exceeds the second threshold, the text message to be identified is at risk of being judged as an illegal text message; otherwise, the text message to be identified will be judged as a trustworthy text message and allowed to be sent to the recipient.
[0071] As can be seen from the above, if the sender needs to send a text message to the receiver, this embodiment of the invention, after determining the text message type, uses the basic character set corresponding to that text message type to identify the text message based on weights. If there are noise words in the text message, the noise words are filtered out (i.e., the noise words in the text message are removed). The text message is then identified using either a white character set or a black character set generated based on the basic character set. During identification, if the number of characters in the text message that do not appear in the white character set exceeds a preset first threshold, the text message is blocked; otherwise, it is allowed to be sent to the receiver. Alternatively, if the number of characters in the text message that appear in the black character set exceeds a preset second threshold, the text message is blocked; otherwise, it is allowed to be sent to the receiver.
[0072] In this embodiment of the invention, to further improve the interception accuracy and ensure that the interception accuracy is within the expected range, thereby reducing the false recognition rate, before generating the target character set according to the first weight threshold and the basic character set in step S32, this embodiment of the invention uses the frequency of occurrence of each character in historical SMS messages in historical released SMS messages and historical intercepted SMS messages, and adopts a backtracking calculation method to correct the value of the first weight threshold. Therefore, based on the interception situation of historical SMS messages using the first weight threshold, the value of the first weight threshold is continuously adjusted and optimized, thereby ensuring that when the obtained target character set identifies the SMS message to be identified, it can intercept more illegal SMS messages and falsely intercept fewer legitimate SMS messages. Based on this, as... Figure 6 As shown, the method provided in this embodiment of the invention further includes the following steps:
[0073] Step S41: Determine the first candidate value of the first weight threshold, and determine the candidate character set of the target character set based on the first candidate value and the basic character set. If it is necessary to optimize the first weight threshold of the black character set, during the screening, characters with a weight lower than the first candidate value in the basic character set are used as the candidate character set of the black character set. If it is necessary to optimize the first weight threshold of the white character set, characters with a weight greater than the first candidate value in the basic character set are used as the candidate character set of the white character set. The first candidate values corresponding to the black character set and the white character set are different, and the first candidate values corresponding to different character sets will be determined based on the frequency of the characters appearing in historical SMS messages.
[0074] Step S42: Use the candidate character set to identify historical SMS messages in order to determine the first proportion of historically blocked SMS messages among those blocked using the candidate character set.
[0075] Step S43: Determine whether the first proportion value reaches the preset first proportion threshold. If so, use the first candidate value as the value of the first weight threshold. If the first proportion value reaches the preset first proportion threshold, it means that the obtained candidate character set can cover more SMS messages that need to be blocked, but will mistakenly block fewer legitimate SMS messages, thereby ensuring the interception accuracy and reducing the false recognition rate.
[0076] In step S42, the step of identifying historical SMS messages using candidate character sets includes: determining whether to intercept historically released and / or intercepted SMS messages based on the number of characters appearing in the candidate character set from the characters in the historically released and intercepted SMS messages. Specifically, if the candidate character set is a candidate white character set of a white character set, the number of characters in the historically released and intercepted SMS messages that do not appear in or do not match the candidate white character set is obtained based on the number of characters appearing in the candidate white character set from the characters in the historically released and intercepted SMS messages. If the number of missing characters exceeds the first candidate threshold of the first number threshold, the historically released and intercepted SMS messages are intercepted; otherwise, they are allowed to send. If the candidate character set is a candidate black character set of a black character set, it is determined whether the number of characters appearing in the candidate black character set from the characters in the historically released and intercepted SMS messages exceeds the second candidate threshold of the second number threshold. If so, the historically released and intercepted SMS messages are intercepted; otherwise, they are allowed to send. Therefore, while using the first candidate value as the value of the first weight threshold, the first number candidate threshold is used as the value of the first number threshold, and the second number candidate threshold is used as the value of the second number threshold, thereby optimizing the first weight threshold.
[0077] like Figure 7 As shown in this embodiment of the invention, the step of recognizing the SMS message to be recognized based on weight using a basic character set further includes the following steps:
[0078] Step S51: Determine the weight of each character to be identified in the SMS message. The weight of the character to be identified is obtained by assigning the weight of the character to be identified the same as that of the character in the basic character set.
[0079] Step S52: Based on the weight of each character to be recognized, determine the average weight of the characters to be recognized, wherein the average weight is obtained by averaging the weights of each character to be recognized.
[0080] Step S53: Determine whether the average weight is less than the predetermined second weight threshold. If so, intercept the SMS message to be identified.
[0081] Therefore, in addition to using the obtained target character set to identify and intercept SMS messages, this embodiment of the invention can also use the weight of each character in the basic character set to obtain the average weight of each character to be identified in a single SMS message. Based on the average weight, the SMS message to be identified is identified. When the average weight is less than the preset second weight threshold, the probability that the single SMS message to be identified is a message that needs to be intercepted is very high, thus the SMS message to be identified is intercepted with high accuracy.
[0082] like Figure 8 As shown, to further improve the interception accuracy, this embodiment of the invention optimizes and adjusts the second weight threshold by backtracking calculations using historical SMS messages obtained within a predetermined time period. Based on this, before step S53 determines whether the average weight is less than the predetermined second weight threshold, the method provided by this embodiment of the invention further includes the following steps:
[0083] Step S61: Determine the second candidate value of the second weight threshold, and sum the weights of the characters in each historical released SMS message and take the average to obtain the first historical average weight, and sum the weights of the characters in each historical blocked SMS message and take the average to obtain the second historical average weight.
[0084] Step S62: Determine whether the first historical average weight and the second historical average weight are less than the second candidate value. If so, intercept the historical released SMS and the historical blocked SMS, and obtain the total number of intercepted historical released SMS and historical blocked SMS. Based on the total number of SMS, determine the second proportion value of the intercepted historical blocked SMS.
[0085] Step S63: Determine whether the second proportion value reaches the preset second proportion threshold. If so, use the second candidate value as the value of the second weight threshold.
[0086] By optimizing and adjusting the second weight threshold through the above backtracking calculation, when using the optimized second weight threshold to identify SMS messages, it can cover and block more illegal SMS messages, but will mistakenly block fewer legitimate SMS messages, thereby improving the interception accuracy.
[0087] As can be seen from the above, the embodiments of the present invention are based on historical SMS sending records. The data source for filtering the basic character set is the characters in historical released and blocked SMS messages, distinguished by industry and template type. This is dynamically adjusted according to the SMS service. By evaluating the frequency of occurrence of each character, the characters in the basic character set are weighted and scored. During identification, the probability of an SMS message being blocked can be determined based on the weight of the characters in the SMS message, resulting in a high interception accuracy. Secondly, to further improve the interception accuracy, noise words are input to filter the SMS messages to be identified. For example, by inputting uncommon names or entire nicknames containing symbols, the SMS messages to be identified are prevented from being mistakenly blocked due to the presence of uncommon names or entire nicknames containing symbols. In addition, since the character weight is related to the obtained weight coefficient, even a small number of SMS messages that need to be blocked can have their character weights increased through the weight value formula. There is no situation where the basic character set remains unchanged despite changes in the SMS service. This allows for flexible handling of the interception of illegal SMS messages in various complex service scenarios, improving the interception accuracy of black and gray market SMS messages.
[0088] The second embodiment of the present invention also provides an electronic device, including: a processor and a memory storing a program, wherein the program includes instructions, which, when executed by the processor, cause the processor to perform the above-described SMS processing method. For details of the SMS processing method, please refer to the content provided in the first embodiment of the present invention, and the embodiments of the present invention will not be described again here.
[0089] The third embodiment of the present invention also provides a non-transitory machine-readable medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above-described SMS processing method. For details of the SMS processing method, please refer to the content provided in the first embodiment of the present invention, and the embodiments of the present invention will not be described again here.
[0090] The fourth embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by the processor of a computer, is used to cause the computer to perform the above-described SMS processing method. For details of the SMS processing method, please refer to the content provided in the first embodiment of the present invention, and the embodiments of the present invention will not be described again here.
[0091] The fifth embodiment of the present invention, based on the above four embodiments, also provides a specific application implementation of the SMS processing method.
[0092] like Figure 2As shown, based on historical released and blocked SMS messages within a predetermined time period (e.g., the last 3 months to 6 months or 1 year), a basic character set is split according to different SMS types. For example, e-commerce marketing SMS messages will correspond to one basic character set, internet finance notification SMS messages will correspond to another, social marketing SMS messages will correspond to yet another, and so on. Secondly, since the content of verification code SMS messages is basically uniform, verification code SMS messages will also correspond to a basic character set. Industry filtering relies on customer-defined industry differentiation rules or intelligent identification of the industry to which historical SMS messages belong through an algorithm model. The underlying principle of this algorithm model is to classify and process a large amount of manually labeled SMS data to determine the industry and template type of the SMS message. Template type filtering relies on the template type selected by the customer when applying for a template. If no template type is specified, the algorithm model intelligently identifies the template type to determine whether the SMS message is a marketing, notification, or verification code type.
[0093] First, this embodiment of the invention needs to count the frequency of each character in historical SMS messages in the corresponding type of SMS. For example, here is a normal e-commerce notification SMS that does not need to be blocked: Your price-locked service for the Lijiang to Shenzhen flight is about to expire, valid until ×× month ×× day ×× hour. Please complete the ticket booking within the validity period. Then there is an illegal SMS disguised as a logistics notification that needs to be blocked: Hello, because you used... Add 132××××××22 as a friend on Alipay and register to receive a message. Taiwan! If a normal SMS message is identified as a notification message from the e-commerce or travel industry, then character statistics are performed as shown in Table 1:
[0094] Character Statistics Table 1
[0095]
[0096] Table 2 shows the character statistics of illegal text messages:
[0097] Character Statistics Table 2
[0098]
[0099] By statistically analyzing the total number of historical released and blocked SMS messages of each SMS type over the past M months in the above manner, a statistical table of normal SMS character sets is obtained, containing a total of x1 normal SMS messages, as shown in Table 3:
[0100] Character set table 3
[0101]
[0102] Based on the characters in character set table 3, a basic character set consisting of all characters from historical SMS messages is obtained, and each character in the basic character set is weighted and scored.
[0103] The method provided in this invention aims to ignore a character that appears frequently in both normal and illegal SMS messages (such as historically allowed SMS messages). Conversely, it highlights a character that appears frequently in illegal SMS messages but infrequently in normal SMS messages. Finally, it marks a character that appears frequently in normal SMS messages but infrequently in illegal SMS messages. Based on this idea, and considering that the volume of normal and illegal SMS messages often differs by orders of magnitude, this invention assigns weights to all characters in the basic character set.
[0104] First, calculate the weighting coefficient ω, ω = Nl / Ni, where Nl is the first total number of characters of historically allowed SMS messages, and Ni is the second total number of characters of historically blocked SMS messages.
[0105] Then the weight W of each character c for:
[0106] W c =A n -ω*B n .
[0107] Then, trigonometric functions are used to normalize the final result:
[0108] W c =arctan(A n -ω*B n ).
[0109] In the formula, A n B is the total frequency of the first occurrence of the character in the historical release SMS messages within a predetermined time period. n This represents the second-highest total frequency of a character appearing in historical blocked SMS messages within a predetermined time period. After normalization, the weights are distributed in the interval (-1, 1). The more frequently a character appears in historical blocked SMS messages (i.e., SMS messages that need to be blocked), the more likely it is to receive a negative score.
[0110] like Figure 5 As shown, there are three identification strategies when using the basic character set to identify and intercept SMS messages. In actual processing, one of these strategies can be selected according to the requirements. The identification strategies include:
[0111] 1. By generating a white character set, SMS messages containing characters not belonging to the white character set can be identified or blocked. In this embodiment of the invention, after assigning weight scores to all characters in the basic character set, based on the actual situation of historical SMS services, all characters in the basic character set with weight scores greater than a first threshold are defined as the white character set. That is, the white character set consists of characters with a low probability of appearing in SMS messages that need to be blocked. When the number of characters not appearing in the white character set does not exceed a pre-set first threshold, the SMS message to be identified can be considered a trustworthy SMS message (i.e., an SMS message that does not need to be blocked). Otherwise, there is a risk of it being judged as an illegal SMS message (i.e., an SMS message that needs to be blocked). Thus, the white character set is used to block SMS messages containing characters exceeding the first threshold that do not belong to the white character set.
[0112] 2. By generating a black character set, SMS messages containing symbols from the black character set are identified or blocked. In this embodiment of the invention, all characters in the basic character set with a weight lower than a second threshold are classified as black character sets. That is, the black character set consists of characters with a high probability of appearing in SMS messages that need to be blocked. In use, the number of characters in the black character set that appear in the SMS message to be identified determines whether the SMS message needs to be blocked. Specifically, if the number of characters in the black character set that appear in the SMS message to be identified exceeds the second threshold, the SMS message to be identified is at risk of being judged as an illegal SMS message and thus blocked; otherwise, the SMS message to be identified will be judged as a trustworthy SMS message and allowed to be sent to the recipient.
[0113] 3. Generate the average weight of a single SMS message to be identified. Calculate the character weight score for each SMS message to be identified. Identify or block the SMS message based on the final weight score. Specifically, assign the weights of characters in the basic character set that are the same as the characters to be identified in the SMS message to the characters to be identified, thus obtaining the weights of each character to be identified in the SMS message. Sum the weights of each character to be identified and take the average to obtain the average weight of a single SMS message to be identified. When the average weight is less than a preset second weight threshold, the probability that the SMS message to be identified is high is that it needs to be blocked, and thus the SMS message to be identified is blocked.
[0114] After determining the white character set for identifying the SMS message based on the industry and template type, this embodiment of the invention provides the following operation flow to improve the interception accuracy when using the white character set for identification:
[0115] 11. Determine candidate values for obtaining the first threshold of the white character set. Select multiple characters in the basic character set whose weight is greater than the first threshold (i.e. characters with a low probability of appearing in illegal text messages) as candidate character sets for the white character set. For example, based on three different candidate values set for the first threshold, select X1, Y1, and Z1 characters with higher weight scores from the basic character set as candidate character sets for the white character set.
[0116] 12. Using a backtracking calculation method, candidate character sets corresponding to X1, Y1, and Z1 characters are used to identify historical SMS messages from the past M1 months. If each historical SMS message contains more than w1 (i.e., the first threshold) characters that are not in the candidate white character set, the corresponding historical SMS message is blocked. Different candidate character sets of the white character set will block x1%, y1%, and z1% of SMS services respectively, corresponding to cover a1%, b1%, and c1% of illegal SMS messages that need to be blocked.
[0117] 13. Adjust the candidate values of the first threshold, M1, w1 and other parameters so that the selection logic of the candidate character set of the white character set can cover and intercept more illegal SMS messages, but will mistakenly intercept fewer legitimate SMS messages. Thus, when using the candidate character set as the final white character set for SMS identification, the interception success rate is high and the expected requirements are met.
[0118] To improve the interception accuracy when using the black character set for recognition, the method provided in this embodiment of the invention performs the following operation process:
[0119] 21. Determine candidate values for obtaining the second threshold of the black character set. Select multiple characters in the basic character set whose weight is less than the second threshold (i.e. characters with a high probability of appearing in the SMS to be intercepted) as the candidate character set of the black character set. Specifically, based on the three different candidate values set for the second threshold, select X2, Y2 and Z2 characters with lower weights from the basic character set as the candidate character set of the black character set.
[0120] 22. Using a backtracking calculation method, candidate character sets corresponding to X2, Y2, and Z2 characters are used to identify historical SMS messages from the past M2 months. If w2 (i.e., the second threshold) or more characters appear in the candidate character sets of the black character set in each historical SMS message (i.e., historical released SMS messages and historical blocked SMS messages), the corresponding historical SMS message is blocked. Different candidate character sets of the black character set will block x2%, y2%, and z2% of SMS services respectively. Among the blocked SMS messages, a2%, b2%, and c2% are illegal SMS messages respectively.
[0121] 23. Adjusting parameters such as the candidate value of the second threshold, M2 months, and w2 allows the selection logic of the candidate character set of the black character set to cover and intercept more illegal SMS messages, but will mistakenly intercept fewer legitimate SMS messages. Thus, when the candidate character set is used as the final black character set for SMS identification, the interception success rate is high, meeting the expected requirements.
[0122] To improve the interception accuracy when identifying SMS messages by calculating the average weight of the SMS messages to be identified, this embodiment of the invention uses a backtracking calculation method to adjust the second weight threshold according to the following process:
[0123] 31. Calculate the average weight of each historical SMS message: Determine the weight of each character in the historical SMS message based on the weight of each character in the basic character set, and sum the weights of each character in the historical SMS message and take the average value to obtain its average weight.
[0124] 32. Assume that the second candidate value of the second weight threshold used for blocking is w3. When the average weight of a historical SMS message is lower than w3, it will be blocked.
[0125] 33. Adjust parameter w3, that is, divide the second weight threshold according to the actual situation, thereby improving the interception accuracy when identifying SMS by calculating the average weight of the SMS to be identified. In other words, ensure that the identification results cover more illegal SMS, but will falsely intercept fewer legitimate SMS, and the interception success rate meets the expected requirements.
[0126] As can be seen, this embodiment of the invention uses historical SMS sending records as the data source for filtering the basic character set. The data source consists of characters from historically released and blocked SMS messages, differentiated by industry and template type. This is dynamically adjusted based on the SMS service. By evaluating the frequency of each character's occurrence, the characters in the basic character set are weighted and scored. During identification, the probability of an SMS message being blocked is determined based on the weight of the characters within it, resulting in high interception accuracy. Secondly, to further improve interception accuracy, noise words are input to filter the SMS messages to be identified. For example, by inputting uncommon names or entire nicknames containing symbols, the SMS messages to be identified are prevented from being mistakenly blocked due to the presence of such names. Furthermore, since the character weights are related to the obtained weight coefficients, even a small number of SMS messages that need to be blocked can have their illegal SMS character weights increased through the weighting formula. This avoids the situation where the basic character set remains unchanged despite changes in the SMS service, allowing for flexible handling of illegal SMS interception in various complex service scenarios.
[0127] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] It should be noted that the term "comprising" and its variations used in the embodiments of the present invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that, unless explicitly indicated otherwise in the context, they should be understood as "one or more".
[0130] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0131] The steps described in the method embodiments provided by this invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of this invention is not limited in this respect.
[0132] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply independence or alternativeity from other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.
[0133] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for processing text messages, comprising: Retrieve historical SMS messages, including historical allowed SMS messages and historical blocked SMS messages; The characters in the historical SMS messages are statistically analyzed to obtain the first total frequency of occurrence of the characters in the historical released SMS messages within a predetermined time period, and the second total frequency of occurrence of the characters in the historical blocked SMS messages within a predetermined time period. Based on the characters in the historical SMS messages, a basic character set is generated. Based on the first total occurrence frequency and the second total occurrence frequency, the weight of each character in the basic character set is determined, and the weight is used to reflect the probability of the character appearing in the SMS message that needs to be blocked; The basic character set is used to identify the SMS message to be identified based on the weight, in order to determine whether the SMS message to be identified needs to be blocked; The step of determining the weight of each character in the basic character set based on the first total frequency of occurrence and the second total frequency of occurrence includes: Based on the first total occurrence frequency, a first total number of characters belonging to the historical released SMS messages in the basic character set is determined, and based on the second total occurrence frequency, a second total number of characters belonging to the historical blocked SMS messages in the basic character set is determined; Calculate the ratio of the first total number of characters to the second total number of characters to obtain the weighting coefficient; The weight of the character is determined based on the first total frequency of occurrence, the second total frequency of occurrence, and the weight coefficient. The step of identifying the SMS message to be identified using the basic character set based on the weights further includes: The weight of each character to be identified in the SMS message to be identified is determined. The weight of the character to be identified is obtained by assigning the weight of the character that is the same as the character to be identified in the basic character set to the character to be identified. Based on the weight of each of the characters to be identified, the average weight of the characters to be identified is determined, wherein the average weight is obtained by averaging the weights of each of the characters to be identified. If the average weight is less than a predetermined second weight threshold, the SMS message to be identified is intercepted.
2. The method according to claim 1, wherein, The step of determining the weight of the character based on the first total frequency of occurrence, the second total frequency of occurrence, and the weight coefficient includes: The first total frequency of occurrence is subtracted from the product of the second total frequency of occurrence and the weight coefficient to obtain the initial weight value. The initial weight value is then normalized to obtain the weight.
3. The method according to claim 1, wherein, Before recognizing the SMS message to be identified based on the weight using the basic character set, the method further includes: The SMS message to be identified is preprocessed using predetermined noise words. During preprocessing, if the noise words are present in the SMS message to be identified, the noise words are removed from the SMS message to be identified. The noise words are words that would be intercepted in the SMS message to be identified when the SMS message to be identified is identified using the basic character set based on the weight.
4. The method according to claim 1, wherein, The steps for identifying the SMS message to be identified using the basic character set based on the weights include: Based on the weight, a first weight threshold is determined, wherein the probability of a character with a weight greater than the first weight threshold appearing in a text message that needs to be blocked is less than the probability of a character with a weight less than the first weight threshold appearing in a text message that needs to be blocked. A target character set is generated based on the first weight threshold and the basic character set, so as to identify the SMS message to be identified using the target character set.
5. The method according to claim 4, wherein, The steps for identifying the SMS message using the target character set include: Based on the number of characters in the target character set that appear in the text message to be identified, it is determined whether the text message to be identified needs to be intercepted.
6. The method according to claim 5, wherein, Before generating the target character set based on the first weight threshold and the base character set, the method further includes: Determine a first candidate value for the first weight threshold, and determine a candidate character set for the target character set based on the first candidate value and the basic character set; The candidate character set is used to identify the historical SMS messages in order to determine a first proportion of the historical SMS messages that were blocked using the candidate character set. Determine whether the first proportion value reaches the preset first proportion threshold. If so, use the first candidate value as the value of the first weight threshold.
7. The method according to claim 6, wherein, The steps for identifying the historical SMS messages using the candidate character set include: Based on the number of characters appearing in the candidate character set from the characters in the historically released SMS messages and the historically blocked SMS messages, it is determined whether the historically released SMS messages and / or the historically blocked SMS messages need to be blocked.
8. The method according to claim 1, wherein, Before determining whether the average weight is less than a predetermined second weight threshold, the method further includes: A second candidate value for the second weight threshold is determined, and the weights of the characters in each historical released SMS message are added together and averaged to obtain a first historical average weight, and the weights of the characters in each historical blocked SMS message are added together and averaged to obtain a second historical average weight. Determine whether the first historical average weight and the second historical average weight are less than the second candidate value. If so, intercept the historical released SMS and the historical blocked SMS, and obtain the total number of the intercepted historical released SMS and the historical blocked SMS, so as to determine the second proportion value of the intercepted historical blocked SMS based on the total number of SMS. Determine whether the second proportion value reaches the preset second proportion threshold. If so, use the second candidate value as the value of the second weight threshold.
9. The method according to claim 1, wherein, Before performing statistical analysis on the characters contained in the historical text messages, the method further includes: The historical text messages are classified according to different text message types, and the characters of the historical text messages belonging to the same text message type are counted, and the basic character set is generated based on the characters of the historical text messages.
10. An electronic device, comprising: A processor, and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 9.
11. A non-transitory machine-readable medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Short message identification method and related equipment
CN111586695A
Short message service (SMS) message segmentation
US20150172883A1