Fraud short message processing method and device, electronic equipment and storage medium
By identifying and dynamically updating the keyword and phrase weights of fraudulent text messages, the problem of fixed rules being unable to identify fraudulent text messages that have changed their tactics has been solved, achieving higher accuracy in identification and processing, and ensuring the security and quality of the communication environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD
- Filing Date
- 2023-11-02
- Publication Date
- 2026-08-04
AI Technical Summary
In existing technologies, fixed rules and vocabulary databases cannot effectively identify and process fraudulent text messages that use different methods, resulting in insufficient accuracy in identification and processing.
By acquiring multiple fraudulent text messages within the target processing period, identifying fraudulent keywords and phrases, dynamically updating the weights in the vocabulary database, determining the degree of fraud based on the weights of the fraudulent keywords and phrases, and then sending the text messages to the corresponding processing platform for processing.
It improves the accuracy of identifying and processing fraudulent text messages, reduces false alarm rates, protects users from fraudulent information, and enhances the security and quality of the communication environment.
Smart Images

Figure CN117580047B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, and in particular to a method, apparatus, electronic device, and storage medium for processing fraudulent text messages. Background Technology
[0002] In the modern information society, the rapid development of communication technology has brought convenience and efficiency. However, at the same time, some criminals use fraudulent text messages to commit fraud, which has caused damage to users' interests. Therefore, it is necessary to identify and deal with fraudulent text messages.
[0003] In existing technologies, static identification strategies are generally used for identification, that is, fixed rules and vocabulary are used to judge fraudulent information. However, since fraudulent text messages often change their fraudulent methods, fixed rules and vocabulary cannot effectively identify and process fraudulent text messages. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for processing fraudulent text messages, which helps improve the accuracy of fraudulent text message identification and processing.
[0005] To address the aforementioned problems, firstly, embodiments of this application provide a method for processing fraudulent text messages, including:
[0006] Obtain multiple fraudulent text messages within the target processing period;
[0007] Identify the fraudulent keywords and fraudulent phrases that include the fraudulent keywords in each of the fraudulent text messages;
[0008] Based on multiple fraudulent text messages within the target processing period, obtain the first weight of each fraudulent keyword and the second weight of each fraudulent phrase;
[0009] For each of the fraudulent text messages, the degree of fraudulentness of the text message is determined based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight.
[0010] For each fraudulent text message, based on the degree of fraud involved, the fraudulent text message is sent to a processing platform corresponding to the degree of fraud. The processing platform is used to process the fraudulent text message using a target processing method.
[0011] Secondly, embodiments of this application provide a device for processing fraudulent text messages, including:
[0012] The SMS acquisition module is used to acquire multiple fraudulent SMS messages within a target processing period.
[0013] The fraud-related vocabulary recognition module is used to identify fraud-related keywords and fraud-related phrases including the fraud-related keywords in each of the fraud-related text messages;
[0014] The vocabulary weight acquisition module is used to acquire a first weight for each fraudulent keyword and a second weight for each fraudulent phrase based on multiple fraudulent text messages within the target processing period.
[0015] The fraud severity determination module is used to determine the degree of fraud in each fraudulent text message based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight.
[0016] The fraudulent SMS distribution module is used to send each fraudulent SMS to a processing platform corresponding to the degree of fraud, based on the degree of fraud in the SMS. The processing platform is used to process the fraudulent SMS using a target processing method.
[0017] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the fraudulent SMS processing method described in embodiments of this application.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, discloses the method for processing fraudulent text messages.
[0019] The fraudulent SMS processing method, apparatus, electronic device, and storage medium provided in this application embodiment, after acquiring multiple fraudulent SMS messages within a target processing period, identifies fraudulent keywords and fraudulent phrases including those keywords in each SMS message. Based on the multiple fraudulent SMS messages within the target processing period, a first weight for each fraudulent keyword and a second weight for each fraudulent phrase are obtained. For each fraudulent SMS message, the degree of fraud is determined based on the fraudulent keywords, first weight, fraudulent phrase, and second weight. Based on the degree of fraud, the fraudulent SMS message is sent to the corresponding processing platform, which processes the SMS message using the target processing method. Since the first weight of the fraudulent keywords and the second weight of the fraudulent phrases can be obtained from multiple fraudulent SMS messages within the target processing period, the determined degree of fraud is adapted to the current target processing period and is not based on static rules. Therefore, the processing is performed by the corresponding processing platform based on the degree of fraud, which can improve the accuracy of fraudulent SMS message identification and processing. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a method for processing fraudulent text messages provided in an embodiment of this application;
[0022] Figure 2 This is a schematic diagram illustrating the distribution of fraudulent text messages in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram illustrating the content analysis of fraudulent text messages in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram illustrating the extraction of fraud-related keywords and phrases in an embodiment of this application;
[0025] Figure 5 This is a schematic diagram illustrating the process of extracting text content from fraudulent text messages in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram illustrating the contextual analysis of fraudulent phrases in an embodiment of this application;
[0027] Figure 7 This is a schematic diagram illustrating the acquisition of fraudulent text messages with duplicate content in an embodiment of this application;
[0028] Figure 8 This is a schematic diagram illustrating the collection of fraudulent text messages with duplicate content in an embodiment of this application;
[0029] Figure 9 This is a schematic diagram illustrating the dynamic updating of fraud-related terms in the vocabulary database in an embodiment of this application;
[0030] Figure 10 This is a schematic diagram of the system for implementing the fraudulent SMS processing method in the embodiments of this application;
[0031] Figure 11 This is a schematic diagram of the structure of a fraudulent SMS processing device provided in an embodiment of this application.
[0032] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] Figure 1 This is a flowchart illustrating a method for handling fraudulent text messages provided in an embodiment of this application. This method can be executed by an electronic device such as a server. Figure 1 As shown, the method includes steps 110 to 150.
[0035] Step 110: Obtain multiple fraudulent text messages within the target processing period.
[0036] Among them, fraudulent text messages are those containing fraudulent or deceptive information that harms the interests of users.
[0037] The reporting platform sends fraudulent text messages reported by users or organizations to electronic devices that execute fraudulent text message processing methods. These devices receive the fraudulent text messages from the reporting platform and extract multiple fraudulent text messages within a target processing period. The electronic devices then process these multiple fraudulent text messages within a single processing period in a unified manner, facilitating the identification of patterns and subsequent action against fraudulent text messages within that period.
[0038] Step 120: Identify the fraudulent keywords and fraudulent phrases including the fraudulent keywords in each of the fraudulent text messages.
[0039] Among them, fraud-related keywords are keywords that involve fraudulent or deceptive information. They can be verbs, such as "click," or nouns.
[0040] For example, the content of each fraudulent text message can be segmented into words, and each segmentation result can be matched with keywords in a vocabulary database. If the match is successful, the segmentation result is identified as a fraudulent keyword. Based on the vocabulary database matching, in order to identify emerging fraud methods, multiple fraudulent text messages within the same target processing period can be compared to identify fraudulent text messages with duplicate content, and then words from the fraudulent text messages with duplicate content can be extracted as fraudulent keywords.
[0041] After identifying the fraudulent keywords in the fraudulent text message, contextual analysis can be performed based on the fraudulent keywords to extract fraudulent phrases that include the fraudulent keywords.
[0042] Step 130: Based on the multiple fraudulent text messages within the target processing period, obtain the first weight of each fraudulent keyword and the second weight of each fraudulent phrase.
[0043] The vocabulary database can store keywords and their corresponding first weights, as well as fraudulent phrases and their corresponding second weights. After identifying the fraudulent keywords and phrases in each fraudulent text message, the first weight of the fraudulent keyword and the second weight of the fraudulent phrase can be obtained from the vocabulary database. Alternatively, the frequency of occurrence of the fraudulent keyword within the target processing period can be counted, and the first weight of the fraudulent keyword can be determined based on this frequency. The frequency of occurrence of the fraudulent phrase within the target processing period can also be counted, and the second weight of the fraudulent phrase can be determined based on this frequency. Alternatively, the initial first weight recorded in the vocabulary database can be updated based on the frequency of occurrence of the fraudulent keyword within the target processing period, and the updated value can be the first weight of the fraudulent keyword. The initial second weight recorded in the vocabulary database can also be updated based on the frequency of occurrence of the fraudulent phrase within the target processing period, and the updated value can be the second weight of the fraudulent phrase.
[0044] Step 140: For each fraudulent text message, determine the degree of fraudulentness of the text message based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight.
[0045] Based on the keywords and their corresponding first weights, as well as the phrases and their corresponding second weights used in each fraudulent text message, fraudulent text messages are identified and classified. By analyzing the fraudulent keywords and their corresponding first weights, and the fraudulent phrases and their corresponding second weights, the degree of fraud in the text messages is determined, and this degree of fraud can be represented by a fraud coefficient.
[0046] For example, one can count the number of fraudulent keywords and the total number of fraudulent text messages, and count the number of each fraudulent keyword and the number of fraudulent text messages in each fraudulent text message. The fraud coefficient can be calculated using the following formula:
[0047]
[0048] Where δ represents the fraud coefficient of the fraudulent text message, i.e., the degree of fraud, and x i This represents the number of the i-th fraudulent keyword in the fraudulent text message, a. i y represents the first weight of the i-th fraudulent keyword. j b represents the number of the j-th fraudulent phrase in the entire fraudulent text message. j This represents the second weight of the j-th fraudulent phrase.
[0049] Step 150: For each fraudulent text message, based on the degree of fraud involved, the fraudulent text message is sent to a processing platform corresponding to the degree of fraud involved. The processing platform is used to process the fraudulent text message using a target processing method.
[0050] Figure 2 This is a schematic diagram illustrating the distribution of fraudulent text messages in an embodiment of this application, such as... Figure 2 As shown, after determining the degree of fraud in each fraudulent text message based on its fraudulent keywords, first weight, fraudulent phrases, and second weight, the fraudulent level of each message is determined, thus defining its corresponding fraud severity level (each level corresponds to a fraud severity range). Fraud severity levels can include, for example, low, medium, high, and very high. The fraudulent text messages are then sent to processing platforms corresponding to their respective fraud severity levels, such as low-level, medium-level, high-level, and very high-level fraudulent text message processing platforms. These platforms then process the messages using targeted methods appropriate to their fraud severity levels, such as blocking the sender's number or preventing that number from sending further messages. By distributing messages to appropriate processing systems based on their fraud severity, interference from fraudulent text messages is effectively reduced. Social media platforms can utilize this strategy to filter out misinformation, fake advertisements, and fraudulent content.
[0051] Based on the severity of fraud, the system categorizes potentially fraudulent SMS messages into different fraud levels, representing varying degrees of potential fraud. The system then precisely distributes these messages to the appropriate processing platforms. High-risk messages may be sent to high-risk platforms, while low-risk messages may be sent to low-risk platforms. This precise distribution process ensures that SMS messages of varying degrees of fraud receive appropriate processing and decision-making. In this way, the system not only reduces false alarm rates and improves accuracy but also protects users from the interference and threats of fraudulent information, enhancing the overall quality and security of the communication environment.
[0052] Steps 110 to 140 above constitute the identification process for fraudulent text messages, while step 150 is the processing process. The identification process analyzes received fraudulent text message messages to determine if they contain features, content, or patterns related to fraud or deception. It utilizes advanced text analysis techniques, pattern recognition, and comparison with a vocabulary database to automatically identify potential fraudulent messages. Once fraudulent content is identified, the processing process distributes the fraudulent text message to different levels of processing platforms for further processing. The processing process focuses on forwarding data and information from a specific source to a predetermined destination. Its core tasks include data transmission, information sharing, and security assurance. During data transmission and processing, it is responsible for data management and security protection, ensuring that data flows to its destination as stipulated. This process plays a crucial role in information transmission, data sharing, and security defense, and its intelligent functions ensure efficient data transmission and secure processing.
[0053] The fraudulent SMS processing method provided in this application, after acquiring multiple fraudulent SMS messages within a target processing period, identifies fraudulent keywords and fraudulent phrases including those keywords in each SMS message. Based on the multiple fraudulent SMS messages within the target processing period, it obtains a first weight for each fraudulent keyword and a second weight for each fraudulent phrase. For each fraudulent SMS message, it determines the degree of fraud based on the fraudulent keywords, first weight, fraudulent phrase, and second weight. Based on the degree of fraud, the fraudulent SMS message is sent to the corresponding processing platform, which processes it using the target processing method. Since the first weight of the fraudulent keywords and the second weight of the fraudulent phrases can be obtained from multiple fraudulent SMS messages within the target processing period, the determined degree of fraud is adapted to the current target processing period and is not based on static rules. Therefore, the processing is carried out by the corresponding processing platform based on the degree of fraud, which can improve the accuracy of fraudulent SMS message identification and processing.
[0054] Based on the above technical solution, the step of identifying fraudulent keywords and fraudulent phrases including the fraudulent keywords in each fraudulent text message includes: for each fraudulent text message, extracting the text content within the fraudulent text message; extracting verbs from the text content and matching the verbs with keywords in a vocabulary database, identifying the successfully matched verbs as the fraudulent keywords; and identifying the fraudulent keywords and their contextual information within the text content as the fraudulent phrases.
[0055] Using advanced text analysis technology, the system performs in-depth content analysis on each text message, extracting key information and phrases, and can also analyze sentiment and theme. This multi-dimensional feature extraction helps the system fully understand the content of the text messages and identify potential fraudulent features, such as fraudulent language and unusual context.
[0056] Figure 3 This is a schematic diagram illustrating the content analysis of fraudulent text messages in an embodiment of this application, such as... Figure 3 As shown, fraudulent text message files are collected from various reporting platforms. After decryption and decompression, the text content of the fraudulent text messages is obtained. Each text message is processed separately. For the currently processed fraudulent text message, the text content is extracted, useless symbols and other information are removed, verbs are identified and extracted from the text content, and the extracted verbs are matched with keywords in the vocabulary database. Verbs that successfully match the keywords in the vocabulary database are identified as fraudulent keywords. Based on the fraudulent keywords and their contextual information, fraudulent phrases are identified. The fraudulent keywords and fraudulent phrases are used together as the fraudulent result keywords.
[0057] Figure 4 This is a schematic diagram illustrating the extraction of fraud-related keywords and phrases in an embodiment of this application, such as... Figure 4 As shown, based on the text content extracted from fraudulent text messages, verbs are extracted, and text containing fraudulent motives within the verbs is captured. This involves matching the verbs with keywords in the vocabulary database. The matching process can be represented as f(x) = lim->verb(d), where f(x) represents the fraudulent keyword, verb(d) represents the verb in the text content, and lim represents the keyword in the vocabulary database. The successfully matched verbs are taken as fraudulent keywords. The content before and after the fraudulent keyword in the fraudulent text message is extracted and concatenated to obtain a fraudulent phrase containing the fraudulent keyword. The process of extracting the fraudulent phrase can be represented as x = f(x) + f(x+1), f(x-1) + f(x), where f(x) represents the fraudulent keyword, f(x+1) represents the content after the fraudulent keyword, f(x-1) represents the content before the fraudulent keyword, and x represents the fraudulent phrase, which is the concatenated verb.
[0058] In addition to extracting fraudulent keywords and phrases, content analysis can also be performed on fraudulent text messages to obtain sentiment analysis results, providing a data basis for determining the degree of fraud in the text messages.
[0059] In the content analysis stage, the system deeply explores the information of fraud-related text messages from multiple perspectives by analyzing aspects such as the text structure, keywords, fraud-related verbs, and context of fraud-related text messages. By identifying keywords and phrases in the text, the system can capture the core content of the text message. At the same time, the analysis of the emotional tendency in the text helps to reveal hidden emotional characteristics and provides clues for the subsequent classification and processing of fraud-related text messages. The system also pays attention to the coherence of the context to ensure an accurate understanding of the overall intention of the text message, rather than just making a surface judgment based on a single word. Through these multi-dimensional analyses, the system can comprehensively grasp the content and meaning of fraud-related text messages, providing strong support for the subsequent judgment of the degree of fraud.
[0060] Based on the above technical solution, the extraction of the text content in the fraud-related text message includes: splitting the text of the fraud-related text message according to the smallest unit to obtain multiple data contents corresponding to the fraud-related text message; encoding each data content with Unicode to obtain the Unicode corresponding to each data content; in the text of the fraud-related text message, removing the data content whose Unicode is outside the range of the text Unicode to obtain the text content of the fraud-related text message.
[0061] Figure 5 It is a processing schematic diagram for extracting the text content in a fraud-related text message in an embodiment of the present application. As Figure 5 shown, first, split the text content of the fraud-related text message. The text of the fraud-related text message can be split according to the smallest unit (which can be a character, such as a single Chinese character, symbol, etc.). Take the text content of the fraud-related text message as the initial variable a, that is, a = text message content, and split the text of the fraud-related text message according to the splitting formula b[n]=a.split('') to obtain multiple data contents b[0], b[1]....b[n]; secondly, encode each data content with Unicode (Universal Character Set) to convert each data content into an integer data to obtain the Unicode corresponding to each data content. The process of Unicode encoding can be represented by the formula int f(x)=(int)b[i], where b[i] represents a single data content, i = 0, 1,...n, and f(x) represents the Unicode corresponding to this data content; finally, obtain the text content of the fraud-related text message through the Unicode. Determine whether the Unicode of each data content is within the range of the text Unicode, that is, determine whether f(x) satisfies u4e00 < f(x) < \u9fff. When the Unicode of the data content is within the range of the text Unicode, determine that this data content is text and append this data content, that is, splice the data content that is text. When the Unicode of the data content is outside the range of the text Unicode, determine that this data content is not text. After removing the data content whose Unicode is outside the range of the text Unicode in this way, the remaining content is the text content of the fraud-related text message.
[0062] After segmenting the text of fraudulent text messages, Unicode encoding is applied. Based on Unicode encoding, the text content of fraudulent text messages can be extracted quickly and accurately.
[0063] Based on the above technical solution, after determining the fraud-related keywords and their contextual information in the text content as the fraud-related phrase, the method further includes: determining the word vector of the fraud-related phrase through a pre-trained language model; determining the synonyms of the fraud-related phrase based on the word vectors; and adding the synonyms to the vocabulary database.
[0064] Figure 6 This is a schematic diagram illustrating the contextual analysis of fraudulent phrases in an embodiment of this application, such as... Figure 6 As shown, for fraudulent phrases extracted from fraudulent text messages, a pre-trained language model is used to obtain their fraudulent word vector representations, thus obtaining the word vectors of the fraudulent phrases. Based on the word vectors of the fraudulent phrases, the semantic relationships and context of the fraudulent phrases can be analyzed by observing their neighboring words or phrases in the context. All words corresponding to the word vectors are determined through the pre-trained language model, and based on the semantic relationships and context of the fraudulent phrases, words with the same semantic relationships and contexts as the fraudulent phrases are obtained from all the words corresponding to the word vectors and used as synonyms of the fraudulent phrases. The synonyms are then added to the vocabulary database.
[0065] By using a pre-trained language model to determine the word vectors of fraudulent phrases, and then using these word vectors to determine the synonyms of the fraudulent phrases, the synonyms are added to the vocabulary database, thus enabling dynamic updates to the vocabulary database and facilitating the discovery of the latest fraudulent words.
[0066] Based on the above technical solution, the step of identifying fraudulent keywords and fraudulent phrases including the fraudulent keywords in each fraudulent text message further includes: when the number of fraudulent text messages in the target processing period is greater than a quantity threshold, acquiring fraudulent text messages with duplicate content among the multiple fraudulent text messages, wherein the quantity threshold is determined based on fraudulent text message data from historical processing periods; and acquiring fraudulent keywords and fraudulent phrases including the fraudulent keywords in the fraudulent text messages with duplicate content.
[0067] By setting a monitoring time window (i.e., the time of one processing cycle), the number of times fraudulent text messages appear within the time window is regularly counted, and the number of occurrences is compared with a quantity threshold. This method can detect abnormal text message activity in real time, quickly identify sudden fraudulent behavior, and improve the awareness of new fraud methods.
[0068] Figure 7 This is a schematic diagram illustrating the acquisition of fraudulent text messages with duplicate content in an embodiment of this application, such as... Figure 7As shown, to monitor the frequency of fraudulent text messages in real time, the number of fraudulent text messages is counted based on a fixed time window (i.e., a processing cycle, such as one hour). By periodically counting the number of text messages, historical data can be generated for all historical processing cycles. This historical data can reflect the fluctuations in fraudulent activity, and based on these fluctuations, a quantity threshold can be determined for each processing cycle within a statistical period (the statistical period includes multiple processing cycles, for example, a statistical period of one day and a processing cycle of one hour). These quantity thresholds can be used as the basis for subsequent analysis. The system compares the number of fraudulent text messages within the target processing cycle with the quantity thresholds to detect whether there is abnormal text message activity within the target processing cycle. When the number of fraudulent text messages within the target processing cycle exceeds the quantity threshold, it is determined that there is abnormal text message activity within the target processing cycle. At this time, the fraudulent text messages with duplicate content among multiple fraudulent text messages within the target processing cycle are obtained, and the content of the fraudulent text messages with duplicate content is taken as the core text content of the fraudulent activity in the target processing cycle.
[0069] Analyzing the content of repeated fraudulent text messages reveals that these messages may contain information about new types of fraudulent activities. The vocabulary database may not contain corresponding keywords, and matching with keywords in the vocabulary database may not yield effective keywords from these fraudulent text messages. In such cases, verbs in the content of these fraudulent text messages can be identified and designated as fraudulent keywords. Furthermore, based on these fraudulent keywords, contextual information can be extracted to obtain fraudulent phrases that include the fraudulent keywords.
[0070] By identifying fraudulent text messages with duplicate content from multiple fraudulent text messages when the number of such messages exceeds a threshold within a target processing period, and by extracting fraudulent keywords and phrases containing those keywords from these duplicate messages, new types of fraudulent activities can be detected promptly. This allows for the rapid discovery of sudden fraudulent activities and further improves the accuracy of fraudulent text message identification and processing.
[0071] Based on the above technical solution, the method further includes: storing fraudulent keywords and phrases obtained from the fraudulent text messages with repeated content into the vocabulary database.
[0072] Keywords and phrases related to fraud obtained from repeated fraudulent text messages may be information related to new types of fraudulent activities. Storing this information in a vocabulary database and dynamically updating the database will facilitate the timely detection of such fraudulent activities.
[0073] Based on the above technical solution, the step of obtaining fraudulent text messages with duplicate content from the multiple fraudulent text messages includes: randomly selecting a preset number of fraudulent text messages from the multiple fraudulent text messages; calculating a hash value for the content of the selected fraudulent text messages; and identifying fraudulent text messages with the same hash value as fraudulent text messages with duplicate content.
[0074] Figure 8 This is a schematic diagram illustrating the collection of fraudulent text messages with duplicate content in an embodiment of this application, such as... Figure 8 As shown, a time series graph can be plotted with the processing period (time window) as the X-axis and the number of fraudulent text messages as the Y-axis. In the time series graph, identify time windows where the number of fraudulent text messages is high, i.e., peak periods. This may indicate that fraudulent activities are more active during certain time periods. Fraudulent text messages with duplicate content can be obtained from peak periods, or duplicate content can be obtained separately for each processing period. Within the target processing period, duplicate content fraudulent text messages can be randomly selected from different data sources (i.e., different reporting platforms). These fraudulent text messages serve as a sample of fraudulent text messages. These data sources can cover various communication channels and social media platforms. Alternatively, random selection can be made from all fraudulent text messages. Extract the sampling center index, i.e., n = random(a), where n represents the sampling center index and a represents the number of fraudulent text messages within the target processing period. Based on the sampling center index, determine the location of a preset number of sampling text messages. The location of the sampling text message can be represented as f[] = |[n+random(a)]-a|...|[(n+random(a)]+a|. Then, extract the corresponding fraudulent text messages based on the sampling text message location. Calculate the hash value of the extracted fraudulent text messages. Determine the fraudulent text messages with the same hash value as those with duplicate content. That is, collect the fraudulent text messages with duplicate content by o = interse(Hash(f[0]..f[n])).
[0075] By extracting a preset number of fraudulent text messages from multiple fraudulent text messages, the amount of data processed can be reduced. By calculating hash values to identify fraudulent text messages with duplicate content, duplicate fraudulent text messages can be obtained relatively quickly.
[0076] Based on the above technical solution, the step of obtaining a first weight for each fraudulent keyword and a second weight for each fraudulent phrase from multiple fraudulent text messages within the target processing period includes:
[0077] For each fraud-related keyword, if the fraud-related keyword does not exist in the vocabulary database, determine the first occurrence number of the fraud-related keyword within the target processing cycle, and determine the first weight of the fraud-related keyword based on the first occurrence number; if the fraud-related keyword exists in the vocabulary database, obtain the first weight of the fraud-related keyword from the vocabulary database.
[0078] For each fraudulent phrase, if the fraudulent phrase does not exist in the vocabulary, determine the second occurrence number of the fraudulent phrase within the target processing cycle, and determine the second weight of the fraudulent phrase based on the second occurrence number; if the fraudulent phrase exists in the vocabulary, obtain the second weight of the fraudulent phrase from the vocabulary.
[0079] By introducing a dynamic weighting mechanism for keywords and phrases, the system regularly updates the criteria for identifying fraudulent terms based on content analysis and statistical analysis of the number of fraudulent text messages within the target processing cycle. This flexibility ensures that the system can adapt to the ever-changing strategies of fraudsters and maintain the accuracy of identifying new types of fraudulent information.
[0080] The dynamic updating of the first weight of fraudulent keywords and the second weight of fraudulent phrases is achieved through the statistical analysis of the number of fraudulent text messages and their content within the target processing period.
[0081] When a fraudulent keyword does not exist in the vocabulary database (i.e., it is not stored there), the first occurrence count of the keyword within the target processing period can be counted, and the first weight of the keyword can be determined based on this first occurrence count. For example, the first occurrence count can be used as the first weight; alternatively, the range of occurrence counts can be determined, and the weight corresponding to that range can be used as the first weight. When a fraudulent keyword exists in the vocabulary database, its first weight can be obtained from the database. Alternatively, the weight of the keyword in the vocabulary database can be updated based on its first occurrence count within the target processing period, and the updated weight can be used as the first weight.
[0082] When a fraudulent phrase does not exist in the vocabulary (i.e., it is not stored in the vocabulary), the second occurrence count of the fraudulent phrase within the target processing period can be counted, and then the second weight of the fraudulent phrase can be determined based on the second occurrence count. For example, the second occurrence count of the fraudulent phrase can be used as the second weight of the fraudulent phrase; alternatively, the range of occurrence counts can be determined, and the weight corresponding to that range can be obtained and used as the second weight of the fraudulent phrase. When a fraudulent phrase exists in the vocabulary, its second weight can be obtained from the vocabulary. Alternatively, the weight of the fraudulent phrase in the vocabulary can be updated based on its second occurrence count within the target processing period, and the updated weight can be used as the second weight of the fraudulent phrase.
[0083] The weight of a fraudulent keyword is determined by its first occurrence, and the weight of a fraudulent phrase is determined by its second occurrence. This allows for flexible adjustment of the weight of fraudulent words, enabling more accurate identification of text messages with varying degrees of fraud.
[0084] Based on the above technical solution, after determining the first weight of the fraud-related keyword according to the first occurrence frequency, the method further includes: storing the fraud-related keyword and the first weight in the vocabulary database;
[0085] After determining the second weight of the fraudulent phrase based on the second occurrence frequency, the method further includes: storing the fraudulent phrase and the second weight in the vocabulary database.
[0086] After determining the first weight of a fraudulent keyword based on its first occurrence frequency, the fraudulent keyword and its corresponding first weight are stored in the vocabulary database. After determining the second weight of a fraudulent phrase based on its second occurrence frequency, the fraudulent phrase and its corresponding second weight are stored in the vocabulary database. This enables dynamic updating of the vocabulary database, which can adapt to new fraud methods and further improve the accuracy of processing fraudulent text messages.
[0087] Figure 9 This is a schematic diagram illustrating the dynamic updating of fraud-related terms in the vocabulary database in this application embodiment, such as... Figure 9 As shown, the system can analyze the extracted fraud-related keywords and the core text content of fraud-related text messages that repeat the same content. It can determine whether the current word (fraud-related keyword or word in the core text content) is a new word that does not exist in the vocabulary database. If the current word is a word that does not exist in the vocabulary database, the vocabulary database is dynamically updated, that is, the word and its corresponding weight are stored in the vocabulary database. If the current word is a word that is already stored in the vocabulary database, no processing is required, and the system continues to judge the next word.
[0088] The system obtains frequency data on fraudulent text messages based on periodic fraud frequency monitoring (monitoring the occurrence frequency of periodic fraudulent text messages). This data reflects the real-time situation of fraudulent activities, including sudden events and fluctuation trends. Simultaneously, the content analysis phase provides the system with information such as fraudulent keywords, fraudulent phrases, and sentiment trends. The system combines the results of periodic fraud frequency monitoring and content analysis to identify fraudulent keywords and phrases, assigning a first weight to each keyword and a second weight to each phrase. By comprehensively analyzing periodic frequency and content information, the system dynamically updates the weights of fraudulent terms (fraudulent keywords and phrases), adjusting these weights based on changes in occurrence frequency and content analysis findings. In this way, the system can flexibly adjust the weights of fraudulent terms according to the latest fraudulent activities and changes in text message content, more accurately identifying text messages with different levels of fraudulent activity. This dynamic update mechanism adapts to constantly changing fraud methods, improves the system's ability to identify fraudulent information, and thus provides users with a healthier communication environment.
[0089] Figure 10 This is a schematic diagram of the system for implementing the fraudulent SMS processing method in this application embodiment, as shown below. Figure 10 As shown, the receiving layer consists of external platforms, such as reporting platforms A, B, C, and D. The collection layer collects fraudulent text messages sent by each reporting platform through the collection area. The parsing layer extracts keywords from the text content of the fraudulent text messages, performs context analysis, and monitors the periodic fraud frequency (the first occurrence of fraudulent keywords, the second occurrence of fraudulent phrases, the number of fraudulent text messages, etc.). The distribution layer updates the weights of fraudulent words in the vocabulary database based on the content analysis results from the analysis area and the periodic fraud frequency monitoring results. Based on the fraudulent words and weights in the fraudulent text messages, it determines the degree of fraud in the messages and then distributes them to the corresponding processing platforms for processing. For example, processing platforms may include low-level processing platforms, medium-level processing platforms, high-level processing platforms, and extremely high-level processing platforms.
[0090] This application's embodiments, through content analysis and dynamically updating the weights of fraudulent terms, can accurately identify fraudulent SMS messages and distribute them precisely to the corresponding processing platforms, thereby reducing false alarm rates and preventing users from being disturbed by fraudulent information. Simultaneously, the implementation of periodic fraud frequency monitoring enables the system to detect and capture sudden fraudulent activities in a short time, allowing for timely countermeasures and ensuring the security and trustworthiness of communications. Dynamically updating the weights of fraudulent terms not only adapts to constantly evolving fraud methods but also helps to continuously improve the system's identification accuracy, thus creating a safer, more trustworthy, and higher-quality communication environment for users. This reduces the interference and threat of fraudulent information to users, thereby having a positive impact on the communications field.
[0091] Figure 11 This is a schematic diagram of the structure of a fraudulent text message processing device provided in an embodiment of this application, as shown below. Figure 11 As shown, the device includes:
[0092] SMS acquisition module 1110 is used to acquire multiple fraudulent SMS messages within the target processing period;
[0093] The fraud-related word recognition module 1120 is used to identify fraud-related keywords and fraud-related phrases including the fraud-related keywords in each of the fraud-related text messages;
[0094] The vocabulary weight acquisition module 1130 is used to acquire a first weight for each fraudulent keyword and a second weight for each fraudulent phrase based on multiple fraudulent text messages within the target processing period.
[0095] The fraud severity determination module 1140 is used to determine the degree of fraud of each fraudulent text message based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight.
[0096] The fraudulent SMS distribution module 1150 is used to send each fraudulent SMS to a processing platform corresponding to the degree of fraud, based on the degree of fraud in the SMS. The processing platform is used to process the fraudulent SMS using a target processing method.
[0097] Optionally, the fraud-related word recognition module includes:
[0098] The text content extraction unit is used to extract the text content from each of the fraudulent text messages.
[0099] The keyword extraction unit is used to extract verbs from the text content, match the verbs with keywords in the vocabulary database, and identify the successfully matched verbs as the fraud-related keywords.
[0100] A phrase extraction unit is used to determine the fraud-related keywords and their contextual information within the text content as the fraud-related phrases.
[0101] Optionally, the text content extraction unit is specifically used for:
[0102] The text of the fraudulent text message is segmented according to the smallest unit to obtain multiple data contents corresponding to the fraudulent text message;
[0103] Each data content is encoded using a Unicode to obtain a Unicode corresponding to each data content.
[0104] In the text of the fraudulent text message, data content whose Unicode is outside the range of Unicode is removed to obtain the text content of the fraudulent text message.
[0105] Optionally, the device further includes:
[0106] The vocabulary update module is used to determine the word vector of the fraudulent phrase through a pre-trained language model; determine the synonyms of the fraudulent phrase based on the word vectors; and add the synonyms to the vocabulary.
[0107] Optionally, the fraud-related word recognition module further includes:
[0108] The duplicate SMS acquisition unit is used to acquire duplicate fraudulent SMS messages among the multiple fraudulent SMS messages when the number of fraudulent SMS messages in the target processing cycle exceeds a quantity threshold. The quantity threshold is determined based on fraudulent SMS message data from historical processing cycles.
[0109] The fraud-related vocabulary extraction unit is used to obtain fraud-related keywords and fraud-related phrases that include the fraud-related keywords from the fraud-related text messages with repeated content.
[0110] Optionally, the device further includes:
[0111] The fraud-related vocabulary storage module is used to store fraud-related keywords and phrases obtained from fraud-related text messages with repeated content into the vocabulary database.
[0112] Optionally, the duplicate SMS acquisition unit is specifically used for:
[0113] A predetermined number of fraudulent text messages are randomly selected from the multiple fraudulent text messages;
[0114] Calculate the hash value of the extracted fraudulent text messages, and identify fraudulent text messages with the same hash value as having duplicate content.
[0115] Optionally, the vocabulary weight acquisition module is specifically used for:
[0116] For each fraud-related keyword, if the fraud-related keyword does not exist in the vocabulary database, determine the first occurrence number of the fraud-related keyword within the target processing cycle, and determine the first weight of the fraud-related keyword based on the first occurrence number; if the fraud-related keyword exists in the vocabulary database, obtain the first weight of the fraud-related keyword from the vocabulary database.
[0117] For each fraudulent phrase, if the fraudulent phrase does not exist in the vocabulary, determine the second occurrence number of the fraudulent phrase within the target processing cycle, and determine the second weight of the fraudulent phrase based on the second occurrence number; if the fraudulent phrase exists in the vocabulary, obtain the second weight of the fraudulent phrase from the vocabulary.
[0118] Optionally, the device further includes:
[0119] The keyword weight storage module is used to store the fraudulent keywords and the first weight into the vocabulary database;
[0120] The phrase weight storage module is used to store the fraudulent phrase and the second weight in the vocabulary database.
[0121] The fraudulent SMS processing device provided in this application embodiment is used to implement the steps of the fraudulent SMS processing method described in this application embodiment. The specific implementation methods of each module of the device are described in the corresponding steps, and will not be repeated here.
[0122] The fraudulent SMS processing device provided in this application, after acquiring multiple fraudulent SMS messages within a target processing period, identifies fraudulent keywords and fraudulent phrases including those keywords in each SMS message. Based on the multiple fraudulent SMS messages within the target processing period, it obtains a first weight for each fraudulent keyword and a second weight for each fraudulent phrase. For each fraudulent SMS message, it determines the degree of fraud based on the fraudulent keywords, first weight, fraudulent phrase, and second weight. Based on the degree of fraud, the fraudulent SMS message is sent to the corresponding processing platform, which processes it using the target processing method. Since the first weight of the fraudulent keywords and the second weight of the fraudulent phrases can be obtained from multiple fraudulent SMS messages within the target processing period, the determined degree of fraud is adapted to the current target processing period and is not based on static rules. Therefore, the processing is performed by the corresponding processing platform based on the degree of fraud, which improves the accuracy of fraudulent SMS message identification and processing.
[0123] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 12As shown, the electronic device 1200 may include one or more processors 1210 and one or more memories 1220 connected to the processors 1210. The electronic device 1200 may also include an input interface 1230 and an output interface 1240 for communicating with another device or system. Program code executed by the processor 1210 may be stored in the memory 1220.
[0124] The processor 1210 in the electronic device 1200 calls the program code stored in the memory 1220 to execute the fraudulent SMS processing method in the above embodiment.
[0125] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the fraudulent SMS processing method as described in this application.
[0126] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0127] The foregoing has provided a detailed description of a method, apparatus, electronic device, and storage medium for processing fraudulent text messages according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
Claims
1. A method for processing fraudulent text messages, characterized in that, include: Obtain multiple fraudulent text messages within the target processing period; Identify the fraudulent keywords and fraudulent phrases that include the fraudulent keywords in each of the fraudulent text messages; Based on multiple fraudulent text messages within the target processing period, obtain the first weight of each fraudulent keyword and the second weight of each fraudulent phrase; For each fraudulent text message, the degree of fraudulent activity is determined according to the following formula, based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight: in, This indicates the degree of fraud in the allegedly fraudulent text message. This indicates the number of the i-th fraudulent keyword in the fraudulent text message. This represents the first weight of the i-th fraudulent keyword. This indicates the number of the j-th fraudulent phrase in the fraudulent text message. The second weight of the j-th fraudulent phrase is represented, and n represents the total number of fraudulent keywords and fraudulent text messages in the fraudulent text messages. For each of the fraudulent text messages, based on the degree of fraud involved, the fraudulent text message is sent to a processing platform corresponding to the degree of fraud, and the processing platform is used to process the fraudulent text message using a target processing method; The step of obtaining a first weight for each fraudulent keyword and a second weight for each fraudulent phrase based on multiple fraudulent text messages within the target processing period includes: For each fraud-related keyword, if the fraud-related keyword does not exist in the vocabulary database, determine the first occurrence number of the fraud-related keyword within the target processing cycle, and determine the first weight of the fraud-related keyword based on the first occurrence number; if the fraud-related keyword exists in the vocabulary database, obtain the first weight of the fraud-related keyword from the vocabulary database. For each fraudulent phrase, if the fraudulent phrase does not exist in the vocabulary, determine the second occurrence number of the fraudulent phrase within the target processing cycle, and determine the second weight of the fraudulent phrase based on the second occurrence number; if the fraudulent phrase exists in the vocabulary, obtain the second weight of the fraudulent phrase from the vocabulary.
2. The method of claim 1, wherein, The process of identifying fraudulent keywords and fraudulent phrases including the fraudulent keywords in each of the fraudulent text messages includes: For each of the fraudulent text messages, extract the text content from the fraudulent text message; Extract verbs from the text content and match them with keywords in the vocabulary database. Verbs that match successfully are identified as the fraud-related keywords. The fraud-related keywords and their contextual information within the text content are identified as the fraud-related phrases.
3. The method of claim 2, wherein, The extraction of text content from the fraudulent text message includes: The text of the fraudulent text message is segmented according to the smallest unit to obtain multiple data contents corresponding to the fraudulent text message; Each data content is encoded using a Unicode to obtain a Unicode corresponding to each data content. In the text of the fraudulent text message, data content whose Unicode is outside the range of Unicode is removed to obtain the text content of the fraudulent text message.
4. The method of claim 2, wherein, After determining the fraudulent keywords and their contextual information within the text content as the fraudulent phrase, the method further includes: The word vectors of the fraudulent phrases are determined using a pre-trained language model; Based on the word vectors, determine the synonyms of the fraudulent phrases; Add the synonyms to the vocabulary database.
5. The method of claim 2, wherein, The step of identifying fraudulent keywords and fraudulent phrases including the fraudulent keywords in each of the fraudulent text messages also includes: When the number of fraudulent text messages exceeds a threshold within the target processing period, fraudulent text messages with duplicate content are obtained from the multiple fraudulent text messages. The threshold is determined based on fraudulent text message data from historical processing periods. Obtain the fraudulent keywords and fraudulent phrases that include the fraudulent keywords from the fraudulent text messages with repeated content.
6. The method of claim 5, wherein, Also includes: Fraudulent keywords and phrases obtained from repeated fraudulent text messages will be stored in the vocabulary database.
7. The method of claim 5, wherein, The step of obtaining fraudulent text messages with duplicate content from the multiple fraudulent text messages includes: A predetermined number of fraudulent text messages are randomly selected from the multiple fraudulent text messages; Calculate the hash value of the extracted fraudulent text messages, and identify fraudulent text messages with the same hash value as having duplicate content.
8. The method according to any one of claims 1 to 7, characterized in that, After determining the first weight of the fraudulent keyword based on the first occurrence frequency, the method further includes: The fraud-related keywords and the first weight are stored in the vocabulary database; After determining the second weight of the fraudulent phrase based on the second frequency of occurrence, the method further includes: The fraudulent phrase and the second weight are stored in the vocabulary database.
9. A fraud message processing apparatus, characterized by comprising: include: The SMS acquisition module is used to acquire multiple fraudulent SMS messages within a target processing period. The fraud-related vocabulary recognition module is used to identify fraud-related keywords and fraud-related phrases including the fraud-related keywords in each of the fraud-related text messages; The vocabulary weight acquisition module is used to acquire a first weight for each fraudulent keyword and a second weight for each fraudulent phrase based on multiple fraudulent text messages within the target processing period. The fraud severity determination module is used to determine the degree of fraud in each fraudulent text message according to the following formula: based on the fraudulent keywords, the first weight, the fraudulent phrase, and the second weight of the text message. in, This indicates the degree of fraud in the allegedly fraudulent text message. This indicates the number of the i-th fraudulent keyword in the fraudulent text message. This represents the first weight of the i-th fraudulent keyword. This indicates the number of the j-th fraudulent phrase in the fraudulent text message. The second weight of the j-th fraudulent phrase is represented, and n represents the total number of fraudulent keywords and fraudulent text messages in the fraudulent text messages. The fraudulent SMS distribution module is used to send each fraudulent SMS to a processing platform corresponding to the degree of fraud, based on the degree of fraud in the SMS. The processing platform is used to process the fraudulent SMS using a target processing method. Specifically, the vocabulary weight acquisition module is used for: For each fraud-related keyword, if the fraud-related keyword does not exist in the vocabulary database, determine the first occurrence number of the fraud-related keyword within the target processing cycle, and determine the first weight of the fraud-related keyword based on the first occurrence number; if the fraud-related keyword exists in the vocabulary database, obtain the first weight of the fraud-related keyword from the vocabulary database. For each fraudulent phrase, if the fraudulent phrase does not exist in the vocabulary, determine the second occurrence number of the fraudulent phrase within the target processing cycle, and determine the second weight of the fraudulent phrase based on the second occurrence number; if the fraudulent phrase exists in the vocabulary, obtain the second weight of the fraudulent phrase from the vocabulary.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the fraudulent SMS processing method according to any one of claims 1 to 8.
11. A computer readable storage medium having stored thereon a computer program, characterized in that, When the program is executed by the processor, it implements the fraudulent SMS processing method according to any one of claims 1 to 8.