Junk short message analysis method and device, equipment and storage medium
Through splicing and word segmentation processing, and using custom spam keyword databases for matching, the problem of insufficient spam keyword recognition in the existing technology is solved, and precise interception of spam messages and protection of user rights is achieved.
Patent Information
- Application Number
- CN202510056617.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively identify the split spam keywords, resulting in the fact that spam messages in long text messages cannot be accurately identified and the content in the signature cannot be effectively identified and judged.
Long short messages are formed by splicing based on the time sequence of short messages, access numbers and preset SMS lengths, and word segmentation is processed after the signature symbol is removed, and spam keyword matching is used using custom spam keyword database and HanLP word segmentation method.
It improves the accuracy of spam identification, realizes accurate interception of spam messages, protects users' rights and interests, improves user experience and satisfaction, and reduces the complaint rate of spam messages.
Smart Images

Figure CN120017624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and in particular to a method, device, equipment and storage medium for analyzing spam text messages. Background Art
[0002] With the rapid development of mobile Internet technology, SMS, as an important medium of communication, has been increasingly used. However, criminals have used SMS platforms and virtual operators to spread spam SMS, which has seriously disrupted the normal lives of users and even caused economic losses. Spam SMS includes advertising information and fraudulent SMS. In particular, the "106 SMS" code, which was originally designed for corporate services, has also been abused by some agents and has become one of the important sources of spam SMS. The low cost of sending spam SMS and its wide range of dissemination have been exploited by the black and gray industries, making it an important channel for spreading illegal fraud SMS.
[0003] The keyword filtering technology in the prior art is to scan and match the content of the text message based on a preset keyword list. Once a text message containing sensitive keywords is found, it is intercepted or marked. However, since the text message content exceeds 140 bytes (70 words, UCS2 encoding), it will be split into multiple short messages according to 134 bytes per message. The words located before and after the splitting node often cannot be identified after the junk keywords are disassembled due to the splitting, resulting in missed judgments when judging junk text messages based on the text message content. Therefore, it is impossible to effectively identify junk text messages in long text messages, so that junk text messages are eventually sent to the user's mobile phone terminal, causing complaints from terminal users and even bringing adverse effects to terminal users.
[0004] In addition, the current SMS signature relies on brackets [], which fail to identify and judge the content in the signature, allowing some agents to exploit the loopholes and achieve their own goals by hiding spam SMS content in the signature. In addition, the text in the signature and the text outside the signature can also form spam keywords, but the existing technology often automatically separates the text in the signature from the text outside the signature because of the brackets in the signature.
[0005] How to properly solve the problem of missing junk keywords through hidden distribution, specifically the distribution in two short messages and inside and outside the brackets of the signature, has become an urgent issue to be solved in the industry. Summary of the invention
[0006] The present invention provides a method, device, equipment and storage medium for analyzing spam text messages, which are used to effectively identify spam text message keywords in text message content, thereby improving the accuracy of identifying spam information, thereby achieving accurate interception of spam text messages.
[0007] According to a first aspect of the present invention, a method for analyzing spam text messages is provided, the method for analyzing spam text messages comprising:
[0008] A long text message consisting of at least two short text messages is spliced based on the time sequence of the short text messages, the access number and the preset text message length;
[0009] After removing the signature symbol in the long text message, the long text message is segmented;
[0010] Performing junk keyword matching on the long text message after word segmentation processing;
[0011] According to the matching result, it is analyzed whether the long text message consisting of at least two short text messages is a junk text message.
[0012] In one embodiment, the long text message composed of at least two short text messages based on the time sequence of the short text messages, the access number and the preset text message length includes:
[0013] Arrange the at least two short messages in chronological order according to the recording time information of the short messages in the time series database;
[0014] Analyzing the access number identifiers and access numbers in the continuous short messages to obtain the short messages constituting any initial long message;
[0015] Determine whether the last short message in the short messages of the initial long message is less than a preset message length or has a message termination mark;
[0016] When the above judgment is true, after removing the long and short header identifiers, the initial long text messages are spliced into one long text message in the chronological order;
[0017] When the above judgment is no, search for a short message with a corresponding access number identifier and access number until a short message shorter than a preset message length or with a message end identifier is found, and the initial long message and the found short message are concatenated into one long message in the chronological order.
[0018] In one embodiment, the long text message composed of at least two short text messages based on the time sequence of the short text messages, the access number and the preset text message length includes:
[0019] storing at least two short messages with long and short header identifiers in a time series database to arrange the at least two short messages with long and short header identifiers in chronological order;
[0020] Confirm the short message that is shorter than the preset length or has a message end mark as the last short message in the initial long message;
[0021] Analyze and obtain the short messages constituting any initial long short message according to the access number identifier and the access number in the last short message and the short message before the last short message that is equal to the preset short message length;
[0022] After removing the long and short header identifiers in the short messages constituting any initial long message, the short messages constituting any initial long message are spliced into one long message in the chronological order.
[0023] In one embodiment, performing junk keyword matching on the long text message after word segmentation processing includes:
[0024] Based on the HanLP word segmentation method, the custom word segmentation is loaded first to perform word segmentation on long text messages;
[0025] The long text message after word segmentation is matched with the preset junk keyword through the preset junk keyword matching model.
[0026] In one embodiment, it further includes:
[0027] Before performing word segmentation processing on the long text message based on the HanLP word segmentation method, a custom junk keyword library is established, and the custom junk keyword library is less than or equal to a preset keyword rule library;
[0028] The user-defined junk keyword library is preferentially used to match the long text message after the word segmentation processing.
[0029] In one embodiment, it further includes:
[0030] A bracket matching algorithm is used to determine the number, length and position of the signatures; if at least two signatures exist in the long text message, the long text message is judged to be a spam text message;
[0031] If the signature position in the long text message is not a preset position, the long text message is marked as a risky text message;
[0032] If the signature length in the long text message exceeds a preset length, the long text message is marked as a risky text message.
[0033] In one embodiment, it further includes:
[0034] Before splicing a long text message composed of at least two short text messages, it is determined whether the text message is a short text message according to the code of the long text message identifier in the text message.
[0035] According to a second aspect of the present invention, there is provided a device for analyzing spam text messages, comprising:
[0036] A splicing module, used for splicing a long text message consisting of at least two short text messages based on the time sequence of the short text messages, the access number and the preset text message length;
[0037] A word segmentation module, used for removing the signature symbol in the long text message and then performing word segmentation processing on the long text message;
[0038] A matching module, used for matching the long SMS after word segmentation with junk keywords;
[0039] The analysis module is used to analyze whether the long text message consisting of at least two short text messages is a junk text message according to the matching result.
[0040] In one embodiment, the splicing module, the word segmentation module, the matching module and the analysis module are controlled to implement any one of the above-mentioned spam text message analysis methods.
[0041] According to a third aspect of the present invention, there is provided an electronic device, the electronic device comprising: a communication interface, a processor, and a memory;
[0042] The memory is used to store program instructions, and when the program instructions are executed by the processor that is communicatively connected to the memory through the communication interface, any of the above-mentioned spam text message analysis methods is implemented.
[0043] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a computer (eg, a processor in the computer), any of the above-mentioned spam text message analysis methods is implemented.
[0044] In summary, the present invention provides a method and device for analyzing spam text messages, the method comprising: splicing a long text message consisting of at least two short text messages based on the time sequence of short text messages, access numbers and preset text message lengths; performing word segmentation processing on the long text message after removing the signature symbol in the long text message; performing spam keyword matching on the long text message after word segmentation processing; and analyzing whether the long text message consisting of at least two short text messages is a spam text message based on the matching result. The technical solution of the present application solves the problem that when splitting a long text message, spam keywords are split into two short text messages, resulting in the inability to identify spam keywords, and also solves the problem that the text in the signature and the regular text in the text message together constitute spam keywords. Effective identification of spam text message keywords in text message content is achieved, thereby improving the accuracy of identifying spam information, thereby achieving accurate interception of spam text messages, effectively protecting the rights and interests of users, improving user experience, and greatly reducing the complaint rate of spam text messages.
[0045] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0046] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 A flowchart of a method for analyzing spam text messages provided by an embodiment of the present invention;
[0049] Figure 2 A flowchart of step S11 of a method for analyzing spam text messages provided by an embodiment of the present invention;
[0050] Figure 3 A flowchart of step S13 of a method for analyzing spam text messages provided by an embodiment of the present invention;
[0051] Figure 4 A structural diagram of a junk text message analysis device provided by an embodiment of the present invention;
[0052] Figure 5 A structural diagram of an electronic device provided by an embodiment of the present invention;
[0053] Figure 6 A schematic diagram of a junk text message processing process provided by an embodiment of the present invention;
[0054] Figure 7 A spam text processing framework diagram provided by an embodiment of the present invention;
[0055] Figure 8 It is a schematic diagram of long text message identification of junk text messages provided by an embodiment of the present invention;
[0056] Fig. 9 It is a schematic diagram of signature extraction of spam text messages provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0058] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0059] like Figure 1 As shown, the present invention provides a method for analyzing junk text messages, and the method for analyzing junk text messages includes:
[0060] In step S11, a long message consisting of at least two short messages is spliced based on the time sequence of the short messages, the access number and the preset message length;
[0061] In step S12, after removing the signature symbol in the long message, the long message is segmented;
[0062] In the embodiment of the present application, only the signature symbol in the long message is removed, but the content of the signature is not removed, so as to match the signature with junk keywords, thereby avoiding the situation where the junk keywords in the signature cause missed recognition.
[0063] In step S13, junk keyword matching is performed on the long text message after word segmentation processing;
[0064] In step S14, based on the matching result, it is analyzed whether the long text message consisting of at least two short text messages is a spam text message.
[0065] In one embodiment, the time sequence of the short messages, the access number and the preset length of the short messages are obtained, and at least two short messages are spliced into a long message based on the time sequence of the short messages, the access number and the preset length of the short messages. Based on the access number identifier and the access number of the short messages, multiple short messages are merged into a complete long message content, such as Figure 8 After the concatenation is completed, the signature symbol in the long text message needs to be removed. The bracket matching algorithm is used to determine the number, length and position of the signature. The signature is usually placed in a specific position in the text message content and may contain special characters, such as square brackets "[ ]", such as Fig. 9 As shown. Removing these symbols is to make the subsequent word segmentation and spam keyword matching more accurate. After removing the signature symbol, the long text message content is segmented. Word segmentation is the process of dividing continuous text into meaningful word sequences, which is helpful for subsequent text analysis. After word segmentation, the long text message is matched with spam keywords. By comparing the word segmentation results with the preset spam keyword library, it is possible to identify whether the text message contains the characteristic words of spam text messages.
[0066] In addition, if there are at least two signatures in a long text message, the long text message is directly judged as a spam text message. The presence of multiple signatures in a text message is a feature of spam text messages. Regular text messages contain only one signature by default. If the signature position in a long text message is not the preset position, or the signature length exceeds the preset length, the long text message is marked as a risky text message. Although the signature position in the text message is not the regular position, or the signature length is too long, it cannot be directly determined to be a spam text message, but the possibility of spam text messages is high.
[0067] Before executing the above steps, it is necessary to determine whether the text message is a short text message based on the code of the long text message identifier in the text message. The long text message needs to be spliced and the signature symbol removed, while the short text message does not.
[0068] Figure 6The following is a schematic diagram of the processing flow of spam text messages. Starting from the reception of text messages, text messages are classified into two types: short text messages and long text messages. For long text messages, a splicing operation is performed to ensure the integrity of the text message content. After the text message content is complete, signature extraction is performed, which involves identifying and separating the signature part in the text message. The accuracy of signature extraction depends on the recognition of "[] characters", "signature length" and "signature position". The text message content enters the word segmentation stage, the purpose of which is to segment the continuous text into meaningful vocabulary units. In order to improve the accuracy of word segmentation, the data structure is dynamically loaded and processed in combination with the custom word library and word segmentation rules. After word segmentation is completed, the matching algorithm is used to compare the word segmentation results with the words in the database to identify the keywords or phrases in the text message. The spam text message is judged based on the results of the matching algorithm. If the text message is judged as spam, it will be marked and enter the further processing stage, including filtering or deletion.
[0069] Figure 7 This is a framework diagram for processing spam text messages. SMS messages are received through two service providers (SP1 and SP2). The core functional modules cover key links such as text recognition, signature extraction, long SMS splicing, and word segmentation processing. The above modules work together to conduct in-depth analysis and processing of SMS content, such as identifying spam SMS or extracting key information. The processed SMS messages are sent to mobile terminals through SMS centers or intercommunication gateways, ensuring that the SMS content received by users has been optimized and screened, which plays an important role in improving SMS processing efficiency, accuracy, and user experience.
[0070] Determine whether the SMS is a short SMS. If so, no subsequent processing is required. If it is a long SMS, concatenate it. Remove the signature symbol from the concatenated long SMS. Perform word segmentation on the long SMS after removing the signature symbol. Perform spam keyword matching on the long SMS after word segmentation. Determine whether the long SMS is a spam SMS or a risky SMS based on the matching result and the signature.
[0071] Table 1 below is the SMS call data structure, which includes six key fields: SMS access number (source), SMS sequence number (msg_id), SMS submission time (submit_time), SMS submission result (result), long SMS identifier (udhi) and SMS content (content). These fields record the source, unique identifier, submission time, result status, whether it is a long SMS and specific content of the SMS. For long SMS, there will be a special identifier and header information, which may produce garbled characters when displayed, but it does not affect the communication of the SMS body content.
[0072] Table 1
[0073]
[0074] The technical solution in this embodiment solves the problem that when splitting a long text message, spam keywords are split into two short text messages, resulting in the inability to identify spam keywords, and also solves the problem that the text in the signature and the regular text in the text message together constitute spam keywords. It realizes the effective identification of spam text message keywords in the text message content, thereby improving the accuracy of spam information identification, thereby realizing accurate interception of spam text messages, effectively protecting the rights and interests of users, improving user experience, and greatly reducing the complaint rate of spam text messages.
[0075] In one embodiment, the continuous short messages belonging to the same long message may be found first, and then the long message may be spliced based on the last short message. In actual application, if two long messages are received at the same time, the short messages split from the long message may be received alternately. Therefore, only when the last short message has a message end mark, it is confirmed that the long message splicing is completed.
[0076] like Figure 2 As shown, step S11 includes the following steps S21-S25:
[0077] In step S21, the at least two short messages are arranged in chronological order according to the recording time information of the short messages in the time series database;
[0078] In step S22, according to the access number identifier and the access number in the short message, the short message constituting any initial long message is analyzed and obtained;
[0079] In step S23, it is determined whether the last short message in the short messages of the initial long message is shorter than a preset message length or has a message termination mark;
[0080] In step S24, when the above judgment is true, after removing the long and short header identifiers, the initial long text messages are spliced into one long text message in the chronological order;
[0081] In step S25, when the above judgment is no, search for a short message with a corresponding access number identifier and access number until a short message shorter than a preset message length or with a message end identifier is found, and the initial long message and the found short message are spliced into one long message in the chronological order.
[0082] In one embodiment, multiple short messages are spliced into a complete long message through time sequence and specific identifiers. At least two short messages are arranged in chronological order according to the recording time information of the short messages in the time series database. Ensure that the splicing order of the short messages is correct, and ensure that the short messages are arranged in the actual time sequence of their sending and receiving when splicing, so as to restore the original order of the long messages. According to the access number identifier and access number in the short message, the short messages constituting the same initial long message are analyzed and obtained. There are cases where multiple users use the same access number, so the short messages belonging to the same initial long message are screened out through the access number identifier and access number of the short message, ensuring that the short messages of the same long message are spliced to avoid mixing with other short messages. Determine whether the last short message in the short messages constituting the initial long message is less than the preset message length or has an information termination identifier. Only when the last short message has an information termination identifier, it is confirmed that the splicing of the long message is completed. Short messages are usually sent in segments, and the last short message contains an termination identifier, indicating that all short messages have been sent, ensuring that the splicing operation is only performed when all short messages are completely received. After confirming that the last short message has a termination mark, remove the long message header mark (the control character in the header of the long message segment) and splice the short messages into a complete long message in chronological order. The purpose of removing the header mark of the long message is to ensure that the spliced text is correct, that is, the spliced text does not contain meaningless special characters.
[0083] In another embodiment, the last short message can be found first, and then the short messages before the last short message can be spliced into a long message. Specifically, step S11 can also include the following steps: at least two short messages with long and short header identifiers are stored in a time series database to arrange at least two short messages with long and short header identifiers in chronological order; a short message with a length less than a preset length or with an information termination identifier is confirmed as the last short message in the initial long message; according to the access number identifier and access number in the last short message and the short message before the last short message that is equal to the preset message length, the short messages constituting any initial long message are analyzed to obtain; after removing the long and short header identifiers in the short messages constituting any initial long message, the short messages constituting any initial long message are spliced into a long message in chronological order.
[0084] During the SMS sending process, short SMS may be segmented and sent at different times. Therefore, by arranging them in time sequence, we can ensure that the SMS splicing is in the order of sending and restore the original complete SMS content. The access number identifier and access number are used to identify which short SMS belong to the same long SMS to prevent the content that does not belong to the same long SMS from being spliced together. The termination identifier ensures that all short SMS have been received. If there is no such identifier, the long SMS has not been fully received and therefore cannot be spliced. During the final splicing, the header identifier of the long SMS must be removed to ensure that the spliced SMS is a continuous and complete text without identifier interference.
[0085] Perform time sorting to ensure that short messages are arranged in the order of sending; use access codes and access numbers to filter out short messages belonging to the same long message; determine whether the last short message has a termination mark to ensure that the splicing operation is only performed after receiving the complete message; splice short messages, remove useless long message header control identifiers, and restore the complete message text. The technical solution in this embodiment can ensure the correct splicing of short messages, and avoid erroneous operations through various condition judgments, ensuring that the spliced long messages are complete and correct.
[0086] In one embodiment, Figure 3 As shown, step S13 includes the following steps S31-S34:
[0087] In step S31, before segmenting the long text message based on the HanLP segmentation method, a custom junk keyword library is established, and the custom junk keyword library is less than or equal to the preset keyword rule library;
[0088] In step S32, the user-defined junk keyword library is preferentially used to match the long text message after the word segmentation processing;
[0089] In step S33, the custom word segmentation is preferentially loaded based on the HanLP word segmentation method, and the long text message is segmented;
[0090] In step S34, the long text message after word segmentation is matched with the preset junk keyword through the preset junk keyword matching model.
[0091] In one embodiment, the HanLP word segmentation method is used for word segmentation, a keyword library is established, and the keyword library is used for spam SMS matching. The long SMS is segmented and matched using a custom spam keyword library to identify whether it is a spam SMS. Before the long SMS is segmented, a custom spam keyword library is established. The size of the custom spam keyword library is less than or equal to the preset keyword rule library, and is used as a keyword library for subsequent matching. The custom spam keyword library can avoid the preset keyword rule library being too large to affect the processing efficiency. At the same time, the custom library is more flexible and specific than the preset keyword library, and is suitable for spam SMS discrimination in specific scenarios. The HanLP word segmentation method is used for long SMS word segmentation. During the word segmentation process, HanLP preferentially loads the custom word segmentation to ensure that specific spam keywords can be effectively processed. The long SMS is divided into smaller text units for subsequent keyword matching. The priority loading of the custom word segmentation is to adapt to special words or terms in specific business scenarios, thereby improving the accuracy of word segmentation. The preset spam keyword matching model is used to match the long SMS after word segmentation with the spam keyword. In the matching process, the custom spam keyword library is used first to match the segmented long text messages. The advantage of using the custom library first to identify the characteristics of spam text messages through keyword matching is that it is optimized for specific business scenarios and can detect potential spam text messages more efficiently.
[0092] Create a custom junk keyword library for word segmentation and then perform keyword matching. Custom word segmentation and custom junk keyword library are adaptable to specific scenarios. The size of the custom junk keyword library is set to be less than or equal to the preset keyword rule library. The custom library is a simplification or supplement to the preset library, which is more targeted and efficient. Prioritize loading the custom word library for word segmentation and use the custom keyword library for matching.
[0093] Before matching, long text messages are segmented. The segmentation process divides the original text message into smaller units for better keyword matching. The quality of the segmentation directly affects the effect of keyword matching, so it is preferred to load the custom word library during the segmentation stage to improve the accuracy of the segmentation and ensure more effective keyword matching.
[0094] The establishment of a custom spam keyword library is the preparatory work for the entire process, which determines the specificity and accuracy of the matching model. The priority loading of the custom word library for HanLP word segmentation is to ensure that specific words can be correctly identified, further improving the accuracy and speed of matching. The priority use of the custom keyword library for matching is to make full use of the previously established custom spam keyword library. The improvement of the matching priority can make the system more effective in detecting specific types of spam text messages. The word segmentation and spam keyword matching operations for long text messages can be performed efficiently and specifically, especially by giving priority to the use of custom word libraries to optimize the entire spam text message identification process.
[0095] In one embodiment, Figure 4 FIG. 1 is a block diagram of a junk text message analysis device according to an exemplary embodiment. Figure 4 As shown, the junk text message analysis device includes a splicing module 41, a word segmentation module 42, a matching module 43 and an analysis module 44.
[0096] The splicing module 41 is used to splice a long text message consisting of at least two short text messages based on the time sequence of the short text messages, the access number and the preset text message length;
[0097] The word segmentation module 42 is used to remove the signature symbol in the long text message and then perform word segmentation on the long text message;
[0098] The matching module 43 is used to match the long text message after the word segmentation process with junk keywords;
[0099] The analysis module 44 is used to analyze whether the long text message consisting of at least two short text messages is a junk text message based on the matching result.
[0100] The block diagram of the device for analyzing spam text messages includes the splicing module 41 , the word segmentation module 42 , the matching module 43 and the analyzing module 44 , which are controlled to execute the spam text message analyzing method described in any of the above embodiments.
[0101] like Figure 5 As shown, the present invention provides an electronic device 500, the electronic device comprising: a communication interface, a processor 501, and a memory 502;
[0102] Among them, the memory 502 is used to store program instructions. When the program instructions are executed by the processor 501 that is communicatively connected to the memory 502 through the communication interface, a long text message consisting of at least two short text messages is spliced based on the chronological order of the short text messages, the access number and the preset text message length; after removing the signature symbol in the long text message, the long text message is segmented; the long text message after the word segmentation is matched with junk keywords; and according to the matching result, whether the long text message consisting of at least two short text messages is a junk text message is analyzed.
[0103] The present invention provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, a long text message consisting of at least two short text messages is spliced based on the time sequence of the short text messages, access numbers and preset text message lengths; after removing the signature symbols in the long text message, the long text message is segmented; junk keyword matching is performed on the long text message after the word segmentation processing; and according to the matching result, whether the long text message consisting of at least two short text messages is a junk text message is analyzed.
[0104] It should be understood that the specific features, operations and details described hereinabove about the method of the present invention may also be similarly applied to the device and system of the present invention, or, vice versa. In addition, each step of the method of the present invention described above may be performed by the corresponding parts or units of the device or system of the present invention.
[0105] It should be understood that each module / unit of the device of the present invention can be implemented in whole or in part by software, hardware, firmware or a combination thereof. Each module / unit can be embedded in the processor of the computer device in the form of hardware or firmware or independent of the processor, or can be stored in the memory of the computer device in the form of software for the processor to call to perform the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.
[0106] In one embodiment, a computer device is provided, which includes a memory and a processor, and the memory stores computer instructions executable by the processor, and the computer instructions instruct the processor to execute each step of the method of the embodiment of the present invention when executed by the processor. The computer device can be a server, a terminal, or any other electronic device with necessary computing and / or processing capabilities in a broad sense. In one embodiment, the computer device may include a processor, a memory, a network interface, a communication interface, etc. connected through a system bus. The processor of the computer device can be used to provide necessary computing, processing and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and an internal memory. An operating system, a computer program, etc. may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the computer device can be used to connect and communicate with external devices through a network. The steps of the method of the present invention are executed by the processor.
[0107] The present invention may be implemented as a computer-readable storage medium having a computer program stored thereon, which causes the steps of the method of an embodiment of the present invention to be executed when executed by a processor. In one embodiment, the computer program is distributed on a plurality of computer devices or processors coupled to a network so that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be performed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be performed by one or more computer devices or processors, and one or more other method steps / operations may be performed by one or more other computer devices or processors. One or more computer devices or processors may perform a single method step / operation, or perform two or more method steps / operations.
[0108] It can be understood by a person skilled in the art that the method steps of the present invention can be completed by instructing related hardware such as a computer device or a processor through a computer program, and the computer program can be stored in a non-temporary computer-readable storage medium, and the steps of the present invention are executed when the computer program is executed. Depending on the circumstances, any reference to memory, storage, database or other media in this article may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (PROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0109] The various technical features described above can be combined arbitrarily. Although all possible combinations of these technical features are not described, any combination of these technical features should be considered to be covered by this specification as long as there is no contradiction in such combination.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for analyzing spam text messages, characterized in that: include: A long text message consisting of at least two short text messages is spliced based on the time sequence of the short text messages, the access number and the preset text message length; After removing the signature symbol in the long text message, the long text message is segmented; Performing junk keyword matching on the long text message after word segmentation processing; According to the matching result, it is analyzed whether the long text message consisting of at least two short text messages is a junk text message.
2. The method for analyzing spam text messages according to claim 1, characterized in that: The long text message composed of at least two short text messages spliced based on the time sequence of the short text messages, the access number and the preset text message length includes: Arrange the at least two short messages in chronological order according to the recording time information of the short messages in the time series database; Analyzing the access number identifiers and access numbers in the continuous short messages to obtain the short messages constituting any initial long message; Determine whether the last short message in the short messages of the initial long message is less than a preset message length or has a message termination mark; When the above judgment is yes, after removing the long and short header identifiers, the initial long text messages are spliced into one long text message in the chronological order; When the above judgment is no, search for a short message with a corresponding access number identifier and access number until a short message shorter than a preset message length or with a message end identifier is found, and the initial long message and the found short message are concatenated into one long message in the chronological order.
3. The method for analyzing spam text messages according to claim 1, characterized in that: The long text message composed of at least two short text messages spliced based on the time sequence of the short text messages, the access number and the preset text message length includes: storing at least two short messages with long and short header identifiers in a time series database to arrange the at least two short messages with long and short header identifiers in chronological order; Confirm the short message that is shorter than the preset length or has a message end mark as the last short message in the initial long message; Analyze and obtain the short messages constituting any initial long short message according to the access number identifier and the access number in the last short message and the short message before the last short message that is equal to the preset short message length; After removing the long and short header identifiers in the short messages constituting any initial long message, the short messages constituting any initial long message are spliced into one long message in the chronological order.
4. The method for analyzing spam text messages according to claim 1, characterized in that: The performing junk keyword matching on the long text message after the word segmentation process includes: Based on the HanLP word segmentation method, the custom word segmentation is loaded first to perform word segmentation on long text messages; The long text message after word segmentation is matched with the preset junk keyword through the preset junk keyword matching model.
5. The method for analyzing spam text messages according to claim 4, characterized in that: Also includes: Before performing word segmentation processing on the long text message based on the HanLP word segmentation method, a custom junk keyword library is established, and the custom junk keyword library is less than or equal to a preset keyword rule library; The user-defined junk keyword library is preferentially used to match the long text message after the word segmentation processing.
6. The method for analyzing spam text messages according to claim 1, characterized in that: Also includes: A bracket matching algorithm is used to determine the number, length and position of signatures; If there are at least two signatures in the long text message, the long text message is judged to be a spam text message; If the signature position in the long text message is not a preset position, the long text message is marked as a risky text message; If the signature length in the long text message exceeds the preset signature length, the long text message is marked as a risky text message.
7. The method for analyzing spam text messages according to claim 1, characterized in that: Also includes: Before splicing a long text message composed of at least two short text messages, it is determined whether the text message is a short text message according to the code of the long text message identifier in the text message.
8. A device for analyzing spam text messages, characterized in that: include: A splicing module, used for splicing a long text message consisting of at least two short text messages based on the time sequence of the short text messages, the access number and the preset text message length; A word segmentation module, used for removing the signature symbol in the long text message and then performing word segmentation processing on the long text message; A matching module, used for matching the long SMS after word segmentation with junk keywords; The analysis module is used to analyze whether the long text message consisting of at least two short text messages is a junk text message according to the matching result.
9. The junk text message analysis device according to claim 8, characterized in that: The splicing module, the word segmentation module, the matching module and the analysis module are controlled to execute the spam text message analysis method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: Communication interface, processor, memory; The memory is used to store program instructions, and when the program instructions are executed by the processor that is communicatively connected to the memory through the communication interface, the electronic device implements the spam text message analysis method according to any one of claims 1 to 7.