Spam message interception method and apparatus, electronic device, and computer readable medium
By acquiring and detecting the text content of the message to be tested, and combining the keywords, hash values, and large language models of historical spam messages, the problem of low accuracy and efficiency in spam message detection is solved, and more efficient spam message interception is achieved.
Patent Information
- Application Number
- CN202411061418.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-08-02
AI Technical Summary
Existing technologies for spam message detection have low accuracy and efficiency, making it difficult to balance false positives and false negatives, and the sample size is insufficient.
By obtaining the text content of the message to be tested, and combining it with the keywords, hash values and large language models of historical spam messages, it is determined whether the message to be tested is spam, and it is blocked when it is determined to be spam.
It improves the accuracy of spam message detection, reduces false positive and false negative rates, solves the problem of insufficient samples, improves detection efficiency, and eliminates the need for manual analysis.
Smart Images

Figure CN119172754B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of computer technology, in particular to a spam message interception method and device, an electronic device and a computer readable medium. BACKGROUND
[0002] With the rapid development of mobile Internet, the problem of spam messages is becoming increasingly serious, which brings many troubles to users. Therefore, it is necessary to detect and intercept spam messages in a timely manner.
[0003] In the prior art, spam message detection technology mainly includes rule-based detection and machine learning-based detection. However, due to the diversification of spam message forms, the lack of sample size of spam messages and other problems, the traditional detection method is difficult to balance false positives and false negatives, and further analysis and judgment are required by manual work, resulting in low accuracy and efficiency of spam message detection. SUMMARY
[0004] Embodiments of the present application provide a spam message interception method and device, an electronic device and a computer readable medium to solve the technical problem of low accuracy and efficiency of spam message detection in the prior art.
[0005] In a first aspect, the embodiments of the present application provide a spam message interception method, which includes: when a message sending request is received, obtaining a to-be-tested message in the message sending request; extracting text content in the to-be-tested message; detecting the text content based on keywords of historical spam messages, hash values of historical spam messages and a large language model to determine whether the to-be-tested message is a spam message; and in response to the to-be-tested message being a spam message, intercepting the to-be-tested message.
[0006] In a second aspect, the embodiments of the present application provide a spam message interception device, which includes: an obtaining unit configured to, when a message sending request is received, obtain a to-be-tested message in the message sending request; an extracting unit configured to extract text content in the to-be-tested message; a detection unit configured to detect the text content based on keywords of historical spam messages, hash values of historical spam messages and a large language model to determine whether the to-be-tested message is a spam message; and an interception unit configured to, in response to the to-be-tested message being a spam message, intercept the to-be-tested message.
[0007] In a third aspect, the embodiments of the present application provide an electronic device, which includes: one or more processors; and a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the embodiments of the first aspect.
[0008] In a fourth aspect, the embodiments of the present application provide a computer readable medium, which stores a computer program. The computer program is executed by a processor to implement the method described in any of the embodiments of the first aspect.
[0009] The method, device, electronic device and computer readable medium provided by the embodiments of the present application can, when receiving a message sending request, first extract the text content in the to-be-tested message in the message sending request, and then detect the text content based on the keywords of the historical spam messages, the hash values of the historical spam messages and the large language model, to determine whether the to-be-tested message is a spam message, so as to intercept the to-be-tested message when the to-be-tested message is a spam message. Since the keywords and hash values of the historical spam messages can accurately determine whether the to-be-tested message is a historical spam message that has occurred before, and the large language model can more accurately capture the semantic features of the spam message, it can accurately determine the spam message that has never occurred before. Therefore, the above method can cope with the problem of diversified forms of spam messages, reduce the false positive rate and the false negative rate of the spam messages, and improve the accuracy of the spam message detection. In addition, the above method can update and supplement the samples according to the detection results, so as to cope with the problem of insufficient samples, and further improve the accuracy of the spam message detection. At the same time, without manual assistance, the efficiency of the spam message detection is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] Other characteristics, objects and advantages of the present application will become more apparent from the following detailed description of non-restrictive embodiments, made with reference to the attached drawings:
[0011] Figure 1 is a flowchart of an embodiment of the spam message interception method of the present application;
[0012] Figure 2 is a schematic diagram of an application scenario of the spam message interception method of the present application;
[0013] Figure 3 is a structural schematic diagram of an embodiment of the spam message interception device of the present application;
[0014] Figure 4 is a structural schematic diagram of an electronic device for implementing the embodiments of the present application. DETAILED DESCRIPTION
[0015] All actions of obtaining signals, information or data in the present application are performed in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization of the owner of the corresponding device.
[0016] The application will be described in further detail below with reference to the drawings and embodiments. It is to be understood that the specific embodiments described herein are merely illustrative of the application and are not intended to limit the application. In addition, it should be noted that, for the purpose of clarity, only the parts of the drawings that are relevant to the application are shown.
[0017] It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict. The application will be described in further detail below with reference to the drawings and embodiments.
[0018] Reference is made to Figure 1 which shows a flow 100 of an embodiment of a spam message interception method according to the present application. The spam message interception method can be applied to various electronic devices with data processing functions. For example, the electronic devices can include, but are not limited to, cloud servers, physical servers, etc. The execution subject of the spam message interception method can be a processor in the electronic devices.
[0019] The spam message interception method includes the following steps:
[0020] Step 101, upon receiving a message sending request, obtaining a to-be-tested message in the message sending request.
[0021] In the embodiment, the message sending request can include the to-be-tested message. The to-be-tested message can be various types of messages, for example, can include but is not limited to short messages, multimedia messages, instant messaging messages, etc.
[0022] In some optional implementations of the embodiment, the spam message interception method can be applied to a first short message center. The to-be-tested message is a short message. The message sending request is sent by a sending terminal to a first mobile device switching center via a sending terminal base station, and is sent by the first mobile device switching center to the first short message center. The first short message center is a short message center (SMSC) of a first operator where the sending terminal is located. The first mobile device switching center is a mobile device switching center (MSC) maintained by the first operator where the sending terminal is located. In practice, the short message center is a network element in a mobile communication network, which is responsible for processing and storing short message service (SMS) messages. The functions of the short message center include routing, forwarding, storing and backing up short messages. The mobile device switching center is a key component in a traditional mobile phone network, and is very important in second-generation and third-generation networks. The mobile device switching center is responsible for processing the routing and switching of mobile phone calls to ensure that calls can be connected within the network or across networks.
[0023] On this basis, in some optional implementations of the embodiment, in response to the to-be-tested message being a normal message, the message sending request is sent to a second short message center, the message sending request is sent to a second mobile device switching center through the second short message center, and the to-be-tested message is sent to a receiving terminal through the second mobile device switching center via a receiving base station; wherein the second short message center is a short message center of a second operator to which the receiving terminal belongs, and the second mobile device switching center is a mobile device switching center maintained by the second operator.
[0024] Step 102, extracting the text content in the to-be-tested message.
[0025] In the embodiment, the execution subject can extract the text content in the to-be-tested message based on the type of the to-be-tested message by using a corresponding text extraction technology.
[0026] In some optional implementations of the embodiment, the extraction of the text content in the to-be-tested message can specifically include: in the case that the to-be-tested message contains text, extracting the text; in the case that the to-be-tested message contains an image, extracting text information from the image by using an optical character recognition technology (OCR); and splicing the text and the text information to generate the text content of the to-be-tested message.
[0027] Specifically, when the text information is extracted from the image by using the optical character recognition technology, the image can be first subjected to brightness detection to detect the dark and bright patterns of multiple regions of the image, and then the character shape is determined. Subsequently, the character shape can be translated into computer text by using a character recognition method (such as a Euclidean space comparison method, a dynamic program comparison method, a neural network-based character comparison method, etc.). In this way, the text information in the image can be obtained.
[0028] It should be noted that if the to-be-tested message includes a video, the optical character recognition technology can be used to extract text information from each video frame, and the extracted text information can be de-duplicated to splice the text content of the to-be-tested message.
[0029] In this way, when the to-be-tested message contains different types of content, the text in each type of content can be completely extracted to obtain the complete text content in the to-be-tested message. Thus, the comprehensiveness and completeness of the text content extraction are improved, which helps to improve the accuracy of the spam message detection.
[0030] Step 103, detecting the text content based on the keywords of the historical spam messages, the hash values of the historical spam messages, and the large language model to determine whether the to-be-tested message is a spam message.
[0031] In this embodiment, the historical spam messages can include, but are not limited to, spam messages detected in various ways within a set historical period (e.g., within the last 1 year, within the last 3 years, within the last one month). The historical spam messages can include multiple messages. The message content of the historical spam messages can include, but is not limited to, promotional information, sensitive information, etc.
[0032] In this embodiment, for each historical spam message, the above execution subject can match the historical spam message and the above text content based on the keywords of the historical spam message to obtain a detection result, which can represent the text matching degree of the text content and the historical spam message. At the same time, the hash value of the historical spam message can be matched with the hash value of the text content to obtain a second detection result, which can represent the similarity of the text content and the historical spam message. At the same time, the above text content can be input into a large language model to obtain a third detection result output by the large language model. Further, the three detection results are combined to determine whether the message to be detected is a spam message.
[0033] The hash value can be calculated using a hash algorithm. The large language model (LLM) is a deep learning model trained based on massive text data. It can not only generate natural language text, but also deeply understand the meaning of the text and process various natural language tasks such as text summarization, question answering, translation, etc. In practical applications, the large language model can be trained based on existing language models or self-developed language models, which are not limited here. As an example, the large language model can use ChatGPT (Chat Generative Pre-trained Transformer) model. Compared with traditional classifiers, using large language models can improve the understanding and generation of natural language, thereby improving the accuracy of spam message detection.
[0034] In some optional implementations of the embodiment, the hash value of the text content and the historical spam message and the similarity determined based on the hash value can be calculated by SSDeep (Similarity Digest Algorithm). SSDeep can be used to compare and identify the similarity of files. SSDeep can divide a file into multiple fixed-size blocks based on a sliding window technique, calculate local hash values for each block, and connect these local hash values to form a global hash value of the file. For each historical spam message, the hash value of the historical spam message is the global hash value calculated by SSDeep for the historical spam message. Unlike traditional hash algorithms, SSDeep can tolerate a certain degree of modification and change in the file, making the detection of spam messages more accurate.
[0035] In some optional implementations of the embodiment, step 103 is specifically performed according to the following steps:
[0036] First, based on the keywords of the historical spam messages, a regular expression is generated, and the regular expression is used to perform regular matching on the text content to obtain a regular matching result (denoted as R r ). In practical applications, a regular expression can be generated based on the keywords of each historical spam message. The text content of the test message can be matched based on the regular expression of each historical spam message. If any regular expression is matched, the regular matching result can be 1; if none of the regular expressions is matched, the regular matching result can be 0.
[0037] Second, the hash value of the text content is determined, and based on the hash value of the text content and the hash value of the historical spam message, the maximum similarity (denoted as R s ) between the text content and the historical spam message is determined. In practical applications, the hash value of the text content of the test message and the hash value of each historical spam message can be calculated for similarity, and thus the maximum value of the similarity calculation result is selected as the maximum similarity. In practical applications, the SSDeep algorithm described above can be used to calculate the hash value and the similarity.
[0038] Third, based on the text content and a preset scoring rule, model input information is generated, the model input information is input into a large language model, and a first score (denoted as R l ) of the text content is obtained.
[0039] As an example, the model input information can be as follows:
[0040] You are now an experienced spam message detection expert. I will now provide you with the message to be tested, and you need to give me the conclusion whether the message to be tested is a spam message. You need to score the content to be tested, the range is 0-10 points, if you are sure that it is a spam message, the score is 10 points, if you are sure that it is not a spam message, score 0. For non-spam messages, the output example is as follows: {"level": 0, "description": "not a spam message, normal message content"}. For spam messages, the output example is as follows: {"level": 10, "description": "is a spam message, the judgment is based on ${answer}"}. The content of the message to be tested is: ${content}.
[0041] It should be noted that the large language model can be pre-trained based on historical spam messages to learn the characteristics of historical spam messages, so as to accurately determine whether the message to be tested is a spam message.
[0042] Fourthly, based on the regular matching result, the maximum similarity and the first score, it is determined whether the message to be tested is a spam message.
[0043] Optionally, the following steps can be performed: first, the average value of the regular matching result and the maximum similarity (i.e. (R r +R s ) / 2) is determined; then, the first score is normalized to obtain the second score (i.e. R l / 10); then, the average value and the second score are weighted and summed to obtain the third score; finally, based on the third score and a preset threshold (which can be denoted as R t ), it is determined whether the message to be tested is a spam message. The third score calculation formula is as follows:
[0044]
[0045] Wherein, D is a balance parameter, i.e. a weight coefficient, which can be set according to needs. As an example, its default value can be set to 0.4.
[0046] The above-mentioned preset threshold can be set according to needs, for example, it can be set to 0.8. If R>R t , it can be determined that the message to be tested is a spam message; otherwise, if R≤R t , it can be determined that the message to be tested is a normal message.
[0047] Step 104, in response to the message to be tested being a spam message, the message to be tested is intercepted.
[0048] In this embodiment, in response to the message to be tested being a spam message, the execution subject can intercept the message to be tested, thereby bringing a more pure communication environment for the user.
[0049] In some optional implementations of this embodiment, after intercepting the message to be tested, the text content can be recorded as historical spam messages, and the keywords and hash values of the text content can be stored. This allows for real-time updates of historical spam messages, increasing the sample size and timeliness of historical spam messages, thereby improving the effectiveness and accuracy of spam message detection.
[0050] The following describes this solution using a specific application scenario as an example. (See also...) Figure 2 This application embodiment can be applied to SMS sending scenarios. All messages in this embodiment can be SMS messages, and the executing entity can be the SMS center (i.e., the first SMS center) of the operator where the sending terminal (i.e., the terminal sending the SMS) is located. First, the sending terminal sends a message sending request to the sending base station. Then, the sending base station sends the message sending request to the first mobile device switching center maintained by the first operator where the sending terminal is located. Afterwards, the first mobile device switching center forwards the message sending request to the first SMS center of the first operator. For the SMS message in the message sending request, the first SMS center first extracts the text content, and then, based on keywords of historical spam SMS messages, hash values of historical spam SMS messages, and a large language model, detects the text content to determine whether the SMS message is spam. If it is spam, it is intercepted, and its relevant information is recorded. Conversely, if it is a normal SMS message, the message sending request is passed to the second SMS center of the second operator where the receiving terminal (i.e., the terminal receiving the SMS) is located. After receiving the message sending request, the second SMS center can send the SMS message to the second mobile device switching center maintained by the second operator, and then the receiving base station sends the SMS message to the receiving terminal.
[0051] The method provided in the above embodiments of this application, upon receiving a message sending request, first extracts the text content of the message to be tested from the message sending request. Then, based on keywords of historical spam messages, hash values of historical spam messages, and a large language model, the text content is detected to determine whether the message to be tested is spam, thereby intercepting it if it is. Since the keywords and hash values of historical spam messages can accurately determine whether the message to be tested is a historical spam message, and the large language model can more accurately capture the semantic features of spam messages, it can accurately identify spam messages that have not appeared in the past. Therefore, the above method can address the problem of diverse forms of spam messages, reduce the false positive rate and false negative rate of spam messages, and improve the accuracy of spam message detection. Furthermore, the above method can update and supplement samples based on the detection results, thereby addressing the problem of insufficient samples and further improving the accuracy of spam message detection. At the same time, no manual analysis is required, improving the efficiency of spam message detection.
[0052] Further reference Figure 3 As an implementation of the methods shown in the figures, this application provides an embodiment of a spam message interception device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0053] like Figure 3 As shown, the spam message interception device 300 of this embodiment includes: an acquisition unit 301, used to acquire a message to be tested in a message sending request when a message sending request is received; an extraction unit 302, used to extract the text content in the message to be tested; a detection unit 303, used to detect the text content based on keywords of historical spam messages, hash values of historical spam messages, and a large language model, to determine whether the message to be tested is a spam message; and an interception unit 304, used to intercept the message to be tested if it is a spam message.
[0054] In some optional implementations of this embodiment, the extraction unit 302 is further configured to: extract the text if the message to be tested contains text; extract text information from the image if the message to be tested contains an image; and concatenate the text and the text information to generate the text content of the message to be tested.
[0055] In some optional implementations of this embodiment, the detection unit 303 is further configured to: generate a regular expression based on keywords of historical spam messages; perform regular expression matching on the text content based on the regular expression to obtain a regular expression matching result; determine the hash value of the text content; determine the maximum similarity between the text content and the historical spam messages based on the hash value of the text content and the hash value of the historical spam messages; generate model input information based on the text content and preset scoring rules; input the model input information into a large language model to obtain a first score for the text content; and determine whether the message to be tested is spam based on the regular expression matching result, the maximum similarity, and the first score.
[0056] In some optional implementations of this embodiment, the detection unit 303 is further configured to: determine the average value of the regular expression matching result and the maximum similarity; normalize the first score to obtain a second score; perform a weighted summation of the average value and the second score to obtain a third score; and determine whether the message to be tested is a spam message based on the third score and a preset threshold.
[0057] In some optional implementations of this embodiment, the device further includes a storage unit for recording the text content as historical spam after intercepting the message to be tested, and storing the keywords and hash values of the text content.
[0058] In some optional implementations of this embodiment, the method is applied to a first SMS center; the message to be tested is an SMS message; the message sending request is sent by the sending terminal to the first mobile device switching center via the sending terminal base station, and then sent by the first mobile device switching center to the first SMS center; wherein, the first SMS center is the SMS center of the first operator where the sending terminal is located, and the first mobile device switching center is the mobile device switching center maintained by the first operator.
[0059] In some optional implementations of this embodiment, the apparatus further includes a sending unit, configured to: in response to the message to be tested being a normal message, send the message sending request to a second SMS center, send the message sending request to a second mobile device switching center through the second SMS center, and send the message to be tested to a receiving terminal via a receiving base station through the second mobile device switching center; wherein, the second SMS center is the SMS center of the second operator where the receiving terminal is located, and the second mobile device switching center is a mobile device switching center maintained by the second operator.
[0060] The apparatus provided in the embodiments of this application, upon receiving a message sending request, first extracts the text content of the message to be tested from the message sending request. Then, based on keywords of historical spam messages, hash values of historical spam messages, and a large language model, it detects the text content to determine whether the message to be tested is spam, thereby intercepting it if it is. Because the keywords and hash values of historical spam messages can accurately determine whether the message to be tested is a historical spam message, and the large language model can more accurately capture the semantic features of spam messages, it can accurately identify spam messages that have not appeared in the past. Therefore, this method can address the problem of diverse spam message formats, reduce the false positive rate and false negative rate of spam messages, and improve the accuracy of spam message detection. Furthermore, this method can update and supplement samples based on the detection results, thereby addressing the problem of insufficient samples and further improving the accuracy of spam message detection. At the same time, it eliminates the need for manual analysis, improving the efficiency of spam message detection.
[0061] The following is for reference. Figure 4 It shows a schematic diagram of the structure of an electronic device used to implement some embodiments of this application. Figure 4The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.
[0062] like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0063] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, disks, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.
[0064] In particular, according to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined above in the methods of some embodiments of this application.
[0065] It should be noted that the computer-readable medium described in some embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0066] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0067] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: upon receiving a message sending request, obtain the message to be tested from the message sending request; extract the text content from the message to be tested; detect the text content based on keywords of historical spam messages, hash values of historical spam messages, and a large language model to determine whether the message to be tested is spam; and, in response to the message to be tested being spam, intercept the message to be tested.
[0068] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++; and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, or it can be connected to an external computer (e.g., via the Internet using an Internet service provider), including local area networks (LANs) or wide area networks (WANs).
[0069] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0070] The units described in some embodiments of this application can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first determining unit, a second determining unit, a selecting unit, and a third determining unit. The names of these units do not necessarily limit the specific unit itself.
[0071] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0072] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.
Claims
1. A method for intercepting spam messages, characterized in that, The method includes: Upon receiving a message sending request, obtain the message to be tested from the message sending request; Extract the text content from the message to be tested; Based on keywords of historical spam messages, hash values of historical spam messages, and a large language model, the text content is detected to determine whether the message to be tested is spam. If the message to be tested is found to be spam, then the message to be tested is intercepted. The method of detecting the text content based on keywords of historical spam messages, hash values of historical spam messages, and a large language model to determine whether the message to be tested is spam includes: Based on keywords from historical spam messages, a regular expression is generated, and the text content is matched using the regular expression to obtain the regular expression matching result. Determine the hash value of the text content, and based on the hash value of the text content and the hash value of historical spam messages, determine the maximum similarity between the text content and the historical spam messages; Based on the text content and the preset scoring rules, model input information is generated, and the model input information is input into the large language model to obtain the first score of the text content; Based on the regular expression matching result, the maximum similarity, and the first score, it is determined whether the message to be tested is spam. The step of determining whether the message to be tested is spam based on the regular expression matching result, the maximum similarity, and the first score includes: Determine the average of the regular expression matching results and the maximum similarity; The first score is normalized to obtain the second score; The third score is obtained by weighted summation of the average value and the second score; Based on the third score and the preset threshold, it is determined whether the message to be tested is spam.
2. The method according to claim 1, characterized in that, The extraction of text content from the message to be tested includes: If the message to be tested contains text, extract the text; When the message to be tested contains an image, text information is extracted from the image using optical character recognition technology. The text and the text information are concatenated to generate the text content of the message to be tested.
3. The method according to claim 1, characterized in that, After intercepting the message to be tested, the method further includes: The text content is recorded as historical spam messages, and the keywords and hash values of the text content are stored.
4. The method according to any one of claims 1-3, characterized in that, The method is applied to a first SMS center; the message to be tested is an SMS message; the message sending request is sent by the sending terminal to the first mobile device switching center via the sending terminal base station, and then sent by the first mobile device switching center to the first SMS center. Wherein, the first SMS center is the SMS center of the first operator where the sending terminal is located, and the first mobile device exchange center is the mobile device exchange center maintained by the first operator.
5. The method according to claim 4, characterized in that, The method further includes: If the message to be tested is a normal message, the message sending request is sent to the second SMS center, which then sends the message sending request to the second mobile device switching center. The second mobile device switching center then sends the message to be tested to the receiving terminal via the receiving base station. Wherein, the second SMS center is the SMS center of the second operator where the receiving terminal is located, and the second mobile device exchange center is the mobile device exchange center maintained by the second operator.
6. A spam message interception device, characterized in that, The device includes: The acquisition unit is used to acquire the message to be tested in the message sending request when a message sending request is received; Extraction unit, used to extract text content from the message to be tested; The detection unit is used to detect the text content based on keywords of historical spam messages, hash values of historical spam messages, and a large language model to determine whether the message to be tested is spam. An interception unit is configured to intercept the message to be tested if the message to be tested is a spam message. The detection unit is further configured to: generate a regular expression based on keywords of historical spam messages; perform regular expression matching on the text content based on the regular expression to obtain a regular expression matching result; determine the hash value of the text content; determine the maximum similarity between the text content and the historical spam messages based on the hash value of the text content and the hash value of the historical spam messages; generate model input information based on the text content and preset scoring rules; input the model input information into a large language model to obtain a first score for the text content; and determine whether the message to be tested is spam based on the regular expression matching result, the maximum similarity, and the first score; wherein, determining whether the message to be tested is spam based on the regular expression matching result, the maximum similarity, and the first score includes: determining the average value of the regular expression matching result and the maximum similarity; normalizing the first score to obtain a second score; performing a weighted summation of the average value and the second score to obtain a third score; and determining whether the message to be tested is spam based on the third score and a preset threshold.
7. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Method for recognizing junk short messages, client, cloud server and system
CN106162584A
Network security system and method based on information identification
CN113887207A
Information identification method and device
CN114282097A
Method and device for recognizing community content spam comments based on large language model
CN118332114A