A method for identifying harmful information in SMS content based on big data
By building a multilingual harmful information recognition model, combining SMS scanning, character feature analysis and traceability technology, the problem of accurate identification and management of harmful information in SMS is solved, and efficient and accurate detection of harmful information and dynamic security management is achieved.
Patent Information
- Application Number
- CN202411607204.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing SMS content recognition methods are difficult to accurately identify and effectively control harmful information, especially in multilingual environments, where traditional methods cannot effectively identify and manage the dissemination of harmful information.
Using a big data-based method, an information recognition model is constructed by collecting harmful information samples from multiple languages, including three modules: the first module performs initial scanning, the second module recognizes suspicious character characteristics of harmful information, and the third module traces the IP address of the sending end of the short message data, and combines the sending feature analysis to dynamically update the user number tag to identify and manage harmful information.
It improves the accuracy and wide applicability of the identification of harmful information, reduces the false alarm rate, can timely detect and respond to potential harmful information, dynamically adjust security monitoring measures, reduce the risk of malicious information proliferation, and ensure the security of the user's communication environment.
Smart Images

Figure CN119907005B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data recognition, and in particular to a method for identifying harmful information in short message content based on big data. Background Art
[0002] With the development and improvement of wireless communication services, while the short message service provides convenient information for users, it also provides a way for the spread of harmful information. Especially for the underage group, facing the intimidation or harassment of harmful information will cause psychological shadows. The traditional method of identifying harmful information in short message content only judges whether it is harmful based on the literal meaning of the short message text, and it is difficult to accurately identify and effectively control the spread of harmful garbage.
[0003] In summary, how to accurately identify and effectively control harmful information is an urgent problem to be solved and optimized for the method of identifying harmful information in short message content based on big data. Summary of the Invention
[0004] The present invention provides a method for identifying harmful information in short message content based on big data, and solves the technical problem of how to accurately identify and effectively control harmful information.
[0005] In order to solve the above technical problems, the present invention provides a method for identifying harmful information in short message content based on big data, and the specific technical solution is as follows:
[0006] A method for identifying harmful information in short message content based on big data includes the following steps:
[0007] S100, collecting a variety of historical samples of harmful information short messages to obtain a historical sample set of harmful information; compiling the harmful information keywords in the sample set into multiple language versions to train an information recognition model and output a harmful information feature recognition result; the information recognition model includes a first module, a second module, and a third module;
[0008] S200, inputting the short message data sent in real time by the sending end into the information recognition model; the first module receives the short message data for scanning and reading to output initial harmful information suspicious short message data;
[0009] S300, transmitting the short message data identified by the first module to the second module, and the second module identifies the harmful information suspicious character features in the initial harmful information suspicious data and obtains the meaning of the suspicious characters to output secondary harmful information suspicious short message data;
[0010] S400, input the SMS data identified by the second module into the third module. The third module traces the IP address of the SMS data sender to obtain the source tracing information of the SMS data sender, and outputs the recognition result of the final harmful information suspicious SMS data according to the source tracing information. When it is recognized that there is SMS data with harmful information, execute step S500;
[0011] S500, conduct a sending feature analysis on the SMS data with harmful information to obtain the feature data of the harmful information sending rule. Based on the feature data of the harmful information sending rule, obtain the currently defined label data of the user number, and update the currently defined label of the user number to obtain a secure user number.
[0012] As a further optimization solution of the present invention, collect multiple historical samples of harmful information SMS to obtain a historical sample set of harmful information. Compile the historical sample set of harmful information into multiple language versions and train an information recognition model to output the recognition result of harmful information features. The information recognition model includes a first module, a second module, and a third module, including:
[0013] Collect historical sample data of various SMS sources with harmful information. The historical sample data includes fraud SMS, terrorist and bloody, pornographic and filthy, and virus links;
[0014] Based on the historical sample data, convert the keywords with harmful information in the historical sample data into multiple target language versions; and conduct semantic verification on the multiple target language versions of multiple keywords to obtain an accurate translation data set;
[0015] For the accurate translation data set, through: to obtain similar words of multiple keywords; where w represents the similar word feature parameter; K i represents the keyword; α represents the similarity threshold. Based on the obtained similar words of multiple keywords, through: ; to obtain a keyword similar semantic expansion set; and combine the keyword similar semantic expansion set with the keyword K i to form a comprehensive sample set; where represents the final multi-language similar keyword expansion set; m represents the number of similar keywords; represents the expansion function;
[0016] Construct the comprehensive sample set into a training set: ; where n i represents the i-th accurate translation data item. Input the training set G into the information recognition model for training to output the recognition result of harmful information features.
[0017] As a further optimization solution of the present invention, the SMS data sent in real time by the sending end is input into the information recognition model; the first module receives the SMS data for scanning and reading to output initial harmful information suspicious SMS data, including:
[0018] The newly sent SMS data is input into the information recognition model, and the first module extracts the SMS data with similar format templates reproduced multiple times to obtain similar template text data;
[0019] Perform pre-screening on the similar template text data D, by setting a format similarity rule set {R i} n i =1; by traversing each template text T in the SMS data j ; for each T j template text data, perform format feature analysis to obtain the matching situation of the template text data format rule R i ; if T j does not conform to the R i format rule, then retain it in the template text data with inconsistent formats, otherwise discard it; to filter out the template text data with unified formats; the template text data with unified formats is the initial harmful information suspicious data.
[0020] As a further optimization solution of the present invention, the SMS data recognized by the first module is transmitted to the second module, and the second module recognizes the harmful information suspicious character features in the initial harmful information suspicious data and obtains the meanings of the suspicious characters to output secondary harmful information suspicious SMS data, including:
[0021] Further read the SMS text characters of the template text data with inconsistent formats to obtain a suspicious character feature data set in the text; the suspicious character feature data contains multi-lingual misleading words;
[0022] Identify and label each suspicious character feature data to obtain a corresponding category data set; the corresponding category data set includes labeling multi-lingual words as fraud, violence and bloodshed, pornographic filth, and horror;
[0023] Based on the suspicious character feature data set, parse the template text data with inconsistent formats according to the context, by: ; to judge that the character data item meaning in the character feature data set is suspicious, where represents the final parsed meaning of the character data item; represents the meaning of the i-th template text segment; represents the inference parameter of the template segment t i ; represents the corresponding category data set; Indicates the inference symbol;
[0024] Based on the obtained judgment result, by: , to obtain the frequency of suspicious characters. In the formula, fi represents the frequency of suspicious characters in the SMS data; C represents the number of occurrences of suspicious characters; W represents the total number of characters in the SMS data; and perform secondary screening on the SMS data with suspicious characters to obtain secondary harmful information suspicious data.
[0025] As a further optimization scheme of the present invention, identify and label each suspicious character feature data to obtain a corresponding category data set; the corresponding category data set includes labeling multi-language words as fraud, violence and blood, pornographic filth, and horror categories, including:
[0026] Based on the corresponding category data set, capture the suspicious words in the SMS, and match the suspicious words in the SMS data with the corresponding category data set to obtain a matching data set; based on the matching data set, establish a sensitive word library for minors;
[0027] Based on the sensitive word library for minors, replace the sensitive words or inappropriate words in the sensitive word library for minors with safe and suitable pronouns for minors to obtain a pronoun data set;
[0028] By comparing the suspicious words in the SMS data with the pronoun data set, and judging the suspicious degree of the pronouns, to analyze the harmful risk of the suspicious pronouns to obtain risk SMS information;
[0029] Perform a similarity detection on the risk SMS information and the sensitive words in the sensitive word library for minors to determine that the pronouns in the SMS data are harmful information data.
[0030] As a further optimization scheme of the present invention, perform a similarity detection on the risk SMS information and the sensitive words in the sensitive word library for minors to determine that the suspicious words in the SMS data are harmful information data, including:
[0031] According to the similarity detection, obtain the matching degree of minor sensitive words; based on the matching degree, perform further fuzzy matching analysis on the risk SMS information to obtain accurate matching data;
[0032] Based on the accurate matching data, determine that the pronouns in the SMS data are harmful information data, and set up an information feedback channel;
[0033] Based on the feedback channel, timely feedback inappropriate information and automatically classify and prioritize the feedback information to update the sensitive word library for minors to obtain a feedback result.
[0034] As a further optimization solution of the present invention, the SMS data after being recognized by the second module is input into the third module, and the third module traces the IP address of the SMS data sender to obtain the origin tracing information of the SMS data sender, and outputs the recognition result of the final harmful information suspicious SMS data according to the origin tracing information, including:
[0035] Trace the IP address according to the metadata of the SMS data; the metadata includes the first eight digits of the sending number, the HTTP request header, and the domain name record information of the SMS link;
[0036] Based on the metadata, obtain the IP address of the SMS data. When the SMS IP address is an overseas address, the degree of suspicion of harmful information is high; and further judge in combination with the text information of the SMS data. When the text information identifies words related to money and personal safety, a warning message is sent to the user number of the receiving end in a timely manner; to identify the final harmful information suspicious SMS data.
[0037] As a further optimization solution of the present invention, perform a sending feature analysis on the SMS data with harmful information to obtain harmful information sending rule feature data; based on the harmful information sending rule feature data, obtain the currently defined label data of the user number, and update the currently defined label of the user number to obtain a secure user number, including:
[0038] Statistically analyze the sending time, sending user information, and sending content of the SMS data with harmful information to obtain harmful information sending feature data;
[0039] Based on the harmful information sending feature data, perform a usage behavior analysis on the user numbers that frequently receive harmful information. The usage behavior of the user number includes: browsing, reading, sending, and receiving SMS data; to obtain usage behavior security analysis data;
[0040] Based on the usage behavior security analysis data, obtain harmful information sending rule feature data; according to the harmful information sending rule feature data, obtain the currently defined label data of the user number; update the currently defined label data of the user number to obtain a secure user number.
[0041] As a further optimization solution of the present invention, according to the harmful information sending rule feature data, obtain the currently defined label data of the user number, including:
[0042] According to the usage behavior security analysis data, by: ; to obtain the abnormal degree of user number usage; where D u represents the abnormal degree of user number usage; f uIt represents the frequency of harmful information received by the user, that is, the number of harmful text messages received within the time window H; μ represents the average frequency of harmful information received by all users; σ represents the standard deviation of the frequency of harmful information received by all users.
[0043] Set a preset threshold range, and use the abnormal degree D of the user number u Compare it with the preset threshold range, and based on the comparison result, judge whether the user number usage behavior is a normal behavior or an abnormal behavior.
[0044] According to the usage behavior of the user number, define labels for the user number to obtain the currently defined label data of the user number; the currently defined label data of the user number includes two labels: normal and abnormal.
[0045] As a further optimization scheme of the present invention, update the currently defined label data of the user number to obtain a secure user number, including:
[0046] Based on the currently defined label data of the user number, obtain the login information, purchase behavior, browsing record or interaction record of the system platform used by the user number to obtain a label source data set;
[0047] Based on the label source data set, provide a prompt for the user to selectively clean the label; so that the user can clean and update the currently defined label data to obtain an initial defined label;
[0048] Based on the initial defined label, make the current user number a secure number.
[0049] The present invention has at least the following beneficial effects: By collecting a variety of historical samples of harmful information text messages, the present invention can construct a set of harmful information samples containing different types and language versions. This process helps the model learn the patterns of harmful information in different languages and cultural backgrounds, thereby improving the accuracy and wide applicability of the model; Enhanced keyword recognition ability: By compiling the harmful information keywords in the sample set into multiple language versions, the model's ability to recognize harmful information can be enhanced, ensuring that it can recognize harmful content in multiple languages and improving the cross-language protection effect.
[0050] Input the text message data into the information recognition model in real time, which can detect potential harmful information in a timely manner, ensure that a response can be made in the shortest time, and avoid the spread of harmful information; Automatically input the text message data into the information recognition model, reducing manual intervention and improving processing efficiency.
[0051] Through the scanning of the first module, potential harmful information text messages can be quickly screened out, and preliminary suspicious data can be marked. This step improves the system's sensitivity to harmful information, enabling subsequent analysis to focus more on potential risks; the work of the first module can quickly identify and mark suspicious data in a large amount of data, ensuring more efficient subsequent processing.
[0052] The second module is not just simple text scanning. It further analyzes the character features in the text message to dig deeper into potential harmful information. This multi-level identification method greatly enhances the detection accuracy; by identifying the meaning of suspicious characters and deeply analyzing their potential risks, it helps to accurately identify real harmful information, reduce the false alarm rate, and improve the credibility of identification.
[0053] By analyzing the sending characteristics of harmful information, the behavior patterns of senders can be identified, and "harmful information sending rules" can be further established. Based on these characteristics, more targeted preventive measures can be taken to pay special attention to and manage high-risk users; by updating the labels of user numbers, the security monitoring measures for users can be dynamically adjusted. For example, if a user frequently sends harmful information, they can be marked as high-risk users, so that their behavior can be more strictly reviewed in the future; based on the sending characteristics and label updates, the risk of malicious information spreading can be effectively reduced, ensuring the security of the user communication environment. Brief Description of the Drawings
[0054] Figure 1 It is a schematic flow diagram of a method for identifying harmful information in text messages based on big data provided by an embodiment of the present invention. Detailed Embodiment
[0055] The following further describes the present application in detail with reference to the drawings. It is necessary to point out here that the following detailed embodiments are only used to further illustrate the present application and cannot be understood as limiting the protection scope of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application according to the above application content.
[0056] A method for identifying harmful information in text messages based on big data provided by this embodiment is specifically implemented as follows:
[0057] As Figure 1 shown, a method for identifying harmful information in text messages based on big data includes the following steps:
[0058] S100, collect historical samples of various harmful information text messages to obtain a set of historical samples of harmful information; compile the harmful information keywords in the sample set into multiple language versions to train an information recognition model and output the recognition results of harmful information features; the information recognition model includes a first module, a second module, and a third module;
[0059] S200, input the text message data sent in real time by the sending end into the information recognition model; the first module receives the text message data for scanning and reading to output initial suspicious text message data of harmful information;
[0060] S300, transmit the text message data recognized by the first module to the second module, and the second module recognizes the suspicious character features of harmful information in the initial suspicious data of harmful information and obtains the meaning of the suspicious characters to output secondary suspicious text message data of harmful information;
[0061] S400, input the text message data recognized by the second module into the third module, and the third module traces the IP address of the text message data sending end to obtain the traceability information of the text message data sending end, and outputs the recognition result of the final suspicious text message data of harmful information according to the traceability information; when text message data containing harmful information is recognized, execute step S500;
[0062] S500, perform a sending feature analysis on the text message data containing harmful information to obtain harmful information sending rule feature data; based on the harmful information sending rule feature data, obtain the currently defined label data of the user number, and update the currently defined label of the user number to obtain a secure user number.
[0063] In this embodiment, by step S100, historical samples of harmful text messages in multiple languages and types are collected to establish a set of harmful information samples containing historical data. Then, the harmful information keywords in these samples are compiled into multiple language versions to build a cross - language information recognition model. The model includes three modules, which are respectively responsible for identifying harmful information at different levels; through the construction of a multilingual sample library and keyword compilation, the model can identify harmful information in a multilingual environment, enhance the ability to identify cross - cultural and multilingual harmful content, help reduce false alarms, and improve the wide applicability and accuracy of the model.
[0064] Based on step S100, the text message data of the sending end is input into the information recognition model in real time through step S200. The first module scans the text message data, extracts and marks the initially identified suspicious information, providing preliminary screening data for the analysis of subsequent modules; the screening of the first module can quickly process a large amount of text message data, reduce the interference of redundant information, filter and mark the suspected harmful information, provide clear data for further analysis, and effectively improve the processing efficiency.
[0065] The data screened by the first module in S300 is transmitted to the second module, which further identifies the harmful character features and meanings in the preliminarily screened data, thereby outputting more in-depth suspicious SMS information; through in-depth analysis at the character level, the second module can identify more concealed harmful information features, further improve the accuracy of identification, reduce false alarms, and enhance the detection credibility of the model.
[0066] The data identified by the second module in step S400 is transmitted to the third module to trace the IP address of the SMS sender, obtain relevant traceability information, confirm the authenticity and risk of the sending source, and finally generate a comprehensive identification result of harmful information; through traceability analysis, the third module can track and mark high-risk sending sources, which helps prevent the further spread of malicious information and improve the overall information security.
[0067] In S500, the sending characteristics of SMS data with harmful information are analyzed to establish and update the "harmful information sending rule feature data". According to the feature data, the user number label is updated to a dynamic security label to ensure accurate marking of high-risk users; by analyzing and updating user labels, user behavior can be dynamically monitored, especially for users who frequently send harmful information. This mechanism effectively reduces the risk of malicious information diffusion and provides continuous security protection for the overall communication environment.
[0068] Through the collaborative work among the above steps, real-time SMS can be quickly detected, malicious information can be promptly identified, and potential hazards (such as fraud information or spam) can be prevented from spreading; the multi-module identification method (from preliminary scanning to in-depth analysis and then to traceability tracking) greatly improves the accuracy of identification, can refine the various features of malicious information and reduce the false alarm rate; through multi-language support and real-time update of user labels, different language environments and changing trends of malicious information can be addressed; not only can information be identified, but also potential malicious users can be monitored and marked, which helps operators or service platforms track user behavior and promptly intervene to protect ordinary users from malicious SMS.
[0069] In a preferred embodiment of the present invention, step S100 may further include the following steps:
[0070] Step S101, collecting historical sample data of various SMS sources with harmful information, where the historical sample data includes fraud SMS, terrorist and bloody, pornographic and filthy, and virus links;
[0071] Step S102, based on the historical sample data, converting the keywords with harmful information in the historical sample data into multiple target language versions; and performing semantic verification on the multiple target language versions of multiple keywords to obtain an accurate translation data set;
[0072] Step S103, for the accurate translation data set, by: , to obtain similar words of multiple keywords; where w represents the similar word feature parameter; K i represents the keyword; α represents the similarity threshold; based on the obtained similar words of the multiple keywords, by: ; to obtain the keyword similar semantic expansion set; and combine the keyword similar semantic expansion set with the keyword K i to form a comprehensive sample set; where represents the final multi-language similar keyword expansion set; m represents the number of similar keywords; represents the expansion function;
[0073] Step S104, construct the comprehensive sample set into a training set: ; where n i represents the i-th accurate translation data item; input the training set G into the information recognition model for training to output the harmful information feature recognition result.
[0074] In the embodiment of the present invention, collecting historical sample data through step S101 is to collect various types of harmful information text messages with representativeness, such as fraud, terrorist bloodshed, pornographic filth, and virus links. These sample data provide keywords, phrases, and content patterns for subsequent processing; the historical sample data provides a rich corpus for model training, ensuring that the model can cover various types of harmful information and improving the recognition accuracy of similar information.
[0075] The harmful keywords extracted from the historical sample data through step S102 will be translated into multiple target languages; ensuring that the translations in different languages retain the original meaning and eliminating the ambiguity in the language translation process. This step will generate an accurate multi-language keyword translation data set; through multi-language keyword expansion and verification, it enhances the model's understanding of multi-language information, ensures the accuracy of keyword translation, and helps in the recognition of harmful information in a multi-language environment.
[0076] Step S103 combines the keyword with the similar word semantic expansion set to generate a comprehensive sample set containing various semantic variants, thereby enhancing the model's recognition flexibility for the same type of information by expanding synonyms, related words, and various semantic variants of the keyword. Even if the harmful information uses different expressions, the model can still recognize it.
[0077] In step S104, the comprehensive sample set is constructed into the training set G. Inputting the training set into the information recognition model, through the training of multi-language and extended semantic samples, enables the model to recognize harmful information in various languages and expressions; specifically including:
[0078] Encode the training set G into sequence data, and input the sequence data into the information recognition model; the information recognition model includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. Transmit the intermediate representation data of each hidden layer to the output layer, and the output layer outputs the recognition result representing the harmful information features of historical samples, specifically as follows:
[0079]
[0080] where W u 、W r 、W c represent weight parameters, b u 、b r 、b c represent bias parameters, represents the dot product, u (t) 、r (t) and c (t) represent the intermediate states of the first, second, and third hidden layers respectively, X (t) represents the t-th data item of the comprehensive sample set, H (t) and H (t-1) represent the intermediate representation data of the t-th and (t - 1)-th comprehensive sample sets respectively, n ≥ t ≥ 1, n represents the total number of data items in the input comprehensive sample set. When t = 1, H (t-1) = X (t) , tanh is the hyperbolic tangent function, and σ represents the sigmoid function;
[0081] Based on the recognition result of the harmful information of the historical comprehensive sample set output by the output layer, output the intermediate representation data of the recognition result of the historical comprehensive sample set;
[0082] In this embodiment, each data item X(t) in the comprehensive sample set is input into the information recognition model. These data items may be arranged in chronological order because in the subsequent description, H (t−1) is mentioned, which represents the data item one time step forward from the current data item X (t) ; the time relationship of the sequence data is retained, and the model can learn and infer through time information, which is suitable for time series data analysis tasks.
[0083] The model structure is an input layer, a hidden layer, and an output layer; when the input layer receives the data X (t) ; the first hidden layer, the second hidden layer, and the third hidden layer sequentially receive the output of the previous layer. Each layer performs weighted sum and bias through the weight W y and the bias b y , and then applies an activation function (the hyperbolic tangent function tanh) to generate intermediate states u (t) 、r (t) and c(t) The intermediate state is transmitted through multiple layers, gradually extracting and expressing the high-order features of the input data.
[0084] The output layer receives the output of the last hidden layer, calculates the weighted sum through weights and biases, and then applies the sigmoid function to output the recognition result; the multiple hidden layers can learn the features at different abstraction levels of the data, improving the expression ability and recognition accuracy of the model.
[0085] The intermediate states u (t) , r (t) and c (t) of the last hidden layer are transmitted to the output layer, the weighted sum is calculated through weights and biases, and the sigmoid function is applied to generate the recognition result; through the transmission of the intermediate states, the model can map the learned complex features to the final recognition decision, thereby improving the accuracy and interpretability of the recognition result; the multiple hidden layer structure allows the model to extract and combine the complex features of the input data layer by layer, enhancing the expression ability of the model.
[0086] The model can process time series data, effectively utilizing the time correlation by retaining the historical information sample set H (t−1) ; through the processing of the sigmoid function in the output layer, the model can generate accurate recognition results and is applicable to multi-class recognition tasks.
[0087] In a preferred embodiment of the present invention, the above step S200 includes:
[0088] Step S201, inputting the newly sent short message data into the information recognition model, and the first module extracts the short message data with a similar format template reproduced multiple times to obtain similar template text data;
[0089] Step S202, performing pre-screening on the similar template text data D, by setting a format similarity rule set {R i} n i = 1; by traversing each template text T j in the short message data; performing format feature analysis on each T j template text data to obtain the matching situation of the template text data format rule R i ; if T j does not conform to the R i format rule, it is retained in the template text data with inconsistent formats, otherwise it is discarded; to filter out the template text data with unified formats; the template text data with unified formats is the initial harmful information suspicious data.
[0090] Specific description: The output layer:
[0091]
[0092] In the formula, represents the feature representation of the newly obtained SMS data sent in real time by the sending end, represents the intermediate representation data of the v-th unit of the data matrix, M represents the set of all units of the data matrix, and W y is the weight parameter, and b y is the bias parameter, representing the sigmoid function.
[0093] In the embodiment of the present invention, new SMS data is input into the information recognition model through step S201. This model identifies whether there is harmful information by analyzing features such as SMS content, format, keywords, etc.; the first module in the model is specifically used to detect and extract similar format templates that appear multiple times in the SMS. Through the template recognition algorithm, the system can identify SMS content with similar formats that appear repeatedly (such as the fixed format of fraud SMS, the standard format of junk advertisements, etc.).
[0094] For example, fraud SMS usually has a fixed format, such as "Dear customer, your account has been frozen. Please click on this link to restore it." This format will appear repeatedly in multiple SMS; Obtaining similar template data: Through the format analysis of SMS data, the system extracts these repeated and similar-format SMS content to form a set of "similar template text data"; By extracting the SMS data of similar templates, the model can quickly screen out potentially risky texts in a large number of SMS, avoiding analyzing all SMS content one by one.
[0095] Template matching can help the model quickly identify harmful information with a fixed format, especially for known fraud and junk information patterns, which greatly improves the accuracy of identification;
[0096] Through step S202, the extracted similar template data is further screened to reduce the computational amount of subsequent processing and the possibility of false alarms. The screening rules are defined by a preset set of format similarity rules {Ri}, and these rules are fixed format template patterns obtained through the analysis of historical data. For example: The format rule R i can be "the SMS starts with a certain keyword, contains a URL link, and ends with a specific word"; for each extracted SMS template text T j , the system will traverse the rule set and match each SMS template with the format rule; the format feature analysis will detect whether T j complies with a specific rule. If the SMS template complies with a certain rule, the SMS is considered to belong to a template with a unified format, otherwise the SMS is considered to have a different format; SMS that complies with the format rule is considered to be harmful information with a unified format (such as the standard format of fraud SMS). These SMS will be regarded as "initial harmful information suspicious data".
[0097] Text messages that do not conform to the format rules will be screened out, which means that these text message formats do not conform to the known harmful information templates and will not be marked as harmful information temporarily; through pre-screening, the system can accurately identify harmful information that conforms to the template according to the known format rules. Text messages that conform to the format rules are likely to be harmful information, while those that do not conform to the rules are excluded, reducing the possibility of misjudgment; it helps to narrow down the scope of suspicious information to be processed and focus the processing on text messages with high risks. In this way, subsequent in-depth analysis can focus on harmful information with high potential; by filtering out template texts with uniform formats, it ensures that only those text messages that truly conform to the characteristics of harmful information are further processed subsequently, avoiding wasting computing resources in a large amount of irrelevant data.
[0098] In a preferred embodiment of the present invention, the above step S300 includes:
[0099] Step S301, further read the text message text characters from the template text data with inconsistent formats to obtain a dataset of suspicious character features in the text; the suspicious character feature data includes multi-lingual misleading words;
[0100] Step S302, identify and label each piece of suspicious character feature data to obtain a corresponding category dataset; the corresponding category dataset includes labeling multi-lingual words as fraud category, violent and bloody category, pornographic and filthy category, and horror category;
[0101] Step S302, based on the dataset of suspicious character features, parse the template text data with inconsistent formats according to the context, through: ; to determine that the character data item meaning in the character feature dataset is suspicious, where represents the final parsed meaning of the character data item; represents the meaning of the i-th template text segment; represents the template segment t i 's inference parameter; represents the corresponding category dataset; represents the inference symbol;
[0102] Step S303, based on the obtained judgment result, through: , to obtain the frequency of the existence of suspicious characters, where fi represents the frequency of the appearance of suspicious characters in the text message data; C represents the number of times the suspicious characters appear; W represents the total number of characters in the text message data; and perform secondary screening on the text message data with suspicious characters to obtain secondary harmful information suspicious data.
[0103] In this embodiment,
[0104] In a preferred embodiment of the present invention, the above step S302 includes:
[0105] Step S3021: Based on the corresponding category dataset, capture the suspicious words in the short message, and match the suspicious words in the short message data with the corresponding category dataset to obtain a matching dataset; based on the matching dataset, establish a minor sensitive word library.
[0106] Step S3022: Based on the minor sensitive word library, replace the sensitive words or inappropriate words in the minor sensitive word library with safe and appropriate pronouns for minors to obtain a pronoun dataset.
[0107] Step S3023: By comparing the suspicious words existing in the short message data with the pronoun dataset and judging the suspicious degree of the pronouns, analyze the harmful risks of the suspicious pronouns to obtain risk short message information.
[0108] Step S3024: Perform a similarity detection on the risk short message information and the sensitive words in the minor sensitive word library to determine that the pronouns in the short message data are harmful information data.
[0109] In this embodiment, through step S301, the short message text characters are read from the template text data with inconsistent formats, and the "suspicious character feature dataset" in the text is identified and extracted. Here, the "suspicious characters" refer to words with misleading or malicious intentions in multiple languages. By deeply analyzing the characteristics of these characters, the system can identify potential risk content; laying a foundation for subsequent risk identification, extracting a character set with specific properties helps improve the accuracy of classification and recognition.
[0110] Through step S302, the "suspicious character feature dataset" obtained in step S301 is analyzed and identified item by item, and classified into different risk categories such as fraud, violence and bloodshed, pornographic filth, and horror. This classification is based on the semantic characteristics of the characters and the speculation of the specific context of multiple languages; through clear marking and classification, the system can more targeted identify potential harmful information, facilitating further analysis and screening processing in subsequent steps.
[0111] Based on step S302, step S303 uses context semantic analysis to combine the information of different template segments in the text to generate the complete meaning of the character features. Based on the analysis of the character feature dataset, judge the meaning of various character data items, and judge whether certain characters are risky; through the context, analyze whether each character data item may be a risk character; for each template segment, speculate its semantics and calculate the possibility; for data items with different category features, infer whether they conform to the risk attributes of certain categories through the marked information; analyzing the semantic relationship and features in the context helps to avoid misjudgment and at the same time improves the ability to identify complex risk content.
[0112] Step S304 calculates the occurrence frequency of suspicious characters in the short message according to the judgment result of step S303; based on the screening of the character occurrence frequency, the secondary harmful information is accurately locked. Through this step, the system can focus on high-risk information, greatly improving the screening accuracy and efficiency.
[0113] In a preferred embodiment of the present invention, the above step S3024 includes:
[0114] Step S30241, according to the similarity detection, to obtain the matching degree of minor sensitive words; based on the matching degree, further fuzzy matching analysis is performed on the risk short message information to obtain accurate matching data;
[0115] Step S30242, based on the accurate matching data, determines that the potential pronouns in the short message data are harmful information data, and sets up an information feedback channel;
[0116] Step S30243, based on the feedback channel, timely feedbacks inappropriate information and automatically classifies and prioritizes the feedback information to update the minor sensitive word library to obtain a feedback result.
[0117] In this embodiment, step S3021 grabs suspicious words in the short message data based on the "corresponding category data set" defined in the previous steps. Suspicious words refer to those words that may contain fraud, violence, pornographic or inappropriate content. By matching with the corresponding category data set, the system compares the suspicious words with predefined categories (such as fraud category, violent bloody category, pornographic filth category, etc.) to determine their belonging categories; based on this matching result, the system further establishes a "minor sensitive word library". This sensitive word library is a specific vocabulary set for harmful information that minors may come into contact with, aiming to filter and manage the information content that minors come into contact with;; by accurately matching the suspicious words with the category data set, the system can automatically identify and classify sensitive information related to minors; establishing the "minor sensitive word library" is a preventive measure that can provide an important word resource library for subsequent processing steps, thus better protecting the information security of minors.
[0118] Step S3022 replaces the sensitive words or inappropriate words in the minor sensitive word library established in step S3021 with "potential pronouns". These potential pronouns refer to alternative words that have a certain degree of concealment but can convey the same or similar meaning in the context, specifically including:
[0119] 1. Game recharge and consumption;
[0120] "Recharge" can be replaced by "obtain items" or "acquire resources";
[0121] Or replace it with "Premium Member", "VIP", or "Special Member";
[0122] Or replace it with "Exclusive Feature" or "Privileged Feature", "Privilege".
[0123] 2. In-game Social Interaction;
[0124] "Clan" or "Guild" can be replaced with "Team" or "Club";
[0125] "PK" or "Arena" can be replaced with "Battle" or "Friendly Battle";
[0126] 3. Violence, Bloodiness, and Bad Culture;
[0127] "Bloodiness" or "Violence" can be replaced with "Intense" or "Tense";
[0128] "Hunt" or "Kill" can be replaced with "Challenge" or "Adventure";
[0129] "Evil" can be replaced with "Opponent" or "Hostile Force".
[0130] 4. Advertising Slogans and Guiding Phrases;
[0131] "Buy Now" can be replaced with "Obtain" or "Unlock";
[0132] "Limited-Time Offer" can be replaced with "Special Event" or "Limited-Time Event".
[0133] 5. Game Duration and Reminders;
[0134] "Addicted" can be replaced with "Focused" or "Engaged"; "Excessive Gaming" can be replaced with "Reasonable Gaming" or "Healthy Gaming".
[0135] 6. Personal Attacks, Vulgarity, and Unhealthy Content;
[0136] "Loser" can be replaced with "Needs Improvement" or "To Be Improved".
[0137] 7. Minor Identification;
[0138] "Minor" can remain unchanged, but ensure privacy and security settings.
[0139] 8. Internet Addiction, Inductive Content;
[0140] "Addicted" can be replaced with "Loves" or "Likes or Must-Play" can be replaced with "Recommended" or "Worth a Try".
[0141] These alternative words are carefully selected to ensure a more friendly and safe information transmission to minors, avoiding direct exposure of sensitive content;; By replacing sensitive words with pronouns, the risk of minors being exposed to inappropriate information can be effectively reduced while not losing the essential conveyance of information; Building a safer barrier for minors' information environment helps improve the accuracy of information filtering and processing.
[0142] Step S3023 compares the suspicious words appearing in the SMS data with the pronoun data set. The design of pronouns is to replace sensitive words, but in actual use, pronouns may also carry certain hidden risks. The system judges whether there are harmful risks based on factors such as the appearance frequency and context of pronouns. This process is achieved by analyzing the suspiciousness of pronouns, that is, judging whether the pronouns still have potential misleading or harmful nature;; Through the analysis of the suspiciousness of pronouns, the system can detect potentially harmful hidden information, especially those risk information expressed in a concealed way; Thus, it helps to further filter or block potential risks to minors and enhance the depth of information security protection.
[0143] Step S3024 detects the similarity between the SMS information after being replaced with pronouns and the sensitive words in the minor sensitive word library. This detection process can judge whether the pronoun still may represent harmful information by calculating text similarity (such as cosine similarity, Jaccard similarity, etc.); When the similarity detection result shows that the similarity between the replacement word of the pronoun and the sensitive word is relatively high, the SMS content is determined to be harmful information and marked as "risk SMS information";; Similarity detection can improve the ability to identify harmful information after being replaced with pronouns. Even if the sensitive word has been replaced with a pronoun, the system can still judge its potential harmfulness through semantic similarity; Thus, it effectively ensures the accuracy of information security processing. Even after being replaced with pronouns, the system can further identify and block the possible impact on minors.
[0144] The above steps together constitute a comprehensive minor information protection mechanism. From the capture, classification, replacement of suspicious words to the final similarity detection, each link minimizes the potential risks of harmful information to minors as much as possible, so as to effectively identify, replace and detect sensitive information and build a multi-level protection system; This process not only improves the ability to identify and filter potential harmful information, but also protects the information security of minors while not losing the accuracy and efficiency of information transmission.
[0145] In a preferred embodiment of the present invention, the above step S400 may include:
[0146] Step S401, trace the IP address according to the metadata of the SMS data; the metadata includes the first eight digits of the sending number, the HTTP request header, and the domain name filing information of the SMS link;
[0147] Step S402, based on the metadata, obtain the IP address of the SMS data. When the SMS IP address is an overseas address, the suspicion degree of harmful information is high; and further judge in combination with the text information of the SMS data; when words related to money and personal safety are recognized in the text information, a warning message is sent to the user number at the receiving end in a timely manner; to identify the final harmful information suspicious SMS data.
[0148] In this embodiment, in step S401, by analyzing the metadata of the SMS, the source IP address of the SMS is traced, providing basic information support for identifying harmful information. Specifically, the following metadata is included; thus helping to initially determine the source country or region of the SMS and providing geographical distribution information of the sender; recording the HTTP request header information during the SMS sending process, which may contain key information such as proxy server information or data packet source, helping to trace the source; the domain name filing information of the SMS link: query the domain name filing of the link included in the SMS content to determine whether the link has a legal filing or is a malicious link, helping to determine the credibility of the SMS source; by analyzing these metadata, the system can initially judge the legality of the SMS source and provide a direction for the next IP address acquisition; thus being able to help lock in potential risky SMS sources in the initial stage. The information contained in the metadata can quickly screen out SMS with obvious abnormal characteristics, thus effectively narrowing the tracing scope; tracing the SMS sending IP address can provide a geographical basis (such as an overseas IP) for the next judgment, helping to judge the risk degree of the SMS; providing a multi-dimensional basic data support, including the comprehensive analysis of the number segment, request header, and domain name filing information, helps to improve the accuracy of tracing the source.
[0149] Step S402 will obtain the IP address based on the metadata provided in step S401. The specific judgment logic is as follows:
[0150] By analyzing the obtained IP address, determine its source. If the IP address belongs to overseas (i.e., non-local or non-domestic IP), the suspicion level of the text message is considered relatively high; further analyze the content of the text message to identify whether it contains sensitive keywords related to money or personal safety, etc. For example, words such as "winning a prize", "transferring money", "account freezing", "emergency contact", etc., which usually appear in fraudulent text messages; when identifying keywords containing potential risks, a warning message will be sent to the receiving user's number to remind the user to be vigilant and prevent being deceived; through the dual judgment of overseas IP and text content, the system can significantly improve the accuracy of identifying harmful information. Text messages with overseas IPs often have relatively high risks. By combining specific analysis of the text content, the misjudgment rate can be effectively reduced; the sending of targeted warning messages can enable users to understand potential risks in real time, thereby strengthening their awareness of prevention, achieving timely warning effects, and reducing the possibility of users being deceived due to information asymmetry; the dual judgment system can provide a more intelligent and proactive security protection mechanism, especially suitable for dealing with cross-border text message fraud and quickly screening out high-risk information.
[0151] In a preferred embodiment of the present invention, the above step S500 includes:
[0152] Step S501, statistically analyze the sending time, sending user information, and sending content of the text message data with harmful information to obtain harmful information sending feature data;
[0153] Step S502, based on the harmful information sending feature data, analyze the usage behavior of the user numbers that frequently receive harmful information. The user number usage behavior includes: browsing, reading, sending, and receiving text message data; to obtain usage behavior security analysis data;
[0154] Step S503, based on the usage behavior security analysis data, to obtain harmful information sending rule feature data; according to the harmful information sending rule feature data, to obtain the currently defined label data of the user number; update the currently defined label data of the user number to obtain a secure user number.
[0155] In this embodiment, step S501 performs data statistics on text messages containing harmful information, specifically including three aspects of data: recording the sending time of each text message containing harmful information; sending user information: obtaining relevant information of the user who sent the harmful information, such as the number, account ID, etc.; collecting the text message content itself, mainly the text content containing harmful information; through the statistics of these data, the system can capture the sending patterns and rules of harmful information, providing a data basis for subsequent analysis; the collection of statistical data provides a basis for analyzing the propagation characteristics of harmful information. For example, whether a certain period is a high-incidence period of harmful information, or whether some users have the behavior of repeatedly sending harmful information, can effectively identify potential risk sources; the harmful information sending characteristic data helps in-depth analysis of user behavior in subsequent steps to identify potential malicious behaviors and information dissemination chains.
[0156] Step S502 will perform further usage behavior analysis based on the harmful information sending characteristic data collected in step S501, especially the "user numbers that frequently receive harmful information". The specific content of the analysis includes: analyzing whether the user frequently browses the harmful information in the text message; checking whether the user habitually opens and reads text messages containing harmful content; whether the user actively sends text messages containing harmful information to others; whether the user frequently receives harmful information; these behaviors are used to evaluate the overall security of the user number. By comprehensively analyzing these behaviors, the system can understand the interaction pattern between the user and harmful information; through detailed usage behavior analysis, the system can identify users who may inadvertently come into contact with and spread harmful information, or users who actively spread harmful content; comprehensive analysis of the user's behavior helps to identify abnormal patterns. For example, some users frequently receive, read, and forward harmful text messages, which may indicate that the user is a source or victim of harmful information; in-depth analysis of this behavior data provides an important basis for subsequent security assessment and labeling of user numbers.
[0157] Step S503 will use the usage behavior security analysis data as a basis to further derive the rule characteristic data of harmful information sending. These rule characteristic data describe the typical patterns of harmful information propagation, such as:
[0158] Sending frequency: Whether the user has the behavior of sending harmful information at a high frequency.
[0159] Propagation range: Whether the user often forwards harmful information to multiple other users.
[0160] Behavior relevance: Whether the user's behavior is associated with some typical harmful information propagation patterns.
[0161] Based on these rule feature data, the system assigns and updates the currently defined tag data for each user number. The tag definitions can be based on risk levels (such as high risk, low risk, etc.), or defined according to the user's behavior patterns (such as malicious spreaders, frequent receivers, etc.); the updated user tag data can provide decision-making support for subsequent security management, risk warning, etc.
[0162] By mining the rule feature data of harmful information transmission, the system can establish a more accurate harmful information propagation model, identify and predict potential harmful information spreaders; updating the user's security tags can reflect the user's behavior changes in real time. If the user's behavior pattern becomes unsafe or abnormal, the tag will be automatically adjusted to provide real-time warnings; this dynamic tag update based on behavior data helps to quickly identify risk users and take corresponding security measures according to the risk level, improving the system's response efficiency; it constitutes a user risk assessment and monitoring system based on behavior data, aiming to identify potential risk users through the analysis of harmful information propagation patterns and implement dynamic security management.
[0163] Based on the above-mentioned behavior analysis and dynamic tag mechanism, it not only improves the accuracy of risk identification, but also helps the system adjust security policies according to the user's real-time behavior, providing more efficient protection for information security.
[0164] In a preferred embodiment of the present invention, the above step S503 includes:
[0165] Step S5031, according to the usage behavior security analysis data, by: ; to obtain the abnormal degree of user number usage; where D u represents the abnormal degree of user number usage; f u represents the frequency of harmful information received by the user, that is, the number of harmful text messages received within the time window H; μ represents the average frequency of harmful information received by all users; σ represents the standard deviation of the frequency of harmful information received by all users;
[0166] Step S5032, set a preset threshold range, compare the abnormal degree of user number usage D u with the preset threshold range, and based on the comparison result, judge whether the user number usage behavior is normal or abnormal;
[0167] Step S5033, according to the usage behavior of the user number, define tags for the user number to obtain the currently defined tag data of the user number; the currently defined tag data of the user number includes two tags: normal and abnormal.
[0168] In this embodiment, step S503 identifies abnormal behavior by calculating the usage abnormality of the user number, and calculates the deviation degree between the harmful information frequency of the user number and the overall average frequency. When D u is a large value, it indicates that the user number has received an abnormally large number of harmful text messages within the time window H, indicating that the number may be under a concentrated attack or being used to spread harmful information; by counting the harmful information frequency, the system can quickly identify user numbers that frequently receive harmful text messages, facilitating further tracking and investigation; this indicator provides a quantitative measure of abnormal behavior for subsequent steps, helping to accurately screen potential victim numbers; it provides dynamic behavior analysis capabilities, facilitating adaptation to changes in harmful information in different time periods, making the detection ability of the system more flexible.
[0169] Step S5032, based on the user abnormality Du calculated in step S5031, in this step, it determines whether the user behavior is abnormal through a preset threshold range:
[0170] The system sets a threshold range to judge the behavior of the user number.
[0171] When Du exceeds the upper threshold, it is marked as abnormal behavior; if it is within the threshold range, it is regarded as normal behavior; when Du is lower than the upper threshold, it is marked that the data has fluctuations and is not accurate enough, and it returns to step S5031 to re-obtain the abnormality. This threshold can be dynamically adjusted according to historical data or different risk strategies to adapt to the security environment requirements of different situations; it effectively distinguishes normal user behavior from abnormal behavior, reduces the false alarm rate, and ensures that only truly suspicious numbers are marked; the dynamically adjusted threshold can optimize the judgment criteria according to the actual situation, making the detection system more flexible and adaptable; by providing a rule-based standard, it helps to efficiently filter out potential abnormal users from a large amount of user data.
[0172] Step S5033, based on the judgment result of step S5032, defines a behavior label for the user number:
[0173] If the behavior of the user number is judged to be normal, the label "normal" is assigned.
[0174] If it is judged to be abnormal behavior, the label "abnormal" is assigned.
[0175] The labeled data is used as a record of the system's definition of the user number behavior, facilitating subsequent statistical analysis, behavior tracking, or taking other further security measures.
[0176] Hierarchical management of user numbers is carried out in the form of tags, which facilitates the identification and monitoring of potential risk numbers; the tag data, as the basic data for subsequent security policies and behavior analysis, can provide effective historical records to help identify behavior patterns or abnormal behavior trends; the management of a large number of numbers is simplified, and the tagged information can more clearly display the behavior status of each number, improving the system operation and maintenance efficiency.
[0177] Step S5034: Based on the currently defined tag data of the user number, obtain the login information, purchase behavior, browsing records, or interaction records of the system platform used by the user number to obtain a tag source data set.
[0178] Step S5035: Based on the tag source data set, provide a prompt for the user to selectively clean tags; so that the user can clean and update the currently defined tag data to obtain an initial defined tag.
[0179] Step S5036: Based on the initial defined tag, make the current user number a safe number.
[0180] In this embodiment, step S5034 starts from the current tag data of the user number, collects the behavior data of the user on the system platform, such as login information, purchase behavior, browsing records, and interaction records, etc., to construct a tag source data set; by recording the login behavior of the user on different system platforms, to identify their activity and platform preferences; analyze the purchase records of the user to judge their consumption habits and frequencies, and identify possible abnormal consumption behaviors; collect the content browsing situation of the user on the platform to understand their interests and access frequencies, and discover abnormal browsing behaviors; including interactions with other users or platform content (such as likes, comments, forwards, etc.), which can reveal the behavior patterns of the user in social or content consumption; the data is integrated into a tag source data set to provide the specific source of the user's behavior, so as to more accurately understand the formation basis of the user tags; thus, it can more comprehensively understand the behavior patterns of the user number, provide a more detailed tag information source; through the establishment of the data set, it helps to identify potential abnormal behaviors and provides a reliable basis for subsequent security judgments; the refined behavior analysis enables the system to more accurately adjust and optimize the tag status of the user, improving the accuracy of the tags.
[0181] Step S5035 provides the user with optional prompts for cleaning tags based on the tag source dataset in step S5034; it will identify tag data that is inconsistent with the user's behavior or is significantly abnormal; the user can choose to "clean" or adjust the tags to delete, modify, or update tag information that no longer conforms to the actual situation; the cleaned tag data will be redefined and updated back to the "initially defined tag" state; through the selective cleaning function, the user has the right to decide whether to retain or update certain tags, enhancing the user's autonomy in data cleaning; maintaining the dynamic update of tag data to ensure that the system can always reflect the user's latest behavior information, improving the authenticity and effectiveness of tag data; effectively preventing mislabeled or outdated tag data from having a negative impact on the user's behavior evaluation and maintaining a high accuracy of system judgment.
[0182] Step S5036 marks user numbers that meet the security standards as "safe numbers" based on the initially defined tag state after user cleaning; the initially defined tags, after being cleaned, reflect the user's true behavior patterns and conform to normal or safe tag data; based on the initial state of these tag data, it is judged whether the user's behavior is normal, and the "safe number" status is marked for user numbers that meet the standards; this mark can be used for subsequent system management to ensure that safe numbers are not subject to further security interference or restrictions; thereby improving the accuracy of the system, avoiding mislabeling risk-free user numbers as abnormal numbers, and enhancing the user experience; enabling focus on processing truly risky user numbers and improving overall management efficiency; the definition of safe numbers provides users with a clearer security identification, facilitating subsequent tracking and management of number status.
[0183] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A method for identifying harmful information in SMS content based on big data, characterized in that, Including the following steps: S100, collect historical samples of various harmful information text messages to obtain a historical sample set of harmful information; compile the harmful information keywords in the sample set into multiple language versions to train an information recognition model and output the recognition results of harmful information features; The information recognition model includes a first module, a second module, and a third module; S200, input the text message data sent in real time by the sender into the information recognition model; The first module receives the text message data for scanning and reading to output initial suspicious text message data of harmful information; S300, transmit the text message data recognized by the first module to the second module, and the second module recognizes the suspicious character features of harmful information in the initial suspicious text message data of harmful information and obtains the meaning of the suspicious characters to output secondary suspicious text message data of harmful information; S400, input the text message data recognized by the second module into the third module, and the third module traces the IP address of the text message data sender to obtain the traceability information of the text message data sender, and outputs the final recognition result of the suspicious text message data of harmful information according to the traceability information; when it is recognized that there is text message data of harmful information, execute step S500; S500, perform a sending feature analysis on the text message data with harmful information to obtain harmful information sending rule feature data; Based on the harmful information sending rule feature data, obtain the currently defined label data of the user number and update the currently defined label of the user number to obtain a secure user number.
2. The method for identifying harmful information in SMS content based on big data according to claim 1, wherein, Collect historical samples of various harmful information text messages to obtain a historical sample set of harmful information; compile the historical sample set of harmful information into multiple language versions and train an information recognition model to output the recognition results of harmful information features; The information recognition model includes a first module, a second module, and a third module, including: Collect historical sample data of various text message sources with harmful information, and the historical sample data includes fraud text messages, terrorist blood and gore, pornographic filth, and virus links; Based on the historical sample data, convert the keywords with harmful information in the historical sample data into multiple target language versions; and perform semantic verification on the multiple target language versions of multiple keywords to obtain an accurate translation data set; Using the accurate translation dataset, by: , to obtain similar words for multiple keywords; where w represents the similar word feature parameter; K i represents the keyword; α represents the similarity threshold; based on the obtained similar words for the multiple keywords, by: ; to obtain the keyword similar semantic expansion set; and combine the keyword similar semantic expansion set with the keyword K i to form a comprehensive sample set; where represents the final multi - language similar keyword expansion set; m represents the number of similar keywords; represents the expansion function; Construct the comprehensive sample set into a training set: ; where n i represents the i-th accurate translation data item; input the training set G into the information recognition model for training to output the recognition result of harmful information features.
3. The method for identifying harmful information in SMS content based on big data according to claim 2, wherein, Input the text message data sent in real time by the sender into the information recognition model; The first module receives the text message data for scanning and reading to output initial suspicious text message data of harmful information, including: Input the newly sent text message data into the information recognition model, and the first module extracts the text message data with a similar format template reproduced multiple times to obtain similar template text data; Perform pre-screening on the similar template text data D by setting a set of format similarity rules {R i} n i = 1; by traversing each template text T in the SMS data j ; for each T j perform format feature analysis on the template text data to obtain the matching situation of the template text data format rule R i ; if T j does not conform to the R i format rule, retain it in the template text data with inconsistent formats, otherwise discard it; to filter out the template text data with unified formats; the template text data with unified formats is the initial harmful information suspicious SMS data.
4. A method for identifying harmful information in SMS content based on big data according to claim 3, characterized in that, Transmit the text message data recognized by the first module to the second module, and the second module recognizes the suspicious character features of harmful information in the initial suspicious text message data of harmful information and obtains the meaning of the suspicious characters to output secondary suspicious text message data of harmful information, including: Further read the short message text characters from the template text data with inconsistent formats to obtain a dataset of suspicious character features in the text; the suspicious character feature data includes multilingual misleading words; Identify and label each suspicious character feature data to obtain a corresponding category dataset; the corresponding category dataset includes labeling multilingual words as fraud category, violent and bloody category, pornographic and filthy category, and horror category; Based on the suspicious character feature data set, the template text data with inconsistent formats is parsed according to the context by: ; to determine that the character data item in the character feature data set is suspicious in meaning, where represents the final parsed meaning of the character data item; represents the meaning of the i-th template text segment; represents the template segment t i 's inference parameter; represents the corresponding category data set; represents the inference symbol; Based on the obtained judgment result, by: , to obtain the existence frequency of suspicious characters, where fi represents the appearance frequency of suspicious characters in the SMS data; C represents the number of occurrences of suspicious characters; W represents the total number of characters in the SMS data; and perform secondary screening on the SMS data with suspicious characters to obtain secondary harmful information suspicious SMS data.
5. A method for identifying harmful information in SMS content based on big data according to claim 4, characterized in that, Identify and label each suspicious character feature data to obtain a corresponding category dataset; the corresponding category dataset includes labeling multilingual words as fraud category, violent and bloody category, pornographic and filthy category, and horror category, including: Based on the corresponding category dataset, capture the suspicious words in the short message, and match the suspicious words in the short message data with the corresponding category dataset to obtain a matching dataset; based on the matching dataset, establish a sensitive word library for minors; Based on the sensitive word library for minors, replace the sensitive words or inappropriate words in the sensitive word library for minors with safe and appropriate pronouns for minors to obtain a pronoun dataset; By comparing the suspicious words existing in the short message data with the pronoun dataset and judging the suspicious degree of the pronouns, analyze the harmful risks of the suspicious pronouns to obtain risk short message information; Perform a similarity detection on the risk short message information and the sensitive words in the sensitive word library for minors to determine that the pronouns in the short message data are harmful information data.
6. The method for identifying harmful information in SMS content based on big data according to claim 5, characterized in that, Perform a similarity detection on the risk short message information and the sensitive words in the sensitive word library for minors to determine that the suspicious words in the short message data are harmful information data, including: According to the similarity detection, obtain the matching degree of minor sensitive words; based on the matching degree, perform a further fuzzy matching analysis on the risk short message information to obtain accurate matching data; Based on the accurate matching data, determine that the pronouns in the short message data are harmful information data, and set up an information feedback channel; Based on the feedback channel, timely feedback inappropriate information and automatically classify and prioritize the feedback information to update the sensitive word library for minors to obtain a feedback result.
7. A method for identifying harmful information in SMS content based on big data according to claim 6, characterized in that, Input the short message data identified by the second module into the third module. The third module traces the IP address of the short message data sending end to obtain the tracing information of the short message data sending end, and outputs the final identification result of harmful information suspicious short message data according to the tracing information, including: Trace the IP address according to the metadata of the short message data; the metadata includes the first eight digits of the sending number, the HTTP request header, and the domain name record information of the short message link; Based on the metadata, obtain the IP address of the short message data. When the short message IP address is an overseas address, the degree of suspicion of harmful information is high; and further judge in combination with the text information of the short message data; when the text information identifies words related to money and personal safety, issue a warning message to the user number at the receiving end in a timely manner; to identify the final harmful information suspicious short message data.
8. A method for identifying harmful information in SMS content based on big data according to claim 7, characterized in that Analyze the sending characteristics of the SMS data with harmful information to obtain the characteristic data of the harmful information sending rules; based on the characteristic data of the harmful information sending rules, obtain the currently defined label data of the user number, and update the currently defined label of the user number to obtain a secure user number, including: Statistically analyze the sending time, sending user information, and sending content of the SMS data with harmful information to obtain the characteristic data of the harmful information sending; Based on the characteristic data of the harmful information sending, analyze the usage behavior of the user numbers that frequently receive harmful information. The usage behavior of the user numbers includes: browsing, reading, sending, and receiving SMS data; to obtain the analysis data of the usage behavior security; Based on the analysis data of the usage behavior security, obtain the characteristic data of the harmful information sending rules; according to the characteristic data of the harmful information sending rules, obtain the currently defined label data of the user number; update the currently defined label data of the user number to obtain a secure user number.
9. A method for identifying harmful information in SMS content based on big data according to claim 8, characterized in that, According to the characteristic data of the harmful information sending rules, obtain the currently defined label data of the user number, including: According to the usage behavior security analysis data, by: ; to obtain the abnormal degree of user number usage; where D u represents the abnormal degree of user number usage; f u represents the frequency of harmful information received by the user, that is, the number of harmful text messages received within the time window H; μ represents the average frequency of harmful information received by all users; σ represents the standard deviation of the frequency of harmful information received by all users; Set a preset threshold range and compare the abnormal degree D of the user number usage u with the preset threshold range. Based on the comparison result, determine whether the user number usage behavior is a normal behavior or an abnormal behavior; Define labels for the user number according to the usage behavior of the user number to obtain the currently defined label data of the user number; the currently defined label data of the user number includes two labels: normal and abnormal.
10. A method for identifying harmful information in SMS content based on big data according to claim 9, characterized in that, Update the currently defined label data of the user number to obtain a secure user number, including: Based on the currently defined label data of the user number, obtain the login information, purchase behavior, browsing record, or interaction record of the system platform used by the user number to obtain the label source data set; Based on the label source data set, provide a prompt for the user to selectively clean the labels; so that the user can clean and update the currently defined label data to obtain the initialized defined label; Based on the initialized defined label, make the current user number a secure number.
Citation Information
Patent Citations
Reciprocating fuzz removing device for cleaning wild peaches
CN111109612A
Network information security detection and analysis system based on user management
CN117014883A