A SMS data intelligent management system and method based on big data

By recording and analyzing user's SMS filtering records, extracting and correcting keyword characteristics, establishing abnormal evaluation rules, and adapting filtering rules, the problem of high error detection rate of existing SMS filtering rules is solved, and the accuracy and user experience of spam recognition are improved.

CN119729374BActive Publication Date: 2025-05-20ZHONGWEI JUDAN DIGITAL TECH (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510242575.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing SMS filtering rules rely on preset keywords and cannot completely cover spam messages, resulting in high error detection rates and affecting the user's information acquisition experience.

Method used

By recording and analyzing user's SMS filtering records, extracting and correcting keyword characteristics, establishing exception evaluation rules, and adapting filtering rules to improve spam SMS recognition accuracy.

Benefits of technology

Reduce false detection, improve the accuracy of spam text message filtering and user's SMS acquisition experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119729374B_ABST
    Figure CN119729374B_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent management system and method for SMS data based on big data, and relates to the technical field of SMS data management. The management method comprises the following steps: after authorization by a user, recording the filtering process of each SMS received by a communication device, and generating a corresponding filtering record; extracting features of the set filtering rules, and classifying the SMS corresponding to the filtering records into categories; establishing an abnormal evaluation rule based on the presentation of each feature in the filtering record to evaluate the filtering situation; capturing each SMS browsing behavior of the user, summarizing each feature of the SMS browsed by the user, and revising the filtering rule; based on the revised filtering rule, extracting corresponding features of the received real-time SMS, performing an abnormal evaluation on the real-time SMS, and judging whether to push the real-time SMS to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of short message data management, and specifically to an intelligent management system and method for short message data based on big data. Background Art

[0002] Junk messages usually refer to messages with the nature of advertising, promotion, fraud, etc. sent to users without their consent or request. These junk messages often involve the collection and use of personal information. Frequently receiving junk messages will cause users to feel troubled and affect their normal life. Therefore, message filtering is set up in message management to help screen and process the received messages in order to identify and block unnecessary content such as junk messages and harassing information.

[0003] Most of the existing message filtering rules filter by preset keywords. Although it can reduce the push of junk messages to a certain extent, the keywords of junk messages will constantly change, resulting in the preset keywords not being able to completely cover all junk messages, or some normal messages may also contain the preset keywords, resulting in the inability to receive normal information. As the frequency of misdetection increases, it will also affect the user's information acquisition and usage experience. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent management system and method for short message data based on big data to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An intelligent management method for short message data based on big data, the management method includes the following steps:

[0006] Step S100: After obtaining user authorization, record the filtering process of each short message received by the communication device to generate corresponding filtering records; based on the filtering results presented by each filtering record, extract the characteristics of the set filtering rules.

[0007] Step S200: Analyze the feature inclusion situation of any filtering record, and classify the types of the short messages corresponding to the filtering record; based on the presentation situation of each feature in the filtering record, establish an abnormal evaluation rule to evaluate the filtering situation.

[0008] Step S300: Capture each short message browsing behavior of the user, analyze the association between the features included in the short messages browsed by the user and the filtering rules; summarize the features of the short messages browsed by the user, and correct the set filtering rules.

[0009] Step S400: Based on the corrected filtering rules, extract corresponding features from the received real-time SMS, and perform anomaly assessment on the real-time SMS; determine whether to push the real-time SMS to the user based on the anomaly assessment result.

[0010] Further, step S100 includes the following steps:

[0011] Step S101: Whenever a communication device of a user receives an SMS, read the content of the received SMS, retrieve the preset SMS filtering rules to judge the content of the SMS, obtain the filtering result of the received SMS, generate a filtering record to store the content of the received SMS and the filtering result; if the received SMS is not pushed to the user after SMS filtering, mark the filtering record of the received SMS as abnormal;

[0012] Step S102: Arbitrarily select a filtering record. If there is an abnormal mark in the filtering record, perform word segmentation on the content of the SMS in the filtering record to generate several keywords;

[0013] Step S103: Obtain all keywords of each filtering record with an abnormal mark, extract any two keywords for similarity comparison. If the obtained similarity exceeds the set similarity threshold, set the two keywords as the same keyword, and generate a corresponding keyword set for each category of the same keyword;

[0014] Step S104: Arbitrarily select a category of keyword sets, obtain the number of keywords included in the keyword set, and set the number of keywords in the i-th category of keyword sets as M i , calculate the quantity proportion α of the i-th category of keyword sets i = M i / A, where A is the number of filtering records with abnormal marks;

[0015] Step S105: Set a quantity proportion threshold as α max , arbitrarily select a filtering record with an abnormal mark, obtain the keyword set where the j-th keyword in the filtering record with the abnormal mark is located, and set the j-th keyword to correspond to the i-th category of keyword sets. If α i > α max , then set the j-th keyword as a feature of the SMS filtering rule. If the quantity proportion of any keyword corresponding to the keyword set is less than the quantity proportion threshold, select the keyword with the largest quantity proportion as a feature of the SMS filtering rule; summarize the features set in all filtering records with abnormal marks to obtain the feature set of the SMS filtering rule; through the adjustment of the SMS filtering rule for recognition, it is beneficial to the calculation of subsequent anomaly assessment and the development of feature adjustment.

[0016] Further, step S200 includes the following steps:

[0017] Step S201: Arbitrarily select a filtering record, extract keywords from the SMS content stored in the filtering record, and compare the similarity of each keyword with any feature in the feature set. If the obtained similarity exceeds the set similarity threshold, match the compared keyword with the keyword set corresponding to the compared feature, and obtain the proportion of the number of the corresponding keyword set;

[0018] Step S202: Obtain the proportion of the number of each keyword set matched by each keyword, sort each keyword set from largest to smallest according to the proportion, select the keyword set with the largest proportion as the first feature type of the filtering record, and so on. The keyword set with the rank of b is used as the b-th feature type of the filtering record; Different levels of feature types represent the influence degree of the corresponding keyword set, and are used as the judgment criteria for preferentially judging whether it is a spam message. The more the number of keywords involved in the SMS belongs to the first feature type, the higher the probability of being judged as a spam message;

[0019] Step S203: Arbitrarily select the keyword set with the rank of b, and the number of keywords corresponding to the filtering record is N b , set the proportion of the number of the keyword set with the rank of b as α b , establish an evaluation model:

[0020]

[0021] where b1 is a positive integer and b1 ∈ (1, e), e is the number of keyword sets to which the filtering record belongs, and N is the total number of keywords extracted from the SMS content of the filtering record; calculate the evaluation value P of the filtering record; Through the proportion of the number and the number of keywords that appear, both are the most direct criteria for judging whether it is a spam message. Based on the proportion of relevant keywords in the SMS content, it can help obtain a more accurate evaluation value and help users accurately identify whether it is a spam message;

[0022] Step S204: If there is an abnormal mark in the filtering record, set the evaluation value P of the filtering record as an abnormal evaluation value; obtain the abnormal evaluation values of all filtering records with abnormal marks, and select the smallest abnormal evaluation value P min as the abnormal evaluation rule for judging whether the SMS is filtered.

[0023] Further, step S300 includes the following steps:

[0024] Step S301: Whenever a user browses a text message, obtain the evaluation value P of the filtering record corresponding to the text message, and set the abnormal evaluation value for determining whether the text message is filtered to P min If P > P min then extract the text content in the text message and perform word segmentation on the text content to obtain a number of keywords;

[0025] Step S302: Compare each keyword with the feature set of the filtering rule to obtain a number of features included in the text message. Arbitrarily select the k-th feature, obtain the keyword set corresponding to the k-th feature, and obtain the proportion α of the number of the keyword set corresponding to the k-th feature k Obtain the number A of filtering records with abnormal marks. According to the formula:

[0026]

[0027] Calculate the corrected proportion α of the number of the keyword set corresponding to the k-th feature ’ k ; Set the proportion threshold to α max If α ’ k < α max then remove the k-th feature from the feature set of the filtering rule; If it is recognized that the user has browsed the text content of the text message identified as spam, it means that there is an error in the current filtering rule or the user has made a misclick behavior, then the proportion of the number of corresponding keywords needs to be adjusted and feature recognition needs to be performed again, which can help reduce the probability of identifying normal text messages as spam;

[0028] Step S303: When a text message is pushed to the user and the user does not browse the pushed text message, extract each keyword of the pushed text message, compare the similarity of any keyword with each feature. If the obtained similarity is less than the set similarity threshold, set the keyword as the expected feature; obtain the feature set and the expected feature set of the pushed text message;

[0029] Step S304: If the feature set of the pushed text message is empty, arbitrarily select an expected feature from the expected feature set, compare the similarity of the expected feature with the keywords in any filtering record, and obtain the number W of the same keywords corresponding to the expected feature sim According to the formula:

[0030]

[0031] where A is the number of filtering records with abnormal marks; calculate the corrected proportion α of the number of the same keywords corresponding to the expected feature ’; If α ’ > α max , then set the expected feature as a new feature of the filtering rule; if the user does not view the pushed SMS, adjustment is also required to help the user reduce the probability of misjudgment caused by new keywords in spam SMS and adaptively improve the recognition accuracy;

[0032] Step S305: Extract features from each SMS browsed by the user and pushed by the communication device to obtain a new feature set and overwrite the feature set of the filtering rule, and correct the original filtering rule.

[0033] Further, step S400 includes the following steps:

[0034] Step S401: When the user's communication device receives a real-time SMS in real time, obtain the SMS content of the real-time SMS to get several keywords, and compare each keyword with the feature set of the corrected filtering rule to generate several features included in the real-time SMS;

[0035] Step S402: Obtain the quantity proportion of the keyword set corresponding to each feature, and sort each feature from largest to smallest according to the quantity proportion; retrieve the evaluation model to evaluate the real-time SMS to obtain an evaluation value of P new ; Obtain the abnormal evaluation value P for judging whether the SMS is filtered min , if P new < P min , then push the real-time SMS to the user, if P new > P min , then filter the real-time SMS.

[0036] To better implement the above method, a smart management system for SMS data is also proposed. The management system includes a historical SMS analysis module, an SMS classification evaluation module, an SMS filtering adjustment module, and an SMS real-time analysis module;

[0037] The historical SMS analysis module is used to record the filtering process of each SMS received by the communication device after user authorization to generate corresponding filtering records; based on the filtering results presented by each filtering record, extract features from the set filtering rules;

[0038] The SMS classification evaluation module is used to analyze the feature inclusion of any filtering record, classify the types of SMS corresponding to the filtering record; based on the presentation of each feature in the filtering record, establish an abnormal evaluation rule to evaluate the filtering situation;

[0039] The SMS filtering adjustment module is used to capture each SMS browsing behavior of the user, analyze the correlation between the features contained in the SMS browsed by the user and the filtering rules; summarize the features of the SMS browsed by the user, and correct the set filtering rules;

[0040] The real-time SMS analysis module is used to extract corresponding features from the received real-time SMS based on the corrected filtering rules, and perform anomaly assessment on the real-time SMS; based on the anomaly assessment result, determine whether to push the real-time SMS to the user.

[0041] Furthermore, the historical SMS analysis module includes a historical filtering collection unit and a filtering rule extraction unit;

[0042] The historical filtering collection unit is used to record the filtering process of each SMS received by the communication device after the user authorizes, and generate corresponding filtering records; the filtering rule extraction unit is used to extract features of the set filtering rules based on the filtering results presented by each filtering record.

[0043] Furthermore, the SMS classification and assessment module includes an SMS type classification unit and an SMS anomaly assessment unit;

[0044] The SMS type classification unit is used to analyze the feature inclusion of any filtering record and classify the types of SMS corresponding to the filtering record; the SMS anomaly assessment unit is used to establish an anomaly assessment rule based on the presentation of each feature in the filtering record to evaluate the filtering situation.

[0045] Furthermore, the SMS filtering adjustment module includes a user behavior analysis unit and an identification rule correction unit;

[0046] The user behavior analysis unit is used to capture each SMS browsing behavior of the user, analyze the correlation between the features contained in the SMS browsed by the user and the filtering rules; the identification rule correction unit is used to summarize the features of the SMS browsed by the user and correct the set filtering rules.

[0047] Furthermore, the real-time SMS analysis module includes a real-time SMS analysis unit and an SMS identification and push unit;

[0048] The real-time SMS analysis unit is used to extract corresponding features from the received real-time SMS based on the corrected filtering rules, and perform anomaly assessment on the real-time SMS; the SMS identification and push unit is used to determine whether to push the real-time SMS to the user based on the anomaly assessment result.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] 1. By adaptively correcting the keywords for judging spam messages, the present invention can retrieve the optimal filtering rules at any time to identify spam messages, reduce the occurrence of misdetection caused by simple judgment of keywords, optimize the filtering of spam messages, and improve the user's SMS acquisition experience.

[0051] 2. The present invention converts simple keyword recognition into abnormal evaluation value calculation, and assigns different evaluation ratios according to the historical recognition situations of different keywords, which can avoid the situation that normal SMS messages are directly filtered because keywords appear in them, and effectively improve the filtering accuracy of spam messages.

[0052] 3. By improving the simple keyword comparison in conventional technical means, the present invention improves the recognition and filtering accuracy of spam messages, better helps users screen valid information, and improves the user's information acquisition experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the steps of a method for intelligent management of SMS data based on big data;

[0054] Figure 2 It is a schematic diagram of the structure of a system for intelligent management of SMS data based on big data. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment: As Figures 1 to 2 shown, the present invention provides a method for intelligent management of SMS data based on big data, and the management method includes the following steps:

[0057] Step S100: After obtaining user authorization, record the filtering process of each SMS message received by the communication device to generate corresponding filtering records; based on the filtering results presented by each filtering record, extract the features of the set filtering rules;

[0058] Among them, step S100 includes the following steps:

[0059] Step S101: Whenever a communication device of a user receives a text message, read the content of the received text message, retrieve a preset text message filtering rule to judge the content of the text message, obtain a filtering result of the received text message, generate a filtering record to store the content of the received text message and the filtering result; if the received text message is not pushed to the user after text message filtering, mark the filtering record of the received text message as abnormal;

[0060] Step S102: Arbitrarily select a filtering record. If there is an abnormal mark in the filtering record, perform word segmentation on the content of the text message in the filtering record to generate a number of keywords;

[0061] Step S103: Obtain all the keywords of each filtering record with an abnormal mark, extract any two keywords for similarity comparison. If the obtained similarity exceeds the set similarity threshold, set the two keywords as the same keyword, and generate a corresponding keyword set for each category of the same keyword;

[0062] Step S104: Arbitrarily select a category of keyword sets, obtain the number of keywords included in the keyword set, and set the number of keywords in the i-th category of keyword sets as M i , calculate the quantity proportion α of the i-th category of keyword sets i =M i / A, where A is the number of filtering records with an abnormal mark;

[0063] Step S105: Set a quantity proportion threshold as α max , arbitrarily select a filtering record with an abnormal mark, obtain the keyword set where the j-th keyword in the filtering record with an abnormal mark is located, and set the j-th keyword to correspond to the i-th category of keyword sets. If α i >α max , then set the j-th keyword as a feature of the text message filtering rule. If the quantity proportion of any keyword corresponding to the keyword set is less than the quantity proportion threshold, select the keyword with the largest quantity proportion and set it as a feature of the text message filtering rule; summarize the features set in all the filtering records with an abnormal mark to obtain a feature set of the text message filtering rule.

[0064] Step S200: Analyze the feature inclusion situation of any filtering record, and classify the types of text messages corresponding to the filtering record; based on the presentation situation of each feature in the filtering record, establish an abnormal evaluation rule to evaluate the filtering situation;

[0065] Among them, Step S200 includes the following steps:

[0066] Step S201: Arbitrarily select a filtering record, extract keywords from the SMS content stored in the filtering record, and compare the similarity of each keyword with any feature in the feature set. If the obtained similarity exceeds the set similarity threshold, match the compared keyword with the keyword set corresponding to the compared feature, and obtain the proportion of the number of the corresponding keyword set;

[0067] Step S202: Obtain the proportion of the number of each keyword set matched by each keyword, sort each keyword set from largest to smallest according to the proportion of the number, select the keyword set with the largest proportion of the number as the first feature type of the filtering record, and so on, and use the keyword set with the rank of b as the b-th feature type of the filtering record;

[0068] Step S203: Arbitrarily select a keyword set with the rank of b, and the number of keywords corresponding to the filtering record is N b , set the proportion of the number of the keyword set with the rank of b as α b , and establish an evaluation model:

[0069]

[0070] where b1 is a positive integer and b1 ∈ (1, e), e is the number of keyword sets to which the filtering record belongs, and N is the total number of keywords extracted from the SMS content of the filtering record; calculate the evaluation value P of the filtering record;

[0071] Example 1: 10 keywords are extracted from a filtering record, among which 5 keywords are respectively matched with 3 keyword sets, 2 of which correspond to the first keyword set, 2 correspond to the second keyword set, and 1 corresponds to the third keyword set; set the proportion of the number of the three keyword sets as 90%, 70% and 50% respectively, and calculate the evaluation value P = 5 / 10×(2 3 ×90% + 2 2 ×70% + 1 1 ×50%) = 0.5×10.5 = 5.25; set the abnormal evaluation value as 5, 5.25 > 5, so the SMS corresponding to the filtering record will be directly filtered;

[0072] Step S204: If there is an abnormal mark in the filtering record, set the evaluation value P of the filtering record as the abnormal evaluation value; obtain the abnormal evaluation values of all filtering records with abnormal marks, and select the smallest abnormal evaluation value P min as the abnormal evaluation rule for judging whether the SMS is filtered.

[0073] Step S300: Capture each SMS browsing behavior of the user, analyze the association between the features contained in the SMS browsed by the user and the filtering rules; summarize the various features of the SMS browsed by the user, and correct the set filtering rules;

[0074] Among them, step S300 includes the following steps:

[0075] Step S301: Whenever the user browses an SMS, obtain the evaluation value P of the filtering record corresponding to the SMS, and set the abnormal evaluation value for judging whether the SMS is filtered as P min , if P > P min , then extract the SMS content in the SMS and perform word segmentation on the SMS content to obtain a number of keywords;

[0076] Step S302: Compare each keyword with the feature set of the filtering rules to obtain the several features contained in the SMS. Arbitrarily select the k-th feature, obtain the keyword set corresponding to the k-th feature, and obtain the proportion α of the number of the keyword set corresponding to the k-th feature; k , obtain the number A of filtering records with abnormal marks. According to the formula:

[0077]

[0078] Calculate the corrected proportion α of the number of the keyword set corresponding to the k-th feature ’ k ; Set the proportion threshold as α max , if α ’ k < α max , then remove the k-th feature from the feature set of the filtering rules;

[0079] Example 2: Set that a normal filtering record contains 1 feature, the proportion of the number of the keyword set corresponding to the feature is 60%, and set the number of filtering records with abnormal marks as 10. Therefore, the corrected proportion is (60% × 10 - 1) / (10 - 1) = 55.55%. If the proportion threshold is 56%, because 55.55% < 56%, the contained feature can be removed from the filtering rules;

[0080] Step S303: When an SMS is pushed to the user and the user does not browse the pushed SMS, extract each keyword of the pushed SMS, compare the similarity of any keyword with each feature. If the obtained similarity is less than the set similarity threshold, set the keyword as the expected feature; obtain the feature set and the expected feature set of the pushed SMS;

[0081] Step S304: If the feature set of the pushed SMS is empty, randomly select an expected feature from the expected feature set, compare the similarity between the expected feature and the keywords in any filtering record, and obtain the number of identical keywords corresponding to the expected feature as W sim , according to the formula:

[0082]

[0083] where A is the number of filtering records with abnormal marks; calculate the corrected quantity proportion of the identical keywords corresponding to the expected feature as α ’ ; if α ’ > α max , then set the expected feature as a new feature of the filtering rule;

[0084] Step S305: Extract features from each SMS browsed by the user and pushed by the communication device, obtain a new feature set to cover the feature set of the filtering rule, and correct the original filtering rule.

[0085] Step S400: Based on the corrected filtering rule, extract corresponding features from the received real-time SMS, and perform anomaly assessment on the real-time SMS; judge whether to push the real-time SMS to the user based on the anomaly assessment result;

[0086] Among them, Step S400 includes the following steps:

[0087] Step S401: When the user's communication device receives a real-time SMS in real time, obtain the SMS content of the real-time SMS to get several keywords, compare each keyword with the feature set of the corrected filtering rule, and generate several features included in the real-time SMS;

[0088] Step S402: Obtain the quantity proportion of the keyword set corresponding to each feature, and sort each feature from largest to smallest according to the quantity proportion; call the evaluation model to evaluate the real-time SMS to obtain an evaluation value of P new ; obtain the anomaly evaluation value P for judging whether the SMS is filtered min , if P new < P min , then push the real-time SMS to the user, if P new > P min , then filter the real-time SMS.

[0089] An intelligent management system for SMS data, the management system includes a historical SMS analysis module, an SMS classification and evaluation module, an SMS filtering and adjustment module, and an SMS real-time analysis module;

[0090] A historical SMS analysis module, which is used to record the filtering process of each SMS received by a communication device after user authorization, generate corresponding filtering records, and extract features of the set filtering rules based on the filtering results presented by each filtering record.

[0091] An SMS classification and evaluation module, which is used to analyze the feature inclusion of any filtering record, classify the types of SMS corresponding to the filtering record, and establish an abnormal evaluation rule to evaluate the filtering situation based on the presentation of each feature in the filtering record.

[0092] An SMS filtering adjustment module, which is used to capture each SMS browsing behavior of a user, analyze the correlation between the features included in the SMS browsed by the user and the filtering rules, summarize the features of the SMS browsed by the user, and correct the set filtering rules.

[0093] An SMS real-time analysis module, which is used to extract corresponding features of the received real-time SMS based on the corrected filtering rules, evaluate the abnormality of the real-time SMS, and determine whether to push the real-time SMS to the user based on the abnormal evaluation result.

[0094] Among them, the historical SMS analysis module includes a historical filtering collection unit and a filtering rule extraction unit.

[0095] The historical filtering collection unit is used to record the filtering process of each SMS received by a communication device after user authorization, generate corresponding filtering records. The filtering rule extraction unit is used to extract features of the set filtering rules based on the filtering results presented by each filtering record.

[0096] Among them, the SMS classification and evaluation module includes an SMS type classification unit and an SMS abnormality evaluation unit.

[0097] The SMS type classification unit is used to analyze the feature inclusion of any filtering record and classify the types of SMS corresponding to the filtering record. The SMS abnormality evaluation unit is used to establish an abnormal evaluation rule to evaluate the filtering situation based on the presentation of each feature in the filtering record.

[0098] Among them, the SMS filtering adjustment module includes a user behavior analysis unit and an identification rule correction unit.

[0099] The user behavior analysis unit is used to capture each SMS browsing behavior of a user and analyze the correlation between the features included in the SMS browsed by the user and the filtering rules. The identification rule correction unit is used to summarize the features of the SMS browsed by the user and correct the set filtering rules.

[0100] Among them, the real-time SMS analysis module includes a real-time SMS analysis unit and an SMS recognition and push unit;

[0101] The real-time SMS analysis unit is used to extract corresponding features from the received real-time SMS based on the corrected filtering rules and perform anomaly evaluation on the real-time SMS; the SMS recognition and push unit is used to determine whether to push the real-time SMS to the user based on the anomaly evaluation result.

[0102] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for intelligent management of SMS data based on big data, characterized in that: The management method comprises the following steps: Step S100: after user authorization, the filtering process of each SMS message received by the communication device is recorded to generate a corresponding filtering record; based on the filtering results presented by each filtering record, feature extraction is performed on the set filtering rules; Step S200: Analyze the feature inclusion of any filtering record, and classify the short messages corresponding to the filtering record into categories; and establish an abnormal evaluation rule to evaluate the filtering situation based on the presentation of each feature in the filtering record; Step S300: Capture each SMS browsing behavior of the user, analyze the correlation between the features contained in the SMS browsed by the user and the filtering rules; summarize the features of the SMS browsed by the user, and modify the set filtering rules; Step S400: Based on the modified filtering rules, extract corresponding features from the received real-time SMS, and perform abnormality assessment on the real-time SMS; and determine whether to push the real-time SMS to the user based on the abnormality assessment result.

2. The method for intelligent management of SMS data based on big data according to claim 1, characterized in that: The step S100 includes the following steps: Step S101: Whenever a user's communication device receives a text message, the text message content of the received text message is read, and a preset text message filtering rule is called to judge the text message content to obtain a filtering result of the received text message, and a filtering record is generated to store the text message content and filtering result of the received text message; if the received text message is not pushed to the user after the text message filtering, the filtering record of the received text message is marked as abnormal; Step S102: randomly selecting a filtering record, and if there is an abnormal mark in the filtering record, performing word segmentation processing on the SMS content in the filtering record to generate a number of keywords; Step S103: Acquire all keywords of each filtering record with an abnormal mark, extract any two keywords for similarity comparison, and if the obtained similarity exceeds a set similarity threshold, set the two keywords as the same keywords, and generate a corresponding keyword set for each type of the same keywords; Step S104: arbitrarily select a keyword set, obtain the number of keywords contained in the keyword set, and set the number of keywords in the i-th keyword set as M. i , calculate the proportion of the number of keyword sets in the i-th category α i =M i / A, where A is the number of filtered records with abnormal marks; Step S105: Set a quantity ratio threshold as α max , randomly select a filter record with an abnormal mark, obtain the keyword set where the jth keyword in the filter record with the abnormal mark is located, set the jth keyword to correspond to the i-th keyword set, if α i >α max , then the j-th keyword is set as a feature of the SMS filtering rule. If the quantity ratio of the keyword set corresponding to any keyword is less than the quantity ratio threshold, then the keyword with the largest quantity ratio is selected as a feature of the SMS filtering rule; the features set in all the filtering records with abnormal marks are summarized to obtain the feature set of the SMS filtering rule.

3. The method for intelligent management of SMS data based on big data according to claim 2, characterized in that: The step S200 includes the following steps: Step S201: randomly select a filtering record, extract keywords from the SMS content stored in the filtering record, and compare the similarity between each keyword and any feature in the feature set. If the obtained similarity exceeds the set similarity threshold, match the compared keyword with the keyword set corresponding to the compared feature, and obtain the quantity ratio of the corresponding keyword set. Step S202: Obtain the proportion of the number of keyword sets matched by each keyword, sort the keyword sets from large to small according to the proportion of the number, select the keyword set with the largest proportion of the number as the first feature type of the filtering record, and so on, select the keyword set with the position b as the bth feature type of the filtering record; Step S203: arbitrarily select a keyword set with position b and the number of keywords corresponding to the filtering record is N. b , set the proportion of the keyword set with position b to α b , build an evaluation model: Wherein, b1 is a positive integer and b1∈(1,e), e is the number of keyword sets to which the filtering record belongs, and N is the total number of keywords extracted from the SMS content of the filtering record; the evaluation value P of the filtering record is calculated; Step S204: If there is an abnormal mark in the filtering record, the evaluation value P of the filtering record is set as the abnormal evaluation value; the abnormal evaluation values ​​of all filtering records with abnormal marks are obtained, and the abnormal evaluation value P with the smallest value is selected. min As an exception assessment rule to determine whether a text message is filtered.

4. The method for intelligent management of SMS data based on big data according to claim 3, characterized in that: The step S300 includes the following steps: Step S301: Whenever a user browses a text message, obtain the evaluation value P of the filtering record corresponding to the text message, and set the abnormal evaluation value for determining whether the text message is filtered to P min , if P>P min , extract the SMS content in the SMS and segment the SMS content to obtain several keywords; Step S302: Compare each keyword with the feature set of the filtering rule to obtain several features contained in the text message, arbitrarily select the kth feature, obtain the keyword set corresponding to the kth feature, and obtain the number of keyword sets corresponding to the kth feature as α k , get the number of filtered records with abnormal marks as A, according to the formula: Calculate the proportion of the corrected number of keyword sets corresponding to the kth feature α ’ k ; Set the quantity ratio threshold to α max , if α ’ k <α max , then the k-th feature is removed from the feature set of the filtering rule; Step S303: when there is a text message pushed to the user, and the user does not browse the pushed text message, each keyword of the pushed text message is extracted, and a similarity comparison is performed between any keyword and each feature. If the obtained similarity is less than a set similarity threshold, the keyword is set as the expected feature; and a feature set of the pushed text message and an expected feature set are obtained; Step S304: If the feature set of the pushed SMS is empty, then select any desired feature from the desired feature set, compare the desired feature with the keywords in any filtering record for similarity, and obtain the number of identical keywords corresponding to the desired feature as W. sim , according to the formula: Where A is the number of filtered records with abnormal marks; the proportion of the number of corrections corresponding to the same keyword of the expected feature is calculated as α ’ ; If α ’ >α max , then the expected feature is set as a new feature of the filtering rule; Step S305: extract features from each SMS message browsed by the user and pushed by the communication device, obtain a new feature set and overwrite the feature set of the filtering rule, and modify the original filtering rule.

5. The method for intelligent management of SMS data based on big data according to claim 4, characterized in that: The step S400 includes the following steps: Step S401: when a user's communication device receives a real-time text message in real time, the text message content of the real-time text message is obtained to obtain a number of keywords, and each keyword is compared with a feature set of the modified filtering rule to generate a number of features contained in the real-time text message; Step S402: Obtain the quantity ratio of the keyword set corresponding to each feature, and sort the features from large to small according to the quantity ratio; call the evaluation model to evaluate the real-time SMS to obtain an evaluation value of P new ; Get the abnormal evaluation value P to determine whether the text message is filtered min , if P new <P min , then push the real-time SMS to the user, if P new >P min , the real-time short message is filtered.

6. An intelligent SMS data management system, used to execute an intelligent SMS data management method based on big data as claimed in any one of claims 1 to 5, characterized in that: The management system includes a historical SMS analysis module, a SMS classification evaluation module, a SMS filtering adjustment module and a SMS real-time analysis module; The historical SMS analysis module is used to record the filtering process of each SMS received by the communication device after user authorization, and generate corresponding filtering records; based on the filtering results presented by each filtering record, feature extraction is performed on the set filtering rules; The SMS classification evaluation module is used to analyze the feature inclusion of any filtering record and classify the SMS corresponding to the filtering record into categories; based on the presentation of each feature in the filtering record, an abnormal evaluation rule is established to evaluate the filtering situation; The SMS filtering adjustment module is used to capture each SMS browsing behavior of the user, analyze the correlation between the features contained in the SMS browsed by the user and the filtering rules; summarize the features of the SMS browsed by the user, and modify the set filtering rules; The SMS real-time analysis module is used to extract corresponding features of the received real-time SMS based on the revised filtering rules and perform abnormality assessment on the real-time SMS; Based on the abnormality assessment result, it is determined whether to push the real-time SMS to the user.

7. The intelligent management system for short message data according to claim 6, characterized in that: The historical SMS analysis module includes a historical filtering collection unit and a filtering rule extraction unit; The historical filtering collection unit is used to record the filtering process of each text message received by the communication device after user authorization and generate corresponding filtering records; the filtering rule extraction unit is used to extract features of the set filtering rules based on the filtering results presented by each filtering record.

8. The intelligent management system for short message data according to claim 6, characterized in that: The SMS classification and evaluation module includes an SMS type classification unit and an SMS anomaly evaluation unit; The SMS type classification unit is used to analyze the feature inclusion of any filtering record and classify the SMS corresponding to the filtering record; the SMS anomaly evaluation unit is used to establish anomaly evaluation rules to evaluate the filtering situation based on the presentation of each feature in the filtering record.

9. The intelligent management system for short message data according to claim 6, characterized in that: The SMS filtering adjustment module includes a user behavior analysis unit and an identification rule correction unit; The user behavior analysis unit is used to capture each SMS browsing behavior of the user and analyze the correlation between the features contained in the SMS browsed by the user and the filtering rules; the identification rule correction unit is used to summarize the various features of the SMS browsed by the user and correct the set filtering rules.

10. The intelligent management system for short message data according to claim 6, characterized in that: The SMS real-time analysis module includes a real-time SMS analysis unit and a SMS identification and push unit; The real-time SMS analysis unit is used to extract corresponding features of the received real-time SMS based on the revised filtering rules and perform abnormality assessment on the real-time SMS; The SMS identification and pushing unit is used to determine whether to push the real-time SMS to the user based on the abnormality assessment result.

Citation Information

Patent Citations

  • Naive Bayesian classification based mobile phone spam short message filtering method and system

    CN103634473A

  • Short message filtering system and method based on artificial intelligence

    CN116887265A