Privacy information sensitivity grading method and system, electronic equipment and storage medium
By obtaining and preprocessing social media hot search data, identifying privacy words and calculating sensitivity parameters, and using the weighted sliding window normalization method, the problem of inaccurate evaluation of privacy information sensitivity in the existing technology is solved, accurate and dynamic sensitivity division is achieved, and privacy protection capabilities are improved.
Patent Information
- Application Number
- CN202510023677.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately evaluate the sensitivity of private information, and cannot effectively deal with the potential impact and dynamic changes after information leakage, and lacks multi-dimensional evaluation standards and reflections on personalized privacy protection needs.
By obtaining hot search data, preprocessing and privacy word recognition, calculating frequency and information influence parameters, and using the weighted sliding window normalization method, the sensitivity level of privacy information is dynamically divided.
It realizes the accurate and dynamic division of the sensitivity of private information, can accurately reflect the true situation of information dissemination in social networks, improves privacy protection capabilities, and provides innovative solutions for privacy information management and data security.
Smart Images

Figure CN119938925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social network analysis, and in particular to a privacy information sensitivity grading method, system, electronic device and storage medium. Background Art
[0002] With the continuous development of information technology, the popularity of the Internet and social media platforms has enabled a large amount of user-generated content to spread rapidly around the world. Especially on social platforms such as Weibo, users frequently share their personal lives and emotional states, and publish content involving sensitive areas such as personal privacy, health, and finance. The leakage of privacy information has become a hot social issue, which not only affects individual privacy protection, but may also bring broader social and legal risks. Therefore, how to automatically identify, classify, and evaluate the sensitivity of privacy information in large-scale data has become an important issue that needs to be solved in the field of data security and privacy protection.
[0003] At present, the automatic identification and classification of private information is usually based on dictionary matching, rule engines or machine learning methods. Although these technologies can identify private information to a certain extent, they have some limitations, such as: the inability to accurately assess the impact of information leakage; the risk assessment of different types of private information is too simple and lacks multi-dimensional assessment standards; the sensitivity level classification lacks dynamic adjustment and fails to consider changes in social and time factors; personalized privacy protection needs are rarely considered, and the privacy sensitivity and preference differences of different users are not effectively reflected. Therefore, how to provide more accurate, comprehensive and intelligent privacy information sensitivity assessment has become the key to improving privacy protection capabilities.
[0004] Existing technical solutions mainly classify privacy information in the following ways:
[0005] Dictionary-based matching method: By pre-building a dictionary containing sensitive information vocabulary, sensitive information in the text is matched. For example, for sensitive information such as ID card number, bank card number, and phone number, a special regular expression or dictionary is built for matching. This method is more effective for identifying known sensitive information, but has poor recognition effect on emerging privacy information (such as network health information, emotional information, etc.).
[0006] Rule engine: detects private information based on manually written rules. Rules usually combine multiple data features, such as words and context, to determine whether the information is private. The advantage of rule engines is that they can be customized according to business needs, but they lack versatility and flexibility and are difficult to deal with complex and changing text information.
[0007] Machine learning methods: Through training models (such as SVM, decision trees, neural networks, etc.), a large amount of labeled data is used to identify private information. In recent years, natural language processing (NLP) models based on deep learning, such as BERT, have been widely used in the identification and classification of private information, with high recognition accuracy. However, these methods usually ignore the potential impact and dynamic changes after information leakage, and it is difficult to accurately assess the true sensitivity of private information.
[0008] Existing private information propagation models often rely too much on static network structures and simple propagation rules, and cannot effectively deal with the dynamic characteristics of information propagation in social networks, especially when dealing with highly sensitive information, which lacks precise control. In addition, existing models are relatively simplified in the diffusion mechanism of sensitive information, and fail to consider the complex interactions between user behavior, social network structure and information sensitivity, resulting in the inability to accurately simulate the propagation of private information. Summary of the invention
[0009] In view of the shortcomings existing in the above problems, the present invention provides a privacy information sensitivity grading method, system, electronic device and storage medium.
[0010] To achieve the above object, the present invention provides a privacy information sensitivity grading method, comprising:
[0011] Get hot search data;
[0012] Preprocessing the hot search data;
[0013] Splitting the continuous Chinese text in the preprocessed hot search data into independent words, and identifying the privacy words;
[0014] Based on the privacy words, calculate the privacy information parameters;
[0015] The private information is classified into levels based on the private information parameters.
[0016] Preferably, a crawler is used to obtain the hot search data on social media, wherein the hot search data includes hot search content, hot search ranking time, popularity, and number of publisher fans.
[0017] Preferably, preprocessing the hot search data includes:
[0018] Performing data cleaning on the hot search data to remove redundant data, erroneous data and irrelevant data in the hot search data;
[0019] Standardizing the cleaned hot search data text;
[0020] The stop words of the hot search data are removed after text normalization.
[0021] Preferably, the jieba tool in python is used to split the words into the independent words, and the SVM model is used to identify the privacy words.
[0022] Preferably, the privacy information parameters include a calculation frequency parameter and an information influence parameter;
[0023] The calculation formula of the frequency parameter is:
[0024]
[0025] Where: N is the sliding time window; i is the information; t k is the time step; ω k is the weight; f i (t k ) is the sliding time window N and the information i at each time step t in the sliding window k The frequency of max (t k ) for each time step t k The maximum frequency of all private information in the
[0026] The calculation formula of the information influence parameter S is:
[0027] S=ω H *H' unit,i +ω I *I i ;
[0028] Where: H is the heat value influence parameter; ω I is the publisher's influence parameter; Hunit,i' is the normalized heat value; I i Influence for publishers.
[0029] Preferably, the formula for the normalized heat value is:
[0030]
[0031] Where: T i Time for being on the trending search list; is the time decay function; H i The popularity of the entry where the private information is located.
[0032] This application also provides a privacy information sensitivity grading system, including:
[0033] Acquisition module, used to obtain hot search data;
[0034] A preprocessing module, used for preprocessing the hot search data;
[0035] A recognition module, used to split the continuous Chinese text in the pre-processed hot search data into independent words and identify the privacy words;
[0036] A calculation module, used for calculating the privacy information parameter based on the privacy words;
[0037] A classification module is used to classify the private information into levels based on the private information parameters.
[0038] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0039] The present invention also provides a storage medium storing a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above method.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] This invention breaks through the limitations of traditional methods in the evaluation of privacy information sensitivity. Through a multi-dimensional evaluation system (including frequency, popularity, publisher influence, etc.) and a weighted sliding window normalization method, the sensitivity classification of privacy information is more accurate, dynamic, and has practical application value. This method can accurately reflect the real situation of information dissemination in social networks, improve the ability of privacy protection, and provide an innovative solution for privacy information management and data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of the privacy information sensitivity grading method of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] Reference Figure 1 The present invention provides a privacy information sensitivity grading method, comprising:
[0045] Get hot search data;
[0046] Specifically, in the data collection and preparation stage, we first use crawlers to obtain the hot search data on social media in the past six months. These data include hot search content, time on the hot search list, popularity, number of publishers, etc., and save the crawling results in a csv file.
[0047] Preprocess the hot search data;
[0048] Specifically, data preprocessing is a key step in the entire privacy information identification and analysis process, which aims to improve the quality and consistency of the original data and lay a solid foundation for subsequent word segmentation, privacy word identification and statistical analysis. The process of this step is as follows:
[0049] Data cleaning: remove redundancy, errors and irrelevant information in the data to ensure data accuracy and consistency. Specifically, use Pandas in Python to first remove duplicate content of the crawled hot search terms, then process the missing values of the processed data, and finally remove noise in the data, such as advertising information, irrelevant content, special characters, etc.
[0050] Text standardization: Remove extra spaces, tabs, and line breaks from the text to ensure the neatness and consistency of the text. Specific work includes unified encoding and removal of extra whitespace characters.
[0051] Remove stop words: Remove common but meaningless words in the text to improve the efficiency and accuracy of subsequent analysis. The specific work is to set up a stop word list and filter stop words based on this list.
[0052] Split the continuous Chinese text in the preprocessed hot search data into independent words and identify the privacy words;
[0053] Specifically, in the word segmentation stage, the preprocessed continuous Chinese text is split into independent words to facilitate the subsequent identification and analysis of privacy words. The jieba tool in Python is used to complete the word segmentation. Because our goal is to identify privacy information, and the default dictionary of jieba cannot fully cover the words involving privacy information, we also need to create a custom dictionary file and use the jieba.load_userdict() function to read it.
[0054] In the privacy word recognition and classification steps, support vector machine (SVM) classification technology is used to ensure high accuracy and reliability. First, a training data set containing various types of privacy information is collected and labeled. Then, the text data is converted into numerical feature vectors through feature extraction methods. Then, these feature vectors are used to train the SVM model and optimize its parameters to improve classification performance. Finally, accurate identification of privacy information is completed.
[0055] Based on the privacy words, calculate the privacy information parameters;
[0056] Specifically, based on the identification results of privacy information, combined with the information dissemination rules in social networks, the parameter value of the privacy information is calculated to lay a solid foundation for the subsequent classification of privacy information levels.
[0057] The sensitivity of privacy information is graded by calculating the frequency parameter F and the information influence parameter S.
[0058] When calculating the frequency parameter F, the weighted sliding window normalization method is used to dynamically reflect the propagation trend and timeliness of private information by combining the frequency of private information within a certain period of time with the time decay factor, thereby improving the accuracy of sensitivity assessment. The specific steps are as follows:
[0059] Determine the sliding time window N and the information i at each time step t in the sliding window k The frequency f i (t k ).
[0060] For each time step t k Define the weight ω k =e -λ*(N-t) , where λ is the attenuation coefficient.
[0061] Calculate the weighted frequency of all time steps and normalize them to get the frequency parameter where f max (t k ) is each time step t k The maximum frequency of all private information in the data.
[0062] When calculating the information influence parameter S, the characteristics of the social network should be comprehensively considered. Therefore, the influence parameter S should be calculated based on the popularity of the term H where the private information is located. i , time of hot search list T i , and the publisher’s influence I i To calculate. Since the speed of information propagation in social networks decays over time, a time decay function e must also be considered -λTi , which is used to more accurately reflect the heat value per unit time. The specific steps are as follows:
[0063] Calculate the heat value of private information i per unit time
[0064] Normalize the heat per unit time to get the normalized heat
[0065] Calculate the influence of the publisher where n iIndicates the number of fans of the publisher of the i-th information.
[0066] Calculate the information influence parameter S = ω H *H′ unit,i +ω I *I i .
[0067] The privacy information is classified into levels based on the privacy information parameters.
[0068] Specifically, in order to facilitate the construction of subsequent models, we need to classify each piece of private information according to the frequency parameter F and information influence parameter S we calculated. We set thresholds f1, f2...fn for F and thresholds s1, s2...sn for S. Classification is based on the threshold range of F and S. The classification criteria are shown in Table 1:
[0069]
[0070] This invention breaks through the limitations of traditional methods in the evaluation of privacy information sensitivity. Through a multi-dimensional evaluation system (including frequency, popularity, publisher influence, etc.) and a weighted sliding window normalization method, the sensitivity classification of privacy information is more accurate, dynamic, and has practical application value. This method can accurately reflect the real situation of information dissemination in social networks, improve the ability of privacy protection, and provide an innovative solution for privacy information management and data security.
[0071] This application also provides a privacy information sensitivity grading system, including:
[0072] Acquisition module, used to obtain hot search data;
[0073] Preprocessing module, used to preprocess hot search data;
[0074] A recognition module is used to split the continuous Chinese text in the pre-processed hot search data into independent words and identify the privacy words;
[0075] A calculation module, used for calculating the privacy information parameters based on the privacy words;
[0076] The classification module is used to classify the privacy information based on the privacy information parameters.
[0077] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0078] The present invention also provides a storage medium storing a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device executes the above method.
[0079] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A privacy information sensitivity grading method, characterized in that: include: Get hot search data; Preprocessing the hot search data; Splitting the continuous Chinese text in the preprocessed hot search data into independent words, and identifying the privacy words; Based on the privacy words, calculate the privacy information parameters; The private information is classified into levels based on the private information parameters.
2. The privacy information sensitivity grading method according to claim 1, characterized in that: Use a crawler to obtain the hot search data on social media, where the hot search data includes hot search content, hot search ranking time, popularity, and number of publisher fans.
3. The privacy information sensitivity grading method according to claim 2, characterized in that: Preprocessing the hot search data includes: Performing data cleaning on the hot search data to remove redundant data, erroneous data and irrelevant data in the hot search data; Standardizing the cleaned hot search data text; The stop words of the hot search data are removed after text normalization.
4. The privacy information sensitivity grading method according to claim 1, characterized in that: The jieba tool in python is used to split the words into independent words, and the SVM model is used to identify the privacy words.
5. The privacy information sensitivity grading method according to claim 4, characterized in that: The privacy information parameters include calculation frequency parameters and information influence parameters; The calculation formula of the frequency parameter is: Where: N is the sliding time window; i is the information; t k is the time step; ω k is the weight; f i (t k ) is the sliding time window N and the information i at each time step t in the sliding window k The frequency of max (t k ) for each time step t k The maximum frequency of all private information in the The calculation formula of the information influence parameter S is: S=ω H *H′ unit,i +oh I *I i ; Where: H is the heat value influence parameter; ω I is the publisher’s influence parameter; H′ unit,i is the normalized heat value; I i Influence for publishers.
6. The privacy information sensitivity grading method according to claim 5, characterized in that: The formula for the normalized heat value is: Where: T i Time for being on the trending search list; is the time decay function; H i The popularity of the entry where the private information is located.
7. The privacy information sensitivity grading method according to claim 6, characterized in that: Classifying each piece of privacy information based on the frequency parameter F and the information influence parameter S includes: Set a plurality of first thresholds and second thresholds for the frequency parameter F and the information influence parameter S respectively; The levels are divided according to the first threshold value and the second threshold value intervals where the frequency parameter F and the information influence parameter S are located.
8. A privacy information sensitivity grading system, characterized in that: include: Acquisition module, used to obtain hot search data; A preprocessing module, used for preprocessing the hot search data; A recognition module, used to split the continuous Chinese text in the pre-processed hot search data into independent words and identify the privacy words; A calculation module, used for calculating the privacy information parameter based on the privacy words; A classification module is used to classify the private information into levels based on the private information parameters.
9. An electronic device, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: It stores a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the method described in any one of claims 1 to 7.