Social network comment privacy leakage detection system
Through the social network comment privacy leak detection system with multi-feature fusion, the difficulty of feature extraction in post and comment correlation detection is solved, efficient privacy leak detection is achieved, and detection accuracy and recall rate are improved.
Patent Information
- Application Number
- CN202510315483.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
The existing social media information privacy protection methods are mainly aimed at original posts, ignoring the correlation between posts and comments, making it difficult to effectively detect privacy leakage in paired texts. The existing technology cannot accurately identify the risk of privacy leakage, especially when there is semantic association, feature extraction is difficult.
A multi-feature fusion detection model is adopted, and through text preprocessing, feature extraction and feature fusion modules, combined with shallow and deep features, a deep neural network is built to conduct privacy detection of social network comment text pairs, including word frequency statistics, similarity calculation, emotion evaluation and entity detection, and online detection is used using SMOTE oversampling and deep neural network training models.
The model's ability to understand paired text association relationships has been significantly improved, and the accuracy rate of 77% and the recall rate of 92% can be achieved, which can better capture the association relationship between text pairs, and solve the problems of difficulty in feature extraction and insufficient correlation analysis in the prior art.
Smart Images

Figure CN120256565A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of information security, and in particular to a social network comment privacy leakage detection system. Background Art
[0002] Most existing methods for protecting information privacy on social media focus on protecting the original posts published by users, while ignoring the comments on the posts. However, by combining the original posts and comments, the privacy information of the poster and the commenter can be inferred. Existing differential privacy leakage detection technologies mainly detect privacy leakage on single texts and cannot effectively handle the correlation problem of paired texts such as posts and comments. Due to the semantic association and contextual dependency between posts and comments, existing technologies have difficulty capturing the complex interactive features between the two, resulting in poor detection results. In addition to the difficulty in feature extraction, it is difficult to extract effective joint features from posts and comments when processing paired texts, especially when there is an implicit semantic association between the two. Existing technologies cannot accurately identify privacy leakage risks. Summary of the invention
[0003] In view of the above-mentioned deficiencies in the prior art, the present invention proposes a social network comment privacy leakage detection system, which extracts features of comment text pairs through a multi-feature fusion detection model, solves the problems of difficulty in associating paired texts and difficulty in feature extraction, and comprehensively combines shallow features and deep features to divide social network comment text pairs into two categories: those involving privacy leakage and those not involving privacy information leakage, thereby achieving accurate detection of user privacy information leakage.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a social network comment privacy leakage detection system, comprising: a text preprocessing module, a network model training module, a text pair privacy classification detection module, a feature extraction module and a feature fusion module, wherein: the text pair preprocessing module preprocesses the text pair to be detected and generates cleaning replacement information and then outputs it to the feature extraction module; the feature extraction module performs word frequency statistics, similarity calculation, sentiment evaluation, entity detection and semantic calculation according to the cleaning replacement information and obtains corresponding five-dimensional feature information and then outputs it to the feature fusion module; the feature fusion module outputs training data information to the network model training module after vector splicing processing, which is used for training a deep neural network; the text pair privacy classification detection module receives the trained network parameters output by the network model training module and the text to be detected input in real time, performs online detection on the fusion features output by the text preprocessing module, the feature extraction module and the feature fusion module and obtains the detection result.
[0006] The present invention relates to a privacy leakage detection method for social network review text pairs with multi-feature fusion based on the above system, including:
[0007] Step 1: After preprocessing the paired text data with existing labels through text cleaning, for the post review text, word frequency statistical features, similarity features, sentiment score features, entity detection features, and deep semantic features are extracted from the overall and individual perspectives through parallel or serial modes; then, through feature fusion processing, the dimension ∑ of the final feature F is obtained. concat F i , Feature categories and metadata are obtained.
[0008] Step 2: Through manual marking and SMOTE oversampling processing of the associated text pairs in the estimated corpus S+T, it is then used to train and construct a deep neural network privacy text pair classification and detection model.
[0009] Step 3: After performing the same processing as in Step 1 on the text pairs composed of the original text of unfamiliar posts and their comments, it is determined whether the text pairs contain privacy-sensitive content through the deep neural network privacy text pair classification and detection model trained in Step 2.
[0010] The above-mentioned text cleaning preprocessing refers to: reducing noise and standardizing the structure of paired unstructured text, including abnormal missing value processing, removing indicator tags, replacing emojis, replacing abbreviated slang, replacing elongated characters, correcting spelling mistakes with Textblob2, removing stop words, and part-of-speech reduction, etc.
[0011] The above-mentioned replacement of emojis means replacing emojis with corresponding related meanings, and this algorithm is actually a replacement of text encoding.
[0012] The above-mentioned text understanding feature extraction refers to: extracting dimensions that can be used for understanding social network review text pairs, including: basic digital statistical word frequency features, text sentiment features, named entity features, and deep semantic features. Among them, the statistical feature uses term frequency-inverse document frequency (TF-IDF), the sentiment feature counts positive, negative, and neutral indices, the named entity counts the entity index contained in the text pair, and the semantic feature uses the paraphrase-mpnet-base-v2 model under the BERT pre-training framework.
[0013] The above-mentioned vector fusion refers to: by comprehensively using shallow features such as word frequency, similarity, entity, and deep features such as sentiment and semantic vectors, and finally outputting a prediction label obtained through a probability space mapping function to determine whether the text pair is related to privacy sensitivity. The five different types of features are concatenated into a fusion feature, and the concatenation is performed by adding the number of channels.
[0014] The neural network mentioned above is a neural network with a three-layer architecture, including: an input layer, two hidden layers, and an output layer, where: the two hidden layers are fully connected layers, using ReLU as the activation function and binary cross entropy loss as the loss function.
[0015] The classifier mentioned above is trained using the backpropagation algorithm, with the default learning rate of 0.001 of the Adam optimizer under the Pytorch framework. The number of training iterations (epoch) is set to 100, and the step size (batch) is set to 64. When using the trained model to predict the unknown text T, the obtained result is the classification result. Technical effects
[0016] Through multi-dimensional feature fusion, the present invention significantly improves the model's ability to understand the correlation relationship between paired texts, solves the problems of difficult feature extraction and insufficient correlation analysis in the prior art. It can better adapt to the particularity of social network texts. On the dataset of social network post review text pairs, an accuracy rate of 77% and a recall rate of 92% are achieved. Compared with the prior art, it can better capture the correlation relationship between text pairs and solve the detection of associated privacy leakage of paired texts. Description of the drawings
[0017] Figure 1 It is a schematic diagram of the principle of the present invention;
[0018] Figure 2 It is a schematic diagram of the system of the present invention;
[0019] Figure 3 It is a flowchart of the embodiment;
[0020] Figure 4 It is a schematic diagram of corpus processing in the embodiment;
[0021] Figure 5 It is a schematic diagram of the effect of the embodiment. Detailed implementation manners
[0022] Such as Figure 2As shown in the figure, this embodiment relates to a privacy leakage detection system for social network review text pairs based on multi-feature fusion, including: a text pair preprocessing module, a feature extraction module, a feature fusion module, a network model training module, and a text pair privacy classification detection module. Among them: the text pair preprocessing module preprocesses the text pairs to be detected, generates cleaning and replacement information, and then outputs it to the feature extraction module; the feature extraction module respectively performs word frequency statistics, similarity calculation, sentiment evaluation, entity detection, and semantic calculation based on the cleaning and replacement information, and obtains corresponding five types of dimensional feature information, and then outputs it to the feature fusion module; the feature fusion module outputs training data information to the network model training module after vector splicing processing for training a deep neural network; the text pair privacy classification detection module receives the trained network parameters output by the network model training module and the real-time input text to be measured, and performs online detection on the fusion features output after passing through the text preprocessing module, the feature extraction module, and the feature fusion module, and obtains the detection result.
[0023] The described text preprocessing module includes: a data reading unit and a text pair cleaning unit. Among them: the data reading unit reads paired texts from the text pair dataset, and the text pair cleaning unit is connected to the data reading unit and processes the obtained texts through three cleaning methods: removal, replacement, and correction.
[0024] The described feature extraction module includes: a statistical feature extraction unit, a similarity calculation unit, a sentiment evaluation unit, an entity detection unit, and a deep semantic extraction unit. Among them: the statistical feature extraction unit calculates the term frequency-inverse document feature by splicing the two sets of original comments in the text pair to form a longer text; the similarity calculation unit gives the sentence-level semantic correlation score feature for the two sets of data in the text pair through a siamese network; the sentiment evaluation unit gives the positive, negative, and neutral sentiment confidence for the two paragraphs of the text pair respectively and connects them as sentiment features; the entity detection unit calculates the probability distribution of the preset entities included in the text pair and accumulates the features; the deep semantic extraction unit receives the spliced text and gives high-dimensional semantic features after data processing; the above five units operate in parallel and connect the calculation results as the data output of this module to the feature fusion module.
[0025] The described feature fusion module includes: a data processing unit and a data marking unit. Among them: the data processing unit is connected to the feature extraction module, receives the text pair feature information, and splices the data to form paired text features; the data marking unit is connected to the data processing unit, receives the fusion features, and gives two types of labels: containing privacy sensitivity and not containing privacy sensitivity. The data output by the data marking unit is connected to the deep neural network training module and provides training data for it.
[0026] The described network model training module includes: a data partitioning unit and a training unit. Specifically: The data partitioning unit receives all the dataset data and is connected to the training unit, cutting the training set and test set data. The training unit trains according to the existing labels given by the data marking unit, outputs data, and saves the best model as the output data of this module, which is connected to the text pair privacy classification detection module and provides a classification detection model for it.
[0027] The described text pair privacy classification detection module includes: a data reading unit and a text pair privacy detection unit. Specifically: The data reading unit receives the unknown text pairs to be detected and is connected to the text preprocessing module. The text pair privacy detection unit is connected to the data processing unit of the feature fusion module, receives the corresponding text feature information, and inputs the data into the trained network model for text classification detection. The classification result is used as the output data of this module to obtain the result of determining whether it contains privacy-sensitive content.
[0028] As Figure 3 and Figure 1 shown, the privacy leakage detection method for social network review text pairs based on multi-feature fusion in this embodiment includes:
[0029] Step 1) Perform text cleaning and preprocessing on the paired text data with existing labels.
[0030] As Figure 4 shown, it is a word cloud diagram of some key information of the used corpus dataset. As shown in Table 1, it is an example of the step-by-step processing of text cleaning and preprocessing.
[0031] Table 1 Technologies and examples of step-by-step processing and cleaning:
[0032] Step 2) For the post review text, extract word frequency statistical features, similarity features, sentiment score features, entity detection features, and deep semantic features from the overall and individual perspectives respectively through parallel or serial modes.
[0033] The described word frequency statistical features are obtained through the following method: Vector extraction is performed on the paired text data preprocessed in Step 1. A pair of post review texts in the training set S are concatenated before and after the string to form an associated long text. The term frequency-inverse document frequency is used to calculate the features, and the word frequency feature tfidf of the text is calculated w,P,C = tf w,P,C ·idf w , where: P is the post, C is the comment, w is the word, tf w,P,C is the word frequency of the word w in the post and comment, and the inverse document frequency of the word w N is the total number of text pairs, df wThe number of text pairs in which the word w appears.
[0034] Set the terms with abnormal deletion frequency, and ignore the terms that appear in more than 70% of the documents and less than 5 documents. Use the calculated words to build a dictionary and a container model, and map the instance text by loading this model. Set it to 1000 dimensions.
[0035] The terms with abnormal deletion frequency mentioned above refer to avoiding allowing all terms (such as overly frequent terms or stop words) to appear, which may reduce the quality and improve the machine processing performance.
[0036] The similarity feature mentioned above is obtained as follows: Take the preprocessed original input post text in step 1 as the first text and its corresponding comment text as the second text. Use the modified Siamese network structure to format the two input texts in dictionary form, and after embedding them into numerical vectors respectively, input them into both sides of the network. Through the last output layer, the fixed lengths of the post text and the comment text are obtained. Map the Manhattan distance to the interval [0, 1] through the sigmoid function, and output the similarity score as 1 dimension. Specifically: Where: is the fixed length of post P, is the fixed length of comment C. The closer the similarity score is to 1, the stronger the correlation.
[0037] The modified Siamese network structure mentioned above refers to a network structure composed of a Siamese twin network extracting paired features and an LSTM network processing sequential data.
[0038] For the LSTM network mentioned above, its embedding parameter is set to 300, the maximum segmentation length is 128, the LSTM activation function uses ReLu, the loss function uses MSE, and the optimizer uses Adam.
[0039] The sentiment score feature mentioned above is obtained as follows: Read the post text one and the comment text two preprocessed in step 1 for sentiment analysis. Use the sentiment analysis tool VADER NLTK based on dictionary and rules to predict and score the sentiment category of the text sentence and map the keywords to a preset dictionary to give numerical scores or weights.
[0040] The sentiment categories mentioned above include three dimensions: positive, negative, and neutral sentiment, and each sentiment is output with a confidence score.
[0041] The confidence score mentioned above refers to: VADER assigns a predefined sentiment score dictionary between -4 (most extremely negative) and 4 (most extremely positive) to the sentiment of each word in a sentence, and adds up each feature score and normalizes it to a confidence score between -1 and 1. The output text pair has 6-dimensional sentiment feature scores.
[0042] The entity detection feature mentioned above is obtained in the following way: Read the preprocessed post text 1 and review text 2 in step 1 for entity detection analysis, use the BERT model bert-base-ner fine-tuned on the standard CoNLL-2003 dataset, and distinguish between upper and lower cases. Calculate the confidence scores of each entity included in this text pair respectively. Before the output layer, to obtain the cumulative entity features, summarize the confidence scores of the identified entities belonging to the same category in the text, and sum the corresponding probability values of the same-category entities. Specifically: Where: j ∈ S, j is the entity type, S is the set of entity types, p(i, j) is the probability score that the i-th entity is recognized as type j, w(i) is the weight of the i-th entity, and f(i, j) is the feature function of the i-th entity belonging to type j, such as whether it is the start position, internal position, etc. It is the sum of the products of the probability scores, weights, and feature functions of all entities recognized as type j; It is the sum of the products of the weights and feature functions of all entities.
[0043] The entity mentioned above refers to: something with clear features that can be distinguished from other things.
[0044] As shown in Table 2, it is a list of entities related to sensitive content using the CoNLL tagging method. The output text pair has 16-dimensional entity detection scores.
[0045] Table 2 Entity types of the CoNLL tagging method: Abbreviation Label Description B-PER Indicates the start of a person's name following another person's name I-PER Person's name B-ORG Indicates the start of an organization's name following another organization's name I-ORG Organization's name B-LOC Indicates the start of a place name following another place name I-LOC Place name B-MISC Indicates the start of a miscellaneous entity following another miscellaneous entity I-MISC Miscellaneous entity
[0046] The deep semantic feature mentioned above is obtained in the following way: Read the <post, review> text pair preprocessed in step 1, splice it as an associated long text for semantic analysis, and use the [SEP] splicing symbol for the model to distinguish between the post and the review. Use the paraphrase-mpnet-base-v2 model similar to BERT, with a maximum cut length of 512, distinguish between upper and lower cases, and perform excellently in the pre-trained model. Finally, map the sentence to a 768-dimensional dense vector space and semantic feature: F(P, C) = BERT([CLS]+P+[SEP]+C+[SEP])[CLS], where: BERT(x) is the output of encoding the input sequence x by BERT, and [CLS] and [SEP] are the start token and the separator token respectively.
[0047] The semantic information mentioned above refers to the specific scenario meaning expressed by the text, rather than the meaning of the text characters.
[0048] Step 3) According to the extracted text pair features, perform feature fusion processing on the word frequency statistical features, similarity features, sentiment score features, entity detection features, and deep semantic features obtained in step 2. To avoid information loss, directly connect the above features in a concatenation (concat) manner to obtain the dimension ∑ of the final feature F. concat F i to obtain the feature categories and metadata, as shown in Table 3.
[0049] Table 3
[0050] Step 4) Manually label the associated text pairs of the estimated corpus S+T.
[0051] The construction of the estimated corpus S+T is carried out by using web crawler technology and social media public interfaces. During the collection, triple filtering is set: language filtering to screen pure English texts; removing texts containing URLs; excluding forwarded content.
[0052] The manual labeling mentioned above means that text pairs containing privacy-sensitive content (non-public) are labeled as 1, and text pairs not containing privacy-sensitive content (public) are labeled as 0. The text pair labels correspond to the final training dataset S and test dataset T of the deep neural network for the feature data.
[0053] The manual labeling mentioned above has the standards shown in Table 4, including: i) the original tweet does not contain private information (p-pub), ii) the original tweet contains private information (p-pri), iii) the comment processed as an independent tweet does not contain private information (c-pub), and iv) the comment contains private information (c-pri). Use i-pub and i-pri for cases where no new private information can be inferred and where new private information can be inferred respectively by combining the original text and the comment, as shown in the examples in Table 5.
[0054] Table 4 Manual Data Evaluation Guidelines and Examples:
[0055] Table 5 Examples of Privacy Leakage in Comments:
[0056] Step 5) Perform SMOTE oversampling on the corpus dataset obtained in Step 4.
[0057] The SMOTE oversampling process refers to: randomly selecting a sample from the minority-class positive samples as the starting point, then randomly selecting one from the 5 nearest neighbor samples of this sample as the reference sample, and generating a new synthetic sample between the starting sample and the reference sample through linear interpolation. Repeat according to the sampling ratio until the number of positive and negative samples reaches 1:1.
[0058] Step 6) Build a deep neural network privacy text pair classification and detection model and train it using the samples processed by oversampling in Step 5?
[0059] The deep neural network privacy text pair classification and detection model includes: an input layer, two hidden layers, and an output layer, where: the two hidden layers are fully connected layers, the activation function of the hidden layer uses the Relu function, the output layer is a fully connected layer that outputs categories to generate a final output vector with two output nodes, and cross-entropy is used as the loss function.
[0060] Step 7) Use the data in the training set S obtained in Step 4 to train the above classifier. The input layer is responsible for receiving the original data, that is, the word vectors of the text pair. After fusing the features, the feature matrix is input into the network. Train using the backpropagation algorithm, with the default learning rate of 0.001 of the Adam optimizer under the Pytorch framework, the number of training iterations epoch set to 100, and the step size batch_size set to 64.
[0061] Step 8) In the online phase, preprocess the text pair composed of the original text of the unfamiliar post and its comments, obtain the feature data of the text pair through feature extraction and fusion, and determine whether the text pair contains privacy-sensitive content through the privacy classification and detection model. Among them, when it is determined that the text pair contains privacy-sensitive content, the model output result is 1, and when it is determined that it does not contain privacy-sensitive content, the model output result is 0.
[0062] After specific experiments, in an environment where the operating system is compatible with 64-bit Ubuntu and Windows 10, the GPU model is Tesla M40 with 16G of video memory, 32-core Inter Xeon CPU E5-2640, and python3.8, transformer4.16, NLTK, Numpy1.21, Torch1.4, by comparing the built DNN with existing machine learning models including LR (Logistic Regression), DT (Decision Tree), RF (Random Forest), SVM (Support Vector Machine), XGBoost, as shown in Table 6 and Table 7, they are respectively the performance of DNN when using a single feature as input, and the feature input with the highest F1 score on each model. By using various feature combinations as input on the dataset, the performance of all models was evaluated in 5-fold cross-validation. According to the combination principle, there are a total of 26 combinations (10 two-feature combinations, 10 three-feature combinations, 5 four-feature combinations, and 1 all-feature combination). As shown in Table 8, it is the comprehensive evaluation of DNN with different feature combinations on the dataset.
[0063] Table 6 Performance of DNN when a single feature is used as input: Accuracy Precision Recall F1 Score Word Frequency Statistical Feature 0.77 0.76 0.79 0.77 Entity Detection Feature 0.66 0.70 0.54 0.61 Sentiment Score Feature 0.63 0.65 0.57 0.61 Similarity Feature 0.56 0.57 0.50 0.53 Semantic Feature 0.85 0.84 0.87 0.85
[0064] Table 7 Feature input with the highest F1 score on each model: Optimal Feature Accuracy Precision Recall F1 Score LR Semantic Feature 0.80 0.76 0.87 0.81 DT Semantic Feature 0.71 0.70 0.74 0.72 RF Semantic Feature 0.85 0.88 0.83 0.85 SVM Semantic Feature 0.85 0.85 0.84 0.85 XGBoost Semantic Feature 0.84 0.83 0.87 0.85 DNN Semantic Feature 0.85 0.84 0.87 0.85
[0065] Table 8 Comprehensive evaluation of DNN with different feature combinations on the dataset:
[0066] As Figure 5 shown, it is the training and testing situation of the deep neural network DNN. The accuracy of the detection and classification model on the training data increased from 64.7% to 85% after 5 rounds of growth, and rose rapidly in 20 rounds and finally reached an accuracy of 96%. At the same time, the loss dropped rapidly from 0.61 to about 0.3 and continued to decline until it approached 0, and the 0-value rounds were fully trained. The training time of DNN is 3 times that of XGBoost, about 21s, the peak memory is 2 / 3 of XGBoost, but it is 2 times that of XGBoost in terms of memory allocation, about 0.17MB. On the test set, the accuracy of the model in multiple rounds exceeded 77%, and was not lower than 70%. The recall rate could not reach 92%. To verify the generalization ability of the model, the deep neural network privacy-sensitive text detection and classifier trained with all the data in the dataset was used, and the text of the real forum-like social network platform under a specific topic was used to further evaluate the model, and its accuracy could reach 70%.
[0067] Compared with the prior art, the present invention addresses the problem of comment privacy leakage in the scenario where both the original post and comments in a social network are texts. Through a multi-dimensional feature fusion technology for social network post-comment text pairs, the present invention extends the feature dimensions of the post and comment text pairs to five categories, including pair similarity, sentiment tendency, entity recognition, semantic association, and user behavior patterns. This multi-dimensional feature fusion method not only enriches the understanding of text pairs but also can accurately identify potential privacy information in text pairs, achieving an accuracy rate of 77% and a recall rate of 92%, effectively solving the deficiencies of existing text classification methods in dealing with related texts. It is easy to integrate into the current social network media platform; it can remind users that their comments may involve privacy-sensitive content related to privacy leakage associated with the original post; this technology can be seamlessly integrated into existing social network platforms without large-scale modifications to the platform architecture. At the same time, by continuously expanding the training data set, the accuracy of the model can be further improved, demonstrating good adaptability and expansion capabilities.
[0068] Those skilled in the art can make partial adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation, and all implementation solutions within its scope are subject to the present invention.
Claims
1. A social network comment privacy leakage detection system, characterized in that, It includes: A text preprocessing module, a network model training module, a text pair privacy classification and detection module, a feature extraction module, and a feature fusion module. Among them: The text pair preprocessing module preprocesses the text pair to be detected, generates cleaning and replacement information, and then outputs it to the feature extraction module; The feature extraction module respectively performs word frequency statistics, similarity calculation, sentiment evaluation, entity detection, and semantic calculation according to the cleaning and replacement information, obtains five types of dimensional feature information, and then outputs it to the feature fusion module; The feature fusion module outputs training data information to the network model training module after vector splicing processing for training a deep neural network; The text pair privacy classification and detection module receives the trained network parameters output by the network model training module and the text to be measured input in real time, and performs online detection on the fusion features output after passing through the text preprocessing module, the feature extraction module, and the feature fusion module to obtain the detection result.
2. The social network comment privacy leakage detection system according to claim 1, characterized in that, The text preprocessing module described above includes: A data reading unit and a text pair cleaning unit. Among them: The data reading unit reads paired texts from the text pair dataset, and the text pair cleaning unit is connected to the data reading unit and processes the obtained texts through three cleaning methods: removal, replacement, and correction.
3. The social network comment privacy leakage detection system according to claim 1, characterized in that, The feature extraction module described above includes: A statistical feature extraction unit, a similarity calculation unit, a sentiment evaluation unit, an entity detection unit, and a deep semantic extraction unit. Among them: The statistical feature extraction unit calculates the term frequency-inverse document feature by splicing the two groups of original text comments in the text pair to form a longer text; The similarity calculation unit gives the sentence-level semantic correlation score feature for the two groups of data in the text pair through a siamese network; The sentiment evaluation unit gives the positive, negative, and neutral sentiment confidence levels for the two texts in the text pair respectively and connects them as sentiment features; The entity detection unit calculates the probability distribution of the preset entities included in the text pair and accumulates the features; The deep semantic extraction unit receives the spliced text and gives high-dimensional semantic features after data processing; The above five units operate in parallel and connect the calculation results as the data output of this module to the feature fusion module.
4. The social network comment privacy leakage detection system according to claim 1, characterized in that, The feature fusion module described above includes: A data processing unit and a data marking unit. Among them: The data processing unit is connected to the feature extraction module, receives the text pair feature information, and splices the data to form paired text features. The data marking unit is connected to the data processing unit, receives the fusion features, and gives two types of labels: containing privacy sensitivity and not containing privacy sensitivity. The data output by the data marking unit is connected to the deep neural network training module and provides training data for it.
5. The social network comment privacy leakage detection system according to claim 1, characterized in that, The network model training module described above includes: A data partitioning unit and a training unit. Among them: The data partitioning unit receives all the dataset data and is connected to the training unit to cut the training set and test set data. The training unit trains according to the existing labels given by the data marking unit, outputs the data, and saves the best model as the output data of this module, which is connected to the text pair privacy classification and detection module and provides a classification and detection model for it.
6. The social network comment privacy leakage detection system according to claim 1, characterized in that, The described text pair privacy classification and detection module includes: a data reading unit and a text pair privacy detection unit, where: the data reading unit receives the unknown text pair to be detected and is connected to the text preprocessing module, and the text pair privacy detection unit is connected to the data processing unit of the feature fusion module, receives the corresponding text feature information, and inputs the data into the trained network model for text classification and detection. The classification result is used as the output data of this module to obtain the result of determining whether the content contains privacy-sensitive information.
7. A method for detecting privacy leakage in social network review texts with multi-feature fusion based on the system described in any one of claims 1-6, characterized in that including: Step 1: After preprocessing the paired text data with existing labels through text cleaning, for the post comment text, word frequency statistical features, similarity features, sentiment score features, entity detection features, and deep semantic features are extracted from the overall and individual perspectives in parallel or serial mode; then The dimension ∑ of the final feature F obtained through feature fusion processing concat F i , and the feature categories and metadata are obtained; Step 2: Through artificial marking and SMOTE oversampling processing of the associated text pairs in the estimated corpus S+T collected using web crawler technology and social media public interfaces, and then used to train and construct the deep neural network privacy text pair classification and detection model; Step 3: After performing the same processing as in Step 1 on the text pairs composed of the original text of the unfamiliar post and its comments, determine whether the text pairs contain privacy-sensitive content through the deep neural network privacy text pair classification and detection model trained in Step 2.
8. The method for detecting privacy leakage of social network review texts with multi-feature fusion according to claim 7, characterized in that, The described text cleaning preprocessing refers to: reducing noise and standardizing the structure of paired unstructured text, including abnormal missing value processing, removing indicator tags, replacing emojis, replacing abbreviated slang, replacing elongated characters, correcting spelling mistakes with Textblob2, removing stop words, and lemmatization.
9. The method for detecting privacy leakage of social network review texts with multi-feature fusion according to claim 7, characterized in that, The described text understanding feature extraction refers to: by extracting the dimensions for understanding text pairs in social network comments, including: basic digital statistical word frequency features, text sentiment features, named entity features, and deep semantic features, where the statistical features use term frequency-inverse document frequency (TF-IDF), the sentiment features count positive, negative, and neutral indices, the named entity statistics the entity index contained in the text pair, and the semantic features use the paraphrase-mpnet-base-v2 model under the BERT pre-training framework.
10. The privacy leakage detection method for social network review texts with multi-feature fusion according to claim 7, characterized in that, The described vector fusion refers to: by comprehensively using shallow features such as word frequency, similarity, entity, and deep features such as sentiment and semantic vectors, finally outputting the predicted label obtained through the probability space mapping function to judge whether the text pair is related to privacy sensitivity. The five different types of features are concatenated into a fusion feature, and the concatenation is performed using the method of adding the number of channels.