False comment perception method based on semantic-emotion double-flow attention fusion
By employing a semantic-sentiment dual-stream parallel processing architecture and sentiment masking modeling, the challenge of sentiment logic modeling in Chinese fake comment identification is solved, achieving high-precision and robust fake comment detection.
Patent Information
- Application Number
- CN202511118959.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to accurately identify fake Chinese comments, especially when dealing with complex Chinese comments. They are unable to effectively model the internal emotional logic and coherence of the comments, resulting in poor recognition performance.
We adopt a semantic-emotion dual-stream parallel processing architecture, extract deep semantic features through a large-scale pre-trained language model, introduce an emotion masking modeling task, and train the recognition model by combining a hierarchical attention mechanism and a joint loss function.
It significantly improves the accuracy and robustness of fake review detection, can identify illogical emotional content in complex Chinese reviews, and does not rely on external information, possessing efficient real-time detection capabilities.
Smart Images

Figure CN120892568A_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a method for detecting fake comments based on semantic-emotion dual-stream attention fusion, involving natural language processing and text classification technology, and particularly a method and system for detecting fake comments online. Background Technology
[0002] In today's digital economy, online reviews have become a crucial bridge connecting consumers and businesses, and a core source of information influencing public opinion and consumer decisions. However, the value of this ecosystem is being severely eroded by the increasing prevalence of fake reviews. Driven by commercial competition or malicious attacks, a large amount of false and misleading content floods the internet, not only harming consumers' right to know but also undermining a fair and trustworthy market environment. Therefore, developing technologies capable of automatically and accurately identifying and detecting fake reviews is of great practical significance for maintaining a healthy online ecosystem.
[0003] Traditional fake review detection techniques face numerous challenges when confronted with ever-evolving deceptive texts, particularly those tailored to the characteristics of the Chinese language. On one hand, the shallow text features relied upon by early methods are no longer sufficient to handle complex linguistic camouflage. On the other hand, while existing models have made progress in general semantic understanding, they lack the specialized ability to model the unique "emotional logic" cues within review texts that hint at their falsity. This is especially true when dealing with the complex and subtly expressive language of the Chinese language, where recognition performance falls short of expectations.
[0004] Currently, mainstream sensing technologies can be broadly categorized as follows:
[0005] (1) Methods based on artificial feature engineering and traditional machine learning
[0006] These methods involve manually designing and extracting a series of quantifiable text features, which are then fed into traditional classifiers such as Support Vector Machines (SVM), Logistic Regression, or Random Forests for training and recognition. These features typically include: text metadata features (such as comment length, proportion of uppercase letters, and rating stars), lexical statistical features (such as the frequency of specific sentiment words and pronouns), and syntactic structural features (such as part-of-speech distribution).
[0007] However, this method has several inherent limitations. First, the feature engineering process heavily relies on the experience of domain experts, consuming significant manpower, and the designed feature set has limited generalization ability. Second, this method treats comments as an unordered "bag of features," completely ignoring word order and contextual relationships between sentences, thus failing to capture the deeper semantic connotations of the text.
[0008] The limitations are particularly evident when facing well-crafted Chinese fake reviews. The authors of deceptive texts can easily circumvent these shallow rules, for example, by adjusting the length of the review or interspersing sentiment words to bypass detection. More importantly, for complex reviews in Chinese that imply true intentions through sentiment shifts or contrasts between sentences (e.g., "The design is perfect, the material and texture are also great, but it broke after a day of use"), such methods are almost unable to understand the underlying logic and true inclination, resulting in a high rate of false positives.
[0009] (2) Methods based on static word embeddings and early deep learning
[0010] With the development of deep learning, researchers began to use models such as Convolutional Neural Networks (CNN) or Long Short-Term Memory Networks (LSTM) to automatically learn text representations. These methods usually map words in text to pre-trained static word vectors (such as Word2Vec, GloVe), and then use neural networks to capture local syntax (CNN) or long-distance dependencies (LSTM) to learn sentence or document representations for final classification.
[0011] Compared to traditional methods, this type of technology has made significant progress in automatic feature extraction and capturing syntax. However, the core "static word embedding" has a major flaw: it cannot solve the polysemy problem, i.e., the same word has the same vector representation in different contexts, which limits the depth of the model's understanding of complex semantics.
[0012] In the fake review detection perception task, this flaw results in the loss of key information. In addition, although LSTM can theoretically capture long dependencies, in practice, for a review containing multiple sentences, multiple sentiment shifts, and logical coherence, standard deep learning models still lack a clear mechanism to model the flow of sentiment and logical coherence. The model may know that the review contains both positive and negative words, but it is difficult to determine whether this sentiment combination is reasonable (such as describing the pros and cons of a product) or illogical and deceptive (such as forcing content to meet the word count or pretending to be a real user).
[0013] (3) Methods based on large-scale pre-trained language models
[0014] Large-scale pre-trained language models (Pre-trained Language Models) represented by BERT have gained strong contextual awareness through self-supervised learning on a vast amount of unlabeled text, greatly promoting the development of natural language processing. In fake review detection, fine-tuning (Fine-tuning) of models such as BERT has become the current mainstream and most advanced baseline method.
[0015] Despite the remarkable achievements, the standard fine-tuning method still has its limitations. The standard fine-tuning process usually only adds a classification head to the final output layer of the model and aims to minimize the classification loss as the only goal. This "end-to-end" optimization method, although direct, also leads to the model becoming a "black box" and the learning process lacks targeted guidance.
[0016] This problem is particularly prominent when dealing with complex and deceptive Chinese reviews. The standard fine-tuned model is not explicitly required to understand and evaluate the "logicality of sentiment development" within the review. For example, a fake review may deliberately create a sentiment shift to appear objective, but the shift may be abrupt and unnatural. The standard BERT model lacks a specialized architecture or optimization goal to capture and utilize this "inconsistency in sentiment logic" when fine-tuning. It may rely solely on certain strong semantic words to make judgments, ignoring deeper clues that reveal its deceptive nature.
[0017] Therefore, how to design a method that can fully utilize the powerful semantic capabilities of pre-trained language models while introducing a new mechanism to explicitly and supervisedly learn and understand the coherence of sentiment logic within the text is a key technical problem that needs to be solved in the field. SUMMARY
[0018] To solve the problem of difficulty in accurately identifying Chinese fake reviews in the prior art, especially the inability to effectively model the coherence of sentiment logic within the review, the present application proposes a fake review perception method that combines deep semantics and sentiment dynamics. This method provides an effective solution for high-precision, strong-robustness automated identification of Chinese fake reviews, and can effectively address the challenges faced by current network information content security.
[0019] The present application proposes a fake review perception method based on a semantic-sentiment dual-flow parallel processing architecture. Through the method of the present application, it can accurately identify fake reviews that are disguised by imitating the tone of real users and deliberately constructing sentiment shifts in complex network public opinion environments, solving the problem of traditional models being easily misled by relying on shallow features or a single semantic dimension.
[0020] In the model construction and training stage, based on the linguistic characteristics of Chinese reviews, the present application first divides a single review text into an ordered sequence of clauses that can reflect the basic units of sentiment. Then, the present application innovatively uses a double-flow parallel architecture to process text information: a semantic feature flow, which processes the complete review text through a large-scale pre-trained language model (such as BERT) to capture global and deep contextual semantic information; a sentiment dynamic feature flow, which generates a sequence of sentiment vectors depicting the sentiment flow trajectory of the review by performing sentiment analysis on each clause. It is particularly crucial that the present application innovatively introduces a masked sentiment modeling (MSM) auxiliary task, which forces the model to learn and understand the internal logic and coherence of sentiment expression by having the model predict the sentiment of a masked clause based on the sentiment of adjacent clauses. Finally, a deep learning main model containing a hierarchical attention mechanism is used to deeply integrate and train the two information flows under the supervision of a joint loss function composed of the main classification task loss and the auxiliary task loss, thereby constructing a recognition model with strong semantic understanding ability and sentiment logic discrimination ability.
[0021] In the recognition and application stage, the present application first performs the same clause processing on a new review text to be detected. Then, through the trained double-flow feature extraction module, the global semantic representation vector and the global sentiment dynamic vector of the review are generated. Finally, the two fused feature vectors are input into the trained deep learning main model for classification and discrimination. The hierarchical attention mechanism is responsible for capturing the complex semantic dependency between clauses in the review, and the sentiment attention mechanism focuses on evaluating the rationality of the overall sentiment flow. The method of the present application can maintain high recognition accuracy and robustness when facing reviews containing complex language phenomena such as sarcasm, metaphor, and sentiment reversal, can perform real-time detection on a large amount of reviews, and does not rely on any external information other than the review text itself.
[0022] To achieve the purpose of the present application, the specific technical steps of the present scheme are as follows: a false review perception method combining deep semantics and sentiment dynamics, comprising the following steps:
[0023] Step (1) constructing and preprocessing a review dataset for model training, which contains a large number of labeled real review and fake review samples;
[0024] Step (2) performing clause processing on the input review text to be detected, dividing a single review text into one or more ordered clauses to form a clause sequence;
[0025] Step (3) uses a double-flow parallel architecture to extract features from the review text and clause sequence, respectively obtaining a deep semantic feature flow and a sentiment dynamic feature flow;
[0026] Step (4) performs a sentiment masking modeling (MSM) auxiliary task based on the sentiment dynamic feature flow obtained in step (3) to train the model's understanding of sentiment context logic;
[0027] Step (5) constructs a deep learning main model containing a hierarchical attention mechanism, a feature fusion network, and a classifier, and uses the dataset of step (1) and a joint loss function to train the main model end-to-end;
[0028] Step (6) saves the false review identification model containing the optimal parameters obtained after sufficient training in step (5);
[0029] Step (7) calls the identification model obtained in step (6) to process and classify any new review text to be detected, and outputs the probability of being a false review.
[0030] Further, in step (1), the specific process of constructing and preprocessing the review dataset is as follows
[0031] (1.1) Collect large-scale public review texts with user labels or platform official labels from mainstream e-commerce platforms, social media websites, and other channels;
[0032] (1.2) Clean the collected data to remove irrelevant characters, HTML tags, and other noise, and based on platform labels, user feedback, or other prior knowledge, preliminarily label the reviews as "real reviews" or "false reviews" to form an original corpus;
[0033] (1.3) Organize artificial cross-validation and fine-tuning of the preliminarily labeled data to ensure the accuracy of the labels, and finally form high-quality training, validation, and test sets for supervised learning.
[0034] Further, in step (2), the method for processing the review text by clauses is as follows:
[0035] (2.1) Analyze the linguistic features of Chinese reviews and find that sentiment turning points within the reviews are often key clues for identifying their authenticity. These turning points are usually marked by specific conjunctions or punctuation marks;
[0036] (2.2) Pre-set a segmentation rule base, which contains common transitional conjunctions in Chinese (such as "but", "but", "however", "however", etc.) and explicit separation punctuation marks (such as periods, question marks, exclamation marks, semicolons, etc.). According to this rule base, the input single review text T is divided into an ordered clause sequence C = {C1, C2,..., C n}.
[0037] The existing technology usually directly inputs the preprocessed text into the model (such as BERT, CNN, GNN, etc.), and lets the model itself learn the relationship within the text. Although this method is feasible, it lacks targeted guidance and needs to spend more effort to learn the logical relationship between sentences. The text segmentation is essentially a supervised and prior knowledge-based feature engineering. This preprocessing for the purpose of "preserving emotional logic" is specifically for the subsequent double-flow model (especially the emotional dynamic flow), which has strong task pertinence. Instead of directly feeding the text to the model, the most important structured information is extracted using linguistic knowledge. This greatly reduces the learning difficulty of the subsequent model, allowing it to focus more on learning the emotional logic between clauses rather than searching for a needle in a haystack from a pile of scattered tokens. This makes the overall model architecture more efficient, accurate, and more innovative.
[0038] Further, in step (3), the feature extraction process of the double-flow parallel architecture is as follows:
[0039] (3.1) Deep semantic feature flow: input the unsegmented, complete original review text T into a large-scale pre-trained language model (such as BERT). The model generates a deep semantic token vector h i for each word in the text through its internal Transformer structure and context self-attention mechanism.
[0040] (3.2) Emotional dynamic feature flow: input each clause C j obtained in step (2) into a pre-trained high-quality sentiment analysis model (which can be an API or another language model), to generate its corresponding sentiment vector S j for each clause. The sentiment vectors of all clauses form a sentiment vector sequence {S1, S2,..., S n}, which explicitly depicts the emotional flow trajectory of the review.
[0041] By generating a sentiment vector for each sentence, we finally get an ordered sequence of sentiment vectors. It explicitly and structurally depicts how the user's sentiment develops and turns in the review. For example, a fake review may have an unnatural pattern of "extremely positive -> abrupt turn -> forced positive". This method clearly captures this pattern in the form of a vector sequence, while traditional methods ignore this process. This is the basis for subsequent "MSM" and sentiment attention analysis, and is one of the most core improvements compared to existing technologies.
[0042] Further, the specific process of training the sentiment masking modeling (MSM) auxiliary task in step (4) is as follows:
[0043] (4.1) This task aims to let the model learn and understand the logical coherence of the sentiment context. Traverse the sentiment vector sequence {S1, S2,..., Sn} generated in step (3.2), and for any index i (from 1 to n-2), construct a training sample; n
[0044] (4.2) Concatenate the sentiment vector S i of the i-th sentence and the sentiment vector S i+2 of the i+2-th sentence in the vector dimension to form a context vector: Context_Vector = [S i ; S i+2 ], which is used as the input of the model;
[0045] (4.3) Set the prediction target of the model as the true sentiment vector S i+1 of the i+1-th sentence which is "masked" in the middle. The model uses a multi-layer perceptron (MLP) as the prediction head, and the output dimension is consistent with the dimension of the sentiment vector;
[0046] (4.4) This is a high-dimensional vector regression task, and the loss function L MSM uses cosine similarity loss, aiming to minimize the directional difference between the predicted vector and the true vector. The loss function formula is as follows:
[0047]
[0048] Where, is the sentiment vector predicted by the model, and S i+1 is the true sentiment vector.
[0049] The task that is most similar to MSM in form is the Masked Language Modeling (MLM) task of BERT itself. But the goal of MLM is to predict the specific token that is covered, which is still learning at the level of language symbols in nature. MSM is a higher level of abstraction. Instead of "should this blank be filled with 'happy' or'satisfied', the model learns that "between the former positive emotion and the latter positive emotion, there should be a positive emotion state in the same direction in the vector space". It learns the relationship between emotional states rather than word collocations. This enables the model to break free from the shackles of surface words and grasp the essence of emotions. Therefore, even if the fake review forges a similar emotional turn with different words, the model trained by MSM may identify the anomaly because the logic of its "emotional vector" is not coherent.
[0050] Further, the specific process of constructing and training the deep learning master model in step (5) is as follows (5.1) constructing the master model architecture. The model consists of hierarchical attention units, feature fusion units and classification units:
[0051] Hierarchical attention unit: contains intra-sentence attention and inter-sentence attention. Intra-sentence attention is used to aggregate word vectors within each clause to form a clause-level semantic representation V j Inter-sentence attention draws on the Transformer Encoder structure to capture long-range dependencies by multi-head self-attention mechanism, residual connection and layer normalization, etc. to interact information of all clause representations V j , and finally output the global semantic representation vector X. At the same time, a sentiment attention unit weights and sums the sequence of sentiment vectors {S1, S2,..., S n} to obtain the global sentiment dynamic vector S.
[0052] Feature fusion unit: concatenate the global semantic representation vector X and the global sentiment dynamic vector S obtained in the previous step in the feature dimension to form the fusion vector [X; S], and input it into a fusion network composed of fully connected layers for deeper information interaction.
[0053] Classification unit: a linear output layer with Sigmoid activation function is connected at the end of the fusion network to output the final probability value of the review being a fake review.
[0054] (5.2) Training with joint loss function. To achieve end-to-end optimization and fully utilize the advantages of multi-task learning, a joint loss function L Total is used to train the entire model. The function is composed of three parts weighted together:
[0055] L Total = L main + α·LMLM + β · L MSM
[0056] where L main is the binary cross-entropy loss of the main task (i.e. true or false classification); L MLM is the mask language modeling loss from the BERT semantic stream, as an auxiliary regularization task; L MSM is the loss of the sentiment masking modeling task from step (4). a and β are hyperparameters for balancing the loss weights of each term.
[0057] The hierarchical attention unit in this paper adopts a hierarchical processing method of "local first, then global", which is very consistent with human reading and understanding habits. This design enables the model to better understand those reviews containing multiple sentiment turns or complex clauses, rather than treating them as a pile of unstructured words, with a depth and accuracy of understanding far beyond the "flat" attention mechanism. In addition, most of the existing models use "single task learning" during the fine-tuning stage, which means that only the final classification loss is optimized. This means that all the efforts of the model are only for one goal, which is easy to cause overfitting, and the learned feature representation may not be general and solid enough. While this paper adopts a joint loss function composed of three parts, which is a very powerful multi-task learning framework. By learning three tasks at the same time, the model is guided to learn a more general, essential and robust internal representation. Because it has been tested by multiple tasks, it has stronger generalization ability and is not easily confused by some superficial and deceptive text features, so the performance on the main task is also improved.
[0058] Further, in step (7), the process of using the trained model for recognition is as follows:
[0059] (7.1) Capture or receive a new review text to be detected;
[0060] (7.2) According to the method of step (2) and step (3), perform sentence processing and double-stream feature extraction on the new review;
[0061] (7.3) Input the extracted features into the trained recognition model saved in step (6);
[0062] (7.4) The model forward propagates and calculates through its internal hierarchical attention, fusion and classification units, and finally outputs a probability value between 0 and 1. A threshold (such as 0.5) can be set, if the output value is greater than the threshold, the review is determined to be a fake review, otherwise it is a real review.
[0063] The whole model is trained and optimized by a joint loss function composed of the main classification task loss, the MLM task loss and the MSM task loss.
[0064] A false review perception system fusing deep semantics and sentiment dynamics, the system comprises:
[0065] A data preprocessing module configured to perform the steps described in steps (1) and (2);
[0066] A feature extraction and model training module configured to perform the steps described in steps (3), (4) and (5); and a recognition and classification module configured to perform the steps described in step (7).
[0067] An electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the false review perception method fusing deep semantics and sentiment dynamics when executing the program.
[0068] A computer readable storage medium having computer instructions stored thereon, wherein the computer instructions are executed by a processor to implement the false review perception method fusing deep semantics and sentiment dynamics
[0069] Compared with the prior art, the advantages of the present application are as follows:
[0070] 1、The present application can more deeply understand the internal structure and emotional logic of Chinese reviews through the semantic and emotional double-flow architecture, innovative MSM task and hierarchical attention mechanism, significantly improving the accuracy and reliability of false review detection perception;
[0071] 2、The technical solution is complete and independent, and the whole perception method only depends on the review text itself and does not depend on any external metadata (such as user information, rating stars, publishing time, IP address, etc.).
[0072] In many scenarios, user profiles, historical behavior data are missing, incomplete, or involve privacy and cannot be obtained. The pure text solution has strong universality and can be easily deployed to any platform with only review text, whether it is e-commerce, social media or news app, achieving "plug and play". In addition, methods such as graph neural networks that rely on user behavior are not effective on new users or "small numbers" (data is sparse), and are easily deceived by fake user behavior;
[0073] 3. Improved Model Interpretability: Although deep learning models are essentially "black boxes," the architectural design provides multiple analytical "windows" for understanding the model's decision-making process. This is something lacking in existing BERT fine-tuning methods. The "emotional dynamic feature stream" generates an ordered sequence of emotion vectors {S1, S2, ..., S}. This means that for any comment judged as fake, analysts can visualize this emotion sequence and intuitively see whether the emotion flow is forced or illogical. This provides valuable clues for manual review and model correction. The "hierarchical attention unit" generates attention weights during computation. By analyzing these weights, it's possible to know which clause in the comment the model focuses on when making decisions. This ability to "provide observation windows" allows users (such as platform moderators) to see part of the basis for the model's decisions, rather than simply accepting a "true / false" conclusion, thus greatly increasing trust in the entire system.
[0074] 4. Deep adaptation to the characteristics of the Chinese language: This solution does not simply apply a powerful general model (BERT) directly to Chinese data, but first uses linguistic knowledge to "preprocess" and "structure" the problem, reflecting a deep understanding of the characteristics of the Chinese language. Attached Figure Description
[0075] Figure 1 This is a schematic diagram of the overall architecture of a fake comment perception model based on an embodiment of the present invention.
[0076] Figure 2 This is a schematic diagram of the task flow based on Emotional Masking Modeling (MSM).
[0077] Figure 3 This is a schematic diagram of the semantic flow attention mechanism.
[0078] In the picture:
[0079] Module 1: Comment Sentence Segmentation Module
[0080] Module 2: Semantic Vector and Sentiment Vector Extraction Module
[0081] Module 3: Attention Mechanisms and Fusion Classification Module. Detailed Implementation
[0082] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0083] Example: Refer to Figure 1 This invention provides a method for detecting fake comments based on joint semantic and sentiment modeling.
[0084] Step 1: Comment clause segmentation (corresponding to Figure 1 Module 1 in
[0085] Obtain an online review text T to be detected. Since the emotional turns within Chinese comments are the key clues for identifying their authenticity, this step first preprocesses the text T. Through a clause segmentation module, using common Chinese conjunctions for turning (such as "but", "however", "nevertheless", etc.) and explicit punctuation marks for separation (such as commas, full stops, exclamation marks, etc.) as boundaries, the text T is segmented into an ordered clause sequence C = {C1, C2,..., C n}.
[0086] Step 2: Semantic vector and sentiment vector extraction (corresponding to Figure 1 Module 2 in
[0087] This step processes the data in a two-stream parallel manner to capture the deep information of the comments from different dimensions.
[0088] (2.1) Semantic vector extraction
[0089] Input the original review text T into a BERT-based deep language model. This model performs the standard masked language modeling (MLM) task, that is, randomly masking some tokens before the input, and then training the model to predict the masked tokens based on the context. This process enables the model to learn rich general language rules. After the training is completed, extract the output of the last layer of the model to obtain the context-aware word vectors h of each token in the text i .
[0090] (2.2) Sentiment vector extraction and sentiment masking modeling (MSM)
[0091] This part is one of the core innovations of this invention that differentiates it from the prior art, aiming to enable the model to go beyond simple semantic understanding and further master the internal logic of emotional expression. In parallel, input each clause C obtained in Step 1 j into a high-quality, pre-trained large language model for sentiment analysis respectively, so as to generate a sentiment vector S for each clause j . These vectors form a sentiment vector sequence {S1, S2,..., S n}, which explicitly depicts the emotional flow trajectory of the comment. At the same time, obtain the sentiment vector S0 of the whole comment for subsequent calculations.
[0092] To solve the problem that the existing model cannot evaluate whether the emotional development is logical, the auxiliary task of emotion masking modeling (MSM) is introduced. The design idea is derived from the cognitive process of human understanding of dialogue: we will naturally use the context to infer the missing part of the tone and emotion. For example, in the context of "the service of this restaurant is good, [...], come again next time", we can probably infer that the [...] part is a positive evaluation. The MSM task is to give the model the logical reasoning ability in the emotional level. The specific implementation is as follows:
[0093] Training sample construction: traverse the above generated emotion vector sequence {S1, S2,..., Sn-1} n}. For each index i (from 1 to n-2) in the sequence, a training sample can be constructed.
[0094] Input and target: the emotion vector S i of the i-th sentence and the emotion vector S i+2 of the i+2-th sentence are spliced in the vector dimension to form a longer context vector: Context_Vector=[S i ;S i+2 ]. The vector is used as the input of the model. The training target of the sample is to predict the real emotion vector S i+1 of the i+1-th sentence which is "masked".
[0095] Prediction head design: a multi-layer perceptron composed of one to two fully connected layers and a nonlinear activation function (such as ReLU, GeLU) is designed as the prediction head. It receives Context_Vector as input and outputs a prediction vector i with the same dimension as the emotion vector S This is a vector regression task in a high-dimensional space, and the cosine similarity loss L MSM is preferably used as the loss function to measure the closeness of the predicted vector and the real vector in direction.
[0096]
[0097] Through the optimization of the MSM task, the model is forced to understand "what kind of emotional sequence is natural and logical", so that it can produce a stronger discrimination signal when facing those false reviews with abrupt emotional transitions and incoherent logic.
[0098] Step three: attention mechanism, feature fusion and classification (corresponding to module 3 in Figure 1 )
[0099] (3.1) Hierarchical semantic attention mechanism
[0100] Intra-sentence attention mechanism: This stage aims to aggregate the word information within each clause to form a clause-level semantic representation. For each clause C j , the set of word vectors contained is {h k | k ∈ C j}. First, the average vector of the clause is calculated as the query vector q j :
[0101]
[0102] Then, the relevance score t j of the query vector q k and each word vector h k in the clause is calculated, and normalized to get the attention weight a k :
[0103] t k = Attention(q j , h k )
[0104]
[0105] The vector O j is obtained by weighted summation, and finally concatenated with the query vector and passed through a feed-forward network (FFN) to generate the initial semantic representation V j of the clause:
[0106]
[0107] V j = FFN([O j ; q j ])
[0108] Inter-sentence attention mechanism: After obtaining the sequence of clause representations V = {V1, V2,..., V n}, an inter-sentence attention module inspired by the Transformer Encoder layer is introduced. This module allows the representations V j of each clause to interact with each other through multi-head self-attention, residual connections, layer normalization, and feed-forward neural networks (FFN), capturing long-range dependencies. Finally, the module outputs a review representation vector X that integrates the global semantic structure.
[0109] (3.2) Sentiment attention mechanism
[0110] Similar to the semantic flow, the sentiment flow also uses an attention mechanism to dynamically aggregate sentiment information across the entire review. The sentiment vector Sj The attention weight α is calculated by comparing it with the global sentiment vector S0 representing the entire comment. j Then, the sentiment vectors of all clauses are weighted and summed to obtain the final global sentiment dynamic vector S:
[0111]
[0112] (3.3) Feature Fusion and Classification: After obtaining the deep semantic representation X and the sentiment dynamic representation S, the two are concatenated along the feature dimension to obtain a highly information-rich fusion vector [X; S]. This vector is fed into a final fusion network and an output layer with a sigmoid activation function for final binary classification prediction, outputting the probability value of whether the comment is true or fake.
[0113] Step 4: Joint Training To achieve end-to-end optimization, this invention employs a joint loss function L. Total It consists of three weighted parts:
[0114] L Total =L main +α·L MLM +β·L MSM
[0115] Among them, L main It is the cross-entropy loss of the main task (i.e., true / false binary classification), L MLM It is the masked language modeling loss from the BERT semantic stream, L MSM This is the loss for a new task (emotional masking modeling) derived from the emotional flow. α and β are hyperparameters used to balance the various losses. By minimizing this total loss, the model is driven to collaboratively learn general language rules, emotional context logic, and the final classification task, thus becoming more powerful and robust.
[0116] Those skilled in the art will understand that the above methods can be embedded in a computer program and executed by related hardware (such as a processor). Therefore, the present invention also provides a fake review detection system, the internal module division of which can be referred to... Figure 1 The invention includes a segmentation module, a feature extraction module (containing a semantic processing unit and a sentiment processing unit), a feature fusion module, and a classification module. The invention also provides an electronic device, such as a server, personal computer, or mobile terminal, which internally includes a memory and a processor, capable of implementing the above method by running a stored program.
[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting fake comments that integrates deep semantics and emotional dynamics, characterized in that, The method includes the following steps: Step (1) Construct and preprocess the comment dataset for model training. The dataset contains a large number of labeled real and fake comment samples. Step (2) performs sentence segmentation on the input comment text to be detected, dividing a single comment text into one or more ordered clauses to form a clause sequence; Step (3) adopts a dual-stream parallel architecture to extract features from the comment text and clause sequence, and obtains deep semantic feature stream and emotional dynamic feature stream respectively; Step (4) Based on the emotional dynamic feature stream obtained in step (3), perform the emotion masking modeling (MSM) auxiliary task to train the model’s ability to understand the emotional context logic; Step (5) Construct a deep learning master model that includes a hierarchical attention mechanism, a feature fusion network and a classifier, and train the master model end-to-end using the dataset from step (1) and a joint loss function; Step (6) Save the fake review identification model with optimal parameters obtained after being fully trained in step (5); Step (7) calls the recognition model obtained in step (6) to process and classify any new comment text to be detected, and outputs the probability that it is a fake comment.
2. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (1), the specific process of constructing and preprocessing the comment dataset is as follows: (1.1) Collect publicly available comment texts with user tags or official platform markings from mainstream e-commerce platforms, social media websites and other channels on a large scale; (1.2) Clean the collected data, remove irrelevant characters and HTML tag noise, and label the comments as "real comments" or "fake comments" based on platform tags, user feedback or other prior knowledge to form the original corpus; (1.3) Organize manual cross-validation and fine-tuning of the initially labeled data to ensure the accuracy of the labels, and finally form high-quality training sets, validation sets and test sets that can be used for supervised learning.
3. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (2), the method for segmenting the comment text into sentences is as follows: (2.1) Analyzing the linguistic features of Chinese comments reveals that the emotional turning points within the comments are often key clues for identifying their authenticity. (2.2) A pre-defined segmentation rule library is established, which includes common transitional conjunctions and explicit delimiters in Chinese. Based on this rule library, the input single comment text T is segmented into an ordered sequence of clauses C = {C1, C2, ..., C...}. n } 4. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (3), the feature extraction process of the dual-stream parallel architecture is as follows: (3.1) Deep Semantic Feature Flow: The complete, unsegmented original comment text T is input into a large-scale pre-trained language model. This model generates a deep semantic word vector h rich in context information for each token in the text through its internal Transformer structure and context self-attention mechanism. i ; (3.2) Emotional dynamic feature flow: Each clause C obtained in step (2) j Each clause is input into a pre-trained, high-quality sentiment analysis model, which generates a corresponding sentiment vector S. j The sentiment vectors of all clauses constitute a sentiment vector sequence {S1, S2, ..., S...} n The sequence explicitly depicts the emotional flow of the comment.
5. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (4), the specific process of training the emotion masking modeling (MSM) auxiliary task is as follows: (4.1) This task aims to enable the model to learn and understand the logical coherence of the sentiment context, traversing the sentiment vector sequence {S1,S2,...,S} generated in step (3.2). n For any index i (from 1 to n-2), construct a training sample; (4.2) The sentiment vector S of the i-th clause i The sentiment vector S of the (i+2)th clause i+2 Concatenate the vectors along the vector dimension to form a context vector: Context_Vector = [S i S i+2 This vector serves as the input to the model; (4.3) Set the prediction target of the model as the true sentiment vector S of the (i+1)th sentence that is "covered" in the middle. i+1 The model uses a multilayer perceptron (MLP) as the prediction head, and its output dimension is consistent with the dimension of the sentiment vector. (4.4) This is a high-dimensional vector regression task, and its loss function L MSM Cosine similarity loss is used to minimize the directional difference between the predicted vector and the true vector. The loss function formula is as follows: in, It is the sentiment vector predicted by the model, S i+1 It is a true vector of emotion.
6. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (5), the specific process of constructing and training the deep learning master model is as follows: (5.1) Construct the main model architecture, which consists of hierarchical attention units, feature fusion units, and classification units: Hierarchical attention unit: Includes intra-sentence attention and inter-sentence attention. Intra-sentence attention is used to aggregate word vectors within each clause to form a clause-level semantic representation V. j Inter-sentence attention borrows from the Transformer Encoder structure, employing multi-head self-attention, residual connections, and layer normalization to process the representation V of all clauses. j Information exchange is performed to capture long-distance dependencies, ultimately outputting a global semantic representation vector X. Simultaneously, an emotion attention unit processes the emotion vector sequence {S1, S2, ..., S...}. n We perform a weighted summation to obtain the global sentiment dynamic vector S. Feature fusion unit: The global semantic representation vector X and the global sentiment dynamic vector S obtained in the previous step are concatenated along the feature dimension to form a fusion vector [X; S]. This fusion vector is then input into a fusion network consisting of fully connected layers for deeper information interaction. The classification unit consists of a linear output layer with a sigmoid activation function connected at the end of the fusion network. This output layer determines the final probability that the review is fake. (5.2) Training is performed using a joint loss function. To achieve end-to-end optimization and fully utilize the advantages of multi-task learning, a joint loss function L is adopted. Total The entire model is trained using a function consisting of three weighted parts: L Total =L main +α·L MLM +β·L MSM Among them, L main It is the binary cross-entropy loss for the main task (i.e., true / false binary classification); L MLM It is a masked language modeling loss derived from the BERT semantic stream, used as an auxiliary regularization task; L MSM The loss is derived from the emotion masking modeling task in step (4), and α and β are hyperparameters used to balance the weights of each loss.
7. The method for detecting fake comments by integrating deep semantics and emotional dynamics according to claim 1, characterized in that, In step (7), the process of using the trained model for recognition is as follows: (7.1) Capture or receive a new comment text to be detected; (7.2) Following the methods in steps (2) and (3), the new comments are processed by sentence segmentation and dual-stream feature extraction; (7.3) Input the extracted features into the recognition model that has been trained and saved in step (6); (7.4) The model forward propagates and calculates through its internal hierarchical attention, fusion and classification units, and finally outputs a probability value between 0 and 1. A threshold of 0.5 is set. If the output value is greater than the threshold, the comment is judged to be a fake comment, otherwise it is a real comment.
8. A fake comment perception system integrating deep semantics and emotional dynamics, characterized in that, The system includes: A data preprocessing module is configured to perform the steps described in steps (1) and (2); A feature extraction and model training module is configured to perform the steps described in steps (3), (4), and (5); a recognition and classification module is configured to perform the steps described in step (7).
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a method for detecting fake comments that integrates deep semantics and emotional dynamics as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, the computer instructions implement a method for detecting fake comments that integrates deep semantics and emotional dynamics as described in any one of claims 1-7.
Citation Information
Cited By
Methods, apparatus, computer equipment, and storage media for analyzing e-commerce user reviews
CN122415143A