Multi-mode self-adaptive online shopping comment authenticity and score credibility evaluation system
By integrating text, images, and user behavior information, a multimodal adaptive online shopping review authenticity and rating credibility assessment system has been developed. This system solves the problems of fake review identification and rating manipulation, achieving high-precision, multi-dimensional review evaluation and improving the accuracy and adaptability of the assessment.
Patent Information
- Application Number
- CN202511643454.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies struggle to achieve high-precision, multi-dimensional, and adaptive review quality assessment in complex e-commerce environments, are unable to effectively identify fake reviews and rating manipulation, and have poor scenario adaptability.
A multimodal adaptive online shopping review authenticity and rating credibility assessment system is adopted. By enhancing text semantic understanding, image authenticity verification and deep user behavior modeling, combined with adaptive weight allocation and two-dimensional prediction modules, it integrates text, image and user behavior information to achieve adaptive weight allocation and accurate assessment.
It enables precise two-dimensional evaluation of the authenticity of online shopping reviews and the credibility of ratings, improving the accuracy and adaptability of the evaluation, and effectively identifying fake reviews and rating manipulation.
Smart Images

Figure CN121544332A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of online shopping review analysis technology, and in particular to a multimodal adaptive online shopping review authenticity and rating credibility assessment system. Background Technology
[0002] With the rapid development of e-commerce and the widespread adoption of online shopping, product reviews have become a core reference for consumers' purchasing decisions. Their authenticity and credibility are directly linked to the credibility of e-commerce platforms and the quality of the user's shopping experience. However, current online shopping review systems face the serious problem of rampant fake reviews, manifested in various forms such as organized review manipulation, malicious negative reviews, cashback incentives for positive reviews, and store employees posting reviews on behalf of customers. Such behavior not only misleads consumers but also severely damages the healthy development of the e-commerce ecosystem.
[0003] Existing technologies are insufficient to meet the demands for high-precision, multi-dimensional, and adaptive review quality assessment in complex e-commerce environments. There is an urgent need to develop an intelligent system that can integrate multimodal information such as text, images, and user behavior to achieve adaptive weight allocation and dual-dimensional assessment of review authenticity and rating credibility. This system would solve technical challenges such as difficulty in identifying fake reviews, weak detection of rating manipulation, and poor scenario adaptability, thereby ensuring the healthy development of the e-commerce ecosystem and protecting consumer rights. Summary of the Invention
[0004] The purpose of this invention is to provide a multimodal adaptive system for evaluating the authenticity of online shopping reviews and the credibility of ratings, in order to solve the problems in the existing technology, such as difficulty in identifying fake reviews, insufficient integration of multimodal information, weak detection capability of rating manipulation, and lack of adaptive adjustment. This system can achieve accurate two-dimensional evaluation of the authenticity of online shopping reviews and the credibility of ratings, thereby improving the accuracy and adaptability of the evaluation.
[0005] To achieve the above objectives, this invention provides a multimodal adaptive online shopping review authenticity and rating credibility evaluation system, comprising: The enhanced text semantic understanding module uses a BERT pre-trained model to semantically encode the comment text, detects tone expression through a regular expression-based pattern matching algorithm, identifies return comments using keyword density analysis and template similarity calculation, and extracts multi-dimensional text features including sentiment tendency, degree of sarcasm, and language standardization by combining sentiment analysis and language quality assessment algorithms. The image authenticity verification module uses a ResNet deep convolutional neural network to extract visual features of user-uploaded images. It identifies Photoshop processing traces through edge consistency detection, noise pattern analysis, and compression trace detection algorithms. It detects image theft by using perceptual hash matching and reverse image search technology, and generates an image credibility score by combining EXIF metadata analysis and shooting environment rationality assessment. The user behavior deep modeling module constructs behavioral feature vectors based on users' historical comment data. It identifies organized fake comment behavior patterns through comment frequency anomaly detection, rating distribution analysis, and content similarity calculation. It evaluates user credibility using time series analysis and social network graph structure analysis algorithms, and uses a multilayer perceptron to encode and reduce the dimensionality of behavioral features. The adaptive weight allocation module calculates cross-modal correlation and dynamically adjusts the fusion weights based on the quality assessment scores of each modality feature through a multi-head attention mechanism, and inputs the fused multimodal features into the two-dimensional prediction module. The dual-dimensional prediction module constructs a rating tendency classifier and a credibility evaluator, respectively. It integrates multimodal information through a weighted feature fusion layer, and performs joint training using a weighted combination of cross-entropy loss function and mean squared error loss function. It also utilizes a multi-task learning framework to simultaneously output the sentiment tendency and authenticity evaluation results of the comments.
[0006] Preferably, the enhanced text semantic understanding module includes: BERT Semantic Encoding Unit: The BERT-Base-Chinese pre-trained model based on the Transformer architecture is adopted. It is pre-trained on a large-scale text corpus and learns language representation through two pre-training tasks: Masked Language Model and Next SentencePrediction. The input comment text is segmented into WordPiece and mapped into a 768-dimensional vector. Semantic information is captured through a 12-layer bidirectional encoder and a multi-head self-attention mechanism. The tone detection unit constructs a detection dictionary containing several ironic and satirical markers, covering categories such as fake praise, irony, implied dissatisfaction, puns and ambiguity, and exaggerated satire. It calculates the sentiment polarity inconsistency score through BERT context analysis and uses a gradient boosting decision tree to quantify the degree of satire. Cashback review identification unit: Establish a cashback inducement keyword dictionary, calculate keyword weight density through TF-IDF algorithm, use text similarity clustering algorithm to detect sentence structure similarity and vocabulary repetition rate of reviews from the same store, use edit distance algorithm and semantic similarity algorithm to analyze the matching degree between reviews and preset templates, and identify organized cashback reviews by combining review time interval, account activity and review style consistency features; Multi-dimensional feature extraction unit: TextCNN convolutional neural network and BiLSTM bidirectional long short-term memory network are used to extract deep semantic features of text. The sentiment lexicon matching algorithm calculates three categories of sentiment tendency scores: positive, negative and neutral, based on the sentiment lexicon and BosonNLP sentiment lexicon. A six-dimensional feature system for authenticity assessment is constructed, including features of real people's reviews, features of store staff posting on behalf of others, features of organized review brushing, and features of malicious negative reviews. Through a multi-task learning framework, the probability distribution of review source type and credibility assessment score are output simultaneously to achieve accurate identification and quantitative assessment of the authenticity of reviews.
[0007] The preferred six-dimensional feature system for authenticity assessment is as follows: The features of real-person comments are as follows: the uniqueness and randomness of users' language habits are evaluated through natural language expression diversity analysis; the naturalness of language expression is detected by lexical diversity index and grammatical structure complexity; users' unique expression preferences and language markers are detected by personalized word recognition algorithm; a personalized lexical feature library is established to identify users' unique language style; the coherence and rationality of emotional expression are evaluated by emotional authenticity analysis; and the naturalness of emotional fluctuations is detected by emotional intensity changes. The system identifies the characteristics of store clerks posting reviews by recognizing the similarity between reviews and standardized templates through templated expressions. It establishes a positive review template library and a customer service terminology dictionary, and uses edit distance and semantic similarity algorithms to identify templated expressions. It analyzes the frequency of use of professional terms to statistically analyze the density of professional vocabulary and technical parameters of products. When the use of professional terms exceeds the average level of ordinary users, it is identified as a store clerk characteristic. It also detects the time clustering pattern of accounts posting a large number of reviews in a short period of time through time concentration analysis. The system identifies organized comment-boosting features by calculating semantic similarity and sentence structure repetition between comments through content homogenization detection, using a text fingerprint algorithm to identify batch-generated similar content, identifying overused extreme adjectives and absolute statements through exaggerated expression detection, establishing a dictionary of exaggerated modifiers to detect abnormal expression patterns, and analyzing the time patterns and frequency anomalies of comment posting through batch posting pattern identification to detect concentrated posting behavior within a short period of time. Malicious negative reviews are characterized by identifying excessively intense negative emotional expressions through extreme negative emotion detection, establishing a negative emotion intensity scoring system to quantify the degree of emotional extremeness, detecting personal attacks and malicious defamation through aggressive language identification, establishing an aggressive vocabulary dictionary containing abusive and threatening language, and identifying implicit malicious attack intentions through malicious slander word density analysis to statistically analyze the frequency of use of slanderous expressions.
[0008] Preferably, the image authenticity verification module includes: The ResNet visual feature extraction unit employs the ResNet-50 deep convolutional neural network architecture to extract five types of visual features from the input image: color histogram, texture, shape, edge, and local binary pattern. It captures image detail information at different resolution levels through a multi-scale feature pyramid network and outputs a high-dimensional feature vector representation. The PS trace detection unit employs a Canny edge detector to extract image edge information, utilizes Hough transform to detect the continuity of lines and curves, identifies splicing boundaries by calculating the degree of abrupt changes in edge gradient directions, and determines the presence of splicing traces when the edge discontinuity threshold is set above 0.15. A noise pattern analysis algorithm is used to decompose the image into multiple frequency sub-bands through wavelet transform, and statistical methods are used to analyze the distribution characteristics of noise in each frequency domain, detecting inconsistencies in noise patterns and traces of manual processing. When the noise variance difference exceeds 0.2, it is determined to be PS processing. A compression trace detection algorithm analyzes the block artifacts and quantization errors generated by JPEG compression, and uses the statistical characteristics of DCT coefficients to detect repeated compression traces. When the compression quality factor difference exceeds 10%, it is determined to be secondary editing processing. Image theft detection unit: Employs perceptual hash matching technology to convert images into 64-bit binary fingerprints using the pHash algorithm. It calculates the similarity between two image fingerprints using Hamming distance, identifying them as similar images when the Hamming distance is less than 5. Combined with reverse image search technology, it uses the SIFT feature point matching algorithm and a search engine interface to retrieve similar images from internet image databases. The source of the image is determined by the number of matching feature points and geometric consistency verification. When the number of matching feature points exceeds 30 and the geometric consistency score is greater than 0.8, it is considered image theft. The EXIF metadata analysis unit extracts EXIF data from image files, verifies the authenticity and consistency of the shooting device through a device fingerprint analysis algorithm, detects the reasonableness of the shooting time and the comment posting time using a timestamp analysis algorithm, analyzes the matching degree between GPS coordinates and the product sales area using a geolocation verification algorithm, and analyzes the technical reasonableness and environmental consistency of shooting parameters through a shooting environment reasonableness evaluation algorithm. When the parameter anomaly exceeds 0.3, it is judged as a suspicious image. Image credibility assessment unit: Construct a credibility assessment system that includes five dimensions: originality, completeness, authenticity, timeliness and matching. Use a weighted average algorithm to calculate the comprehensive credibility score and output a normalized credibility score in the range of 0-1.
[0009] Preferably, the user behavior deep modeling module includes: The user history behavior data collection unit constructs a user profile database, collects users' recent complete comment trajectory data, and establishes a 32-dimensional user behavior feature vector, including time feature dimension, frequency feature dimension, content feature dimension, interaction feature dimension, product feature dimension, rating feature dimension, social feature dimension, and anomaly feature dimension. Standardized feature vectors are constructed through preprocessing. Comment frequency anomaly detection unit: It adopts an anomaly detection algorithm based on sliding time window, sets four time window granularities, establishes a baseline of normal user comment frequency through a Poisson distribution model, triggers an anomaly warning when the comment frequency exceeds 3 times the standard deviation of the baseline, uses the Isolation Forest algorithm to detect outliers in the time series, identifies dense bursts of comments in a short period of time through the DBSCAN density clustering algorithm, and calculates the deviation between the comment posting time and the distribution of normal user behavior by combining the chi-square test. Rating Distribution Anomaly Analysis Unit: The unit uses a mixed model of Beta and normal distribution to fit the user's historical rating patterns, evaluates the degree of rating distribution anomaly through the Kolmogorov-Smirnov test, establishes the expected rating distribution matrix under the five-star rating system, and marks anomalies when the chi-square distance between the user rating distribution and the expected distribution exceeds a threshold. The rating diversity is calculated through rating entropy value, and the rating transition probability is analyzed using a Markov chain model to identify rating manipulation behavior. Content similarity calculation unit: Doc2Vec document vectorization is used to convert user history comments into semantic vectors. Cosine similarity is used to calculate the semantic similarity between comments. A similarity threshold of 0.85 is set. When the proportion of similar comments exceeds 30%, it is judged as an abnormal content duplication. The edit distance algorithm and the longest common subsequence algorithm are used to detect the similarity of text strings. The SimHash algorithm is used to generate fingerprint features. The similarity of comment fingerprints is calculated by Hamming distance. Multi-level content similarity is comprehensively evaluated by combining TF-IDF term frequency weight. Organized Fake Review Behavior Pattern Identification Unit: Establish an organized review-boosting behavior feature identification system that includes time features, content features, rating features, social features, and product features, including 12 specific feature indicators, and use a random forest classifier to comprehensively evaluate organized review-boosting behavior; Time Series Behavior Analysis Unit: The ARIMA autoregressive integral moving average model is used to model the time series of user comments. The difference order is determined by the ADF stationarity test. An ARIMA(p,d,q) model is established to predict the future comment behavior patterns of users. An LSTM long short-term memory network is used to capture the long-term dependencies of user behavior. The Prophet time series prediction algorithm is used to analyze the trend, periodicity and seasonality of user behavior. A change point detection algorithm is used to identify sudden changes in user behavior patterns. Social Network Graph Structure Analysis Unit: Constructs a multi-layered social network graph based on user attention relationships, interactive behaviors, and joint evaluations of products; establishes a complex network model including node attributes and edge attributes; calculates user authority scores using the PageRank algorithm; identifies abnormal clustering patterns using the Community Detection algorithm; performs community partitioning using the Louvain algorithm; evaluates user network influence through centrality analysis; and calculates user node clustering coefficients to assess the authenticity of social relationships. User Credibility Comprehensive Evaluation Unit: Establish a user credibility evaluation model based on the AHP (Analytic Hierarchy Process) method, construct a hierarchical structure including six dimensions: behavioral consistency, content authenticity, time reasonableness, social health, rating reliability, and historical credibility. Use fuzzy comprehensive evaluation method to handle the quantification of qualitative indicators, establish a five-level credibility scoring system, and learn the probabilistic dependencies between features through a Bayesian network model. The multilayer perceptron encoding dimensionality reduction unit adopts a deep neural network architecture containing an input layer, hidden layers, and an output layer. The activation function is the ReLU function, and batch normalization is used to accelerate convergence. The Adam optimizer is used for gradient descent, and the dropout rate is set to prevent overfitting. The PCA algorithm is used to perform dimensionality reduction preprocessing on the high-dimensional features. Finally, the 32-dimensional original features are encoded into 8-dimensional latent feature representations, including user activity factor, content quality factor, time pattern factor, social influence factor, rating stability factor, abnormal behavior factor, credibility factor, and risk warning factor.
[0010] Preferably, the adaptive weight allocation module includes: The multimodal feature quality assessment unit calculates the quality score of each modality, representing the degree of authenticity and credibility, and normalizes it to the 0-1 range; The multi-head attention cross-modal association computation unit employs a 12-head parallel attention mechanism, with each attention head set to 64 dimensions. Text features, image features, and user behavior features are mapped to query vectors, key vectors, and value vectors respectively through linear transformations. Attention weights are calculated by the dot product of the query vector and the key vector, and after scaling and normalization, they are multiplied by the value vector to obtain the output features. Cross-modal correlation is quantified by calculating the mutual information between features of different modalities, and the kernel density estimation method is used to calculate the degree of information sharing between continuous variables to measure the statistical dependence between modalities. A 3×3 cross-modal attention matrix is constructed to represent the association strength between three pairs of modalities: text and image, text and behavior, and image and behavior. Intermodal interaction features are calculated through bilinear pooling operations, and a learnable weight matrix is used to model the nonlinear association between modalities. The dynamic weight normalization unit introduces a learnable temperature parameter, which is adaptively adjusted through the backpropagation algorithm. The temperature parameter controls the smoothness of the exponential normalization function, and the adaptive control of weight allocation is achieved by dynamically adjusting the temperature value. A gating mechanism is adopted to adjust the weights of each mode by combining the quality score, cross-modal correlation score and attention weight. The gating unit controls the information flow by learning parameters and bias terms. Finally, the fused weight is obtained by normalizing the product of the temperature-adjusted weight and the gating output. The gradient descent weight optimization unit employs an adaptive moment estimation optimization algorithm, defining a composite loss function that includes classification loss, regularization loss, and consistency loss. The optimization process is achieved by calculating the gradient of the loss function with respect to the parameters and updating the parameters.
[0011] Preferably, the two-dimensional prediction module includes: The rating tendency classification unit adopts a multi-layer perceptron architecture based on an attention mechanism. It takes the fused multimodal feature vector as input and performs nonlinear transformation through three fully connected layers to construct a five-class classification system for the rating prediction task. The output layer uses the Softmax normalization function to map the features to the probability distribution of each rating level. The classification unit identifies rating tendency by learning key semantic features in the review text. It combines image quality rating and user historical rating behavior patterns, and uses a weighted voting mechanism to comprehensively predict the final rating level by integrating multimodal information. The confidence of the output probability distribution is calibrated by temperature scaling technology. The credibility assessment unit is built on a deep neural network regression model and adopts a 4-layer deep network structure with residual connections. Each layer includes batch normalization and Dropout regularization techniques. The input layer receives a comprehensive feature vector containing text authenticity features, image authenticity features, and user behavior credibility features. It learns a continuous numerical representation of comment authenticity through nonlinear mapping. The output layer uses the Sigmoid activation function to constrain the prediction results to the range of 0 to 1 to represent the credibility score, where 0 represents completely untrustworthy and 1 represents completely trustworthy. The assessment unit constructs a six-dimensional feature system for authenticity assessment and simultaneously outputs the probability distribution of 4 types of comment source types and the comprehensive credibility assessment score through a multi-task learning framework. The weighted feature fusion unit is designed with an adaptive weight learning mechanism to integrate feature representations of text, images, and user behavior. The importance weights of each modality feature are calculated through a gated fusion unit. The gating mechanism adopts an attention-based weight allocation strategy. The calculation formula is that the attention weight is equal to the dot product of the feature vector and the learnable parameter matrix, after Tanh activation, and then the weight vector. The output feature vector of the fusion unit is set to 1024 dimensions, which includes global semantic information, local detailed features, and cross-modal correlation features. Residual connections are used to preserve the original feature information, and layer normalization technology is used to stabilize the training process. The joint loss function design unit uses the cross-entropy loss function to measure the difference between the predicted probability distribution and the true label distribution for the rating and classification task. The cross-entropy loss is calculated by taking the negative logarithm of the predicted probability at the corresponding position of the true label. For the credibility regression task, the mean squared error loss function is used to calculate the squared difference between the predicted value and the true credibility score. The overall loss function adopts a weighted linear combination form, and the weight coefficients are dynamically adjusted through the performance of the validation set. The multi-task learning unit adopts a hard parameter shared architecture. The bottom feature extraction network is a shared layer, and the top layers are constructed with two task-specific output heads for rating classification and credibility assessment, respectively. The shared layer contains 4 layers of convolutional neural networks and 2 layers of recurrent neural networks. The task-specific layers are optimized for classification and regression tasks, respectively. The learning progress of the two tasks is balanced through an alternating training strategy, and an adaptive adjustment mechanism for task weights is set. The learning unit outputs a 5-dimensional probability vector representing the rating tendency distribution and a 1-dimensional continuous value representing the credibility score. The post-processing module converts the probability distribution into the expected rating value and combines it with the credibility score to provide a confidence quantification for the final prediction result, achieving accurate prediction of both the authenticity of the review and the rating tendency in two dimensions.
[0012] Therefore, the present invention adopts the above-mentioned multimodal adaptive online shopping review authenticity and rating credibility evaluation system, which solves the problems of difficulty in identifying fake reviews, insufficient integration of multimodal information, weak rating manipulation detection capability and lack of adaptive adjustment in the prior art, and realizes accurate two-dimensional evaluation of online shopping review authenticity and rating credibility, thereby improving the evaluation accuracy and scenario adaptability.
[0013] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall system architecture according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the enhanced text semantic understanding module structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the image authenticity verification module structure of the present invention; Figure 4 This is a schematic diagram of the user behavior deep modeling module structure according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the adaptive weight allocation module structure according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the two-dimensional prediction module structure according to an embodiment of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0016] Figure 1 This is a schematic diagram of the overall architecture of the multi-modal adaptive online shopping review authenticity and score credibility evaluation system of the present invention, including an enhanced text semantic understanding module, an image authenticity verification module, a user behavior deep modeling module, an adaptive weight allocation module, and a two-dimensional prediction module; In this embodiment, the system is deployed in a distributed cloud server cluster environment and is connected in real time with the review systems of major e-commerce platforms through standardized API interfaces, capable of handling real-time analysis tasks of tens of thousands of review data per second, providing 7×24-hour uninterrupted review quality monitoring services for the platform. The entire system adopts a microservices architecture design, and asynchronous communication is carried out between each functional module through a message queue, ensuring the high availability and scalability of the system, and at the same time supporting horizontal scaling to meet the processing requirements during peak business periods. The specific implementation is as follows: Figure 2 This is a schematic diagram of the structure of the enhanced text semantic understanding module, including: BERT semantic encoding unit: Adopting the BERT-Base-Chinese pre-training model, pre-trained based on Chinese Wikipedia, Baidu Encyclopedia, news corpus, and e-commerce review data, taking the reviews of clothing products on e-commerce platforms as the evaluation object, the input review text is "The quality of the clothes is quite good, but the price is a bit expensive, but it's still worth it". Perform WordPiece tokenization on the review text to obtain a sub-word sequence, map the review text after WordPiece tokenization to a 768-dimensional vector, capture semantic information through 12 layers of bidirectional encoders and multi-head self-attention mechanisms, and be able to capture the uncertain emotion hidden in "quite good" and the price-sensitive information in "a bit expensive", and obtain the global semantic representation vector of the entire sentence review through the [CLS] special token, providing a high-quality semantic basis for subsequent authenticity analysis.
[0017] Irony and sarcasm expression recognition unit: Construct a detection dictionary of 186 irony and sarcasm marker words, covering 47 words of false praise type, 39 words of irony type, 35 words of implied dissatisfaction type, 28 words of pun and ambiguity type, and 37 words of hyperbole and sarcasm type. These words seemingly express affirmation or praise, but often carry ironic or dissatisfied implicit meanings in specific contexts. Calculate the emotional polarity inconsistency score through BERT context analysis. For example, for the review "This mobile phone is really great, but it broke after using it for one day", the system can recognize the ironic meaning of "really great" in this context, and quantify the irony degree as a high irony score of 0.85 through a gradient boosting decision tree, thus accurately judging that this review has the expression characteristics of irony and sarcasm; Cashback Review Recognition Unit: A comprehensive dictionary of 420 cashback inducement keywords has been established, including core keywords such as "five-star review," "contact customer service," "show pictures," and "cashback." Taking a review of an electronic product as an example, "The product quality is good, five-star review, contact customer service for more discounts, remember to show pictures!" The TF-IDF algorithm is used to calculate the keyword weight density. It is found that the concentrated occurrence of keywords such as "five-star review," "contact customer service," and "show pictures" causes the density value to reach 0.31, exceeding the set threshold of 0.25, and is therefore judged as a cashback-inducing review. At the same time, for batch input scenarios, the system detects multiple structurally similar reviews from the same store within the same time period, such as templated expressions like "I received the item, the quality is very good, the logistics is fast, five-star review." When the number of reviews with a similarity of more than 0.75 reaches 5, the system automatically identifies it as organized cashback review behavior. Multi-dimensional Feature Extraction Unit: TextCNN convolutional neural network and BiLSTM bidirectional long short-term memory network are used to extract text features. Sentiment scores are calculated based on CNKI and BosonNLP sentiment dictionaries. A six-dimensional feature system is constructed, encompassing genuine reviews, employee-posted reviews, organized review manipulation, and malicious negative reviews. The system also outputs the probability distribution of review sources. For example, a cosmetic review stating, "After a month of use, the effect is acceptable, the packaging is beautiful, and the price is good. Recommended purchase," is analyzed and found to have natural language expression features, a moderate sentiment intensity (0.72 positive sentiment score), and good language standardization, classifying it as a genuine review with a credibility score of 0.89. Conversely, for the review "The product is superb, the quality is first-class, the service attitude is very good, the logistics speed is fast, the packaging is beautiful, the price is extremely high, highly recommended purchase," the system identifies features such as excessive use of positive adjectives, templated expressions, and a lack of personalized details, classifying it as an organized review manipulation feature with a credibility score of only 0.23.
[0018] Figure 3 This is a schematic diagram of the image authenticity verification module structure of the present invention, including: The ResNet visual feature extraction unit employs the ResNet-50 deep convolutional neural network architecture. This network addresses the vanishing gradient problem in deep networks through a residual learning mechanism. It consists of one 7×7 convolutional layer, four residual block groups, and one fully connected layer. Each residual block achieves identity mapping through skip connections. Taking a user-uploaded mobile phone image as an example, the original input image size is 1080×1920×3. It is first adjusted to a standard 224×224×3 input format. After network processing, five types of visual features are extracted, including color histogram features (RGB distribution features, color saturation, brightness distribution), texture features (local binary mode, gray-level co-occurrence matrix), shape features (edge density, contour complexity), and edge features (gradient magnitude, orientation distribution). After processing by a multi-scale feature pyramid network, the system can capture image detail information at different resolution levels, ultimately outputting a 2048-dimensional high-dimensional feature vector representation, providing rich visual semantic information for subsequent realism analysis. The PS trace detection unit employs the Canny edge detector to extract image edge information and utilizes Hough transform to detect the continuity of lines and curves. Taking a clothing product image as an example, the image shows traces of PS processing, specifically replacing the model's head. The system calculates the abrupt change in the edge gradient direction of the neck region and finds that the edge discontinuity reaches 0.18, exceeding the set threshold of 0.15, successfully identifying the splicing boundary. Simultaneously, the noise pattern analysis algorithm decomposes the image into multiple frequency sub-bands using wavelet transform, detecting a significant difference in noise distribution between the model's head region and body region, with a noise variance difference reaching 0.25, exceeding the judgment threshold of 0.2, further confirming the existence of PS processing. The compression trace detection algorithm analyzes the block effect produced by JPEG compression and finds that the compression quality factor of the head region differs from that of the body region by 15%, exceeding the threshold standard of 10%, indicating that the image has undergone multiple editing and recompression processes.
[0019] Image theft detection unit: Employing perceptual hash matching technology, the system generates a 64-bit binary fingerprint using the pHash algorithm. Taking an image from an electronic product review as an example, the system extracts low-frequency features from the image using discrete cosine transform, generating a hash fingerprint of "A3B5C7D9E1F2A4B6C8DA2B4C6D8E0F1A3B5C7D9E1F2A4B6C8DA2B4C6D8E0F1". This fingerprint is then compared with 10 million product images in the image database. The system finds that the hash fingerprint of a product image from an official online store has a Hamming distance of only 3, less than the set similarity threshold of 5, initially indicating image theft. Further detection using the SIFT feature point matching algorithm reveals 127 matching feature points between the two images, far exceeding the 30-point threshold, and the geometric consistency score reaches 0.92, exceeding the 0.8 confidence threshold. Ultimately, it is confirmed that the review image is a product image stolen from official channels, not a photo taken by the user after actual purchase. The EXIF metadata analysis unit extracts 15 key parameters from the image file, including detailed information such as the shooting device model, shooting time, GPS location, aperture value, shutter speed, ISO sensitivity, focal length, and white balance mode. Taking a cosmetic review image as an example, the EXIF data extracted by the system shows that the shooting device is "Canon EOS 5D Mark IV", the shooting time is "2024:03:15 14:23:07", the GPS coordinates are "39.9042°N, 116.4074°E", the aperture value is F / 2.8, the shutter speed is 1 / 60 second, and the ISO sensitivity is 400. Device fingerprint analysis shows that the camera model is consistent with the shooting devices used in the user's historical reviews, reflecting the continuity of device use; timestamp analysis shows that the time interval between the shooting time and the posting of the review is only 2 hours, which is consistent with normal usage scenarios; geolocation verification shows that the GPS coordinates are located in Chaoyang District, Beijing, which matches the main sales area of the product. Based on a comprehensive analysis of the technical rationality and environmental consistency of various parameters, the system determined that the parameter anomaly of the image was only 0.12, far below the suspicious threshold of 0.3, indicating that the image has a high degree of authenticity.
[0020] The image credibility assessment unit constructs a credibility assessment system encompassing five dimensions: originality, completeness, authenticity, timeliness, and matching. Taking a home furnishing review image as an example, the originality assessment, based on image theft detection results, found that the image was taken by the user, with a uniqueness score of 0.93; the completeness assessment, based on Photoshop trace detection results, showed that the image was not modified, with a completeness score of 0.87; the authenticity assessment, based on EXIF metadata verification, confirmed that the shooting parameters were authentic and valid, with an authenticity score of 0.91; the timeliness assessment showed that the shooting time highly matched the review posting time, with a timeliness score of 0.95; and the matching assessment, based on product feature matching analysis, confirmed that the image content was highly consistent with the review text description, with a matching score of 0.89. A weighted average algorithm was used to calculate the comprehensive credibility score, with the weights for each dimension being originality 0.25, completeness 0.25, authenticity 0.20, timeliness 0.15, and matching 0.15, respectively. The final comprehensive credibility score for this image was 0.90, indicating that the image has a high level of credibility.
[0021] Figure 4 This is a schematic diagram of the user behavior deep modeling module structure according to an embodiment of the present invention, including: The user history behavior data collection unit constructs a user profile database containing eight basic dimensions, including user ID, comment timestamp, product category, rating value, comment length, number of likes, number of replies, and number of reports. It collects complete comment trajectory data of users over the past 24 months and establishes a 32-dimensional user behavior feature vector. Taking user "UserID_12345" as an example, this user posted 186 comments in the past 24 months. The time feature dimension extracted by the system shows that the comment time distribution entropy value is 0.73 (indicating that the comment time distribution is relatively random), the average comment interval is 3.8 days, the active period is mainly concentrated between 19:00-22:00 in the evening (concentration is 0.65), and the periodic comment intensity is 0.31. The frequency feature dimension shows that the average number of comments per day is 0.26, the comment burst coefficient is 1.2 (indicating occasional concentrated comment behavior), the longest silence period is 15 days, and the average duration of continuous activity is 4.2 days. The content feature dimension shows that the average number of words in the comments is 127, the sentiment polarity variance is 0.34 (indicating diverse sentiment expression), the topic diversity index is 0.78 (involving multiple product categories), and the language complexity score is 0.66. The rating feature dimension shows that the average rating is 4.2 points, the rating standard deviation is 0.87, the proportion of extreme ratings (1 point or 5 points) is 23%, and the consistency between rating and text sentiment is 0.84.
[0022] The comment frequency anomaly detection unit employs a sliding time window-based anomaly detection algorithm, setting four time window granularities: 1 hour, 6 hours, 24 hours, and 7 days. Taking a suspicious user "UserID_67890" as an example, this user posted 12 comments within a time window from 9:00 AM to 10:00 AM on a certain day. Historical data shows that the normal baseline for the number of comments per hour is 0.8 ± 1.2, meaning the current frequency exceeds the baseline by 4.2 standard deviations, triggering a high-risk anomaly warning. Analysis of the user's comment time series using the IsolationForest algorithm reveals a clear comment burst pattern at multiple time points, including March 15th, March 22nd, and April 8th. Each burst involves posting 10-15 comments in a short period, followed by a relatively long period of silence. This pattern is significantly different from the commenting behavior of normal users. Further analysis using the DBSCAN density clustering algorithm reveals that the user's commenting behavior forms five distinct clusters along the time dimension, each corresponding to a concentrated commenting activity, indicating organized or task-oriented commenting behavior characteristics.
[0023] The rating distribution anomaly analysis unit uses a mixed model of Beta and normal distributions to fit users' historical rating patterns, and assesses the degree of anomaly in the rating distribution using the Kolmogorov-Smirnov test. An expected rating distribution matrix is established under a five-star rating system, where 1 star accounts for 5%, 2 stars for 8%, 3 stars for 15%, 4 stars for 35%, and 5 stars for 37%. Taking user "UserID_54321" as an example, this user's rating distribution is 1 star 2%, 2 stars 3%, 3 stars 5%, 4 stars 15%, and 5 stars 75%, with a chi-square distance of 18.7 from the expected distribution, far exceeding the set threshold of 12.0, indicating a significant tendency for rating manipulation. Rating entropy calculations reveal that this user's rating diversity is only 0.31, far below the average of 0.72 for normal users, indicating an overly simplistic rating pattern. Markov chain model analysis of rating transition probabilities reveals that the probability of this user switching from a 4-star rating to a 5-star rating is as high as 0.89, while the probability of a similar transition for normal users is only 0.34, further confirming the existence of rating manipulation behavior.
[0024] The content similarity calculation unit uses Doc2Vec document vectorization technology to convert user history comments into 300-dimensional semantic vectors, and calculates the semantic similarity between comments using cosine similarity. Taking user "UserID_98765" as an example, the user's comments include "The product quality is very good, the logistics speed is very fast, the packaging is exquisite, and it is worth recommending," "The product quality is good, the express delivery is fast, the packaging is very good, and it is recommended to buy," and "The product quality is great, the delivery is fast, the packaging is very careful, and it is worth buying." Semantic similarity analysis reveals that 85% of the comments have a similarity exceeding the set threshold of 0.85, far higher than the average level of 30% for normal users. Edit distance detection shows that these comments also have a string-level similarity of 0.78, indicating obvious templated expression characteristics. The 64-bit fingerprint feature generated using the SimHash algorithm shows that the Hamming distance of multiple comments from this user is less than 8, indicating that the comment content is highly repetitive, lacks personalized expression, and conforms to the characteristic pattern of mass production or template filling.
[0025] The organized fake review behavior pattern identification unit establishes a feature identification system for organized review brushing behavior, which includes five major categories: time features, content features, rating features, social features, and product features. The organized review brushing behavior is comprehensively evaluated by a random forest classifier. Taking a suspected organized review-boosting account as an example, this account exhibits the following characteristics in terms of time: abnormally concentrated comment periods (90% of comments are concentrated in specific time periods on weekdays), regular comment intervals (commenting activity every 2-3 days), and bursts of mass comments (up to 27 comments per day); in terms of content characteristics, it exhibits highly consistent comment length (95% of comments are between 80-120 characters), templated emotional expression (using similar sentence structures), and professional product descriptions (frequent use of product terminology); in terms of rating characteristics, it exhibits a clear tendency towards extreme ratings (89% are 5-star reviews), and the ratings are highly consistent with the emotional content but lack supporting details; in terms of social characteristics, it exhibits high activity as a new account (registered for only 3 months but already published 200+ comments), lack of interactive behavior (rarely receiving likes or replies), and abnormal friend relationships (following too many merchants but lacking genuine interaction). Through comprehensive evaluation using a random forest classifier, the probability of this account being identified as an organized review-boosting account reaches 0.94, indicating an extremely high risk level.
[0026] The multilayer perceptron encoding dimensionality reduction unit employs a deep neural network architecture comprising an input layer (32-dimensional), hidden layers (64-32-16 neurons), and an output layer (8-dimensional). The ReLU activation function is used, and batch normalization is employed to accelerate convergence. The Adam optimizer is used for gradient descent, and a dropout rate of 0.3 is set to prevent overfitting. Taking the aforementioned user "UserID_12345" as an example, its original 32-dimensional behavioral feature vector is encoded into an 8-dimensional latent feature representation after network processing: user activity factor 0.67 (indicating moderate activity level), content quality factor 0.78 (indicating high-quality comment content), time regularity factor 0.52 (indicating relatively regular comment timing), social influence factor 0.34 (indicating moderate social influence), rating stability factor 0.71 (indicating relatively stable rating behavior), abnormal behavior factor 0.19 (indicating low risk of abnormal behavior), credibility factor 0.82 (indicating high account credibility), and risk warning factor 0.23 (indicating low risk warning level). This dimensionality reduction representation retains the core features of user behavior while significantly reducing computational complexity, providing efficient feature input for subsequent multimodal fusion.
[0027] Figure 5 This is a schematic diagram of the adaptive weight allocation module structure according to an embodiment of the present invention, including: Multimodal quality assessment unit: Calculates quality scores for text, images, and behavior separately; quality scores are calculated based on four dimensions: real-person review recognition, store clerk posting detection, organized review manipulation recognition, and malicious negative review detection. Taking a 3C electronics product review as an example: "I've used it for half a month, and overall it's pretty good, except the battery life is average, the photo quality is satisfactory, and the price-performance ratio is acceptable." This review has a real-person review recognition score of 0.87, a store clerk posting detection score of 0.12, an organized review manipulation recognition score of 0.08, a malicious negative review detection score of 0.05, and a comprehensive text quality score of 0.85. Regarding image modal features, taking user-uploaded product usage scenario images as an example, the originality assessment score is 0.92, the completeness assessment score is 0.88, the relevance assessment score is 0.91, the authenticity assessment score is 0.89, and the overall image quality score is 0.90. Regarding user behavior modal features, taking a real user as an example, the behavior authenticity score is 0.83, the time reasonableness score is 0.79, the interaction naturalness score is 0.76, the pattern consistency score is 0.81, and the overall user behavior quality score is 0.80.
[0028] Multi-head attention computation unit: Employing 12 parallel attention heads, the three modal features are mapped to query / key / value vectors, cross-modal correlations are calculated, and a cross-modal attention matrix is constructed. Taking multimodal data from a cosmetic review as an example, the text feature vector is a 768-dimensional semantic representation, the image feature vector is a 2048-dimensional visual representation, and the user behavior feature vector is an 8-dimensional behavioral encoding. Through parallel computation by 12 attention heads, system analysis reveals a strong correlation between the phrase "very good moisturizing effect" mentioned in the text and the delicate texture of the product shown in the image, with a correlation score of 0.82. The user's historical cosmetic purchase behavior pattern also shows a high degree of consistency with the professional description in the current review, with a correlation score of 0.78. The constructed 3×3 cross-modal attention matrix shows that the text-image correlation strength is 0.82, the text-behavior correlation strength is 0.76, and the image-behavior correlation strength is 0.71, indicating a strong mutual support relationship among the three modalities, which improves the credibility of the review's authenticity. Dynamic weight normalization unit: A learnable temperature parameter τ is introduced, initially set to 1.0, and adaptively adjusted through backpropagation. The temperature parameter controls the smoothness of the exponential normalization function. When the temperature approaches 0, the weight distribution tends to be hard-allocated, highlighting the dominant mode; when the temperature approaches infinity, the weight distribution tends to be uniformly distributed, balancing the contributions of each mode. Adaptive control of weight allocation is achieved by dynamically adjusting the temperature value. Taking a controversial product review as an example, the system detected a text quality score of 0.65, indicating slightly stiff language; an image quality score of 0.92, indicating high image authenticity; and a user behavior quality score of 0.71, indicating normal user history behavior. With a Softmax value of 0.8, the weights after Softmax normalization are 0.28 for the text modality, 0.48 for the image modality, and 0.24 for the user behavior modality. The system tends to rely more on image information for judgment; while when... When the weight is 1.5, the weight allocation becomes 0.31 for text modality, 0.38 for image modality, and 0.31 for user behavior modality, presenting a more balanced distribution pattern. Through the adjustment of the gating mechanism, the system finally determines the fusion weight as 0.32 for text modality, 0.43 for image modality, and 0.25 for user behavior modality, which takes into account both the high credibility of the image and maintains an effective balance of multimodal information.
[0029] Gradient descent weight optimization unit: The Adam optimization algorithm is used, with a learning rate set to 0.001 and a momentum parameter... =0.9, second moment parameter =0.999. The composite loss function is defined as consisting of three components: classification loss, regression loss, and regularization loss. In a batch of training, taking a training sample containing 1000 comments as an example, the classification loss uses cross-entropy loss to calculate the difference between the predicted label and the true label, resulting in... =0.234; The regression loss is calculated using the mean squared error loss to determine the difference between the predicted value and the true confidence score. =0.089; Regularization loss through Norm constraints are used to constrain model weight parameters to prevent overfitting, resulting in... =0.018. The overall loss function is the weighted sum of all loss terms, and its calculation formula is as follows: ; Substituting the above values, = 0.6×0.234+0.4×0.089+0.001×0.018 ≈ 0.177. By limiting the upper bound of the gradient norm to 1.0 through gradient clipping, the gradient explosion problem is effectively prevented. After 500 training epochs, the system's weight allocation strategy gradually converged, and the prediction accuracy of multimodal fusion improved from the initial 85.2% to 92.6%, verifying the effectiveness of the adaptive weight optimization mechanism.
[0030] Figure 6 This is a schematic diagram of the two-dimensional prediction module structure according to an embodiment of the present invention, including: The rating tendency classification unit employs a multilayer perceptron architecture based on an attention mechanism. It takes the fused multimodal feature vectors as input and performs nonlinear transformation through three fully connected layers. The number of neurons in each layer is set to 512, 256, and 128, respectively. The ReLU activation function is used to enhance the model's nonlinear expressive power. Taking a mobile phone product review as an example: "The phone's design is beautiful, the system runs smoothly, but the charging speed is a bit slow, the camera effect is average, and overall the price-performance ratio is acceptable," after processing by the rating tendency classification unit, the system outputs the following probability distributions for each rating level: 1 star probability 0.02, 2 stars probability 0.08, 3 stars probability 0.25, 4 stars probability 0.52, and 5 stars probability 0.13. The final predicted rating is 4 stars (highest probability), which highly matches the moderate preference sentiment expressed in the review: "Overall, the price-performance ratio is acceptable." The classification unit learns key semantic features in the review text such as "relatively smooth", "a bit slow", and "mediocre", combines clear product images uploaded by users (image quality score 0.87) and users' historical rational rating behavior patterns (rating stability factor 0.76), and adopts a weighted voting mechanism to integrate multimodal information. The confidence level of the output probability distribution is calibrated to 0.89 through temperature scaling technology.
[0031] The credibility assessment unit, built upon a deep neural network regression model, employs a 4-layer deep network structure with residual connections. Each layer incorporates batch normalization and Dropout regularization techniques to prevent overfitting. Taking a suspected fake review as an example: "This product is amazing! Superb quality, excellent service, fast shipping, beautiful packaging, and incredibly cost-effective. Highly recommended!", after inputting this review into the credibility assessment unit, the network analysis revealed excessive use of extremely positive vocabulary, a lack of specific usage details and personalized experience descriptions, exaggerated emotional expression, and a clear sales pitch bias. Combined with image analysis, it was found that the user-uploaded image was an official promotional image (originality score only 0.15). User behavior analysis showed that the account had a short registration period but abnormally high activity (risk warning factor 0.78). After comprehensive evaluation, the credibility assessment unit output a authenticity score of only 0.23, indicating that the review is highly likely to be fake. The evaluation unit also constructed a six-dimensional feature system for authenticity assessment. The review performed as follows in each dimension: real person review feature 0.21 (language expression lacks naturalness), store clerk posting feature 0.82 (clearly promotional nature), organized review brushing feature 0.79 (high degree of content template), malicious negative review feature 0.05 (no malicious attack). Finally, the source of the review was determined to be organized review brushing, with a confidence level of 0.89.
[0032] Weighted Feature Fusion Unit: An adaptive weight learning mechanism is designed to integrate feature representations from three modalities: text, image, and user behavior. A gating fusion unit calculates the importance weights of each modality's features. Taking a clothing product review as an example, text features show that the user described their personal experience in detail, including fabric texture, size fit, and color (text quality score 0.84); image features show that the user uploaded multiple photos of the actual wearing effect from various angles (image quality score 0.91); and user behavior features show that the user has a history of purchasing clothing products and maintains a consistent review style (behavior quality score 0.79). A gating attention mechanism is used to assign modal weights, outputting a 1024-dimensional fused feature; the calculation formula is as follows: ; After Softmax normalization, modal weights are assigned: 0.35 for text, 0.42 for image, and 0.23 for behavior; the output is a 1024-dimensional fused feature, which is stabilized by residual connections and layer normalization.
[0033] Multi-task learning unit: Employing a hard parameter-shared architecture, the bottom-level feature extraction network is a shared layer. The top layers construct two task-specific output heads: rating classification and credibility assessment. Alternating training balances task progress, outputting a 1-5 star rating probability distribution with a credibility score in the 0-1 range. Taking the processing of a home appliance review as an example, the input multimodal fusion features first pass through the shared feature extraction layer to extract general feature representations including semantic understanding, sentiment analysis, and authenticity judgment. These are then fed into the two task-specific output heads for specialized processing. The rating classification head outputs a 5-dimensional probability vector [0.03, 0.12, 0.28, 0.45, 0.12] through a 3-layer fully connected network, representing the probability distribution of each rating level, predicting a 4-star rating. The credibility assessment head outputs a continuous value of 0.87 through a 4-layer deep network, representing the credibility score of the review. An alternating training strategy balances the learning progress of the two tasks. When the loss of the rating classification task decreases slowly, the system automatically increases the training weight of that task from 0.6 to 0.7, ensuring coordinated optimization of the two tasks.
[0034] Joint Loss Unit: Employs cross-entropy loss and mean squared error loss, and introduces L2 regularization; in a certain batch of training, taking a training sample containing 500 labeled comments as an example, the cross-entropy loss for the rating classification task is calculated as follows: ,get =0.287; the mean squared error loss for the credibility regression task is calculated as follows: 2 ,get =0.156. The overall loss function adopts a weighted linear combination form, and the weight coefficients are dynamically adjusted based on the validation set performance. Currently, they are set to a classification loss weight of 0.6 and a regression loss weight of 0.4. Regularization is used to prevent overfitting; the regularization coefficient is set to 0.001, and the final loss function is... = 0.177. By limiting the upper bound of the gradient norm to 1.0 through gradient clipping, the gradient explosion problem during training is effectively prevented, ensuring the stability and convergence of model training.
[0035] The post-processing unit converts the rating probability distribution into the expected rating value, combines it with the credibility score to quantify the prediction confidence, and outputs a two-dimensional evaluation result. The learning unit improves the model's generalization ability by jointly optimizing two related tasks, utilizes the correlation between tasks to achieve knowledge transfer, and ultimately outputs both the rating tendency distribution and the credibility score. The post-processing module converts the probability distribution into the expected rating value of 4.1, and combines it with the credibility score of 0.87 to provide a reliable confidence quantification for the final prediction result, achieving accurate two-dimensional prediction of both review authenticity and rating tendency.
[0036] Therefore, this invention employs the aforementioned multimodal adaptive online shopping review authenticity and rating credibility assessment system. By integrating advanced deep learning technology and multimodal information processing capabilities, it achieves high-precision identification of fake reviews and accurate assessment of rating credibility. This provides important technical support and solutions for the healthy development of the e-commerce industry and the protection of consumer rights. It solves problems in existing technologies such as difficulty in identifying fake reviews, insufficient integration of multimodal information, weak detection capabilities for rating manipulation, and lack of adaptive adjustment. It achieves accurate dual-dimensional assessment of the authenticity of online shopping reviews and the credibility of ratings, improving assessment accuracy and scenario adaptability.
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-modal adaptive e-commerce review authenticity and rating credibility evaluation system, characterized in that, Comprise: Enhanced text semantic understanding module, using BERT pre-training model to encode the semantic of review text, detecting tone expression based on regular expression pattern matching algorithm, using keyword density analysis and template similarity calculation to identify cashback comments, combined with sentiment analysis and language quality evaluation algorithm to extract multi-dimensional text features including sentiment orientation, sarcasm degree, language standardization; Image authenticity verification module, using ResNet deep convolutional neural network to extract visual features of user uploaded images, detecting PS processing traces through edge consistency detection, noise pattern analysis and compression trace detection algorithm, using perceptual hashing matching and reverse image search technology to detect plagiarism behavior, combined with EXIF metadata analysis and shooting environment rationality evaluation to generate image credibility score; User behavior deep modeling module, based on user historical review data to construct behavior feature vector, through comment frequency anomaly detection, score distribution analysis and content similarity calculation to identify organized fake comment behavior patterns, using time series analysis and social network graph structure analysis algorithm to evaluate user credibility, and using multilayer perception to encode and reduce dimension of behavior features; Adaptive weight allocation module, based on the quality evaluation score of each modal feature, through multi-head attention mechanism to calculate cross-modal correlation and dynamically adjust the fusion weight, and input the fused multi-modal features into the double-dimensional prediction module; Double-dimensional prediction module, respectively constructing score tendency classifier and credibility evaluator, integrating multi-modal information through weighted feature fusion layer, using weighted combination of cross-entropy loss function and mean square error loss function for joint training, and using multi-task learning framework to output sentiment tendency and authenticity evaluation results of comments at the same time.
2. The multi-modal adaptive e-commerce review veracity and rating credibility evaluation system according to claim 1, characterized in that, The enhanced text semantic understanding module comprises: BERT semantic encoding unit: using BERT-Base-Chinese pre-training model based on Transformer architecture, pre-training based on large-scale text corpus, learning language representation through Masked Language Model and Next Sentence Prediction two pre-training tasks, mapping the input review text into 768-dimensional vector after WordPiece tokenization, capturing semantic information through 12-layer bidirectional encoder and multi-head self-attention mechanism; Tone detection unit: build a detection dictionary containing several irony and sarcasm keywords, covering false praise, irony, implied dissatisfaction, double-meaning ambiguity, and exaggerated sarcasm, calculate the sentiment polarity inconsistency score through BERT context analysis, and use gradient boosting decision tree to quantify the sarcasm degree; Cashback comment recognition unit: establish a cashback keyword dictionary, calculate the keyword weight density through TF-IDF algorithm, detect the sentence structure similarity and word repetition rate of comments in the same store through text similarity clustering algorithm, analyze the matching degree of comments and preset templates using edit distance algorithm and semantic similarity algorithm, combined with comment time interval, account activity and comment style consistency features to identify organized cashback comments; Multi-dimensional feature extraction unit: adopt TextCNN convolutional neural network and BiLSTM bidirectional long short term memory network to extract deep semantic features of text, calculate positive, negative and neutral sentiment tendency scores based on sentiment dictionary and BosonNLP sentiment dictionary through sentiment dictionary matching algorithm; Construct a six-dimensional feature system for authenticity evaluation, including real review features, employee-generated features, organized review features, malicious review features, and output comment source type probability distribution and credibility evaluation score through multi-task learning framework, realize accurate identification and quantitative evaluation of review authenticity.
3. The multi-modal adaptive e-commerce review veracity and rating credibility evaluation system according to claim 2, characterized in that, The six-dimensional feature system for authenticity evaluation is as follows: Real review features: through natural language expression diversity analysis to evaluate the uniqueness and randomness of user language habits, use vocabulary diversity index and syntax structure complexity to detect the naturalness of language expression, detect user-specific expression preferences and language markers through personalized word recognition algorithm, establish personalized vocabulary feature library to identify user's unique language style, evaluate the coherence and rationality of emotional expression through emotional authenticity analysis, and detect the naturalness of emotional fluctuations using emotional intensity change; Employee-generated features: detect the similarity between comments and standardized templates through template expression recognition, establish a good review template library and customer service vocabulary dictionary, identify template expressions using edit distance and semantic similarity algorithms, and detect the time aggregation pattern of account batch publishing through professional term usage frequency analysis and statistical density of commodity professional vocabulary and technical parameters when professional terms are used more than the average level of ordinary users; Organized review features: calculate semantic similarity and sentence structure repetition between comments through content homogenization detection, identify similar content generated in batches using text fingerprint algorithm, identify the overuse of extreme adjectives and absolute expressions through exaggerated expression detection, establish a dictionary of exaggerated modifiers to detect abnormal expression patterns, and analyze the time pattern and frequency anomaly of comment publishing through batch publishing mode recognition; Malicious review features: identify excessive negative emotional expression through extreme negative emotion detection, establish a negative emotion intensity scoring system to quantify the degree of extreme emotion, detect personal attacks and malicious defamation content through aggressive language recognition, establish an aggressive vocabulary dictionary containing abusive and threatening language, and analyze the usage frequency of malicious slander expressions to identify implicit malicious attack intentions.
4. The multi-modal adaptive e-commerce review truth and rating trustworthiness evaluation system, according to claim 1, wherein, Image authenticity verification module includes: ResNet visual feature extraction unit: adopts ResNet-50 deep convolutional neural network architecture to extract five types of visual features from the input image, including color histogram, texture, shape, edge and local binary pattern, capture image detail information at different resolution levels through multi-scale feature pyramid network, and output high-dimensional feature vector representation; The PS trace detection unit: the Canny edge detector is used to extract the edge information of the image, the Hough transform is used to detect the continuity of straight lines and curves, the edge gradient direction is calculated to identify the splicing boundary, and the edge discontinuity threshold is set to 0.15 or more to determine whether there is a splicing trace; the noise pattern analysis algorithm is used to decompose the image into multiple frequency subbands through wavelet transform, and the statistical method is used to analyze the distribution characteristics of noise in each frequency domain, and the inconsistency of noise pattern and artificial processing trace are detected, and when the noise variance difference is more than 0.2, it is determined that the PS processing exists; the compression trace detection algorithm is used to analyze the blocking effect and quantization error generated by JPEG compression, and the DCT coefficient statistical characteristics are used to detect repeated compression traces, and when the compression quality factor difference is more than 10%, it is determined that there is secondary editing processing; The stolen image detection unit: the perceptual hash matching technology is used to convert the image into a 64-bit binary fingerprint through the pHash algorithm, and the similarity of the fingerprints of two images is calculated through the Hamming distance, and when the Hamming distance is less than 5, it is determined that the images are similar; combined with the reverse image search technology, the SIFT feature point matching algorithm and the search engine interface are used to search for similar images in the Internet image database, and the image source is verified through the number of matched feature points and geometric consistency, and when the number of matched feature points is more than 30 and the geometric consistency score is greater than 0.8, it is determined that the image is stolen; The EXIF metadata analysis unit extracts the EXIF data in the image file, verifies the authenticity and consistency of the shooting device through the device fingerprint analysis algorithm, detects the rationality of the shooting time and the comment publishing time through the timestamp analysis algorithm, analyzes the matching degree of the GPS coordinates and the commodity sales area through the geographic location verification algorithm, and analyzes the technical rationality and environmental consistency of the shooting parameters through the shooting environment rationality evaluation algorithm, and when the parameter abnormality is more than 0.3, it is determined that the image is suspicious; The image credibility evaluation unit: a credibility evaluation system including originality, integrity, authenticity, timeliness and matching is constructed, a weighted average algorithm is used to calculate the comprehensive credibility score, and a normalized credibility score in the range of 0-1 is output.
5. The multi-modal adaptive e-commerce review truth and rating trustworthiness evaluation system, according to claim 1, wherein, The user behavior deep modeling module includes: The user historical behavior data acquisition unit constructs a user portrait database, acquires the complete comment trajectory data of the user in the recent period, establishes a 32-dimensional user behavior feature vector, including time feature dimension, frequency feature dimension, content feature dimension, interaction feature dimension, commodity feature dimension, score feature dimension, social feature dimension and abnormal feature dimension, and constructs a standardized feature vector through preprocessing; Review frequency anomaly detection unit: Adopt sliding time window-based anomaly detection algorithm, set four time window granularities, establish user normal review frequency baseline through Poisson distribution model, trigger abnormal early warning when review frequency exceeds 3 times standard deviation of baseline, use Isolation Forest algorithm to detect time series outliers, identify review intensive burst mode in short time through DBSCAN density clustering algorithm, combine chi-square test to calculate deviation degree of review publishing time from normal user behavior distribution; Review score distribution anomaly analysis unit: Adopt Beta distribution and normal distribution mixed model to fit user historical score pattern, evaluate score distribution anomaly degree through Kolmogorov-Smirnov test, establish expected score distribution matrix under five-star rating system, mark as abnormal when chi-square distance between user score distribution and expected distribution exceeds threshold, calculate score diversity through score entropy, analyze score transition probability through Markov chain model, identify score manipulation behavior; Content similarity calculation unit: Adopt Doc2Vec document vectorization to convert user historical reviews into semantic vectors, calculate semantic similarity between reviews through cosine similarity, set similarity threshold to 0.85, determine content repetition anomaly when similar reviews account for more than 30%, use edit distance algorithm and longest common subsequence algorithm to detect text string similarity, generate fingerprint features through SimHash algorithm, calculate review fingerprint similarity through Hamming distance, combine TF-IDF word frequency weight for multi-level content similarity comprehensive evaluation; Organized fake review behavior pattern recognition unit: Establish organized review manipulation behavior feature recognition system containing time feature class, content feature class, score feature class, social feature class, and commodity feature class, including 12 specific feature indicators, use random forest classifier to comprehensively evaluate organized review manipulation behavior; Time series behavior analysis unit: Adopt ARIMA autoregressive integrated moving average model to model user review time series, determine difference order through ADF stationarity test, establish ARIMA(p,d,q) model to predict user future review behavior pattern, use LSTM long short-term memory network to capture long-term dependence of user behavior, use Prophet time series prediction algorithm to analyze trend, periodicity and seasonality characteristics of user behavior, use change point detection algorithm to identify sudden changes in user behavior pattern; Social network graph structure analysis unit: Construct multi-layer social network graph based on user follow relationship, interaction behavior and common evaluation of commodities, establish complex network model containing node attributes and edge attributes, calculate user authority score through PageRank algorithm, identify abnormal aggregation mode through Community Detection algorithm, divide communities through Louvain algorithm, evaluate user network influence through centrality analysis, calculate user node clustering coefficient to evaluate social relationship authenticity; User credibility comprehensive evaluation unit: a user credibility evaluation model based on AHP hierarchical analysis method is established, a hierarchical structure containing behavior consistency, content authenticity, time rationality, social health degree, score reliability, and historical credit degree is constructed, a fuzzy comprehensive evaluation method is used to process the quantitative problem of qualitative indicators, a five-level credibility scoring system is established, and a Bayesian network model is used to learn the probability dependence relationship between each feature; The multi-layer perception coding dimension reduction unit adopts a deep neural network architecture including an input layer, a hidden layer, and an output layer, uses a ReLU function as the activation function, accelerates convergence through Batch Normalization batch normalization technology, uses an Adam optimizer for gradient descent, sets a Dropout rate to prevent overfitting, uses a PCA algorithm for dimension reduction preprocessing of high-dimensional features, and finally encodes 32-dimensional original features into 8-dimensional latent feature representations, including user activity factor, content quality factor, time regularity factor, social influence factor, score stability factor, abnormal behavior factor, credibility factor, and risk warning factor.
6. The multi-modal adaptive e-commerce review truth and rating trustworthiness evaluation system, according to claim 1, wherein, The adaptive weight distribution module includes: The multi-modal feature quality evaluation unit calculates the quality score of each modality, representing the degree of authenticity and credibility, and normalizes it to the 0-1 interval; The multi-head attention cross-modal correlation calculation unit adopts a 12-head parallel attention mechanism, each attention head is set to 64 dimensions, the text features, image features, and user behavior features are respectively mapped to query vectors, key vectors, and value vectors through linear transformation, the attention weight is calculated through the dot product of the query vector and the key vector, and after scaling and normalization processing, the output feature is obtained by multiplying the value vector; the cross-modal correlation is quantified by calculating the mutual information between different modal features, the kernel density estimation method is used to calculate the information sharing degree between continuous variables, and the statistical dependence relationship between modalities is measured; a 3x3 cross-modal attention matrix is constructed to represent the correlation strength between the three pairs of modalities: text and image, text and behavior, and image and behavior, the interaction features between modalities are calculated through bilinear pooling operation, and the non-linear correlation between modalities is modeled using a learnable weight matrix; The dynamic weight normalization unit introduces a learnable temperature parameter, which is adaptively adjusted through backpropagation algorithm, the temperature parameter controls the smoothness of the exponential normalization function, and the adaptive control of weight distribution is realized by dynamically adjusting the temperature value; the quality score, cross-modal correlation score, and attention weight are used to adjust the weight of each modality through a gating mechanism, the gating unit controls the information flow through learning parameters and bias terms, and finally the fused weight is normalized through the product of the temperature-adjusted weight and the gating output; The gradient descent weight optimization unit adopts an adaptive matrix estimation optimization algorithm, defines a composite loss function including classification loss, regularization loss, and consistency loss, and updates the parameters by calculating the gradient of the loss function with respect to the parameters during the optimization process.
7. The multi-modal adaptive e-commerce review truth and rating trustworthiness evaluation system, according to claim 1, wherein, The double-dimensional prediction module includes: The score tendency classification unit adopts a multi-layer perceptron architecture based on an attention mechanism, takes the fused multi-modal feature vector as input, performs non-linear transformation through 3 fully connected layers, constructs a 5-classification system for score prediction task, and uses the Softmax normalization function in the output layer to map the features to the probability distribution of each score level. The classification unit identifies the score tendency by learning the key semantic features in the review text, combines the image quality score and the user's historical scoring behavior pattern, uses a weighted voting mechanism to integrate multi-modal information to predict the final score level, and calibrates the confidence of the output probability distribution through temperature scaling technology; The credibility evaluation unit is constructed based on a deep neural network regression model, adopts a 4-layer deep network structure with residual connection, each layer contains batch normalization and Dropout regularization technology, the input layer receives a comprehensive feature vector containing text authenticity features, image authenticity features, and user behavior credibility features, learns a continuous numerical representation of review authenticity through non-linear mapping, and the output layer uses the Sigmoid activation function to constrain the prediction result in the interval of 0 to 1 to represent the credibility score, where 0 represents completely untrustworthy and 1 represents completely trustworthy. The evaluation unit constructs a six-dimensional feature system for authenticity evaluation and simultaneously outputs the probability distribution of four types of comment sources and the comprehensive credibility evaluation score through a multi-task learning framework; The weighted feature fusion unit designs an adaptive weight learning mechanism to integrate the feature representations of text, image, and user behavior, calculates the importance weight of each modal feature through a gating fusion unit, and the gating mechanism adopts an attention-based weight allocation strategy, the calculation formula is attention weight equals to the dot product of feature vector and learnable parameter matrix after Tanh activation and weight vector, the fusion unit outputs a 1024-dimensional feature vector containing global semantic information, local detail features, and cross-modal correlation features, uses residual connection to preserve original feature information, and stabilizes the training process through layer normalization technology; The joint loss function design unit adopts cross-entropy loss function for score classification task to measure the difference between predicted probability distribution and true label distribution, the cross-entropy loss calculation method is to take the negative logarithm of the predicted probability at the position corresponding to the true label, and mean square error loss function is used for credibility regression task to calculate the square difference between predicted value and true credibility score, the overall loss function adopts weighted linear combination form, and the weight coefficient is dynamically adjusted through validation set performance. The multi-task learning unit adopts a hard parameter sharing architecture, a bottom layer feature extraction network is a shared layer, and a top layer respectively constructs a score classification and a credibility evaluation two task-specific output heads, the shared layer includes 4 layers of convolutional neural networks and 2 layers of recurrent neural networks, the task-specific layer is optimized for the classification and regression tasks respectively, the learning progress of the two tasks is balanced through an alternating training strategy, a task weight adaptive adjustment mechanism is set, the learning unit outputs a 5-dimensional probability vector representing a score tendency distribution and a 1-dimensional continuous value representing a credibility score, and through a post-processing module, the probability distribution is converted into an expected score value, combined with the credibility score, a confidence quantification is provided for the final prediction result, and a two-dimensional accurate prediction of the review authenticity and the score tendency is realized.
Citation Information
Patent Citations
Method for filtering cosmetic internet false comments based on semantic analysis
CN115687774A
Knowledge-enhanced user multi-modal online comment quality evaluation method and system
CN116881689A
User comment authenticity verification method and device and medium
CN120894037A
A computer-implemented method and system for analyzing and evaluating user reviews
WO2017051425A1
Cited By
Knowledge graph-based agricultural product online sales perception quality analysis method and system
CN122115083A
Multimodal sarcasm recognition method, system, storage medium and electronic device
CN122471018B