Mobile e-commerce false comment detection method and system based on large language model
By extracting and transforming fake comment texts with specific features, combined with the fine-tuning and distillation technology of large language models, the problem of false comment recognition in mobile e-commerce platforms is solved, and high accuracy and efficient false comment detection is achieved.
Patent Information
- Application Number
- CN202510314769.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
It is difficult to accurately identify complex and flexible false comments in mobile e-commerce platforms, especially when facing comments with deep meanings or emotional tendencies, and it is difficult to make accurate judgments, and at the same time lack customized learning capabilities for specific fields.
Generate false comment text by designing prompt words, build a training data set, extract numerical features including emotional polarity features, text readability features, part-of-speech distribution features, text theme distribution features, type-mark ratio features, and type-mark ratio features, and convert them into text classification descriptions based on feature rules, build a rule data set, fine-tune training on the large language model, obtain the first largest language model, and distillate it to the second largest language model with smaller model parameters for false comment detection.
It realizes accurate identification of false comments in the mobile e-commerce environment, improves the accuracy of identification of false comments, is efficient and scalable, and can capture the unique language styles and expression methods in the mobile e-commerce field.
Smart Images

Figure SMS_11 
Figure SMS_13 
Figure FDA0005315640130000011
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and data analysis, and specifically to a method and system for detecting fake reviews in mobile e-commerce based on large language models. Background Art
[0002] Consumers' evaluations of products or services on different e-commerce platforms have become an important factor affecting purchase intention. The pursuit of good reviews by merchants, and the use of fake reviews by some unscrupulous merchants or competitors to hype or attack products, have made the problem of fake reviews on online platforms increasingly serious.
[0003] A large number of studies and technical means have been invested in the detection of fake reviews. Technically, it mainly relies on natural language processing and machine learning technologies, and can be identified by combining manually designed features with classification algorithms. With the rapid development of deep learning technologies, applying large language models to mobile e-commerce platforms, by capturing the deep semantic features of texts to establish text and associated entity relationships, it is possible to identify deeply forged fake content, providing the possibility for fine-grained and semantic mobile e-commerce review analysis.
[0004] However, existing methods often have the following limitations: First, existing detection methods mainly screen reviews through keyword matching or predefined rule libraries. For example, reviews containing obvious promotional words or frequently repeated derogatory terms are marked. However, as the forms of fake reviews become more and more flexible and complex, this method is difficult to cover rich and diverse review expressions, and is even less able to handle implicit terms or deliberately changed text patterns that may appear in reviews; Second, there is insufficient understanding of complex semantics. When faced with reviews with deep meanings or emotional tendencies, it is difficult to accurately judge the authenticity of the reviews; Third, there is a lack of customized learning ability for specific domains. Currently, many detection methods mainly rely on models trained on large-scale general datasets, and when these models face reviews in specific domains: e-commerce, catering, tourism, they often have difficulty capturing the unique language styles, word usage habits, and expression methods of these domains. General models may miss these key domain features during generalization, thus affecting the accuracy of distinguishing fake reviews from real feedback. During actual detection, it may show a relatively high false positive rate, affecting the authenticity of review discrimination.
[0005] Therefore, there is an urgent need to invent a large language model detection method specifically for the mobile e-commerce environment that can accurately identify fake reviews and has both high efficiency and scalability. Summary of the Invention
[0006] To solve the above problems raised in the background art, the present invention provides a method and system for detecting fake reviews in mobile e-commerce based on large language models.
[0007] The technical solution of the present invention is as follows:
[0008] A method for detecting fake reviews of mobile e-commerce based on a large language model comprises the following steps:
[0009] S1. Obtain the real review text of the user on the product, generate false review text based on the designed prompt words, integrate the real review text and false review text, and construct a training data set;
[0010] S2. Clean and deduplicate the data in the training data set in turn to obtain standard data;
[0011] S3, based on the semantic paragraphs, the standard data is divided into sentences and words in turn, and a number of sentences and words are obtained accordingly, and numerical features are extracted based on the sentences and words. The numerical features include: sentiment polarity features, text readability features, part of speech distribution features, text topic distribution features, and type-token ratio features;
[0012] S4. Based on the set feature rules, the numerical features are described in text classification according to the feature value ranges in which they are located, a number of feature texts are obtained, and the several feature texts are combined to obtain a rule set;
[0013] Fuse the training data set and the rule set to construct a rule data set;
[0014] S5. Fine-tune the large language model using the rule data set to obtain the first large language model;
[0015] S6. Distill the first largest language model into a second largest language model with smaller model parameters, and perform false review detection on new product review texts.
[0016] Specifically, the extraction of the text readability feature in S3 is to consider complex vocabulary and professional terms on the basis of the total number of words, the total number of sentences, the total number of syllables, and the total number of words to obtain a readability score, and the formula is as follows:
[0017]
[0018] Among them, C complex is the penalty coefficient for complex words, C technical is the penalty coefficient of professional terms,
[0019] The design of the prompt words in S1 includes role prompt words, task prompt words, and prompt word feature constraints.
[0020] Furthermore, the role prompt is: You are an online shopping review generation assistant; the task prompt is: Please generate several reviews based on product information, where the product information includes: product category, product name, price, and product features;
[0021] The prompt feature constraints are: The generated reviews need to meet the following characteristics: The review length is within a specific range; the review contains exaggerated or untrue usage experiences; it simulates the writing habits of real users and includes colloquial expressions; the review contains a scoring tendency and emotional intensity.
[0022] The false review text is generated based on the design prompt. Each false review text needs to include the following elements: purchase experience, usage feeling, and specific detail description, while paying attention to the change in review length; each false review needs to consider simultaneously: word diversity, emotional expression intensity, specific degree of detail description, and authenticity degree.
[0023] During the distillation process of S6:
[0024] Take the first probability distribution of the real review text or false review text output by the first large language model as the soft label;
[0025] Use the soft label as the target label for training the second large language model, and at the same time use the training dataset to train the second large language model to output the second probability distribution;
[0026] Calculate the KL divergence loss based on the first probability distribution and the second probability distribution, calculate the cross-entropy loss based on the second probability distribution and the target label, and update the model parameters of the second large language model based on the KL divergence loss and the cross-entropy loss.
[0027] The training dataset D constructed in S1 text is expressed as follows:
[0028] D text = real review text S 1 {text content published by users, true / false label, publication time} + false review text S 2 {generated text content, true / false label, publication time}.
[0029] In S5, the rule dataset is used to fine-tune and train the large language model. To inject a low-rank adapter into the large language model, the rank parameter of the low-rank adapter is updated according to the gradient value of the loss function. The loss function used is: where y i is the output corresponding to the i-th text sample, P is the probability function, M fine-tuned is the first large language model after fine-tuning training, and N is the batch size.
[0030] The data cleaning in S2 is as follows: filtering out emoji and garbled characters in the data, further unifying full-width or half-width characters in the data, performing simplified and traditional Chinese character conversion in sequence, standardizing the case, and at the same time removing comments with less than 10 characters.
[0031] The present invention also provides a false review detection system for mobile e-commerce based on a large language model, including:
[0032] A data acquisition and generation module: used to obtain the real review text of users for goods, generate false review text based on the designed prompts, integrate the real review text and the false review text, and construct a training data set;
[0033] A data processing module: used to sequentially perform data cleaning and data deduplication on the data in the training data set to obtain standard data;
[0034] A feature extraction module: used to sequentially perform sentence splitting and word segmentation on the standard data based on semantic paragraphs, obtain a number of sentences and words respectively, and extract numerical features based on the number of sentences and words. The numerical features include: sentiment polarity feature, text readability feature, part-of-speech distribution feature, text topic distribution feature, type-token ratio feature;
[0035] A feature conversion module: used to perform text classification description on the numerical features according to their corresponding feature value ranges based on the set feature rules, obtain a number of feature texts, combine the number of feature texts to obtain a rule set; fuse the training data set and the rule set to construct a rule data set;
[0036] A model fine-tuning module: used to fine-tune and train the large language model using the rule data set to obtain a first large language model;
[0037] A false review detection module: used to distill the first large language model into a second large language model with smaller model parameters to detect false reviews for new product reviews.
[0038] The beneficial effects of the present invention are as follows:
[0039] 1. The present invention generates false review text by designing prompts, constructs a training data set, extracts five types of numerical features including sentiment polarity feature, text readability feature, part-of-speech distribution feature, text topic distribution feature, and type-token ratio feature, and then performs text classification description on the numerical features according to their corresponding feature value ranges based on the set feature rules, converts the numerical features into text natural language features, obtains a rule set, and further constructs a rule data set, which can capture the unique language style, word usage habits, and expression methods in the field of mobile e-commerce, better distinguish real feedback and false reviews, and improve the accuracy of false review identification.
[0040] 2. Fine-tune and train a large language model using a regular dataset to obtain a first large language model, and distill the first large language model into a second large language model with smaller model parameters, which can still maintain good performance without using numerical features, ensuring efficient detection in the mobile e-commerce environment. Detailed implementation manners
[0041] The exemplary implementation manners of the present disclosure will be described in more detail below.
[0042] Embodiment
[0043] This embodiment provides a method for detecting fake reviews in mobile e-commerce based on a large language model, including the following steps:
[0044] S1. Obtain the real review text of the product by the user, generate fake review text based on the designed prompt words, and integrate the real review text and the fake review text to construct a training dataset.
[0045] In this embodiment, the real review text comes from the public review data of the actual e-commerce platform. On the premise of following relevant laws and platform rules, obtain the review text of various products by the user. Use the BeautifulSoup library in Python to obtain the user reviews on the product detail page.
[0046] Furthermore, generate fake review text based on the designed prompt words. The design of the prompt words should meet the requirements of generating a large number of diverse online shopping review data and be as close as possible to the characteristics of artificial writing in terms of content.
[0047] The design of the prompt words includes role prompt words, task prompt words, and prompt word feature constraints. Among them, the role prompt words are: You are an online shopping review generation assistant; the task prompt words are: Please generate several reviews based on the product information, and the product information includes: product category, product name, price, product features.
[0048] The prompt word feature constraints are: the generated reviews need to meet the following characteristics: (1) the review length is within a specific range, for example, set the review length between 20 - 100 words; (2) the review contains exaggerated or untrue usage experiences; (3) simulate the writing habits of real users, including colloquial expressions; (4) the review contains a scoring tendency and emotional intensity.
[0049] At the same time, each fake review text needs to include the following elements: purchase experience, usage feeling, specific detail description, and pay attention to the change of review length; each fake review needs to consider at the same time: word diversity, emotional expression intensity, specific degree of detail description, authenticity degree.
[0050] Taking digital products as an example, the prompt words can be designed as:
[0051] Role prompt: You are a professional digital product review generation assistant.
[0052] Task prompt: Please generate 10 fake reviews based on the following product information. Product information includes: Product category: wireless headphones; Product name: True Wireless Noise Cancelling Headphones Pro; Price: 1,299 yuan; Product features: active noise reduction, 40 hours of battery life, Bluetooth 5.3, Hi-Res certification.
[0053] Prompt word feature constraints: (1) Review length: 30-80 words; (2) Reviews contain fictitious and unreasonable usage scenarios, exaggerate product performance, and add unrealistic comparative experiences; (3) Simulate the writing habits of real users and include colloquial expressions; (4) Reviews contain rating tendencies and emotional intensity, with 80% positive reviews and 20% negative reviews.
[0054] The review must include: specific usage scenarios, technical parameters, subjective experience description, and comparison with other brands. Time range: January 1, 2023 to present.
[0055] Furthermore, the acquired real review text and the generated fake review text are fused to construct a training dataset D text , which is expressed as follows:
[0056] D text =Real review text S 1 {Text content posted by the user, true or false label, release time} + false comment text S 2 {Generated text content, true or false labels, release time}.
[0057] S2. Clean and deduplicate the data in the training data set in turn to obtain standard data.
[0058] The main purpose of step S2 is to clean and deduplicate the data in the training data set and filter out valid data. The data cleaning process includes: filtering emoticons and garbled characters in the data, unifying the full-width or half-width characters in the data, converting simplified and traditional Chinese characters, normalizing uppercase and lowercase characters, and removing comments with less than 10 characters.
[0059] The process of data deduplication is as follows: convert the cleaned data into text vectors, calculate the similarity between text vectors based on cosine similarity, set the similarity value greater than or equal to the set similarity threshold as a duplicate comment, and only retain the comment with the earlier time. In addition, the comments that obviously do not contain any substantive evaluation content are marked as "invalid comments" and removed. Finally, the standard data is obtained.
[0060] S3. The standard data is divided into sentences and words based on semantic paragraphs, and a number of sentences and words are obtained accordingly. Numerical features are extracted based on the sentences and words. The numerical features include: sentiment polarity features, text readability features, part-of-speech distribution features, text topic distribution features, and type-token ratio features.
[0061] In step S3, the standard data is segmented based on semantic paragraphs, and the long text is split into multiple sentences, so as to facilitate the subsequent statistics of the part of speech of each sentence and the subsequent statistics of the missing rate of verbs or other parts of speech in each sentence.
[0062] Then, each sentence is segmented to obtain a continuous word sequence. Based on the segmentation, each word obtained after the segmentation is assigned a corresponding part-of-speech tag. For example, for the sentence "This mobile phone has good performance", after the segmentation, "this", "model", "mobile phone", "performance", "good" will be obtained, and then each word is assigned a corresponding part-of-speech tag, such as DT (qualifier), Q (quantifier), N (noun), V (verb), V (adjective), and finally: this / DT, model / Q, mobile phone / N, performance / N, good / V.
[0063] The main purpose of step S3 is to extract numerical features based on a number of sentences and words. The numerical features of the present invention include: sentiment polarity features, text readability features, part-of-speech distribution features, text topic distribution features, and type-token ratio features. The following introduces each numerical feature one by one:
[0064] For sentiment polarity features, we mainly use the existing large-scale Chinese pre-trained RoBERTa model to extract the sentiment polarity features in the words after word segmentation, and map them to the [0,1] interval through the sigmoid function to obtain the sentiment polarity score.
[0065] Regarding the text readability feature, the traditional text readability is mainly evaluated based on the sentence length and the average number of words in each sentence. On the basis of the traditional text readability evaluation, the present invention adds the processing of complex vocabulary and professional terms to more accurately evaluate the readability of Chinese text and obtain the readability score. The formula is as follows:
[0066]
[0067] Among them, C complex is the penalty coefficient for complex words, C technical is the penalty coefficient of professional terms,
[0068] Among them, complex words are defined as words containing ≥3 characters and uncommon combinations. These complex words have a negative impact on the readability of the text. Therefore, a penalty coefficient for complex words is added during calculation. The penalty coefficient of complex words represents the proportion of complex words in the total number of words. The higher the proportion, the worse the readability of the text and the greater the penalty coefficient.
[0069] Considering the influence of professional terms, a professional term library for mobile e-commerce is defined, which is specifically used to identify and mark terms related to the field. If a word outside the professional term library for mobile e-commerce appears in the comment, it will be considered an external word, and its impact on the readability of the text is calculated through the penalty coefficient of professional terms. Texts with a higher proportion of external words usually contain too many professional terms, making it difficult for ordinary readers to understand. According to the calculated readability score, if the score is less than 30, the text is determined to have low readability. The setting of 30 takes into account the reading comprehension ability of users. If the text is difficult to understand, its readability is poor.
[0070] Regarding the part-of-speech distribution characteristics, in this embodiment, mainly calculate the noun density, adjective concentration, lack rate of verb-object structure, text theme distribution characteristics, type-token ratio index, etc. The specific expressions are as follows: Part-of-speech distribution characteristics
[0071] (1) The noun density is the proportion of nouns in a sentence, which is obtained by calculating the number of nouns in the sentence divided by the total number of words. The calculation formula is:
[0072] The noun density reflects the usage frequency of nouns in a sentence. Generally speaking, texts with a higher noun density may be more descriptive.
[0073] (2) The adjective concentration represents the concentration of adjectives in a sentence, which measures the frequency of adjectives appearing in the sentence. The calculation formula is:
[0074] The adjective concentration reflects the evaluative or descriptive intensity of the text. Sentences with a high adjective concentration often contain more subjective colors and emotional expressions.
[0075] (3) Lack rate of verb-object structure: The verb-object structure refers to a phrase composed of a verb and a noun. Whether a sentence makes sense is closely related to the integrity of the verb-object structure. Texts with a higher lack rate of verb-object structure may have incomplete grammar or unclear expressions.
[0076] The calculation formula is:
[0077] Regarding the text theme distribution characteristics, a theme probability distribution vector is generated for each comment. The LDA model will generate a theme probability distribution vector for each comment, and the formula is expressed as follows:
[0078]
[0079] Among them, p i represents the probability that the comment d belongs to the i-th topic, and K represents the number of topics.
[0080] Then calculate the topic dispersion entropy value H:
[0081] To make the result in the interval [0, 1], define the normalized dispersion entropy value H norm as:
[0082]
[0083] Among them, H norm being 0 means it is completely concentrated on a single topic, and H norm being 1 means the topic distribution is completely uniform.
[0084] When judging, if the value is a low dispersion entropy value (H norm <0.5), it indicates that the comment is on a small number of topics and tends to be judged as a false comment feature.
[0085] For the type-token ratio index (TTR), it is an index used to measure the lexical diversity in a text. It reflects the usage of different words in a text and can help analyze the richness of the language. False comments have limited lexical richness and often have a low TTR value.
[0086] The basic formula of TTR is:
[0087] Among them: V is the number of different words in the text, and N is the total number of all words in the text.
[0088] S4. Based on the set feature rules, classify and describe the numerical features according to their corresponding eigenvalue ranges to obtain several feature texts, and combine the several feature texts to obtain a rule set;
[0089] Fuse the training data set and the rule set to construct a rule data set.
[0090] In step S4, through the set feature rules, convert the numerical features into more interpretable text classification descriptions to obtain several feature texts.
[0091] For the numerical features: sentiment polarity feature, text readability feature, part-of-speech distribution feature, text topic distribution feature, type-token ratio feature, the set feature rules are respectively:
[0092] (1) Emotional polarity feature: The range of emotional polarity scores is [0.9–1.0]: "Extremely strong emotional expression"; [0.7–0.89]: "Strong emotional tendency"; [0.4–0.69] → "Moderate emotional intensity"; [0.0–0.39] → "Weak or contradictory emotional expression".
[0093] (2) Text readability feature: Readability score < 30: "Low readability"; Readability score range [30–60]: "Moderate readability"; Readability score > 60: "High readability".
[0094] (3) Text theme distribution feature: Dispersion entropy value < 0.3: "Highly concentrated on a single theme"; Dispersion entropy value range [0.3–0.5]: "Moderate theme concentration"; Dispersion entropy value > 0.5: "Themes are dispersed".
[0095] (4) Type-Token Ratio feature (TTR): TTR < 0.4: "High lexical repetition rate"; TTR value range 0.4–0.6: "Moderate lexical diversity"; TTR > 0.6: "Lexically rich".
[0096] (5) Part-of-speech distribution feature: Adjective proportion > 40%: "Exaggerated description"; Adjective proportion is 20%–40%: "Regular description"; Adjective proportion < 20%: "Lack of subjective evaluation".
[0097] For example, for the review text "This mobile phone is amazing. The photos are super clear and the battery life is extremely good!", after the numerical feature extraction in step S3, the obtained numerical features are as follows:
[0098] Emotional polarity feature: 0.95, Text readability feature: 93.7, Part-of-speech distribution feature: Adjective proportion 40%, Text theme distribution feature: Dispersion entropy value 0.3, Type-Token Ratio feature: 0.6.
[0099] In this example, too much numerical feature data is not conducive to the subsequent large language model understanding the internal meaning of the numerical values. Therefore, the numerical features are described in natural language according to their corresponding eigenvalue ranges, and we get:
[0100] Emotional polarity feature: This review shows an extremely strong positive emotion; Expression feature: The text is clearly expressed and easy to understand, the content is highly concentrated on the product's advantages, and the language expression is natural and fluent; Language pattern: The frequency of using adjectives is significantly higher than that of ordinary reviews, and the density of praise words is relatively large.
[0101] Then, several feature texts are combined to obtain a rule set. The training data set and the rule set are fused to construct a rule data set, enabling the large language model to receive both text and feature information during processing.
[0102] S5. Fine-tune and train the large language model using the rule dataset to obtain the first large language model.
[0103] The function of step S5 is to fine-tune the large language model so that it can adapt to data in a specific domain. The large language model used in this embodiment is Qwen2.5-72B. By injecting a low-rank adapter into the large language model, fine-tune and train the large language model using the rule dataset, and update the rank parameters of the low-rank adapter according to the gradient value of the loss function until the preset number of iterations is reached, and the fine-tuning training is completed. Define the forward propagation as y = M Qwen2.5 (x) + A T ·B T ·x, A ∈ R d×r , B ∈ R r×d , where M Qwen2.5 represents the large language model used, x represents the vector input to the large language model, A and B both represent matrices of specific dimensions, T represents the transpose of the matrix, d represents the dimension of the hidden layer of the large language model, r represents the rank of the low-rank adapter, and r < d to ensure the low-rank property and reduce computational complexity. In the present invention, the value of r is specified in the range of 400 - 1200.
[0104] During the fine-tuning training process, define the fine-tuning task as identifying false comments in the training dataset, and update the rank parameters of the low-rank adapter according to the gradient value of the loss function. The loss function used is: where, y i is the output corresponding to the i-th text sample, P is the probability function, M fine-tuned is the first large language model after fine-tuning training, and N is the batch size. In each iteration cycle, continuously update the rank parameters of the low-rank adapter until the preset number of iterations is reached, and the fine-tuning training is completed to obtain the first large language model.
[0105] At the same time, use the pre-prepared validation data to evaluate the large language model, and calculate the accuracy, precision, recall rate, and F1 score. The accuracy is used to measure the proportion of all comments that are correctly classified. The precision is the proportion of actual false comments in the set predicted as false comments. The recall rate is the proportion of actual false comments that are correctly predicted. The F1 score is used to comprehensively measure the detection ability and accuracy of the large language model for false comments to evaluate the current training progress.
[0106] S6. Distill the first large language model into a second large language model with smaller model parameters to detect false comments in new product review texts.
[0107] In this embodiment, the first large language model is distilled into a second large language model with smaller model parameters. Qwen2.5-1.5B with smaller parameters is used as the second large language model, which is suitable for various small business scenarios and has lower resource utilization. During the distillation process, the second large language model can learn the knowledge of the first large language model and accept only text comments as input, making as accurate a judgment as possible without using features.
[0108] During the distillation process: The first probability distribution of the true or false review text output by the first large language model is used as the soft label; the soft label is used as the target label for training the second large language model, and at the same time, the training data set is used to train the second large language model to output the second probability distribution; the KL divergence loss is calculated based on the first probability distribution and the second probability distribution, the cross-entropy loss is calculated based on the second probability distribution and the target label, and the model parameters of the second large language model are updated based on the KL divergence loss and the cross-entropy loss. The formula is as follows:
[0109] Loss = α·CrossEntropyLoss(P student , Y true ) + (1 - α)·KLDivLoss(P student , P teacher ),
[0110] where Loss is the loss function when updating the model parameters of the second large language model, α is a weight coefficient, P student is the second probability distribution, Y true is the target label, P teacher is the first probability distribution, CrossEntropyLoss is the cross-entropy loss, and KLDivLoss is the KL divergence loss.
[0111] Finally, the second large language model is actually deployed for the mobile e-commerce environment to detect false reviews of new product review texts and obtain the detection results of whether the reviews are true or false.
[0112] In this embodiment, the method for detecting false reviews of mobile e-commerce based on large language models used in the present invention is also compared with three other different methods, namely SVM, BERT, and Few-shot. The comparison results are shown in Table 1.
[0113] Table 1 Detection Results of False Reviews in Mobile E-commerce
[0114] Method Accuracy Precision Recall F1 Score SVM 71.3% 74.6% 70.2% 72.3% BERT 82.5% 80.3% 76.6% 78.4% Few-shot 58.7% 57.1% 53.8% 55.4% The present invention 87.1% 85.8% 83.4% 84.6%
[0115] As can be seen from Table 1, the accuracy of the traditional machine learning method Few-shot is the lowest, only 58.7%, and the performance of SVM is also poor, with an accuracy of 71.3%. This indicates that there are still significant limitations in the detection methods without fine-tuning training when identifying false reviews. The performance of BERT is better than that of SVM and Few-shot, reaching an accuracy of 82.5%. Through data processing, feature extraction, feature transformation, and model fine-tuning, the accuracy of the present invention reaches 87.1% when detecting false reviews. At the same time, it is superior to other methods in various evaluation indicators such as precision, recall, and F1 score. Therefore, the present invention is better able to perform the task of detecting false e-commerce reviews.
[0116] The present invention also provides a mobile e-commerce false review detection system based on a large language model, including:
[0117] A data acquisition and generation module: used to obtain the real review text of users for goods, generate false review text based on the designed prompt words, integrate the real review text and the false review text, and construct a training data set;
[0118] A data processing module: used to sequentially perform data cleaning and data deduplication on the data in the training data set to obtain standard data;
[0119] A feature extraction module: used to sequentially perform sentence splitting and word segmentation on the standard data based on semantic paragraphs, respectively obtain a number of sentences and words, and extract numerical features based on the number of sentences and words. The numerical features include: sentiment polarity features, text readability features, part-of-speech distribution features, text topic distribution features, type-token ratio features;
[0120] A feature transformation module: used to perform text classification descriptions on the numerical features according to their corresponding feature value ranges based on the set feature rules, obtain a number of feature texts, combine the number of feature texts to obtain a rule set; fuse the training data set and the rule set to construct a rule data set;
[0121] A model fine-tuning module: used to fine-tune and train the large language model using the rule data set to obtain a first large language model;
[0122] A false review detection module: used to distill the first large language model into a second large language model with smaller model parameters to detect false reviews of new product reviews.
Claims
1. A method for detecting fake reviews in mobile e-commerce based on a large language model, characterized in that: The following steps are involved: S1. Obtain the real review text of the user on the product, generate false review text based on the designed prompt words, integrate the real review text and false review text, and construct a training data set; S2. Clean and deduplicate the data in the training data set in turn to obtain standard data; S3, based on the semantic paragraphs, the standard data is divided into sentences and words in turn, and a number of sentences and words are obtained accordingly, and numerical features are extracted based on the sentences and words. The numerical features include: sentiment polarity features, text readability features, part of speech distribution features, text topic distribution features, and type-token ratio features; S4. Based on the set feature rules, the numerical features are described in text classification according to the feature value ranges in which they are located, a number of feature texts are obtained, and the several feature texts are combined to obtain a rule set; Fuse the training data set and the rule set to construct a rule data set; S5. Fine-tune the large language model using the rule data set to obtain the first large language model; S6. Distill the first largest language model into a second largest language model with smaller model parameters, and perform false review detection on new product review texts.
2. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The extraction of the text readability feature in S3 is to consider complex vocabulary and professional terms on the basis of the total number of words, the total number of sentences, the total number of syllables, and the total number of words to obtain a readability score, and the formula is as follows: Among them, C complex is the penalty coefficient for complex words, C technical is the penalty coefficient of professional terms, 3. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The design of the prompt words in S1 includes role prompt words, task prompt words, and prompt word feature constraints.
4. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 3 is characterized in that: The role prompt is: You are an online shopping review generation assistant; The task prompt is: Please generate several comments based on product information. Product information includes: product category, product name, price, and product features; The feature constraints of the prompt word are: the generated comments need to meet the following features: The length of the review is within a specific range; the review contains exaggerated or unrealistic user experiences; the review simulates the writing habits of real users and contains colloquial expressions; the review contains rating tendencies and emotional intensity.
5. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 3 is characterized in that: The fake review text is generated based on the designed prompt words, and each fake review text must contain the following elements: purchase experience, usage experience, specific details description, and pay attention to the change in the length of the review; each fake review needs to consider at the same time: word diversity, emotional expression intensity, specific degree of detail description, and degree of authenticity.
6. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The S6 during distillation: The first probability distribution of the real review text or the fake review text output by the first language model is used as a soft label; The soft label is used as the target label for training the second largest language model, and the training data set is used to train the second largest language model to output a second probability distribution; The KL divergence loss is calculated based on the first probability distribution and the second probability distribution, the cross entropy loss is calculated based on the second probability distribution and the target label, and the model parameters of the second largest language model are updated based on the KL divergence loss and the cross entropy loss.
7. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The training dataset D constructed in S1 text It is expressed as follows: D text =Real comment text S1{text content posted by the user, true or false label, release time}+false comment text S2{generated text content, true or false label, release time}.
8. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The S5 uses a regular data set to fine-tune the large language model. In order to inject a low-rank adapter into the large language model, the rank parameter of the low-rank adapter is updated according to the gradient value of the loss function. The loss function used is: Among them, y i is the output corresponding to the i-th text sample, P is the probability function, M fine-tuned To fine-tune the first language model after training, N is the batch size.
9. The method for detecting false reviews of mobile e-commerce based on a large language model according to claim 1, characterized in that: The data cleaning in S2 is as follows: filtering out emoticons and garbled characters in the data, further unifying the full-width or half-width characters in the data, performing simplified and traditional Chinese conversions in sequence, standardizing uppercase and lowercase characters, and removing comments with less than 10 characters.
10. A mobile e-commerce fake review detection system based on a large language model, characterized in that: include: Data acquisition and generation module: used to obtain users' real comments on products, generate fake comments based on designed prompt words, integrate real and fake comments, and build training data sets; Data processing module: used to clean and deduplicate the data in the training data set in order to obtain standard data; Feature extraction module: used to perform sentence and word segmentation on the standard data based on semantic paragraphs, and obtain a number of sentences and words accordingly, and extract numerical features based on the sentences and words. The numerical features include: sentiment polarity features, text readability features, part-of-speech distribution features, text topic distribution features, and type-token ratio features; Feature conversion module: used to classify and describe the numerical features in text according to the feature value range based on the set feature rules, obtain several feature texts, combine several feature texts to obtain a rule set; fuse the training data set and the rule set to construct a rule data set; Model fine-tuning module: used to fine-tune the large language model using the rule data set to obtain the first large language model; Fake review detection module: used to distill the first largest language model into the second largest language model with smaller model parameters, and detect fake reviews for new product reviews.
Citation Information
Patent Citations
False comment detection method and system based on semi-supervised learning model
CN110580341A
False comment detection method and system based on feature fusion and screening, and medium
CN114492423A
Comment generation model training method and device and information generation method and device
CN117591948A
Method for detecting deceptive e-commerce reviews based on sentiment-topic joint probability
US20210027016A1
Cited By
Automatic detection and control method and system for false information and comments based on large model and fine tuning
CN120372064A
Multi-mode irony detection method based on large language model
CN120952006A