Method and device for processing commodity comment data, electronic equipment and storage medium
By generating a set of positive reviews and detecting the product status in the review images, the problem of covert order-brushing in e-commerce transactions is solved, and effective filtering and authenticity judgment of product review data is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MASHANG CONSUMER FINANCE CO LTD
- Filing Date
- 2023-02-03
- Publication Date
- 2026-07-28
AI Technical Summary
Existing technologies are insufficient to effectively identify covert commercial fraudulent activities in e-commerce transactions, especially those that lure users into creating fake positive reviews by promising cash back, making it difficult to guarantee the validity of product review data.
By acquiring product review data and determining product categories, a positive review set is generated. The product review data is then matched with the review images, and the product status in the images is checked to determine the validity of the reviews.
Effectively filter out invalid reviews, prevent fraudulent order practices, and improve the authenticity and accuracy of product review data.
Smart Images

Figure CN116150137B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device, and readable storage medium for processing product review data. Background Technology
[0002] When shopping online, some merchants may use manual or technical means to forge or fabricate orders and positive product reviews to lure consumers into purchasing goods that do not meet their quality or expectations. Alternatively, some user reviews may be subjective or inaccurate. For example, some users may take blurry photos or even photograph the wrong subject when writing product reviews. Therefore, to ensure the validity of product review data, it is necessary to analyze and test the product review data.
[0003] In related technologies, big data and graph theory are commonly used to statistically analyze transaction information, product browsing and order data, store customer service communication information, and product review data. This method can identify and process orders and reviews fabricated by illegal companies, organizations, and professionals. However, due to the wide variety of invalid reviews, this method cannot effectively identify all types of invalid reviews. For example, this method cannot effectively detect and handle covert commercial order-brushing activities involving a wide range of participants and relatively normal e-commerce transaction behavior, especially those that lure users into forging positive reviews by promising cash back. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and readable storage medium for processing product review data, which identifies and processes invalid reviews by detecting product review data.
[0005] Firstly, this application provides a method for processing product review data, including:
[0006] Obtain multiple product review data corresponding to the target product;
[0007] The product category of the target product and the set of positive reviews corresponding to the product category are determined. The set of positive reviews includes multiple reviews that give positive evaluations of products belonging to the product category.
[0008] The multiple product review data are matched with the positive review set, and the successfully matched product review data is selected as the product review data to be detected;
[0009] Obtain comment images corresponding to the product comment data to be detected, and detect the status of the product contained in the comment images;
[0010] If the product in the comment image is in a preset state, the product comment data to be detected is determined to be an invalid comment.
[0011] Secondly, this application provides a device for processing product review data, comprising:
[0012] The acquisition module is suitable for acquiring multiple product review data corresponding to a target product;
[0013] The processing module is adapted to determine the product category of the target product and a set of positive reviews corresponding to the product category, wherein the set of positive reviews includes multiple reviews that give positive evaluations of products belonging to the product category;
[0014] The matching module is adapted to match the multiple product review data with the positive review set, and filter the successfully matched product review data as product review data to be detected;
[0015] The detection module is adapted to acquire a comment image corresponding to the product comment data to be detected, and to detect the status of the product contained in the comment image;
[0016] The judgment module is adapted to determine that the product review data to be detected is invalid if the status of the product contained in the review image is a preset status.
[0017] Thirdly, this application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the above-described method.
[0018] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-described method when executed by a processor / processor core.
[0019] In the embodiments provided in this application, multiple product review data corresponding to the target product are obtained, and the product category of the target product is determined. Then, a set of positive reviews corresponding to the product category is obtained. The multiple product review data are then matched with the set of positive reviews, and the successfully matched product review data is filtered as the product review data to be detected. Correspondingly, review images corresponding to the product review data to be detected are obtained, and the state of the product contained in the review image is detected. If the state of the product contained in the review image is a preset state, the product review data to be detected is determined to be an invalid review. Therefore, on the one hand, this application first determines the product category of the target product and obtains a set of positive reviews corresponding to the product category. Then, it matches multiple product review data with the set of positive reviews, and filters the successfully matched product review data as the product review data to be detected. Compared with product review data that has not been matched with a set of positive reviews, the method of matching multiple product review data with a set of positive reviews corresponding to the product category can effectively filter out reviews that give positive evaluations of product quality. On the other hand, this application obtains review images corresponding to the product review data to be detected and detects the state of the product contained in the review images. If the state of the product contained in the review images is a preset state, the product review data to be detected is determined to be an invalid review. In summary, since the state of the product contained in the review images can reflect specific information about the product, such as whether it has been opened, whether the actual product matches the purchase record, etc., judging whether the product images of positive reviews are in a preset state can effectively filter out invalid reviews. In practice, the product review data is first matched with a set of positive reviews to filter out reviews that positively evaluate product quality. Then, the corresponding review images are obtained. Finally, the validity of a review is determined by checking if the product image in the review image is in a preset state. If the review image is determined to be in a preset state, the review is considered invalid.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the embodiments of the present application to explain the application and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed example embodiments described with reference to the accompanying drawings, in which:
[0022] Figure 1A flowchart illustrating a method for processing product review data provided in this application embodiment;
[0023] Figure 2 A flowchart illustrating another method for processing product review data provided in this application embodiment;
[0024] Figure 3 A schematic diagram of a product review data processing device provided in this application embodiment;
[0025] Figure 4 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions of this application, exemplary embodiments of this application are described below in conjunction with the accompanying drawings, including various details of the embodiments of this application to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] Where there is no conflict, the various embodiments of this application and the features thereof may be combined with each other.
[0028] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0030] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0031] The method for processing product review data according to embodiments of this application can be executed by electronic devices such as terminal devices or servers. Terminal devices can be in-vehicle devices, user devices, mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants, handheld devices, computing devices, wearable devices, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Specifically, the method can be implemented by a processor calling a computer program stored in memory.
[0032] In related technologies, big data and graph theory are commonly used to statistically analyze transaction information, product browsing and order data, store customer service communication information, and product review data. This allows for the identification and handling of orders and reviews fabricated by illegal companies, organizations, and professionals. However, current methods are ineffective in detecting and addressing covert commercial order-brushing activities involving a wide range of participants and relatively normal e-commerce transactions, especially those that lure users into creating fake positive reviews with promises of cashback. Furthermore, the wide range of content contained in product reviews, including text, images, and videos, makes it difficult for existing big data and graph theory methods to achieve the desired results. To address the aforementioned issues, this application's embodiments, on the one hand, determine the product category of the target product and obtain a set of positive reviews corresponding to that product category. Then, multiple product review data are matched against this set of positive reviews, and the successfully matched reviews are selected as the product review data to be detected. Compared to product review data that has not been matched against a set of positive reviews, matching multiple product review data against a set of positive reviews corresponding to the product category effectively filters out reviews that positively evaluate the product quality. On the other hand, review images corresponding to the product review data to be detected are obtained, and the state of the product contained in the review images is detected. If the state of the product contained in the review image is a preset state, the product review data to be detected is determined to be invalid. In short, since genuine product reviews can only be obtained after the product has been opened and used, judging whether the product image of a positive review is in a preset state can effectively filter out invalid reviews. In practice, the product review data is first matched with a set of positive reviews to filter out reviews that positively evaluate product quality. Then, the corresponding review images are obtained. Finally, the validity of the review is determined by checking if the product status in the review image matches a preset state. If the review image is determined to be in a preset state, the review is considered invalid and suspected of being fraudulent.
[0033] Figure 1A flowchart illustrating a method for processing product review data, provided as an embodiment of this application. (See attached diagram.) Figure 1 The method includes:
[0034] Step S110: Obtain multiple product review data corresponding to the target product.
[0035] The target product is a pre-selected product used to check whether its review data is valid. The product review data consists of multiple reviews for the target product; preferably, positive reviews with images are selected as the corresponding product review data. Specifically, the target product can be identified in various ways, such as selecting products with few views but many positive reviews on shopping websites, or products that have sold in large quantities but never received negative reviews. These products are initially judged as potentially having fraudulent reviews (i.e., invalid reviews) and will be further evaluated later. In short, the target products include products suspected of fraudulent reviews and products whose reviews may be invalid for various other reasons.
[0036] Therefore, by obtaining multiple product review data corresponding to the target product, products suspected of fraudulent order placement can be initially screened out and their product reviews can be obtained. The product review data will then be further verified.
[0037] Step S120: Determine the product category of the target product and the set of positive reviews corresponding to the product category. The set of positive reviews includes multiple reviews that give positive evaluations of products belonging to the product category.
[0038] The product category refers to the class to which the target product belongs, specifically, it can be digital products, personal care, clothing, cosmetics, food, fresh produce, etc. Users will use different words to describe different types of products. For example, when reviewing "clothing," they will use words like "comfortable" and "durable" to give a positive review, while when reviewing "fresh produce," they will use words like "fresh" and "delicious" to give a positive review. The positive review set refers to a collection of words used to store words commonly used to give positive reviews of the quality of products in different product categories. These words can be stored using linked lists, databases, or other similar formats.
[0039] In one alternative implementation, the collection of positive reviews corresponding to a product category is obtained in the following way:
[0040] First, for multiple product categories, obtain historical review data with multiple positive reviews corresponding to each product category. The product category refers to the type of product to which it belongs, specifically including digital products, personal care, clothing, cosmetics, food, fresh produce, etc. The historical review data with multiple positive reviews can be obtained from a historical review database or manually retrieved directly from the review sections of popular products on major shopping websites. "Positive review type" specifically means that the user gave a positive review. Therefore, by obtaining multiple positive reviews from multiple product categories, a diverse range of reviews for different products can be collected, facilitating subsequent processing.
[0041] Then, each historical review data corresponding to the product category undergoes word segmentation and part-of-speech tagging. Based on the part of speech of each segmented word, positive review units composed of sentiment feature words and descriptive object words are extracted from the historical review data. Word segmentation involves dividing a complete review into multiple different words; part-of-speech tagging involves assigning different parts of speech to the segmented words according to grammatical rules, specifically nouns, adjectives, adverbs, pronouns, etc. Sentiment feature words are words that express emotional characteristics, such as adverbs like "very," "extremely," and "quite," and adjectives like "comfortable," "beautiful," and "affordable." Descriptive object words are words that specifically refer to objects, such as pronouns like "it," "this," and "that," and nouns like "watch," "camera," and "bag." Furthermore, the positive review unit, composed of sentiment feature words and descriptive object words, expresses the positive emotional inclination of reviewers towards the product in the historical review data. Therefore, by processing multiple historical comment data and extracting sentiment feature words and descriptive object words, it is possible to effectively locate the object being commented on and the sentiment tendency of the comment, and combine positive sentiment feature words and descriptive object words into positive comment units.
[0042] Finally, based on the multiple positive review units corresponding to each product category, a category-specific positive review set is generated. This set consists of multiple positive review units, which are the smallest units constituting a complete evaluation. For example, if only biased reviews are included without describing the object, the review will be incomplete because the subject of the review is unclear. Furthermore, including too much irrelevant information will negatively impact subsequent matching results. Therefore, mining positive review units can balance the completeness of the reviews with the accuracy of subsequent matching.
[0043] In one optional implementation, when extracting positive comment units composed of sentiment feature words and descriptive object words from historical comment data based on the part-of-speech of each word segment, it is achieved in the following way:
[0044] First, word segments with the first specified part of speech are extracted as sentiment feature words. The first specified part of speech includes adjectives and / or adverbs; sentiment feature words refer to words that can express emotional characteristics. In a specific example, adverbs such as "very," "extremely," and "quite," as well as adjectives such as "comfortable," "beautiful," and "affordable," are extracted as sentiment feature words from historical comment data that has undergone word segmentation.
[0045] Next, using sentiment feature words as the central words, a forward search and / or backward search are performed in the current historical comment data to extract the word segments with the second specified part of speech as descriptive object words. The second specified part of speech includes nouns and / or pronouns; the descriptive object words refer to words that can specifically refer to objects. In a specific example, pronouns represented by "it," "this," and "that," and nouns represented by "watch," "camera," and "bag," can all be used as descriptive object words. Furthermore, the process of "forward search" and "backward search" is as follows: after identifying the sentiment feature words in the historical comment data, based on the comment statements containing those sentiment feature words, and using those sentiment feature words as the origin, a search is performed sequentially from the beginning and end of the comment statements to select nouns and / or pronouns as descriptive object words.
[0046] Finally, the sentiment feature words and descriptive words are combined into a positive comment unit. In a specific example, "comfortable" in "The shoes are so comfortable!" has been identified as the sentiment feature word. Based on this word, we search for nouns and / or pronouns in "so comfortable", "wearing", "shoes", and "ah" at the beginning and end of the sentence, and determine that "shoes" is a noun. Then, "shoes" and "comfortable" are combined into a positive comment unit, representing the meaning of "the shoes are comfortable", which carries a praising sentiment.
[0047] In summary, this implementation aims to extract the descriptive object and the biased comments made on the descriptive object through part-of-speech tags, and combine the two into a positive comment unit to facilitate subsequent matching operations.
[0048] In one alternative implementation, when extracting the searched words with the second specified part-of-speech tag into descriptive object words, it is done in the following way:
[0049] First, the word segments with the second specified part of speech are matched against a preset filter word list. If a match fails, the word segments with the second specified part of speech are extracted as descriptive words. The filter word list stores descriptive words unrelated to product quality, including logistics-related terms and / or after-sales service-related terms. Therefore, filtering out reviews unrelated to quality, such as those related to logistics and after-sales service, can improve the accuracy of the solution. In a specific example, suppose the historical review data is "The shoes are so comfortable, and the delivery was so fast!", with the sentiment features "comfortable" and "fast". During forward and / or backward searches, the words with the second specified part of speech are "shoes" and "delivery". When matched against the preset filter word list, "delivery" matches as a logistics-related term and is therefore filtered. The opposite of "delivery," "fast," is also filtered. "Shoes," however, does not match the preset filter word list and is extracted as a descriptive word. Therefore, by using a filter word list, words unrelated to describing product quality can be filtered out, thus improving the accuracy of the solution.
[0050] In one optional implementation, when generating a set of positive reviews corresponding to a product category based on multiple positive review units for each product category, the following method is used:
[0051] First, for each positive review unit in the current product category, calculate the first occurrence frequency of the positive review unit in the current product category and the second occurrence frequency of the positive review unit in all product categories. In a specific example, the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is used to calculate the first occurrence frequency of the positive review unit in the current product category and the second occurrence frequency of the positive review unit in all product categories. The TF-IDF algorithm is an inverted document probability calculation method. Its main idea is that if a word or phrase appears frequently in one article but rarely in other articles, it is considered to have good category discrimination ability and is suitable for classification. Specifically, when calculating the first occurrence frequency of the positive review unit in the current product category and the second occurrence frequency of the positive review unit in all product categories, the sentiment feature words and descriptive object words in the positive review unit are counted as a whole. For example, if there are 100 positive reviews for a product category, and 60 of them are positive review units consisting of sentiment words and descriptive words as a whole, then the frequency of occurrence is 60%.
[0052] Then, based on the comparison between the first and second occurrence frequencies, several positive review units are selected and added to the category-specific positive review set corresponding to the current product category. Specifically, when the first occurrence frequency of a positive review unit in the current product category is greater than its second occurrence frequency across all product categories, it indicates that the positive review unit has a high category differentiation ability. Preferably, the ratio of the first to second occurrence frequencies of each positive review unit is calculated; the larger the ratio, the higher the differentiation ability of the positive review unit. The results are then sorted from largest to smallest, and the positive review units corresponding to the top few values are added to the category-specific positive review set corresponding to the current product category.
[0053] Therefore, after generating positive comment units, the idea of using the term frequency inverse document frequency (TF-IDF) algorithm to select positive comment units with high discriminative power and form a category-specific positive comment set effectively improves the accuracy of matching.
[0054] Step S130: Match multiple product review data with the positive review set, and filter the successfully matched product review data as the product review data to be tested.
[0055] Among them, the product review data consists of multiple reviews of the products to be detected for fraudulent orders, which are obtained in advance. The positive review set of the category has also been generated in the manner described above. Matching the two can filter out the positive review data in the review data and mark it as product review data that needs further detection.
[0056] Step S140: Obtain the comment image corresponding to the product comment data to be detected, and detect the status of the product contained in the comment image.
[0057] In this context, the comment images are the accompanying images in the product comment data to be detected. By detecting these comment images, the product's status can be determined. The product status includes: the product's usage status, the actual quantity of the product, and the product's photographic status. In short, any status that can characterize a product can be used as the product status to be detected in this embodiment; this application does not limit the specific connotation of the product status. For example, the product's usage status can be determined by whether the product's packaging is intact. The product's packaging status specifically includes: unopened packaging (obviously the product is not used), unopened packaging (obviously the product is not used), and opened packaging (indicating the product has been used). The actual quantity of the product is used to indicate whether the actual quantity matches the purchase data. The photographic status of the product is used to indicate whether the comment image contains global information about the product. For example, the comment image corresponding to valid comment data should contain global information about the product, rather than only containing partial information. In addition, the shooting status of the product can also be used to characterize information such as the clarity of the product in the review image. For example, the product in the review image corresponding to valid review data should be clearly photographed, not blurry, and it can also indicate whether the product in the review image is the same as the target product, thus avoiding the appearance of fake images.
[0058] In one alternative implementation, the detection of the status of a product contained in a review image is achieved in the following way:
[0059] First, the review image is input into a first image classification model. This first image classification model is a multi-classification model, and its output includes: various types of product packaging, and no product packaging. The first image classification model uses a pre-trained neural network model to identify whether the review image contains packaging or a package.
[0060] Then, if the review image contains product packaging based on the output of the first image classification model, the packaging type of the product packaging is determined. The input to the first image classification model is the product review image, and after processing by the neural network model, it outputs the packaging type in the review image, such as foam packaging, cardboard boxes, plastic boxes, bags, or no packaging.
[0061] Next, a second image classification model corresponding to the packaging type is selected from multiple second image classification models, and the comment image is input into the selected second image classification model. Each second image classification model corresponds to one packaging type, and each second image classification model is a binary classification model. The output of the second image classification model includes: an output representing the unopened state and an output representing the opened state. There are multiple second image classification models, each corresponding to a different packaging type from the multiple outputs of the first image classification model, and each second image classification model specifically identifies whether a single packaging type has been opened. Packaging types include foam packaging, cardboard boxes, plastic boxes, and bags. Since the feature points for detecting whether different packaging containers are opened are different, training different second image classification models for different containers can significantly improve the detection accuracy.
[0062] Finally, the output of the second image classification model is used to determine whether the product in the review image is in an unopened state. The second image classification model outputs only two values: one representing an unopened state and the other representing an opened state. In a specific example, 1 and 0 can be used to represent the output: if the output is 1, the container has been opened; if the output is 0, the container is unopened.
[0063] In summary, this implementation aims to use two pre-trained neural network models to identify the type of packaging and whether the packaging has been opened, thereby detecting whether the product in the review image has been opened and used, and thus determining whether the review is a fake positive review.
[0064] Therefore, by examining the images in the comments, it is possible to determine whether the product in the image has been opened. If the product in the comment image is unopened but the text indicates a positive review, then the comment can be considered a fake positive review to some extent, which is a form of covert order-brushing behavior.
[0065] Step S150: If the product status contained in the comment image is a preset status, the product comment data to be detected is determined to be an invalid comment.
[0066] As mentioned above, the product status can indirectly reflect the validity of a review. Therefore, if the product in the review image corresponding to a positive review is in a preset status, the review data to be detected can be determined as invalid. The specific type of the preset status can be set according to actual business needs. For example, if the product status includes the product's usage status, the preset status can be an unopened state; if the product status includes the actual quantity of the product, the preset status can be a state where the actual quantity does not match the purchased quantity; if the product status includes the product's photographic state, the preset status can be a state that does not contain global product information or a state where the clarity is below a preset clarity threshold, etc. In short, the type of preset status can be flexibly set according to business needs. If the product status in the review image matches the preset status, the review data to be detected is determined to be invalid. Therefore, invalid reviews in this embodiment can include various situations such as fake reviews, erroneous reviews, and low-value reviews (due to insufficient clarity). In short, any review whose authenticity is questionable or whose content is inaccurate can be considered invalid.
[0067] In one optional implementation, the preset state mainly refers to the product being unopened, and invalid reviews mainly refer to fake reviews. For example, if the product is unopened, it indicates that the positive review was made without using the product, suggesting fake reviews. Furthermore, in a specific example, if it is detected that the product in the review image corresponding to the positive review is not unopened, or the product in the review image is not the actual product being sold, or the image is too blurry to identify whether it is the actual product being sold, or the order only purchased 2 items but the review image shows 3 items, then it can be concluded that the positive review is fake and involves fake reviews.
[0068] In one alternative implementation, after determining that the product review data to be detected is invalid, the product detected as a fraudulent review can be processed using at least one of the following two implementation methods:
[0069] In a first optional implementation, a first preset processing is performed on product review data determined to be of the fraudulent type. This first preset processing includes: deleting product review data, reducing the ranking weight of the product review data, and / or adding alert messages. Specifically, for product reviews identified as fraudulent, the first preset processing reduces the likelihood of the review being seen by the buyer by deleting the product review data, reducing the ranking weight of the product review data, and adding alert messages to the review, or alerts the buyer that the review is fraudulent.
[0070] In summary, this implementation method is mainly aimed at single product reviews, and uses multiple methods to process the fake review to prevent it from misleading buyers.
[0071] In the second optional implementation, a second preset processing is performed on the target product; wherein the second preset processing includes: reducing search exposure and sending a warning message to the store where the target product is located. Specifically, the second preset processing targets products or stores whose product reviews are determined to be fraudulent, by reducing the search exposure of the product or store to decrease the chances of the product being viewed by buyers, or by sending a warning message to the store where the product is located, alerting it to the detection of fraudulent activity.
[0072] In summary, this implementation method primarily targets products or stores. Unlike the first method's handling of reviews, this method reduces pageviews through various means to penalize fraudulent transactions. Therefore, both methods have their advantages. In practice, they can be implemented individually or in combination. When combined, they can address both malicious fraudulent reviews and the products and stores engaging in such activities, thus preventing undue influence on ordinary buyers and maintaining market fairness.
[0073] In one alternative implementation, the second preset processing on the target product is performed in the following way:
[0074] The number of product review data of the type of fraudulent order corresponding to the target product is determined. If the number meets the preset conditions, a second preset processing is performed on the target product. The preset conditions include: the number of product review data of the type of fraudulent order is greater than a preset quantity threshold, and / or, the proportion of product review data of the type of fraudulent order in the total number of reviews exceeds a preset proportion threshold.
[0075] The statement "If a positive review was given, but the product in the review image is unopened" can only preliminarily determine if the review involved fraudulent activity, but it cannot definitively confirm whether the product itself was involved in fraudulent activity. Therefore, at least one of the following two methods is needed for further determination:
[0076] In the first method, if the number of product reviews for products suspected of being fraudulent exceeds a preset threshold, then the product is deemed to have indeed engaged in fraudulent activities. Specifically, if the number of fraudulent reviews exceeds the preset threshold, it indicates that the product has too many fraudulent reviews, thus qualifying as fraudulent activity. In short, this method is primarily suitable for identifying fraudulent activities when the total number of reviews for a product is relatively small.
[0077] In the second method, if the proportion of product reviews containing fraudulent orders exceeds a preset threshold, the product is deemed to have engaged in fraudulent order activity. Specifically, if the number of fraudulent reviews exceeds the preset threshold, it indicates that the proportion of fraudulent reviews for that product is too high, thus qualifying as fraudulent activity. In summary, this method is primarily suitable for identifying fraudulent activity when a product has a large number of total reviews.
[0078] Therefore, both methods have their advantages and can be implemented individually or in combination. When implemented together, the total number of product reviews allows for flexible judgment of fraudulent activity. Punitive measures are only taken against a product when a large number of such reviews appear, thus avoiding misjudgments and maintaining fairness in the market.
[0079] In one alternative implementation, when performing the first preset processing on product review data determined to be of the fraudulent order type, it is achieved in the following way:
[0080] First, obtain user order data corresponding to product review data of the "brushing" type.
[0081] Then, check whether the user order data contains repurchase order data. If not, perform the first preset processing on the product review data of the order-brushing type.
[0082] The purpose of obtaining user order data is to detect fraudulent activity. If a user who submits a fake review has no prior order history for the product, they are considered a new customer. Since the product in the positive review photos is unopened, the user can be identified as a fraudulent customer. Conversely, if a user's order history shows multiple previous purchases of the product, they are considered a returning customer. Since they have previously purchased the same product, the unopened condition in the positive review photos is understandable, and the user is not considered a fraudulent customer.
[0083] In summary, for a review that has been identified as fraudulent, further investigation can be conducted to determine if the corresponding user has made a repeat purchase. If so, it indicates that the user may not be engaging in fraudulent activities. This can help verify fraudulent reviews and avoid misjudgments.
[0084] In summary, the implementation method described in this application involves matching product review data with a set of positive reviews to filter out product review data that gives positive feedback on product quality. Then, it obtains the review images corresponding to the filtered product review data and determines whether the product status contained in the review images is a preset status. If it is determined that the product status contained in the image of a positive review is a preset status, it indicates that the review is invalid and is suspected of being a fraudulent review.
[0085] For ease of understanding, Figure 2 A flowchart illustrating another method for processing product review data provided in this application embodiment is shown below. Figure 2 In the scenario corresponding to the embodiment, it is necessary to detect whether the product status is "opened" and whether invalid reviews are "brushed" reviews. (Refer to...) Figure 2 The method includes:
[0086] Step S210: Collect a set of positive comments.
[0087] Step S220: Collect and filter product review data with images, keeping only the reviews whose text content matches the positive review set of their respective categories, and extract the images corresponding to the filtered reviews.
[0088] Step S230: Determine whether the image contains product packaging.
[0089] In step S230, if it is determined that the image does not contain product packaging, it indicates that the express package has been opened, the review is written based on user experience, and does not belong to fake reviews, and step S260 is executed; if it is determined that the image contains product packaging, step S240 is executed.
[0090] Step S240: Determine whether the product packaging has been opened.
[0091] In step S240, if it is determined that the product is unopened, it is determined to be a fake review, and step S250 is executed; if it is determined that the product has been opened, the review is a normal review and does not belong to a fake review, and step S260 is executed.
[0092] Step S250: The comment is judged as fraudulent.
[0093] Step S260: The comment is determined to be non-fraudulent.
[0094] The process of determining whether an image contains product packaging and whether the packaging has been opened requires a neural network-based image classification model. Before use, the neural network model needs to be pre-trained by inputting images of various types of product packaging, such as parcels, bottles, bags, and boxes. The specific training method is as follows: First, collect images of the various types of packaging and label them accordingly. When collecting images, it is important to collect as many related types as possible; for example, parcel images should include various categories such as shrink-wrapped, boxed, and bagged parcels, and bottle images should include images of various materials and shapes. Then, the generated labeled images are augmented using image data augmentation methods including but not limited to translation, flipping, horizontal / vertical / tilting stretching or squeezing, rotation, inversion, scaling, lighting and shadow processing, color replacement, and random noise addition to generate various augmented copies. Finally, the original and augmented images are used to train the image classification model using a deep learning model.
[0095] In this example, the image classification model used can employ image pre-training techniques such as DeepCluster and outer classification layers to train category classification using the generated image data. DeepCluster is an unsupervised image representation learning method that references the AlexNet network structure, employing a five-layer convolutional structure for deep feature extraction. The five convolutional layers have 96, 256, 384, 384, and 256 channels, respectively. Furthermore, each convolutional layer is followed by a Batch Normalization and a ReLU function, and then three fully connected layers with a 0.5 Dropout applied between layers. To further decouple image color and semantics and prevent color interference with semantic learning, DeepCluster applies Sobel filtering to the input image to separate the color signal. This method uses a large amount of image data and, based on K-means clustering, continuously optimizes the feature extraction capability of the five convolutional layers through comprehensive calculations of clustering results under different label systems and cross-entropy, ensuring excellent results for the final feature representation across various clustering label systems. Next, the pre-trained DeepCluster model is used to extract deep features from the images, and then further classifies these features into images. Since the third and fourth layers of the AlexNet network trained using the DeepCluster method perform better than the fifth layer, the image feature representation output from the fourth layer is used in this example. Specifically, the output of the fourth convolutional layer is used as the image feature representation, followed by a linear layer for a linear mapping of the matrix features, and then two linear layers. The first linear layer performs the final classification, and the output dimension N of the second linear layer represents the number of target categories. The classification loss function can use the MSE loss, and the optimizer function during model training can use the Adam optimizer.
[0096] In one alternative implementation, the image classification model is trained in the following way:
[0097] First, prepare the training data. Collect online review image data, label and separate specific image sets based on image content, and then label them. The label content of each image set is the image type, including parcels, bottles, bags, boxes, etc. For ease of explanation, this implementation assumes only three container types: bottles, bags, and boxes. Therefore, the final output layer of the image classification model is N = 5 (parcels, bottles, bags, boxes, and others), and the number of images in each type is no less than 1000. The category labels can be: 0, Other; 1, Parcels; 2, Bottles; 3, Bags; 4, Boxes.
[0098] Next, data augmentation is performed. For each image in each category, image data augmentation methods including but not limited to translation, flipping, horizontal / vertical / skewed stretching or squeezing, rotation, inversion, scaling, lighting and shadow manipulation, color replacement, and random noise addition are applied to obtain augmented image copies. These copies have the same label as the original images. The purpose of these operations is to enrich the data as much as possible, allowing the model to use diverse data during training. For example, some users upload photos taken vertically, so image rotation increases the model's adaptability to this situation; some users upload photos showing only half of the packaging instead of the whole image; in this case, translation allows the model to better adapt to this situation; some users upload photos taken against the light, and lighting and shadow manipulation increases the model's adaptability to this type of data; furthermore, users upload packaging packages with varying colors, and random color replacement and random noise addition increase the model's adaptability to this type of data. In short, any data augmentation method that helps the model's receptive range can be used.
[0099] Specifically, to prevent data anomalies caused by excessive transformation, the following parameters are selected:
[0100] Image translation: The image is moved vertically and horizontally without changing more than half of its original size, and the background that is left blank after translation is uniformly colored, such as black.
[0101] Image flipping: including horizontal and vertical flipping;
[0102] Image stretching: The horizontal or vertical resolution stretching range is 0.5 to 2 times the original size;
[0103] Image lighting: Randomly adjusts the image's brightness, contrast, color saturation, sharpness, and other parameters, with an adjustment range between 0.5 and 2 times;
[0104] Image rotation: The image is rotated clockwise within a range of 0-360 degrees, with the center point of the image as the axis during rotation;
[0105] Image noise enhancement: Randomly add Gaussian noise or salt and pepper noise to the image;
[0106] Image Color: Randomly replace color blocks in the image. For example, replace all blue with red; alternatively, you can count the RGB values appearing in the image, select the top N colors with the most pixels, and randomly construct RGB values to replace them.
[0107] Next, data preprocessing is performed. For images, following the DeepCluster processing flow, images are first scaled down to 255x255 square images, then cropped to 224x224 square images using center cropping. The RGB values are then regularized to a value space with mean values of [0.485, 0.456, 0.406] and standard deviations of [0.229, 0.224, 0.225]. For the label data, each image's label is represented as a one-hot vector of size 5. For example, if the wrapper class label is 1, the corresponding vector is 01000. Finally, 10% of the samples are extracted from the overall dataset as a validation dataset.
[0108] Finally, model training is performed. During forward computation, some parameters of the DeepCluster pre-trained model are frozen and not used in training. Then, pre-processed image data is input into the model in batches. The model output layer outputs N=5 predicted classification labels and the probability of each predicted label. Then, the classification results are compared with the actual label vector of the image to calculate the MSE loss. During backpropagation, the loss function is encapsulated in the Adam optimizer, and backward propagation loss is performed. The learning rate is gradually adjusted and gradients are calculated to adjust the values of each model parameter. Every 5000 steps, classification is validated on the validation set using the current model parameters, and the classification loss on the validation set is recorded. When the model iterates until the loss on the validation set no longer decreases after three consecutive iterations, the model is considered to have converged, and the trained image classification neural network is obtained.
[0109] It should be noted that the above classification model is only one implementation of the image classification model in this application, and is intended to illustrate one possible execution method of this application. This application does not limit the network structure and training organization method of a specific image classification model, and many excellent image neural classifier structures are not outside the scope of this application.
[0110] In summary, in the embodiments provided in this disclosure, multiple product review data corresponding to the target product are obtained, and the product category of the target product is determined. Then, a set of positive reviews corresponding to the product category is obtained. The multiple product review data are then matched with the set of positive reviews, and the successfully matched product review data is filtered as the product review data to be detected. Correspondingly, a review image corresponding to the product review data to be detected is obtained, and the state of the product contained in the review image is detected. If the state of the product contained in the review image is a preset state, then the product review data to be detected is determined to be an invalid review. Therefore, on the one hand, this application first determines the product category of the target product and obtains a set of positive reviews corresponding to that category. Then, it matches multiple product review data with this set of positive reviews, filtering out the successfully matched reviews as the product review data to be tested. Compared to product review data that has not been matched with a set of positive reviews, this method of matching multiple product review data with a set of positive reviews corresponding to the product category can effectively filter out reviews that positively evaluate the product quality. On the other hand, this application obtains review images corresponding to the product review data to be tested and checks the state of the product contained in the review images. If the state of the product contained in the review image is a preset state, then the product review data to be tested is determined to be invalid. In short, since genuine product reviews can only be obtained after the product has been opened and used, judging whether the product images of positive reviews are in a preset state can effectively filter out invalid reviews. In practice, the product review data is first matched with a set of positive reviews for the product category. This filters out product reviews that positively evaluate product quality. Then, the corresponding review images are obtained. Finally, the validity of the review is determined by checking if the product status in the review image matches a preset state. If the review image is determined to be in a preset state, the review is considered invalid and suspected of being fraudulent.
[0111] Figure 3 This is a schematic diagram of the structure of a product review data processing device 30 provided in an embodiment of this disclosure, as shown below. Figure 3 As shown, it includes:
[0112] Module 31 is suitable for acquiring multiple product review data corresponding to the target product;
[0113] Processing module 32 is adapted to determine the product category of the target product and the category positive review set corresponding to the product category, wherein the positive review set includes multiple reviews that give positive evaluations of products belonging to the product category;
[0114] The matching module 33 is adapted to match the multiple product review data with the category positive review set, and filter the successfully matched product review data as product review data to be detected;
[0115] Detection module 34 is adapted to acquire a comment image corresponding to the product comment data to be detected, and to detect the status of the product contained in the comment image;
[0116] The judgment module 35 is adapted to determine that the product review data to be detected is invalid if the status of the product contained in the review image is a preset status.
[0117] Optionally, the processing module is specifically adapted to:
[0118] For multiple product categories, obtain historical comment data with multiple positive comment types corresponding to the product categories;
[0119] Each historical review data corresponding to the product category is processed by word segmentation and part-of-speech tagging. Based on the part of speech of each word segment, positive review units composed of sentiment feature words and descriptive object words are extracted from the historical review data.
[0120] Based on the multiple positive review units corresponding to each product category, a category positive review set corresponding to the product category is generated.
[0121] Optionally, the processing module is specifically adapted to:
[0122] Extract the words whose part of speech is the first specified part of speech into sentiment feature words;
[0123] Based on the aforementioned sentiment feature words, perform forward search and / or backward search in the current historical comment data, and extract the word segments with the second specified part of speech as descriptive object words;
[0124] The emotional feature words and the descriptive object words are combined into a positive comment unit;
[0125] Wherein, the first specified part of speech includes: adjectives and / or adverbs; the second specified part of speech includes: nouns and / or pronouns.
[0126] Optionally, the processing module is specifically adapted to:
[0127] The word segment with the second specified part of speech is matched with the preset filter word list. If the match fails, the word segment with the second specified part of speech is extracted as the description object word.
[0128] The filter term list is used to store descriptive objects that are not related to product quality, including: logistics-related terms and / or after-sales service-related terms.
[0129] Optionally, the processing module is specifically adapted to:
[0130] For each positive review unit in the current product category, calculate the first occurrence frequency of the positive review unit in the current product category and the second occurrence frequency of the positive review unit in all product categories;
[0131] Based on the comparison between the first occurrence frequency and the second occurrence frequency, several positive comment units are selected and added to the category positive comment set corresponding to the current product category.
[0132] Optionally, the detection module is specifically adapted to:
[0133] The comment image is input into a first image classification model; wherein, the first image classification model is a multi-classification model, and the output of the first image classification model includes: product packaging of various packaging types, and no product packaging;
[0134] If the comment image contains product packaging based on the output of the first image classification model, the packaging type of the product packaging is determined.
[0135] Select a second image classification model corresponding to the packaging type from multiple second image classification models, and input the comment image into the selected second image classification model; wherein, each second image classification model corresponds to a packaging type, and each second image classification model is a binary classification model, then the output of the second image classification model includes: an output result for representing the unopened state, and an output result for representing the opened state;
[0136] Based on the output of the second image classification model, determine whether the product in the comment image is in an unopened state.
[0137] Optionally, the determination module is specifically adapted to:
[0138] For product review data identified as fraudulent, a first preset processing step is performed; wherein, the first preset processing step includes: deleting product review data, reducing the sorting weight of product review data, and / or adding alarm notification information; or,
[0139] A second preset process is performed on the target product; wherein, the second preset process includes: reducing search exposure and sending a warning message about fraudulent orders to the store where the target product is located.
[0140] Figure 4 A block diagram of an electronic device provided in an embodiment of this disclosure, with reference to... Figure 4The electronic device includes: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to perform the above-mentioned product review data processing method.
[0141] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the above-described method for processing product review data. The computer-readable storage medium may be volatile or non-volatile.
[0142] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described method for processing product review data.
[0143] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0144] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0145] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0146] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (SA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0147] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0148] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0149] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0150] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0152] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A method for processing product review data, characterized in that, include: Obtain multiple product review data corresponding to the target product; The product category of the target product and the corresponding set of positive reviews are determined. The set of positive reviews includes multiple reviews that give positive evaluations of products belonging to the product category. The set of positive reviews is composed of positive review units extracted from historical positive review data corresponding to the product category. Each positive review unit consists of sentiment feature words and their corresponding descriptive object words. The sentiment feature words are used to characterize the sentiment type of the review, and the descriptive object words are used to characterize the object of the review. The multiple product review data are matched with the positive review set, and the successfully matched product review data is selected as the product review data to be detected; Obtain comment images corresponding to the product comment data to be detected, and detect the status of the product contained in the comment images; If the product in the review image is in an unopened state, the product review data to be detected is determined to be an invalid review.
2. The method according to claim 1, characterized in that, The set of positive reviews corresponding to the product category is generated in the following way: For multiple product categories, obtain historical comment data with multiple positive comment types corresponding to the product categories; Each historical review data corresponding to the product category is processed by word segmentation and part-of-speech tagging. Based on the part of speech of each word segment, the positive review unit is extracted from the historical review data. Based on the multiple positive review units corresponding to each product category, a category positive review set corresponding to the product category is generated.
3. The method according to claim 2, characterized in that, The step of extracting positive comment units composed of sentiment feature words and descriptive object words from the historical comment data based on the part-of-speech of each word segment includes: Extract the words whose part of speech is the first specified part of speech into sentiment feature words; Based on the aforementioned sentiment feature words, perform forward search and / or backward search in the current historical comment data, and extract the word segments with the second specified part of speech as descriptive object words; The emotional feature words and the descriptive object words are combined into a positive comment unit; Wherein, the first specified part of speech includes: adjectives and / or adverbs; the second specified part of speech includes: nouns and / or pronouns.
4. The method according to claim 3, characterized in that, The step of extracting the searched words with the second specified part of speech as descriptive object words specifically includes: The word segment with the second specified part of speech is matched with the preset filter word list. If the match fails, the word segment with the second specified part of speech is extracted as the description object word. The filter term list is used to store descriptive objects that are not related to product quality, including: logistics terms and / or after-sales service terms.
5. The method according to claim 2, characterized in that, The step of generating a category-specific positive review set corresponding to each product category based on multiple positive review units includes: For each positive review unit in the current product category, calculate the first occurrence frequency of the positive review unit in the current product category and the second occurrence frequency of the positive review unit in all product categories; Based on the comparison between the first occurrence frequency and the second occurrence frequency, several positive comment units are selected and added to the category positive comment set corresponding to the current product category.
6. The method according to claim 1, characterized in that, The detection of the status of the product contained in the comment image includes: The comment image is input into a first image classification model; wherein, the first image classification model is a multi-classification model, and the output of the first image classification model includes: product packaging of various packaging types, and no product packaging; If the comment image contains product packaging based on the output of the first image classification model, the packaging type of the product packaging is determined. Select a second image classification model corresponding to the packaging type from multiple second image classification models, and input the comment image into the selected second image classification model; wherein, each second image classification model corresponds to a packaging type, and each second image classification model is a binary classification model, then the output of the second image classification model includes: an output result for representing the unopened state, and an output result for representing the opened state; Based on the output of the second image classification model, determine whether the product in the comment image is in an unopened state.
7. The method according to any one of claims 1-6, characterized in that, The invalid reviews include those involving fraudulent order placement; after determining that the product review data to be detected is invalid, the process further includes: For product review data identified as fraudulent, a first preset processing step is performed; wherein, the first preset processing step includes: deleting product review data, reducing the sorting weight of product review data, and / or adding alarm notification information; or, A second preset process is performed on the target product; wherein, the second preset process includes: reducing search exposure and sending a warning message about fraudulent orders to the store where the target product is located.
8. A device for processing product review data, characterized in that, include: The acquisition module is suitable for acquiring multiple product review data corresponding to a target product; The processing module is adapted to determine the product category of the target product and the set of positive reviews corresponding to the product category. The set of positive reviews includes multiple reviews that give positive evaluations of products belonging to the product category. The set of positive reviews is composed of positive review units extracted from historical positive review data corresponding to the product category. Each positive review unit consists of sentiment feature words and their corresponding descriptive object words. The sentiment feature words are used to characterize the sentiment type of the review, and the descriptive object words are used to characterize the object of the review. The matching module is adapted to match the multiple product review data with the positive review set, and filter the successfully matched product review data as product review data to be detected; The detection module is adapted to acquire a comment image corresponding to the product comment data to be detected, and to detect the status of the product contained in the comment image; The judgment module is adapted to determine that the product review data to be detected is invalid if the product contained in the review image is in an unopened state.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the product review data processing method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for processing product review data as described in any one of claims 1-7.