Methods, apparatus, computer equipment, and storage media for analyzing e-commerce user reviews
Patent Information
- Application Number
- CN202610874426.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0004]本申请实施例提供了电商用户口碑的分析方法、装置、计算机设备及存储介质,可以解决现有技术中分析电商用户口碑的分析结果不够准确的技术问题
[0015]This application provides a method, apparatus, computer device, and storage medium for analyzing e-commerce user reviews. The method includes: responding to a review analysis command, acquiring an initial data stream of user reviews on an e-commerce platform and scanning it to obtain a first scan result; based on the real-time attention in the initial data stream, performing a first allocation processing of collection resources on products in the high-frequency attention set of the first scan result to obtain a second scan result, wherein the first allocation processing does not change the overall collection resources consumed, and the higher the real-time attention, the more collection resources are invested; performing dual-stream parallel semantic analysis and false information identification processing on the content of the initial data stream, fusing them to obtain a fused content feature containing information on the credibility of the reviews, wherein the fused content feature incorporates user sentiment and false information; based on the fused content feature, performing a second allocation processing of collection resources on the products corresponding to the second scan result to obtain a third scan result, wherein the second allocation processing does not change the overall collection resources consumed, and the allocated collection resources are associated with the content feature; and performing review analysis based on the third scan result to obtain a user review analysis result. This application not only supports dynamic data adaptation through data collection and allocation to ensure data real-time performance, but also integrates a novel analysis framework with semantic awareness and fake data identification mechanisms to ensure data stability without altering the overall collection resources. This enhances the credibility of the scanned data in multiple ways, thereby improving the accuracy of the analysis results when analyzing e-commerce user reviews.
Smart Images

Figure CN122415143B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing, and in particular to methods, apparatus, computer equipment, and storage media for analyzing user reviews in e-commerce. Background Technology
[0002] Monitoring user reviews on e-commerce platforms is a crucial aspect of brand management and market analysis in the digital age. By automating the processing of user-generated content from various e-commerce, social media, and content platforms, the aim is to extract and analyze public attitudes and opinions towards products, services, or brands. This technology typically involves multiple branches of technology, including web data collection, natural language processing, data mining, and visualization. Its core lies in the efficient and accurate interpretation of massive amounts of unstructured text data.
[0003] However, existing technologies have significant shortcomings in practical applications. Data structures, interface protocols, and anti-scraping strategies vary greatly across different platforms. Traditional fixed-period data collection methods struggle to guarantee data real-time performance and stability. The high concurrency of massive amounts of comment data and batch-processing-based data processing architectures lead to severe delays in analysis results. Conventional sentiment analysis models are insufficient in understanding complex linguistic phenomena such as irony and internet slang, and struggle to effectively identify mass-produced, patterned fake comments, significantly reducing the accuracy of the analysis results and leaving room for improvement. Therefore, existing technologies suffer from inaccurate analysis results for e-commerce user reviews. Summary of the Invention
[0004] This application provides a method, apparatus, computer equipment, and storage medium for analyzing e-commerce user reviews, which can solve the technical problem that the analysis results of e-commerce user reviews are not accurate enough in the prior art.
[0005] In a first aspect, embodiments of this application provide a method for analyzing e-commerce user reviews, which includes: In response to the word-of-mouth analysis command, the system obtains the initial data stream of user reviews on the e-commerce platform and scans it to obtain the first scan result. Based on the real-time attention in the initial data stream, the first allocation process of collection resources is performed on the products in the high-frequency attention set in the first scan result to obtain the second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The content of the initial data stream is subjected to dual-stream parallel semantic analysis and false information identification processing, and the resulting fused content features include information on the credibility of the comments. The fused content features incorporate user sentiment and information on false word-of-mouth. Based on the fused content features, a second allocation process of collection resources is performed on the product corresponding to the second scan result to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. Based on the third scan result, a reputation analysis is performed to obtain the user reputation analysis result.
[0006] In some embodiments, the real-time attention is the real-time comment increment. Based on the real-time attention in the initial data stream, a first allocation process of collection resources is performed on the products in the high-frequency attention set in the first scan result to obtain a second scan result, including: The initial data stream whose real-time comment increment exceeds a preset increment threshold is aggregated and divided in the semantic space to obtain a high-frequency attention set. A second scanning result is obtained by increasing the sampling frequency of products in the high-frequency attention set and decreasing the sampling frequency of products outside the high-frequency attention set, wherein the increased frequency and the decreased frequency are equal in terms of acquisition resources.
[0007] In some embodiments, the process of performing dual-stream parallel semantic analysis and fake news detection on the content of the initial data stream, and fusing them to obtain fused content features containing information on the credibility of comments, includes: The initial data stream is subjected to feature obfuscation to obtain an obfuscated stream with controlled noise, wherein the feature obfuscation includes replacing synonyms and / or transposing characters; The initial data stream is then subjected to standard cleaning processing to obtain an explicit stream with random noise removed. The obfuscated stream is input into a preset sentiment model for sentiment analysis, and the sentiment features of the comment are output. The explicit stream is input into a preset adversarial detection model for fake pattern recognition, and the output is the fake probability feature of the comment. The fake pattern recognition includes extracting text length and / or information entropy. The sentiment features and the false probability features are fused to obtain fused content features that include information on the credibility of the comments.
[0008] In some embodiments, after performing feature fusion processing on the sentiment features and the false probability features to obtain fused content features containing information on the credibility of comments, the process includes: Based on the fused content features, a credibility score representing the credibility of the comment is calculated; Each comment is filtered based on a preset credibility threshold and the credibility level to obtain the target comments; The products contained in the target reviews are considered as popular products, and a popular product set including multiple of the popular products is obtained. The items that intersect with the high-frequency attention set and the popular item set are identified as public goods; A third scanning result is obtained by increasing the collection resources invested in the public goods and decreasing the collection resources for other goods besides the public goods, wherein the increase and decrease correspond to equal collection resources.
[0009] In some embodiments, the public goods originate from multiple platforms. After determining that the goods intersecting the high-frequency attention set and the popular goods set are public goods, the process includes: Based on the multi-dimensional characteristics of public goods across multiple platforms, a cross-platform product association table composed of temporary entity links is generated. If at least two temporary entity links in the cross-platform product association table are ambiguous, the temporary entity links involved in the ambiguity will be marked as pending arbitration. The ambiguity refers to the fact that the product name dimension is consistent, but other dimensions are inconsistent. Remove the temporary entity links in the pending arbitration status from the cross-platform product association table; Aggregate the products linked to temporary entities awaiting arbitration to identify high-risk products and prompt for manual intervention.
[0010] In some embodiments, multi-dimensional features include product name, price, and product brand. Based on the multi-dimensional features of public goods across multiple platforms, a cross-platform product association table consisting of temporary entity links is generated, including: From the product name, price, and brand of public goods on multiple platforms, the first letter of the pinyin, numerical features, and attribute roots are extracted to obtain multidimensional features of the products. If the product name, price, and brand dimensions are consistent across different platforms, and other inconsistent dimensions do not conflict, then a temporary entity link is established based on the consistent dimensions. Aggregate all the temporary entity links to obtain the cross-platform product association table.
[0011] In some embodiments, the sentiment feature is a sentiment vector matrix representing the intensity of sentiment polarity, the false probability feature is a false probability matrix, and based on the fused content features, a credibility level representing the credibility of the comment is calculated, including: The emotion vector matrix and the false probability matrix are input into a preset contradiction arbitrator; Calculate the difference between 1 and the false probability corresponding to the false probability matrix; If the maximum value of the emotional polarity intensity in the emotional vector matrix is greater than a preset first threshold, and the false probability is greater than a preset second threshold, then a preset penalty is determined to be applied to the difference, and the preset penalty weight is less than 1. The product of the difference and the preset penalty is calculated to obtain the credibility of each comment, which represents the credibility level of the comment.
[0012] Secondly, embodiments of this application also provide an e-commerce user reputation analysis device, which includes: The acquisition unit is used to respond to the word-of-mouth analysis command, acquire the initial data stream of user reviews on the e-commerce platform, and scan it to obtain the first scan result; The first allocation unit is used to perform a first allocation process on the high-frequency attention set of products in the first scan result based on the real-time attention in the initial data stream, and to obtain a second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The dual-stream parallel processing unit is used to perform dual-stream parallel semantic analysis and false information identification on the content of the initial data stream, and fuse them to obtain fused content features containing information on the credibility of comments. The fused content features incorporate user sentiment and false word-of-mouth. The second allocation unit is used to perform a second allocation process of collection resources on the product corresponding to the second scan result based on the fused content features to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. The reputation analysis unit is used to perform reputation analysis based on the third scan result to obtain user reputation analysis results.
[0013] Thirdly, embodiments of this application also provide a computer device for analyzing e-commerce user reviews, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-mentioned method.
[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.
[0015] This application provides a method, apparatus, computer device, and storage medium for analyzing e-commerce user reviews. The method includes: responding to a review analysis command, acquiring an initial data stream of user reviews on an e-commerce platform and scanning it to obtain a first scan result; based on the real-time attention in the initial data stream, performing a first allocation processing of collection resources on products in the high-frequency attention set of the first scan result to obtain a second scan result, wherein the first allocation processing does not change the overall collection resources consumed, and the higher the real-time attention, the more collection resources are invested; performing dual-stream parallel semantic analysis and false information identification processing on the content of the initial data stream, fusing them to obtain a fused content feature containing information on the credibility of the reviews, wherein the fused content feature incorporates user sentiment and false information; based on the fused content feature, performing a second allocation processing of collection resources on the products corresponding to the second scan result to obtain a third scan result, wherein the second allocation processing does not change the overall collection resources consumed, and the allocated collection resources are associated with the content feature; and performing review analysis based on the third scan result to obtain a user review analysis result. This application not only supports dynamic data adaptation through data collection and allocation to ensure data real-time performance, but also integrates a novel analysis framework with semantic awareness and fake data identification mechanisms to ensure data stability without altering the overall collection resources. This enhances the credibility of the scanned data in multiple ways, thereby improving the accuracy of the analysis results when analyzing e-commerce user reviews. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the e-commerce user reputation analysis method provided in this application embodiment; Figure 2 A schematic block diagram of an e-commerce user reputation analysis device provided in an embodiment of this application; Figure 3 A schematic block diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0020] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0021] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0022] Monitoring user reviews on e-commerce platforms is a crucial aspect of brand management and market analysis in the digital age. By automating the processing of user-generated content from various e-commerce, social media, and content platforms, the aim is to extract and analyze public attitudes and opinions towards products, services, or brands. This technology typically involves multiple branches of technology, including web data collection, natural language processing, data mining, and visualization. Its core lies in the efficient and accurate interpretation of massive amounts of unstructured text data.
[0023] However, existing technologies have significant shortcomings in practical applications. Data structures, interface protocols, and anti-scraping strategies vary greatly across different platforms. Traditional fixed-period data collection methods struggle to guarantee data real-time performance and stability. The high concurrency of massive amounts of comment data and batch-processing-based data processing architectures lead to severe delays in analysis results. Conventional sentiment analysis models are insufficient in understanding complex linguistic phenomena such as irony and internet slang, and struggle to effectively identify mass-produced, patterned fake comments, significantly reducing the accuracy of the analysis results and leaving room for improvement. Therefore, existing technologies suffer from inaccurate analysis results when analyzing e-commerce user reviews.
[0024] To address the aforementioned issues, embodiments of this application provide a method, apparatus, computer equipment, and storage medium for analyzing e-commerce user reviews.
[0025] This application addresses the challenge of handling massive, high-concurrency comment data, supporting dynamic data adaptation and coordinated collection strategies. It boasts real-time streaming processing capabilities, reducing latency and ensuring data timeliness. Furthermore, it integrates a novel analytical framework combining semantic awareness and fake review detection mechanisms, enhancing language understanding, improving the robustness of sentiment analysis, and increasing the detection rate of fake reviews, thus ensuring data stability. It effectively identifies batched data and patterned fake reviews, enhancing the accuracy of the analysis results. Therefore, it improves the accuracy of analysis results when analyzing e-commerce user reviews.
[0026] In some embodiments, the method for analyzing e-commerce user reviews is applied to a computer device for analyzing e-commerce user reviews. This computer device can be a terminal or a server. The terminal can be a smartphone, tablet computer, PDA, or laptop computer, etc.
[0027] Figure 1 This is a flowchart illustrating the e-commerce user reputation analysis method provided in this application embodiment. Figure 1 As shown, the method includes the following steps S110-S150: S110, In response to the word-of-mouth analysis command, obtain the initial data stream of user reviews on the e-commerce platform and scan it to obtain the first scan result; When conducting e-commerce reputation monitoring, this application responds to reputation analysis commands by employing distributed web crawling technology to capture publicly available user review data from various platforms, obtaining an initial data stream. Unlike existing technologies that use fixed-period crawling methods, this application does not collect data at fixed intervals. The specific collection strategy is dynamically adjusted based on the characteristics of the initial data stream, such as the content of the initial data stream or the real-time level of attention to the data stream.
[0028] Specifically, because the scanning strategy for each scan of the initial data stream in this application is not fixed, for example, the collection resources invested each time are different. Collection resources can be reflected in the sampling frequency; a higher sampling frequency indicates more sampling resources invested. For truly valuable evaluation data streams, more sampling resources are devoted, and crawler collection at a high sampling frequency can uncover more evaluation information, obtain a more realistic understanding of user experience, and make the evaluation and analysis results of user reputation more objective. Therefore, as an example, the first scan result implicitly contains the basic sampling frequency representing the collection resources invested in each product.
[0029] As an example, when acquiring the initial data stream, a uniform scan is performed across three dimensions—product category, platform, and region—at the corresponding basic sampling frequency to generate the first scan result. This first scan result includes the product category, e-commerce platform, and the region where the product is located. These scan results reflect the basic review information encompassed by the user review.
[0030] S120. Based on the real-time attention in the initial data stream, perform a first allocation process of collection resources on the products in the high-frequency attention set in the first scan result to obtain a second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. Based on the comment hotspots reflected in the initial data stream, adjust the investment of resources for the next data collection, focus on more valuable user comments, and keep up with the dynamic changes in comment trends.
[0031] The initial data stream's comment trending topics are quantified using real-time attention metrics. Higher real-time attention indicates more popular comments and greater user discussion. The higher-performing segments of real-time attention correspond to the high-frequency attention set. These high-performing comments are grouped together to obtain the high-frequency attention set. An initial allocation of data collection resources is then targeted at this high-frequency attention set. During this initial allocation, resources for collecting comments with higher real-time attention are increased, while overall data collection resources remain unchanged. This is because while increasing resources for this segment, resources for other segments are reduced, allowing for a targeted adjustment of the data collection strategy without altering the existing resource constraints.
[0032] In some embodiments, real-time attention is the real-time comment increment, and S120 includes S1201-S1202:
[0033] S1201. The initial data stream with real-time comment increments exceeding a preset increment threshold is aggregated and divided in the semantic space to obtain a high-frequency attention set; Specifically, real-time attention is reflected in the number of real-time comments. The greater the increase in real-time comments, the higher the real-time attention.
[0034] The system monitors and identifies the real-time increase in product comments within the initial data stream. When the real-time increase in comments for a product exceeds a preset dynamic threshold within a set time window, it determines that the product has high real-time attention and belongs to the currently detected comment hotspots. Then, it aggregates products in the semantic space that have real-time comment increases exceeding the preset dynamic threshold to obtain a high-frequency attention set. This high-frequency attention set is used for subsequent resource borrowing and lending adjustments based on collected resources.
[0035] S1202. Increase the sampling frequency of the products in the high-frequency attention set and decrease the sampling frequency of products outside the high-frequency attention set to obtain a second scanning result, wherein the increased frequency and the decreased frequency are equal in terms of acquisition resources.
[0036] By increasing the sampling frequency of frequently viewed products, we can allocate more resources to reputation analysis for these popular products.
[0037] In some embodiments, the sampling frequency of items outside the high-frequency attention set is also reduced. The increased frequency and the reduced frequency are equal in terms of acquisition resources. The advantages of this acquisition strategy of increasing and decreasing the frequency are twofold: first, it does not change the overall acquisition resource consumption; second, it allows for targeted focus on more valuable review items.
[0038] S130. Perform parallel semantic analysis and fake product identification on the content of the initial data stream, and fuse them to obtain fused content features containing information on the credibility of comments. The fused content features incorporate user sentiment and information on fake word-of-mouth. Step S120 above involves real-time incremental comments from the initial data stream, allocating collected resources based on statistical analysis (first allocation processing). Step S130 delves deeper into the specific data stream content, performing a second allocation processing of collected resources based on the content level. This step, through in-depth semantic and content-level analysis, yields fused content features. By performing parallel dual-stream processing on the initial data stream content, semantic analysis is performed separately to fit the specific context, obtaining user sentiment information; and false information is identified separately to determine the extent of misleading reviews. Finally, the semantic analysis results and false information identification results are fused to obtain fused content features that characterize the credibility of the comments.
[0039] In some embodiments, S130 includes S1301-S1305: S1301. The initial data stream is subjected to feature obfuscation processing to obtain an obfuscated stream with controlled noise, wherein the feature obfuscation processing includes replacing synonyms and / or transposing characters; The terms explicit flow and obfuscated flow have different meanings in various fields such as software security, code protection, and fluid mechanics. Simply put, obfuscated flow refers to a deliberately complex information transmission process for the purpose of hiding or protecting information, while explicit flow usually refers to direct and clear data transmission.
[0040] The initial data stream is subjected to feature obfuscation and standard cleaning processes through different pipelines. For example, the initial data stream is sent to a preset first cleaning pipeline, where controlled noise is introduced to obfuscate its features. The initial data stream is then sent to a preset second cleaning pipeline, where random noise is removed to perform standard cleaning.
[0041] As an example, in the first pipeline, synonym substitution and character transposition are performed on the initial data stream, introducing controlled perturbations into the original text to disrupt the machine-generated pattern and generate a distorted stream.
[0042] Feature obfuscation can be represented as: .
[0043] Where C'i represents the i-th comment text in the obfuscated stream generated after processing; Ci represents the i-th original comment text in the initial data stream; T represents a pre-defined domain synonym dictionary; Rs represents the synonym substitution rate (e.g., 0.05~0.15), which is the probability of a noun or adjective being replaced by a synonym. Pt represents the character transpose probability (e.g., 0.01~0.05), used to rearrange the internal character order of high-frequency words.
[0044] S1302. Perform standard cleaning processing on the initial data stream to obtain an explicit stream with random noise removed; Standard data cleaning typically refers to a series of standardized operations performed on the raw data stream before it enters the analysis or storage system to ensure data quality, consistency, and availability. Standard data cleaning includes missing value handling, deduplication, outlier detection, data type correction, noise filtering, unit and normalization, and business rule validation.
[0045] The initial data stream is fed into the second cleaning pipeline, and standard cleaning operations are performed on the initial data stream, such as removing HTML tags and meaningless stop words, unifying capitalization, removing extra spaces, converting the text into a token sequence that the model can recognize, removing random noise, preserving the original syntactic structure, and generating an explicit stream.
[0046] S1303. Input the obfuscated stream into a preset sentiment model for sentiment analysis, and output the sentiment features of the comment. Subsequently, the obfuscated stream is input into the pre-trained aspect-level sentiment model for sentiment analysis, outputting sentiment features representing the sentiment tendencies of each dimension. In e-commerce reputation monitoring, distributed web crawling technology is used to collect data at a pre-defined frequency. After pre-processing and data cleaning, the collected data is fed into a deep learning natural language processing model for sentiment analysis, determining whether the comments are positive, negative, or neutral, thus obtaining the sentiment features of the comments.
[0047] The input, confused stream, and corresponding token sequence are converted into numerical feature representations. Deep learning models like BERT generate context-dependent vectors for each token through multi-layer Transformer encoders and extract semantic vectors representing the entire sentence. The extracted feature vectors are then input into a classification layer to calculate the probability distribution of the text belonging to each sentiment category. For the three-category classification of positive, neutral, and negative, the model outputs three probability values, the sum of which is 1. Based on a predefined category mapping rule, the index corresponding to the highest probability is converted into a readable sentiment label, such as converting "positive" to "positive" (positive).
[0048] S1304. Input the explicit stream into a preset adversarial detection model to perform fake pattern recognition, and output the fake probability features of the comment, wherein the fake pattern recognition includes extracting text length and / or information entropy; Adversarial detection models are defensive models designed to counter adversarial attacks. Their core objective is to identify specially crafted input data intended to deceive or mislead machine learning models—these are adversarial examples. Adversarial detection models act as a crucial line of defense, identifying and intercepting such attacks.
[0049] An explicit stream is input into an adversarial detection model to identify false patterns and output false probability features. By extracting statistical features such as text length and information entropy, false probability features representing the probability of false activity are output.
[0050] In the adversarial detection phase, the cleaned comment data, i.e., the explicit stream, is input into a pre-defined adversarial detection model. The core task of the model is to identify fake patterns, specifically by extracting several quantifiable features of the comments, such as text length, number of characters or words, and information entropy. Information entropy is an indicator that measures the randomness or disorder of text. By analyzing the differences between these features and the distribution of real comments, the model calculates the probability that each comment is fake content, and finally outputs a fake probability feature between 0 and 1. The higher the probability value, the greater the likelihood that the comment is judged to be fake or machine-generated.
[0051] S1305. Perform feature fusion processing on the emotional features and the false probability features to obtain fused content features that include information on the credibility of the comments.
[0052] Based on the sentiment features and false probability features obtained from dual-stream parallel processing, feature fusion processing is performed to obtain fused content features, which can characterize the credibility of the current comment.
[0053] S140. Based on the fused content features, perform a second allocation process of collection resources on the product corresponding to the second scan result to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. After performing feature fusion processing on the sentiment features and the false probability features to obtain fused content features containing information on the credibility of the comments, S140 includes steps S1401 to S1405: S1401. Based on the fused content features, calculate the credibility level that characterizes the credibility of the comment; In some embodiments, the emotional feature is an emotional vector matrix representing the intensity of emotional polarity, and the false probability feature is a false probability matrix. S1401 includes steps A1 to A4: A1. Input the emotion vector matrix and the false probability matrix into the preset contradiction arbitrator; Input the two matrices, the emotion vector matrix and the false probability matrix, into the preset contradiction arbitrator.
[0054] Conflict arbitrators are often designed as a monitoring program or arbitration engine to handle conflicts. Conflict detection monitors all agent plans and resource requests in real time, triggering an immediate action upon detecting overlaps or mutual exclusions; it then performs evidence reasoning, requiring conflicting parties to provide evidence and confidence scores for their judgments, thus entering the evidence reasoning phase; the conflict arbitrator makes a final decision based on a pre-set priority matrix or an advanced reasoning model, and must ensure full traceability and auditability throughout the process.
[0055] A2. Calculate the difference between 1 and the false probability corresponding to the false probability matrix; Ps represents the false probability of the corresponding comment in the false probability matrix, and the difference is... .
[0056] A3. If the maximum value of the emotional polarity intensity in the emotional vector matrix is greater than a preset first threshold, and the false probability is greater than a preset second threshold, then it is determined that a preset penalty will be applied to the difference, and the preset penalty weight is less than 1. The maximum value Em of the emotional polarity intensity is extracted from the sentiment vector matrix. If Em is greater than a first threshold and the false probability Ps is greater than a second threshold, it indicates that the comment has both extreme sentiment and high false probability characteristics, triggering a penalty. For example, the first threshold is set to 0.9 and the second threshold is set to 0.8.
[0057] If a penalty is triggered, the conflict arbitrator receives the sentiment vector matrix and the false probability matrix, and then calculates the credibility comprehensive score Sc, using the following formula: Where Wp represents the preset penalty weight, such as 0.4~0.6, and Sc represents the overall credibility score of a single comment.
[0058] As an example, if no penalty is triggered, Wp can be set to 1.
[0059] A4. Calculate the product of the difference and the preset penalty to obtain the credibility of each comment, which represents the credibility of the comment.
[0060] The product of the difference and the preset penalty is used to obtain the credibility score for each comment, which represents the credibility level of the comment. The greater the credibility, the more credible the comment.
[0061] The sentiment vector matrix and credibility score, which represent the credibility of the comments, are obtained. The sentiment vector matrix and credibility score are then correlated with the corresponding comments in the standardized data stream to generate data to be stored. The data to be stored is stored in a distributed search engine and a distributed data warehouse respectively to generate real-time retrieval data and a full historical dataset.
[0062] S1402. Based on a preset credibility threshold and the credibility, each comment is filtered to obtain the target comments; S1403. The products contained in the target reviews are taken as popular products and combined to obtain a popular product set that includes multiple popular products; Target comments are obtained by filtering comments whose Sc is below the credibility threshold. The products contained in these target comments are designated as popular products, generating a hot product set, which is then fed back to the resource scheduling layer. For products at the intersection of the hot product set and the high-frequency attention set, a borrowing enhancement coefficient is applied to further increase their collection quota.
[0063] S1404. Determine that the items intersecting the high-frequency attention set and the popular item set are public goods;
[0064] High-frequency attention sets are derived from the statistical analysis results of the initial data stream. For example, high-frequency attention sets are obtained by aggregating and filtering based on the statistical real-time comment increments.
[0065] The popular product set is derived from content analysis of the initial data stream, such as semantic sentiment analysis and fake product identification, and is selected from the popular product set.
[0066] The former analyzes users' focus from a attentional perspective, while the latter analyzes what users truly care about from the psychological level of their expressions. The intersection of these two sets can represent what users are truly concerned about, and is more focused. Such an intersection sample is more representative of users' true evaluations for subsequent reputation analysis.
[0067] The goods that intersect the high-frequency attention set and the popular goods set are denoted as public goods.
[0068] The public goods are sourced from multiple platforms. Following S1404, steps B1 through B4 are included: B1. Based on the multi-dimensional characteristics of public goods across multiple platforms, generate a cross-platform product association table composed of temporary entity links; As an example, public goods can originate from different platforms, and goods across multiple platforms can be represented using a cross-platform goods association table.
[0069] Multi-dimensional features include product name, price, and brand. B1 includes steps C1 through C3: C1. Extract the first letter of the pinyin, numerical features, and attribute roots from the product name, price, and brand of public goods on multiple platforms to obtain multi-dimensional product features; By comparing vectors from different platforms, if at least three hash feature dimensions are consistent and there are no conflicts in the brand or core attribute dimensions, temporary entity links are established and aggregated to form a cross-platform product association table.
[0070] For example, obtain the product names, prices, and brand names from each platform, and extract the first letter of the pinyin, numerical features, and attribute roots of each product name to generate a multidimensional sparse vector. Determine whether the multidimensional sparse vectors from different sources meet the matching condition that at least three dimensions are consistent and all inconsistent dimensions are non-conflicting. If so, establish temporary entity links. Aggregate all the temporary entity links to generate a cross-platform product association table.
[0071] For example, multi-dimensional feature extraction for products involves extracting the following from the product name, price, and brand: the first letter of the pinyin, such as "a mobile phone," to obtain the letter 'a'; extracting numerical features from the price and specifications; and extracting attribute roots, such as keywords like "smart" and "Bluetooth." The original text or numerical values are then converted into comparable structured feature vectors.
[0072] Based on the public goods that are successfully aligned in the two sets, update the borrowing quota in the resource lending adjustment step to increase its scanning frequency.
[0073] Extract the first letter of the pinyin, numbers, and attribute roots of product information from various platforms, and generate multidimensional sparse vectors such as 256-dimensional binary vectors through hash function mapping to achieve multidimensional feature mapping.
[0074] C2. If the product name dimension, price dimension, and product brand dimension are consistent across different platforms, and other inconsistent dimensions do not conflict, then a temporary entity link is established based on the consistent dimension. For example, two product records on different platforms may be identical in terms of product name, price, and brand, but differ in other aspects such as model number. However, they can be mapped, and the inconsistencies do not conflict. Once the conditions are met, a candidate pair of temporary entity links is created based on the consistent dimensions. High-confidence candidate matching pairs are then generated.
[0075] A balance is sought between matching accuracy and recall, prioritizing high-confidence candidate pairs while leaving room for error in subsequent aggregation.
[0076] The reason why the product name, price, and brand dimensions need to be consistent is that the product name is the most direct basis for users to identify the product. Although different platforms may name the same product differently, after extracting the first letter of the pinyin and attribute roots, the core identifier will converge to a consistency. Mandating consistency in the name dimension can significantly filter out different types of products, such as mobile phones and phone cases. Price is an important quantitative characteristic of a product. Although there may be promotional price differences between different platforms, the normal price range of the same product is usually within the same order of magnitude or a fixed percentage. Requiring consistency in the price dimension can effectively exclude products with significantly different models and configurations. The product brand is the fundamental identifier of product ownership. The same product on different platforms must belong to the same brand. Requiring consistency in the brand can quickly eliminate cross-brand mismatches; for example, "phone A" and "phone B" cannot be associated even if their names are similar.
[0077] C3. Aggregate all the temporary entity links to obtain the cross-platform product association table.
[0078] Temporary entity links may include matching evidence such as name, price, or brand consistency. These matching pairs may be redundant, conflicting, or isolated. For example, the same product may be linked multiple times through different combinations of dimensions; for instance, A matches B, B matches C, and A and C may also match directly. Conflicting matches may occur due to data noise or flawed matching rules; for example, A matches B, and A also matches C, but B and C are clearly different. Some matching pairs only appear between two platforms and cannot be directly extended to other platforms.
[0079] Therefore, it is necessary to merge, deduplicatize, and validate all temporary entity links to ultimately generate a cross-platform product association table and output usable cross-platform product mapping relationships.
[0080] B2. If at least two temporary entity links in the cross-platform product association table are ambiguous, the temporary entity links involved in the ambiguity will be marked as pending arbitration. The ambiguity refers to the fact that the product name dimension is consistent, but other dimensions are inconsistent. Scanning the cross-platform product association table, if abnormal one-to-many links are found—that is, multiple temporary links from different brands or with mutually exclusive attributes pointing to the same target product—these links are marked as pending arbitration and logically removed. The automatic matching algorithm may generate links between different products due to similar names. This ambiguity is captured by checking the consistency of other dimensions and isolated with the pending arbitration mark, rather than being directly deleted or forcibly merged.
[0081] High-risk product labels are displayed to prompt human intervention and ensure that data analysis is not contaminated.
[0082] B3. Remove the temporary entity links in the pending arbitration status from the cross-platform product association table; Instead of retaining temporary entity links that are pending arbitration, remove them directly from the cross-platform product association table. This ensures the reliability of the association table, as downstream services such as price comparison and review aggregation rely on high-precision matching. Retaining disputed links may lead to incorrect price comparisons or mixed reviews, harming the user experience. After removal, the main table only contains unambiguous, high-confidence matches, facilitating fast queries.
[0083] B4. Aggregate the corresponding products of temporary entities in the pending arbitration status to obtain high-risk products and prompt manual intervention.
[0084] Temporary entity links that are pending arbitration and have been removed are aggregated and processed to identify high-risk products, ensuring they are not mixed with correct matches. This provides clear boundaries for manual review, prompting human intervention for high-risk products and resolving semantic conflicts that the algorithm cannot determine. For example, products with the same name but different specifications (one is the official standard version, the other a bundled version) require human judgment to determine whether they should be considered the same product entity. The results of manual review can be fed back into the matching model for subsequent training or rule adjustments. In the realm of high-value products, erroneous associations can lead to serious business consequences, making manual intervention essential.
[0085] S1405. Increase the collection resources invested in the public goods and decrease the collection resources for other goods besides the public goods to obtain a third scanning result, wherein the increase and decrease correspond to the same amount of collection resources.
[0086] For public goods, alignment computing resources are prioritized in the queue, and a gain coefficient is added to successfully aligned public goods to continuously optimize their acquisition frequency. Simultaneously, acquisition resources for other goods besides the public goods are reduced to obtain the third scan result. Acquisition resources are then increased to match the reduction, ensuring that overall acquisition resources remain unchanged.
[0087] The dual-stream parallel semantic analysis and fake review architecture enhances the depth of semantic understanding and the accuracy of fake reputation identification. By performing parallel feature obfuscation and standardization cleaning on the raw data stream, and combining the results of sentiment analysis and fake pattern recognition with a contradiction arbitrator, it effectively distinguishes between the complex emotional expressions of real users and patterned fake comments, providing a more credible and pure data foundation for decision-making. Through effective aggregation of cross-platform data and multi-dimensional correlation analysis, it solves the information silo problem caused by data heterogeneity between platforms.
[0088] S150. Based on the third scan result, a reputation analysis is performed to obtain the user reputation analysis result.
[0089] The process involves conducting word-of-mouth analysis to obtain user feedback results, ultimately outputting a structured report that includes: overall satisfaction, key product advantages, key product issues, product attribute ratings, word-of-mouth trends, and risk warnings. For example, if 72% are positive, 15% are neutral, and 13% are negative, then the overall satisfaction level is good.
[0090] For example, the main advantages of a certain mobile phone in word-of-mouth reviews are that 45% of the mentions mention its amazing screen display and 38% mention its fast charging speed.
[0091] For example, regarding the main issues, 22% of the negative mentions mentioned occasional system lag, and 18% mentioned excessive noise in nighttime photos.
[0092] For example, the attribute ratings are three stars for the battery, five stars for the screen, and four stars for the system.
[0093] For example, the word-of-mouth trend shows that the negative rate has risen from 10% to 18% in the past month, mainly due to the battery consumption issues of the new version.
[0094] For example, a risk warning indicated that negative reviews for product X surged by 300% between May 1st and 3rd, with 60% of these reviews having a false probability greater than 0.8, suggesting a possible malicious attack.
[0095] This application provides a method, apparatus, computer device, and storage medium for analyzing e-commerce user reviews. The method includes: responding to a review analysis command, acquiring an initial data stream of user reviews on an e-commerce platform and scanning it to obtain a first scan result; based on the real-time attention in the initial data stream, performing a first allocation processing of collection resources on products in the high-frequency attention set of the first scan result to obtain a second scan result, wherein the first allocation processing does not change the overall collection resources consumed, and the higher the real-time attention, the more collection resources are invested; performing dual-stream parallel semantic analysis and false information identification processing on the content of the initial data stream, fusing them to obtain a fused content feature containing information on the credibility of the reviews, wherein the fused content feature incorporates user sentiment and false information; based on the fused content feature, performing a second allocation processing of collection resources on the products corresponding to the second scan result to obtain a third scan result, wherein the second allocation processing does not change the overall collection resources consumed, and the allocated collection resources are associated with the content feature; and performing review analysis based on the third scan result to obtain a user review analysis result. In this application, not only is data dynamic adaptation supported through collection and allocation to ensure data real-time performance, but a novel analysis framework integrating semantic awareness and fake data identification mechanisms is also used to ensure data stability without altering the overall collection resources. This enhances the credibility of the scanned data in multiple ways, thereby improving the accuracy of the analysis results when analyzing e-commerce user reviews.
[0096] Figure 2 This is a schematic block diagram of an e-commerce user reputation analysis device provided in an embodiment of this application. Figure 2 As shown, corresponding to the above-described method for analyzing e-commerce user reviews, this application also provides an e-commerce user review analysis device 600. This e-commerce user review analysis device 600 includes a unit for executing the above-described method for analyzing e-commerce user reviews, and can be configured in terminals such as desktop computers, tablet computers, and laptops. For details, please refer to... Figure 2 The e-commerce user reputation analysis device 600 includes an acquisition unit 601, a first allocation unit 602, a dual-stream parallel processing unit 603, a second allocation unit 604, and a reputation analysis unit 605, wherein: The acquisition unit 601 is used to respond to the word-of-mouth analysis command, acquire the initial data stream of user reviews on the e-commerce platform, and scan it to obtain the first scan result; The first allocation unit 602 is used to perform a first allocation process of collection resources on the products of the high-frequency attention set in the first scan result based on the real-time attention in the initial data stream, so as to obtain a second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The dual-stream parallel processing unit 603 is used to perform dual-stream parallel semantic analysis and false information identification on the content of the initial data stream, and fuse them to obtain fused content features containing information on the credibility of comments. The fused content features incorporate user sentiment and false word-of-mouth. The second allocation unit 604 is used to perform a second allocation process of collection resources on the product corresponding to the second scan result based on the fused content features to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. The reputation analysis unit 605 is used to perform reputation analysis based on the third scan result to obtain user reputation analysis results.
[0097] In some embodiments, the real-time attention is the real-time comment increment. The first allocation unit 602, based on the real-time attention in the initial data stream, performs a first allocation process of collection resources on the products in the high-frequency attention set of the first scan result to obtain a second scan result, specifically used for: The initial data stream whose real-time comment increment exceeds a preset increment threshold is aggregated and divided in the semantic space to obtain a high-frequency attention set. A second scanning result is obtained by increasing the sampling frequency of products in the high-frequency attention set and decreasing the sampling frequency of products outside the high-frequency attention set, wherein the increased frequency and the decreased frequency are equal in terms of acquisition resources.
[0098] In some embodiments, the dual-stream parallel processing unit 603 performs dual-stream parallel semantic analysis and fake news detection on the content of the initial data stream, fusing them to obtain fused content features containing information on the credibility of comments, specifically for: The initial data stream is subjected to feature obfuscation to obtain an obfuscated stream with controlled noise, wherein the feature obfuscation includes replacing synonyms and / or transposing characters; The initial data stream is then subjected to standard cleaning processing to obtain an explicit stream with random noise removed. The obfuscated stream is input into a preset sentiment model for sentiment analysis, and the sentiment features of the comment are output. The explicit stream is input into a preset adversarial detection model for fake pattern recognition, and the output is the fake probability feature of the comment. The fake pattern recognition includes extracting text length and / or information entropy. The sentiment features and the false probability features are fused to obtain fused content features that include information on the credibility of the comments.
[0099] In some embodiments, after performing feature fusion processing on the sentiment feature and the false probability feature to obtain fused content features containing information on the credibility of the comments, the second allocation unit 604 is specifically used for: Based on the fused content features, a credibility score representing the credibility of the comment is calculated; Each comment is filtered based on a preset credibility threshold and the credibility level to obtain the target comments; The products contained in the target reviews are considered as popular products, and a popular product set including multiple of the popular products is obtained. The items that intersect with the high-frequency attention set and the popular item set are identified as public goods; A third scanning result is obtained by increasing the collection resources invested in the public goods and decreasing the collection resources for other goods besides the public goods, wherein the increase and decrease correspond to equal collection resources.
[0100] In some embodiments, the public goods originate from multiple platforms. After determining that the goods intersecting the high-frequency attention set and the popular goods set are public goods, the second allocation unit 604 is specifically used for: Based on the multi-dimensional characteristics of public goods across multiple platforms, a cross-platform product association table composed of temporary entity links is generated. If at least two temporary entity links in the cross-platform product association table are ambiguous, the temporary entity links involved in the ambiguity will be marked as pending arbitration. The ambiguity refers to the fact that the product name dimension is consistent, but other dimensions are inconsistent. Remove the temporary entity links in the pending arbitration status from the cross-platform product association table; Aggregate the products linked to temporary entities awaiting arbitration to identify high-risk products and prompt for manual intervention.
[0101] In some embodiments, the multi-dimensional features include product name, price, and product brand. Based on the multi-dimensional features of public goods across multiple platforms, a cross-platform product association table composed of temporary entity links is generated. The second allocation unit 604 is specifically used for: From the product name, price, and brand of public goods on multiple platforms, the first letter of the pinyin, numerical features, and attribute roots are extracted to obtain multidimensional features of the products. If the product name, price, and brand dimensions are consistent across different platforms, and other inconsistent dimensions do not conflict, then a temporary entity link is established based on the consistent dimensions. Aggregate all the temporary entity links to obtain the cross-platform product association table.
[0102] In some embodiments, the sentiment feature is a sentiment vector matrix representing the intensity of sentiment polarity, the false probability feature is a false probability matrix, and a credibility level representing the credibility of the comment is calculated based on the fused content features. The second allocation unit 604 is specifically used for: The emotion vector matrix and the false probability matrix are input into a preset contradiction arbitrator; Calculate the difference between 1 and the false probability corresponding to the false probability matrix; If the maximum value of the emotional polarity intensity in the emotional vector matrix is greater than a preset first threshold, and the false probability is greater than a preset second threshold, then a preset penalty is determined to be applied to the difference, and the preset penalty weight is less than 1. The product of the difference and the preset penalty is calculated to obtain the credibility of each comment, which represents the credibility level of the comment.
[0103] In summary, the e-commerce user reputation analysis device 600 in this embodiment of the application, in response to a reputation analysis command, acquires the initial data stream of user reviews on the e-commerce platform and scans it to obtain a first scan result; based on the real-time attention in the initial data stream, it performs a first allocation processing of collection resources on the products in the high-frequency attention set of the first scan result to obtain a second scan result, wherein the first allocation processing does not change the overall collection resources consumed, and the higher the real-time attention, the more collection resources are invested; the content of the initial data stream is subjected to dual-stream parallel semantic analysis and false identification processing, and fused to obtain a fused content feature containing information on the credibility of the reviews, wherein the fused content feature integrates user sentiment and false reputation; based on the fused content feature, a second allocation processing of collection resources is performed on the products corresponding to the second scan result to obtain a third scan result, wherein the second allocation processing does not change the overall collection resources consumed, and the allocated collection resources are associated with the content feature; based on the third scan result, reputation analysis is performed to obtain the user reputation analysis result. This application not only supports dynamic data adaptation through data collection and allocation to ensure data real-time performance, but also integrates a novel analysis framework with semantic awareness and fraud detection mechanisms to ensure data stability without altering the overall data collection resources. This multifaceted approach enhances the credibility of the scanned data, thereby improving the accuracy of the analysis results when analyzing e-commerce user reviews.
[0104] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned e-commerce user reputation analysis device and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0105] The aforementioned e-commerce user reputation analysis device can be implemented as a computer program, which can, for example... Figure 3 It runs on the computer device shown.
[0106] Please see Figure 3 , Figure 3 This is a schematic block diagram of a computer device 700 provided in an embodiment of this application. The computer device 700 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0107] See Figure 3 The computer device 700 includes a processor 702, a memory, and a network interface 705 connected via a system bus 701. The memory may include a non-volatile storage medium 703 and internal memory 704.
[0108] The non-volatile storage medium 703 may store an operating system 7031 and a computer program 7032. The computer program 7032 includes program instructions that, when executed, cause the processor 702 to perform an e-commerce user reputation analysis method.
[0109] The processor 702 provides computing and control capabilities to support the operation of the entire computer device 700.
[0110] The internal memory 704 provides an environment for the execution of the computer program 7032 in the non-volatile storage medium 703. When the computer program 7032 is executed by the processor 702, the processor 702 can execute a method for analyzing e-commerce user reviews.
[0111] This network interface 705 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 700 to which the present application is applied. The specific computer device 700 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0112] The processor 702 is used to run the computer program 7032 stored in the memory to perform the following steps: In response to the word-of-mouth analysis command, the system obtains the initial data stream of user reviews on the e-commerce platform and scans it to obtain the first scan result. Based on the real-time attention in the initial data stream, the first allocation process of collection resources is performed on the products in the high-frequency attention set in the first scan result to obtain the second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The content of the initial data stream is subjected to dual-stream parallel semantic analysis and false information identification processing, and the resulting fused content features include information on the credibility of the comments. The fused content features incorporate user sentiment and information on false word-of-mouth. Based on the fused content features, a second allocation process of collection resources is performed on the product corresponding to the second scan result to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. Based on the third scan result, a reputation analysis is performed to obtain the user reputation analysis result.
[0113] It should be understood that in the embodiments of this application, the processor 702 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0114] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0115] Therefore, this application also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the following steps: In response to the word-of-mouth analysis command, the system obtains the initial data stream of user reviews on the e-commerce platform and scans it to obtain the first scan result. Based on the real-time attention in the initial data stream, the first allocation process of collection resources is performed on the products in the high-frequency attention set in the first scan result to obtain the second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The content of the initial data stream is subjected to dual-stream parallel semantic analysis and false information identification processing, and the resulting fused content features include information on the credibility of the comments. The fused content features incorporate user sentiment and information on false word-of-mouth. Based on the fused content features, a second allocation process of collection resources is performed on the product corresponding to the second scan result to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. Based on the third scan result, a reputation analysis is performed to obtain the user reputation analysis result.
[0116] The storage medium can be any computer-readable storage medium that can store program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0117] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0119] The steps in the methods of this application embodiment can be adjusted, merged, or deleted according to actual needs. The units in the apparatus of this application embodiment can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0121] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for analyzing user reviews in e-commerce, characterized in that, The methods for analyzing e-commerce user reviews include: In response to the word-of-mouth analysis command, the system obtains the initial data stream of user reviews on the e-commerce platform and scans it to obtain the first scan result. Based on the real-time attention in the initial data stream, the first allocation process of collection resources is performed on the products in the high-frequency attention set in the first scan result to obtain the second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. The initial data stream content is subjected to dual-stream parallel semantic analysis and false identification processing. The sentiment features obtained from the semantic analysis and the false probability features obtained from the false identification processing are fused to generate fused content features. The fused content features incorporate user sentiment and false word-of-mouth. Based on the fused content features, the credibility of the review can be calculated. Based on the fused content features, a second allocation process of collection resources is performed on the product corresponding to the second scan result to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. Based on the third scan result, a reputation analysis is performed to obtain the user reputation analysis result.
2. The method according to claim 1, characterized in that, The real-time attention level is the real-time comment increment. Based on the real-time attention level in the initial data stream, the products in the high-frequency attention set in the first scan result undergo a first allocation process of collection resources to obtain a second scan result, including: The initial data stream whose real-time comment increment exceeds a preset increment threshold is aggregated and divided in the semantic space to obtain a high-frequency attention set. A second scanning result is obtained by increasing the sampling frequency of products in the high-frequency attention set and decreasing the sampling frequency of products outside the high-frequency attention set, wherein the increased frequency and the decreased frequency are equal in terms of acquisition resources.
3. The method according to claim 1, characterized in that, The initial data stream content is subjected to dual-stream parallel semantic analysis and false positive detection processing. The sentiment features obtained from the semantic analysis are fused with the false positive probability features obtained from the false positive detection processing to generate fused content features, including: The initial data stream is subjected to feature obfuscation to obtain an obfuscated stream with controlled noise, wherein the feature obfuscation includes replacing synonyms and / or transposing characters; The initial data stream is then subjected to standard cleaning processing to obtain an explicit stream with random noise removed. The obfuscated stream is input into a preset sentiment model for sentiment analysis, and the sentiment features of the comment are output. The explicit stream is input into a preset adversarial detection model for fake pattern recognition, and the output is the fake probability feature of the comment. The fake pattern recognition includes extracting text length and / or information entropy. The sentiment features and the false probability features are fused to obtain fused content features that include information on the credibility of the comments.
4. The method according to claim 3, characterized in that, After performing feature fusion processing on the sentiment features and the false probability features to obtain fused content features containing information on the credibility of the comments, the following steps are included: Based on the fused content features, a credibility score representing the credibility of the comment is calculated; Each comment is filtered based on a preset credibility threshold and the credibility level to obtain the target comments; The products contained in the target reviews are considered as popular products, and a popular product set including multiple of the popular products is obtained. The items that intersect with the high-frequency attention set and the popular item set are identified as public goods; A third scanning result is obtained by increasing the collection resources invested in the public goods and decreasing the collection resources for other goods besides the public goods, wherein the increase and decrease correspond to equal collection resources.
5. The method according to claim 4, characterized in that, The public goods originate from multiple platforms. After determining that the goods intersect between the high-frequency attention set and the popular goods set are public goods, the following steps are included: Based on the multi-dimensional characteristics of public goods across multiple platforms, a cross-platform product association table composed of temporary entity links is generated. If at least two temporary entity links in the cross-platform product association table are ambiguous, the temporary entity links involved in the ambiguity will be marked as pending arbitration. The ambiguity refers to the fact that the product name dimension is consistent, but other dimensions are inconsistent. Remove the temporary entity links in the pending arbitration status from the cross-platform product association table; Aggregate the products linked to temporary entities awaiting arbitration to identify high-risk products and prompt for manual intervention.
6. The method according to claim 5, characterized in that, Multi-dimensional features include product name, price, and brand. Based on these multi-dimensional features of public goods across multiple platforms, a cross-platform product association table composed of temporary entity links is generated, including: From the product name, price, and brand of public goods on multiple platforms, the first letter of the pinyin, numerical features, and attribute roots are extracted to obtain multidimensional features of the products. If the product name, price, and brand dimensions are consistent across different platforms, and other inconsistent dimensions do not conflict, then a temporary entity link is established based on the consistent dimensions. Aggregate all the temporary entity links to obtain the cross-platform product association table.
7. The method according to claim 4, characterized in that, The sentiment feature is a sentiment vector matrix representing the intensity of sentiment polarity, and the false probability feature is a false probability matrix. Based on the fused content features, the credibility of the comment is calculated, including: The emotion vector matrix and the false probability matrix are input into a preset contradiction arbitrator; Calculate the difference between 1 and the false probability corresponding to the false probability matrix; If the maximum value of the emotional polarity intensity in the emotional vector matrix is greater than a preset first threshold, and the false probability is greater than a preset second threshold, then a preset penalty is determined to be applied to the difference, and the preset penalty weight is less than 1. The product of the difference and the preset penalty is calculated to obtain the credibility of each comment, which represents the credibility level of the comment.
8. A device for analyzing user reviews in e-commerce, characterized in that, The e-commerce user reputation analysis device includes: The acquisition unit is used to respond to the word-of-mouth analysis command, acquire the initial data stream of user reviews on the e-commerce platform, and scan it to obtain the first scan result; The first allocation unit is used to perform a first allocation process on the high-frequency attention set of products in the first scan result based on the real-time attention in the initial data stream, and to obtain a second scan result. The first allocation process does not change the overall collection resources consumed. The higher the real-time attention, the more collection resources are invested. A dual-stream parallel processing unit is used to perform dual-stream parallel semantic analysis and false identification processing on the content of the initial data stream. The sentiment features obtained from the semantic analysis are fused with the false probability features obtained from the false identification processing to generate fused content features. The fused content features incorporate user sentiment and false word-of-mouth. Based on the fused content features, the credibility of the review can be calculated. The second allocation unit is used to perform a second allocation process of collection resources on the product corresponding to the second scan result based on the fused content features to obtain a third scan result. The second allocation process does not change the overall collection resources consumed, and the allocated collection resources are associated with the content features. The reputation analysis unit is used to perform reputation analysis based on the third scan result to obtain user reputation analysis results.
9. A computer device for analyzing e-commerce user reviews, characterized in that, The method includes a memory, a processor, and an e-commerce user reputation analysis program stored in the memory and executable on the processor. The processor executes the e-commerce user reputation analysis program to implement the steps of the e-commerce user reputation analysis method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a program for implementing an e-commerce user reputation analysis method, which is executed by a processor to implement the steps of the e-commerce user reputation analysis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Comment emotion stream data simulation generation method
CN117494885A
False comment perception method based on semantic-emotion double-flow attention fusion
CN120892568A