Method for carrying out qualitative and quantitative analysis on mass e-commerce data by utilizing large model

Through large models and self-built label models, the problem of low efficiency of massive e-commerce data analysis is solved, efficient and accurate data processing and intuitive user interaction are achieved, and enterprises are supported to operate in a refined manner.

CN120541118APending Publication Date: 2025-08-26SHENZHEN SKIEER INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510698437.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

There is a lack of effective methods in the existing technology to conduct flexible and comprehensive qualitative and quantitative analysis of massive e-commerce data, resulting in high consumption of computing resources and users need to recalculate each time they analyze the needs of demand change, which is difficult to meet the needs of enterprises for refined operations.

Method used

The large model is used to semantic recognition of e-commerce data, and the self-built tag model is used to summarize tags. Users select analysis instructions through the cloud listening message page, call the big model interface, return qualitative and quantitative analysis results, and visual output and database management of the data.

Benefits of technology

It realizes fast and accurate e-commerce data analysis, improves data processing efficiency and quality, intuitive and convenient user interaction, supports corporate decision-making, and facilitates data entry into the database for query and management, and taps potential value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541118A_ABST
    Figure CN120541118A_ABST
Patent Text Reader

Abstract

The invention provides a method for carrying out qualitative and quantitative analysis on mass e-commerce data by utilizing a large model, and relates to the technical field of e-commerce data analysis. According to the method and the device, the e-commerce data is collected, and semantic recognition and label induction are performed by utilizing the sentiment analysis large model and the self-built label model, so that mass e-commerce data can be quickly and accurately analyzed and classified, and the efficiency and the quality of data processing are improved; a user can screen messages according to requirements through a cloud message listening page, open a page interacting with a large model, select qualitative and quantitative analysis instructions, interact with the large model in an intuitive and convenient manner, and obtain a required data analysis result; the data and the corresponding cue words are assembled into the request object to call the large model, so that the powerful analysis capability of the large model can be fully utilized, more accurate qualitative and quantitative analysis results are returned, and more powerful support is provided for enterprise decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of e-commerce data analysis, and in particular to a method for performing qualitative and quantitative analysis on massive e-commerce data using a large model. Background Art

[0002] In the customer experience management industry, unlocking the value of e-commerce data is key to ensuring companies have a deeper understanding of user profiles and behavioral motivations. Flexible and comprehensive qualitative and quantitative analysis can help companies deeply interpret data after the traffic dividend fades, and is a good helper for companies to shift from extensive growth to refined operations. Currently, there are no good solutions on the market for qualitative and quantitative analysis of massive e-commerce data. The difficulty lies in the complex data structure, which requires a lot of time to semantically understand and label the data. It also consumes a lot of computing resources, and users' focus on the data varies each time they use it. Every time the analysis needs change, the data must be recalculated. Therefore, we have made improvements to this and proposed a method for using large models to perform qualitative and quantitative analysis of massive e-commerce data. Summary of the Invention

[0003] The present invention provides a method for qualitative and quantitative analysis of massive e-commerce data using a large model, including: Collect e-commerce data; use sentiment analysis models to perform semantic recognition on the collected e-commerce data; build a self-built labeling model, use the organized label system as input to the model, and then perform label induction and recognition on the data, label it, and store the labeled data in the database; After entering the homepage, users can filter messages based on their needs through the Cloud Listening message page; open the page that interacts with the large model, select qualitative and quantitative analysis instructions, and call the Cloud Listening data query interface to query the stored data; Assemble the data and the corresponding prompt words into a request object and call the large model; the large model returns the results of qualitative and quantitative analysis, and finally ends.

[0004] As the preferred technical solution of this application, the collection of e-commerce data includes: real-time collection of user comment data, transaction record data, and product details page data of the e-commerce platform through the API interface.

[0005] As a preferred technical solution of this application, the collection of e-commerce data also includes: synchronizing user behavior log data in the e-commerce database through a scheduled task.

[0006] As the preferred technical solution of this application, the sentiment analysis large model is a pre-trained model fine-tuned based on the e-commerce field corpus, including BERT-e-commerce version, GPT-4-e-commerce fine-tuning model or Llama-2-sentiment analysis dedicated model.

[0007] As the preferred technical solution of this application, the semantic recognition includes sentiment polarity classification of the text, extraction of key sentiment words and sentiment intensity scoring.

[0008] As the preferred technical solution of this application, the process of constructing the self-built label model includes: S1: Define the initial labeling system based on e-commerce business needs. When defining the initial labeling system, organize e-commerce business experts and data analysts to conduct discussions to ensure that the labeling system can cover all aspects of the e-commerce business. S2: Through a semi-supervised learning method, a label classifier is trained using a small amount of manually labeled data to automatically expand the labels of unlabeled data. The expansion accuracy of the label classifier is evaluated using the following formula: A=(T / V)×100%; Where: A is the label expansion accuracy; T is the number of correct expanded labels verified by humans; V is the total number of labels verified by humans; During the evaluation process, the cross-validation method is used to divide the dataset into multiple subsets, which are used as training sets and validation sets in turn; S3: Use a large model to resolve ambiguity in label conflict data and optimize the labeling system.

[0009] As the preferred technical solution of this application, the filtering conditions for the user to filter messages through the Yunting message page include: time range, product category, sentiment score threshold, and tag combination; in terms of filtering condition setting, two methods of fuzzy matching and exact matching are provided, and users can choose according to actual needs; at the same time, users are supported to save commonly used filtering condition combinations.

[0010] As the preferred technical solution of this application, when assembling the data and the corresponding prompt words into a request object, the sensitive data is desensitized, including coding the user ID, hiding some digits of the mobile phone number, and blurring the address information; and the request object is encrypted and transmitted to the large model interface through the AES encryption algorithm. During the encryption process, a secure key management mechanism is adopted and the encryption key is replaced regularly; at the same time, the integrity of the encrypted request object is checked.

[0011] As the preferred technical solution of this application, after the large model returns the analysis results, it also includes: Visualize the results into various types of charts. During the visualization process, choose the appropriate chart type based on the data characteristics and analysis requirements. Supports users to perform interactive operations on visual charts, including zooming, panning, and filtering; Support users to export reports in Excel or PDF format.

[0012] As the preferred technical solution of this application, different storage partitions and indexing strategies are set for e-commerce data from different sources and types during data warehousing. For frequently accessed data, it is stored in a high-performance solid-state hard disk partition, and a multi-level index is established, including one based on product ID, timestamp, and sentiment level. For historical archived data, it is stored in a large-capacity mechanical hard disk partition, and compressed storage is used.

[0013] Compared with the prior art, the present invention has the following beneficial effects: In the scheme of this application: 1. This application collects e-commerce data and uses a large sentiment analysis model and a self-built labeling model for semantic recognition and label induction. This allows for rapid and accurate analysis and classification of massive e-commerce data, improving the efficiency and quality of data processing. 2. Users can filter messages based on their needs through the Cloud Listening Messages page, open the page for interacting with the big model, select qualitative and quantitative analysis instructions, and interact with the big model in an intuitive and convenient way to obtain the required data analysis results; 3. Assembling the data and corresponding prompt words into a request object to call the big model can fully utilize the powerful analytical capabilities of the big model to return more accurate qualitative and quantitative analysis results, providing stronger support for corporate decision-making; 4. The analyzed data is stored in the database to facilitate subsequent query and management. Enterprises can call historical data for analysis and comparison at any time to explore the potential value behind the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of the method provided in this application for qualitative and quantitative analysis of massive e-commerce data using a large model; Figure 2 A schematic diagram of the message interface provided for this application; Figure 3 Schematic diagram of the screening panel provided for this application; Figure 4 A schematic diagram of the reply to the original sound of the summary message provided for this application; Figure 5 A schematic diagram of the input box provided for this application; Figure 6 Large model code diagram provided for this application; Figure 7 The Cloud Listening Assistant page display image provided for this application; Figure 8 Another display image of the Cloud Listening Assistant page provided for this application. DETAILED DESCRIPTION

[0015] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0016] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions therein may be combined with each other.

[0017] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0018] Example 1, please refer to Figure 1 ,include: Collect e-commerce data; use sentiment analysis models to perform semantic recognition on the collected e-commerce data; build a custom labeling model, use the organized label system as input to the model, then perform label induction and identification on the data, label it, and store the labeled data; achieve structured processing and deep semantic analysis of e-commerce data, and provide a standardized data foundation for subsequent analysis; After entering the homepage, users can filter messages according to their needs through the Yunting message page; open the page that interacts with the large model, select qualitative and quantitative analysis instructions, call the Yunting data query interface, and query the stored data; the visual filtering page and standardized instruction calling process enable non-technical users to quickly locate target data, avoid the inefficiency of manual retrieval, and shorten the data analysis cycle. The data and corresponding prompt words are assembled into a request object, and the large model is called; the large model returns the results of qualitative and quantitative analysis, and finally ends; the natural language processing capabilities of the large model can deeply analyze data associations and automatically generate conclusions such as trend analysis and problem attribution, which is more efficient and comprehensive than traditional statistical methods.

[0019] Furthermore, the collection of e-commerce data includes: real-time collection of user comment data, transaction record data, and product details page data from e-commerce platforms through API interfaces; real-time collection can promptly reflect market dynamics (such as new product reviews and sudden public opinion), helping companies quickly respond to changes in consumer demand and avoid decision-making biases caused by lagging analysis.

[0020] Furthermore, collecting e-commerce data also includes: synchronizing user behavior log data (including click, add to cart, favorite, and return records) in the e-commerce database through scheduled tasks; scheduled synchronization of log data can build user behavior portraits, reveal potential consumption habits, and provide long-term data support for product optimization and marketing strategy formulation.

[0021] Furthermore, the large sentiment analysis model is a pre-trained model fine-tuned based on the e-commerce field corpus, including the BERT-e-commerce version, the GPT-4-e-commerce fine-tuned model or the Llama-2-sentiment analysis-specific model; the model fine-tuned based on the e-commerce corpus can accurately identify industry-specific words (such as "large size" and "high cost performance"), avoiding the semantic misjudgment of the general model, and improving the accuracy of sentiment classification.

[0022] Furthermore, semantic recognition includes the classification of the text's sentiment polarity (positive / neutral / negative), the extraction of key sentiment words (such as "satisfactory", "poor quality", "slow logistics") and sentiment intensity scoring (quantitative value of 0-1 points); sentiment polarity classification, keyword extraction and intensity scoring form a multi-dimensional analysis framework. For example, "negative + slow logistics + 0.2 points" can accurately locate service shortcomings, facilitating targeted improvements by enterprises.

[0023] Furthermore, the process of building a self-built label model includes: S1: Define an initial labeling system based on e-commerce business needs (including product category labels, user behavior labels, sentiment level labels, and problem type labels (such as "quality issues" and "service issues"). When defining the initial labeling system, organize e-commerce business experts and data analysts to conduct discussions to ensure that the labeling system covers all aspects of the e-commerce business. S2: Through a semi-supervised learning method, a label classifier is trained using a small amount of manually labeled data to automatically expand the labels of unlabeled data. The expansion accuracy of the label classifier is evaluated using the following formula: A=(T / V)×100%; Where: A is the label extension accuracy; T is the number of correctly extended labels verified by humans; V is the total number of labels verified by humans (V ≥ 100 to ensure statistical significance); During the evaluation process, the cross-validation method is used to divide the dataset into multiple subsets, which are used as training sets and validation sets in turn; S3: Use a large model to resolve ambiguity in label conflict data (for example, a comment triggers both the "poor quality" and "fast logistics" labels) and optimize the labeling system. The large model resolves label conflicts, avoids logical contradictions, and improves the rigor of the labeling system.

[0024] Furthermore, the filtering conditions for users to filter messages through the Yunting message page include: time range (such as the past 7 days / 30 days), product category (such as beauty / 3C), sentiment score threshold (such as only filtering negative comments with a sentiment score of less than 0.3), and label combination (such as "return record + quality problem label"); in terms of setting filtering conditions, two methods are provided: fuzzy matching and exact matching, and users can choose according to actual needs; at the same time, users are supported to save commonly used filtering condition combinations for quick use next time; filtering conditions such as time, category, sentiment score, and label combination can accurately filter invalid information, and the fuzzy / exact matching mode adapts to different scenarios. Saving commonly used combinations reduces repeated operations and improves query efficiency.

[0025] Furthermore, when assembling the data and the corresponding prompt words into a request object, the sensitive data is desensitized, including coding the user ID, hiding some digits of the mobile phone number, and blurring the address information; and the request object is encrypted using the AES encryption algorithm and transmitted to the large model interface. During the encryption process, a secure key management mechanism is adopted, and the encryption key is replaced regularly to ensure the security of data transmission; at the same time, the encrypted request object is integrity checked to prevent data tampering during transmission.

[0026] Furthermore, after the large model returns the analysis results, it also includes: Visualize the results into various chart types (such as sentiment distribution pie charts, keyword word clouds, trend line charts, bar charts, etc.). During the visualization process, select the appropriate chart type based on the data characteristics and analysis requirements, and customize the chart's color, font, and annotations. The chart type adapts to different data characteristics, making the conclusions intuitive and easy to understand. Supports interactive operations on visualization charts, including zooming, panning, and filtering. Zooming, panning, and filtering allow users to dynamically explore data details, while the report export function facilitates cross-departmental collaboration and improves decision-making efficiency. Support users to export reports in Excel or PDF format.

[0027] Furthermore, during data warehousing, different storage partitions and indexing strategies are set for e-commerce data from different sources and types. Frequently accessed data is stored in high-performance solid-state drive partitions, and multi-level indexes are established, including those based on product ID, timestamp, and sentiment level, to speed up data queries. Historical archived data is stored in large-capacity mechanical hard drive partitions, using compressed storage to reduce storage costs. High-frequency data is stored in SSDs and multi-level indexes are established, which can shorten query response time to milliseconds to meet real-time analysis needs; historical data is compressed and stored on mechanical hard drives to reduce storage costs while ensuring data traceability and meeting compliance archiving requirements.

[0028] Example 2: Log in to the Yunting system; select the message page, such as Figure 2 As shown; select the conditions you want to filter from the filter panel, and then click Query, as shown Figure 3 As shown; Click on the Cloud Listening Assistant to call up the window for interacting with the large model and send a message summary instruction, such as Figure 4 As shown; Enter the qualitative analysis and quantitative analysis request in the input box and click Send. Figure 5 As shown; Query e-commerce data messages according to the request, assemble them together with the prompt words and send them to the big model, such as Figure 6 As shown; The qualitative and quantitative analysis results are displayed on the Cloud Listening Assistant page. Figure 7 and Figure 8 shown.

[0029] In the present invention, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection, or communication; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0030] Obviously, the embodiments described above are only some embodiments of the present invention, rather than all embodiments. The preferred embodiments of the present invention are given in the accompanying drawings, but they do not limit the patent scope of the present invention. The present invention can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive. Although the present invention has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present invention specification and drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present invention.

Claims

1. A method for qualitative and quantitative analysis of massive e-commerce data using a large model, characterized by: include: Collect e-commerce data; Use sentiment analysis models to perform semantic recognition on collected e-commerce data; Build your own labeling model, use the organized labeling system as the input to the large model, then perform label induction and identification on the data, label it, and store the labeled data in the database; After entering the homepage, users can filter messages according to their needs through the Cloud Listening Messages page; Open the page for interacting with the large model, select the qualitative and quantitative analysis instructions, call the Yunting data query interface, and query the stored data; Assemble the data and the corresponding prompt words into a request object and call the large model; the large model returns the results of qualitative and quantitative analysis, and finally ends.

2. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The e-commerce data collection includes: collecting user review data, transaction record data, and product detail page data of the e-commerce platform in real time through the API interface.

3. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The collecting of e-commerce data further includes: synchronizing user behavior log data in the e-commerce database through a scheduled task.

4. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The sentiment analysis model is a pre-trained model fine-tuned based on the e-commerce corpus, including the BERT-e-commerce version, the GPT-4-e-commerce fine-tuned model, or the Llama-2-sentiment analysis-specific model.

5. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The semantic recognition includes sentiment polarity classification of text, extraction of key sentiment words and sentiment intensity scoring.

6. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The process of building the self-built label model includes: S1: Define the initial labeling system based on e-commerce business needs. When defining the initial labeling system, organize e-commerce business experts and data analysts to conduct discussions to ensure that the labeling system can cover all aspects of the e-commerce business. S2: Through a semi-supervised learning method, a label classifier is trained using a small amount of manually labeled data to automatically expand the labels of unlabeled data. The expansion accuracy of the label classifier is evaluated using the following formula: A=(T / V)×100%; Where: A is the label expansion accuracy; T is the number of correct expanded labels verified by humans; V is the total number of labels verified by humans; During the evaluation process, the cross-validation method is used to divide the dataset into multiple subsets, which are used as training sets and validation sets in turn; S3: Use a large model to resolve ambiguity in label conflict data and optimize the labeling system.

7. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: The filtering conditions for users to filter messages through the Yunting message page include: time range, product category, sentiment score threshold, and tag combination; in terms of filtering condition setting, two methods of fuzzy matching and exact matching are provided, and users can choose according to actual needs; at the same time, users are supported to save commonly used filtering condition combinations.

8. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: When assembling the data and the corresponding prompt words into a request object, the sensitive data is desensitized, including coding the user ID, hiding some digits of the mobile phone number, and blurring the address information; the request object is encrypted using the AES encryption algorithm and transmitted to the large model interface. During the encryption process, a secure key management mechanism is adopted and the encryption key is replaced regularly; at the same time, the integrity of the encrypted request object is checked.

9. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: After the large model returns the analysis results, it also includes: Visualize the results into various types of charts. During the visualization process, choose the appropriate chart type based on the data characteristics and analysis requirements. Supports users to perform interactive operations on visual charts, including zooming, panning, and filtering; Support users to export reports in Excel or PDF format.

10. The method for qualitative and quantitative analysis of massive e-commerce data using a large model according to claim 1, characterized in that: During the data warehousing process, different storage partitions and indexing strategies are set for e-commerce data from different sources and types. Frequently accessed data is stored in high-performance solid-state drive partitions, and multi-level indexes are established, including those based on product ID, timestamp, and sentiment level. Historical archived data is stored in large-capacity mechanical hard drive partitions using compressed storage.

Citation Information

Patent Citations

  • Data information label processing method of large language model

    CN117453921A

  • Real-time data query method and system based on multi-level index

    CN118673043A

  • Adaptive commodity classification method based on AI identification of large model

    CN119884380A

  • Refined user viewpoint extraction and analysis method based on large model

    CN119961423A

  • Dify platform-based intelligent big data summarization system for dialogue interaction

    CN120030150A