Tobacco retail shop-oriented market report automatic generation method

By constructing data sets in the tobacco retail field and fine-tuning large language models using LoRA technology, and combining prompt words and CoT thinking chains for content generation, the problem that the existing technology cannot be directly applied to the tobacco retail field is solved, and the effect of efficient and automatic generation of market reports is achieved.

CN120106019APending Publication Date: 2025-06-06SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311635951.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art cannot be directly applied in the case of professional terms involving the field of tobacco retail, and the newly generated knowledge cannot be processed in a timely manner, resulting in inefficient market report generation.

Method used

By building new data sets, we can enhance the processing ability of professional data embedded in the tobacco retail field, use fine-tuning technology based on specific data sets to improve the vectorization ability of sentence embedding, and combine LoRA technology to fine-tune large language models, and use prompt words and CoT thinking chains to generate inference content.

Benefits of technology

It has achieved efficient and automatic generation of market reports including news summary, dialogue summary, product review classification, product review emotion recognition and statistical data in the tobacco retail field, improving the efficiency and accuracy of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106019A_ABST
    Figure CN120106019A_ABST
Patent Text Reader

Abstract

A tobacco retail store-oriented market report automatic generation method comprises the following steps of: constructing and training a statement embedding model, carrying out LoRA-based large language model fine adjustment on input data, and carrying out inference content generation and report filling according to Prompt and CoT thinking chain-based inference content generation. According to the method, the processing capability of embedding statements into professional data in the tobacco retail field is enhanced by constructing a new data set. Through fine tuning and statement embedding based on a specific data set, the text can be vectorized better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of natural language processing, in particular to an automatic generation method of market reports for tobacco retail stores. Background Art

[0002] Market reports can integrate multiple traditional natural language processing tasks. However, traditional technologies can only handle a single task and cannot integrate multiple tasks. Large language models can achieve various tasks that originally required targeted training or fine-tuning. It can be applied to tobacco retail stores and help relevant personnel quickly understand market information. Summary of the invention

[0003] In view of the shortcomings of the existing technology that cannot be directly applied when involving professional terminology in the field and cannot process the newly generated knowledge in a timely manner, this paper proposes a method for automatically generating market reports for tobacco retail stores, which enhances the processing ability of sentence embedding in the professional data of the tobacco retail field by constructing a new data set. Through fine-tuning based on a specific data set, sentence embedding can better vectorize text.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to an automatic generation method of market reports for tobacco retail stores. After constructing and training a sentence embedding model, the input data is fine-tuned based on a large language model of LoRA, and then the inference-based content generation and report filling based on prompt words and CoT thinking chains are performed. Technical Effects

[0006] The present invention constructs a data set for the tobacco retail field and trains a sentence embedding model; constructs a data set for the tobacco retail field and uses the LoRA method to fine-tune the large language model; and implements a report automatic generation technology based on sentence embedding and the large language model. Compared with the prior art, the present invention automatically generates reports for the tobacco retail field that include news summaries, conversation summaries, product review classification, product review sentiment recognition, and various statistical data. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is a flow chart of the present invention;

[0008] Figure 2 It is a schematic diagram of the embodiment system;

[0009] Figure 3 It is a schematic diagram of the effect of the embodiment. DETAILED DESCRIPTION

[0010] like Figure 2As shown, the present invention relates to an automatic market report generation system for tobacco retail stores, including: a data processing layer, a data storage layer and a content generation layer, wherein: the data processing layer processes the relevant text data required for report content generation, and finally generates a text block of a certain size; the data storage layer converts the text block into a low-dimensional semantic vector through a sentence embedding model, and stores the vector in a vector database; the content generation layer matches and retrieves relevant text blocks from the vector database according to the content generation target, and provides them to the large language model in the form of prompt words, so that the large language model can generate content.

[0011] like Figure 1 As shown, this embodiment involves a method for automatically generating a market report for tobacco retail stores. After constructing and training a sentence embedding model, the input data is fine-tuned based on a large language model of LoRA, and then the inference-based content generation and report filling based on prompt words (Prompt) and CoT thinking chain are performed. The specific steps include:

[0012] Step 1) Cut the database corpus into chunks, and then give a certain number of chunks as context when a specific application calls LLM.

[0013] Step 2) Fine-tune the large language model using the tobacco retail corpus constructed by the present invention: add the professional terms in the tobacco field to the Tokenizer of the model, and fine-tune the model to adapt to the new Tokenizer. Since the number of parameters of the large language model is relatively large, the video memory capacity of the personal computer cannot support the fine-tuning of all parameters of the large model. Therefore, the present invention will use LoRA technology for fine-tuning.

[0014] Step 3) Generate a report using the fine-tuned large language model, including:

[0015] 3.1) Construct a series of prompt words to complete classification, clustering, summary generation, and sentiment analysis tasks, and finally output modules such as cigarette brand investigation and analysis, consumer behavior data collection, investigation and analysis, and price factors that consumers pay attention to when shopping.

[0016] 3.2) Input the prompt words and related data into the large language model and improve the model's reasoning ability through CoT thinking chain;

[0017] 3.3) After the model outputs the results of each module, it is integrated through relevant interfaces to form a report on the tobacco retail industry.

[0018] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.

Claims

1. A method for automatically generating market reports for tobacco retail stores. It is characterized in that After building and training the sentence embedding model, the input data is fine-tuned based on the LoRA large language model, and then the inference-based content generation and report filling based on prompts and CoT thinking chains are performed.

2. The method for automatically generating a market report for tobacco retail stores according to claim 1, Its characteristics are that include: Step 1) Cut the database corpus into chunks, and then give a certain number of chunks as context when a specific application calls LLM; Step 2) Fine-tune the large language model using the tobacco retail domain corpus: Add tobacco domain professional terms to the model’s Tokenizer and fine-tune the model using LoRA technology to adapt to the new Tokenizer; Step 3) Generate a report using the fine-tuned large language model, including: 3.1) Construct a series of prompt words to complete classification, clustering, summary generation, and sentiment analysis tasks, and finally output cigarette brand investigation and analysis, consumer behavior data collection investigation and analysis, and price factors that consumers pay attention to when shopping and analysis modules; 3.2) Input the prompt words and related data into the large language model and improve the model's reasoning ability through CoT thinking chain; 3.3) After the model outputs the results of each module, it is integrated through relevant interfaces to form a report on the tobacco retail industry.

3. A market report automatic generation system for tobacco retail stores implementing the method of claim 1 or 2, It is characterized in that include: Data processing layer, data storage layer and content generation layer, wherein: the data processing layer processes the relevant text data required for report content generation and generates text blocks; The data storage layer converts text blocks into low-dimensional semantic vectors through a sentence embedding model and stores the vectors in a vector database. The content generation layer matches and retrieves relevant text blocks from the vector database according to the content generation goals and provides them to the large language model in the form of prompt words, allowing the large language model to generate content.