Intelligent number asking method suitable for sales data query scene

By employing intelligent data query methods, combined with voice interaction and multimodal analysis, the problem of low data query efficiency in cigarette sales scenarios has been solved, achieving efficient and accurate data processing and display, and improving business intelligence and scientific decision-making.

CN121579639APending Publication Date: 2026-02-27SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511750184.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

The existing BI-driven model suffers from problems such as lack of data skills among business personnel, resource constraints in data teams, and information lag for decision-makers when dealing with the complexity and variability of cigarette sales scenarios. This results in low data query efficiency and an inability to respond to business needs in a timely manner.

Method used

Employing intelligent questioning methods and combining technologies such as speech synthesis, speech recognition, and semantic understanding, it achieves data analysis of natural language interaction. Through multimodal analysis and pluggable model access, it supports various business scenarios and provides platform-based services, including semantic parsing, intent classification, business scenario modeling, data center services, and knowledge base retrieval, enabling intelligent data processing and display.

Benefits of technology

It improved data query efficiency, reduced manual time from 3 hours to 5 minutes, increased the accuracy of query intent recognition to over 90%, and improved customer classification and labeling efficiency by 80%, thus realizing business intelligence and scientific decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579639A_ABST
    Figure CN121579639A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent number asking method suitable for a sales data query scene, which belongs to the technical field of data processing and comprises the following steps of: (1) inputting a natural language; (2) dialogue type number asking; (3) accessing a plug-in model; (4) carrying out multi-modal analysis; and (5) carrying out platform service. According to the intelligent data asking system architecture, an industry knowledge graph is used as a cognitive footstone, a large language model is used as an interaction engine, and dynamic data association is used as execution blood vessel three-in-one. The core pain point of'water and soil disability 'in vertical industry application of a general AI model is solved, the method is a key for really changing the AI from'chatting' to'drying ', and unprecedented agility and accuracy are provided for data-driven decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data processing technology and artificial intelligence technology, and in particular to an intelligent data query method suitable for sales data query scenarios. Background Technology

[0002] The existing BI-driven model reveals the following three core flaws when dealing with the complexity and variability of cigarette sales scenarios:

[0003] (1) The "digital divide" among business personnel

[0004] Marketing managers, brand managers, and other frontline business personnel, while possessing the deepest understanding of the business, generally lack data skills such as SQL. Faced with complex data dashboards, they often only see "what" but struggle to delve into "why." When it's necessary to analyze the impact of a specific marketing campaign on the surrounding market on an ad-hoc basis, they cannot independently and quickly find answers from massive amounts of data, relying heavily on the support of the data department.

[0005] The root cause lies in the tool-oriented rather than problem-oriented approach. Traditional BI tools require users to adapt to the tool's logic (drag and drop dimensions, configure metrics), rather than the tool understanding the user's natural language questions. This paradigm shift has become the biggest bottleneck to business agility in today's increasingly fragmented and immediacy-driven data demands.

[0006] (2) The "Resource Trap" for Data Teams

[0007] Data analysts and IT staff are bogged down in repetitive tasks of "data extraction - table creation - report writing." Statistics show that over 80% of their time is spent on data cleaning, caliber alignment, and report development, leaving very little time for in-depth analysis and insightful analysis. Faced with a constant stream of ad-hoc and exploratory requests from business departments, data teams are often overwhelmed and their responses are severely delayed.

[0008] Traditional reports and dashboards are "static," serving known and routine issues. Business decision-making, however, demands are "dynamic" and unpredictable. This fundamental supply-demand imbalance leads to a rigid data supply system that cannot keep pace with market changes.

[0009] (3) The "information lag" of the decision-making level

[0010] In decision-making scenarios such as business analysis meetings, a follow-up question from leadership ("Compared to the same period last year, how does the 'Golden Leaf' brand perform in out-of-province markets in Province A and Province B?") often fails to elicit an immediate and accurate answer. The data team may need to spend hours or even days after the meeting to provide a report, missing the optimal opportunity for decision-making.

[0011] The system lacks semantic understanding capabilities for human-computer dialogue. Existing systems cannot understand natural language questions from leadership, which contain industry jargon and complex logic. Data queries must be mediated by technical personnel, which is not only time-consuming but also degrades the accuracy and timeliness of information during transmission, causing data-driven decision-making to fail at critical moments. Summary of the Invention

[0012] To address the above technical problems, this invention provides an intelligent data query method suitable for sales data query scenarios.

[0013] The technical solution of this invention is:

[0014] A smart data query method suitable for sales data query scenarios, including

[0015] (1) Natural language input

[0016] By combining multiple intelligent technologies such as speech synthesis, speech recognition, semantic understanding, and semantic disambiguation, it enables voice-based conversational question analysis, making analysis accessible to everyone.

[0017] (2) Dialogue-style questioning

[0018] Data analysis based on natural language interaction via chat conversations intelligently identifies user intent and uses multi-turn question-and-answer sessions. This makes data acquisition and analysis easy for everyone, enabling everyone to use data.

[0019] (3) Plug-in model integration

[0020] It supports adaptation to multiple underlying models, allowing you to choose the appropriate model based on your specific business needs and scale.

[0021] (4) Multimodal analysis

[0022] Based on multiple basic large models and multiple business-level data mining small models, a data intelligent question analysis model engine platform is built, which supports single chart question counting, dashboard-type question counting, and prediction / classification question counting to meet a variety of business scenarios.

[0023] (5) Platform-based services

[0024] To provide accurate, reliable, and intervention-enabled analytical capabilities, the platform offers engineering services such as Prompt service management, table metadata synonym fine-tuning, user intervention, and chart switching.

[0025] Furthermore,

[0026] It includes the following underlying logic modules:

[0027] (1) Semantic parsing

[0028] Leveraging the natural language understanding capabilities of the 14B large model, user questions are rewritten to optimize their structure and expression, making them more suitable for the model's processing requirements; multimodal recall is performed based on the rewritten questions; and embedding and reordering models are used to process and sort the recalled information.

[0029] (2) Intention Classification

[0030] By using extraction tools to extract dimensions and metrics, we can more accurately understand user needs and provide key data support for subsequent business scenario modeling, data center services, and knowledge base retrieval.

[0031] (3) Business Scenario Modeling text2API

[0032] The extracted text undergoes API schema checks, an HTTP request is sent to query the data, and a visualization is returned. If the conditions are not met, clarification information can be output or the process can be reverted to the data center service text2Metric.

[0033] (4) Data Center Service text2SQL

[0034] We conduct in-depth data mining targeting industry characteristics, achieving full-domain data penetration through terminology extraction and data source identification; we use large models for problem planning, and through repeated iterations with LLM plugins and examples, we generate executable code and dynamic prompts based on business knowledge; finally, we use large models for reasoning analysis to achieve streaming data output.

[0035] (5) Knowledge base retrieval text2KB

[0036] By conducting in-depth data retrieval using data metrics, retrieving relevant information using a knowledge base, and generating the final result using a large model;

[0037] (6) Graphic and text display

[0038] Based on the data, it can intelligently select the type of bar chart, pie chart, or table for display; and intelligently select the most appropriate display type based on the specific data characteristics and analysis needs.

[0039] Furthermore,

[0040] (1) First, define the data analysis API interface, design and implement it fully according to the business situation, and form the API user manual;

[0041] (2) The user inputs natural language, and the LLM transforms the user's input into a call to the API tool, including the name of the API and the extracted parameters;

[0042] (3) Call the specified API based on the LLM response to obtain the returned data;

[0043] (4) As needed, the returned data is appended to the user input and then given to the LLM again, which outputs the final analysis results to the customer.

[0044] Furthermore,

[0045] 1. Application Scenario Implementation Path

[0046] Data collection: Deploy data capture tools at various sales terminals, trading platforms, and price monitoring systems, or connect to big data centers to collect market operation data in real time;

[0047] Data processing: Cleaning, organizing and processing the collected data, extracting key indicators, and analyzing trends in purchasing, sales and inventory, market share and price fluctuations through data mining techniques;

[0048] Intelligent analysis: Utilizing large-scale model analysis and reasoning capabilities, it conducts in-depth analysis of market data to uncover potential problems and business opportunities;

[0049] Speech synthesis: Converts processed data reports into speech using natural language generation technology, making the report content more similar to human language expression. Supports several speech styles and pronunciations to meet the preferences of different users.

[0050] 2. Implementation Path of Intelligent Data Query

[0051] Database construction: Based on the sales data, the database and table structure were rebuilt to align with the understanding of the larger model;

[0052] Text2SQL model training: Collect relevant natural language query samples and corresponding SQL statements, as well as industry basic knowledge, fine-tune the basic model, and evaluate the fine-tuning effect;

[0053] Intelligent agent development: First, semantic understanding of the user's query query is performed to extract key information and generate corresponding SQL query statements; then, the SQL is sent to the database for execution to obtain query results; finally, the query results returned by the database are converted into natural language and replied to marketing personnel through an AI assistant; testing is conducted to ensure accurate return of results in various query scenarios; based on test feedback and user feedback, the accuracy and response speed of the text2sql model are continuously optimized.

[0054] 3. Intelligent Analysis Implementation Path

[0055] Build an API knowledge base: Organize the API interfaces related to sales data analysis, and label the functions implemented by the APIs and their input and output parameters; if some analysis requirements do not have existing API interfaces, then API interfaces need to be developed.

[0056] Large model fine-tuning: Collect industry-specific knowledge related to API parameters, perform professional-level fine-tuning of the original basic model, and comprehensively evaluate the performance after adjustment;

[0057] Intelligent Agent Development: Utilizing large-scale models to analyze user queries, through word segmentation, part-of-speech tagging, and semantic understanding, key information in instructions is identified to clarify user intent; based on user intent, the API knowledge base is matched to select the appropriate API, and the parameters mentioned by the user are parsed; the corresponding API is called to obtain data results; large-scale models are used to apply big data analytics algorithms for in-depth data mining and analysis; furthermore, large-scale models leverage visualization technology to transform analysis results into intuitive and easy-to-understand market analysis reports and dynamic data dashboards, using a combination of charts and text to clearly present the data analysis reports required by clients;

[0058] Testing and Optimization: Test the agent to ensure it can accurately return results in various data analysis scenarios; continuously optimize the accuracy of semantic understanding of the large model based on test feedback and user feedback.

[0059] The beneficial effects of this invention are

[0060] To address the current challenges of data-driven approaches, achieve intelligent business operations, scientific decision-making, and a systematic knowledge base, thereby driving data-driven business transformation and a leap forward in work methods:

[0061] 1. Based on the basic capabilities of natural language processing for querying, industry knowledge is added to enhance semantic understanding capabilities, improving the accuracy of query intent recognition to over 90%;

[0062] 2. Improve data query efficiency, reducing manual data report processing time from 3 hours to 5 minutes;

[0063] 3. Assists customers in classifying and labeling, improving the efficiency of customer classification labeling by 80%. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the workflow of the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0066] This invention provides an intelligent data query method suitable for sales data inquiry scenarios. Conversational data analysis uses natural language to interact with the machine for data analysis. Regardless of whether users have professional skills or background knowledge, they can easily interact with the computer to perform queries and analyses. This makes data analysis more intuitive and convenient, improving the user's interactive experience and efficiency.

[0067] (1) Natural language input

[0068] By combining multiple intelligent technologies such as speech synthesis, speech recognition, semantic understanding, and semantic disambiguation, it enables voice-based conversational question analysis, making analysis accessible to everyone.

[0069] (2) Dialogue-style questioning

[0070] Data analysis based on natural language interaction via chat conversations intelligently identifies user intent and uses multi-turn question-and-answer sessions. This makes data acquisition and analysis easy for everyone, enabling everyone to use data.

[0071] (3) Plug-in model integration

[0072] It supports adaptation to multiple underlying models, allowing you to choose the appropriate model based on your specific business needs and scale.

[0073] (4) Multimodal analysis

[0074] Based on multiple basic large models and multiple business-level data mining small models, a data intelligent question analysis model engine platform is built, which supports single chart question counting, dashboard-type question counting, and prediction / classification question counting to meet a variety of business scenarios.

[0075] (5) Platform-based services

[0076] To provide accurate, reliable, and intervention-enabled analytical capabilities, the platform offers engineering services such as Prompt service management, table metadata synonym fine-tuning, user intervention, and chart switching.

[0077] It mainly includes the following 6 underlying logic modules:

[0078] (1) Semantic parsing

[0079] Leveraging the natural language understanding capabilities of the 14B large model, user queries are rewritten to optimize their structure and expression, making them more suitable for the model's processing requirements. Based on the rewritten queries, multimodal recall is performed to ensure the comprehensiveness and diversity of information. Embedding and re-ranking models are used to process and sort the recalled information.

[0080] (2) Intention Classification

[0081] Utilize extraction tools to extract dimensions (such as time, organization, etc.) and metrics (such as sales volume, sales revenue, and unit price). This allows for a more accurate understanding of user needs, providing crucial data support for subsequent business scenario modeling, data center services, and knowledge base retrieval.

[0082] (3) Business Scenario Modeling (text2API)

[0083] The extracted text undergoes API schema checks, an HTTP request is sent to query the data, and the visualization results are returned. If the conditions are not met, clarification information can be output or the process can be redirected to the data center service (text2Metric).

[0084] (4) Data Center Service (text2SQL)

[0085] We conduct in-depth data mining targeting the characteristics of the tobacco industry, achieving full-domain data penetration through terminology extraction and data source identification. A large-scale model is used for problem planning, and executable code is generated through repeated iterations using LLM plugins and examples. Dynamic prompts are generated incorporating business knowledge. Finally, the large-scale model (DeepSeek coder model) is used for inference analysis, reducing manual intervention and achieving streaming data output.

[0086] (5) Knowledge base retrieval (text2KB)

[0087] By conducting in-depth data retrieval using data metrics, relevant information is retrieved using a knowledge base, and the final result is generated using a large model.

[0088] (6) Graphic and text display

[0089] Based on the data, the system can intelligently select from various display types, such as bar charts, pie charts, and tables. The large-scale model can intelligently choose the most suitable display type based on specific data characteristics and analytical needs. Whether it's a straightforward bar chart, a vivid pie chart, or a detailed and clear table, the system can flexibly select and apply the appropriate format according to the nature of the data and the purpose of the analysis. This intelligent display method not only makes the data presentation more accurate and intuitive but also greatly improves the efficiency and effectiveness of data analysis.

[0090] To ensure priority availability in certain scenarios, optimizations should be considered in the following aspects.

[0091] (1) Select or fine-tune a suitable large model

[0092] For specific business needs and data characteristics, carefully select or finely tune basic large-scale models to ensure that the models can more accurately fit the application scenarios of the professional field, thereby improving the efficiency and accuracy of business processing.

[0093] In-depth analysis of business scenarios: First, a comprehensive understanding and analysis of the business processes, target users, operating environment, and expected results is essential. This step is fundamental to ensuring that the selected model meets actual needs.

[0094] Assess the data: Conduct a detailed evaluation of the size, quality, diversity, and distribution of the existing data. Data is the foundation for model training and optimization; therefore, a deep understanding of the data is crucial for model selection.

[0095] Choosing the right large model: Based on the business scenario and data evaluation results, select the most suitable one from among many large models. Consider factors such as model performance, generalization ability, training cost, and whether it is open source.

[0096] Model fine-tuning strategy: After selecting the base model, a fine-tuning strategy is formulated based on business characteristics and data features. This may include adjusting the model structure, optimizing training parameters, and retraining using domain-specific corpora.

[0097] Experimentation and Validation: During the fine-tuning process, multiple experiments were conducted, and methods such as cross-validation and A / B testing were used to evaluate the model's performance in business scenarios to ensure that the fine-tuning effect could meet expectations.

[0098] Continuous optimization: After the model is deployed, feedback data is continuously collected, model performance is monitored, and the model is constantly adjusted and optimized according to business development needs and environmental changes to maintain its leading position and adaptability in specific business scenarios.

[0099] (2) Optimization of prompt words:

[0100] The text2sql Prompt consists of several core parts:

[0101] Instructions: For example, "You are an SQL generation expert. Please refer to the following table structure and directly output the SQL statement without any further explanation."

[0102] Data Structure (Table Schema): Similar to a "vocabulary" in language translation. This refers to the database table structure you will be using. Since large models cannot directly access the database, you need to assemble the data structure into a Prompt. This typically includes table names, column names, column types, column meanings, and primary / foreign key information.

[0103] User questions: Questions expressed in natural language, such as: "Statistics on the sales volume of each specification last month."

[0104] Few-shot: This is an optional feature, and a common technique in project planning. It serves as a reference sample to guide the generation of the SQL query for the larger model.

[0105] Other tips: Any other instructions you deem necessary. For example, specifying that expressions are not allowed in the generated SQL, or requiring column names to be in the form of "table.column".

[0106] About text2API

[0107] (1) First, it is necessary to define a good data analysis API interface (such as the open API of the existing BI system). It is necessary to fully design and implement it according to their respective business situations, and form an API "instruction manual".

[0108] (2) The user inputs natural language, and the system uses LLM to convert the user's input into a call to the API tool, including the name of the API and the extracted parameters.

[0109] (3) Call the specified API based on the LLM response to obtain the returned data.

[0110] (4) As needed, the returned data is appended to the user input and then given to the LLM again, which outputs the final analysis results to the customer.

[0111] The 14B large model can be flexibly deployed on local or cloud servers, greatly lowering the barrier to entry. At the same time, its high flexibility allows the model to be fine-tuned and optimized according to specific needs, thus better adapting to various application scenarios.

[0112] The main model selection is as follows:

[0113] bge-large-zh: A Chinese-specific text embedding model that converts text into 1024-dimensional dense vectors (Embeddings) to capture deep semantic information of the text.

[0114] bge-reranker-large: Optimizes the retrieval results of the Embedding model by reordering the results to form an end-to-end optimized link.

[0115] Qwen2.5-14B-Instruct is a general-purpose multimodal task model that balances performance and cost. It can accurately parse user instructions, generate structured output, support multi-step logical reasoning and task decomposition, and improve the accuracy of knowledge base question answering.

[0116] Qwen2.5-Coder-14B-Instruct: is a large-scale language model specifically for open-source code. It can generate complete functions or modules based on natural language descriptions and combine with a visual model to achieve joint generation of "text-code-graph".

[0117] Application scenarios

[0118] Yesterday's Market Preview

[0119] Every day, the system's AI automatically generates yesterday's market analysis report and converts it into audio. Marketers can listen to the market analysis report at any time by clicking on the audio.

[0120] (1) Scene restoration

[0121] At 7 a.m., marketing staff tell the AI ​​assistant, "Report yesterday's market sales." The AI ​​assistant can then access real-time market data tailored to each marketing staff member's role, including yesterday's market performance, sales activity, market prices, and alerts. It transforms complex data into concise reports, delivered via voice. For example, the AI ​​assistant might say, "Yesterday's sales increased by 15% compared to the previous week, but the market price of a certain product fluctuated; we suggest you pay attention."

[0122] Marketers can easily stay informed about the latest market trends while brushing their teeth and washing their face.

[0123] (2) Implementation path

[0124] Data collection: Deploy data capture tools at various sales terminals, trading platforms, and price monitoring systems, or connect to big data centers to collect market operation data in real time, including cigarette purchase, sales, and inventory data, agreement and plan execution data, order and contract progress data, retailer order data, market price and sales data, etc.

[0125] Data processing: Cleaning, organizing, and processing the collected data, extracting key indicators, and analyzing trends such as purchasing, sales, inventory, market share, and price fluctuations through data mining techniques.

[0126] Intelligent analytics: Leveraging large-scale model analysis and reasoning capabilities, it conducts in-depth analysis of market data to uncover potential problems and business opportunities. For example, it performs attribution analysis on sales growth to predict future market trends and provide a basis for marketing strategies.

[0127] Speech synthesis: Converts processed data reports into speech using natural language generation technology, making the report content more similar to human language. It also supports multiple voice styles and pronunciations to meet the preferences of different marketers.

[0128] Intelligent Questioning

[0129] Asking questions using natural language allows you to quickly obtain data results; what you ask is what you get.

[0130] (1) Scene restoration

[0131] Marketers can directly ask the AI ​​assistant, "How much has competitor XXX sold nationwide?"

[0132] AI Assistant: "As of yesterday, the total sales volume of XXX nationwide was XXXX boxes, an increase of XX boxes year-on-year, representing a growth of XX%."

[0133] (2) Implementation path

[0134] Database construction: Based on the cigarette sales data of Project No. 1, the database and table structure were reconstructed to conform to the understanding of the large model.

[0135] Text2SQL model training: Collect a large number of natural language query samples and corresponding SQL statements related to cigarette sales, as well as basic knowledge of the tobacco industry, to fine-tune the basic model and evaluate the fine-tuning effect.

[0136] Intelligent agent development: First, semantic understanding is performed on the user's query question to extract key information and generate a corresponding SQL query. Then, the SQL is sent to the database for execution to obtain the query results. Finally, the query results returned by the database are converted into natural language and responded to marketing personnel through an AI assistant.

[0137] Testing and Optimization: Test the system to ensure accurate results are returned under various query scenarios. Continuously optimize the accuracy and response speed of the text2sql model based on test feedback and user feedback.

[0138] (3) Indicator Dimensions

[0139] The deep learning-based contextual understanding module, unlike the operational limitations of traditional BI tools, is deeply integrated with business scenarios, allowing users to ask conversational questions using natural language across multiple dimensions and indicators, and can intelligently identify ambiguous semantics.

[0140] The metrics and query dimensions supported by the metrics-based Q&A platform are as follows:

[0141]

[0142] Intelligent Analysis

[0143] Integrate marketing data from multiple systems, and generate intuitive and easy-to-understand market analysis reports and dynamic data dashboards through big data analysis and visualization technology.

[0144] (1) Scene restoration

[0145] A marketing professional is going to visit a commercial company in city X. In order to understand the company's overall sales and the sales performance of our company's brand in that company, the professional says to the AI ​​assistant: "Analyze the sales performance of the company in city X."

[0146] AI Assistant: "Sales Analysis Report of XX City Company"

[0147] I. Overall Sales Situation

[0148] As of yesterday, cumulative sales reached XXX cartons, compared to the same period last year…; sales revenue…; sales per carton…; high-priced cigarette sales…; pure premium cigarette sales…

[0149] XX City ranks X in sales and X in growth rate within XX Province.

[0150] The leading brands in XX City are XXX (sales volume XXX, market share XX%), XXX (sales volume XXX, market share XX%), and (sales volume XXX, market share XX%).

[0151] II. Sales performance of the XX brand

[0152] As of yesterday, cumulative sales reached XXX cartons, compared to the same period last year…; sales revenue…; sales per carton…; high-priced cigarette sales…; pure premium cigarette sales…

[0153] Brand XX ranks X in sales volume in XX city, accounting for X%; its year-on-year increase ranks X, and its growth rate ranks X.

[0154] The main selling product specifications are XXX (sales volume XXX, percentage of sales XX%), XXX (sales volume XXX, percentage of sales XX%), and (sales volume XXX, percentage of sales XX%).

[0155] (2) Implementation path

[0156] Build an API knowledge base: Organize the API interfaces related to sales data analysis, and label the functions implemented by each API, as well as their input and output parameters. If some analysis requirements do not have existing API interfaces, then API interface development will be necessary.

[0157] Large-scale model fine-tuning: This involves collecting fundamental tobacco industry knowledge related to API parameters, performing professional-level fine-tuning on the original basic model, and comprehensively evaluating the performance after adjustments. This initiative aims to significantly enhance the large-scale model's understanding of the API and improve the accuracy of parameter identification.

[0158] Intelligent Agent Development: Utilizing large-scale models to analyze user queries, through word segmentation, part-of-speech tagging, and semantic understanding, key information in the instructions is identified, clarifying the user's intent. Based on the user's intent, the system matches the API knowledge base, selects the appropriate API, and parses the parameters mentioned by the user, including time range, analysis dimensions, and metrics. The corresponding API is then called to obtain data results. Large-scale models are used to apply big data analytics algorithms for in-depth data mining and analysis (e.g., using time series analysis to predict the company's future sales trends; calculating key indicators such as the sales share and growth rate of key brands within the company). Furthermore, large-scale models can leverage visualization technology to transform the analysis results into intuitive and easy-to-understand market analysis reports and dynamic data dashboards, using a combination of charts (bar charts, line charts, pie charts, etc.) and text to clearly present the data analysis reports required by the client.

[0159] Testing and Optimization: Test the agent to ensure accurate results in various data analysis scenarios. Continuously optimize the accuracy of the large model's semantic understanding based on test feedback and user feedback.

[0160] (3) Report Template

[0161] A comprehensive comparison of local data with national averages and historical data highlights differences and unique characteristics. Taking a problem-oriented approach, it delves into the contradictions and shortcomings behind the data, conducting root cause analysis using attribution indicators to provide summary recommendations. The core structure is as follows:

[0162] 1) Report Analysis Section

[0163] The questions are listed in a bullet-point format, with each question forming an independent paragraph for easy reading. Each question is data-driven, including data comparisons (such as growth rate, percentage, and difference) and benchmark references (such as "below benchmark city"). Each analysis is expressed in a formulaic "keywords + data + conclusion" format, further refining the question.

[0164] The analysis modules are as follows:

[0165] Problem Type Analysis points Comparison Dimensions Market performance Sales volume, unit value, sales progress Compared with the province and benchmark cities Market Structure The proportion of cigarettes from within and outside the province Benchmarking Brands and Products Brand differentiation, high-end cigarettes Brand competition, category growth Industrial cooperation Highlights and shortcomings of cooperation among industrial enterprises Increment, contribution rate

[0166] 2) Summary and Recommendations Section

[0167] By matching each problem with a corresponding solution, a closed loop of "problem-solution" is formed, enhancing the relevance of the solution.

[0168] Each suggestion includes specific goals and benchmarks, making them highly actionable.

[0169] The strategy is presented in layers, covering multiple dimensions such as products, channels, brands, and partnerships, encompassing the entire business chain from market performance to brand and supply chain.

[0170] Key data (such as growth rate and difference) are highlighted in bold or red to facilitate quick identification of key information.

[0171] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An intelligent query method suitable for a sales data query scenario, characterized by comprising: (1) natural language input Realize voice dialogue type query analysis by combining intelligent technologies such as speech synthesis, speech recognition, semantic understanding and semantic disambiguation; (2) dialogue query Realize natural language interaction data analysis based on Chat conversation mode, intelligently identify user intent, and multiple rounds of question and answer; (3) plug-in model access Support adaptation to several kinds of underlying models, and select according to actual business and scale; (4) multi-modal analysis Based on several kinds of basic large models and several kinds of business-level data mining small models, build a model engine platform for data intelligent query analysis, support single chart query, support dashboard type query, support prediction / classification query to meet various business scenarios; (5) platform service Provide Prompt service management, synonym fine-tuning of table metadata, user intervention, and engineering services for chart switching.

2. The method of claim 1, characterized by comprising the following underlying logic modules: (1) semantic analysis With the natural language understanding ability of 14B large model, the user's question is rewritten to optimize its structure and expression, so that it is more in line with the processing requirements of the model; according to the rewritten question, multi-modal recall is performed; embedded models and reordering models are used to process and sort the recalled information; (2) intent classification Use extraction tools to extract dimensions and indicators to more accurately understand user needs, providing key data support for subsequent business scenario modeling, data center services and knowledge base retrieval; (3) business scenario modeling text2API Check the extracted text with API schema, send HTTP request, query data, and return visualization results; if the conditions are not met, output clarification information or fall back to data center service text2Metric; (4) data center service text2SQL Deeply mine data according to industry characteristics, realize global data penetration through term extraction and data source judgment; Use large models for problem planning, combine LLM plug-ins and examples to generate executable code through repeated iteration, and generate dynamic prompts combined with business knowledge; finally, use large models for reasoning analysis to realize streaming data output; (5) knowledge base retrieval text2KB Through data index deep retrieval, use knowledge base to recall related information, and use large models to generate final results; (6) graphic display Intelligently select column chart, pie chart and table type for display combined with data; intelligently select the most suitable display type according to specific data characteristics and analysis requirements.

3. The method of claim 2, characterized by: (1) first define the data analysis API interface, fully design and implement according to the respective business situation, and form the API instruction manual; (2) the user inputs natural language, and the LLM converts the user's input question into a call to the API tool, including the name of the API and the extracted parameters; (3) call the specified API according to the response of the LLM, and get the returned data; ​ ​ (4) According to the situation, the returned data is attached to the user input again, and the LLM is given to output the final response to the customer's analysis results.

4. The method of claim 3, wherein, Application scenario implementation path Data collection: Deploy data scraping tools at various sales terminals, trading platforms, and price monitoring systems, or interface with big data centers to collect market operation data in real time; Data processing: Clean, organize, and process the collected data, extract key indicators, and analyze trends in sales, market share, and price fluctuations through data mining techniques; Intelligent analysis: Use large model analysis and reasoning capabilities to conduct in-depth analysis of market data to identify potential problems and business opportunities; Speech synthesis: Convert the processed data report into speech using natural language generation technology to make the report content more close to human language expression.

5. The method of claim 4, wherein, Support several voice styles and pronunciations to meet the preferences of different personnel.

6. The method of claim 3, wherein, Intelligent question implementation path Database construction: Rebuild the database and table structure to meet the understanding of the large model based on sales data; text2sql model training: Collect relevant natural language query samples and corresponding sql statements, as well as industry basic knowledge, fine-tune the basic large model, and evaluate the fine-tuning effect; Agent development: First, perform semantic understanding on the user's query sentence, extract key information, and generate the corresponding sql query statement; then, send the sql to the database for execution and obtain the query result; finally, convert the database returned query result into natural language and reply to the marketing personnel through the AI assistant.

7. The method of claim 6, wherein, Test to ensure accurate results under various query scenarios; based on test feedback and user usage feedback, continuously optimize the accuracy and response speed of the text2sql model.

8. The method of claim 3, wherein, Intelligent analysis implementation path Build API knowledge base: Sort out API interfaces related to sales data analysis, and label the functions and input / output parameters of API implementation; if there is no ready-made API interface for part of the analysis requirements, API interface development is needed; Large model fine-tuning: Collect industry basic knowledge related to API parameters and fine-tune the original basic model at a professional level to comprehensively evaluate the performance of the adjusted model; Agent development: Use the large model to analyze user problems, identify key information in the instructions through word segmentation, part-of-speech tagging, and semantic understanding, and clarify user intent; According to the user's intention to match in the API knowledge base, select the appropriate API, and parse the parameters mentioned by the user; call the corresponding API to obtain the data result; use the large model to use big data analysis algorithm to deeply mine and analyze the data; in addition, the large model uses visualization technology to convert the analysis result into an intuitive and easy-to-understand market analysis report and dynamic data dashboard, using a combination of charts and text to clearly display the data analysis report required by the customer; Testing and optimization: test the agent to ensure that it can accurately return results in various data analysis scenarios; according to the test feedback and user feedback, continuously optimize the accuracy of the large model semantic understanding.

Citation Information

Patent Citations

  • Interactive number asking agent system based on large language model

    CN120216656A

  • System for realizing business intelligent data question answering based on semantic recognition

    CN120492570A

Cited By

  • Conversational model construction method, device and equipment

    CN122047523A

  • Industrial Internet AI Data Insight Method and System Based on Dynamic Skill System

    CN122309579A

  • Industrial internet ai data insight method and system based on dynamic skill system

    CN122309579B