Intelligent data query method and device based on large model and medium
Through the intelligent data query method driven by large-scale models, the problems of cross-system accurate query and multi-dimensional data visualization generation are solved, and natural language query and efficient and secure data services are realized for non-technical personnel.
Patent Information
- Application Number
- CN202510359503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology is difficult to realize natural language-driven cross-system accurate query, and traditional data query tools have a high threshold for non-technical personnel, so they cannot understand the business background and generate visual combination analysis of multi-source heterogeneous data.
The intelligent data query method based on the big model is adopted, and the target interface group is determined through preset business line classification rules and keyword matching algorithms, the interface call parameters are analyzed, and standardized data results are generated in combination with business logic rules, and multi-dimensional visual results are automatically generated.
It lowers the technical threshold for data query, realizes real-time convergence and visual analysis of cross-system data, improves data comprehensibility and security, and supports efficient, accurate and secure intelligent data services.
Smart Images

Figure CN120296126A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large model application technology, and in particular to a large model-based intelligent data query method, device and medium. Background Art
[0002] As enterprises’ digital transformation deepens, the scale of business system data is growing exponentially. How to efficiently integrate scattered data and lower the threshold for data use has become a core challenge to improving enterprise decision-making efficiency. Especially in diversified business fields such as retail and manufacturing, systems such as sales, warehousing, and supply chain are usually deployed independently and have heterogeneous data structures, forming "data islands." Although opening up data access rights through API interfaces has become a mainstream solution, joint queries of cross-system data still rely on manually written interface call logic, which is difficult to meet real-time and flexible business needs.
[0003] Current enterprise data query tools generally use structured query language (SQL) or visual chart configuration, which has significant limitations: on the one hand, non-technical personnel need to master professional knowledge such as database table structure and field mapping relationship to complete effective query, resulting in a large number of business personnel unable to obtain data independently; on the other hand, existing tools only support accurate retrieval of preset conditions, and cannot understand vague semantic descriptions such as "reasons for recent sales fluctuations in North China", making it difficult to mine the associated value behind the data. Some improvement plans try to realize the conversion of natural language to query statements through rule engines, but they rely on limited predefined templates, and the parsing error rate increases significantly when faced with complex business semantics (such as "statistical caliber includes return orders"), and cannot dynamically adapt to new interfaces or changed business logic.
[0004] In addition, traditional solutions have serious flaws in data result presentation and interpretation. The raw data returned by the interface lacks business background supplements (such as data sources and calculation rules), and users need to consult technical personnel to understand its meaning; and the fusion display of multi-source data usually only supports basic tables or single charts, and cannot automatically generate trend analysis, comparison views and other visualization combinations based on query intent. This inefficient interaction mode not only prolongs the data usage chain, but may also lead to decision-making errors caused by information misunderstanding.
[0005] Therefore, how to use intelligent means to connect enterprise data silos, achieve natural language-driven cross-system precision query, and automatically generate business-interpretable multidimensional data results has become a technical problem that technical personnel in this field urgently need to solve. Summary of the invention
[0006] The embodiments of the present application provide an intelligent data query method, device, and medium based on a large model to solve the following technical problems: how to break through enterprise data islands through intelligent means, achieve accurate cross-system query driven by natural language, and automatically generate multi-dimensional data results that are business-explainable.
[0007] In a first aspect, the embodiments of the present application provide an intelligent data query method based on a large model. The method includes: receiving a natural language query text input by a user, and determining a target interface group from a pre-associated interface group through a preset business line classification rule and keyword matching algorithm; inputting the natural language input text into a pre-trained large model to parse a parameter set required for interface invocation; where the parameter set includes at least two of the following: a time range field, a business entity field, and a statistical dimension field; invoking an API interface in the target interface group according to the parameter set to obtain original business data; performing context supplementation on the original business data based on business logic rules to generate a standardized data result including a business background description; reconstructing the standardized data result according to a display format preset by the user, and outputting a combined visualization result including charts, tables, and text descriptions.
[0008] In an implementation manner of the present application, determining a target interface group from a pre-associated interface group through a preset business line classification rule and keyword matching algorithm specifically includes: grouping interfaces according to business line labels, and binding at least a first preset number of business keywords to each group of interfaces; calculating the semantic similarity between the user input text and the keywords of each interface group using the TF-IDF algorithm; selecting the interface group with a similarity exceeding a preset similarity threshold as the target interface group.
[0009] In an implementation manner of the present application, inputting the natural language input text into a pre-trained large model to parse a parameter set required for interface invocation specifically includes: identifying time description words in the natural language and mapping them to preset time enumeration values; when extracting the business entity field, preferentially matching the entity list associated with the interface group, and if not hit, retrieving in the business database through a fuzzy matching algorithm; after identifying the statistical dimension field, performing standardization processing on the statistical dimension field to convert it into a region_code field required by the interface.
[0010] In an implementation manner of the present application, performing context supplementation on the original business data based on business logic rules to generate a standardized data result including a business background description specifically includes: extracting an associated description from a pre-configured business rule library according to the type of the invoked API interface; where the associated description includes a statistical caliber, a data source, and a calculation method; using the large model to convert the associated description into a natural language description and inserting it at the head of the data result to generate a standardized data result.
[0011] In one implementation of the present application, after reconstructing the standardized data results according to the display format preset by the user and outputting a combined visualization result including charts, tables, and text descriptions, the method further includes: when the user inputs query content containing a first preset keyword, automatically generating a line chart and year-on-year or month-on-month calculation data; highlighting key indicators in the calculation data, and generating a summary text description not exceeding a second preset quantity through a large model.
[0012] In one implementation of the present application, the method further includes: fusion processing of multi-interface data, specifically including: when the user's query involves cross-business-line data, parallelly invoking APIs of at least a third preset quantity of interface groups; associating multi-source data according to a preset primary key field; wherein, the time field is aligned using UNIX timestamps, and the business entity field is mapped through a unified coding table.
[0013] In one implementation of the present application, before invoking the API interface in the target interface group according to the parameter set to obtain the original business data, the method further includes: detecting whether the time range field exceeds the maximum supported range of the interface, and if it exceeds the limit, automatically splitting it into multiple sub-requests for batch processing; performing permission verification on the business entity field and shielding sensitive data interfaces that the user has no access permission to.
[0014] In one implementation of the present application, the method further includes: when the user inputs a statistical requirement containing a second preset keyword, automatically applying a moving average algorithm or a sorting algorithm to the data returned by the interface; adding a calculation process traceability identifier to the derivative indicators generated by the large model, and displaying the original data and formula after responding to a click operation.
[0015] In a second aspect, an intelligent data query device based on a large model provided by an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an intelligent data query method based on a large model as described in any one of the above.
[0016] In a third aspect, a non-volatile computer storage medium for intelligent data query based on a large model provided by an embodiment of the present application stores computer-executable instructions, and when the computer-executable instructions are executed, an intelligent data query method based on a large model as described in any one of the above is implemented.
[0017] An intelligent data query method, device, and medium based on a large model provided by an embodiment of the present application have the following beneficial effects:
[0018] Through the deep integration of large model semantic parsing and business logic, this application significantly reduces the technical threshold for enterprise data query. Non-technical personnel can directly input natural language descriptions (such as "Compare the sales volume and return rate in East China in the past three months"), and the system automatically parses the time range, business entity, and statistical dimension fields, and accurately matches the target API interface group, completely getting rid of the limitations of relying on SQL or manual configuration of interface calls in traditional solutions. Combining the TF-IDF algorithm with the interface group screening mechanism of business keywords, it can lock cross-system data sources within milliseconds, and at the same time solve the problem of business entity naming ambiguity through fuzzy matching and standardized coding conversion, greatly improving the data query efficiency by more than that.
[0019] Regarding the integration problem of multi-source heterogeneous data, this application adopts parallel API call and primary key dynamic association technology to achieve real-time integration of data across business lines. Through UNIX timestamp alignment and unified coding table mapping, it ensures the spatio-temporal consistency of data in systems such as sales, warehousing, and supply chain; based on the context supplement mechanism of the business rule library, it automatically adds statistical caliber descriptions to the original data (such as "sales amount includes tax and excludes returns"), combined with the visual combination (trend chart, comparison table, text summary) generated by the large model, greatly improving the comprehensibility of complex data and effectively avoiding decision-making misjudgments caused by the lack of background information.
[0020] In addition, this application has made breakthroughs in terms of security and interpretability. Through the permission verification module, it automatically shields sensitive interfaces to ensure that users only access authorized data; the batch processing mechanism for the over-limit time range takes into account both interface stability and query integrity. The design of the calculation process traceability identifier for derived indicators enables users to trace the original data and algorithm formulas, enhancing the credibility of the results. Compared with traditional solutions, this application shortens the data query response time and supports tens of millions of concurrent requests per day, providing core technical support for enterprises to build a real-time, accurate, and secure intelligent data service system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0022] Figure 1 is a flowchart of an intelligent data query method based on a large model provided by an embodiment of this application;
[0023] Figure 2 is a schematic internal structure diagram of an intelligent data query device based on a large model provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0025] The embodiments of this application provide an intelligent data query method, device, and medium based on a large model to solve the following technical problems: how to break through enterprise data islands through intelligent means, achieve accurate cross-system queries driven by natural language, and automatically generate multi-dimensional data results that can be interpreted by the business.
[0026] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the drawings.
[0027] Figure 1 It is a flowchart of an intelligent data query method based on a large model provided for the embodiments of this application. As Figure 1 shown, an intelligent data query method based on a large model provided by the embodiments of this application specifically includes the following steps:
[0028] Step 101: Receive the natural language query text input by the user, and determine the target interface group from the pre-associated interface groups through the preset business line classification rules and keyword matching algorithms.
[0029] In an embodiment of this application, to implement intelligent data query based on a large model, after receiving the natural language query text input by the user, it is first necessary to determine the target interface group from the pre-associated interface groups through the preset business line classification rules and keyword matching algorithms.
[0030] Specifically, group the interfaces according to the business line labels, and bind at least the first preset number of business keywords to each group of interfaces; use the TF-IDF algorithm to calculate the semantic similarity between the user input text and the keywords of each interface group; select the interface group with a similarity exceeding the preset similarity threshold as the target interface group.
[0031] In this embodiment, the business line classification rules are dynamically generated based on the enterprise organizational structure and the attribution relationship of data sources, and each business line corresponds to at least one independent business system. For example, in the scenario of a retail enterprise, the preset business lines include three categories: "Supply Chain Management", "Store Sales", and "Online Mall", and the corresponding ERP, POS, and OMS system API interfaces are mounted under each business line. During technical implementation, the data processor adds business line tags to each interface through a configuration platform to form a pre-associated interface group. The association logic of the interface group includes: the interfaces within the same business line share the definition of data primary keys and support access through a unified authentication token.
[0032] The keyword matching algorithm is applied in the interface group screening stage, and its core lies in establishing the mapping relationship between business semantics and interface functions. Specifically, each pre-associated interface group needs to be bound with at least 10 business keywords, and these keywords are generated through the following two methods: 1) Extract high-frequency business entity nouns from the interface documentation, such as "Inventory Turnover Rate" and "Order Fulfillment Rate"; 2) Based on the clustering analysis of historical query logs, mine the scenario-based expressions commonly used by users. For example, "Promotion Effect Analysis" corresponds to the interfaces related to commodity sales volume, discount rate, and customer unit price. During the implementation process, the keyword library is automatically expanded quarterly through the NLP model, and the keyword binding operation needs to be compulsorily completed when adding a new interface group.
[0033] When receiving the query text input by the user (such as "Please output the list of slow-moving products in each store in the East China region in the past three months"), the system first performs word segmentation and stop word filtering to extract the effective words "East China region", "three months", "store", and "slow-moving products". Subsequently, the TF-IDF algorithm is used to calculate the semantic similarity between these words and the keywords of each interface group, and the interface group that matches the "Store Operations" business line is preferentially selected. During this process, the algorithm dynamically adjusts the weight distribution: the matching weight of geographical words (such as "East China") is reduced by 15%, and the matching weight of time words (such as "three months") is increased by 20% to adapt to the time-sensitive characteristics of the business line data. If there is more than one interface group with a similarity exceeding 0.7, the system automatically triggers a secondary confirmation mechanism, such as returning a prompt "Please confirm whether you need to query online mall or physical store data", and locks the target interface group after the user makes a selection.
[0034] The establishment of the pre-associated interface group includes a data source adaptation layer, and its technical implementation includes: 1) Protocol encapsulation of the access interfaces of heterogeneous databases, and unified conversion into the RESTful API format; 2) Adding a metadata description file to each interface to record attributes such as parameter structure, return fields, and business line attribution. For example, for the inbound details interface of the supply chain management business line, its metadata includes the required parameter "warehouse code" and the optional parameter "product category", and it is marked that the data returned by this interface includes extended fields such as "batch number" and "quality inspection status". This design enables the parameter parsing module in subsequent steps to dynamically adapt to the data specifications of different interfaces.
[0035] In the exception handling scenario, when the user input text cannot match any interface group (such as querying "store passenger flow" but no Internet of Things device interface is configured), the system executes a degradation strategy: 1) Recommend a similar interface group based on the business line label, and prompt "Currently, store sales data can be provided"; 2) Record the unrecognized keywords and trigger an alarm to drive the data handler to expand the interface coverage range subsequently. This mechanism effectively solves the contradiction between business requirement iteration and interface development lag, and ensures system availability.
[0036] Step 102: Input the natural language input text into a pre-trained large model to parse the parameter set required for interface invocation.
[0037] First of all, it should be noted that the parameter set in this embodiment includes at least two of the following: time range field, business entity field, and statistical dimension field.
[0038] In an embodiment of the present application, inputting the natural language input text into a pre-trained large model to parse the parameter set required for interface invocation specifically includes: identifying the time description words in the natural language and mapping them to preset time enumeration values; when extracting the business entity field, first match the entity list associated with the interface group, and if not hit, retrieve in the business database through a fuzzy matching algorithm; after identifying the statistical dimension field, perform standardization processing on the statistical dimension field to convert it into the region_code field required by the interface.
[0039] In this embodiment, the pre-trained large model adopts a multi-task joint fine-tuning architecture. Its base model is a pre-trained language model based on Transformer (such as BERT, GPT-3), and it is domain-adapted on a business corpus composed of enterprise historical query logs and API call records. Three types of training objectives are introduced in the fine-tuning stage: 1) The time expression recognition task, which learns the mapping relationship between fuzzy expressions such as "end of the quarter" and "last week" and specific date ranges; 2) The business entity disambiguation task, which establishes the association between expressions such as "East China Region" and "Store Number A001" and the primary key of the database; 3) The statistical intent classification task, which identifies deep requirements such as "trend analysis" and "TOP10 ranking". For example, when the user inputs "Show me the five worst-selling products last month", the model needs to parse the time range (from the 1st to the end of last month), the business entity (list of product SKUs), and the statistical dimension (take the fifth in reverse order by sales).
[0040] In this embodiment, the parsing of the time range field is implemented through a dynamic context awareness mechanism. The large model first identifies explicit time description words in the input text (such as "the last three months", "Q2 2023"). If there is a fuzzy expression (such as "year to date"), the date range is dynamically calculated in combination with the query trigger time. Specifically: The model builds a preset time enumeration value mapping table, for example, "quarter" corresponds to the natural quarter division (January - March is Q1), and it supports the custom configuration of the enterprise fiscal year (such as the retail industry defines November - January as the promotion season); for expressions that are not clearly defined (such as "recently"), default values are set according to the characteristics of the business line: the supply chain business line defaults to the last 7 days, and the sales business line defaults to the last 30 days; the parsing result is converted to the ISO 8601 standard format (such as "2023-07-01T00:00:00 / 2023-09-30T23:59:59") to adapt to the date parameter specifications of different interfaces.
[0041] The extraction of the business entity field adopts a hierarchical matching strategy. First, based on the target interface group determined in step 101, load the associated entity list (such as the store interface group binds all store codes and names), and lock the target through exact matching first. For example, when the user inputs "Sales data of Hangzhou West Lake Store in August", it directly matches the store code HZ-XH-001. If not hit (such as the user inputs "flagship store" but there are "Beijing Flagship Store" and "Shanghai Flagship Store" in the list), then start fuzzy matching: Build an entity alias library based on the business database (such as "flagship store" corresponds to stores at store level 1); Use the Levenshtein distance algorithm to calculate the similarity between the input text and the alias library, and filter candidate entities with a similarity threshold greater than 0.8; When the number of candidates exceeds 3, call the large model to generate a clarification question (such as "Do you need to query the data of the Beijing, Shanghai or Guangzhou flagship store?").
[0042] The standardization of statistical dimension fields depends on the interface metadata description file. The large model matches the identified statistical intent (such as "aggregate by region" and "compare by product category") with the dimension fields supported by the interface and performs unified unit conversion. For example: when the interface requires the region_code format to be a two-digit administrative division code (such as "East China" corresponds to HD), and the user enters "Jiangsu, Zhejiang and Shanghai regions", the model will break it down into "Jiangsu (JS)", "Zhejiang (ZJ)" and "Shanghai (SH)" and generate multi-value parameters; for scenarios where units do not match (such as the interface returns "sales" in units of 10,000 yuan, and the user queries "meta-level details"), the scaling factor parameter (scale_factor = 0.0001) is automatically added.
[0043] Step 103: Call the API interface in the target interface group according to the parameter set to obtain original business data.
[0044] In this embodiment, the parameter set calling process adopts a dynamic routing and load balancing mechanism, the core of which is to automatically adapt the calling protocol according to the interface metadata description and ensure stability in high-concurrency scenarios. When the target interface group is determined (such as the warehousing detail interface group under the supply chain management business line), the system first obtains its access endpoint (Endpoint), authentication method (OAuth 2.0 / Basic Auth) and parameter specifications from the interface registration center. For example, the warehousing system API requires the time range parameter to be named date_range and in the format of start_date, end_date, while the sales system uses begin_time and end_time as independent parameters. At this time, the system automatically reconstructs the parameter name and structure according to the interface definition to achieve accurate mapping of unified parameter sets to heterogeneous interfaces.
[0045] In one embodiment of the present application, the method also includes: fusion processing of multi-interface data, specifically including: when the user query involves cross-business line data, calling the API of at least a third preset number of interface groups in parallel; associating multi-source data according to a preset primary key field; wherein the time field is aligned with UNIX timestamp, and the business subject field is mapped through a unified coding table.
[0046] In this embodiment, cross-business line data calls are implemented through a parallel request processing engine. When a user query involves multiple business lines (such as "comparing the East China supply chain warehousing efficiency with the online mall order fulfillment rate"), the system simultaneously initiates API calls to the supply chain management interface group and the online mall interface group. The specific implementation includes:
[0047] 1. Interface parallel triggering: Create an independent thread pool, assign a dedicated thread to each interface group, set a timeout threshold (5 seconds by default) to prevent a single interface delay from blocking the overall process;
[0048] 2. Data Association Preprocessing: Inject an association identifier at the request level (e.g., correlation_id = 20230915-QUERY001) to ensure that cross-system data for the same query can be traced in subsequent steps;
[0049] 3. Heterogeneous Data Alignment: Map the fields returned by different interfaces according to the preset primary key rules. For example, the warehouse_id returned by the supply chain interface and the store_code of the mall interface are associated as the same physical warehouse through a unified coding table, and the time fields are uniformly converted to UNIX timestamps (e.g., 2023-09-15 14:30:00 → 1694766600) to eliminate time zone and format differences.
[0050] In an embodiment of the present application, before calling the API interfaces in the target interface group according to the parameter set to obtain the original business data, the method further includes: detecting whether the time range field exceeds the maximum supported range of the interface, and if it exceeds the limit, automatically splitting it into multiple sub-requests for batch processing; performing permission verification on the business entity field and shielding sensitive data interfaces that the user has no access rights to.
[0051] In this embodiment, the time range compliance check includes: reading the max_date_range attribute defined in the interface metadata (e.g., the supply chain interface supports a maximum 90-day query), and when the user's requested time span exceeds this limit, the system automatically splits it into a sequence of sub-requests by week / month. For example, if the user queries the "inventory turnover rate for the whole year of 2023" and the interface only supports a single 90-day query, it is split into 4 calls (Q1-Q4), and the result sets are merged at the data return stage. The business entity permission check includes: based on the RBAC (Role-Based Access Control) model, querying the data permission table bound to the user's role in real time. For example, an ordinary salesperson can only access the store data with region_code = HD in the area they are responsible for. If the request contains region_code = HB, this part of the parameter is automatically filtered, and the unauthorized attempt is recorded in the log. The sensitive interface shielding includes: for interfaces marked with security_level = high (such as the cost price interface), directly returning a "The interface is temporarily unavailable" prompt during unauthorized periods (such as non-financial accounting periods) to avoid the risk of data leakage.
[0052] In an embodiment of the present application, when obtaining the original business data, the original data standardization process is completed at the interface response stage, including: 1. Protocol conversion: Convert the data returned by non-RESTful interfaces such as SOAP / GraphQL into a unified JSON format through an adapter; 2. Null value filling: According to the default value configuration in the business rule library, intelligently complete the missing fields. For example, when the warehousing interface does not return inbound_quantity, it is automatically filled with 0 and the is_estimated=true flag is added; 3. Exception fusing: When an interface returns 5xx errors continuously for 3 times, temporarily remove it from the target interface group and trigger an alarm to notify the operation and maintenance personnel.
[0053] Step 104: Perform context supplementation on the original business data based on business logic rules to generate a standardized data result including business background description.
[0054] In an embodiment of the present application, context supplementation is performed on the original business data based on business logic rules to generate a standardized data result including business background description, specifically including: extracting associated descriptions from a pre-configured business rule library according to the type of API interface called; where the associated descriptions include statistical caliber, data source, and calculation method; using a large model to convert the associated descriptions into natural language descriptions and inserting them at the head of the data result to generate a standardized data result.
[0055] In this embodiment, the business logic rules are dynamically loaded through a configurable rule engine, and its core lies in associating technical metadata with business semantics to form a system for enhancing data interpretability. The rule engine is synchronized bidirectionally with the business rule library, and the rule library stores three types of key information according to the interface dimension: 1) Statistical caliber description, such as "sales calculation includes the prepayment of cancelled orders"; 2) Data lineage, marking the original data source (such as "supply chain data comes from the ECC module of the SAP system"); 3) Calculation logic formula, such as "gross profit margin = (sales - cost) / sales * 100%". Rule updates are completed through a visual configuration interface. For example, when the finance department adjusts the statistical caliber, it can check the "deduct freight" option and specify the effective time in the interface, and the system automatically generates a versioned rule file.
[0056] Among them, the context supplementation process is divided into two collaborative stages: 1. Rule matching and extraction: According to the API interface identifier called in step 103 (such as / api / sales / detail), retrieve the matching set of associated descriptions from the rule library. For example, when the interface return field contains return_amount (return amount), automatically append the rule that "return orders within the current data statistics period need to be submitted within 7 days after shipment to be counted"; 2. Semantic conversion and embedding: The large model receives the original data and associated rules and generates a natural language description that conforms to the business scenario. During the conversion process, the model adjusts the expression granularity according to the user role: when outputting to management, a macro summary is used (such as "The sales volume in Q3 of 2023 increased by 12% year-on-year, and the main growth came from the new product launch in the East China region"), while detailed descriptions are provided to operation personnel (such as "Sales volume statistics include online mall GMV and offline store POS transaction data, excluding group purchase and wholesale business").
[0057] In addition, the context supplementation process is divided into two collaborative stages: Rule matching and extraction: According to the API interface identifier called in step 103 (such as / api / sales / detail), retrieve the matching set of associated descriptions from the rule library. For example, when the interface return field contains return_amount (return amount), automatically append the rule that "return orders within the current data statistics period need to be submitted within 7 days after shipment to be counted"; Semantic conversion and embedding: The large model receives the original data and associated rules and generates a natural language description that conforms to the business scenario. During the conversion process, the model adjusts the expression granularity according to the user role: when outputting to management, a macro summary is used (such as "The sales volume in Q3 of 2023 increased by 12% year-on-year, and the main growth came from the new product launch in the East China region"), while detailed descriptions are provided to operation personnel (such as "Sales volume statistics include online mall GMV and offline store POS transaction data, excluding group purchase and wholesale business").
[0058] Step 105: Reconstruct the standardized data results according to the display format preset by the user, and output a combined visual result including charts, tables, and text descriptions.
[0059] In this embodiment, the display format reconstruction process uses a dynamic rendering engine, the core of which is to deeply integrate standardized data with user portraits and business scenarios to generate an interactive multimodal analysis report. The system pre-sets multiple display templates (such as business analysis dashboards, supply chain monitoring views). Users can set default preferences through a personalized configuration interface. For example, when selecting the "mobile-first" template, the chart size and text layout are automatically optimized to fit the mobile screen, while the "executive briefing" template forces the summary mode to be enabled and detailed data to be hidden. For scenarios without a pre-set format, the system intelligently recommends display solutions based on data characteristics: time series data is preferentially rendered as a trend line chart, categorical comparison data uses a stacked bar chart, and geographical distribution data triggers a heat map component.
[0060] In an embodiment of the present application, after reconstructing the standardized data result according to the display format preset by the user and outputting a combined visualization result including charts, tables, and text descriptions, the method further includes: when the user inputs query content containing a first preset keyword, automatically generating a line chart and year-on-year or month-on-month calculation data; highlighting key indicators in the calculation data, and generating a summary text description not exceeding a second preset number through a large model.
[0061] In an embodiment of the present application, the method further includes: when the user inputs a statistical requirement containing a second preset keyword, automatically applying a moving average algorithm or a sorting algorithm to the data returned by the interface; adding a calculation process traceability identifier to the derivative indicators generated by the large model, and displaying the original data and formula after responding to a click operation.
[0062] In this embodiment, the intelligent generation of visualization elements is achieved through the following collaborative mechanism:
[0063] 1. Chart type decision-making: When the user inputs keywords such as "trend" and "growth rate", call the time series analysis engine to automatically generate a line chart with a smooth curve and add year-on-year / month-on-month auxiliary lines (such as using a dotted line to mark the data of the same period last year). For example, in the scenario of "analyzing the sales trend in East China in the past six months", the horizontal axis of the line chart is divided by month, and the vertical axis shows both the absolute sales amount and the month-on-month percentage;
[0064] 2. Statistical derivative processing: When statistical requirement words such as "TOP10" and "ranking" are detected, apply a quick sorting algorithm to the data returned by the interface, generate a horizontal bar chart, and mark the numerical range with a gradient color. For complex calculations such as moving average and standard deviation, add a formula explanation floating window below the chart, and the user can click on the formula icon to expand the detailed derivation process;
[0065] 3. Multi-dimensional drill-down: The row / column fields of table data support dynamic switching. For example, in the "Quarterly Sales Table of Each Store", clicking on the region title can drill down to the city dimension, and right-clicking on the product category can select "Group by Supplier". Each interaction operation triggers the large model to generate new analysis conclusions in real time, such as "The sales proportion of mobile phone categories in the Hangzhou region in Q3 has increased to 35%, mainly due to the first launch of new products by Supplier A."
[0066] Among them, the generation rule of the text description adopts a three-layer progressive structure: First-level summary: The large model extracts key features such as data extreme values and inflection points, and generates a core conclusion of no more than two lines. For example, "The procurement cost of servers in September increased by 22% month-on-month, mainly affected by the shortage of DRAM chips."; Second-level analysis: Combining industry knowledge in the business rule library, provide attribution analysis for abnormal fluctuations. For example, append "Note: The price index of the global DRAM market increased by 18% in Q3, see Annex 'Semiconductor Industry Weekly Report No. 37'" in the cost increase description; Third-level suggestion: When a continuous negative trend is identified (such as the inventory turnover rate has decreased for 3 consecutive months), call the decision tree model to generate optimization suggestions, such as "It is recommended to transfer the redundant inventory in the East China warehouse to the South China warehouse, which is expected to reduce the warehousing cost by 3%."
[0067] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiment of this application also provides an intelligent data query device based on a large model, and its structure is as Figure 2 shown.
[0068] Figure 2 It is a schematic diagram of the internal structure of an intelligent data query device based on a large model provided by the embodiment of this application. As Figure 2 shown, the device includes:
[0069] At least one processor 201;
[0070] And, a memory 202 communicatively connected to at least one processor;
[0071] Among them, the memory 202 stores instructions executable by at least one processor, and the instructions are executed by at least one processor 201, so that at least one processor 201 can:
[0072] Receive the natural language query text input by the user, and determine the target interface group from the pre-associated interface groups through the preset business line classification rules and keyword matching algorithms; input the natural language input text into a pre-trained large model to parse the parameter set required for interface calls; wherein, the parameter set includes at least two of the following: time range field, business entity field, statistical dimension field; call the API interfaces in the target interface group according to the parameter set to obtain the original business data; perform context supplementation on the original business data based on the business logic rules to generate a standardized data result including business background descriptions; reconstruct the standardized data result according to the display format preset by the user, and output a combined visualization result including charts, tables and text descriptions.
[0073] Some embodiments of the present application provide a Figure 1 non-volatile computer storage medium for intelligent data query based on a large model, storing computer-executable instructions, and the computer-executable instructions are set as:
[0074] Receive the natural language query text input by the user, and determine the target interface group from the pre-associated interface groups through the preset business line classification rules and keyword matching algorithms; input the natural language input text into a pre-trained large model to parse the parameter set required for interface calls; wherein, the parameter set includes at least two of the following: time range field, business entity field, statistical dimension field; call the API interfaces in the target interface group according to the parameter set to obtain the original business data; perform context supplementation on the original business data based on the business logic rules to generate a standardized data result including business background descriptions; reconstruct the standardized data result according to the display format preset by the user, and output a combined visualization result including charts, tables and text descriptions.
[0075] The embodiments in the present application are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0076] The systems and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0077] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0078] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0081] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0082] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0083] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0084] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0085] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An intelligent data query method based on a large model, characterized in that, The method includes: Receiving a natural language query text input by a user, and determining a target interface group from a pre-associated interface group through a preset business line classification rule and a keyword matching algorithm; Inputting the natural language input text into a pre-trained large model to parse a parameter set required for interface invocation; wherein, the parameter set includes at least two of the following: a time range field, a business entity field, and a statistical dimension field; Invoking an API interface in the target interface group according to the parameter set to obtain original business data; Performing context supplementation on the original business data based on business logic rules to generate a standardized data result including a business background description; Reconstructing the standardized data result according to a display format preset by the user, and outputting a combined visualization result including charts, tables, and text descriptions.
2. The intelligent data query method based on a large model according to claim 1, wherein Determining a target interface group from a pre-associated interface group through a preset business line classification rule and a keyword matching algorithm, specifically including: Grouping interfaces according to business line tags, and binding at least a first preset number of business keywords to each group of interfaces; Calculating the semantic similarity between the user input text and the keywords of each interface group using the TF-IDF algorithm; Selecting the interface group with a similarity exceeding a preset similarity threshold as the target interface group.
3. An intelligent data query method based on a large model according to claim 1, characterized in that Inputting the natural language input text into a pre-trained large model to parse a parameter set required for interface invocation, specifically including: Identifying time description words in natural language and mapping them to a preset time enumeration value; When extracting the business entity field, preferentially matching the entity list associated with the interface group, and if not hit, retrieving it in the business database through a fuzzy matching algorithm; After identifying the statistical dimension field, performing standardization processing on the statistical dimension field to convert it into a region_code field required by the interface.
4. An intelligent data query method based on a large model according to claim 1, characterized in that, Performing context supplementation on the original business data based on business logic rules to generate a standardized data result including a business background description, specifically including: Extracting an associated description from a pre-configured business rule library according to the type of the invoked API interface; wherein, the associated description includes a statistical caliber, a data source, and a calculation method; Using the large model to convert the associated description into a natural language description and inserting it at the head of the data result to generate the standardized data result.
5. An intelligent data query method based on a large model according to claim 1, characterized in that, After reconstructing the standardized data result according to a display format preset by the user and outputting a combined visualization result including charts, tables, and text descriptions, the method further includes: When the user input contains a query content with a first preset keyword, automatically generating a line chart and year-on-year or month-on-month calculation data; Highlighting key indicators in the calculation data, and generating a summary text description not exceeding a second preset number through the large model.
6. An intelligent data query method based on a large model according to claim 1, characterized in that, The method further includes: Fusion processing of multi-interface data, specifically including: When the user query involves cross-business line data, concurrently invoking APIs of at least a third preset number of interface groups; Associating multi-source data according to a preset primary key field; wherein, the time field is aligned using UNIX timestamps, and the business entity field is mapped through a unified coding table.
7. An intelligent data query method based on a large model according to claim 1, characterized in that, Before calling the API interfaces in the target interface group according to the parameter set to obtain the original business data, the method further includes: Detect whether the time range field exceeds the maximum supported range of the interface. If it exceeds the limit, it is automatically split into multiple sub-requests for batch processing; Perform permission verification on the business entity field to shield sensitive data interfaces that the user has no access rights to.
8. An intelligent data query method based on a large model according to claim 1, characterized in that The method further includes: When the user inputs a statistical requirement containing a second preset keyword, automatically apply a moving average algorithm or a sorting algorithm to the data returned by the interface; Add a calculation process traceability identifier to the derivative indicators generated by the large model, and display the original data and formula after responding to the click operation.
9. An intelligent data query device based on a large model, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a method according to any one of claims 1-8.
10. A non-volatile computer storage medium for intelligent data query based on a large model, storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, a method according to any one of claims 1-8 is implemented.
Citation Information
Cited By
Construction method of intelligent question answering system based on lightweight large model
CN120705279A
A method for constructing an intelligent question and answer system based on a lightweight large model
CN120705279B
Dify-based natural language data query system and method
CN120929476A
Intelligent number asking method and system based on large language model
CN121117053A
Data processing method and device based on large model, medium, equipment and product
CN121166900A