A large model-based intelligent data analysis method and system

By extracting semantic constraints from the large-scale intelligent data analysis system, determining whether there are matching existing reports in the report database, and directly generating query results, the problem of low query efficiency in existing technologies is solved, achieving efficient query response and accurate data presentation.

CN120994811BActive Publication Date: 2026-02-10SHANGHAI JINGKUN COMPUTER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511508259.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-10
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing intelligent data analysis technologies based on large models fail to effectively utilize existing report resources when matching user query needs with existing reports, resulting in low query efficiency.

Method used

By collecting natural language query requests, semantic parsing is performed to extract semantic constraints, and it is determined whether there are matching existing reports in the report database. If they exist, the query results are generated directly; otherwise, the results are generated through semantic transformation methods, and finally, a visual chart is generated.

Benefits of technology

It significantly improves query response time, enhances overall query efficiency, ensures that there is no need for repeated data parsing and processing when existing reports exist, and reduces the impact of environmental interference on the accuracy of voice query content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994811B_ABST
    Figure CN120994811B_ABST
Patent Text Reader

Abstract

The application relates to a large model-based intelligent data analysis method and system, and relates to the field of intelligent data analysis, which comprises the following steps: collecting a natural language query requirement; performing semantic analysis on the natural language query requirement to extract semantic constraint conditions; judging whether there is a matched inventory report in a preset report database based on the semantic constraint conditions; if there is a matched inventory report, obtaining a query result according to the semantic constraint conditions; if there is no matched inventory report, obtaining a query result based on a preset semantic conversion method; generating a visual chart based on the query result and outputting the visual chart to a preset chart viewing terminal. The application has the effect of improving query efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent data analysis, in particular to a large model-based intelligent data analysis method and system. BACKGROUND

[0002] Large model-based intelligent data analysis refers to the process of analyzing user needs and automatically completing data query, analysis and result presentation by means of the semantic understanding, knowledge reasoning and data processing capabilities of large-scale pre-training models.

[0003] At present, large model-based intelligent data analysis technology has been gradually applied in various fields. In related schemes, there have been attempts to analyze user natural language queries through large models, directly generate data processing instructions and call underlying data interfaces to obtain results. At the same time, some systems will automatically generate basic charts based on analysis results to assist users in understanding data.

[0004] However, the existing scheme fails to effectively combine the preset stock report resources. When the user query demand matches the existing report, it still relies on the large model to analyze and process data from scratch, thereby reducing the query efficiency. SUMMARY

[0005] In order to improve the query efficiency, the present application provides a large model-based intelligent data analysis method and system.

[0006] In the first aspect, the present application provides a large model-based intelligent data analysis method, which adopts the following technical scheme:

[0007] A large model-based intelligent data analysis method, comprising:

[0008] Collecting natural language query requirements;

[0009] Performing semantic analysis on the natural language query requirements to extract semantic constraint conditions;

[0010] Based on the semantic constraint conditions, it is determined whether there is a matching stock report in the preset report database;

[0011] If there is a matching stock report, the query result is obtained according to the semantic constraint conditions;

[0012] If there is no matching stock report, the query result is obtained based on the preset semantic conversion method;

[0013] Based on the query result, a visual chart is generated and output to a preset chart viewing terminal.

[0014] By adopting the technical scheme, the natural language query requirement of a user is first collected, and core semantic constraint conditions are extracted through semantic analysis; then, based on the semantic constraint conditions, a report database is searched to determine whether there is a stock report that can meet the requirement, if there is, the query result is directly generated based on the stock report combined with the constraint conditions, without the large model analyzing and processing data from scratch; if there is no matching stock report, a result is generated through a semantic conversion method. Finally, a visual chart is generated and output to a terminal, so that the stock report resources are preferentially reused, the query response time is greatly shortened, and the overall query efficiency is significantly improved.

[0015] Optionally, the method further comprises:

[0016] determining whether the semantic constraint conditions contain a preset hierarchical query feature;

[0017] if the hierarchical query feature is contained, obtaining semantic hierarchical logic based on the semantic constraint conditions and a preset multidimensional data model;

[0018] combining the semantic hierarchical logic and the multidimensional data model to obtain hierarchical model data dimensions;

[0019] calling corresponding hierarchical data from the hierarchical model data dimensions, and performing inter-hierarchical association analysis based on the corresponding hierarchical data to generate hierarchical association results;

[0020] updating the visual chart with the hierarchical association results and the query result, and outputting the updated visual chart to a chart viewing terminal.

[0021] Optionally, the semantic conversion method comprises:

[0022] when there is no matching stock report, matching database fields and semantic logic relationships from a preset business scenario semantic library based on the semantic constraint conditions;

[0023] combining the semantic constraint conditions, the database fields, the semantic logic relationships, and a preset database specification statement to generate a specification query statement;

[0024] performing syntax checking and data access permission checking on the specification query statement;

[0025] if the checking passes, querying from a preset original business database based on the specification query statement to obtain a query result;

[0026] if the checking does not pass, reporting a return requirement correction prompt.

[0027] Optionally, the method further comprises a method of judging whether there is a matching stock report in a preset report database based on the semantic constraint conditions:

[0028] extracting the business index and the data dimension based on the semantic constraint condition;

[0029] generating a demand feature vector based on the business index and the data dimension;

[0030] combining the demand feature vector with a preset report model feature library to calculate a vector matching degree;

[0031] if the vector matching degree is not lower than a preset matching degree threshold, it is determined that there is a matched inventory report;

[0032] if the vector matching degree is lower than the preset matching degree threshold, it is determined that there is no matched inventory report, and a preset semantic conversion method is executed to obtain a query result.

[0033] Optionally, the method further comprises a step after collecting the natural language query demand:

[0034] collecting a voice receiving signal of a preset voice receiving device;

[0035] when the voice receiving signal contains a preset human voice signal, extracting an exit query content and a specific frequency band auxiliary sound from the voice receiving signal;

[0036] obtaining an actual frequency band of the auxiliary sound based on the specific frequency band auxiliary sound;

[0037] obtaining a standard frequency band of the auxiliary sound based on the exit query content;

[0038] performing matching degree calculation on the actual frequency band of the auxiliary sound and the standard frequency band of the auxiliary sound to obtain a frequency band matching degree;

[0039] if the frequency band matching degree reaches a preset frequency band matching degree threshold, it is confirmed that the exit query content collection is accurate, and the step of performing semantic analysis on the natural language query demand to extract the semantic constraint condition is executed;

[0040] if the frequency band matching degree does not reach the preset frequency band matching degree threshold, a voice content verification exception prompt is reported.

[0041] Optionally, the method further comprises a processing method of the specific frequency band auxiliary sound:

[0042] collecting a mixed audio signal of the voice receiving device;

[0043] performing frequency domain analysis on the mixed audio signal to obtain a human voice frequency band, an environmental noise and an auxiliary sound frequency band;

[0044] determining a frequency band separation threshold based on the human voice frequency band and the auxiliary sound frequency band;

[0045] performing frequency band separation processing on the mixed audio signal based on the frequency band separation threshold to obtain a pure auxiliary sound frequency band signal;

[0046] Extract specific frequency components from the pure auxiliary audio frequency band signal and calculate the auxiliary amplitude characteristics of each frequency component;

[0047] Effective frequency components are selected by comparing the auxiliary amplitude features with the preset auxiliary amplitude threshold.

[0048] The effective frequency components are matched with a preset frequency mapping table to obtain the mapping semantic constraints. When the mapping semantic constraints are consistent with the semantic constraints, the semantic parsing step is executed.

[0049] Optional, also includes:

[0050] Collect interactive feedback data from visual charts;

[0051] The feedback feature vector is obtained based on interactive feedback data and semantic constraints.

[0052] Optimize parameters in response to feedback feature vectors to generate query results;

[0053] The parameters are iteratively adjusted based on the query results to optimize the query results;

[0054] The optimized query results are displayed in a visual chart, which is then pushed to the chart viewing terminal. The optimization parameters are stored in the preset model optimization library.

[0055] Optional, also includes:

[0056] Extracting voiceprint features of individuals based on human voice audio;

[0057] The speaker's voiceprint features are matched with a pre-stored user voiceprint database to determine the speaker's identity.

[0058] Query access permission levels based on the speaker's identity;

[0059] Combine access permission levels and query results to perform permission filtering to obtain filtered query results;

[0060] The system generates visual charts based on the filtered query results and outputs them to the chart viewing terminal.

[0061] Optional, also includes:

[0062] Extracting emotional features from human voice audio;

[0063] Based on speech emotion features to obtain the emotional state of the query;

[0064] When the emotional state is a preset emergency state, a priority processing prompt is reported, and specific priority information is determined based on the person's voiceprint characteristics;

[0065] In response to priority prompts, specific priority information is processed.

[0066] Secondly, this application provides an intelligent data analysis system based on a large model, employing the following technical solution:

[0067] A large-model-based intelligent data analysis system includes:

[0068] The data acquisition module is used to collect natural language query requests.

[0069] The memory is used to store the program that implements an intelligent data analysis method based on a large model;

[0070] The processor is used to load and execute programs stored in memory.

[0071] In summary, this application includes at least one of the following beneficial technical effects:

[0072] 1. By first collecting users' natural language query requirements, and then extracting core semantic constraints through semantic parsing, the system searches the report database based on these constraints to determine if any existing reports meet the requirements. If so, it quickly generates query results directly from the existing reports and constraints, eliminating the need for a large model to parse and process data from scratch. If no matching existing reports exist, the system then generates results through semantic transformation. Finally, it generates visual charts and outputs them to the terminal. By prioritizing the reuse of existing report resources, it significantly shortens query response time and greatly improves overall query efficiency.

[0073] 2. After collecting natural language query requests, the system first acquires the voice signal received by the voice receiving device. When human voice signals are detected, the spoken query content and auxiliary sound in a specific frequency band are extracted simultaneously. The accuracy of the voice acquisition is determined by calculating the matching degree between the actual frequency band of the auxiliary sound and the standard frequency band based on the spoken query content. If the matching degree is satisfactory, it indicates that the voice signal has not been severely interfered with, confirming that the spoken content acquisition is correct, and proceeding to the subsequent semantic analysis step. If the matching degree is unsatisfactory, a voice verification anomaly is immediately reported to avoid deviations in subsequent query results due to errors in the acquired content. This overcomes the limitations of traditional methods that rely solely on human voice signals to judge acquisition quality. Through secondary verification using auxiliary sound in a specific frequency band, the system significantly reduces the impact of environmental interference and signal distortion on the accuracy of the voice query content, ensuring that the input natural language query truly reflects the user's intent. This improves the accuracy of pre-processing intelligent data analysis and reduces invalid queries caused by content errors.

[0074] 3. By first acquiring mixed audio signals and separating the human voice, environmental noise, and auxiliary sound audio segments through frequency domain analysis, and determining the separation threshold based on the frequency band characteristics of the human voice and auxiliary sounds, a pure auxiliary sound audio segment signal is extracted, effectively eliminating environmental noise interference. Specific frequency components are extracted from the pure auxiliary sound signal, and their amplitude characteristics are analyzed. Valid frequency components that meet the preset amplitude threshold are selected, and then the valid frequency components are converted into corresponding semantic constraints using a frequency mapping table. Only when the mapping constraint matches the semantic constraint parsed from the spoken content is the accuracy of the voice query confirmed, and subsequent steps are executed. This significantly reduces the risk of single-voice recognition errors, making it particularly suitable for voice query scenarios in noisy environments, providing a more reliable pre-emptive guarantee for the accuracy of intelligent data analysis. Attached Figure Description

[0075] Figure 1 This is a flowchart of a method for intelligent data analysis based on a large model;

[0076] Figure 2 This is a flowchart of the steps following the collection of natural language query requests;

[0077] Figure 3 This is a flowchart of a method for processing auxiliary sound in a specific frequency band. Detailed Implementation

[0078] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0079] Reference Figure 1 This application discloses an intelligent data analysis method based on a large model, comprising the following steps:

[0080] S10: Collect natural language query requests.

[0081] Natural language query requests refer to data analysis requests made by users in their everyday conversational language.

[0082] Natural language query requests are collected through the system's preset input interfaces, including text input interfaces and voice input interfaces. Text input is obtained directly from the user's input text, while voice input is obtained by converting speech into text through an automatic speech recognition model. The input interfaces are developed and adapted by those skilled in the art based on the system's application scenarios, and will not be elaborated here.

[0083] S11: Perform semantic parsing on natural language query requirements to extract semantic constraints.

[0084] Semantic constraints refer to the key conditions that limit the scope, dimensions, and logic of data analysis extracted after semantic parsing of natural language query requirements.

[0085] The semantic constraints are obtained by processing natural language query requirements through a pre-set semantic parsing model. This model is fine-tuned based on large language models such as BERT and can realize entity recognition and relation extraction. The model training data is collected by those skilled in the art from natural language query samples and corresponding labeled constraints in various business scenarios. After labeling and cleaning, the data is used for model training. After training, the data is deployed to the system. Inputting natural language will automatically output the semantic constraints, which will not be elaborated here.

[0086] S12: Determine whether a matching existing report exists in the preset report database based on semantic constraints.

[0087] A report database is a database that pre-stores standardized data analysis reports that have already been generated. The reports need to be associated with metadata such as business scenarios, data dimensions, and calculation logic, and support quick retrieval and matching based on metadata.

[0088] The report database is a structured database built by those skilled in the art, and standardized reports generated in historical business are entered into the database in a preset format; the fields of the metadata table are defined by technical personnel in combination with business analysis needs, and the entry process is completed through batch import tools or manual entry, which will not be described in detail here.

[0089] Existing reports refer to standardized reports that have been pre-generated in the report database and can directly meet or adapt to user query needs.

[0090] The stock reports are initially determined by matching the metadata and semantic constraints in the report database. When technical personnel enter reports, they need to fill in the metadata completely. The system filters out candidate reports through keyword matching, and then determines whether they are matching stock reports after subsequent vector matching verification. This will not be elaborated on here.

[0091] S13: If a matching stock report exists, the query results are obtained based on the semantic constraints.

[0092] The query results refer to structured data results that conform to semantic constraints. The query results are generated by first performing precise field matching between the core elements of the semantic constraints and the metadata of the existing reports to determine the corresponding data fields in the existing reports. Then, the underlying data of the existing reports is extracted and calculated in a targeted manner. During the extraction process, the system automatically verifies the consistency of the data format, ultimately generating structured data results that conform to the semantic constraints. The field mapping rules and logical processing logic of the script are pre-configured by technical personnel based on the data structure of the existing reports, and automated testing ensures accurate extraction of target data under different constraints; these details are not elaborated here.

[0093] If a matching existing report exists, it means that the user's query requirements are highly compatible with the existing standardized reports in the report database in terms of business indicators, data dimensions, and logical relationships. There is no need to recalculate from the original business data; the target information can be extracted directly from the existing report to quickly obtain the query results.

[0094] S14: If no matching stock report exists, the query results will be obtained based on the preset semantic transformation method.

[0095] Semantic transformation refers to a standardized process that transforms semantic constraints into executable database query statements and calculates query results from the original business data when there are no matching existing reports in the report database.

[0096] The specific semantic transformation methods will be explained in detail in S30 to S34, and will not be repeated here.

[0097] If no matching existing report exists, it means that the user's query requirements and the existing standardized reports in the report database are not highly compatible in terms of business indicators, data dimensions, and logical relationships, and data cannot be extracted directly. It is necessary to first use semantic transformation methods to convert natural language semantics into database executable statements, and then calculate the query results from the original business data.

[0098] S15: Generate a visual chart based on the query results and output it to the preset chart viewing terminal.

[0099] Visual charts are media that present query results in an intuitive, graphical form.

[0100] Visual charts are generated through the system's integrated visualization tools. Technical personnel preset the corresponding rules for "query result type → chart type". The system automatically matches the chart type based on the field type of the query result, calls the tool's API to generate the chart, and the chart style is configured according to the preset visual specifications, which will not be elaborated here.

[0101] A chart viewing terminal refers to a device or platform used to display visual charts.

[0102] The chart viewing terminal is pre-configured by those skilled in the art and selected and adapted according to the system application scenario, which will not be elaborated here.

[0103] Also includes:

[0104] S20: Determine whether the semantic constraints contain preset hierarchical query features.

[0105] Hierarchical query features refer to features that contain multi-level progressive analysis requirements in semantic constraints, such as "first break down sales by region, then break down sales by product category in each region" or "first calculate total annual profit, then break down quarterly profit and departmental profit", which need to be met through multi-level association analysis.

[0106] The hierarchical query features are identified through a pre-set keyword matching library. Technical personnel compile commonly used keywords for hierarchical queries (such as "first...then...", split, layer by layer, by...then by...") to build a keyword library. The system performs text segmentation based on semantic constraints. If the segmentation results contain keywords from the library and conform to the "hierarchical progression" grammatical structure, it is determined that the hierarchical query features are included. The keyword library needs to be updated regularly in accordance with business query habits, which will not be elaborated here.

[0107] S21: If hierarchical query features are included, semantic hierarchical logic is obtained based on semantic constraints and a preset multidimensional data model.

[0108] A multidimensional data model is a structured data model built based on business scenarios, which includes multiple analytical dimensions (such as time, region, product, and customer) and business indicators (such as sales, profit, and order volume).

[0109] The multidimensional data model is designed by those skilled in the art based on business needs using a star schema or snowflake schema, and is constructed using a "fact table + dimension table"; the model relationships are defined by those skilled in the art according to business data logic, supporting multidimensional queries and hierarchical decomposition, which will not be elaborated here.

[0110] Semantic hierarchical logic refers to the hierarchical analysis order and association rules determined based on hierarchical query features and multidimensional data models.

[0111] The semantic hierarchical logic is determined by combining the hierarchical keywords in the semantic constraints with the dimensional hierarchical relationship of the multidimensional data model. Technical personnel preset the "coarse and fine granularity" of each dimension in the multidimensional data model (e.g., regional dimension: national > provincial > municipal, time dimension: year > quarter > month). The system matches the granularity order of the dimensions in the model with the keywords in the hierarchical query features to generate hierarchical logic from coarse to fine. The logic must ensure that each level of dimension can be associated in the model, which will not be elaborated here.

[0112] If the query features are hierarchical, it means that the user's query needs are not based on a single-dimensional basic data analysis, but rather require multi-level progressive decomposition to uncover data relationships. Therefore, it is necessary to first obtain the semantic hierarchical logic for subsequent steps.

[0113] S22: Combine semantic hierarchical logic and multidimensional data model to obtain hierarchical model data dimensions.

[0114] The data dimension of a hierarchical model refers to the exclusive data dimension of each level that is separated from a multidimensional data model based on semantic hierarchical logic.

[0115] The hierarchical model data dimensions are obtained by extracting the corresponding dimensions and related indicators for each level from the multidimensional data model according to the hierarchical order of semantic hierarchical logic. The technical personnel have defined the "dimension-indicator" relationship during model design. When the system extracts dimensions, it simultaneously matches the related indicators and determines the granularity of each level's dimensions, ultimately obtaining independent and associative hierarchical model data dimensions for each level. This ensures that subsequent levels of data can be decomposed and correlated layer by layer, which will not be elaborated upon here.

[0116] S23: Retrieve the corresponding level data from the hierarchical model data dimension, and perform inter-level correlation analysis based on the corresponding level data to generate hierarchical correlation results.

[0117] Corresponding level data refers to the original data that belongs to a specific level and is retrieved from the data dimensions of the hierarchical model.

[0118] The corresponding hierarchical data is obtained by retrieving data from the multidimensional model of the data warehouse according to the hierarchical order of the data dimensions in the hierarchical model. The retrieval logic is implemented through SQL queries. The method of using SQL queries is common knowledge in this field and will not be elaborated here.

[0119] Hierarchical association results refer to the results obtained after performing association analysis on data at each level.

[0120] The hierarchical association results are obtained by processing the intermediate datasets at each level through a preset association analysis algorithm. The algorithm includes summary calculation, trend comparison, and attribution analysis. The algorithm logic is implemented by technical personnel based on business analysis requirements and will not be elaborated here.

[0121] S24: Combine the hierarchical association results with the query results to update the visualization chart, and output the updated visualization chart to the chart viewing terminal.

[0122] Add the hierarchical association results to the original chart to generate a combined chart containing the hierarchical association conclusions, and then push it to the chart viewing terminal.

[0123] Semantic transformation methods include:

[0124] S30: When no matching existing report exists, the database fields and semantic logical relationships are matched from the preset business scenario semantic library based on semantic constraints.

[0125] A business scenario semantic library refers to a pre-built knowledge base that stores the corresponding rules of "natural language semantics - database fields - logical relationships" under different business scenarios.

[0126] The business scenario semantic library is compiled by professionals in collaboration with business personnel, who sort out the core semantics, corresponding database fields, and logical relationships of each business scenario. The data is organized into a structured data table in the format of "scenario-semantics-fields-logic" and stored in the database. The contents of the library need to be updated regularly, and the update cycle is set in advance by professionals. When a new business scenario or field is added, the technical personnel will supplement and enter the corresponding rules, which will not be elaborated here.

[0127] A database field is a column name used to store specific data.

[0128] Database fields are directly obtained from the table structure of the business database. When creating a business table, the database administrator defines the field name, data type, and meaning. The naming rules and data type selection of the fields are common knowledge in the field. When building a business scenario semantic library, the technical personnel associate and match natural language semantics with these fields. The matching rules are set in advance by those skilled in the art to ensure accurate mapping of semantics to fields, which will not be elaborated here.

[0129] Semantic logical relationships refer to the data calculation, filtering, and comparison logic implicit in natural language query requirements.

[0130] Semantic logical relationships are obtained by technicians converting natural language logic into executable syntax or functions for the database. These conversion rules are stored in a business scenario semantic library. The specific content of the conversion rules is set in advance by those skilled in the art, and the system automatically calls the corresponding logic when matching semantics, which will not be elaborated here.

[0131] S31: Combine semantic constraints, database fields, semantic logical relationships, and preset database specification statements to generate a standard query statement.

[0132] Database standard statements refer to query statement templates that conform to database syntax standards.

[0133] Database standard statements are standardized statement templates written by technical personnel based on database type and common query scenarios, and will not be elaborated here.

[0134] A standard query statement is a query statement that can be executed directly, generated by filling a database standard query statement template based on semantic constraints, database fields, and semantic logical relationships.

[0135] The standardized query statement is generated through the system's "parameter filling" logic. First, the database fields and semantic logical relationships corresponding to the semantic constraints are matched from the business scenario semantic library. Then, these parameters are filled into the database standardized statement template according to their positions. During the filling process, table associations and data format conversions are automatically handled. The data format conversion rules are set in advance by those skilled in the art. After generation, the syntax integrity is initially verified, which will not be elaborated here.

[0136] S32: Perform syntax validation and data access permission checks on the standard query statements.

[0137] Syntax validation is performed using an SQL syntax checking tool. The tool selection is common knowledge in the field, and the tool automatically detects syntax errors in the standard query statement. Permission checks are performed by comparing the user permission table. The configuration rules of the user permission table are set in advance by those skilled in the art. If the statement contains fields or tables that the user does not have permission to access, the permission check is deemed to have failed. This will not be elaborated on here.

[0138] S33: If the verification passes, the query results are obtained from the preset original business database based on the standard query statement.

[0139] If the verification passes, the query results corresponding to the standardized query statement must be retrieved directly from the original business database. The original business database stores the original business data. Once the standardized query statement is obtained, logical calculations must be performed in conjunction with the original business data in the original business database to derive the query results. The specific logical calculation process is common knowledge to those skilled in the art and will not be elaborated upon here. The original business database is formed by those skilled in the art writing the original business data in advance, and will not be elaborated upon here either.

[0140] S34: If the verification fails, a request for correction will be sent back.

[0141] If the validation fails, it indicates that there is an exception in the standard query statement that prevents it from being executed, and it needs to be reported and a correction prompt should be returned.

[0142] It also includes a method based on semantic constraints to determine whether a matching existing report exists in a pre-defined report database:

[0143] S40: Extract business metrics and data dimensions based on semantic constraints.

[0144] Business metrics refer to the core analytical indicators that users focus on in their query needs.

[0145] Business metrics are extracted using the entity recognition function of the semantic parsing model. During model training, technicians label "metric entities". The rules for entity labeling are set in advance by those skilled in the art. The model can automatically identify and output such entities from semantic constraints. If there are ambiguous expressions, they are corrected by synonym mapping in the business scenario semantic library. The synonym mapping rules are set in advance by those skilled in the art to ensure the accuracy of the extracted metrics, which will not be elaborated here.

[0146] The data dimension for demand extraction refers to the analytical perspective used to limit or break down business metrics in user query requirements.

[0147] The data dimensions extracted from requirements are consistent with the logic for extracting business metrics, and are extracted through the entity recognition function of the semantic parsing model, which will not be elaborated here.

[0148] S41: Extract data dimensions based on business metrics and requirements to generate a requirement feature vector.

[0149] The demand feature vector refers to the numerical vector transformed from the extracted business indicators and demand extraction data dimensions.

[0150] The demand feature vector is generated using a combination of encoding and embedding. Business indicators and discrete dimensions are encoded using One-Hot encoding, with encoding rules set in advance by those skilled in the art. Continuous or fine-grained dimensions are converted into low-dimensional vectors using a pre-trained embedding model, with the selection of the embedding model set in advance by those skilled in the art. The encoding rules and embedding model are determined by technical personnel, and the embedding model is pre-trained based on business data to ensure that the vector accurately reflects the feature differences, which will not be elaborated here.

[0151] S42: Combine the requirement feature vector with the preset report model feature library to calculate the vector matching degree.

[0152] The report model feature library refers to a database that stores the feature vectors of "business indicators - data dimensions extracted from requirements" of existing reports. Each existing report corresponds to a unique feature vector.

[0153] When technicians enter existing reports into the report database, they simultaneously extract the business indicators and data dimensions of the reports according to the requirements. They then generate feature vectors using the "encoding + embedding" method of S41 and store them in a dedicated feature library, namely the report model feature library. The vector generation logic is completely consistent with the requirement feature vectors. The parameter settings of the generation logic are set in advance by those skilled in the art to ensure the fairness of subsequent matching, which will not be elaborated here.

[0154] Vector matching degree refers to the degree of similarity between the demand feature vector and the existing report feature vectors in the report model feature library.

[0155] Vector matching degree is calculated using the cosine similarity algorithm. The implementation details of the cosine similarity algorithm are common knowledge in the field. The algorithm is implemented by technical personnel by writing code. The system traverses all vectors in the feature library of the report model, calculates the similarity with the required vector one by one, and outputs the matching degree of all vectors. It will not be elaborated here.

[0156] S43: If the vector matching degree is not lower than the preset matching degree threshold, then it is determined that there is a matching existing report.

[0157] The matching threshold refers to a preset critical value for determining whether a requirement matches an existing report. The matching threshold is set in advance by those skilled in the art and will not be elaborated upon here.

[0158] If the vector matching degree is not lower than the matching degree threshold, it means that the feature vector corresponding to the user's query needs is highly compatible with the feature vector of a certain existing report in the report model feature library in terms of core analysis elements, and it can be determined that there is a matching existing report.

[0159] S44: If the vector matching degree is lower than the preset matching degree threshold, it is determined that there is no matching stock report, and the preset semantic transformation method is executed to obtain the query result.

[0160] If the vector matching degree is lower than the matching degree threshold, it means that the feature vector corresponding to the user's query requirement does not meet the preset standard in terms of the core analysis elements of the feature vectors of all existing reports in the report model feature library. It is necessary to determine that there are no matching existing reports and execute step S14 to obtain the query results through semantic transformation.

[0161] Reference Figure 2 It also includes steps following the collection of natural language query requests:

[0162] S50: Collect the voice reception signal from the preset voice receiving device.

[0163] A voice receiving device is a hardware device used to collect users' verbal query requests.

[0164] The voice receiving device is selected and adapted by those skilled in the art according to the system application scenario, and will not be described in detail here.

[0165] Voice reception signal refers to the raw audio electrical signal collected by the voice receiving device.

[0166] The voice signal is obtained by the voice receiving device acquiring the ambient audio in real time and converting the sound vibrations into electrical signals.

[0167] S51: When the voice received signal contains a preset human voice signal, extract the spoken query content and specific frequency band auxiliary sound from the voice received signal.

[0168] Human voice signal refers to the audio signal produced by human voice in the speech reception signal. The human voice signal is preset by those skilled in the art and will not be described in detail here.

[0169] Verbal query content refers to the query needs expressed verbally by users, extracted from human voice signals.

[0170] The spoken query content is transcribed into text using an automatic speech recognition model. The method of speech-to-text conversion is common knowledge in this field and will not be elaborated upon here.

[0171] Specific frequency band auxiliary sound refers to auxiliary audio with a specific frequency range that is preset to verify the accuracy of the oral query content. It is collected synchronously with the user's oral speech and is used to assist in verification.

[0172] The auxiliary sound in a specific frequency band is emitted synchronously by the auxiliary sound-generating device associated with the system control. The frequency range is preset by the technicians, and the set value of the frequency range is set in advance by those skilled in the art to ensure that it does not overlap with the human voice frequency band and is easy to separate. It is stored synchronously with the human voice signal during acquisition. The storage path is set in advance by those skilled in the art as a reference signal for subsequent verification, which will not be elaborated here.

[0173] S52: Obtain the actual frequency band of the auxiliary sound based on a specific frequency band auxiliary sound.

[0174] The actual frequency band of auxiliary sound refers to the actual frequency range of the auxiliary sound extracted from the voice reception signal, reflecting the characteristics of the auxiliary sound actually collected.

[0175] The actual frequency band of the auxiliary sound is obtained by the system performing frequency domain analysis on the auxiliary sound signal in a specific frequency band. Technicians write frequency domain analysis algorithms, and the parameters of the frequency domain analysis are set in advance by those skilled in the art. The frequency spectrum of the signal is extracted, and the frequency range in which the energy is concentrated in the spectrum is identified. This range is the actual frequency band of the auxiliary sound. During the analysis process, frequency interference from human voices and noise is eliminated. The logic for interference elimination is set in advance by those skilled in the art to ensure accurate frequency band extraction, which will not be elaborated here.

[0176] S53: Obtain the auxiliary sound standard frequency band based on the verbal query content.

[0177] The standard frequency band for auxiliary sound refers to the reference frequency range determined by performing frequency domain analysis on the standard auxiliary sound samples associated with the user's verbal query content. It is used to characterize the standard features that auxiliary sound should have in this type of verbal query scenario.

[0178] The standard frequency band for auxiliary audio is determined by technicians during the system deployment phase, after standardizing a pre-built standard auxiliary audio sample library for different types of spoken query content. First, typical auxiliary audio samples corresponding to each type of spoken query content are collected. Using the same frequency domain analysis algorithm as the actual frequency band extraction in S52, frequency intervals with concentrated and stable energy are extracted from the samples. After multiple repeatability verifications to remove outliers, these intervals are solidified as the standard frequency band for the corresponding spoken query content. The coverage of the sample library, the parameter configuration of the analysis algorithm, and the fault tolerance threshold of the standard frequency band are all set by those skilled in the art based on the actual application scenario, serving as the benchmark for subsequent matching and verification with the actual frequency band of the auxiliary audio; these details are not elaborated here.

[0179] S54: Calculate the matching degree between the actual frequency band of the auxiliary sound and the standard frequency band of the auxiliary sound to obtain the frequency band matching degree.

[0180] Frequency band matching degree refers to the degree of overlap between the actual frequency band of the auxiliary sound and the standard frequency band. The frequency band matching degree can be obtained using a matching degree calculation formula. This formula is common knowledge in the field and will not be elaborated upon here.

[0181] S55: If the frequency band matching degree reaches the preset frequency band matching degree threshold, then the accuracy of the oral query content collection is confirmed, and the step of performing semantic parsing on the natural language query requirements to extract semantic constraints is executed.

[0182] The frequency band matching threshold is a critical value used to determine whether the auxiliary sound has been accurately acquired. The frequency band matching threshold is set in advance by those skilled in the art and will not be elaborated here.

[0183] If the frequency band matching degree reaches the frequency band matching degree threshold, it indicates that the oral query content is accurately collected, and the subsequent steps of semantic parsing of the natural language query requirements to extract semantic constraints can be performed, namely S11 to S15.

[0184] S56: If the frequency band matching degree does not reach the preset frequency band matching degree threshold, a voice content verification error prompt will be reported.

[0185] If the frequency band matching degree does not reach the frequency band matching degree threshold, it indicates that the oral query content is not collected accurately, and an abnormal voice content verification prompt should be reported.

[0186] Reference Figure 3 It also includes methods for processing auxiliary sound in specific frequency bands:

[0187] S60: Acquires mixed audio signals from the voice receiving device.

[0188] Mixed audio signals refer to unfiltered raw audio signals that contain human voices, environmental noise, and auxiliary sounds in specific frequency bands.

[0189] The mixed audio signal is obtained by directly acquiring all sound components on site using a voice receiving device.

[0190] S61: Perform frequency domain analysis on the mixed audio signal to obtain the human voice audio, ambient noise and auxiliary sound audio segments.

[0191] Human voice audio refers to audio signals that are separated from mixed audio signals and contain only human voices.

[0192] Human voice audio is obtained through frequency domain analysis combined with bandpass filtering. The value of the bandpass filtering range is set in advance by those skilled in the art. The audio components of this frequency band are extracted from the mixed signal. At the same time, residual noise in the frequency band is removed by a noise reduction algorithm. The parameters of the noise reduction algorithm are set in advance by those skilled in the art to obtain pure human voice audio, which will not be elaborated here.

[0193] Environmental noise refers to irrelevant audio signals in a mixed audio signal, excluding human voices and auxiliary sounds in specific frequency bands, such as wind noise, background talking, and equipment electrical noise, which need to be removed through frequency band separation.

[0194] Environmental noise is obtained by identifying non-human voice and non-auxiliary sound frequency bands through frequency domain analysis. Technicians set noise identification rules, the specific content of which is set in advance by those skilled in the art. The components of these frequency bands are extracted from the mixed signal and marked as environmental noise for subsequent filtering processing, which will not be elaborated here.

[0195] The auxiliary sound frequency band refers to the frequency range corresponding to the auxiliary sound in a specific frequency band, and it is the basis for separating the auxiliary sound from the mixed audio signal.

[0196] The auxiliary sound frequency band is obtained by extracting the energy concentration range in the mixed signal and conforming to the preset auxiliary sound frequency range through frequency domain analysis. Technicians write frequency band identification algorithms, and the sensitivity parameters of frequency band identification are set in advance by those skilled in the art. The range of energy peaks in the location signal that are close to the standard frequency of the auxiliary sound is the auxiliary sound frequency band, which will not be elaborated here.

[0197] S62: Determine the frequency band separation threshold based on the human voice frequency band and the auxiliary audio frequency band.

[0198] The frequency band separation threshold refers to the critical frequency value used to distinguish between the human voice frequency band, the auxiliary sound frequency band, and the environmental noise frequency band.

[0199] The frequency band separation threshold is determined based on the upper limit of the human voice frequency band and the upper / lower limit of the auxiliary sound frequency band. The technicians set the threshold as (lower limit of the human voice frequency band + upper limit of the auxiliary sound frequency band) / 2. The calculation formula of the threshold is set in advance by the technicians to ensure that the threshold is in the blank area between the human voice and auxiliary sound frequency bands to avoid mutual interference during separation. This will not be elaborated here.

[0200] S63: Perform frequency band separation processing on the mixed audio signal based on the frequency band separation threshold to obtain a pure auxiliary audio band signal.

[0201] Pure auxiliary sound audio band signal refers to an audio signal that contains only auxiliary sound in a specific frequency band after frequency band separation processing to remove human voice and environmental noise.

[0202] The pure auxiliary audio band signal is obtained by separating the frequency bands according to the threshold, using a bandpass filtering algorithm to extract the auxiliary audio band components in the mixed signal, and using a notch filter to remove any residual human voices in the frequency band. Then, the signal amplitude is normalized by gain adjustment, and finally, a pure auxiliary audio band signal without interference is obtained. The frequency band parameters of the bandpass filter, the frequency parameters of the notch filter, and the target value of amplitude normalization are set in advance by those skilled in the art and will not be described in detail here.

[0203] S64: Extract specific frequency components from the pure auxiliary audio frequency band signal and calculate the auxiliary amplitude characteristics of each frequency component.

[0204] A specific frequency component refers to a signal component with one or more representative frequency points extracted from a pure auxiliary audio frequency band signal.

[0205] Specific frequency components are obtained by performing spectral analysis on the pure auxiliary sound audio band signal. Technicians write peak detection algorithms to identify the top 3 to 5 frequency points with the highest energy in the spectrum. These frequency points are the specific frequency components, ensuring that the extracted components can represent the core characteristics of the auxiliary sound. The sensitivity of peak detection and the number of frequency points are set in advance by those skilled in the art, and will not be elaborated here.

[0206] Auxiliary amplitude characteristics refer to the amplitude of a specific frequency component.

[0207] The auxiliary amplitude feature is obtained by calculating the amplitude value of a specific frequency component in the spectrum. The system reads the energy value corresponding to each frequency point, converts it into amplitude units, and records the amplitude data of each specific frequency component as the basis for subsequent screening of effective components. The unit conversion coefficient is set in advance by those skilled in the art and will not be elaborated here.

[0208] S65: Compare the auxiliary amplitude features with the preset auxiliary amplitude threshold to filter out the effective frequency components.

[0209] The auxiliary amplitude threshold is a critical amplitude value used to determine whether a specific frequency component is valid. The auxiliary amplitude threshold is set in advance by those skilled in the art and will not be elaborated here.

[0210] Effective frequency components refer to specific frequency components that meet the auxiliary amplitude threshold requirements.

[0211] The effective frequency components are obtained by comparing the auxiliary amplitude characteristics of each specific frequency component with the auxiliary amplitude threshold. The system retains frequency components with amplitude ≥ the threshold and removes components with amplitude < the threshold, thus obtaining the core effective frequency components. The logic for interference determination is set in advance by those skilled in the art and will not be elaborated here.

[0212] S66: Match the effective frequency components with the preset frequency mapping table to obtain the mapping semantic constraints. When the mapping semantic constraints are consistent with the semantic constraints, execute the semantic parsing step.

[0213] A frequency mapping table is a pre-defined table of correspondences between effective frequency components and semantic constraints, used to transform auxiliary sound features into semantic constraints.

[0214] The frequency mapping table is established by those skilled in the art based on the different values ​​of the effective frequency components and common semantic constraints. The specific content of the mapping relationship is set in advance by those skilled in the art and organized into a structured table in the format of "frequency value - semantic constraint type - constraint content". It is stored in the system database and the mapping relationship is updated periodically according to business needs. The triggering conditions for the update are set in advance by those skilled in the art, which will not be elaborated here.

[0215] The semantic constraints are semantic constraints obtained by matching the effective frequency components with the frequency mapping table. They are used for cross-validation with the semantic constraints extracted from the oral query content to ensure that the constraints are accurate.

[0216] The mapping semantic constraints are obtained by matching each effective frequency component with a frequency mapping table. For example, the effective component 1.01kHz matches "time dimension: month" and 1.02kHz matches "regional dimension: East China". After integration, a complete mapping semantic constraint is formed. If the constraint is consistent with the constraint extracted from the spoken content, the semantic parsing step is executed. The logic of consistency judgment is set in advance by those skilled in the art and will not be elaborated here.

[0217] Also includes:

[0218] S70: Collects interactive feedback data from visual charts.

[0219] Interactive feedback data refers to the data generated when users interact with the generated visualizations. This data is obtained in real-time by the system's front-end module, which records user actions.

[0220] S71: Obtain the feedback feature vector based on interactive feedback data and semantic constraints.

[0221] Feedback feature vectors are quantitative vectors formed by combining interactive feedback data with initial semantic constraints, used to describe users' optimization needs for query results.

[0222] The feedback feature vector is generated using a combination of encoding and embedding. The operation type in the interactive feedback data is encoded using One-Hot encoding. The operation parameters and initial semantic constraints are converted into low-dimensional vectors using a pre-trained embedding model. The two types of vectors are concatenated to obtain the feedback feature vector, ensuring that the vector can fully reflect the user's interaction intent. The encoding dimension and the parameters of the embedding model are set in advance by those skilled in the art and will not be elaborated here.

[0223] S72: Optimize parameters in response to feedback feature vectors to generate query results.

[0224] Query result optimization parameters refer to parameters used to adjust query results.

[0225] The optimization parameters for the query results are generated by the system response feedback feature vector. Technical personnel preset the corresponding rules for "feedback feature vector type - optimization parameters" (such as "filtering operation + time parameter" - "adjust time dimension granularity"). The system outputs specific optimization parameters according to the vector matching rules. The parameter format must be compatible with the subsequent query result adjustment logic. The specific content and format standards of the corresponding rules are set in advance by those skilled in the art and will not be elaborated here.

[0226] S73: Optimize the query results by iteratively adjusting the parameters based on the query results.

[0227] Optimizing query results refers to adjusting the original query results by applying optimized parameters to obtain new results that better meet user interaction needs.

[0228] The optimized query results are obtained by re-executing the query logic according to the optimized parameters. The specific iterative adjustment method is common knowledge in this field and will not be elaborated here.

[0229] S74: Push the visualization chart corresponding to the optimized query results to the chart viewing terminal and store the optimization parameters in the preset model optimization library.

[0230] The model optimization library is a database that stores optimization parameters for query results and corresponding scenario information. It is used for automatic optimization of subsequent similar query needs to avoid repeated interactions.

[0231] The model optimization library is a structured database built by those skilled in the art, with a field structure of "user ID - original requirement characteristics - optimization parameters - application effect". The system stores the parameters and corresponding scenario information for each optimization in the library, and periodically analyzes the reusability of parameters through algorithms to improve the efficiency of subsequent optimizations. The definition of fields and the selection of algorithms are set in advance by those skilled in the art, and will not be elaborated here.

[0232] The optimized query results are used to generate a visual chart in the same way as S15, and then pushed to the chart viewing terminal. At the same time, the optimization parameters are stored in the model optimization library.

[0233] Also includes:

[0234] S80: Extracts voiceprint features of people based on human voice audio.

[0235] Voiceprint features refer to the unique voiceprint biometric features extracted from human voice audio.

[0236] Voiceprint features are extracted from human audio using a voiceprint extraction algorithm. The voiceprint extraction algorithm is common knowledge in this field and will not be elaborated upon here. The collection and processing of voiceprints were legal and authorized by the user.

[0237] S81: Match the speaker's voiceprint features with the pre-stored user voiceprint database to determine the speaker's identity.

[0238] The user voiceprint database refers to a database that pre-stores the correspondence between the "voiceprint features and identity information" of system users.

[0239] The user voiceprint database is created by collecting voiceprint samples from registered users of the system using techniques skilled in the art. The voice signals are purified through preprocessing algorithms, and stable voiceprint feature vectors are extracted using a voiceprint feature extraction model. These vectors are then associated with user identity information and stored in the database. If user information changes subsequently, technicians update the corresponding entries in the database through a pre-defined interface. The sample collection duration, feature extraction model parameters, and database storage format are all set by technicians based on system performance requirements and serve as the baseline data source for identity matching in S81; these details are not elaborated upon here.

[0240] Speaker identity refers to the specific identity information of the user making the voice query, determined by matching the extracted voiceprint features with the user voiceprint database. The user voiceprint database stores the correspondence between speaker identities and voiceprint features.

[0241] S82: Query access permission level based on the speaker's identity.

[0242] Access permission levels refer to the hierarchical levels of access permissions to data, preset according to user identity.

[0243] Access permission levels are obtained by querying a pre-defined user permission table. This table records the mapping between a system user's unique identifier and their corresponding access permission level, along with the specific access scope and validity period for that level. These records are pre-configured by those skilled in the art based on business needs and stored in a structured data table, enabling quick retrieval of corresponding permission level information using the user's identifier. Further details are omitted here.

[0244] S83: Combine access permission levels and query results to perform permission filtering to obtain filtered query results.

[0245] Filtered query results refer to the results obtained by removing sensitive information and limiting the data range of the original query results based on the user's access permission level.

[0246] The filtering of query results obtains filtered query results that comply with user permissions by removing sensitive information outside the user's permissions and restricting the data range. The logic of removal and restriction is set in advance by those skilled in the art and will not be elaborated here.

[0247] S84: Generate a visual chart based on the filtered query results and output it to the chart viewing terminal.

[0248] This step is the same as S15 above, and will not be repeated here.

[0249] Also includes:

[0250] S90: Extract speech emotion features from human voice audio.

[0251] Voice emotion features refer to the features extracted from human voice audio that reflect the user's emotional state.

[0252] Voice emotion features are obtained from human voice audio through emotion feature extraction algorithms, which are common knowledge in this field and will not be elaborated here.

[0253] S91: Obtain the emotional state of the query based on speech emotion features.

[0254] The emotional state of a query refers to the emotional state of a user during a query, determined based on the emotional characteristics of their voice. This includes an urgent state (fast speech, high pitch, loud volume), a calm state (stable speech, moderate pitch, no abrupt pauses), and a questioning state (frequent pauses in speech, large fluctuations in pitch), which is used to determine whether the query needs to be prioritized.

[0255] The emotional state query is obtained by processing speech emotional features through a pre-defined emotional classification model. Technical personnel first supplement the data with business-related speech emotional samples based on publicly available speech emotional feature datasets and the system application scenarios, then fine-tune the basic classification model to adapt it to the emotional recognition needs of the business scenarios.

[0256] During model training, technicians set parameters such as training rounds and learning rate (the parameter values ​​are set in advance by technicians based on the model's convergence effect). The voice emotion feature vector (such as a multi-dimensional vector composed of speech rate, pitch, and volume) is used as input, and the corresponding emotion state category (urgent, calm, questioning, etc.) is output. At the same time, technicians preset the confidence threshold for emotion classification (such as 80%, if it is lower than the threshold, it is judged as "emotion state is unclear"). Only when the confidence of the emotion category output by the model meets the standard is it finally determined as the query emotion state to ensure the reliability of the judgment result. This will not be elaborated here.

[0257] S92: When the query emotion state is a preset emergency state, report the priority prompt and determine the specific priority information based on the person's voiceprint characteristics.

[0258] An emergency state refers to a user's emotional state that requires priority handling by the system. Emergency states are pre-defined by those skilled in the art and will not be elaborated upon here.

[0259] Specific priority information refers to relevant information that needs to be processed first.

[0260] First, based on the user's voiceprint characteristics, the system matches the corresponding speaker's identity in a pre-stored user voiceprint database (the matching logic is the same as S81, and will not be elaborated here). Then, the system retrieves a preset "User Identity - Emergency Needs Priority Handling Rule Table", which is pre-set by those skilled in the art and records the priority handling direction for users with different identities in emergency situations.

[0261] For example, “Urgent needs of management users – prioritize outputting core indicator summary data + concise visualization charts” and “Urgent needs of ordinary employees – prioritize ensuring data accuracy and output basic detailed data”; finally, based on the matched user identity, the corresponding priority processing items are extracted from the rule table and integrated into specific priority information.

[0262] When the emotional status is queried as "urgent," it indicates that the user's need is quite urgent and should be reported for priority processing. Further priority information should be determined for subsequent steps.

[0263] S93: Respond to processing priority prompts to process specific priority information.

[0264] When a priority processing prompt is received, it should be processed according to the specific priority information to ensure that urgent needs are handled efficiently.

[0265] Based on the same inventive concept, embodiments of the present invention provide an intelligent data analysis system based on a large model, comprising:

[0266] The data acquisition module is used to collect natural language query requests, voice reception signals, mixed audio signals, and interactive feedback data.

[0267] The memory is used to store the program that implements an intelligent data analysis method based on a large model;

[0268] The processor is used to load and execute programs stored in memory.

[0269] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0270] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A smart data analysis method based on a large model, characterized in that, include: Collect natural language query requests; Semantic parsing of natural language query requirements is performed to extract semantic constraints; Based on semantic constraints, determine whether there are matching existing reports in the preset report database. If a matching existing report exists, the query results are obtained based on semantic constraints. If no matching existing report exists, the query results will be obtained based on the preset semantic transformation method. The query results are used to generate visual charts, which are then output to a preset chart viewing terminal. The semantic transformation method includes: When no matching existing report exists, database fields and semantic logical relationships are matched from a preset business scenario semantic library based on semantic constraints. Combine semantic constraints, database fields, semantic logical relationships, and preset database specification statements to generate standardized query statements; Perform syntax validation and data access permission checks on standard query statements; If the verification passes, the query results will be obtained from the preset original business database based on the standard query statement; If the verification fails, a request for revision will be submitted. It also includes steps that follow the collection of natural language query requests: Collect the voice signal received by the preset voice receiving device; When the received voice signal contains a preset human voice signal, extract the spoken query content and specific frequency band auxiliary sound from the received voice signal; The actual frequency band of the auxiliary sound is obtained by using auxiliary sound in a specific frequency band. The auxiliary sound standard frequency band is obtained based on the verbal query content; The matching degree between the actual frequency band of the auxiliary sound and the standard frequency band of the auxiliary sound is calculated to obtain the frequency band matching degree. If the frequency band matching degree reaches the preset frequency band matching degree threshold, it is confirmed that the oral query content is collected accurately, and the step of performing semantic parsing on the natural language query requirements to extract semantic constraints is executed. If the frequency band matching degree does not reach the preset frequency band matching degree threshold, a voice content verification error message will be reported.

2. The intelligent data analysis method based on a large model according to claim 1, characterized in that, Also includes: Determine whether the semantic constraints contain preset hierarchical query features; If hierarchical query features are included, semantic hierarchical logic is obtained based on semantic constraints and a pre-defined multidimensional data model. Combining semantic hierarchical logic and multidimensional data models to obtain the hierarchical model data dimensions; Retrieve data from the corresponding level of the hierarchical model data dimension, and perform inter-level correlation analysis based on the corresponding level data to generate hierarchical correlation results; The hierarchical association results are used to update the visualization chart, and the updated visualization chart is output to the chart viewing terminal.

3. The intelligent data analysis method based on a large model according to claim 1, characterized in that, It also includes a method based on semantic constraints to determine whether a matching existing report exists in a pre-defined report database: Extract business metrics and data dimensions based on semantic constraints; Data dimensions are extracted based on business metrics and requirements to generate a requirement feature vector; The vector matching degree is calculated by combining the demand feature vector with the preset report model feature library; If the vector matching degree is not lower than the preset matching degree threshold, then it is determined that there is a matching existing report. If the vector matching degree is lower than the preset matching degree threshold, it is determined that there is no matching existing report, and the preset semantic transformation method is executed to obtain the query result.

4. The intelligent data analysis method based on a large model according to claim 1, characterized in that, Also includes: Collect interactive feedback data from visual charts; The feedback feature vector is obtained based on interactive feedback data and semantic constraints. Optimize parameters in response to feedback feature vectors to generate query results; The parameters are iteratively adjusted based on the query results to optimize the query results; The optimized query results are displayed in a visual chart, which is then pushed to the chart viewing terminal. The optimization parameters are stored in the preset model optimization library.

5. The intelligent data analysis method based on a large model according to claim 1, characterized in that, It also includes methods for processing auxiliary sound in specific frequency bands: Acquire mixed audio signals from the voice receiving device; Frequency domain analysis of the mixed audio signal is performed to obtain the audio segments of human voice, ambient noise, and auxiliary sound; The frequency band separation threshold is determined based on the human voice frequency band and the auxiliary sound frequency band; The mixed audio signal is subjected to frequency band separation processing based on the frequency band separation threshold to obtain a clean auxiliary audio band signal; Extract specific frequency components from the pure auxiliary audio frequency band signal and calculate the auxiliary amplitude characteristics of each frequency component; Effective frequency components are selected by comparing the auxiliary amplitude features with the preset auxiliary amplitude threshold. The effective frequency components are matched with a preset frequency mapping table to obtain the mapping semantic constraints. When the mapping semantic constraints are consistent with the semantic constraints, the semantic parsing step is executed.

6. The intelligent data analysis method based on a large model according to claim 5, characterized in that, Also includes: Extracting voiceprint features of individuals based on human voice audio; The speaker's voiceprint features are matched with a pre-stored user voiceprint database to determine the speaker's identity. Query access permission levels based on the speaker's identity; Combine access permission levels and query results to perform permission filtering to obtain filtered query results; The system generates visual charts based on the filtered query results and outputs them to the chart viewing terminal.

7. The intelligent data analysis method based on a large model according to claim 6, characterized in that, Also includes: Extracting emotional features from human voice audio; Based on speech emotion features to obtain the emotional state of the query; When the emotional state is a preset emergency state, a priority processing prompt is reported, and specific priority information is determined based on the person's voiceprint characteristics; In response to priority prompts, specific priority information is processed.

8. An intelligent data analysis system based on a large model, characterized in that, include: The data acquisition module is used to collect natural language query requests. A memory for storing a program that implements a large-model-based intelligent data analysis method as described in any one of claims 1 to 7; The processor is used to load and execute programs stored in memory.

Citation Information

Patent Citations

  • Speech feature processing method and device, equipment and medium

    CN120340477A

  • Interactive AI report generation method and system based on intelligent semantic driving

    CN120541091A