Vehicle network interaction data query method based on large language model
Patent Information
- Application Number
- CN202510995235.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-21
AI Technical Summary
在车网互动场景中,充电行为分析和电网负荷预测等业务需频繁查询海量充换电数据,但运营人员缺乏SQL技能,现有技术方案效率低且门槛高。
采用基于大语言模型的数据查询方法,通过获取数据库连接信息、处理用户查询请求、生成和校验SQL查询语句,并利用DeepSeek-V3模型进行结果可视化呈现,实现端到端的自动化流程。
显著提升查询效率,降低用户门槛,提高实时决策能力,并通过可视化结果降低解读复杂数据的难度。
Smart Images

Figure CN120994683A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an intelligent query method for natural language to SQL based on a large language model, and belongs to the technical field of data intelligent query of vehicle network interaction. BACKGROUND
[0002] The core task of natural language to structured query language (Text2SQL) is to automatically convert the data retrieval requirements expressed in daily language by a user into an accurate database query statement (SQL). For example, when the user inputs "display all active customer list", the system can convert it into a standard SQL statement such as "SELECT * FROM customers WHERE status='active'" SQL command. The typical application process is: the user inputs a query text and specifies the table structure information of the target database, the system calls the Text2SQL service component to parse the input, generates the corresponding SQL statement, then executes the statement to access the database, and finally presents the query result to the user. At present, the main technical paths to solve the Text2SQL challenge are divided into two categories: the rule-based method relies on the pre-defined pattern matching and conversion rules of experts, assembles the SQL statement by parsing the keywords and structure in the natural language, and has the advantages of high accuracy in specific scenarios, but the disadvantages are that a large amount of manpower is needed to build the rule library and it is difficult to adapt to complex and variable or unforeseen query sentence patterns; the machine learning-based method uses algorithm models (especially deep neural networks) to automatically generate queries by learning the mapping relationship between natural language queries and standard SQL in large-scale labeled data, and has the advantages of stronger generalization ability to cope with diversified user requests, but the disadvantage is the high dependence on high-quality and large-scale training sample sets. SUMMARY
[0003] The technical problem to be solved by the application is that the existing technology has the following limitations in the vehicle network interaction scene: charging behavior analysis, power grid load prediction and other businesses need to frequently query massive charging and battery swapping data, but the operation personnel often lack SQL skills.
[0004] In order to solve the above technical problems, the technical scheme of the application discloses a vehicle network interaction data query method based on a large language model, characterized in that it comprises the following steps:
[0005] Step 1, obtain database connection information, read related vehicle network interaction table structure schema information, sort out business-related SQL table creation statements, and assemble them into prompts;
[0006] Step 2, process the user's query request for the database, comprising the following steps:
[0007] Step 201, a user submits a natural language query;
[0008] Step 202, standardizing the original natural language query;
[0009] Step 203, encoding the preprocessed natural language query into a high-dimensional vector;
[0010] Step 3, generating a prompt for the large language model, the prompt is encapsulated, the Chinese name of the prompt is "prompt word", and the Few-shot method is used to help the large language model understand the task by providing examples, the examples include input and expected output;
[0011] Step 4, deploying the DeepSeek-V3 model using the Dify platform, establishing a RESTful API service interface to realize the acquisition of the execution result of the large language model;
[0012] Step 5, input the encoded text into the pre-trained DeepSeek-V3 model, the DeepSeek-V3 model outputs one or more possible SQL query statements, performs syntax checking on the generated SQL statements, and returns the processed SQL query statements to the user;
[0013] Step 6, according to the returned SQL query statement, execute the corresponding database operation, get the charging operation index result;
[0014] Step 7, use DeepSeek-V3 model to parse data and generate ECharts chart configuration, render the visualization result and present it to the user.
[0015] Preferably, in step 1, the database is a relational database, which contains multiple tables, each table defines core business fields through UNIQUE KEY and duplicate key, and logical connection between tables is realized through semantic association fields.
[0016] Preferably, in step 1, the database connection information includes database URL, username and password.
[0017] Preferably, the standardization processing in step 202 includes: text normalization; intent feature extraction, the feature extraction process is based on massive SQL question and answer corpus to fine-tune the model, and natural language processing technology is used to realize accurate identification of query intent.
[0018] Preferably, in step 203, the preprocessed query is encoded into a high-dimensional vector using the text-embedding-v3 model.
[0019] Compared with the prior art, the present application has the following advantages:
[0020] 1、Query efficiency is significantly improved
[0021] The traditional scheme relies on manual writing of SQL statements, which requires a lot of time for syntax debugging and logic verification, and has high skill requirements for the operator. The present application introduces a natural language interaction mechanism, and the user can directly describe the demand in daily language, and the system automatically analyzes and generates the query statement through the large language model and vector retrieval technology, realizes the transformation from "manual coding" to "intelligent dialogue", the response speed presents order of magnitude improvement, and the real-time decision-making ability is significantly enhanced.
[0022] 2、Analysis result visualization and readability optimization
[0023] The original data table output by the traditional scheme has the problems of high information density and high interpretation threshold. The present application relies on the result generation capability of the large language model to automatically convert multi-dimensional data into rich text reports with symbolic annotations, emoticons and logical segmentation, and combines SQL generation and chart rendering of DeepSeek-V3 to make the complex data mode intuitive, effectively reducing the user's cognitive load. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a flowchart of the existing Text2SQL technology;
[0025] Figure 2 is a data query method flowchart for realizing natural language to SQL based on a large language model provided by the present application;
[0026] Figure 3 is an effect diagram of the method of the present application. DETAILED DESCRIPTION
[0027] The present application will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present application and not to limit the scope of the present application. In addition, it should be understood that after reading the content taught by the present application, those skilled in the art can make various modifications or modifications to the present application, and these equivalent forms also fall within the scope defined by the appended claims of the present application.
[0028] To solve the data query obstacle of non-technical personnel in the field of vehicle-network interaction, the vehicle-network interaction data query method based on a large language model provided by the present application can realize:
[0029] 1、The vehicle-network interaction data query method based on a large language model reduces the query threshold
[0030] 2、Deeply adapt to the field term understanding ability of vehicle-network interaction
[0031] 3、End-to-end "query-execution-visualization" automation process
[0032] 4. Avoid model training and complex rule definition, improve implementation efficiency.
[0033] To achieve the above purpose, the embodiment of the application discloses a vehicle network interaction data query method based on a large language model, as shown in the figure, which specifically includes the following steps. Figure 2
[0034] Step 1: Obtain database connection information
[0035] Firstly, the application needs a relational database, such as a Doris database and the database connection information, wherein the database connection information includes database URL, username and password, so as to ensure that the system can successfully connect to the target database.
[0036] The relational database contains multiple tables, each table defines core business fields (such as user ID, timestamp, etc.) through UNIQUE KEY and duplicate key, and logically connects tables through semantic association fields (for example: the user_id field in the user_behavior table forms an association semantic with the user_id in the user_profile table).
[0037] In step 1, the relevant vehicle network interaction table structure schema information is read. This step is very important, because if the information arrangement effect is not good, it will directly affect the correctness of the execution result of the large language model. After many test experiments, the application selects the optimal scheme to encapsulate the table structure information: the business-related SQL table creation statement is sorted out, assembled into the prompt, and the large model clearly knows which business tables the following questions are based on to generate SQL.
[0038] Step 2: User query processing
[0039] This stage mainly processes the user's query request for the database, and the specific process is as follows:
[0040] (1) Query input: As the initial link of the process, the user can submit a natural language query through various interactive ways (such as text input, voice instruction, etc.).
[0041] (2) Query preprocessing: Standardize the original query, including: removing stop words, special symbols and numbers.
[0042] Normalize the text (such as converting to lowercase, extracting stems). Intention feature extraction (identify the operation type in the query, such as filtering conditions, aggregation functions, sorting requirements, etc.), which is based on a large amount of SQL question and answer corpus to fine-tune the model, and uses natural language processing technology to accurately identify the query intention.
[0043] (3) Vectorization representation: use the text-embedding-v3 model to encode the preprocessed query into a high-dimensional vector. This advanced embedding technology can effectively capture semantic associations, providing a machine-understandable numerical representation for subsequent SQL conversion.
[0044] Step 3: Generate prompt for large language model.
[0045] Prompt is a short piece of text that guides the large model to generate a specific type of output. They can be in the form of questions, instructions, examples, etc., providing context and direction for the model. Through the clever design of Prompt, it can guide the model to generate more accurate, targeted answers and creative output. High-quality Prompt is crucial for generating high-quality content. A good Prompt can clearly guide the model to generate accurate and targeted output, while a low-quality Prompt may lead to confusion, irrelevance or low-quality results.
[0046] The invention encapsulates the Prompt and names it "prompt" in Chinese. The invention uses the Few-shot method to help the model understand the task by providing a small number of labeled samples (examples). These examples usually include input and expected output, so that the model can better understand the nature and requirements of the task. For the target task, since the model first sees good examples, it can better understand human intentions and standards for what type of answer is needed. Whether in supervised learning or unsupervised learning training process, Prompt can help the model better understand the intention of the input and respond accordingly. In addition, Prompt can also improve the explainability and accessibility of the model. Simply put, Prompt is like providing an "hint" or "guide" to the AI model, helping it better understand and complete the task. The encapsulation of Prompt has undergone multiple rounds of experiments to achieve the best effect. This process ensures the rigor, formality and professionalism of the text, while meeting language standards.
[0047] Prompt encapsulation examples are as follows:
[0048] You are a car network interactive data query generator, your main goal is to assist users as much as possible to convert the input text into a SQL statement, the SQL spelling process is strictly prohibited to use non-existent fields, please return the SQL statement directly and ensure the correctness of the SQL syntax, pay attention to not appear explanatory statements, also do not appear word description. The table name and table field generated are from the following table: {}
[0049] Examples are as follows:
[0050] Q: Statistics on charging volume in each administrative district yesterday.
[0051] SELECT district AS administrative district,SUM(elect)AS power FROM charge_detail WHERE d=DATE_SUB(curdate(),INTERVAL 1DAY)GROUP BY district
[0052] Q: Statistics on the cumulative number of charging facilities connected in April 2025
[0053] SELECT count(case when station_type=1then 1else null end) public, count(case when station_type<>1then 1else null end) private FROM access_equipmentsWHERE d='202504'and station_type<>50
[0054] Step 4: Deploy the DeepSeek-V3 model using the Dify platform and establish a RESTful API service interface to obtain the execution results of the large language model.
[0055] This invention selects the DeepSeek-V3 model as the deep learning model. The DeepSeek-V3 model boasts high accuracy in natural language processing, capable of accurately classifying, labeling, and parsing text. By utilizing the DeepSeek-V3 model, higher accuracy, greater efficiency, and more intelligent services can be achieved, thus better meeting the needs of enterprises and users.
[0056] Step 5: Input the encoded text into the pre-trained large language model, and the model will output one or more possible SQL query statements.
[0057] Perform syntax validation on the generated SQL statement to avoid grammatically incorrect SQL queries and ensure that the generated SQL query statement is valid. Return the processed SQL query statement to the user.
[0058] Step 6: Based on the returned SQL query, execute the corresponding database operations to obtain the charging operation indicator results;
[0059] Step 7: Finally, use the DeepSeek-V3 model to parse the data and generate ECharts chart configurations, render the visualization results, and present them to the user, as shown below. Figure 3 As shown.
Claims
1. A method for querying vehicle-to-everything (V2X) interactive data based on a large language model, characterized in that, Includes the following steps: Step 1: Obtain database connection information, read the relevant vehicle network interaction table structure schema information, organize the business-related SQL table creation statements, and assemble them into the prompt. Step 2: Process user query requests to the database, including the following steps: Step 201: The user submits a natural language query; Step 202: Standardize the original natural language query; Step 203: Encode the preprocessed natural language query into a high-dimensional vector; Step 3: Generate a prompt for the large language model. The prompt is encapsulated and named "prompt words" in Chinese. It uses a Few-shot approach to provide examples to help the large language model understand the task. The examples include input and expected output. Step 4: Deploy the DeepSeek-V3 model using the Dify platform and establish a RESTful API service interface to obtain the execution results of the large language model; Step 5: Input the encoded text into the pre-trained DeepSeek-V3 model. The DeepSeek-V3 model outputs one or more possible SQL query statements, performs syntax validation on the generated SQL statements, and returns the processed SQL query statements to the user. Step 6: Based on the returned SQL query statement, execute the corresponding database operations to obtain the charging operation indicator results; Step 7: Use the DeepSeek-V3 model to parse the data and generate ECharts chart configuration, render the visualization results and present them to the user.
2. The method for querying vehicle-to-everything (V2X) interactive data based on a large language model as described in claim 1, characterized in that, In step 1, the database is a relational database containing multiple tables. Each table defines core business fields using UNIQUE KEY and duplicate key, and the tables are logically connected through semantic association fields.
3. The method for querying vehicle-to-everything (V2X) interactive data based on a large language model as described in claim 1, characterized in that, In step 1, the database connection information includes the database URL, username, and password.
4. The method for querying vehicle-to-everything (V2X) interactive data based on a large language model as described in claim 1, characterized in that, The standardization process in step 202 includes: normalizing the text; extracting intent features. The feature extraction process is based on fine-tuning the model using a massive SQL question-and-answer corpus, and natural language processing technology is used to achieve accurate identification of query intent.
5. The vehicle-to-everything (V2X) interactive data query method based on a large language model as described in claim 1, characterized in that, In step 203, the preprocessed query is encoded into a high-dimensional vector using the text-embedding-v3 model.
Citation Information
Patent Citations
Model development and deployment method based on cloud native micro-service
CN113961174A
Data query method for realizing conversion from natural language to SQL (Structured Query Language) based on large language model
CN117708161A
Multi-LLM model integrated reasoning method for Text2SQL task
CN120218227A