Intelligent assistant efficiency optimization method and system and storage medium
By combining semantic matching algorithms, RAG algorithms, predefined scenario matching and LLM models, existing smart assistants solve the problems of intelligent bottlenecks, performance efficiency dilemmas, scenario coverage vulnerabilities and poor data processing performance in the electronics manufacturing industry, and efficient and flexible user request response and data-driven decision-making are achieved.
Patent Information
- Application Number
- CN202411812598.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
Existing smart assistants face intelligent bottlenecks, performance efficiency dilemmas, scenario coverage vulnerabilities and poor data processing performance in the electronics manufacturing industry, and it is difficult to meet the complex and changeable business needs and the needs of efficient data-driven decision-making.
Through a multi-step intelligent processing process, including a combination of semantic matching algorithm, RAG algorithm, predefined scene matching and LLM model, accurately adapted intent SQL statements are generated to optimize the response accuracy and flexibility of user requests.
It significantly improves the response accuracy and flexibility of smart assistants, enhances user trust and usage satisfaction, expands application scope and business coverage capabilities, meets complex and changeable business needs, and improves the effectiveness of data-driven decision-making.
Smart Images

Figure CN120011606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic manufacturing, and in particular to a method, system and storage medium for optimizing the performance of an intelligent assistant. Background Art
[0002] The wave of digital transformation and intelligent manufacturing is sweeping in. As a key pillar of the national economy, the electronic manufacturing industry is accelerating its intelligent transformation. Intelligent assistants have been applied to a certain extent in many core business areas of the industry, such as supply chain, production, order and material management. However, an analysis of the existing intelligent assistant technology paradigm shows that it mainly relies on intent recognition technology to identify scenarios, configurable templates to generate standard responses, and slot filling mode to handle multiple rounds of interactions. Although it has made some achievements in improving management efficiency, it is gradually facing difficulties in the face of increasing business complexity and increasing dynamic diversity of scenarios.
[0003] First, the bottleneck of intelligence is prominent. The predefined templates and slot-filling question-and-answer mode are rigid, and the response to complex and changing business needs is slow. They lack flexible adaptability and are difficult to accurately capture the deep intentions of users, resulting in limited interactive intelligence. Second, the performance efficiency dilemma is aggravated. The expansion of scenarios and the upgrade of complexity have caused algorithm overload, a sharp drop in computing performance, redundant template accumulation, lack of system flexibility, greatly reduced user experience, and frequent problems such as delayed operation and delayed feedback. Third, the scene coverage loopholes are significant. There is a deviation between the preset scene template system and the actual business panorama of the electronics manufacturing industry. Emerging businesses and special needs are often outside the scope of system processing, resulting in no effective response to some user queries and obstruction of business processes. Fourth, data processing efficiency is poor. In complex data query and analysis scenarios, the traditional architecture has a low degree of automation and intelligence, over-reliance on manual intervention, insufficient timeliness and accuracy of decision support, and it is difficult to meet the data-driven decision-making needs of enterprises, hindering the scientific improvement of production and operation decisions and efficient collaborative development. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a method, system and storage medium for optimizing the performance of an intelligent assistant.
[0005] According to the first aspect of the present application, a method for optimizing the performance of an intelligent assistant is proposed, the method comprising:
[0006] S1, obtain user request data;
[0007] S2, based on the semantic matching algorithm, calculating the semantic similarity of the user request data, and in response to the semantic similarity being greater than a first threshold, using the RAG algorithm to search the knowledge base and generate reply information;
[0008] S3, in response to the semantic similarity being less than or equal to the first threshold, enter the predefined scene matching calculation, if the judged scene matching confidence is greater than the second threshold, obtain the key elements of the user request data, splice and generate an SQL statement that accurately adapts to the query intent according to the predefined logical template and SQL grammar rules, and generate a reply message;
[0009] S4, in response to the scenario matching confidence being less than or equal to the second threshold, converting the user request data into the intended SQL statement 2 based on the LLM model 3, and generating reply information;
[0010] S5, output reply information.
[0011] In the above technical solution, the accuracy and flexibility of the intelligent assistant's response to user requests are significantly improved through a multi-step intelligent processing flow. Whether it is through precise semantic matching and RAG algorithm retrieval of the knowledge base, or predefined scenario matching and LLM model conversion to generate SQL statements, different types of user requests can be effectively processed, thereby increasing the probability of users obtaining accurate information and enhancing users' trust in and satisfaction with the intelligent assistant. This diversified processing method also enables the intelligent assistant to adapt to complex and changing business environments, reducing the number of unanswered or incorrect answers due to a single request type or limited processing logic, greatly expanding the application scope and business coverage of the intelligent assistant, and laying a solid foundation for its in-depth application in various industries.
[0012] Furthermore, the method further includes a refusal identification step disposed between step S1 and step S2, and in response to identifying that the user request data is sensitive information, the operation is directly terminated. Adding the refusal identification step effectively ensures the security and compliance of the system.
[0013] Further, step S2 includes the following sub-steps:
[0014] S21, collecting data pairs in business scenarios, where the data pairs include business questions and corresponding structured query statements;
[0015] S22, using embedding tools, accurately encodes business problems into storage semantic vectors, and stores the storage semantic vectors in a vector database in an orderly manner, and associates the corresponding business data with the knowledge base architecture;
[0016] S23, in response to obtaining the user request data, reusing the embedding tool to convert the user request data into a request semantic vector;
[0017] S24, using a semantic matching algorithm to calculate the similarity between the requested semantic vector and the stored semantic vector, and in response to the similarity being greater than a first threshold, filtering out business data pairs corresponding to the most similar top 3 knowledge vectors;
[0018] S25, based on the business logic and user request context, splices into prompt1, and inputs prompt1 into LLM model 1 to obtain reply information.
[0019] In the above technical solution, by collecting data pairs in business scenarios and encoding business problems into semantic vectors and storing them in the vector database, deep integration and rapid retrieval of business knowledge are achieved. When a user makes a request, it can be quickly converted into a request semantic vector and the similarity with the stored semantic vector is calculated to screen out the most relevant knowledge pairs. Compared with traditional keyword matching or simple text matching methods, this matching method based on semantic vectors has higher accuracy and recall rate, and can more accurately understand user intentions and provide relevant response information. The prompt1 input LLM model based on business logic and user request context splicing further optimizes the response generation process, so that the generated responses are not only based on accurate knowledge retrieval, but also combined with the model's powerful language generation capabilities to generate more natural and fluent responses that conform to user reading habits and have depth and logic, effectively improving the interactive experience and communication effect between users and intelligent assistants.
[0020] Furthermore, step S3 includes the following sub-steps:
[0021] S31, collecting scene data pairs, using the scene classification model to mine the mapping logic between text features of the scene data pairs and business scene labels, and learning text feature combination patterns to identify business scene categories;
[0022] S32, formulating a dedicated SQL query statement for a predefined business scenario according to the rules, data structure, and query requirements, and storing it in a business scenario database;
[0023] S33, in response to obtaining the user request data, using the scenario classification model to calculate the scenario matching confidence of the business scenario category to which the user request data belongs;
[0024] S34, in response to the scene matching confidence being greater than the second threshold, extracting key elements of the user request data, and improving the corresponding exclusive SQL query statement, and splicing and generating an intent SQL statement 1 adapted to the query;
[0025] S35, retrieve the corresponding intention SQL statement 1 from the database, execute data query, and generate reply information.
[0026] Furthermore, step S34 also includes extracting the key elements of the user's requested data, and then determining whether multiple rounds of extraction are required. If the determination is "yes", multiple rounds of inquiries are performed using LLM model 2 until a complete intended SQL statement 1 is obtained.
[0027] Furthermore, the exclusive SQL query statement includes a "where" clause part, and the key elements are filled in the parameter positions corresponding to the "where" clause part.
[0028] In the above technical solution, by collecting scene data pairs and using the scene classification model to learn the mapping logic between text features and business scenario labels, the intelligent assistant can quickly and accurately determine the business scenario category to which the user request belongs. Exclusive SQL query statements formulated for different scenarios are stored in the database. After a successful match, they can be directly spliced to generate intent SQL statements for data query, avoiding repeated analysis and query construction processes when dealing with common business problems, and greatly shortening the response time. At the same time, this processing method based on scene classification enables the intelligent assistant to provide more professional and customized response information for different business scenarios, meet the specific information needs of different departments or different business processes within the enterprise, and enhance the actual application value and decision-making support capabilities of the intelligent assistant in the enterprise business process, and promote the efficiency and intelligence level of the enterprise's business operations.
[0029] Furthermore, after generating the first and second intention SQL statements in step S3 and / or step S4, the process also includes executing data query and determining whether result analysis is required. If the determination is “yes”, data analysis is performed based on the Multi-agent architecture to generate response information. This in-depth analysis can transform the original query results into information with insight and decision support value.
[0030] Furthermore, the value range of the first threshold and the second threshold is [90%, 98%]. Within this value range, the higher threshold ensures the accuracy of both semantic matching and scene matching, avoids the introduction of a large amount of irrelevant information due to overly loose matching, and ensures the quality and relevance of the reply information.
[0031] In a second aspect, the present application proposes a system for optimizing the performance of an intelligent assistant, the system comprising:
[0032] A data acquisition module, configured to acquire user request data;
[0033] A RAG module is configured to calculate the semantic similarity of the user request data based on a semantic matching algorithm, and in response to the semantic similarity being greater than a first threshold, retrieve the knowledge base using the RAG algorithm to generate reply information;
[0034] The scene matching module is configured to enter the predefined scene matching calculation in response to the semantic similarity being less than or equal to the first threshold, and if the judged scene matching confidence is greater than the second threshold, obtain the key elements of the user request data, and splice and generate an SQL statement that accurately adapts to the query intent according to the predefined logic template and SQL grammar rules, and generate a reply message;
[0035] An LLM model three generation module is configured to convert the user request data into an intent SQL statement two based on the LLM model three in response to the scene matching confidence being less than or equal to a second threshold, and generate reply information;
[0036] The reply output module is configured to output reply information.
[0037] In a third aspect, the present application proposes a terminal device comprising a processor, a memory, and a computer program stored in the memory, wherein the computer program is executed by the processor to implement a method for optimizing the performance of an intelligent assistant as described above.
[0038] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, a method for optimizing the performance of an intelligent assistant as described above is implemented.
[0039] Compared with the prior art, the beneficial results of the present invention are:
[0040] 1. The present invention uses the excellent natural language understanding and generation capabilities of the large language model, abandons the limitations of traditional slot-filling inquiries, and achieves a more intelligent, natural and smooth multi-round dialogue mode. It can deeply understand and accurately respond to the complex and diverse needs of users, greatly improving the intelligence level of the smart assistant and making user interaction more humane, efficient and convenient.
[0041] 2. On the one hand, the present invention, with the optimized query matching mechanism and the carefully crafted system architecture, can quickly and accurately match common queries with predefined scenarios, greatly enhancing system reliability and response speed, and significantly optimizing user experience; on the other hand, through the detailed definition and in-depth segmentation of various business scenarios in the electronics manufacturing industry, the application scope of the system is effectively broadened, which can meet the unique and personalized needs of different enterprises, and greatly improve the adaptability of the system and the breadth of business coverage.
[0042] 3. With the help of Agent automated execution and multi-agent data analysis mechanism, the present invention enables efficient and intelligent processing of complex data query and analysis tasks, greatly reduces manual intervention, significantly improves work efficiency and strengthens data-driven decision-making effectiveness; at the same time, it integrates an advanced content filtering engine that can automatically and accurately identify and intercept inappropriate content, effectively ensure system security and compliance, effectively avoid legal and ethical risks, and maintain the system's professional image and user trust foundation.
[0043] 4. The present invention adopts a modular architecture design, which enables the system to be flexibly adjusted and upgraded according to the dynamic changes in business needs, which not only improves the sustainability of the system, but also brings a more substantial return on investment to the enterprise; and the synergistic effect of intelligent interaction, multi-round dialogue and rapid response mechanism greatly optimizes the user experience, enhances user satisfaction and usage stickiness, and effectively promotes the widespread application of the system and improves market recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present invention. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.
[0045] Figure 1 is a flow chart of a method for optimizing the performance of an intelligent assistant according to an embodiment of the present invention;
[0046] Figure 2 is a flow chart of an intelligent assistant performance optimization algorithm according to an embodiment of the present invention;
[0047] Figure 3 is a flowchart of a solution selection method for optimizing the performance of an intelligent assistant according to an embodiment of the present invention;
[0048] Figure 4 is a flowchart of a multi-agent linear structure processing data analysis task of an intelligent assistant performance optimization method according to an embodiment of the present invention;
[0049] Figure 5 is a flowchart of a large language model training method for optimizing the performance of an intelligent assistant according to an embodiment of the present invention;
[0050] Figure 6 is a Lora structure diagram of a smart assistant performance optimization method according to an embodiment of the present invention;
[0051] Figure 7 is a structural difference diagram of Lora and Qlora according to a method for optimizing the performance of an intelligent assistant according to an embodiment of the present application;
[0052] Figure 8 is a framework diagram of an intelligent assistant performance optimization system according to an embodiment of the present application;
[0053] Fig. 9 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION
[0054] In the following detailed description, reference is made to the accompanying drawings, which form a part of the detailed description and are shown by illustrative specific embodiments in which the present invention can be put into practice. To this end, directional terms, such as "top", "bottom", "left", "right", "up", "down", etc., are used with reference to the orientation of the figures described. Because the parts of the embodiments can be positioned in several different orientations, directional terms are used for the purpose of illustration and are by no means limiting. It should be understood that other embodiments may be utilized or logical changes may be made without departing from the scope of the present invention. Therefore, the following detailed description should not be adopted in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0055] The present invention provides a display screen assembly method. Figure 1 A flow chart of a display screen assembly method according to an embodiment of the present invention is shown. As shown in the figure, the method includes:
[0056] S101, obtaining user request data.
[0057] S102, based on a semantic matching algorithm, calculating the semantic similarity of the user request data, and in response to the semantic similarity being greater than a first threshold, using a RAG algorithm to search a knowledge base and generate reply information.
[0058] In a specific embodiment, step S102 includes the following sub-steps:
[0059] S1021, collecting data pairs in business scenarios, where the data pairs include business questions and corresponding structured query statements;
[0060] S1022, using embedding tools, accurately encode business problems into storage semantic vectors, and store the storage semantic vectors in a vector database in an orderly manner, and associate the corresponding business data with the knowledge base architecture;
[0061] S1023, in response to obtaining the user request data, reusing the embedding tool to convert the user request data into a request semantic vector;
[0062] S1024, using a semantic matching algorithm to calculate the similarity between the requested semantic vector and the stored semantic vector, and in response to the similarity being greater than a first threshold, filtering out business data pairs corresponding to the most similar top 3 knowledge vectors;
[0063] S1025, based on the business logic and the user request context, splice into prompt1, and input prompt1 into LLM model 1 to obtain the reply information
[0064] Specifically, when building a knowledge base, first collect a large number of business data pairs, such as <Help me find out which work orders are not scheduled? , select * from where status = 'Not scheduled'>, then use a large language model embedding tool such as text-embedding-ada-002 to encode the questions in the data pair into semantic vectors and store them in a vector database such as milvus. After all business data are encoded and stored, the knowledge base is built. When a user request is passed in, the same embedding model is first used to semantically encode the user request to obtain a vector representation, and then the vector is used in milvus for similarity matching calculation. When the semantic similarity is greater than the first threshold, the most similar top3 knowledge is extracted, and prompt1 similar to "Based on the given three pieces of knowledge, help me answer the user's question" is spliced. Finally, these data are input into a large language model such as chatglm, so that chatglm can generate and provide the user with the corresponding reply.
[0065] S103, in response to the semantic similarity being less than or equal to the first threshold, enter the predefined scene matching calculation. If the judged scene matching confidence is greater than the second threshold, obtain the key elements of the user request data, and splice and generate an SQL statement that accurately adapts to the query intent according to the predefined logical template and SQL grammar rules, and generate a reply message.
[0066] In a specific embodiment, step S103 includes the following sub-steps:
[0067] S1031, collecting scene data pairs, using the scene classification model to mine the mapping logic between text features of the scene data pairs and business scene labels, and learning text feature combination patterns to identify business scene categories;
[0068] S1032, formulating a dedicated SQL query statement for a predefined business scenario according to the rules, data structure, and query requirements, and storing it in a business scenario database;
[0069] S1033, in response to obtaining the user request data, using the scenario classification model to calculate the scenario matching confidence of the business scenario category to which the user request data belongs;
[0070] S1034, in response to the scene matching confidence being greater than the second threshold, extract the key elements of the user request data, improve the corresponding exclusive SQL query statement, and splice to generate the intended SQL statement one that adapts to the query. After extracting the key elements of the user request data, determine whether multiple rounds of extraction are required. If the judgment is "yes", use LLM model 2 to perform multiple rounds of inquiries until the intended SQL statement one is complete.
[0071] S1035, retrieve the corresponding SQL statement from the database to execute data query and generate reply information.
[0072] Specifically, in the initial stage, we comprehensively collect various scenario data covering multiple business fields, such as <Which orders are waiting to be placed?, Shipment management customer order status> and <Are there any obsolete materials in the current inventory for more than 720 days?, Inventory management obsolete material report>, etc. Then, we use these rich data to train the scenario classification model (such as bert-classifier) to enable it to accurately identify different scenarios; then, we carefully configure the SQL statements that are suitable for each specific scenario and save them properly in the database. When the user asks a question, we immediately use the trained model to classify the user's question text, match the corresponding SQL statement based on the classification result to execute the data acquisition task, and finally return the acquired data results to the user accurately, so as to efficiently meet the user's needs for specific scenario data and realize accurate information interaction services.
[0073] S104, in response to the scene matching confidence being less than or equal to the second threshold, converting the user request data into intent SQL statement two based on LLM model three, and generating reply information.
[0074] In a specific embodiment, prompt2 is carefully designed first, such as generating executable SQL statements based on a given data schema and user input and avoiding irrelevant information output. When the user enters a question, the data table information and prompt2 are input into the large language model chatglm, which directly returns the generated SQL statement. There are many optimization techniques when designing prompt2. First, the task instructions should be clear, that is, the instructions should be clear and accurate during design so that the model can understand the user's intention, such as "Please generate an article about artificial intelligence" instead of simply "artificial intelligence"; secondly, sufficient contextual information should be incorporated into prompt2, covering topics, target audiences, format requirements, etc., to help the model generate more relevant outputs; finally, diversity testing is required, that is, using different prompt2s to test model responses to explore the optimal design, because the model has different generation quality for different prompt2s, so continuous optimization is performed to improve the effect and accuracy of SQL statement generation.
[0075] S105, output reply information.
[0076] In a specific embodiment, the method further includes a refusal identification step provided between step S101 and step S102, and in response to identifying that the user request data is sensitive information, the operation is terminated directly.
[0077] Specifically, a user query content review mechanism is established to accurately identify the query content submitted by the user to determine whether it is a chatty expression or involves sensitive information. Once it is determined that the user's query content belongs to such chat or sensitive information categories, the system will immediately start the rejection procedure and will not proceed with the subsequent processing process, so as to ensure the compliance, security and efficiency of the interactive environment, and avoid the waste of system resources and negative impact on the overall interactive experience due to the intervention of irrelevant or bad information.
[0078] In a specific embodiment, after generating the intended SQL statement one and / or the intended SQL statement two in step S103 and / or step S104, it also includes executing a data query and determining whether result analysis is required. If the determination is "yes", a data analysis is performed based on the Multi-agent architecture to generate a reply message.
[0079] Continue to refer Figure 2 , Figure 2 A flowchart of an intelligent assistant performance optimization algorithm according to an embodiment of the present invention is shown. As shown in the figure, the algorithm includes:
[0080] Step 201: The user inputs a question. The user inputs the query content in the form of text or voice to start the entire processing flow.
[0081] Step 202, refusal to answer identification. With the help of technical means based on keyword and semantic analysis, the content filtering engine can comprehensively monitor the interactive content in real time, timely and accurately intercept queries such as sensitive content such as chat, and the refusal to answer mechanism will give appropriate prompts or directly reject such sensitive queries. In this way, not only the overall security and compliance of the system are significantly improved, and potential legal and ethical risks are effectively avoided, but also the professional image and user trust of the system are effectively maintained, thereby comprehensively improving the user experience and laying a solid foundation for the stable, reliable and continuous operation of the system.
[0082] Step 203, determine the solution selection. First, perform semantic matching calculation. In response to the semantic similarity being greater than the first threshold, match the knowledge base and execute step 212; if the semantic similarity is less than or equal to the first threshold, enter the predefined scene matching calculation. If the calculated scene matching confidence is greater than the second threshold, execute step 209; if the calculated scene matching confidence is less than the second threshold, it is equivalent to a user input problem, which does not match the scene and the knowledge base, and execute step 204. Specifically, the system performs multiple rounds of matching judgments in sequence, first trying to match the RAG module, and using the semantic matching algorithm to accurately calculate the similarity between the user input and the knowledge base content. If the similarity exceeds 95%, the match is successful and the knowledge base is retrieved and a reply is generated according to the RAG module; if it is less than or equal to 95%, enter the predefined scene matching process. When the scene matching probability is higher than 96%, the predefined scene module processing is enabled; if this probability is not reached, the large language model (LLM) is started to generate the SQL solution.
[0083] Step 204, LLM (generate SQL using LLM). In the no-match scenario and knowledge base scenario, LLM converts the query into precise executable SQL statements based on the input natural language query semantic logic and its powerful language understanding and code generation capabilities.
[0084] Step 205, return the query result. After executing the database query with the SQL statements generated in step 204, step 209 and step 211, the system extracts and organizes the query results, and prepares for subsequent processing or direct feedback to the user. This step focuses on the accuracy of the result screening and format specification to ensure that the data quality meets the system interaction requirements and user expectations.
[0085] Step 206, determine whether result analysis is required. If "yes", execute step 207, if "no", execute step 208. The system intelligently determines whether the query results require in-depth analysis to support decision-making. If it is determined that analysis is required (such as involving business trend insights, problem diagnosis scenarios), the process goes to step 207; if no further analysis is required (such as simple data query), go directly to step 208 to return the final result. This judgment dynamically adjusts the processing path based on the depth of user needs and the complexity of the business scenario.
[0086] Step 207, multi-agent data analysis. In scenarios that require in-depth analysis, the multi-agent data analysis module deploys professional agent teams (such as sales trend analysis, production efficiency evaluation agents, etc.) on demand, and efficiently interacts and collaborates according to preset collaborative rules. Each agent uses professional knowledge and data analysis algorithms to deeply explore data features, rules and potential values, and after multiple rounds of iterative processing, generates analysis reports with decision-making reference value, drives data intelligence into decision-making wisdom, and improves the scientificity and foresight of enterprise operation decisions.
[0087] Step 208, the final result is returned. The system directly obtains the result in step 205 or the result after in-depth analysis in step 207 or accurately replies to the RAG module in step 212, and encapsulates and processes it in a standardized format and interactive protocol, accurately feeds back to the user, and successfully completes the query processing flow, providing users with accurate and valuable information feedback, improving the user interaction experience and the practical value of the system.
[0088] Step 209, extract key information and piece it together into SQL. Each predefined scenario configuration corresponds to the SQL template, including the previous main select part and the conditional where part, where the where part includes element slots filled with key parameters. When matching predefined scenarios, the system accurately extracts key elements (such as business entities, attributes, and operation instructions) from user input, and fills the extracted key elements into the corresponding element slots in the SQL template according to the predefined logical template and SQL grammar rules, and splices to generate an SQL statement that accurately adapts to the query intent, clarifying the direction for subsequent data queries or operations. This process requires a deep understanding of the relationship between business semantics and database architecture to ensure SQL validity and performance optimization.
[0089] Step 210, determine whether multiple rounds are required. If "yes", execute step 211, if "no", execute step 205. For predefined scenario processing, the system determines whether multiple rounds of interaction are required based on the complexity of user intent and SQL integrity. If the elements extracted from the key information cannot meet the SQL template defined in the predefined scenario, the SQL construction requires additional information (such as associating multiple tables for querying, complex condition screening), and it is determined that multiple rounds are required (go to step 211); if the SQL generated once can meet the query (such as a simple single-table query), go directly to step 205 to return the query result. This mechanism improves the flexibility and accuracy of the system in handling complex business interactions.
[0090] Step 211, LLM conducts multiple rounds of inquiries until the SQL is complete. When it is determined that multiple rounds of interaction are required, LLM generates targeted inquiries to guide users to complete information based on the previous interaction context and the logical loopholes of the unfinished SQL fragments, and continuously iterates and optimizes the SQL statement until it is completely executable. The process integrates natural language understanding and database semantic mapping technology, while ensuring a smooth user interaction experience, improving the success rate of complex query processing and optimizing the quality of system intelligent services.
[0091] In a specific embodiment, the specific application examples of steps 209 to 211 are as follows:
[0092] In response to receiving the user's query request "Help me find out which work orders are not completed", the system first performs a scenario classification operation on the user query, and determines that it belongs to the specific scenario category of "work order detailed information query". In this predefined scenario, the corresponding SQL template is "select wo_id, status, num where create_time = %s and status = %s". The "where" clause of the SQL template sets two parameters, corresponding to the two key elements of work order time and work order status. When extracting key information elements of the user query, only the "unfinished" work order status related element is obtained, while the work order time element is missing. The system needs to interact with the user through multiple rounds of inquiries to guide the user to further clarify the scope of the work orders he wants to query. After obtaining the specific work order query time information, the corresponding information is filled into the parameter position corresponding to the SQL template, thereby forming a complete executable SQL statement, and then executing the SQL query to obtain the required data. If the user's query is "Help me find out which work orders are not completed today", when extracting key information, the clear work order time element "today" and the work order status element "not completed" can be directly obtained. At this time, these key information can be directly filled into the corresponding parameter position of the SQL template, without going through multiple rounds of query processes, and a complete SQL statement can be quickly generated for subsequent data query operations, so as to achieve the purpose of accurately obtaining the corresponding work order data according to user needs.
[0093] Step 212, RAG. Based on the knowledge base matching scenario, the RAG module integrates the reasoning ability of the large language model and the knowledge retrieval advantage of the knowledge base. First, the user input is encoded into a semantic vector, and similar knowledge fragments are quickly retrieved in the vector database. Then, combined with the semantic understanding and generation capabilities of the large language model, accurate replies are synthesized based on the retrieved knowledge, providing users with knowledge-rich and logically coherent answers, enhancing the depth and breadth of the system's knowledge services, especially when processing professional knowledge queries.
[0094] Further, the judgment scheme selection method of step 203 refers to Figure 3 , Figure 3 The flowchart of the scheme selection according to the intelligent assistant performance optimization method of the present invention is shown as follows: Figure 3 As shown, the options include:
[0095] Step 301, calculation of semantic similarity of knowledge base. After receiving user input information, the system uses semantic analysis technology and vector space model and other algorithm tools to compare the user query statement with a large number of pre-stored knowledge texts in the knowledge base one by one, accurately quantify the degree of semantic association, and calculate the semantic similarity value of the knowledge base.
[0096] Step 302, determine whether the similarity is greater than 95%. When it is judged as "yes", directly execute step 307 and return the result. If it is judged as "no", continue to execute step 303. Make key judgments based on the similarity value obtained in step 301. If the similarity exceeds 95%, it indicates that the user query is highly matched with the specific knowledge in the knowledge base, and the system immediately jumps to step 307 to start the RAG module processing flow. The RAG module integrates the advantages of knowledge retrieval and language model reasoning, accurately locates similar knowledge in the knowledge base and generates high-quality responses, ensuring that the knowledge supply is accurate and efficient, and meets user needs. If the similarity does not reach 95%, the process proceeds to step 303, starts the predefined scene matching calculation phase, and continues to explore the optimal processing path.
[0097] Step 303, predefined scenario matching calculation. When the knowledge base matching is unsuccessful, the system switches to the predefined scenario matching mode. Using the deep neural network model and business rule engine, the user query text features, vocabulary patterns and potential business intentions are deeply analyzed, and the matching degree is calculated one by one with the system's predefined massive business scenarios (covering key scenarios in multiple fields such as supply chain, production, and orders). The confidence value of the scenario matching is generated to provide a quantitative basis for subsequent accurate decision-making. The accuracy of this link is related to the system's adaptability to complex business scenarios and processing efficiency.
[0098] Step 304, determine whether the scene matching confidence is greater than 96%. If it is "yes", directly execute step 306 and return the result. If it is judged as "no", continue to execute step 305. Based on the confidence value generated in step 303, key judgments are made. If the confidence is higher than 96%, it means that the user query is accurately adapted to a predefined scenario, and the system quickly enters step 306 to implement predefined business logic processing. At this stage, the system efficiently organizes and processes data resources to generate response results based on scenario-specific business logic rules (such as order status screening and related data integration rules in order query scenarios). If the confidence does not meet the standard, the process goes to step 305, and the large language model (LLM) is enabled to generate SQL to obtain data, expanding the system's ability to process unfamiliar or complex queries.
[0099] Step 305: LLM generates SQL to obtain data. In low-confidence scenarios, the system uses LLM's powerful semantic understanding and code generation capabilities to accurately convert user natural language queries into executable SQL statements. Based on the deep learning of electronic manufacturing business semantics and database architecture knowledge, LLM builds accurate data query logic, drives the database query engine to efficiently retrieve and extract target data, and provides data support for subsequent processing.
[0100] Step 306, predefined business logic processing. This step is initiated in the high-confidence predefined scenario matching scenario. The system strictly follows the predefined business process and rule templates, orderly schedules database query, data processing, business logic operations and other operations, deeply integrates multi-source data resources, such as summarizing order details according to order management scenario rules, associating production progress and material supply data, etc., to generate structured response results that match business logic and meet user needs, ensure business processing consistency, accuracy and reliability, and improve the quality of professional services of the system.
[0101] Step 307, RAG. This step is triggered when the semantic similarity of the knowledge base exceeds 95%. The RAG module first uses the embedding model to encode the user query into a high-dimensional semantic vector, and efficiently retrieves neighboring semantic knowledge fragments in the massive vector knowledge base. Then, it relies on the language model to deeply understand and integrate the semantic associations of knowledge fragments, and generates semantically coherent and knowledge-rich response content according to the user's demand context, realizing the synergy of accurate knowledge positioning and intelligent content generation, effectively enhancing the depth, breadth and accuracy of the system's knowledge service, and optimizing the user's knowledge acquisition experience, especially in dealing with complex knowledge query scenarios.
[0102] In a specific embodiment, the multi-agent data analysis step is combined with step 207, referring to Figure 4 , Figure 4 FIG. 4 shows a multi-agent linear structure processing data analysis task flow chart of the intelligent assistant performance optimization method according to the present invention, as shown in FIG. Figure 4 As shown in the figure, the data analysis task process includes:
[0103] Step 401, table data. Such table data originates from multiple business scenarios in the electronic manufacturing industry (such as production reports, sales data, inventory lists, etc.).
[0104] Step 402, data processing agent, using data processing agent to collect and clean data. On the one hand, focus on data collection tasks accurately, and screen and extract valid data subsets that meet analysis requirements from massive data elements in the table according to the established data requirement framework and potential analysis goals; on the other hand, start a refined cleaning process for possible quality defects in the collected data, such as format disorder, missing value dispersion, and outlier interference.
[0105] Step 403, data analysis agent, using data analysis agent to perform data analysis. The data refined by the data processing agent flows into the data analysis agent link. For example, mining product sales trends, seasonal fluctuation correlations and customer purchasing behavior patterns in sales data; exploring production efficiency bottlenecks, quality defect roots and equipment operation performance trends in production reports, generating multi-dimensional analysis insights and preliminary conclusions, providing data-driven basis for decision-making, helping to accurately grasp business trends and explore potential optimization space.
[0106] Step 404, the data analysis result check agent is used to verify the result. Steps 403 and 404 are executed repeatedly until the big model considers the result reliable. The big model is essentially a GPT model. Its calculation logic is to calculate the probability of the subsequent text according to the previous text. When the calculated probability of the subsequent text reaches a specific threshold, it considers the result reliable or the number of cycles reaches a manually specified threshold. If it is manually specified, it is mainly based on comprehensive consideration of business performance and efficiency, and step 405 is executed.
[0107] Step 405, analyzing the results.
[0108] In a specific embodiment, in combination Figure 5-Figure 7 , Figure 5 FIG. 4 shows a flowchart of a large language model training according to the present invention. Figure 5 As shown in the figure, the pre-training process uses the business text data of the electronic manufacturing industry as the core training data source, such as "Huawei has achieved end-to-end high-quality delivery through continuous management changes, reduced operating costs, improved efficiency, and achieved an indispensable role in process system. This document first introduces the value and connotation of Huawei's process system, the depth and breadth of the process system, what the process system does, and how the process system does it." In terms of training methods, a self-supervised learning strategy is adopted. Specifically, the GPT model is provided with the upper part of a text, and is required to predict the subsequent content based on the existing text information. Through this pre-training mode, the model can deeply mine and learn the relevant business knowledge of the electronic manufacturing industry, continuously optimize its own knowledge system and language comprehension ability, thereby providing a solid model foundation and knowledge reserve for subsequent applications in related business fields, and effectively improving the task processing efficiency and semantic understanding accuracy in the electronic manufacturing industry scenario.
[0109] Specifically, the fine-tuning stage in the training of the large language model mainly uses training data consisting of business question-answer pairs, such as business data in the form of (help me check the inventory, select * from stock), using the supervised fine-tuning (SFT) method, and using algorithms such as lora and qlora to accelerate training. Figure 6, Figure 6 The Lora structure diagram according to the present invention is shown as Figure 6 As shown in the figure, lora (Low-Rank Adaptation) is a technical means for efficient fine-tuning of large pre-trained language models (such as GPT, BERT, etc.). Its principle is not to directly update all the parameters of the pre-trained model (corresponding to the left part of the lora structure diagram), but to introduce low-rank (low-dimensional) trainable adaptation matrices (A and B) at specific key positions of the model, usually the weight matrix in the attention mechanism. Since the number of parameters of these adaptation matrices is greatly reduced compared to the original model parameters, the purpose of accelerating training is achieved. Continue to refer to Figure 7 , Figure 7 The structural difference diagram of Lora and Qlora according to the present invention is shown, as shown in Figure 7 As shown in the figure, Qlora is an extension of Lora technology. Its optimization method is similar to Lora. The difference is that Qlora integrates quantization technology. On the basis of fine-tuning with Lora, it implements quantization processing on the basic model, that is, converting high-precision parameters to low-precision (such as 8 bits or lower), thereby reducing the memory usage of the model and improving the calculation speed, which is especially suitable for resource-constrained environments. Through SFT training, the model can show stronger capabilities and adaptability in processing specific businesses (such as text to SQL statements, etc.).
[0110] Example 1
[0111] The query results are as follows:
[0112] Order Number Total price including tax Confirm delivery date Order status state SO241112006 6349.47 2024 / 11 / 25 No work order opened Reviewed SO241105005 293.8 2024 / 12 / 5 No work order opened Reviewed SO241030004 2260 2024 / 12 / 10 No work order opened Reviewed SO241106003 13609.72 2024 / 12 / 27 No work order opened Reviewed SO241118002 12537.35 2024 / 12 / 11 No work order opened Reviewed SO241120001 4654.47 2024 / 11 / 29 No work order opened Reviewed
[0113] Data analysis results:
[0114] After further analysis of the pre-processed manufacturing data, the following are the key insights and potential issues we have discovered:
[0115] Key insights: Order volume trend analysis shows that the number of orders in November and December 2024 increased significantly, which may be related to seasonal demand or promotional activities. Price distribution analysis shows that the product line covers different market positioning and meets the needs of customers of different sizes, with a price range from RMB 2,260 to RMB 13,609.72. Currently, the status of all orders is "not started" and "reviewed", indicating that there are no orders in production and all orders are waiting for production to start.
[0116] Potential issues: The order backlog is significant, and long-term backlogs may affect customer satisfaction and the company's reputation. Delivery dates are concentrated in specific months, and the company needs to ensure peak production capacity to avoid production delays. Data analysis is limited, lacking detailed information such as order content, customer information, and production progress, which affects in-depth analysis.
[0117] In summary, the current data analysis reveals that the company faces challenges in order processing and production, especially during the peak season. It is recommended that the company further optimize the order processing and production processes, improve operational efficiency, and enhance customer satisfaction.
[0118] Further references Figure 8 As an implementation of the above method, in a second aspect, the present application provides an embodiment of a system 800 for optimizing the performance of an intelligent assistant. Figure 1 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices. The system 800 includes a data acquisition module 801, a RAG module, a scene matching module 803, an LLM model three generation module 804 and a response output module 805 which are connected to each other in communication, wherein:
[0119] The data acquisition module 801 is configured to acquire user request data;
[0120] The RAG module 802 is configured to calculate the semantic similarity of the user request data based on a semantic matching algorithm, and in response to the semantic similarity being greater than a first threshold, search the knowledge base using the RAG algorithm to generate reply information;
[0121] The scene matching module 803 is configured to enter the predefined scene matching calculation in response to the semantic similarity being less than or equal to the first threshold, and if the judged scene matching confidence is greater than the second threshold, obtain the key elements of the user request data, splice and generate an SQL statement 1 that accurately adapts to the query intent according to the predefined logic template and SQL grammar rules, and generate a reply message;
[0122] LLM model three generating module 804, configured to convert the user request data into the intended SQL statement two based on LLM model three in response to the scene matching confidence being less than or equal to the second threshold, and generate reply information;
[0123] The reply output module 805 is configured to output reply information.
[0124] In a specific embodiment, first, the data acquisition module 801 obtains the query content input by the user in the form of text or voice, and then enters the refusal to answer recognition stage. The system accurately judges the user's query content. If it is determined to be in the category of idle chat or sensitive information, it will directly refuse to answer; then the solution selection operation is carried out. According to the user input content, it is matched with the RAG module 802 first, and the similarity calculation is implemented by semantic matching. When the similarity is greater than 95%, it is determined that the match is successful and the RAG module 802 is enabled to generate a reply. If the similarity is less than or equal to 95%, the system executes the scene matching module 803 to match the predefined scene. Once the scene probability is greater than 96%, it directly enters the scene matching module 803. If it is less than or equal to 96%, the LLM model three generation module 804 is used to directly generate the SQL solution. However, although this solution is relatively simple, it is highly dependent on the performance of the large model itself. In the scene matching stage, the system uses the scene matching module 803 to accurately identify the predefined scene to which the query belongs. The processing flow is then triggered. For predefined scenarios, the intelligent question-answering processing module is entered, and multiple rounds of dialogues are carried out with the help of a large language model to meet user needs. In the case of knowledge base matching, the RAG module 802 is called to perform knowledge base matching and obtain query results. For scenarios that do not support the RAG module 802 and the scenario matching module 803, the LLM model three generation module 804 is called to convert the query into an SQL statement to execute data query or operation.
[0125] Furthermore, for data analysis scenarios, the multi-agent data analysis module is called to deploy professional agents for in-depth analysis and provide decision support. During the entire interaction process, the security and filtering module monitors the interactive content in real time and intercepts inappropriate queries in a timely manner. Finally, the system feeds back the results to the user after processing by the corresponding module, successfully completing an interactive operation, effectively ensuring the accuracy, efficiency, security and adaptability of the system operation to user needs.
[0126] In a third aspect, the present application proposes a terminal device, including a processor, a memory, and a computer program stored in the memory, wherein the computer program is executed by the processor to implement any of the above display screen assembly methods.
[0127] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, any of the above display screen assembly methods is implemented.
[0128] Reference below Fig. 9 , which shows a schematic diagram of the structure of a computer system 900 of a terminal device or server suitable for implementing an embodiment of the present application. Fig. 9 The terminal device or server shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0129] like Fig. 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the computer system 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0130] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed.
[0131] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable medium of the present application can be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus or a device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus or a device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0132] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0133] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0134] Obviously, those skilled in the art can make various modifications and changes to the embodiments of the present invention without departing from the spirit and scope of the present invention. In this way, if these modifications and changes are within the scope of the claims of the present invention and their equivalents, the present invention is also intended to cover these modifications and changes. The word "comprising" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are recorded in mutually different dependent claims does not indicate that the combination of these measures cannot be used to profit. Any reference numerals in the claims should not be considered to limit the scope.
Claims
1. A method for optimizing the performance of an intelligent assistant, characterized in that: The method comprises: S1, obtain user request data; S2, calculating the semantic similarity of the user request data based on a semantic matching algorithm, and in response to the semantic similarity being greater than a first threshold, searching a knowledge base using a RAG algorithm to generate reply information; S3, in response to the semantic similarity being less than or equal to the first threshold, entering into a predefined scene matching calculation, and if the judged scene matching confidence is greater than a second threshold, obtaining the key elements of the user request data, splicing and generating an SQL statement 1 that accurately adapts to the query intent according to the predefined logic template and SQL grammar rules, and generating the reply information; S4, in response to the scenario matching confidence being less than or equal to a second threshold, converting the user request data into an intended SQL statement 2 based on the LLM model 3, and generating the reply information; S5, output the reply information.
2. A method for optimizing the performance of an intelligent assistant according to claim 1, characterized in that: The method further includes a refusal identification step disposed between the step S1 and the step S2, in response to identifying that the user request data is sensitive information, directly terminating the operation or issuing a reminder message.
3. The method for optimizing the performance of an intelligent assistant according to claim 1, characterized in that: The S2 step includes the following sub-steps: S21, collecting data pairs in business scenarios, where the data pairs include business questions and corresponding structured query statements; S22, using an embedding tool, accurately encode the business problem into a storage semantic vector, and store the storage semantic vector in a vector database in an orderly manner, and associate the corresponding business data with the knowledge base architecture; S23, in response to obtaining the user request data, reusing the embedding tool to convert the user request data into a request semantic vector; S24, using the semantic matching algorithm to calculate the similarity between the request semantic vector and the stored semantic vector, and in response to the similarity being greater than a first threshold, filtering out business data pairs corresponding to the most similar top 3 knowledge vectors; S25, based on the business logic and the user request context, splice into prompt1, and input the prompt1 into LLM model 1 to obtain the reply information.
4. The method for optimizing the performance of an intelligent assistant according to claim 1, characterized in that: The S3 step includes the following sub-steps: S31, collecting scene data pairs, using the scene classification model to mine the mapping logic between the text features of the scene data pairs and the business scene labels, and learning the text feature combination pattern to distinguish the business scene category; S32, formulating a dedicated SQL query statement for a predefined business scenario according to the rules, data structure, and query requirements, and storing it in a business scenario database; S33, in response to obtaining the user request data, using the scenario classification model to calculate the scenario matching confidence of the business scenario category to which the user request data belongs; S34, in response to the scene matching confidence being greater than the second threshold, extracting key elements of the user request data, and improving the corresponding dedicated SQL query statement, and splicing and generating the first intent SQL statement adapted to the query; S35, retrieve the corresponding intended SQL statement 1 from the database, execute data query, and generate the reply information.
5. A method for optimizing the performance of an intelligent assistant according to claim 4, characterized in that: The step S34 also includes extracting the key elements of the user request data, determining whether multiple rounds of extraction are required, and if the determination is "yes", performing multiple rounds of inquiries using the LLM model 2 until the complete intended SQL statement 1 is obtained.
6. The method for optimizing the performance of an intelligent assistant according to claim 4, characterized in that: The exclusive SQL query statement includes a "where" clause part, and the key elements are filled into the parameter positions corresponding to the "where" clause part.
7. The method for optimizing the performance of an intelligent assistant according to claim 1, characterized in that: After the S3 step and / or the S4 step generates the intended SQL statement one and / or the intended SQL statement two, it also includes executing data query and determining whether result analysis is required. If the determination is "yes", the reply information is formed after data analysis based on the Multi-agent architecture.
8. The method for optimizing the performance of an intelligent assistant according to claim 1, characterized in that: The value range of the first threshold and the second threshold is [90%, 98%].
9. A system for optimizing the performance of an intelligent assistant, characterized in that: The system is implemented based on a method for optimizing the performance of an intelligent assistant according to any one of claims 1 to 8, and the system includes: A data acquisition module, configured to acquire user request data; A RAG module is configured to calculate the semantic similarity of the user request data based on a semantic matching algorithm, and in response to the semantic similarity being greater than a first threshold, retrieve a knowledge base using the RAG algorithm to generate reply information; A scene matching module is configured to enter a predefined scene matching calculation in response to the semantic similarity being less than or equal to a first threshold, and if the judged scene matching confidence is greater than a second threshold, obtain key elements of the user request data, and splice and generate an SQL statement that accurately adapts to the query intent according to a predefined logical template and SQL grammar rules, and generate the reply information; An LLM model three generating module, configured to convert the user request data into an intended SQL statement two based on the LLM model three in response to the scene matching confidence being less than or equal to a second threshold, and generate the reply information; The reply output module is configured to output the reply information.
10. A computer-readable storage medium, wherein a computer program is stored in the medium, and when the computer program is executed by a processor, a method for optimizing the performance of an intelligent assistant as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Data result acquisition method and system based on large model analysis
CN121144349A
Intelligent question and answer method and system, medium, equipment and program product
CN121233705A
Personalized content generation method and device, electronic equipment and storage medium
CN121681896A