A large model question answering optimization system and method combining functional functions and slot filling
By combining functional functions and slot filling, the problem of insufficient answer accuracy and complex problem processing capabilities in the large language model in the question and answer system is solved, and more accurate user intention understanding and more professional function calls are achieved, which improves the consistency and user experience of question and answers.
Patent Information
- Application Number
- CN202510122166.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The existing large language models have insufficient answer accuracy, consistency and complex problem-solving capabilities in the question-and-answer system, and the general model is not trained for specific fields or scenarios, resulting in the inability to generate answers that meet user requirements.
A question-answer optimization system combining functional functions and slot filling, through user intention analysis modules, intention matching judgment modules, slot filling processing modules, function function processing modules and question-answer display modules, identify and fill key information in the query, and call corresponding functional functions or execute specific logic to complete tasks.
It effectively solves the problem that the large language model is stubborn and unable to understand user intentions in the Q&A system, improves the accuracy of user's intention understanding and the professionalism of functional functions, maintains the consistency and consistency of Q&A, and provides a smoother user experience.
Smart Images

Figure CN119557410B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent question - answering systems, and particularly relates to a large - model question - answering optimization system and method combining functional functions and slot filling. Background Art
[0002] Human - computer interaction is the science that studies the interaction relationship between a system and a user. The system can be various machines, or computerized systems and software. Through human - computer interaction, various artificial - intelligence systems can be realized, such as intelligent customer - service systems, voice - control systems, large models, etc.; Frequently Asked Questions (FAQ) is a standard system for most products, which helps users find questions by themselves to reduce the cost of artificial customer service. The common FAQ systems on the market are question - answering - type FAQ systems. In a chat dialog box, users directly consult questions, and the system provides an answer to a similar question based on keywords or the latest similarity algorithm. Among them, the Large Language Model (LLM, also known as the large model) has become one of the core technologies in the field of natural language processing. These models are trained on large - scale text data, can capture rich language rules and knowledge information, and thus possess powerful natural - language understanding and generation capabilities, and are widely applied to various question - answering systems.
[0003] However, although large language models have achieved remarkable results in many natural - language processing tasks, their performance in question - answering systems, especially in answering professional questions, still needs to be improved. This is because traditional question - answering systems usually rely on manually designed rules or specific knowledge bases to answer questions, while large language models attempt to automatically learn the ability to answer questions from a vast amount of text data. This difference leads to some challenges for large language models in question - answering systems, such as the accuracy and consistency of answers and the ability to handle complex questions. Moreover, general large language models have not been trained for a specific domain or usage scenario, so general large language models cannot generate answers that meet user requirements in a specific domain. In addition, in existing systems, when dealing with complex queries, problems such as inaccurate slot filling and low efficiency of functional - function calls often occur, resulting in an unsmooth question - answering process and poor user experience. Summary of the Invention
[0004] Aiming at the above - mentioned deficiencies of the prior art, this application provides a large - model question - answering optimization system and method combining functional functions and slot filling.
[0005] In the first aspect, this application proposes a large - model question - answering optimization system combining functional functions and slot filling, including a user - intention analysis module, an intention - matching judgment module, a slot - filling processing module, a functional - function processing module, and a question - answering display module:
[0006] The user intention analysis module is used to build an intention intelligent analysis model and obtain the problem description information input by the user, input the problem description information into the intention intelligent analysis model for user intention analysis and calculation, and obtain an intention intelligent judgment result;
[0007] The intention matching and judgment module is used to compare the intention intelligent judgment result with the preset intention types in the intention library to determine whether it matches the preset intention types in the intention library. When the match is successful, it means that the preset intention is hit, and the processing flow of the slot filling processing module is executed. When the match fails, it means that the preset intention is not hit, and the processing flow of the function function processing module is executed;
[0008] The slot filling processing module is used to enable the slot pool and call the slot template matching the preset intention type, perform slot filling according to the content to be filled in the slot template, and generate a first type of optimized response statement adapted to the problem description information according to the filling result;
[0009] The function function processing module is used to call the preset corresponding function function according to the intention intelligent judgment result, parse the problem description information according to the function function, input the parsed information into the large model for keyword extraction, and generate a second type of optimized response statement adapted to the problem description information based on the keyword extraction result;
[0010] The question and answer display module is used to display the first type of optimized response statement and the second type of optimized response statement to the user.
[0011] In some optional implementation manners of some embodiments, the user intention analysis module includes an intention intelligent analysis model construction unit, which is used to build the intention intelligent analysis model through an embedding layer, a self-att layer, a FeedForward layer, and a Layernorm layer, expressed as:
[0012] y = hidden + MLP(LayerNorm(hidden)) + Attention(LayerNorm(hidden))
[0013] Among them, the calculation of the embedding layer is:
[0014]
[0015] Among them, x is the id representation obtained by mapping each character through the vocabulary, A is the weight matrix of the embedding layer, b is the bias matrix of the embedding layer, hidden represents the result obtained after the current input data passes through the embedding layer. Next, hidden passes through the self-att layer and the FeedForward layer respectively. Attention() represents the calculation process of the self-att layer, MLP() represents the calculation process of the FeedForward layer, LayerNorm() represents the normalization of hidden, and y represents the result of intent intelligent judgment.
[0016] In some optional implementation manners of some embodiments, the user intent analysis module further includes an intelligent analysis and judgment unit, which is used to input the problem description information into the intent intelligent analysis model for user intent analysis and calculation, including the calculation process of the self-att layer:
[0017] The input hidden first passes through the layernorm layer for normalization:
[0018] The calculation of the layernorm layer is:
[0019]
[0020] Among them, represents the mean of the current input data, represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being zero, and the value is 0.0001, represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer;
[0021] The output of the layernorm layer passes through three linear layers to obtain different vector representations: q, k, v:
[0022]
[0023] Among them, is the output result after passing through the layernorm layer, , 2, 3 are the weight matrices of the query layer, key layer, and value layer respectively, and b1, b2, and b3 are the bias matrices of the query layer, key layer, and value layer respectively;
[0024] Next, the self-attention attention of q, k, and v is calculated:
[0025]
[0026] Among them, m and n respectively represent the m-th character and the n-th character; d is a user-defined hyperparameter, exp is the exponential function with the natural constant e as the base, N represents the character length of the current text, is the attention score between the m-th character and the n-th character, is the attention score matrix of the current m-th character and all characters in the current text. The attention score matrices of all characters are saved as hidden;
[0027] It also includes the calculation process of the FeedForward layer:
[0028] The input hidden first undergoes normalization through the layernorm layer:
[0029] The calculation of the layernorm layer is:
[0030]
[0031] Among them, hidden represents the current input data, represents the mean of the current input data, represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being zero, with a value of 0.0001, represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer;
[0032] Next, the calculation of the FeedForward layer is performed:
[0033]
[0034] Among them, the calculation formula of the GeLU activation function is:
[0035]
[0036] Among them, tanh is the activation function, and the formula is:
[0037]
[0038] Among them, is the output result after passing through the layernorm layer, , 2 are respectively different weight matrices, and b1, b2 are respectively different bias matrices, is the final output calculated by the multi-head matrix calculation layer, represents the calculation output of the first layer network of the FeedForward layer, represents the calculation output of the intermediate activation function layer of the FeedForward layer.
[0039] In some alternative implementation manners of some embodiments, obtaining an intention intelligent judgment result through the intelligent analysis and judgment unit includes:
[0040] After the FeedForward layer calculation is completed, it is then calculated through the softmax layer. The softmax layer represents a normalization layer, and the calculation is as follows:
[0041]
[0042] Represents the final output of the softmax layer. Represents An exponential function with the natural constant e as the base. i and j respectively represent the i-th and j-th inputs, and ∑ represents accumulation.
[0043] The calculation result of the softmax layer is a matrix with the same dimension as the number of intention categories. Select the index where the maximum value in the matrix is located, and use this index to find the corresponding intention in the list storing intention categories, obtaining the intention value predicted by the model, which corresponds to the intention intelligent judgment result.
[0044] In some alternative implementation manners of some embodiments, the intention matching and judgment module includes an intention library construction unit and an intention matching unit;
[0045] The intention library construction unit is used to preset multiple intention types according to the business scenario, save the preset intention types and name them as the intention library to complete the construction of the intention library;
[0046] The intention matching unit is used to match the intention intelligent judgment result with the information in the intention library through a preset matching rule. If the match is successful, it means that the preset intention is hit; if the match fails, it means that the preset intention is not hit.
[0047] In some alternative implementation manners of some embodiments, the slot filling processing module includes: a slot template calling unit, a slot preliminary filling unit, a filling completeness judgment unit, a first type of rhetorical question construction unit, and a first type of reply statement construction unit;
[0048] The slot template calling unit is used to enable the slot pool and call a slot template that matches the preset intention type;
[0049] The slot preliminary filling unit is used to call the large model to extract keywords from the question description information according to the content to be filled in the slot template by means of prompt extraction, input the extracted keywords into the large model for key information extraction, and supplement the slots in the slot template according to the key information extraction result to obtain a preliminary filling result;
[0050] The filling completeness judgment unit is configured to judge whether the key information in the slot is missing according to the preliminary filling result. If it is missing, the first type of rhetorical question input unit is executed; if not, the first type of response sentence input unit is executed.
[0051] The first type of rhetorical question construction unit is configured to construct a question and rhetorically ask the user for the missing information when it is determined that there is still missing information after the key information in the slot is supplemented.
[0052] The first type of response sentence construction unit is configured to directly input the filled slot template information into the large model for direct response when it is determined that the key information in the slot has been supplemented.
[0053] In some optional implementation manners of some embodiments, the functional function processing module includes a functional function encapsulation unit, a search and parsing unit, and a second type of optimized response sentence construction unit:
[0054] The functional function encapsulation unit is configured to encapsulate various query functions into functions called through APIs in advance and store them in a repository.
[0055] The search and parsing unit is configured to call the large model and extract keywords from the problem description information by means of prompt word extraction, input the extracted keywords into the large model for key information extraction, and parse the problem description information according to the key information extraction result to obtain a search and parsing result.
[0056] The second type of optimized response sentence construction unit is configured to select a corresponding functional function from the repository for API call according to the search and parsing result, extract parsing information from the search and parsing result according to the called functional function, convert the parsing information extraction result into natural language through the large model, and generate a second type of optimized response sentence according to the converted natural language.
[0057] In some optional implementation manners of some embodiments, the functional function processing module further includes a request construction and sending unit and a response receiving and parsing unit;
[0058] The request construction and sending unit is configured to construct a request that meets the requirements of the corresponding API according to the key information extracted from the search and parsing result, including setting request headers, request methods, request bodies, and URL parameters; and send the constructed request to the corresponding API server through the HTTP protocol.
[0059] The response receiving and parsing unit is used to process and return a response containing the key information in the search parsing result after the API server receives a request, return the result data in JSON format, parse the result data, and extract the information corresponding to the search parsing result.
[0060] In a second aspect, the present application proposes a large model question and answer optimization method combining functional functions and slot filling, including the following steps;
[0061] Construct an intent intelligent analysis model and obtain the question description information input by the user, input the question description information into the intent intelligent analysis model for user intent analysis calculation, and obtain an intent intelligent judgment result;
[0062] Compare the intent intelligent judgment result with the preset intent types in the intent library to determine whether it matches the preset intent types in the intent library. When the match is successful, it means that the preset intent is hit, and the slot filling processing flow is executed. When the match fails, it means that the preset intent is not hit, and the functional function processing flow is executed;
[0063] Enable the slot pool and call the slot template matching the preset intent type, perform slot filling according to the content to be filled in the slot template, and generate the first type of optimized response statement adapted to the question description information according to the filling result;
[0064] Call the preset corresponding functional function according to the intent intelligent judgment result, parse the question description information according to the functional function, input the parsed information into the large model for keyword extraction, and generate the second type of optimized response statement adapted to the question description information based on the keyword extraction result;
[0065] Display the first type of optimized response statement and the second type of optimized response statement to the user.
[0066] In a third aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0067] Advantages of the present invention:
[0068] The present invention identifies and fills in key information (such as time, location, person, etc.) in the query through slot filling, combines the query intention to call the corresponding functional function or execute specific logic to complete the task, effectively solving the problems often occurring in the application scenarios of existing large language models, such as stubbornly sticking to one's own opinions, repeatedly giving the same wrong answers to multiple questions about the same problem and being unable to understand the user's intention. This enables the large model to more accurately understand the intention of the user's question when processing the user's question, and more professionally call the functional function to answer the user's question. It can maintain coherence and consistency in multi-round Q&A, provide a more fluent user experience, and optimize the overall Q&A effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is the system schematic diagram of the device.
[0070] Figure 2 is the overall flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0072] In a first aspect, the present application proposes a large model Q&A optimization system combining functional functions and slot filling, as Figure 1 shown, including a user intention analysis module, an intention matching judgment module, a slot filling processing module, a functional function processing module, and a Q&A display module:
[0073] The user intention analysis module is used to build an intention intelligent analysis model and obtain the problem description information input by the user, input the problem description information into the intention intelligent analysis model for user intention analysis calculation, and obtain an intention intelligent judgment result;
[0074] The intention matching judgment module is used to compare the intention intelligent judgment result with the preset intention types in the intention library to judge whether it matches the preset intention types in the intention library. When the match is successful, it means that the preset intention is hit, and the processing flow of the slot filling processing module is executed. When the match fails, it means that the preset intention is not hit, and the processing flow of the functional function processing module is executed;
[0075] The slot filling processing module is used to enable the slot pool and call the slot template matching the preset intention type, perform slot filling according to the content to be filled in the slot template, and generate a first type of optimized response statement adapted to the problem description information according to the filling result;
[0076] The functional function processing module is used to call a preset corresponding functional function according to the intelligent judgment result of the intention, parse the problem description information according to the functional function, input the parsed information into a large model for keyword extraction, and generate a second type of optimized response statement adapted to the problem description information based on the keyword extraction result;
[0077] The question and answer display module is used to display the first type of optimized response statement and the second type of optimized response statement to the user.
[0078] In some optional implementation manners of some embodiments, the user intention analysis module includes an intention intelligent analysis model construction unit, which is used to construct the intention intelligent analysis model through an embedding layer, a self-att layer, a FeedForward layer, and a Layernorm layer, expressed as:
[0079] y = hidden + MLP(LayerNorm(hidden)) + Attention(LayerNorm(hidden))
[0080] Among them, the calculation of the embedding layer is:
[0081]
[0082] Among them, x is the id representation obtained by mapping each character through the vocabulary, A is the weight matrix of the embedding layer, b is the bias matrix of the embedding layer, hidden represents the result obtained after the current input data passes through the embedding layer. Next, hidden passes through the self-att layer and the FeedForward layer respectively. Attention() represents the calculation process of the self-att layer, MLP() represents the calculation process of the FeedForward layer, LayerNorm() represents the normalization of hidden, and y represents the intelligent judgment result of the intention.
[0083] In some optional implementation manners of some embodiments, the user intention analysis module further includes an intelligent analysis and judgment unit, which is used to input the problem description information into the intention intelligent analysis model for user intention analysis calculation, including the calculation process of the self-att layer:
[0084] The input hidden first passes through the layernorm layer for normalization:
[0085] The calculation of the layernorm layer is:
[0086]
[0087] Among them, represents the mean of the current input data, represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being zero, and the value is 0.0001, represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer;
[0088] The output of the layernorm layer passes through three linear layers to obtain different vector representations: q, k, v:
[0089]
[0090] Among them, is the output result after passing through the layernorm layer, , 2, 3 are the weight matrices of the query layer, key layer, and value layer respectively, and b1, b2, b3 are the bias matrices of the query layer, key layer, and value layer respectively;
[0091] Next, the self-attention calculation of q, k, v is performed:
[0092]
[0093] Among them, m and n represent the m-th character and the n-th character respectively; d is a user-defined hyperparameter, exp is the exponential function with the natural constant e as the base, and N represents the character length of the current text, is the attention score between the m-th character and the n-th character, is the attention score matrix of the current m-th character and all characters in the current text, and the attention score matrix of all characters is saved as hidden;
[0094] It also includes the calculation process of the FeedForward layer:
[0095] The input hidden first passes through the layernorm layer for normalization:
[0096] The calculation of the layernorm layer is:
[0097]
[0098] Among them, hidden represents the current input data, represents the mean of the current input data, represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being zero, with a value of 0.0001. represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer;
[0099] Next, perform the calculation of the FeedForward layer:
[0100]
[0101] Among them, the calculation formula of the GeLU activation function is:
[0102]
[0103] Among them, tanh is the activation function, and the formula is:
[0104]
[0105] Among them, is the output result after passing through the layernorm layer, , are respectively different weight matrices, and b1, b2 are respectively different bias matrices. is the final output calculated by the multi-head matrix calculation layer, represents the calculation output of the first layer network of the FeedForward layer, represents the calculation output of the intermediate activation function layer of the FeedForward layer.
[0106] In some optional implementation manners of some embodiments, obtaining the intention intelligent judgment result through the intelligent analysis judgment unit includes:
[0107] After the calculation of the FeedForward layer is completed, it is then calculated through the softmax layer. The softmax layer represents the normalization layer, and the calculation is:
[0108]
[0109] represents the final output of the softmax layer, represents the exponential function with the natural constant e as the base. i and j respectively represent the i-th and j-th inputs, and ∑ represents the accumulation;
[0110] The calculation result of the softmax layer is a matrix with the same dimension as the number of intention categories. Select the index where the maximum value in the matrix is located, and use this index to find the corresponding intention in the list storing the intention categories to obtain the intention value predicted by the model, corresponding to the intention intelligent judgment result.
[0111] In some alternative implementation manners of some embodiments, the intention matching and judging module includes an intention library construction unit and an intention matching unit;
[0112] The intention library construction unit is configured to preset multiple intention types according to the business scenario, save the preset intention types and name them as the intention library, and complete the construction of the intention library;
[0113] The intention matching unit is configured to match the intelligent intention judgment result with the information in the intention library through a preset matching rule. If the match is successful, it indicates that the preset intention is hit; if the match fails, it indicates that the preset intention is not hit.
[0114] Among them, the matching process between the intention recognition result and the intention library can be implemented in the following several ways:
[0115] 1. Matching based on vector similarity: First, represent the intelligent intention judgment result and the intentions in the intention library in vector form (for example, through word embedding models such as BERT or Word2Vec). Then calculate the similarity between the intelligent intention judgment result vector and each intention vector in the intention library (such as cosine similarity, Mahalanobis distance, etc.). Judge whether there is a match according to the similarity score. If the similarity score is higher than the preset threshold, it is considered that the intelligent intention judgment result hits an intention in the intention library; otherwise, it is considered not to hit.
[0116] 2. Matching based on text matching: Directly match the intelligent intention judgment result with the text in the intention library. A text matching model (such as TextCNN, Transformer, etc.) can be used to calculate the similarity between the two. If the matching degree is higher than the set threshold, it is considered that the intelligent intention judgment result matches an intention in the intention library successfully; otherwise, it does not hit. 3. Keyword matching: Extract the keywords in the intelligent intention judgment result and match them with the preset keywords in the intention library. If the keywords in the intelligent intention judgment result are the same as or highly relevant to the keywords in the intention library, it is considered a successful match; otherwise, it does not hit.
[0117] In some alternative implementation manners of some embodiments, the slot filling processing module includes: a slot template calling unit, a slot preliminary filling unit, a filling completeness judgment unit, a first type of rhetorical question sentence construction unit, and a first type of reply sentence construction unit;
[0118] The slot template calling unit is configured to enable the slot pool and call the slot template that matches the preset intention type;
[0119] The slot preliminary filling unit is used to call the large model to extract keywords from the problem description information according to the content to be filled in the slot template by means of prompt extraction, input the extracted keywords into the large model for key information extraction, and supplement the slots in the slot template according to the key information extraction results to obtain a preliminary filling result;
[0120] When the preset intention is hit: Call the slot pool, select the slot template corresponding to the currently hit intention, and the sample of the hit slot template is as follows:
[0121] "Book movie tickets": {
[0122] "Movie name": __,
[0123] "Cinema name": _,
[0124] "Time": _,
[0125] "Quantity": _,
[0126] "Seat location": _,
[0127] }
[0128] Next, call the large model to extract keywords from the user query (problem description information) according to the content to be filled in the slot by means of prompt. The prompt constructed for inputting into the large model is:
[0129] "The current user input query is: Book a movie ticket for the XX (movie name) show this afternoon. Please extract the following information from this query: "Movie name", "Cinema name", "Time", "Quantity", "Seat location"".
[0130] Input the defined keyword extraction prompt into the large model for key information extraction. Then, supplement the slots in the slot template according to the extraction results; judge whether it is necessary to ask the user a follow-up question according to the supplementary situation;
[0131] The filling completeness judgment unit is used to judge whether the key information in the slot is missing according to the preliminary filling result. If it is missing, execute the first type of follow-up question input unit; if it is not missing, execute the first type of reply statement input unit;
[0132] Among them, judging whether there is any missing key information in the slot after supplementation can be achieved through the following methods:
[0133] 1. Integrity check based on the slot template:
[0134] Slot templates usually preset the key information fields that need to be filled. After supplementing the key information, it is possible to judge whether there is any omission by checking whether these fields are filled correctly. For example, if a slot template contains multiple fields (such as time, place, person, etc.), then it is necessary to check one by one whether these fields have been filled.
[0135] 2. Confidence based on model prediction:
[0136] During the slot filling process, the confidence output by the model can be used to judge whether the information is complete. For example, if the confidence of a certain slot is lower than the preset threshold, it is considered that the information in this slot may be incomplete. This method relies on the prediction ability of the model and is usually implemented in combination with deep learning models (such as BERT, CNN, etc.).
[0137] 3. Rule-based check:
[0138] Judge whether the slot information is complete through preset rules. For example, if a certain field in the slot template is a required item but the model fails to extract the relevant information, it is considered that this field is missing. This method is simple and efficient, but depends on the completeness of the rules.
[0139] The first type of rhetorical question construction unit is used to construct questions and rhetorically ask the user for the missing information when it is determined that there is still a missing part after the key information in the slot has been supplemented;
[0140] The first type of reply statement construction unit is used to directly input the filled slot template information into the large model for direct reply when it is determined that the key information in the slot has been supplemented.
[0141] In some optional implementation manners of some embodiments, the functional function processing module includes a functional function encapsulation unit, a search and parsing unit, and a second type of optimized response statement construction unit:
[0142] The functional function encapsulation unit is used to encapsulate various query functions in advance into functions called through APIs and store them in the repository;
[0143] The search and parsing unit is used to call the large model and extract keywords from the problem description information by means of prompt extraction, input the extracted keywords into the large model for key information extraction, and parse the problem description information according to the key information extraction result to obtain the search and parsing result;
[0144] The second type of optimized response statement construction unit is used to select a corresponding function from the repository for API call according to the search and parsing result, extract parsing information from the search and parsing result according to the called function, convert the parsing information extraction result into natural language through the large model, and generate a second type of optimized response statement according to the converted natural language.
[0145] When the preset intention is not hit: various functions such as querying air quality index, querying temperature, and hot search are encapsulated in advance into functions that can be called through APIs as preset function, and stored in the repository. When the intention extracted by the large model does not hit the preset intention, we perform keyword matching to determine whether it hits the preset function in the repository. Suppose the user's input is: "What is the air quality index in Tianjin today?"
[0146] 1. When hitting the function in the repository, first perform input parsing. Use the large model to obtain the keywords in the input query by constructing a prompt, and extract the key information from the text, that is, "Tianjin" as the location and "today" as the time. Plus the previously recognized intention: "query air quality index" as the output of the parsing;
[0147] 2. Determine the function or API to call
[0148] The large model selects the "air quality index query" api interface in the repository according to "query air quality index" in the parsing result for call, and decides to call an API for querying air quality index to obtain data.
[0149] 3. Construct a request and call the API
[0150] ① Construct a request
[0151] The large model constructs a request that meets the requirements of the air quality index API according to the extracted key information (location and time). This usually includes setting request headers (such as Content-Type), request methods (such as GET or POST), request bodies (for POST requests), and URL parameters (such as location and time as query parameters).
[0152] For example, the constructed URL can be:
[0153] "https: / / api.weather.com / data / 2.5 / weather?q= Tianjin&appid=YOUR_API_KEY&units=metric&lang=zh_cn"
[0154] ② Send the request
[0155] The large model sends the constructed request to the server of the air quality index API through the HTTP protocol (GET or POST request).
[0156] 4. Process the API response
[0157] ① Receive the response
[0158] After receiving the request, the air quality index API server processes and returns a response containing air quality index information. The result data in JSON format is returned.
[0159] ② Parse the response
[0160] The large model parses the response data and extracts the air quality index information that the user cares about, such as temperature, humidity, wind direction, etc.
[0161] 5. Generate a reply
[0162] Construct a prompt with the parsed air quality index information and input it into the large model. The large model converts the parsed air quality index information into natural language and generates an easy-to-understand second type of optimized response statement, such as: "The weather in Tianjin is sunny today, with a temperature of 30°C, a humidity of 60%, and a southwest wind of level 2."
[0163] In some optional implementation manners of some embodiments, the functional function processing module further includes a request construction and sending unit and a response receiving and parsing unit;
[0164] The request construction and sending unit is used to construct a request that meets the requirements of the corresponding API according to the key information extracted from the search and parsing results, including setting the request header, request method, request body, and URL parameters; and sending the constructed request to the corresponding API server through the HTTP protocol;
[0165] The response receiving and parsing unit is used to, after the API server receives the request, process and return a response containing the key information in the search and parsing results, and return the result data in JSON format;
[0166] Parse the result data and extract the information corresponding to the search and parsing results.
[0167] In a second aspect, the present application proposes a large model question and answer optimization method combining functional functions and slot filling, as Figure 2 shown, including the following steps;
[0168] S100: Build an intent intelligent analysis model and obtain the problem description information input by the user. Input the problem description information into the intent intelligent analysis model for user intent analysis and calculation to obtain an intent intelligent judgment result;
[0169] S200: Compare the intelligent intent judgment result with the preset intent types in the intent library to determine whether it matches the preset intent types in the intent library. When the match is successful, it indicates that the preset intent is hit, and the slot filling processing flow is executed. When the match fails, it indicates that the preset intent is not hit, and the function function processing flow is executed;
[0170] S300: Enable the slot pool and call the slot template that matches the preset intent type. Perform slot filling according to the content to be filled in the slot template, and generate the first type of optimized response statement adapted to the problem description information according to the filling result;
[0171] S400: Call the preset corresponding function function according to the intelligent intent judgment result, parse the problem description information according to the function function, input the parsed information into the large model for keyword extraction, and generate the second type of optimized response statement adapted to the problem description information based on the keyword extraction result;
[0172] S500: Display the first type of optimized response statement and the second type of optimized response statement to the user.
[0173] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0174] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.
[0175] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0176] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0177] In the embodiments provided in this disclosure, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0178] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0179] In addition, the functional units in each embodiment of this disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0180] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0181] The above are only the preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, without departing from the premise of the technical solution of the present invention, several modified and improved technical solutions should also be regarded as falling within the scope protected by this claim book.
Claims
1. A large-model question-answering optimization system combining functional functions and slot filling, characterized by: It includes user intention analysis module, intention matching judgment module, slot filling processing module, function processing module and question and answer display module: The user intention analysis module is used to build an intention intelligent analysis model and obtain the question description information input by the user, input the question description information into the intention intelligent analysis model to perform user intention analysis calculation, and obtain the intention intelligent judgment result. The user intention analysis module includes an intention intelligent analysis model construction unit, which is used to build the intention intelligent analysis model through the embedding layer, self-att layer, FeedForward layer and Layernorm layer, which is expressed as: y=hidden+MLP(LayerNorm(hidden))+Attention(LayerNorm(hidden)) Among them, the calculation of the embedding layer is: Among them, x is the id obtained by mapping each character into the vocabulary, A is the weight matrix of the embedding layer, b is the bias matrix of the embedding layer, hidden represents the result obtained after the current input data passes through the embedding layer, then hidden passes through the self-att layer and FeedForward layer respectively, Attention() represents the calculation process of the self-att layer, MLP() represents the calculation process of the FeedForward layer, LayerNorm() represents the normalization of hidden, and y represents the result of intelligent judgment of intent; The intention matching judgment module is used to compare the intention intelligent judgment result with the preset intention type in the intention library to determine whether it matches the preset intention type in the intention library. When the match is successful, it means that the preset intention is hit, and the processing flow of the slot filling processing module is executed. When the match fails, it means that the preset intention is not hit, and the processing flow of the function function processing module is executed; The slot filling processing module is used to enable the slot pool and call the slot template that matches the preset intent type, fill the slot according to the content to be filled in the slot template, and generate a first type of optimized response statement adapted to the problem description information according to the filling result; The function processing module is used to call the preset corresponding function according to the intention intelligent judgment result, parse the problem description information according to the function function, input the parsed information into the big model for keyword extraction, and generate a second type of optimized response statement suitable for the problem description information based on the keyword extraction result; The question and answer display module is used to display the first type of optimized response statements and the second type of optimized response statements to the user.
2. The system according to claim 1, characterized in that: The user intention analysis module also includes an intelligent analysis and judgment unit, which is used to input the problem description information into the intention intelligent analysis model to perform user intention analysis calculation, including the calculation process of the self-att layer: The input hidden is first normalized by the layernorm layer: The calculation of the layernorm layer is: in, Represents the mean of the current input data, Represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being 0, and its value is 0.0001. represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer; the output of the layernorm layer passes through three linear layers to obtain different vector representations: q, k, v: in, is the output result after the layernorm layer, , 2, 3 are the weight matrices of the query layer, key layer, and value layer, respectively; b1, b2, and b3 are the bias matrices of the query layer, key layer, and value layer, respectively; Next, calculate the self-attention for q, k, and v: Where m and n represent the mth and nth characters respectively; d is a custom hyperparameter, exp is an exponential function with the natural constant e as the base, and N represents the character length of the current text. is the attention score between the mth character and the nth character, is the attention score matrix of the current mth character and all the characters in the current text, and the attention score matrix of all characters is saved as hidden; It also includes the calculation process of the FeedForward layer: The input hidden is first normalized by the layernorm layer: The calculation of the layernorm layer is: Among them, hidden represents the current input data, Represents the mean of the current input data, Represents the variance of the input data hidden, is a hyperparameter to prevent the denominator from being 0, and its value is 0.0001. represents the weight matrix of the LN layer, and β represents the bias weight of the LN layer; Next, the calculation of the FeedForward layer is performed: Among them, the calculation formula of the GeLU activation function is: Among them, tanh is the activation function, and the formula is: in, is the output result after the layernorm layer, , 2 are different weight matrices, b1 and b2 are different bias matrices, Calculate the final output for the multi-head matrix calculation layer, Represents the calculation output of the first layer network of the FeedForward layer, Represents the calculation output of the intermediate activation function layer of the FeedForward layer.
3. The system according to claim 2, characterized in that: The intelligent analysis and judgment unit obtains the intention intelligent judgment result, including: after the FeedForward layer is calculated, it is calculated by the softmax layer, and the softmax layer represents the normalization layer, which is calculated as: Represents the final output of the softmax layer. represent An exponential function with the natural constant e as the base, i and j represent the i-th and j-th inputs respectively, and ∑ represents accumulation; The calculation result of the softmax layer is a matrix with the same dimension as the number of intent categories. The index with the largest value in the matrix is selected, and the corresponding intent is found in the list of intent categories using the index to obtain the intent value predicted by the model, which corresponds to the intelligent judgment result of the intent.
4. The system according to claim 3, characterized in that: The intention matching judgment module includes an intention library construction unit and an intention matching unit; The intent library construction unit is used to preset multiple intent types according to business scenarios, save the preset intent types and name them as intent libraries, and complete the intent library construction; The intention matching unit is used to match the intention intelligent judgment result with the information in the intention library through preset matching rules. If the match is successful, it means that the preset intention is hit; if the match fails, it means that the preset intention is not hit.
5. The system according to claim 4, characterized in that: The slot filling processing module includes: a slot template calling unit, a slot preliminary filling unit, a filling completeness judgment unit, a first type of rhetorical question sentence construction unit and a first type of reply sentence construction unit; The slot template calling unit is used to enable the slot pool and call the slot template matching the preset intent type; The slot preliminary filling unit is used to call the large model to extract keywords from the problem description information according to the content to be filled in the slot template by means of prompt word extraction, input the extracted keywords into the large model to extract key information, and supplement the slots in the slot template according to the key information extraction results to obtain a preliminary filling result; The filling completeness judgment unit is used to judge whether the key information in the slot is missing according to the preliminary filling result, and if it is missing, execute the first type of rhetorical question input unit; if it is not missing, execute the first type of reply sentence input unit; The first type of rhetorical question sentence construction unit is used to construct a question sentence for the missing information and ask the user a rhetorical question when it is determined that the key information in the slot is still missing after being supplemented; The first type of reply statement construction unit is used to directly input the supplemented slot template information into the large model for direct reply when it is determined that the key information in the slot has been supplemented.
6. The system according to claim 5, characterized in that: The function processing module includes a function encapsulation unit, a search and analysis unit, and a second type of optimized response statement construction unit: The function encapsulation unit is used to encapsulate various query functions into functions called by API in advance and store them in the storage repository; The search and analysis unit is used to call the big model and extract keywords from the problem description information by means of prompt word extraction, input the extracted keywords into the big model to extract key information, and parse the problem description information according to the key information extraction result to obtain the search and analysis result; The second type of optimized response statement construction unit is used to select a corresponding function function from the repository to make an API call according to the search and analysis results, extract parsing information from the search and analysis results according to the called function function, convert the parsing information extraction results into natural language through the large model, and generate a second type of optimized response statement according to the converted natural language.
7. The system according to claim 6, characterized in that: The functional function processing module also includes a request construction sending unit and a response receiving parsing unit; The request construction and sending unit is used to construct a request that meets the requirements of the corresponding API according to the key information extracted from the search and analysis results, including setting the request header, request method, request body and URL parameters; and send the constructed request to the corresponding API server through the HTTP protocol; The response receiving and parsing unit is used to process and return a response containing key information in the search and parsing result after the API server receives the request, return the result data in json format, parse the result data, and extract information corresponding to the search and parsing result.
8. A large-model question-answering optimization method combining functional functions and slot filling, characterized in that: The steps include: Construct an intention intelligent analysis model and obtain the question description information input by the user, input the question description information into the intention intelligent analysis model to perform user intention analysis and calculation, and obtain the intention intelligent judgment result, wherein the intention intelligent analysis model is constructed through the embedding layer, self-att layer, FeedForward layer and Layernorm layer, which is expressed as: y=hidden+MLP(LayerNorm(hidden))+Attention(LayerNorm(hidden)) Among them, the calculation of the embedding layer is: Among them, x is the id obtained by mapping each character into the vocabulary, A is the weight matrix of the embedding layer, b is the bias matrix of the embedding layer, hidden represents the result obtained after the current input data passes through the embedding layer, then hidden passes through the self-att layer and FeedForward layer respectively, Attention() represents the calculation process of the self-att layer, MLP() represents the calculation process of the FeedForward layer, LayerNorm() represents the normalization of hidden, and y represents the result of intelligent judgment of intent; The intention intelligent judgment result is compared with the preset intention type in the intention library to determine whether it matches the preset intention type in the intention library. When the match is successful, it means that the preset intention is hit, and the slot filling process is executed. When the match fails, it means that the preset intention is not hit, and the function function process is executed; Enable the slot pool and call the slot template that matches the preset intent type, fill the slot according to the content to be filled in the slot template, and generate a first type of optimized response statement that is suitable for the problem description information according to the filling result; Calling a preset corresponding function according to the intention intelligent judgment result, parsing the problem description information according to the function function, inputting the parsed information into the big model for keyword extraction, and generating a second type of optimized response statement suitable for the problem description information based on the keyword extraction result; The first type of optimized response statements and the second type of optimized response statements are displayed to the user.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.
Citation Information
Patent Citations
Artificial intelligence-based search intention recognition method, device, equipment and storage medium
CN113707300A
Futures question and answer-oriented user intention identification method and system
CN114218392A