Intention understanding tool calling method based on large model function call technology
By building a structured function interface scheduling mechanism through a large language model, the problems of inaccurate intention understanding and function calls in human-computer interaction are solved, and deep semantic analysis, structured parameter filling and closed-loop output of response content are achieved, which improves the intelligence and reliability of the interactive system.
Patent Information
- Application Number
- CN202510755055.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing human-computer interaction methods face problems such as inaccurate intent understanding, function matching errors, incomplete parameter filling, and call failures in scenarios with diverse semantic expressions, complex task scenarios, and frequent multi-round context changes. They lack deep semantic processing and closed-loop response design.
A large language model is used to build a structured function interface scheduling mechanism, and intent recognition is performed through a semantic vector model combined with a multi-level semantic fusion mechanism. Multi-factor similarity calculation and a triple scoring mechanism are introduced for function matching and parameter extraction. A standardized call request format is constructed, and closed-loop output of the response content is achieved.
It significantly improves the accuracy of intent recognition and context coherence, increases the success rate of function calls and the consistency of interactive responses, and enhances user experience and system intelligence.
Smart Images

Figure CN120633858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large-scale language model interaction technology, and in particular to a method for calling an intent understanding tool based on large-scale model function call technology. Background Art
[0002] Against the backdrop of the rapid development of artificial intelligence, natural language processing technology has been widely used in various intelligent interaction scenarios, including intelligent customer service, virtual assistants, intelligent question-and-answer systems, and automated office systems. In these applications, how to accurately understand user intent and precisely call back-end tools has become the key to improving interaction efficiency and intelligence. However, existing human-computer interaction methods mostly rely on template matching, keyword triggering, rule engines, and other methods to complete intent parsing and function docking. Such methods have certain practicality in static scenarios, but in the face of actual needs with diverse semantic expressions, complex task scenarios, and frequent multi-round context changes, they often exhibit problems such as one-sided understanding, matching errors, incomplete parameter filling, and call failures, which seriously restrict the intelligence and practicality of the interactive system.
[0003] In recent years, as the capabilities of large language models have continued to increase, they have demonstrated remarkable abilities in understanding complex natural language expressions and extracting key semantic information. Some studies have begun to explore the introduction of large models into the intent understanding and tool calling process, proposing a new approach of using text to drive function matching and instruction execution, in which the "function call" mechanism becomes a key link. The so-called function call refers to building a structured function definition in a large model so that the model can not only understand user semantics, but also map the understanding results into a specific function call format and carry the corresponding parameters to execute external call tasks. Compared with traditional methods, this approach has obvious advantages in model uniformity, versatility, and scalability.
[0004] Despite this, the current large-model tool scheduling method based on function call still has a series of technical bottlenecks. First, in the intent parsing stage, existing methods are mainly based on pure text matching or one-way prediction, lacking a deep processing mechanism that combines context, structured expression and multi-level semantic reasoning, resulting in inaccurate identification of intent types and operation objects. Secondly, in the function matching process, most studies rely on hard-coded matching rules or vector retrieval strategies, lacking the ability to integrate parameter constraints and function description semantics, thus affecting the accuracy of function call selection. In addition, in the parameter extraction link, due to the ambiguity, redundancy and uncertainty of user instruction expressions, existing methods are often unable to accurately extract structured parameter content that meets the constraints while retaining the integrity of context semantics, resulting in missing or incorrect filling of call parameters.
[0005] Furthermore, the construction of function call interface formats mostly relies on static template splicing, failing to fully integrate with intent recognition results and parameter extraction processes, making it difficult to achieve a smooth transition from semantics to call behavior. Regarding response generation, some methods directly convert model outputs into user responses, neglecting the coordinated processing of intent structure information and tool execution results. This results in ambiguous responses, disorganized structure, and a lack of specificity. Furthermore, the lack of a closed-loop response design prevents the system from fully reflecting the causal relationship between call logic and execution results through natural language feedback, impacting the user experience.
[0006] Therefore, how to provide an intent understanding tool calling method based on large-model function call technology to achieve deep semantic analysis of natural language instructions, precise matching of function descriptions, structured parameter filling, interface format standard construction and closed-loop output of response content has become an urgent problem that technical personnel in this field need to solve. Summary of the Invention
[0007] One purpose of the present invention is to propose an intent understanding tool calling method based on large model function call technology. The present invention makes full use of the semantic understanding ability of the large language model, the structured function interface scheduling mechanism and the parameter intelligent filling algorithm, and describes in detail the full-process mapping strategy from natural language instructions to function call execution. It has the advantages of high instruction understanding accuracy, high tool call success rate and complete interactive response closed loop.
[0008] According to an embodiment of the present invention, a method for calling an intent understanding tool based on a large model function call technology includes the following steps:
[0009] S1. Obtain the user's natural language instructions and preprocess them to obtain structured instruction fragments;
[0010] S2. Based on the structured instruction fragment and the contextual dialogue information, a semantic vector for intent recognition is constructed, and the semantic vector is input into the large language model to obtain an intent recognition result including the target intent type and the operation object;
[0011] S3. Based on the intent recognition result, filter out function description information that matches the target intent type from the function definition library, where the function description information includes function identifier, parameter names, and parameter restrictions;
[0012] S4. Based on the intent recognition results, under the parameter restriction constraints of the function description information, extract the parameter content from the semantic vector and construct a structured parameter value set;
[0013] S5. Construct a call request format that complies with the function call interface specification based on the function identifier and the structured parameter value set, input the call request format into the function call interface of the large language model, and execute the corresponding tool call task;
[0014] S6. Receive the function call result, and combine the operation object and target intent type in the intent recognition result to generate the corresponding natural language response text, and output it to the user interface to complete the instruction response closed loop.
[0015] Optionally, the natural language instructions include declarative requests, interrogative queries, operational commands, and multiple rounds of contextual dialogue content.
[0016] Optionally, the preprocessing includes word segmentation, part-of-speech tagging, named entity recognition and dependency parsing.
[0017] Optionally, the S2 specifically includes:
[0018] S21. Based on the structured instruction fragment, position encoding and semantic embedding are performed on each word to construct a semantic matrix M0. The semantic matrix is defined as:
[0019] M0=[e1+p1;e2+p2;…;e n +p n ];
[0020] Where n is the number of words in the structured instruction fragment, e i Represents the word vector of the i-th word, p i represents the position information embedding vector of the i-th word, […] represents the matrix formed by row splicing;
[0021] S22. Extract the context tensor C from the context dialogue information and interactively fuse it with the semantic matrix M0 to generate a semantic vector V. The semantic vector is calculated using the following formula:
[0022] V=Sigmoid(W1·Attn(M0,C)+b1)+tanh(W2·FFN(M0)+b2);
[0023] Among them, Attn(M0, C) represents the multi-head attention calculation result performed with M0 as the query and C as the key-value pair, FFN(M0) represents the feedforward neural network mapping performed on M0, W1 and W2 are linear transformation matrices respectively, b1 and b2 are bias terms, Sigmoid represents the nonlinear activation function, and tanh represents the hyperbolic tangent function;
[0024] S23. Input the semantic vector V into the large language model that has completed the pre-training process, and output the intent recognition result through maximum probability decoding. The intent recognition result is represented as a tuple (T, O), where T represents the target intent type label and O represents the instruction fragment of the operation object.
[0025] Optionally, the intent recognition result is generated based on a dual-branch decoding mechanism constructed based on semantic vectors. The first branch uses a fully connected network to perform multi-class label classification on the semantic vector to output the target intent type. The second branch uses a conditional random field to perform sequence labeling on the semantic vector to extract the operation object, and introduces a cross-attention weight adjustment factor in the output layer to realize the collaborative reasoning generation process between the intent type and the operation object.
[0026] Optionally, the S3 specifically includes:
[0027] S31, receive the intent recognition result tuple (T, O), and compare the target intent type T with the function type label set L in each function description information in the function definition library k Perform matching screening and build a matching function set F m , the matching function set consists of all matching functions that make T=L k The function number f that holds true k composition;
[0028] S32, matching function set F m Each function number f in k , read the corresponding function description information D k , the function description information is defined as a triple D k =(I k ,A k ,R k ), where I k Indicates the function identification string, A k Represents an ordered sequence of parameter names, R k Represents a set of constraints consisting of parameter restriction tensors;
[0029] S33, calculate the semantic matching score S between the semantic vector V and each function description information k , filter the maximum matching function f * , the matching score is defined as:
[0030]
[0031] Among them, φ(I k ),φ(A k,j ) respectively represent I k ,A k,j Encoded as a vector, cos(·) represents cosine similarity, Ak,j Indicates the name of the jth parameter of the kth function, |A k | represents the length of the parameter sequence, ρ(R k ,V) represents the constraint consistency scoring function between the parameter constraint condition set and the semantic vector, γ1, γ2, γ3 are normalized weighting coefficients, satisfying γ1+γ2+γ3=1;
[0032] S34, the maximum matching function f * The corresponding function description information triple D * =(I * ,A * ,R * ) as the structural basis for parameter extraction and calling.
[0033] Optionally, the S4 specifically includes:
[0034] S41, according to the function description information triple D * =(I * ,A * ,R * ) and semantic vector V, for the ordered sequence A of parameter names * =[a1,a2,…,a m Each parameter a in ] j Construct the corresponding extraction vector query pair (a j ,V), and j∈{1,2,...,m};
[0035] S42, based on parameter restriction condition set R * =[r1,r2,…,r m ], calculate any parameter a j The extraction score q in the semantic vector V j , and determine the parameter value x j , the extraction score is defined as:
[0036] q j =δ1·cos(V,φ(a j ))+δ2·λ(V,r j )+δ3·μ(φ(a j ),r j );
[0037] Among them, φ(a j ) means to change the parameter name a j Encoded as a vector, λ(V,r j ) represents the semantic vector V and the constraint tensor r j The conformity score, μ(φ(a j ),r j ) represents the parameter name vector and the constraint condition tensor rj The consistency score of , δ1, δ2, δ3 are normalized weighting factors, satisfying δ1+δ2+δ3=1;
[0038] S43. Determine all parameter values and form a structured parameter value set X as a parameter content input basis for the function call request.
[0039] Optionally, the parameter value x j Based on the semantic vector V, parameter name vector φ(a j ) and the constraint tensor r j Construct a triple matching mechanism to maximize the extraction score q j The objective function realizes parameter value positioning, where the parameter value x j A nested two-stage extraction strategy is adopted to determine . In the first stage, candidate segments are extracted from the semantic vector based on the semantic attention mechanism. In the second stage, constraint consistency screening is performed in combination with the constraint tensor. The best segment that meets the semantic similarity threshold and the constraint confidence interval is selected from the candidate segments as the final parameter value.
[0040] Optionally, the S5 specifically includes:
[0041] S51, based on function identification I * and the structured parameter value set X, and combine the two into a call field triplet P = (I * ,A * ,X), where A * =[a1,a2,…,a m ] is an ordered sequence of parameter names;
[0042] S52. According to the preset function call interface standard format, the call field triplet P is normalized and encoded to generate a standardized call request format R. The call request format is defined as:
[0043]
[0044] Among them, ω(I * ) means to use the function identifier string I * Convert to a format string that meets the prefix requirements of the calling interface, η(a j ,x j ) means to change the parameter name a j With parameter value x j Convert to a key-value pair structure string, θ(·) means converting all key-value pair structure strings into JSON format, Indicates the parameter key value splicing and aggregation operation, m represents the number of parameters;
[0045] S53, the call request format R is used as input, and is passed to the large language model function call interface, triggering the function identifier I * The corresponding tool calls the task execution process.
[0046] Optionally, the S6 specifically includes:
[0047] S61. Receive the execution result string returned by the function call interface, record it as response content Y, and input it together with the intent recognition result (T, O) as common content;
[0048] S62: Concatenate the response content Y, the target intent type T, and the operation object O in sequence into a complete semantic input string, input it into the large language model, and generate a natural language response text;
[0049] S63. Perform language integrity review and context consistency judgment on the natural language response text, filter out redundant expression content that does not meet the semantic closed loop requirements, and output the finalized response text to the user interface to complete the instruction response closed loop based on the function call.
[0050] The beneficial effects of the present invention are:
[0051] First, the present invention constructs a semantic vector model that combines structured instruction fragments with contextual semantic information, and introduces a multi-level semantic fusion mechanism, so that the large language model can more accurately identify the target intention type and operation object when processing user natural language instructions. It solves the problems of vague intention recognition and disconnected context understanding caused by the diversity of semantic expression in the existing technology, and significantly improves the accuracy of semantic understanding and context coherence.
[0052] Secondly, the present invention proposes a function description information structure centered around function identifiers, parameter names, and parameter constraints, and performs function screening and matching based on a multi-factor similarity calculation formula, thus achieving semantically driven optimization of the function selection process. During the parameter extraction process, the present invention introduces a triple scoring mechanism and a two-stage extraction strategy, effectively ensuring the logical consistency and structural matching between the extracted parameter content and the function constraints. This solves the problems of inaccurate parameter positioning and high call failure rates in traditional methods, significantly improving the reliability and adaptability of tool calls.
[0053] Finally, in the process of constructing the function call interface format and generating natural language responses, this paper proposes a standardized request format assembly method based on function identifiers and structured parameter value sets. By integrating execution results with the intent structure to generate natural language feedback, this method achieves an integrated process of user intent expression, tool call execution, and closed-loop output of response results. This process exhibits strong semantic consistency and execution transparency, significantly improving the user interaction experience and the overall intelligence level of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0055] Figure 1 This is a flowchart of the method for calling the intent understanding tool based on the large model function call technology proposed by the present invention;
[0056] Figure 2 This is a schematic diagram of semantic vector construction and intent recognition for the intent understanding tool calling method based on the large model function call technology proposed in the present invention;
[0057] Figure 3 This is the interface call request construction and tool execution flow chart of the intent understanding tool calling method based on the large model function call technology proposed by the present invention. DETAILED DESCRIPTION
[0058] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0059] refer to Figure 1-3 The calling method of the intent understanding tool based on the large model function call technology includes the following steps:
[0060] S1. Obtain the user's natural language instructions and preprocess them to obtain structured instruction fragments;
[0061] S2. Based on the structured instruction fragment and the contextual dialogue information, a semantic vector for intent recognition is constructed, and the semantic vector is input into the large language model to obtain an intent recognition result including the target intent type and the operation object;
[0062] S3. Based on the intent recognition result, filter out function description information that matches the target intent type from the function definition library, where the function description information includes function identifier, parameter names, and parameter restrictions;
[0063] S4. Based on the intent recognition results, under the parameter restriction constraints of the function description information, extract the parameter content from the semantic vector and construct a structured parameter value set;
[0064] S5. Construct a call request format that complies with the function call interface specification based on the function identifier and the structured parameter value set, input the call request format into the function call interface of the large language model, and execute the corresponding tool call task;
[0065] S6. Receive the function call result, and combine the operation object and target intent type in the intent recognition result to generate the corresponding natural language response text, and output it to the user interface to complete the instruction response closed loop.
[0066] The present invention constructs a function call process based on a large model, starting from natural language input, and completing intent analysis, function matching, parameter extraction, call execution and response output in sequence, forming a closed loop of semantic understanding and tool calling, significantly improving user command response efficiency and interactive intelligence.
[0067] In this embodiment, the natural language instructions include declarative requests, interrogative queries, operational commands, and multiple rounds of contextual dialogue content.
[0068] By supporting declarative requests, interrogative queries, operational commands, and multiple rounds of contextual content, the present invention enables the model to robustly identify intent types in multiple types of language input scenarios, thereby expanding the adaptability of semantic input and enhancing the coverage of interactive scenarios.
[0069] In this embodiment, the preprocessing includes word segmentation, part-of-speech tagging, named entity recognition and dependency parsing.
[0070] By introducing word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis in the preprocessing stage, the present invention effectively extracts structured semantic features in the language, improves the accuracy of semantic vector construction, and provides a solid foundation for subsequent intent recognition and parameter extraction.
[0071] In this embodiment, S2 specifically includes:
[0072] S21. Based on the structured instruction fragment, position encoding and semantic embedding are performed on each word to construct a semantic matrix M0. The semantic matrix is defined as:
[0073] M0=[e1+p1;e2+p2;…;e n +p n ];
[0074] Where n is the number of words in the structured instruction fragment, e i Represents the word vector of the i-th word, p i represents the position information embedding vector of the i-th word, […] represents the matrix formed by row splicing;
[0075] S22. Extract the context tensor C from the context dialogue information and interactively fuse it with the semantic matrix M0 to generate a semantic vector V. The semantic vector is calculated using the following formula:
[0076] V=Sigmoid(W1·Attn(M0,C)+b1)+tanh(W2·FFN(M0)+b2);
[0077] Among them, Attn(M0, C) represents the multi-head attention calculation result performed with M0 as the query and C as the key-value pair, FFN(M0) represents the feedforward neural network mapping performed on M0, W1 and W2 are linear transformation matrices respectively, b1 and b2 are bias terms, Sigmoid represents the nonlinear activation function, and tanh represents the hyperbolic tangent function;
[0078] S23. Input the semantic vector V into the large language model that has completed the pre-training process, and output the intent recognition result through maximum probability decoding. The intent recognition result is represented as a tuple (T, O), where T represents the target intent type label and O represents the instruction fragment of the operation object.
[0079] The present invention constructs a fusion mechanism of semantic matrix and context tensor, adopts multi-head attention and feedforward neural network to jointly generate semantic vectors, and combines it with a large model decoding structure, which significantly improves the model's ability to capture multi-level semantic structures and achieves more accurate intent labeling and operation object recognition.
[0080] In this embodiment, the intention recognition result is generated based on a dual-branch decoding mechanism constructed based on semantic vectors. The first branch uses a fully connected network to perform multi-class label classification on the semantic vector to output the target intent type. The second branch uses a conditional random field to perform sequence labeling on the semantic vector to extract the operation object, and introduces a cross-attention weight adjustment factor in the output layer to realize the collaborative reasoning generation process between the intent type and the operation object.
[0081] The present invention adopts a dual-branch decoding structure to decouple the intent type from the operation object, and introduces a cross-attention mechanism to enhance the collaborative reasoning ability between the two, so that the recognition results of the intent structure have higher stability and consistency, effectively improving the problem of low recognition accuracy.
[0082] In this embodiment, S3 specifically includes:
[0083] S31, receive the intent recognition result tuple (T, O), and compare the target intent type T with the function type label set L in each function description information in the function definition library k Perform matching screening and build a matching function set F m , the matching function set consists of all matching functions that make T=L kThe function number f that holds true k composition;
[0084] S32, matching function set F m Each function number f in k , read the corresponding function description information D k , the function description information is defined as a triple D k =(I k ,A k ,R k ), where I k Indicates the function identification string, A k Represents an ordered sequence of parameter names, R k Represents a set of constraints consisting of parameter restriction tensors;
[0085] S33, calculate the semantic matching score S between the semantic vector V and each function description information k , filter the maximum matching function f * , the matching score is defined as:
[0086]
[0087] Among them, φ(I k ),φ(A k,j ) respectively represent I k ,A k,j Encoded as a vector, cos(·) represents cosine similarity, A k,j Indicates the name of the jth parameter of the kth function, |A k | represents the length of the parameter sequence, ρ(R k ,V) represents the constraint consistency scoring function between the parameter constraint condition set and the semantic vector, γ1, γ2, γ3 are normalized weighting coefficients, satisfying γ1+γ2+γ3=1;
[0088] S34, the maximum matching function f * The corresponding function description information triple D * =(I * ,A * ,R * ) as the structural basis for parameter extraction and calling.
[0089] The present invention realizes the accurate screening of function identifiers by constructing a function description information triple and a matching score calculation formula based on multi-factor weighting, improves the semantic fit of function scheduling, and effectively solves the problems of large function selection deviation and frequent miscalls in traditional methods.
[0090] In this embodiment, the S4 specifically includes:
[0091] S41, according to the function description information triple D * =(I * ,A * ,R * ) and semantic vector V, for the ordered sequence A of parameter names * =[a1,a2,…,a m Each parameter a in ] j Construct the corresponding extraction vector query pair (a j ,V), and j∈{1,2,...,m};
[0092] S42, based on parameter restriction condition set R * =[r1,r2,…,r m ], calculate any parameter a j The extraction score q in the semantic vector V j , and determine the parameter value x j , the extraction score is defined as:
[0093] q j =δ1·cos(V,φ(a j ))+δ2·λ(V,r j )+δ3·μ(φ(a j ),r j );
[0094] Among them, φ(a j ) means to change the parameter name a j Encoded as a vector, λ(V,r j ) represents the semantic vector V and the constraint tensor r j The conformity score, μ(φ(a j ),r j ) represents the parameter name vector and the constraint condition tensor r j The consistency score of , δ1, δ2, δ3 are normalized weighting factors, satisfying δ1+δ2+δ3=1;
[0095] S43. Determine all parameter values and form a structured parameter value set X as a parameter content input basis for the function call request.
[0096] The present invention constructs a parameter extraction scoring function based on a triple scoring mechanism, accurately extracts structured parameter values through consistency scoring with semantic vectors and constraint conditions, and improves the accuracy of the parameter parsing process and the stability of the call input.
[0097] In this embodiment, the parameter value x j Based on the semantic vector V, parameter name vector φ(a j ) and the constraint tensor r jConstruct a triple matching mechanism to maximize the extraction score q j The objective function realizes parameter value positioning, where the parameter value x j A nested two-stage extraction strategy is adopted to determine . In the first stage, candidate segments are extracted from the semantic vector based on the semantic attention mechanism. In the second stage, constraint consistency screening is performed in combination with the constraint tensor. The best segment that meets the semantic similarity threshold and the constraint confidence interval is selected from the candidate segments as the final parameter value.
[0098] The present invention further introduces a nested two-stage mechanism in the parameter extraction stage, combines the attention mechanism with the constraint consistency screening strategy, and can achieve high-confidence parameter positioning under complex language expressions, thereby improving the call success rate and diversity semantic adaptability.
[0099] In this embodiment, the S5 specifically includes:
[0100] S51, based on function identification I * and the structured parameter value set X, and combine the two into a call field triplet P = (I * ,A * ,X), where A * =[a1,a2,…,a m ] is an ordered sequence of parameter names;
[0101] S52. According to the preset function call interface standard format, the call field triplet P is normalized and encoded to generate a standardized call request format R. The call request format is defined as:
[0102]
[0103] Among them, ω(I * ) means to use the function identifier string I * Convert to a format string that meets the prefix requirements of the calling interface, η(a j ,x j ) means to change the parameter name a j With parameter value x j Convert to a key-value pair structure string, θ(·) means converting all key-value pair structure strings into JSON format, Indicates the parameter key value splicing and aggregation operation, m represents the number of parameters;
[0104] S53, the call request format R is used as input, and is passed to the large language model function call interface, triggering the function identifier I * The corresponding tool calls the task execution process.
[0105] The present invention proposes a structured function call format construction process, which strictly combines the function identifier, parameter name and value encoding into JSON format to ensure the consistency and executableness of the call structure and improve the stability and versatility of the call system.
[0106] In this embodiment, S6 specifically includes:
[0107] S61. Receive the execution result string returned by the function call interface, record it as response content Y, and input it together with the intent recognition result (T, O) as common content;
[0108] S62: Concatenate the response content Y, the target intent type T, and the operation object O in sequence into a complete semantic input string, input it into the large language model, and generate a natural language response text;
[0109] S63. Perform language integrity review and context consistency judgment on the natural language response text, filter out redundant expression content that does not meet the semantic closed loop requirements, and output the finalized response text to the user interface to complete the instruction response closed loop based on the function call.
[0110] The present invention collaboratively generates natural language feedback through function call results and intent structures, combines integrity and context consistency verification, outputs accurate and clear user response information, realizes a closed loop of the interaction process, and enhances the system's interpretability and user experience.
[0111] Example 1:
[0112] In order to verify the feasibility of the present invention in implementation, the present invention is applied to a natural language intelligent office assistant system, which is used to assist users in calling a variety of back-end automation tools through natural language instructions, including functions such as schedule management, email sending, data statistics and document generation. In traditional voice assistants or dialogue systems, the instructions expressed by users often have semantic ambiguity, unclear context, missing parameters or non-standard expressions, resulting in inaccurate function calls or high call failure rates. The present invention introduces an intent understanding tool calling method based on large model function call technology, integrates and verifies it in this system, aiming to improve the natural language instruction comprehension ability, function call accuracy and overall interactive experience.
[0113] In actual application scenarios, users may issue commands such as "Help me schedule a meeting with Zhang Wei next Monday morning." Traditional systems are prone to problems such as "time parsing errors" and "contact identification failures," ultimately failing to correctly call the schedule creation interface. After introducing the method of the present invention, the system first structures the natural language instruction, extracts key fragments through lexical and syntactic analysis, generates a semantic matrix and context tensor, and further constructs a semantic vector. With the support of the semantic vector, the large language model successfully identifies the intent type as "create schedule" and the operation object as "meeting." Combined with the function description information registered in the function definition library, it accurately matches the "createScheduleEvent" function through a multi-factor matching scoring mechanism. Subsequently, the parameter extraction module successfully extracts the two structured parameter contents of "time = next Monday morning" and "object = Zhang Wei" based on the function parameter constraints, and constructs a standardized call format to pass to the large model's function call interface. After the system successfully calls, it returns the execution result "Schedule added". The system combines this information with the intent result and outputs a natural language response: "A meeting has been scheduled for you and Zhang Wei next Monday morning."
[0114] Through actual measurements in the office assistant system, the system received and processed a total of 1,120 natural language instructions, covering task types including sending emails, setting reminders, adding schedules, generating reports, organizing meeting minutes, etc. Among them, 960 instructions were non-template inputs, that is, they were not expressed according to a fixed structure and contained a large number of natural language variants. Comparing the execution results of the traditional keyword trigger + rule function call scheme with the method of the present invention, it was found that the latter had significant improvements in intent recognition accuracy, function call success rate, parameter parsing completeness and user satisfaction. The specific performance is as follows: the intent recognition accuracy increased from 86.3% to 97.9%, the function call success rate increased from 82.7% to 96.8%, and the parameter parsing completeness increased from 78.4% to 95.2%. At the same time, after introducing the method of the present invention, the user satisfaction score for the interactive system increased from 4.1 to 4.8 (out of 5 points).
[0115] Furthermore, an analysis of execution time revealed that despite the introduction of extensive semantic modeling and multi-round structured processing, the large-scale model inference mechanism based on quantitative optimization improved the average response time from 1.97 seconds to 1.62 seconds, demonstrating enhanced real-time performance. In a multi-round dialogue support test, the proposed method maintained an accurate call rate of 92.5% across five consecutive rounds of dialogue, while the accuracy of traditional methods dropped to 72.8%, demonstrating strong context retention and reasoning capabilities.
[0116] The following table summarizes the comparative data between the present invention and the traditional method in real systems:
[0117] Table 1 Performance comparison between the semantic understanding tool calling method based on function call and the traditional method
[0118] Indicator Category Traditional method data value Data value of the method of the present invention Performance improvement rate Intent recognition accuracy 86.3% 97.9% +13.4% Function call success rate 82.7% 96.8% +17.1% Parameter parsing completeness rate 78.4% 95.2% +16.8% User satisfaction rating 4.1 / 5 4.8 / 5 +0.7 Average response time 1.97 seconds 1.62 seconds -17.8% Multi-round dialogue call accuracy 72.8% 92.5% +19.7% Number of rollback calls with errors 146 times 31 times -78.7% Unable to parse input ratio 9.2% 1.4% -84.8% Parameter value error rate 12.5% 3.7% -70.4%
[0119] The above data fully demonstrate that the method proposed in the present invention is highly feasible and has practical application value in the automated tool calling scenario driven by natural language understanding, and is particularly suitable for intelligent interactive systems with complex semantics, multi-tool access, and multi-round context concurrency.
[0120] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. The intention understanding tool calling method based on the large model function call technology is characterized by: The steps include: S1. Obtain the user's natural language instructions and preprocess them to obtain structured instruction fragments; S2. Based on the structured instruction fragment and the contextual dialogue information, a semantic vector for intent recognition is constructed, and the semantic vector is input into the large language model to obtain an intent recognition result including the target intent type and the operation object; S3. Based on the intent recognition result, filter out function description information that matches the target intent type from the function definition library, where the function description information includes function identifier, parameter names, and parameter restrictions; S4. Based on the intent recognition results, under the parameter restriction constraints of the function description information, extract the parameter content from the semantic vector and construct a structured parameter value set; S5. Construct a call request format that complies with the function call interface specification based on the function identifier and the structured parameter value set, input the call request format into the function call interface of the large language model, and execute the corresponding tool call task; S6. Receive the function call result, and combine the operation object and target intent type in the intent recognition result to generate the corresponding natural language response text, and output it to the user interface to complete the instruction response closed loop.
2. The method for calling an intention understanding tool based on a large model function call technology according to claim 1 is characterized in that: The natural language instructions include declarative requests, interrogative queries, operational commands, and multi-round contextual dialogue content.
3. The method for calling an intention understanding tool based on a large model function call technology according to claim 1, characterized in that: The preprocessing includes word segmentation, part-of-speech tagging, named entity recognition and dependency parsing.
4. The method for calling an intent understanding tool based on a large model function call technology according to claim 1, characterized in that: The S2 specifically includes: S21. Based on the structured instruction fragment, position encoding and semantic embedding are performed on each word to construct a semantic matrix M0. The semantic matrix is defined as: M0=[e1+p1;e2+p2;...;e n +p n ]: Where n is the number of words in the structured instruction fragment, e i Represents the word vector of the i-th word, p i represents the position information embedding vector of the i-th word, […] represents the matrix formed by row splicing; S22. Extract the context tensor C from the context dialogue information and interactively fuse it with the semantic matrix M0 to generate a semantic vector V. The semantic vector is calculated using the following formula: V=Sigmoid(W1·Attn(M0,C)+b1)+tanh(W2·FFN(M0)+b2); Among them, Attn(M0, C) represents the multi-head attention calculation result performed with M0 as the query and C as the key-value pair, FFN(M0) represents the feedforward neural network mapping performed on M0, W1 and W2 are linear transformation matrices respectively, b1 and b2 are bias terms, Sigmoid represents the nonlinear activation function, and tanh represents the hyperbolic tangent function; S23. Input the semantic vector V into the large language model that has completed the pre-training process, and output the intent recognition result through maximum probability decoding. The intent recognition result is represented as a tuple (T, O), where T represents the target intent type label and O represents the instruction fragment of the operation object.
5. The method for calling an intention understanding tool based on a large model function call technology according to claim 4 is characterized in that: The intent recognition result is generated based on a dual-branch decoding mechanism constructed based on semantic vectors. The first branch uses a fully connected network to perform multi-class label classification on the semantic vectors to output the target intent type. The second branch uses a conditional random field to perform sequence labeling on the semantic vectors to extract the operation objects, and introduces a cross-attention weight adjustment factor in the output layer to realize the collaborative reasoning generation process between the intent type and the operation object.
6. The method for calling an intention understanding tool based on a large model function call technology according to claim 1, characterized in that: The S3 specifically includes: S31, receive the intent recognition result tuple (T, O), and compare the target intent type T with the function type label set L in each function description information in the function definition library k Perform matching screening and build a matching function set F m , the matching function set consists of all matching functions that make T=L k The function number f that holds true k composition; S32, matching function set F m Each function number f in k , read the corresponding function description information D k , the function description information is defined as a triple D k =(I k ,A k ,R k ), where I k Indicates the function identification string, A k Represents an ordered sequence of parameter names, R k Represents a set of constraints consisting of parameter restriction tensors; S33, calculate the semantic matching score S between the semantic vector V and each function description information k , filter the maximum matching function f * , the matching score is defined as: Among them, φ(I k ),φ(A k,j ) respectively represent I k ,A k,j Encoded as a vector, cos(·) represents cosine similarity, A k,j Indicates the name of the jth parameter of the kth function, |A k | represents the length of the parameter sequence, ρ(R k ,V) represents the constraint consistency scoring function between the parameter constraint condition set and the semantic vector, γ1, γ2, γ3 are normalized weighting coefficients, satisfying γ1+γ2+γ3=1; S34, the maximum matching function f * The corresponding function description information triple D * =(I * ,A * ,R * ) as the structural basis for parameter extraction and calling.
7. The method for calling an intention understanding tool based on a large model function call technology according to claim 1, characterized in that: The S4 specifically includes: S41, according to the function description information triple D * =(I * ,A * ,R * ) and semantic vector V, for the ordered sequence A of parameter names * =[a1,a2,…,a m Each parameter a in ] j Construct the corresponding extraction vector query pair (a j ,V), and j∈{1,2,...,m}; S42, based on parameter restriction condition set R * =[r1,r2,…,r m ], calculate any parameter a j The extraction score q in the semantic vector V j , and determine the parameter value x j , the extraction score is defined as: q j =δ1·cos(V,φ(a j ))+δ2·λ(V,r j )+δ3·μ(φ(a j ),r j ); Among them, φ(a j ) means to change the parameter name a j Encoded as a vector, λ(V,r j ) represents the semantic vector V and the constraint tensor r j The conformity score, μ(φ(a j ),r j ) represents the parameter name vector and the constraint condition tensor r j The consistency score of , δ1, δ2, δ3 are normalized weighting factors, satisfying δ1+δ2+δ3=1; S43. Determine all parameter values and form a structured parameter value set X as a parameter content input basis for the function call request.
8. The method for calling an intention understanding tool based on a large model function call technology according to claim 7, characterized in that: The parameter value x j Based on the semantic vector V, parameter name vector φ(a j ) and the constraint tensor r j Construct a triple matching mechanism to maximize the extraction score q j The objective function realizes parameter value positioning, where the parameter value x j A nested two-stage extraction strategy is adopted to determine . In the first stage, candidate segments are extracted from the semantic vector based on the semantic attention mechanism. In the second stage, constraint consistency screening is performed in combination with the constraint tensor. The best segment that meets the semantic similarity threshold and the constraint confidence interval is selected from the candidate segments as the final parameter value.
9. The method for calling an intention understanding tool based on a large model function call technology according to claim 1, characterized in that: The S5 specifically includes: S51, based on function identification I * and the structured parameter value set X, and combine the two into a call field triplet P = (I * ,A * ,X), where A * =[a1,a2,…,a m ] is an ordered sequence of parameter names; S52. According to the preset function call interface standard format, the call field triplet P is normalized and encoded to generate a standardized call request format R. The call request format is defined as: Among them, ω(I * ) means to use the function identifier string I * Convert to a format string that meets the prefix requirements of the calling interface, η(a j ,x j ) means to change the parameter name a j With parameter value x j Convert to a key-value pair structure string, θ(·) means converting all key-value pair structure strings into JSON format, Indicates the parameter key value splicing and aggregation operation, m represents the number of parameters; S53, the call request format R is used as input, and is passed to the large language model function call interface, triggering the function identifier I * The corresponding tool calls the task execution process.
10. The method for calling an intention understanding tool based on a large model function call technology according to claim 1, characterized in that: The S6 specifically includes: S61. Receive the execution result string returned by the function call interface, record it as response content Y, and input it together with the intent recognition result (T, O) as common content; S62: Concatenate the response content Y, the target intent type T, and the operation object O in sequence into a complete semantic input string, input it into the large language model, and generate a natural language response text; S63. Perform language integrity review and context consistency judgment on the natural language response text, filter out redundant expression content that does not meet the semantic closed loop requirements, and output the finalized response text to the user interface to complete the function call-based instruction response closed loop.
Citation Information
Cited By
Intelligent data report generation method and system based on MCP protocol
CN121052228A
Intelligent self-adaptive RAG enhancement mechanism question answering method and system
CN121412354A
Text abstract generation method for deep semantic understanding
CN121579686A
Method for controlling NAS (Network Attached Storage) equipment based on large model and NAS equipment
CN121681489A
An Embedded Processor Natural Language Debugging System and Method Based on JTAG Real-Time Data Exchange Channel
CN122414132A