Text processing method, device, equipment and storage medium based on large language model
By classifying the input text and adapting the processing objects, the problem of insufficient stability and maintainability of large language models in text processing is solved, and low-cost and efficient text processing capabilities are achieved to meet the needs of complex and new input texts.
Patent Information
- Application Number
- CN202411253061.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-09-06
AI Technical Summary
Existing large language models suffer from insufficient stability, poor maintainability, and insufficient low-cost scalability during text processing. Especially when faced with complex input text and new text types, they are unable to meet the needs of enterprise distributed business systems.
A text processing method based on a large language model is adopted. By performing the first and second classifications on the input text, the processing object and text category are determined, and the adapted target tool or intelligent agent is called for processing to form a logical chain or interactive topology diagram to ensure the accuracy and reliability of text processing.
It achieves stable, reliable and low-cost processing of input text under different text categories, improves the accuracy and efficiency of text processing, adapts to changes in enterprise business systems, and supports flexible processing of multiple text types.
Smart Images

Figure CN119415646B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of AI (Artificial Intelligence), specifically to technical fields such as NLP (Natural Language Processing), LLM (Large Language Model), and deep learning. It can be used in application fields such as generative search, intelligent document editing, intelligent assistants, and intelligent e-commerce, and especially to text processing methods, devices, electronic devices, and storage media based on large language models. Background Art
[0002] With the rapid development of NLP and AI technologies, computers have become able to understand and generate natural language text, making it possible to automatically process user input text (such as query text and user questions). For example, automated text processing technology can be used to automatically process input text. Currently, automated text processing technology is widely used in search engines, intelligent customer service, chatbots, and other fields, enabling rapid response and processing of user input text. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for text processing based on a large language model.
[0004] According to one aspect of the present disclosure, a text processing method based on a large language model is provided, comprising:
[0005] Performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text;
[0006] Performing a second classification on the input text to obtain a text category;
[0007] Based on the text category, calling the processing object to process the input text to obtain a response text;
[0008] The input text is replied based on the response text.
[0009] According to another aspect of the present disclosure, a text processing apparatus based on a large language model is provided, comprising:
[0010] A first classification module, configured to perform a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text;
[0011] A second classification module is used to perform a second classification on the input text to obtain a text category;
[0012] A calling module, configured to call the processing object to process the input text based on the text category to obtain a response text;
[0013] A reply module is used to reply to the input text based on the response text.
[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text processing method based on the large language model proposed in the above aspect of the present disclosure.
[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium of computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the text processing method based on the large language model proposed in the above aspect of the present disclosure.
[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the text processing method based on the large language model proposed in the above aspect of the present disclosure.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0022] Figure 1 A flowchart of a text processing method based on a large language model provided in the first embodiment of the present disclosure;
[0023] Figure 2 A flowchart of a text processing method based on a large language model provided in the second embodiment of the present disclosure;
[0024] Figure 3 A schematic diagram of an assembly method or combination method of the atomic tool provided in an embodiment of the present disclosure;
[0025] Figure 4A flowchart of a text processing method based on a large language model provided in the third embodiment of the present disclosure;
[0026] Figure 5 A flowchart of a text processing method based on a large language model provided in the fourth embodiment of the present disclosure;
[0027] Figure 6 A flowchart of a text processing method based on a large language model provided in the fifth embodiment of the present disclosure;
[0028] Figure 7 A flowchart of a text processing method based on a large language model provided in Example 6 of the present disclosure;
[0029] Figure 8 A flowchart of a text processing method based on a large language model provided in Example 7 of the present disclosure;
[0030] Figure 9 A flowchart of a text processing method based on a large language model provided in the eighth embodiment of the present disclosure;
[0031] Figure 10 A flowchart of a text processing method based on a large language model provided in Example 9 of the present disclosure;
[0032] Figure 11 A schematic diagram of the structure of the low-code tool platform provided in an embodiment of the present disclosure;
[0033] Figure 12 A schematic diagram of the interaction process between a tool-type intelligent agent and a tool provided in an embodiment of the present disclosure;
[0034] Figure 13 A schematic diagram of the interaction relationship between intelligent agents provided in an embodiment of the present disclosure;
[0035] Figure 14 A schematic diagram of the interaction relationship between atomic tools provided in an embodiment of the present disclosure;
[0036] Figure 15 This is a structural diagram of a text processing device based on a large language model provided in the tenth embodiment of the present disclosure;
[0037] Figure 16 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0038] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0039] Large models are currently booming, and large language models (LLMs) in particular have demonstrated powerful capabilities. In the field of text processing, the industry also has many intelligent processing solutions that combine LLMs, but none of these solutions are simultaneously stable, maintainable, and low-cost.
[0040] Stability: During text processing, every step must produce stable and reliable analysis results. However, a typical drawback of LLM is the existence of "illusions," which poses a challenge to the stability of analysis results.
[0041] Maintainability: When processing input text, different business parties (i.e., business teams) across an enterprise's distributed business systems participate. These parties possess the most core expertise in their respective businesses. If these expertise cannot be personally accumulated by these parties, then even the most well-implemented intelligent processing system will be unmaintainable. This means that if the business system changes, the knowledge or experience in the intelligent processing system cannot be updated promptly by the business parties, rendering the intelligent processing system ineffective.
[0042] Low-cost: When new text types are introduced, only new tools or new agents need to be built at a low cost, without making any changes to existing tools or agents. In other words, new text types must make the intelligent processing system more robust without affecting the processing capabilities of existing text types.
[0043] In related technologies, the following solutions are mainly used to process input text:
[0044] The first one uses an intelligent agent platform to process input text.
[0045] However, the agent platform is responded by a single agent, and the input text is only distributed to a single agent, which cannot process complex input text.
[0046] The second method is to use multi-agent intelligent operation and maintenance to process input text.
[0047] Although this approach uses multi-agent interaction to process input text, the logic for handling existing input text may need to be modified when encountering new input forms, which means it lacks maintainability. Furthermore, it fails to address how to enable friendly participation from different business parties, failing to demonstrate business friendliness.
[0048] The third method is to use multi-scenario intelligent operation and maintenance to realize the interactive mode between intelligent agents to process input text.
[0049] However, this approach uses function_call when invoking the tool and employs a "reflection" mechanism to address instability caused by model illusions. Furthermore, it fails to address how to enable friendly participation from different business parties, failing to demonstrate business friendliness.
[0050] Therefore, in response to at least one of the above-mentioned problems, the present disclosure proposes a text processing method, apparatus, device and storage medium based on a large language model to simultaneously meet the above-mentioned requirements of "stability", "maintainability" and "low cost" when processing various types of input texts.
[0051] The following describes the text processing method, apparatus, device, and storage medium based on a large language model according to an embodiment of the present disclosure with reference to the accompanying drawings. Before describing the embodiment of the present disclosure in detail, for ease of understanding, the following common technical terms are first introduced:
[0052] LLM is a type of natural language processing model based on deep learning. Its main features are huge model parameters and complex neural network structure, strong language understanding, context perception and language generation capabilities. It can automatically learn useful feature representations from input data and generate relevant text.
[0053] Figure 1 This is a flowchart of the text processing method based on the large language model provided in the first embodiment of the present disclosure.
[0054] The embodiment of the present disclosure uses the example of the text processing method based on a large language model being configured in a text processing device based on a large language model. The text processing device can be applied to any electronic device so that the electronic device can perform text processing functions.
[0055] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a mobile phone, tablet computer, personal digital assistant, wearable device, etc., which are hardware devices with various operating systems, touch screens and / or display screens.
[0056] like Figure 1As shown, the text processing method based on the large language model may include the following steps S101 to S104:
[0057] Step S101 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0058] The input text is input or provided by the user; the input method of the input text may include but is not limited to: touch input (such as sliding, clicking, etc.), keyboard input, voice input, etc.
[0059] The processing objects for processing the input text include but are not limited to: tools, agents, etc. The tools include but are not limited to: script commands, APIs (Application Programming Interfaces), functions, etc.
[0060] In the embodiment of the present disclosure, a text classification technology may be used to perform a first classification on the input text to obtain a classification result, wherein the classification result is used to indicate a processing object of the input text.
[0061] Step S102: Perform a second classification on the input text to obtain a text category.
[0062] The text category will vary depending on the specific application scenario, business requirements, and characteristics of the text content of the input text.
[0063] For example, when the input text is a question asked by the user, the text categories include but are not limited to: clear questions, vague questions, complex questions, common questions, emotional questions, etc.
[0064] Exemplarily, when the input text is a query text, the text categories include but are not limited to: functional query (eg, querying resource usage of a distributed business system), detailed query, summary query, fuzzy query, precise query, etc.
[0065] For example, in the case where the input text is a fault problem entered in response to a fault in a distributed business system, the text categories include but are not limited to: query loss problems (such as, for the user's input query, there are no search results, irrelevant search results, duplicate results (that is, there are a large number of duplicate entries in the search results)), long-tail problems, Core problems (core system problems), etc.
[0066] In the embodiment of the present disclosure, the input text may be subjected to a second classification based on a text classification technology to obtain the text category to which the input text belongs.
[0067] Step S103: Based on the text category, the processing object is called to process the input text to obtain a response text.
[0068] In the embodiment of the present disclosure, based on the text category, the processing object indicated by the classification result can be called to perform targeted processing on the input text to obtain a response text.
[0069] For example, when the input text is a query text, the response text may be a specific query result; when the input text is a user question, the response text may be an answer, solution steps or solution; when the input text is a fault problem, the response text may be the root cause of the fault.
[0070] Step S104: reply to the input text based on the response text.
[0071] In the embodiment of the present disclosure, a timely reply can be made to the input text based on the response text, so that the user can obtain the response text in time, thereby reducing the user's waiting time.
[0072] The text processing method based on a large language model in the embodiment of the present disclosure can implement targeted processing of different input texts by calling a processing object adapted to the input text based on the text category to which the input text belongs, thereby improving the accuracy and rationality of text processing and improving the user experience.
[0073] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0074] In order to clearly illustrate how any embodiment of the present disclosure processes the input text based on the text category and calls the processing object indicated by the classification result to obtain the response text, the present disclosure also proposes a text processing method based on a large language model.
[0075] Figure 2 This is a flowchart of the text processing method based on the large language model provided in the second embodiment of the present disclosure.
[0076] like Figure 2 As shown, the text processing method based on the large language model may include the following steps S201 to S205:
[0077] Step S201 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0078] Step S202: Perform a second classification on the input text to obtain a text category.
[0079] The explanation of steps S201 to S202 can be found in the relevant description of any embodiment of the present disclosure and will not be repeated here.
[0080] Step S203 : in response to the processing object being a tool, obtaining a target tool adapted to the text category; wherein the target tool is assembled from at least one atomic tool.
[0081] It should be noted that for different text categories, the adapted target tools may be different to improve the pertinence and accuracy of different input text processing.
[0082] In an embodiment of the present disclosure, when the processing object of the input text is a tool, a target tool that is compatible with the text category can be obtained, where the target tool is assembled from at least one atomic tool. For example, the target tool can be obtained by assembling at least one atomic tool that is compatible with the text category by dragging and dropping.
[0083] In any embodiment of the present disclosure, the target tool adapted to the text category can be obtained by following steps A to D:
[0084] Step A: In response to the processing object being a tool, query whether there is an assembly tool associated with the text category to which the input text belongs. If so, execute step B; if not, execute steps C to D.
[0085] The assembly tool is obtained by assembling at least one atomic tool associated with a text category in a tool library within a historical period.
[0086] Step B: Use the assembly tool as the target tool.
[0087] In an embodiment of the present disclosure, when there is an assembly tool associated with the text category to which the input text belongs, the assembly tool can be directly used as a target tool adapted to the text category.
[0088] That is to say, the text category to which the input text belongs is an existing known text category, and the execution subject of the present disclosure has processed other texts belonging to the same text category as the input text within a historical period. When processing the other texts, the execution subject obtains at least one atomic tool associated with the same text category from the tool library, assembles the at least one atomic tool to obtain an assembly tool, and uses the assembly tool to process the other texts.
[0089] Moreover, in order to improve the processing efficiency of subsequent texts, the executing entity may also establish an association relationship between the same text category and the assembly tool. Thus, in the present disclosure, in order to improve the processing efficiency of the input text, the above association relationship may be queried to determine the assembly tool associated with the text category to which the input text belongs, and the assembly tool may be used as a target tool adapted to the text category to which the input text belongs.
[0090] Step C: Obtain an atomic tool adapted to the text category to which the input text belongs; wherein the adapted atomic tool includes: an atomic tool created based on the input text in response to the first configuration instruction, and / or an atomic tool in a tool library.
[0091] As an example, when there is no assembly tool associated with the text category to which the input text belongs, relevant personnel (such as administrators, business personnel, etc.) can specify an atomic tool that is adapted to the text category; wherein the atomic tool that is adapted to the text category includes: manually created atomic tools (i.e., atomic tools created based on the input text in response to a first configuration instruction triggered by relevant personnel), and / or atomic tools created in the tool library.
[0092] Step D: Assemble the atomic tools adapted to the text category to obtain the target tool.
[0093] As a possible implementation method, in response to configuration instructions triggered by relevant personnel, atomic tools adapted to the text category can be assembled to obtain the target tool.
[0094] As an example, in actual application scenarios, there are many types of text categories. In order to support these scenarios, the assembly or combination of tools can include: Figure 3 (a) Figure 3 (b) and Figure 3 (c) shows three types, among which, Figure 3 (a) refers to the scenario of serial assembly of multiple atomic tools, Figure 3 (b) refers to the scenario of selecting which atomic tool to trigger based on the conditions. Figure 3 (c) refers to the scenario where multiple atomic tools are assembled in parallel. In the present disclosure, relevant personnel can complete the assembly of atomic tools by "dragging and dropping" in the UI (User Interface).
[0095] The "Data Whiteboard" can be considered a central data storage or message queue for transferring data between different atomic tools. The Data Whiteboard acts as a neutral, accessible data exchange center, allowing any atomic tool to read or write data on it.
[0096] In summary, when the text category to which the input text belongs is a known or existing text category, directly based on the association relationship between the text category and the assembly tool, the target tool associated with or adapted to the text category to which the input text belongs is determined from the assembly tool. This can not only achieve rapid processing of input texts of known text categories based on existing business experience, thereby improving the processing efficiency of input texts, but also achieve tool reuse without the need to create duplicate tools, thereby saving manpower and time costs. In the case where the text category to which the input text belongs is a newly added text category, the target tool adapted to the newly added text category is configured through configuration, so that the input text of the newly added text category can be processed based on the target tool, which can improve the accuracy and effectiveness of text processing.
[0097] Step S204: calling the target tool to process the input text to obtain a response text.
[0098] In the embodiment of the present disclosure, a target tool may be called to process the input text to obtain a response text.
[0099] Step S205: reply to the input text based on the response text.
[0100] For explanation of step S205 , please refer to the relevant description in any embodiment of the present disclosure and will not be repeated here.
[0101] The text processing method based on a large language model in an embodiment of the present disclosure obtains a target tool adapted to the text category and calls the target tool to process the input text when the classification result indicates that the processing object of the input text is a tool, which can improve the accuracy and effectiveness of text processing.
[0102] In order to clearly illustrate how to call a target tool to process input text in any embodiment of the present disclosure, the present disclosure also proposes a text processing method based on a large language model.
[0103] Figure 4 This is a flowchart of the text processing method based on the large language model provided in the third embodiment of the present disclosure.
[0104] like Figure 4 As shown, the text processing method based on the large language model may include the following steps S401 to S408:
[0105] Step S401 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0106] Step S402: Perform a second classification on the input text to obtain a text category.
[0107] Step S403 : in response to the processing object being a tool, obtaining a plurality of target tools adapted to the text category.
[0108] There are multiple target tools that are compatible with the text category, and each target tool is assembled from at least one atomic tool that is compatible with the text category.
[0109] The explanation of steps S401 to S403 can be found in the relevant description of any embodiment of the present disclosure and will not be repeated here.
[0110] Step S404: determining a calling sequence among the multiple target tools, and calling the multiple target tools in sequence based on the calling sequence.
[0111] The order of calling multiple target tools adapted to the text category to which the input text belongs may be pre-specified, or may be manually specified by relevant personnel, and the embodiments of the present disclosure do not impose any limitation on this.
[0112] In the embodiment of the present disclosure, multiple target tools may be called sequentially based on the calling order among the multiple target tools.
[0113] Step S405 : In response to the currently called target tool being the first called target tool, the first target tool is called to process the input text to obtain a processing result of the first target tool.
[0114] In an embodiment of the present disclosure, when the currently called target tool is the first called target tool, the first target tool may be called to process the input text to obtain a processing result of the first target tool.
[0115] Step S406 , in response to the currently called target tool being a non-first called target tool, the non-first target tool is called, and the processing result of the previously called target tool is processed to obtain the processing result of the non-first target tool.
[0116] In an embodiment of the present disclosure, when the currently called target tool is not the first called target tool, the non-first target tool can be called to process the processing result of the previously called target tool to obtain the processing result of the non-first target tool.
[0117] Step S407: Determine the response text according to the processing result of the last called target tool.
[0118] In the embodiment of the present disclosure, the response text for replying to the input text may be determined according to the processing result of the last called target tool.
[0119] Step S408: reply to the input text based on the response text.
[0120] For explanation of step S408, please refer to the relevant description in any embodiment of the present disclosure, which will not be repeated here.
[0121] The text processing method based on a large language model of the embodiment of the present disclosure calls multiple target tools in sequence to collaboratively process the input text based on the calling order between multiple target tools. This can ensure that the processing of each target tool is based on the processing result of the previous target tool, forming a complete logical chain. This sequentiality ensures the consistency and accuracy of text processing.
[0122] In order to clearly illustrate how any embodiment of the present disclosure processes the input text based on the text category and calls the processing object indicated by the classification result to obtain the response text, the present disclosure also proposes a text processing method based on a large language model.
[0123] Figure 5 This is a flowchart of the text processing method based on the large language model provided in the fourth embodiment of the present disclosure.
[0124] like Figure 5 As shown, the text processing method based on the large language model may include the following steps S501 to S506:
[0125] Step S501 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0126] Step S502: Perform a second classification on the input text to obtain a text category.
[0127] For explanations of steps S501 to S502 , reference may be made to the relevant descriptions in any embodiment of the present disclosure and will not be repeated here.
[0128] Step S503, in response to the processing object being an agent, obtain an agent interaction topology diagram adapted to the text category; wherein, the nodes in the agent interaction topology diagram are used to indicate agents, and the edges between the nodes are used to indicate the interaction relationship between agents, and the agents include tool-type agents and / or non-tool-type agents bound to tools.
[0129] Tools include but are not limited to script commands, APIs, functions, etc.
[0130] In an embodiment of the present disclosure, when the classification result indicates that the processing object of the input text is an agent, an agent interaction topology diagram adapted to the text category to which the input text belongs can be obtained; wherein, the nodes in the agent interaction topology diagram are used to indicate agents, and the edges between the nodes are used to indicate the interaction relationship between agents, wherein the agents include tool-type agents and / or non-tool-type agents bound to tools.
[0131] It should be noted that for different text categories, the adapted intelligent agent interaction topology diagram can be different to improve the pertinence and accuracy of different input text processing.
[0132] Step S504: sequentially call the agents indicated by the nodes in the agent interaction topology diagram to process the input text.
[0133] In the embodiment of the present disclosure, the agents indicated by each node in the agent interaction topology diagram can be called in sequence to process the input text to obtain the execution results of the agents indicated by each node.
[0134] In any embodiment of the present disclosure, the agents indicated by each node in the agent interaction topology diagram can be called in sequence. When the currently called agent is the agent indicated by the root node in the agent interaction topology diagram, it can be determined whether the agent indicated by the root node is a tool-type agent. If so, the input text can be processed by the agent indicated by the root node (i.e., the tool-type agent) and the tool bound to the agent indicated by the root node to obtain the processing result of the agent indicated by the root node; if not, the input text can be directly processed by the agent indicated by the root node (i.e., the non-tool-type agent) to obtain the processing result of the agent indicated by the root node. Therefore, different processing strategies are adopted for tool-type agents and non-tool-type agents to carry out targeted processing on the input text, which can improve the pertinence, rationality and accuracy of the agent processing.
[0135] When the currently called agent is the agent indicated by the non-root node in the agent interaction topology diagram, it can be determined whether the agent indicated by the non-root node is a tool-type agent. If so, the processing results of the agent indicated by the parent node of the non-root node are processed through the agent indicated by the non-root node (i.e., the tool-type agent) and the tool bound to the agent indicated by the non-root node to obtain the processing results of the agent indicated by the non-root node; if not, the processing results of the agent indicated by the parent node of the non-root node are directly processed through the agent indicated by the non-root node (i.e., the non-tool-type agent) to obtain the processing results of the agent indicated by the non-root node.
[0136] In summary, the processing results of the agents indicated by the child nodes in the agent interaction topology diagram are dependent on the processing results of the agents indicated by their parent nodes. Through clear dependencies or interactions, the logical clarity of the text processing process can be ensured. Each agent only starts working after its parent node has completed processing. This reduces error handling caused by sequence errors or data inconsistencies, allowing each agent to operate in the correct context, thereby improving the reliability and accuracy of text processing. In addition, the dependencies between agents can be flexibly configured, which makes it easier for the system to add new agents or modify the dependencies between existing agents when faced with new text types or processing requirements, thereby improving the system's flexibility and scalability.
[0137] Step S505: Determine the response text according to the processing result of the agent indicated by the leaf node in the agent interaction topology diagram.
[0138] In the embodiment of the present disclosure, the processing results of the agents indicated by the leaf nodes in the agent interaction topology diagram can be integrated to determine the response text used to reply to the input text.
[0139] Step S506: reply to the input text based on the response text.
[0140] For explanation of step S506, please refer to the relevant description in any embodiment of the present disclosure, which will not be repeated here.
[0141] The text processing method based on the large language model in the embodiment of the present disclosure has high accuracy and reliability because the leaf node is located at the very end of the agent interaction topology diagram, and its processing result is often obtained based on the accumulated information and processing logic of all previous nodes (including its parent node). Therefore, according to the processing result of the agent indicated by the leaf node, the response information used to reply to the input text is determined, which can improve the accuracy and reliability of the text reply.
[0142] In order to clearly illustrate how the agents indicated by the nodes in the agent interaction topology diagram are called in sequence to process the input text in the above embodiment, the present disclosure also proposes a text processing method based on a large language model.
[0143] Figure 6 This is a flowchart of the text processing method based on the large language model provided in the fifth embodiment of the present disclosure.
[0144] like Figure 6 As shown, the text processing method based on the large language model may include the following steps S601 to S612:
[0145] Step S601 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0146] Step S602: Perform a second classification on the input text to obtain a text category.
[0147] Step S603: In response to the processing object being an agent, obtaining an agent interaction topology map adapted to the text category.
[0148] Among them, the nodes in the agent interaction topology diagram are used to indicate the agents, and the edges between the nodes are used to indicate the interaction relationship between the agents. The agents include tool-type agents and / or non-tool-type agents bound to tools.
[0149] Step S604: sequentially call the agents indicated by the nodes in the agent interaction topology diagram.
[0150] For explanations of steps S601 to S604 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0151] Step S605, in response to the currently called agent being the agent indicated by the root node in the agent interaction topology diagram, determine whether the agent indicated by the root node is a tool-type agent. If so, execute step S606; if not, execute step S607.
[0152] It should be noted that step S606 and step S607 are two parallel implementation methods. In actual application, only one needs to be executed.
[0153] Step S606: Process the input text through the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain the processing result of the agent indicated by the root node.
[0154] In an embodiment of the present disclosure, when the currently called agent is the agent indicated by the root node in the agent interaction topology diagram, and the agent indicated by the root node is a tool-type agent, the input text can be collaboratively processed by the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain the processing result of the agent indicated by the root node.
[0155] In any embodiment of the present disclosure, the following steps A to E may be used to process the input text to obtain the processing result of the agent indicated by the root node:
[0156] Step A: Obtain the input parameter description and function description associated with the tool bound to the agent indicated by the root node.
[0157] The input parameter description is used to describe the input parameter description of the tool. For example, the input parameter description of the tool can be described in markdown format. For example, the input parameter description of the tool can be marked as input_param_description.
[0158] The function description is used to describe the function of the tool. For example, the function description of the tool can be marked as description.
[0159] As an example, different interfaces can be built in the tool, and the input parameter description and function description of the tool can be obtained by calling the above interfaces.
[0160] For example, the following interfaces can be built in the tool:
[0161] Interface name: name
[0162] Input: None.
[0163] Output: Tool name.
[0164] Interface name: description / / tool function description
[0165] Input: None.
[0166] Output: Description of the tool's functionality.
[0167] Interface name: input_param_description / / tool input parameter description
[0168] Input: None.
[0169] Output: Description of tool input parameters in markdown format.
[0170] Interface name: run_with_params / / Tool execution results
[0171] Input: parameter list in json format.
[0172] Output: tool execution results.
[0173] Step B: Determine tool call parameters based on input text, input parameter description and function description.
[0174] In an embodiment of the present disclosure, the input text, input parameter description and function description can be combined to determine the calling parameters of the tool bound to the intelligent agent indicated by the root node, which are recorded as tool calling parameters in the present disclosure.
[0175] As a possible implementation method, a large language model can be used to process the input text, input parameter description and function description to obtain tool calling parameters.
[0176] As an example, first, at least one first reference example associated with the agent indicated by the root node can be obtained or queried, wherein the first reference example is used to indicate a conversion method for converting natural language information into input parameters (or parameter list) of a tool bound to the agent indicated by the root node. Afterwards, a first prompt message (Prompt) can be generated based on each first reference example, input parameter description, function description and input text. For example, each first reference example, input parameter description, function description and input text can be assembled to obtain a first prompt message, wherein the first prompt message is used to prompt the large language model to perform a call parameter generation task, and then the large language model can be called to process the first prompt message to obtain the tool call parameters.
[0177] Exemplarily, each first reference example associated with the agent indicated by the root node can be obtained by calling the agent interface.
[0178] For example, an agent interface may include:
[0179] Interface name: name
[0180] Input: None.
[0181] Output: Agent name.
[0182] Interface name: description / / Functional description of the agent
[0183] Input: None.
[0184] Output: A description of the agent's functionality.
[0185] Interface name: binding_tool_name / / Tool name bound to the agent
[0186] Input: None.
[0187] Output: The name of the bound tool.
[0188] Interface name: set_tool_input_examples / / Set tool input parameter examples (referred to as the first reference example in this disclosure)
[0189] Input: string. A series of reference examples showing how to convert natural language information into input parameter lists for tools bound to the agent.
[0190] Output: None.
[0191] Interface name: tool_input_examples / / Tool input parameter examples (referred to as the first reference example in this disclosure)
[0192] Input: None.
[0193] Output: Corresponding to the reference examples set by the set_tool_input_examples interface, these reference examples are returned.
[0194] In summary, the first prompt can serve as prior information or task information, indicating the task to be performed by the large language model, which can improve the prediction accuracy of the large language model. Furthermore, the first prompt can include numerous reference examples to indicate how to convert natural language information (such as colloquial input text) into tool input parameters, which can improve the stability of the large language model's generation process.
[0195] Step C: Based on the tool calling parameters, call the tool bound to the agent indicated by the root node to obtain the execution result (or output result) of the tool, which is recorded as the tool execution result in this disclosure.
[0196] Step D: Convert the tool execution results to obtain natural language text that matches the input format required by the agent indicated by the root node.
[0197] In the embodiment of the present disclosure, the tool execution result can be converted to obtain a natural language text (or colloquial text, colloquial result) that matches the input format required by the agent indicated by the root node.
[0198] As a possible implementation method, a large language model can be used to convert the tool execution results to obtain the natural language text required by the intelligent agent indicated by the root node.
[0199] As an example, first, the output format description associated with the tool bound to the agent indicated by the root node can be obtained, and at least one second reference example associated with the agent indicated by the root node can be obtained, wherein the second reference example is used to indicate the processing method of the output result of the tool bound to the agent indicated by the root node to the input data of the agent indicated by the root node. Afterwards, the second prompt information can be generated according to each second reference example, tool execution result, output format description and function description. For example, the various second reference examples, tool execution results, output format description and function description can be assembled to obtain the second prompt information, wherein the second prompt information is used to prompt the large language model to perform the text conversion task, and then the large language model can be called to process the second prompt information to obtain the natural language text required by the agent indicated by the root node.
[0200] For example, the output format description associated with the tool can be obtained through an interface built in the tool. For example, the following interface can also be built in the tool:
[0201] Interface name: output_param_description / / Output format description of the tool
[0202] Input: None.
[0203] Output: Tool output parameter description described in markdown format.
[0204] Exemplarily, the second reference examples associated with the agent indicated by the root node can be obtained by calling the agent interface. For example, the agent interface can also include:
[0205] Interface name: set_tool_output_examples / / Set the output format example of the tool (referred to as the second reference example in this disclosure)
[0206] Input: String. A series of reference examples showing how to convert the tool's output into the natural language text (also known as spoken text or spoken results) required by the agent.
[0207] Output: None.
[0208] Interface name: tool_output_examples / / Tool output format example (referred to as the second reference example in this disclosure)
[0209] Input: None.
[0210] Output: Corresponding to the reference examples set by the set_tool_output_examples interface, these reference examples are returned.
[0211] Interface name: query / / Input text in spoken form entered by the user
[0212] Input: None.
[0213] Output: Colloquial version of input text.
[0214] In summary, the second prompt can serve as prior information or task information, indicating the task to be performed by the large language model, which can improve the prediction accuracy of the large language model. Furthermore, the second prompt can include a large number of reference examples to indicate how to convert the tool's output (such as structured results in JSON format) into the natural language text (such as spoken text) required by the agent, which can improve the stability of the large language model's result generation process.
[0215] Step E: Use the agent indicated by the root node to process the natural language text to obtain the processing result of the agent indicated by the root node.
[0216] Step S607: Process the input text through the agent indicated by the root node to obtain the processing result of the agent indicated by the root node.
[0217] In an embodiment of the present disclosure, when the currently called agent is the agent indicated by the root node in the agent interaction topology diagram, and the agent indicated by the root node is a non-tool-type agent, the input text can be directly processed by the agent indicated by the root node (i.e., the non-tool-type agent) to obtain the processing result of the agent indicated by the root node.
[0218] Step S608, in response to the currently called agent being the agent indicated by the non-root node in the agent interaction topology diagram, determine whether the agent indicated by the non-root node is a tool-type agent. If so, execute step S609; if not, execute step S610.
[0219] It should be noted that step S609 and step S610 are two parallel implementation methods. In actual application, only one needs to be executed.
[0220] Step S609, the processing result of the agent indicated by the non-root node is processed by the agent indicated by the non-root node and the tool bound to the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
[0221] In the embodiment of the present disclosure, when the currently called agent is the agent indicated by the non-root node in the agent interaction topology diagram, and the agent indicated by the non-root node is a tool-type agent, the agent indicated by the non-root node and the tool bound to the agent indicated by the non-root node can be used to process the processing result of the agent indicated by the parent node of the non-root node to obtain the processing result of the agent indicated by the non-root node. The implementation principle is similar to that of step S606 and will not be repeated here.
[0222] As an example, the input parameter description and function description associated with the tool bound to the agent indicated by the non-root node can be obtained, and the tool calling parameters can be determined based on the processing results, input parameter description and function description of the agent indicated by the parent node of the non-root node. Afterwards, the tool bound to the agent indicated by the non-root node can be called based on the tool calling parameters to obtain the tool execution result. Then, the tool execution result can be converted to obtain a natural language text that matches the input format required by the agent indicated by the non-root node. Finally, the natural language text can be processed by the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
[0223] Step S610 , the processing result of the agent indicated by the parent node is processed by the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
[0224] In an embodiment of the present disclosure, when the currently called agent is the agent indicated by the non-root node in the agent interaction topology diagram, and the agent indicated by the non-root node is a non-instrumental agent, the processing results of the agent indicated by the parent node of the non-root node can be directly processed through the agent indicated by the non-root node (i.e., the non-instrumental agent) to obtain the processing results of the agent indicated by the non-root node.
[0225] Step S611, determining the response text according to the processing result of the agent indicated by the leaf node in the agent interaction topology diagram.
[0226] In the embodiment of the present disclosure, the processing results of the agents indicated by the leaf nodes in the agent interaction topology diagram can be integrated to determine the response text used to reply to the input text.
[0227] Step S612: reply to the input text based on the response text.
[0228] For explanation of step S612, please refer to the relevant description in any embodiment of the present disclosure and will not be repeated here.
[0229] The text processing method based on the large language model of the embodiment of the present disclosure, combined with the tool-type intelligent agent and the tools bound to it, collaboratively processes the input text, which can not only improve the stability of text processing, but also improve the accuracy of text processing.
[0230] In order to clearly illustrate any embodiment of the present disclosure, the present disclosure also proposes a text processing method based on a large language model.
[0231] Figure 7 This is a flowchart of the text processing method based on the large language model provided in Example 6 of the present disclosure.
[0232] like Figure 7 As shown, the text processing method based on the large language model may include the following steps S701 to S708:
[0233] Step S701 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0234] Step S702: Perform a second classification on the input text to obtain a text category.
[0235] Step S703: In response to the processing object being an agent, a plurality of agent interaction topology graphs adapted to the text category are obtained.
[0236] Among them, the nodes in the agent interaction topology diagram are used to indicate the agents, and the edges between the nodes are used to indicate the interaction relationship between the agents. The agents include tool-type agents and / or non-tool-type agents bound to tools.
[0237] For explanations of steps S701 to S703 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0238] Step S704: determine the interaction order between the multiple agent interaction topology diagrams, and call the multiple agent interaction topology diagrams in sequence based on the interaction order.
[0239] Among them, the interaction order of multiple intelligent agent interaction topology diagrams adapted to the text category can be pre-specified, or can be manually specified by relevant personnel, and the embodiments of the present disclosure do not limit this.
[0240] In an embodiment of the present disclosure, multiple agent interaction topology graphs can be called in sequence based on the interaction order between the multiple agent interaction topology graphs.
[0241] Step S705, in response to the currently called intelligent agent interaction topology diagram being the first called intelligent agent interaction topology diagram, the intelligent agents indicated by each node in the first intelligent agent interaction topology diagram are called in sequence to process the input text to obtain the processing result of the first intelligent agent interaction topology diagram.
[0242] In the disclosed embodiment, when the currently called agent interaction topology is the first called agent interaction topology, the agents indicated by the nodes in the first agent interaction topology may be called in sequence to process the input text to obtain the processing result of the first agent interaction topology; wherein the processing result of the first agent interaction topology may be: the processing result of the agent indicated by the leaf node in the first agent interaction topology. The implementation principle is similar to steps S604 to S610 and will not be repeated here.
[0243] Step S706, in response to the fact that the currently called intelligent agent interaction topology map is not the first called intelligent agent interaction topology map, the intelligent agents indicated by each node in the non-first intelligent agent interaction topology map are called in sequence, and the processing results of the previously called intelligent agent interaction topology map are processed to obtain the processing results of the non-first intelligent agent interaction topology map.
[0244] In an embodiment of the present disclosure, if the currently called agent interaction topology diagram is not the first called agent interaction topology diagram, the agents indicated by each node in the non-first agent interaction topology diagram may be called in sequence, and the processing result of the previously called agent interaction topology diagram may be processed to obtain the processing result of the non-first agent interaction topology diagram. The processing result of the non-first agent interaction topology diagram may be: the processing result of the agent indicated by the leaf node in the non-first agent interaction topology diagram.
[0245] Step S707, determining the response text according to the processing result of the last called intelligent agent interaction topology diagram.
[0246] In the embodiment of the present disclosure, the processing results of the agents indicated by the various leaf nodes in the last called agent interaction topology diagram can be integrated to determine the response text used to reply to the input text.
[0247] Step S708: reply to the input text based on the response text.
[0248] For explanation of step S708, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0249] The text processing method based on a large language model in the embodiment of the present disclosure calls the agents in the multiple agent interaction topology diagrams in sequence to collaboratively process the input text based on the interaction order between the multiple agent interaction topology diagrams. This can ensure that the processing of each agent is based on the processing result of the previous agent, forming a complete logical chain. This sequentiality ensures the consistency and accuracy of text processing.
[0250] In order to clearly illustrate how to obtain an agent interaction topology map adapted to the text category to which the input text belongs in the above embodiment, the present disclosure also proposes a text processing method based on a large language model.
[0251] Figure 8 This is a flowchart of the text processing method based on the large language model provided in Example 7 of the present disclosure.
[0252] like Figure 8 As shown, the text processing method based on the large language model may include the following steps S801 to S810:
[0253] Step S801 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0254] Step S802: Perform a second classification on the input text to obtain a text category.
[0255] For explanations of steps S801 to S802 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and will not be repeated here.
[0256] Step S803 , in response to the processing object being an intelligent agent, determines whether the text category is a newly added text category, if so, executes steps S806 to S808 , if not, executes steps S804 to S805 .
[0257] It should be noted that steps S806 to S808 and steps S804 to S805 are two parallel implementation methods. In actual application, only one needs to be executed.
[0258] Step S804 : Based on the known correspondence between text categories and business objects, determine a first business object corresponding to the text category from a plurality of business objects.
[0259] The number of first business objects (such as business teams, business parties) may be one or more, and this is not limited in the embodiment of the present disclosure.
[0260] In an embodiment of the present disclosure, when the text category to which the input text belongs is a newly added text category, the first business object corresponding to the text category can be determined from multiple business objects based on the correspondence between known text categories and business objects (such as business teams, business parties).
[0261] Step S805: Determine an agent interaction topology map that is adapted to the text category based on the agent interaction topology map maintained by the first business object.
[0262] Among them, the nodes in the agent interaction topology diagram are used to indicate the agents, and the edges between the nodes are used to indicate the interaction relationship between the agents. The agents include tool-type agents and / or non-tool-type agents bound to tools.
[0263] As an example, the agent interaction topology diagram maintained by each first business object can be used as the agent interaction topology diagram adapted to the text category to which the input text belongs.
[0264] Step S806 : In response to the second configuration instruction, a second business object adapted to the text category is determined from the plurality of business objects, and a corresponding relationship between the text category and the second business object is established.
[0265] The number of the second business objects (such as business teams, business parties) may be one or more, and this is not limited in the embodiment of the present disclosure.
[0266] In an embodiment of the present disclosure, when the text category to which the input text belongs is a newly added text category, a second business object that is adapted to the text category can be determined from multiple business objects in response to a configuration instruction (referred to as a second configuration instruction in the present disclosure) triggered by a relevant person (such as an administrator). That is, the relevant person manually selects the second business object that is adapted to the text category to which the input text belongs, and establishes a correspondence between the text category and the second business object.
[0267] Step S807: dispatch the input text to the second business object so that the second business object updates the corresponding maintained agent interaction topology diagram.
[0268] In an embodiment of the present disclosure, the input text may be dispatched to the second business object so that the second business object can update the agent interaction topology map maintained by itself based on existing business experience.
[0269] As a possible implementation method, a scheduling task can be sent to the second business object based on the input text and the text category to which it belongs, so that the second business object executes the scheduling task and updates the agent interaction topology map maintained by the second business object.
[0270] The scheduling task can be performed by following steps a to c:
[0271] Step a: Determine whether there is an agent in the agent library that is adapted to the text category to which the input text belongs. If so, execute step c; if not, execute steps b to c.
[0272] Step b: If there is no agent adapted to the text category in the agent library, then in response to the third configuration instruction triggered by the second business object, an agent adapted to the text category is created based on the input text.
[0273] As an example, the second business object can manually determine whether there is an agent in the agent library that is adapted to the text category to which the input text belongs. If not, the second business object can manually create an agent that is adapted to the text category based on the input text.
[0274] Step c: Based on the agent that matches the text category to which the input text belongs, the agent interaction topology graph maintained by the second business object is updated. For example, a new node is added to the agent interaction topology graph maintained by the second business object, indicating the agent that matches the text category to which the input text belongs. Alternatively, the node positions, dependencies between nodes, or interactions within the agent interaction topology graph maintained by the second business object may be updated.
[0275] As an example, the agent interaction topology diagram maintained by the second business object can be updated based on the agent adapted to the text category to which the input text belongs by dragging and dropping.
[0276] It should be noted that when the processing of an input text requires the participation of different business objects (or business parties), the business objects hold the most core experience of their business. If these experiences cannot be precipitated by the business objects themselves, then even if the intelligent processing system for processing input texts is implemented well, it is not maintainable. For example, when the distributed business system changes, the business objects cannot update the knowledge or experience in the intelligent processing system in a timely manner, so that the intelligent processing system is no longer effective. In the present disclosure, for newly added text categories, new logic or intelligent entities can be introduced in an "addition" manner to achieve business maintainability. Moreover, the new logic or intelligent entity is introduced by the business objects (or business parties) who maintain the distributed business system, which can reflect business friendliness.
[0277] As another possible implementation, the scheduling task can also be performed using the following steps d to f:
[0278] Step d: Determine whether there is an atomic tool associated with an intelligent agent that is adapted to the text category to which the input text belongs among the multiple atomic tools in the tool library. If so, execute step f; if not, execute steps e to f.
[0279] Step e: If there is no atomic tool associated with an intelligent agent that is adapted to the text category to which the input text belongs in the tool library, then according to the fourth configuration instruction, create at least one atomic tool associated with an intelligent agent that is adapted to the text category, and add the at least one atomic tool to the tool library.
[0280] As an example, the second business object can manually determine whether there is an atomic tool associated with an intelligent agent that is adapted to the text category to which the input text belongs in the tool library. If not, the second business object can manually create at least one atomic tool associated with an intelligent agent that is adapted to the text category based on the input text, and add the created atomic tools to the tool library.
[0281] Step f: In response to the fifth configuration instruction, assemble at least one atomic tool associated with the agent adapted to the text category to which the input text belongs to obtain a tool bound to the adapted agent.
[0282] As an example, at least one atomic tool associated with the agent adapted to the text category can be assembled by dragging and dropping, etc., to obtain a tool bound to the agent adapted to the text category.
[0283] In summary, when new text categories emerge, we only need to build new tools or new intelligent agents in a low-cost manner. At the same time, we do not make any changes to the original tools or intelligent agents. This can make the intelligent processing system for processing input text more powerful without any operation on the solution capabilities of existing text categories.
[0284] Step S808: Determine an agent interaction topology map adapted to the text category based on the updated agent interaction topology map maintained by the second business object.
[0285] As an example, the updated agent interaction topology graph maintained by each second business object can be used as the agent interaction topology graph adapted to the text category to which the input text belongs.
[0286] Step S809: sequentially call the agents indicated by the nodes in the agent interaction topology diagram to process the input text to obtain a response text.
[0287] Step S810: reply to the input text based on the response text.
[0288] For explanations of steps S809 to S810, reference may be made to the relevant descriptions in any embodiment of the present disclosure and will not be repeated here.
[0289] The text processing method based on a large language model of the embodiment of the present disclosure, when the text category to which the input text belongs is an original or known text category, directly determines the first business object corresponding to the text category to which the input text belongs from multiple business objects based on the correspondence between the known text category and the business object, and uses the intelligent agent interaction topology map maintained by the first business object as the intelligent agent interaction topology map adapted to the text category. This can realize the rapid processing of input text of a known text category based on existing business experience, thereby improving the processing efficiency of the input text.
[0290] In the case where the text category to which the input text belongs is a newly added text category, a second business object that is adapted to the newly added text category is determined from multiple business objects through a visual configuration method, and the intelligent agent interaction topology map maintained by itself is updated based on the visual configuration method through the second business object, so that the input text of the newly added text category can be processed based on the updated intelligent agent interaction topology map, which can improve the accuracy and effectiveness of text processing.
[0291] In order to clearly illustrate how the input text is first classified in the above embodiment, the present disclosure further proposes a text processing method based on a large language model.
[0292] Figure 9 This is a flowchart of the text processing method based on the large language model provided in the eighth embodiment of the present disclosure.
[0293] like Figure 9 As shown, the text processing method based on the large language model may include the following steps S901 to S906:
[0294] Step S901: Acquire input text and acquire at least one third reference example; wherein the third reference example includes a first reference text and a processing object of the first reference text.
[0295] For explanation of the input text, please refer to the relevant description in the above embodiment, which will not be repeated here.
[0296] The processing objects used to process the first reference text include but are not limited to: tools, agents, etc. The tools include but are not limited to: script commands, APIs, functions, etc.
[0297] Each third reference example is preset.
[0298] Step S902: Generate third prompt information based on the third reference example and the input text; wherein the third prompt information is used to prompt the large language model to perform the text classification task.
[0299] In the embodiment of the present disclosure, each third reference example can be assembled with the input text to obtain third prompt information, wherein the third prompt information is used to prompt the large language model to perform the text classification task.
[0300] Step S903: calling the large language model to process the third prompt information to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0301] In the embodiment of the present disclosure, a large language model may be called to process the third prompt information to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0302] Step S904: perform a second classification on the input text to obtain a text category.
[0303] Step S905: Based on the text category, the processing object is called to process the input text to obtain a response text.
[0304] Step S906: reply to the input text based on the response text.
[0305] The explanation of steps S904 to S906 can be found in the relevant description of any embodiment of the present disclosure and will not be repeated here.
[0306] In the text processing method based on a large language model in the disclosed embodiments, the third prompt information can serve as prior information or task information, indicating the task to be performed by the large language model, thereby improving the prediction accuracy of the large language model. Furthermore, the third prompt information can list a large number of reference examples to indicate the processing targets of different reference texts. This allows the large language model to perform classification only within the scope of the processing targets indicated by the third prompt information, rather than classifying all processing targets. This can eliminate classification illusions, thereby improving the reliability of the classification results and, in other words, enhancing the stability of the large language model's generation process.
[0307] In order to clearly illustrate how the second classification of the input text is performed in the above embodiment, the present disclosure further proposes a text processing method based on a large language model.
[0308] Figure 10 This is a flowchart of the text processing method based on the large language model provided in the ninth embodiment of the present disclosure.
[0309] like Figure 10 As shown, the text processing method based on the large language model may include the following steps S1001 to S1007:
[0310] Step S1001 , performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text.
[0311] For explanation of step S1001, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0312] Step S1002 : Acquire multiple fourth reference examples; wherein the fourth reference examples include the second reference text and the text category to which the second reference text belongs.
[0313] The text category to which the second reference text belongs may vary depending on the specific application scenario, business requirements, and characteristics of the text content of the second reference text.
[0314] Each fourth reference example is preset.
[0315] Step S1003: Generate fourth prompt information based on multiple fourth reference examples and the input text; wherein the fourth prompt information is used to prompt the large language model to perform the text classification task.
[0316] In an embodiment of the present disclosure, multiple fourth reference examples and input text may be assembled to obtain fourth prompt information, wherein the fourth prompt information is used to prompt the large language model to perform a text classification task.
[0317] Step S1004: calling the large language model to process the fourth prompt information to obtain the text category to which the input text belongs.
[0318] In an embodiment of the present disclosure, a large language model may be called to process the fourth prompt information to obtain the text category to which the input text belongs.
[0319] Step S1005: Based on the text category, the processing object is called to process the input text to obtain a response text.
[0320] Step S1006: reply to the input text based on the response text.
[0321] For explanations of steps S1005 to S1006 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0322] In the text processing method based on a large language model in the disclosed embodiments, the fourth prompt information can serve as prior information or task information, indicating the task to be performed by the large language model, thereby improving the prediction accuracy of the large language model. Furthermore, the fourth prompt information can list a large number of reference examples to indicate existing text categories, allowing the large language model to classify only within the text categories indicated by the fourth prompt information, rather than across all text categories. This can eliminate classification illusions, thereby improving the reliability of the classification results and, in other words, enhancing the stability of the large language model's generation process.
[0323] In any embodiment of the present disclosure, the following can be used: Figure 11 The low-code tool platform shown is used to implement the above-mentioned method embodiments, wherein the first layer (tool layer) and the second layer (scenario layer) in the low-code tool platform are the core two layers of this disclosure.
[0324] The first layer (the tool layer, also known as the tool library or tool market) completes tool construction through low-code methods. The individual tools in the tool layer are also called atomic tools.
[0325] Figure 11Clickhouse in the tool layer is a completely column-based distributed database management system; the callgraph platform is a distributed service tracking system used to monitor and track the call relationships between components in a distributed business system; the tower platform is a project management tool designed for collaboration among business teams. It provides tasks management, project progress tracking, team collaboration and other functions to help business teams advance projects more efficiently; NL2SQL is a technology that converts natural language into SQL (Structured Query Language) queries (Natural Language to Structured Query Language).
[0326] The second layer (the scenario layer, also known as the assembly layer) assembles the atomic tools from the first layer into larger tools that can process relatively complex input text. Furthermore, this layer can create intelligent agents and construct a topological graph of interactions between them to achieve the goal of processing complex text using natural language processing.
[0327] The low-code tool platform has at least the following advantages:
[0328] Part one: stable execution capability.
[0329] 1.1. Fine-grained agent.
[0330] The simpler the responsibility of an intelligent agent, the more stable its execution. Therefore, in this disclosure, when constructing an intelligent agent, it is avoided to be too large, and an intelligent agent can only complete one task. If the intelligent agent completes too many things, the intelligent agent is split.
[0331] 1.2. Tool-type intelligent agent.
[0332] The reason agents suffer from unstable execution is that they rely on LLMs. Tools, however, have hard-coded logic and thus do not face execution stability issues. In other words, tool execution is the most stable because, in this disclosure, tools can also "act as" agents, possessing the ability to interact with the outside world through natural language. This type of agent can be called a tool-type agent.
[0333] The specific steps are as follows:
[0334] (1) Interface built in the tool
[0335] Interface name: name / / name of the tool
[0336] Input: None.
[0337] Output: Tool name.
[0338] Interface name: description / / tool function description
[0339] Input: None.
[0340] Output: Description of the tool's functionality.
[0341] Interface name: input_param_description / / tool input parameter description
[0342] Input: None.
[0343] Output: Description of tool input parameters in markdown format.
[0344] Interface name: output_param_description / / Output format description of the tool
[0345] Input: None.
[0346] Output: Tool output parameter description described in markdown format.
[0347] Interface name: run_with_params / / Call the tool based on the parameters, the execution result of the tool
[0348] Input: parameter list in json format.
[0349] Output: tool execution results.
[0350] (2) Agent interface / / Agent name
[0351] Interface name: name
[0352] Input: None.
[0353] Output: Agent name.
[0354] Interface name: description / / Functional description of the agent
[0355] Input: None.
[0356] Output: A description of the agent's functionality.
[0357] Interface name: binding_tool_name / / The name of the tool bound to the agent
[0358] Input: None.
[0359] Output: The name of the bound tool.
[0360] Interface name: set_tool_input_examples / / Set tool input parameter examples
[0361] Input: a string, i.e., a series of examples showing how to convert natural language information into a list of input parameters for a tool bound to an agent, i.e., a string indicating how to convert natural language information into input parameters for a tool;
[0362] Output: None.
[0363] Interface name: tool_input_examples / / Tool input parameter examples
[0364] Input: None.
[0365] Output: The examples corresponding to the set_tool_input_examples interface setting are returned.
[0366] Interface name: set_tool_output_examples / / Set tool output format examples
[0367] Input: a string, i.e., a series of examples showing how to convert the output of the tool into the spoken input data of the agent, i.e., indicating the specific processing method of the output of the tool into the spoken input data;
[0368] Output: None.
[0369] Interface name: tool_output_examples / / Tool output format examples
[0370] Input: None.
[0371] Output: corresponds to the examples set by the set_tool_output_examples interface, and returns these examples.
[0372] Interface name: query / / user-entered spoken input text
[0373] Input: None.
[0374] Output: Colloquial version of the input text.
[0375] As an example, you can Figure 12 The interactive process shown in the figure is used to process the query input by the user, where: Figure 12 The middle gray area refers to the tool execution engine, which is responsible for coordinating function calls between tools and agents and interacting with LLM. Figure 12 The area to the left of is a tool-type agent. Figure 12 The area to the right of is the tool, and the two are bound through the binding_tool_name (binding tool name) interface of the intelligent body.
[0376] Figure 12 It consists of two parts, the upper and lower parts. Figure 12 The upper part mainly converts the input text (query) of the colloquial description into tool call parameters. Specifically, the tool execution engine calls the tool-type agent and tool-related functions (such as the tool-type agent's tool_input_examples, the tool's description, and the tool's input_param_description) to obtain the corresponding information and assemble it to obtain the prompt. Then, the tool execution engine uses the prompt to request the LLM and obtains the various parameters of the tool described in Json format returned by the LLM (i.e., the tool call parameters). Finally, the tool execution engine can call the tool's run_with_param interface based on these parameters, execute the tool, and obtain the tool's execution result (structured result in Json format).
[0377] Figure 12 The second half of the tool mainly completes the conversion of the tool's execution results (structured results) into the spoken results required by the tool-type agent, that is, converting the tool's execution results into spoken results that match the input format required by the tool-type agent. Specifically, first, the tool execution engine calls the tool-type agent and tool-related functions (such as the tool's description, the tool's output_param_description, and the tool-type agent's tool_output_examples) to obtain the corresponding information and assemble it to obtain the prompt. Then, the tool execution engine uses the prompt to request the LLM and obtains the spoken result required by the tool-type agent returned by the LLM. Finally, the tool execution engine returns the spoken result to the agent for processing to obtain the processed result output by the agent.
[0378] 1.3. Enhance stability by conveniently introducing examples.
[0379] During the execution of the above-mentioned tool-based intelligent agent, a large number of examples were listed when constructing the two prompts to further improve the stability of the LLM result generation process.
[0380] The second part is maintainability.
[0381] 2.1. Align organizational structure.
[0382] The interaction architecture of tool-based agents follows the division of business teams. First, there is a master entry agent at the top level. Under the master entry agent, there are more agents divided according to each business team. For each business team, the business team has a portal agent that connects to the master entry agent. Moreover, the portal agent of the business team can include more sub-agents for division of labor and collaboration, such as Figure 13 shown.
[0383] exist Figure 13 In the architecture shown, each business team maintains an agent within its own business scope. Furthermore, the agent can be split to the appropriate granularity for each business, and different agents interact with each other using natural language.
[0384] The general entry agent is responsible for understanding the original query input by the user and organizing the entry agents of each business team to complete the task.
[0385] 2.2. New text type query routing.
[0386] When a new text-type query appears, there is no need to modify the existing agent's already established logic, but only to add new logic. Specifically:
[0387] (1) Add an example of a query for a new text category to the main entry agent, indicating to which business team the query corresponding to the text category should be routed, or through which business teams' collaboration to complete it.
[0388] (2) If there is currently a lack of the necessary agent to complete the query within the agent maintained by the business team, a new agent is created.
[0389] The third part is low cost.
[0390] 3.1. Low-cost tool assembly.
[0391] The tool layer contains a wide variety of atomic tools, which are the foundation for building scenarios. To support the ability to "build scenarios at low cost," each atomic tool can be designed as follows:
[0392] (1) The input format and output format are fixed, both are a key-value map.
[0393] (2) Multiple atomic tools can be connected through Figure 14 The "data whiteboard" shown transfers data.
[0394] The "Data Whiteboard" can be considered a central data storage or message queue for transferring data between different atomic tools. The Data Whiteboard acts as a neutral, accessible data exchange center, allowing any atomic tool to read or write data on it. Transferring data through the "Data Whiteboard" has at least the following advantages:
[0395] Decoupling: Atomic tools do not need to communicate directly, but communicate indirectly through the "data whiteboard", which reduces the coupling between atomic tools and improves the maintainability and scalability of the system.
[0396] Asynchronous processing: Asynchronous data processing is supported. Atomic tools can publish results to the "data whiteboard" after completing the current task without having to wait for responses from other atomic tools.
[0397] Concurrent processing: Multiple atomic tools can read and process data from the "data whiteboard" at the same time, improving overall processing efficiency.
[0398] Once atomic tools support the above two capabilities, they can be assembled through a data whiteboard to form more complex tools. This can be accomplished through the drag-and-drop method of the UI (User Interface).
[0399] 3.2. Low-cost assembly of intelligent agents.
[0400] Similar to tools, agents can be assembled by dragging and dropping them on the UI. However, unlike tool assembly, tools communicate with each other through deterministic APIs, while assembled agents interact with each other through natural language.
[0401] In summary, the solution provided by this disclosure has at least the following advantages:
[0402] 1. Stability: By maintaining the fine-grained nature of the agent, implementing a tool-based agent, and using an example-based approach (i.e., setting examples), we ensure the stability requirements during the agent's execution.
[0403] 2. Maintainability: Align the agent architecture with the organizational structure and introduce new logic through incremental processing of new text types to achieve business maintainability.
[0404] 3. Low Cost: We implemented a "data whiteboard" for data exchange between tools and supported three assembly modes, enabling drag-and-drop assembly of tools. Agents are assembled in a similar manner. The logic within agents is fully input using business natural language. These methods enable low-cost creation of tools and agents.
[0405] With the above Figures 1 to 10 Corresponding to the text processing method based on the large language model provided in the embodiment, the present disclosure also provides a text processing device based on the large language model. Figures 1 to 10 The text processing method based on the large language model provided in the embodiment corresponds to the text processing method based on the large language model. Therefore, the implementation of the text processing method based on the large language model is also applicable to the text processing device based on the large language model provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0406] Figure 15 This is a structural diagram of a text processing device based on a large language model provided in the tenth embodiment of the present disclosure.
[0407] like Figure 15 As shown, the text processing device 1500 based on a large language model may include: a first classification module 1510 , a second classification module 1520 , a calling module 1530 and a reply module 1540 .
[0408] The first classification module 1510 is used to perform a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object of the input text;
[0409] A second classification module 1520 is used to perform a second classification on the input text to obtain a text category;
[0410] The calling module 1530 is used to call the processing object to process the input text based on the text category to obtain a response text;
[0411] The reply module 1540 is configured to reply to the input text based on the response text.
[0412] In a possible implementation of an embodiment of the present disclosure, the calling module 1530 is used to: in response to the processing object being a tool, obtain a target tool adapted to the text category; wherein the target tool is assembled from at least one atomic tool; and call the target tool to process the input text to obtain a response text.
[0413] In a possible implementation of an embodiment of the present disclosure, module 1530 is called to: in response to the processing object being a tool, query whether there is an assembly tool associated with the text category, wherein the assembly tool is obtained by assembling at least one atomic tool associated with the text category in the tool library within a historical period; if so, the assembly tool is used as the target tool; if not, an atomic tool adapted to the text category is obtained; wherein the adapted atomic tool includes: an atomic tool created based on the input text in response to the first configuration instruction, and / or an atomic tool in the tool library; the adapted atomic tools are assembled to obtain the target tool.
[0414] In a possible implementation of the embodiment of the present disclosure, there are multiple target tools, and the calling module 1530 is used to: determine the calling order among the multiple target tools, and call the multiple target tools in sequence based on the calling order; in response to the currently called target tool being the first called target tool, the first target tool is called to process the input text to obtain the processing result of the first target tool; in response to the currently called target tool being not the first called target tool, the non-first target tool is called to process the processing result of the previously called target tool to obtain the processing result of the non-first target tool; and determine the response text based on the processing result of the last called target tool.
[0415] In a possible implementation of the embodiment of the present disclosure, module 1530 is called to: in response to the processing object being an agent, obtain an agent interaction topology diagram adapted to the text category; wherein the nodes in the agent interaction topology diagram are used to indicate the agent, and the edges between the nodes are used to indicate the interaction relationship between the agents, and the agents include tool-type agents and / or non-tool-type agents bound to the tools; the agents indicated by each node in the agent interaction topology diagram are called in turn to process the input text; and the response text is determined based on the processing results of the agents indicated by the leaf nodes in the agent interaction topology diagram.
[0416] In a possible implementation of the embodiment of the present disclosure, the calling module 1530 is used to: sequentially call the agents indicated by the nodes in the agent interaction topology diagram; in response to the currently called agent being the agent indicated by the root node in the agent interaction topology diagram, and the agent indicated by the root node being a tool-type agent, the input text is processed by the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain the processing result of the agent indicated by the root node; in response to the agent indicated by the root node being a non-tool-type agent, the input text is processed by the agent indicated by the root node to obtain the processing result of the agent indicated by the root node. the processing result of the agent; in response to the currently called agent being the agent indicated by the non-root node in the agent interaction topology diagram, and the agent indicated by the non-root node being a tool-type agent, the processing result of the agent indicated by the parent node of the non-root node is processed through the agent indicated by the non-root node and the tool bound to the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node; in response to the agent indicated by the non-root node being a non-tool-type agent, the processing result of the agent indicated by the parent node is processed through the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
[0417] In a possible implementation of the embodiment of the present disclosure, the calling module 1530 is used to: determine the tool calling parameters based on the input text and the input parameter description and function description associated with the tool bound to the intelligent agent indicated by the root node; based on the tool calling parameters, call the tool bound to the intelligent agent indicated by the root node to obtain the tool execution result; convert the tool execution result to obtain a natural language text that matches the input format required by the intelligent agent indicated by the root node; use the intelligent agent indicated by the root node to process the natural language text to obtain the processing result of the intelligent agent indicated by the root node.
[0418] In a possible implementation of an embodiment of the present disclosure, the calling module 1530 is used to: obtain at least one first reference example associated with the intelligent agent indicated by the root node; wherein the first reference example is used to indicate a conversion method for converting natural language information into input parameters of a tool bound to the intelligent agent indicated by the root node; generate first prompt information based on each first reference example, input parameter description, function description and input text; wherein the first prompt information is used to prompt the large language model to perform a call parameter generation task; and use the large language model to process the first prompt information to obtain tool call parameters.
[0419] In a possible implementation of the embodiments of the present disclosure, module 1530 is called to: obtain an output format description associated with a tool bound to the agent indicated by the root node; obtain at least one second reference example associated with the agent indicated by the root node; wherein the second reference example is used to indicate a method for processing the output result of the tool bound to the agent indicated by the root node to the input data of the agent indicated by the root node; generate a second prompt message based on each second reference example, tool execution result, output format description and function description; wherein the second prompt message is used to prompt the large language model to perform a text conversion task; and the large language model is used to process the second prompt message to obtain the natural language text required by the agent indicated by the root node.
[0420] In a possible implementation of an embodiment of the present disclosure, module 1530 is called to: determine whether the text category is a newly added text category; if not, based on the correspondence between known text categories and business objects, determine a first business object corresponding to the text category from multiple business objects; and determine an agent interaction topology map adapted to the text category based on the agent interaction topology map maintained by the first business object.
[0421] In a possible implementation of an embodiment of the present disclosure, calling module 1530 is also used to: if yes, then in response to a second configuration instruction, determine a second business object that is adapted to the text category from multiple business objects, and establish a correspondence between the text category and the second business object; dispatch the input text to the second business object so that the second business object updates the corresponding maintained intelligent agent interaction topology map; and determine an intelligent agent interaction topology map that is adapted to the text category based on the updated intelligent agent interaction topology map maintained by the second business object.
[0422] In one possible implementation of the embodiments of the present disclosure, the first classification module 1510 is used to: obtain at least one third reference example; wherein the third reference example includes a first reference text and a processing object of the first reference text; generate third prompt information based on the third reference example and the input text; wherein the third prompt information is used to prompt the large language model to perform a text classification task; and call the large language model to process the third prompt information to obtain a classification result.
[0423] In one possible implementation of the embodiments of the present disclosure, the second classification module 1520 is used to: obtain at least one fourth reference example; wherein the fourth reference example includes a second reference text and a text category to which the second reference text belongs; generate fourth prompt information based on the fourth reference example and the input text; wherein the fourth prompt information is used to prompt the large language model to perform a text classification task; and call the large language model to process the fourth prompt information to obtain a text category.
[0424] The text processing device based on the large language model in the embodiment of the present disclosure can implement targeted processing of different input texts by calling a processing object adapted to the input text based on the text category to which the input text belongs, thereby improving the accuracy and rationality of text processing and improving the user experience.
[0425] In order to implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text processing method based on the large language model proposed in any of the above embodiments of the present disclosure.
[0426] In order to implement the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the text processing method based on a large language model proposed in any of the above embodiments of the present disclosure.
[0427] In order to implement the above embodiments, the present disclosure further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the text processing method based on the large language model proposed in any of the above embodiments of the present disclosure.
[0428] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0429] Figure 16 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. The electronic device may include the server and client in the above-mentioned embodiments. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0430] like Figure 16 As shown, the electronic device 1600 includes a computing unit 1601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1602 or a computer program loaded from a storage unit 1607 into a RAM (Random Access Memory) 1603. Various programs and data required for the operation of the device 1600 can also be stored in the RAM 1603. The computing unit 1601, the ROM 1602, and the RAM 1603 are connected to each other via a bus 1604. An I / O (Input / Output) interface 1605 is also connected to the bus 1604.
[0431] Various components in device 1600 are connected to I / O interface 1605, including an input unit 1606, such as a keyboard and mouse; an output unit 1607, such as various types of displays and speakers; a storage unit 1608, such as a magnetic disk and optical disk; and a communication unit 1609, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1609 allows device 1600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0432] Computing unit 1601 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of computing unit 1601 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphic Processing Units), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, DSPs (Digital Signal Processors), and any appropriate processors, controllers, microcontrollers, etc. Computing unit 1601 performs the various methods and processes described above, such as the aforementioned large language model-based text processing method. For example, in some embodiments, the aforementioned large language model-based text processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1600 via ROM 1602 and / or communication unit 1609. When the computer program is loaded into RAM 1603 and executed by computing unit 1601, one or more steps of the large language model-based text processing method described above may be performed. Alternatively, in other embodiments, computing unit 1601 may be configured to perform the large language model-based text processing method described above in any other appropriate manner (e.g., via firmware).
[0433] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0434] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0435] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0436] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0437] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0438] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0439] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0440] According to the technical solution of the embodiment of the present disclosure, it is possible to implement targeted processing of different input texts by calling a processing object adapted to the input text based on the text category to which the input text belongs, thereby improving the accuracy and rationality of text processing and improving the user experience.
[0441] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0442] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A text processing method based on a large language model, comprising: Performing a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object for the input text; wherein the processing object includes a tool or an intelligent agent; wherein the tool is obtained by assembling at least one atomic tool in response to an interactive operation triggered on a user interface, wherein the interactive operation includes a drag operation; and wherein the intelligent agent is used to implement a single function or perform a single task; Performing a second classification on the input text to obtain a text category; Based on the text category, calling the processing object to process the input text to obtain a response text; wherein the text category has a corresponding business object, the business object is used to maintain an agent interaction topology map adapted to the text category, and the agent indicated by each node in the agent interaction topology map is used to process the input text adapted to the text category; The input text is replied based on the response text.
2. The method according to claim 1, wherein The step of calling the processing object to process the input text based on the text category to obtain a response text includes: In response to the processing object being a tool, obtaining a target tool adapted to the text category; wherein the target tool is assembled from at least one atomic tool; The target tool is called to process the input text to obtain the response text.
3. The method according to claim 2, wherein: In response to the processing object being a tool, obtaining a target tool adapted to the text category includes: In response to the processing object being a tool, querying whether there is an assembly tool associated with the text category, wherein the assembly tool is obtained by assembling at least one atomic tool associated with the text category in a tool library within a historical period; If yes, the assembly tool is used as the target tool; If not, obtaining an atomic tool adapted to the text category; wherein the adapted atomic tool includes: an atomic tool created based on the input text in response to the first configuration instruction, and / or an atomic tool in the tool library; The adapted atomic tools are assembled to obtain the target tool.
4. The method according to claim 2, wherein: There are multiple target tools, and calling the target tools to process the input text to obtain the response text includes: Determining a calling order among a plurality of target tools, and calling the plurality of target tools in sequence based on the calling order; In response to the currently called target tool being the first called target tool, calling the first target tool to process the input text to obtain a processing result of the first target tool; In response to the currently called target tool being a non-first called target tool, calling the non-first target tool, processing the processing result of the previously called target tool, and obtaining the processing result of the non-first target tool; The response text is determined according to the processing result of the last called target tool.
5. The method according to claim 1, wherein The step of calling the processing object to process the input text based on the text category to obtain a response text includes: In response to the processing object being an agent, obtaining an agent interaction topology map adapted to the text category; wherein nodes in the agent interaction topology map are used to indicate agents, and edges between the nodes are used to indicate interaction relationships between agents, and the agents include tool-type agents and / or non-tool-type agents bound to tools; Invoking the agents indicated by the nodes in the agent interaction topology diagram in sequence to process the input text; The response text is determined according to the processing result of the agent indicated by the leaf node in the agent interaction topology diagram.
6. The method according to claim 5, wherein: The step of sequentially calling the agents indicated by the nodes in the agent interaction topology diagram to process the input text includes: Invoking the agents indicated by the nodes in the agent interaction topology diagram in sequence, in response to the currently invoked agent being the agent indicated by the root node in the agent interaction topology diagram, and the agent indicated by the root node being a tool-type agent, processing the input text by the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain a processing result of the agent indicated by the root node; In response to the agent indicated by the root node being a non-tool-type agent, processing the input text by the agent indicated by the root node to obtain a processing result of the agent indicated by the root node; In response to the currently called agent being the agent indicated by the non-root node in the agent interaction topology diagram, and the agent indicated by the non-root node being a tool-type agent, the processing result of the agent indicated by the parent node of the non-root node is processed by the agent indicated by the non-root node and the tool bound to the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node; In response to the agent indicated by the non-root node being a non-tool-type agent, the processing result of the agent indicated by the parent node is processed by the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
7. The method according to claim 6, wherein: The step of processing the input text by the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain the processing result of the agent indicated by the root node includes: Determine tool call parameters based on the input text and the input parameter description and function description associated with the tool bound to the agent indicated by the root node; Based on the tool calling parameters, calling the tool bound to the agent indicated by the root node to obtain the tool execution result; Converting the tool execution result to obtain a natural language text that matches the input format required by the agent indicated by the root node; The natural language text is processed using the agent indicated by the root node to obtain a processing result of the agent indicated by the root node.
8. The method according to claim 7, wherein: The determining of tool call parameters according to the input text and the input parameter description and function description associated with the tool bound to the agent indicated by the root node includes: Obtaining at least one first reference example associated with the agent indicated by the root node; wherein the first reference example is used to indicate a conversion method for converting natural language information into input parameters of a tool bound to the agent indicated by the root node; Generate first prompt information based on each of the first reference examples, the input parameter description, the function description, and the input text; wherein the first prompt information is used to prompt the large language model to execute a call parameter generation task; The first prompt information is processed using the large language model to obtain the tool calling parameter.
9. The method according to claim 7, wherein: The step of converting the tool execution result to obtain a natural language text that matches the input format required by the agent indicated by the root node includes: Obtain an output format description associated with a tool bound to the agent indicated by the root node; Obtain at least one second reference example associated with the agent indicated by the root node; wherein the second reference example is used to indicate a method for processing output results of a tool bound to the agent indicated by the root node into input data of the agent indicated by the root node; Generate second prompt information based on each of the second reference examples, the tool execution result, the output format description, and the function description; wherein the second prompt information is used to prompt the large language model to perform the text conversion task; The second prompt information is processed using the large language model to obtain the natural language text required by the agent indicated by the root node.
10. The method according to claim 5, wherein The obtaining of an agent interaction topology diagram adapted to the text category includes: Determining whether the text category is a newly added text category; If not, determining a first business object corresponding to the text category from a plurality of business objects based on a known correspondence between the text category and the business object; According to the agent interaction topology diagram maintained by the first business object, an agent interaction topology diagram adapted to the text category is determined.
11. The method according to claim 10, wherein: The step of obtaining an agent interaction topology diagram adapted to the text category further includes: If so, in response to a second configuration instruction, determining a second business object adapted to the text category from the plurality of business objects, and establishing a corresponding relationship between the text category and the second business object; Dispatching the input text to the second business object so that the second business object updates the corresponding maintained agent interaction topology graph; According to the updated agent interaction topology map maintained by the second business object, an agent interaction topology map adapted to the text category is determined.
12. The method according to any one of claims 1 to 11, wherein The first classification of the input text to obtain a classification result includes: Acquire at least one third reference example; wherein the third reference example includes a first reference text and a processing object of the first reference text; Generate third prompt information based on the third reference example and the input text; wherein the third prompt information is used to prompt the large language model to perform a text classification task; The large language model is called to process the third prompt information to obtain the classification result.
13. The method according to any one of claims 1 to 11, wherein The second classification of the input text to obtain a text category includes: Acquire at least one fourth reference example; wherein the fourth reference example includes a second reference text and a text category to which the second reference text belongs; Generate fourth prompt information based on the fourth reference example and the input text; wherein the fourth prompt information is used to prompt the large language model to perform a text classification task; The large language model is called to process the fourth prompt information to obtain the text category.
14. A text processing device based on a large language model, comprising: A first classification module is configured to perform a first classification on the input text to obtain a classification result; wherein the classification result is used to indicate a processing object for the input text; wherein the processing object includes a tool or an intelligent agent; wherein the tool is obtained by assembling at least one atomic tool in response to an interactive operation triggered on a user interface, wherein the interactive operation includes a drag operation; and wherein the intelligent agent is used to implement a single function or perform a single task; A second classification module is used to perform a second classification on the input text to obtain a text category; A calling module is configured to call the processing object based on the text category to process the input text and obtain a response text; wherein the text category has a corresponding business object, the business object is configured to maintain an agent interaction topology map adapted to the text category, and the agent indicated by each node in the agent interaction topology map is configured to process the input text adapted to the text category; A reply module is used to reply to the input text based on the response text.
15. The device according to claim 14, wherein The calling module is used to: In response to the processing object being a tool, obtaining a target tool adapted to the text category; wherein the target tool is assembled from at least one atomic tool; The target tool is called to process the input text to obtain the response text.
16. The device according to claim 15, wherein The calling module is used to: In response to the processing object being a tool, querying whether there is an assembly tool associated with the text category, wherein the assembly tool is obtained by assembling at least one atomic tool associated with the text category in a tool library within a historical period; If yes, the assembly tool is used as the target tool; If not, obtaining an atomic tool adapted to the text category; wherein the adapted atomic tool includes: an atomic tool created based on the input text in response to the first configuration instruction, and / or an atomic tool in the tool library; The adapted atomic tools are assembled to obtain the target tool.
17. The device according to claim 15, wherein There are multiple target tools, and the calling module is used to: Determining a calling order among a plurality of target tools, and calling the plurality of target tools in sequence based on the calling order; In response to the currently called target tool being the first called target tool, calling the first target tool to process the input text to obtain a processing result of the first target tool; In response to the currently called target tool being a non-first called target tool, calling the non-first target tool, processing the processing result of the previously called target tool, and obtaining the processing result of the non-first target tool; The response text is determined according to the processing result of the last called target tool.
18. The device according to claim 14, wherein The calling module is used to: In response to the processing object being an agent, obtaining an agent interaction topology map adapted to the text category; wherein nodes in the agent interaction topology map are used to indicate agents, and edges between the nodes are used to indicate interaction relationships between agents, and the agents include tool-type agents and / or non-tool-type agents bound to tools; Invoking the agents indicated by the nodes in the agent interaction topology diagram in sequence to process the input text; The response text is determined according to the processing result of the agent indicated by the leaf node in the agent interaction topology diagram.
19. The device according to claim 18, wherein The calling module is used to: Invoking the agents indicated by the nodes in the agent interaction topology diagram in sequence, in response to the currently invoked agent being the agent indicated by the root node in the agent interaction topology diagram, and the agent indicated by the root node being a tool-type agent, processing the input text by the agent indicated by the root node and the tool bound to the agent indicated by the root node to obtain a processing result of the agent indicated by the root node; In response to the agent indicated by the root node being a non-tool-type agent, processing the input text by the agent indicated by the root node to obtain a processing result of the agent indicated by the root node; In response to the currently called agent being the agent indicated by the non-root node in the agent interaction topology diagram, and the agent indicated by the non-root node being a tool-type agent, the processing result of the agent indicated by the parent node of the non-root node is processed by the agent indicated by the non-root node and the tool bound to the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node; In response to the agent indicated by the non-root node being a non-tool-type agent, the processing result of the agent indicated by the parent node is processed by the agent indicated by the non-root node to obtain the processing result of the agent indicated by the non-root node.
20. The device according to claim 19, wherein The calling module is used to: Determine tool call parameters based on the input text and the input parameter description and function description associated with the tool bound to the agent indicated by the root node; Based on the tool calling parameters, calling the tool bound to the agent indicated by the root node to obtain the tool execution result; Converting the tool execution result to obtain a natural language text that matches the input format required by the agent indicated by the root node; The natural language text is processed using the agent indicated by the root node to obtain a processing result of the agent indicated by the root node.
21. The device according to claim 20, wherein The calling module is used to: Obtaining at least one first reference example associated with the agent indicated by the root node; wherein the first reference example is used to indicate a conversion method for converting natural language information into input parameters of a tool bound to the agent indicated by the root node; Generate first prompt information based on each of the first reference examples, the input parameter description, the function description, and the input text; wherein the first prompt information is used to prompt the large language model to execute a call parameter generation task; The first prompt information is processed using the large language model to obtain the tool calling parameter.
22. The device according to claim 20, wherein The calling module is used to: Obtain an output format description associated with a tool bound to the agent indicated by the root node; Obtain at least one second reference example associated with the agent indicated by the root node; wherein the second reference example is used to indicate a method for processing output results of a tool bound to the agent indicated by the root node into input data of the agent indicated by the root node; Generate second prompt information based on each of the second reference examples, the tool execution result, the output format description, and the function description; wherein the second prompt information is used to prompt the large language model to perform the text conversion task; The second prompt information is processed using the large language model to obtain the natural language text required by the agent indicated by the root node.
23. The apparatus according to claim 18, wherein The calling module is used to: Determining whether the text category is a newly added text category; If not, determining a first business object corresponding to the text category from a plurality of business objects based on a known correspondence between the text category and the business object; According to the agent interaction topology diagram maintained by the first business object, an agent interaction topology diagram adapted to the text category is determined.
24. The device according to claim 23, wherein The calling module is further used to: If so, in response to a second configuration instruction, determining a second business object adapted to the text category from the plurality of business objects, and establishing a corresponding relationship between the text category and the second business object; Dispatching the input text to the second business object so that the second business object updates the corresponding maintained agent interaction topology graph; According to the updated agent interaction topology map maintained by the second business object, an agent interaction topology map adapted to the text category is determined.
25. The device according to any one of claims 14 to 24, wherein The first classification module is used to: Acquire at least one third reference example; wherein the third reference example includes a first reference text and a processing object of the first reference text; Generate third prompt information based on the third reference example and the input text; wherein the third prompt information is used to prompt the large language model to perform a text classification task; The large language model is called to process the third prompt information to obtain the classification result.
26. The device according to any one of claims 14 to 24, wherein The second classification module is used to: Acquire at least one fourth reference example; wherein the fourth reference example includes a second reference text and a text category to which the second reference text belongs; Generate fourth prompt information based on the fourth reference example and the input text; wherein the fourth prompt information is used to prompt the large language model to perform a text classification task; The large language model is called to process the fourth prompt information to obtain the text category.
27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the text processing method based on a large language model according to any one of claims 1 to 13.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the text processing method based on a large language model according to any one of claims 1 to 13.
29. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the text processing method based on a large language model according to any one of claims 1 to 13.
Citation Information
Patent Citations
Intelligent agent scheduling method, system and equipment based on large language model and medium
CN118132227A
Text processing method and device, electronic equipment and storage medium
CN118227868A