Interaction method and device based on artificial intelligence, intelligent agent, equipment and medium
By grouping tools and using matching filtering strategies for fine screening, the problem of tools and user requests not matching when the big model is combined with tools is solved, and interaction efficiency and feedback accuracy are improved.
Patent Information
- Application Number
- CN202510337120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-24
AI Technical Summary
When combining big models with tools, it becomes a difficult point to ensure that the called tools meet the expected results. The existing technology is difficult to effectively solve the problem of mismatch between tools and user requests, resulting in the final feedback not meeting user needs.
By grouping massive tools and dividing them into multiple tool ranges, using tool features to narrow the filter range in advance, increasing the difficulty of primary screening. Then, use a filtering strategy that matches the scope of the tool for fine screening, and formulate different screening strategies for the tool characteristics of tools within different tool scopes to improve the flexibility and pertinence of screening.
It improves interaction efficiency and user personalized needs, ensures the matching degree between tool calls and user requests, and thus improves the accuracy and user experience of feedback results.
Smart Images

Figure CN120197704A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as human-computer interaction, large models, and tool invocation. Specifically, it relates to an interaction method, device, intelligent agent, electronic device, storage medium, and program product based on artificial intelligence. Background Art
[0002] With the rapid development of computer technology and artificial intelligence technology, combining large models with various tools can provide different types of services such as natural language processing, image processing, and speech synthesis, thereby simplifying the application development process and improving development efficiency. However, how to ensure that the invoked tools meet the expected effects has become a difficult point in the combination of large models and tools. Summary of the Invention
[0003] This disclosure provides an interaction method, device, intelligent agent, electronic device, storage medium, and program product based on artificial intelligence.
[0004] According to one aspect of this disclosure, there is provided an interaction method based on artificial intelligence, including: determining a tool range of a tool to be invoked based on the intent information represented by the input content; using a screening strategy matching the above tool range to determine a target tool for processing the above input content from multiple tools belonging to the above tool range; obtaining an intermediate result by invoking the above target tool to process the above input content; and inputting the above intermediate result and the above input content into an interaction large model to obtain a feedback result for the above input content.
[0005] According to another aspect of this disclosure, there is provided an interaction device based on artificial intelligence, including: a primary screening module for determining a screening range of a tool to be invoked based on the intent information represented by the input content; a fine screening module for using a screening strategy matching the above tool range to determine a target tool for processing the above input content from multiple tools belonging to the above tool range; an intermediate processing module for obtaining an intermediate result by invoking the above target tool to process the above input content; and a feedback module for inputting the above intermediate result and the above input content into an interaction large model to obtain a feedback result for the above input content.
[0006] According to another aspect of this disclosure, there is provided an intelligent agent based on artificial intelligence, configured to execute the method as in this disclosure.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method as described above.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which an artificial intelligence-based interaction method and apparatus can be applied according to an embodiment of the present disclosure;
[0013] Figure 2 Schematically shows a flowchart of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;
[0014] Figure 3 Schematically shows a flowchart of determining a tool range according to an embodiment of the present disclosure;
[0015] Figure 4 Schematically shows a tool screening schematic diagram using a simplified screening strategy according to an embodiment of the present disclosure;
[0016] Figure 5 Schematically shows a tool screening schematic diagram using a complex screening strategy according to an embodiment of the present disclosure;
[0017] Figure 6 Schematically shows a block diagram of an artificial intelligence-based interaction apparatus according to an embodiment of the present disclosure;
[0018] Figure 7 Schematically shows a block diagram of an artificial intelligence-based intelligent agent according to an embodiment of the present disclosure; and
[0019] Figure 8 FIG. schematically shows a block diagram of an electronic device suitable for implementing an artificial intelligence-based interaction method according to an embodiment of the present disclosure. DETAILED IMPLEMENTATION MANNER
[0020] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0021] Function Call, also known as tool call. In response to receiving a user request, a tool can be called, and the large model uses the result returned by the tool for human-computer interaction operations.
[0022] With the rapid development of computer technology, the types of tools and the functions that can be achieved are becoming more and more powerful. For example, information retrieval, database operations, knowledge search and reasoning, system operations, etc. can be performed. In addition, user requests are also diverse. For example, chatting, playing chess, making short videos, etc. Sometimes, the reply content finally fed back to the user does not meet the user's needs due to the mismatch between the called tool and the user request.
[0023] In view of this, embodiments of the present disclosure provide an artificial intelligence-based interaction method, apparatus, intelligent agent, electronic device, storage medium, and program product. By grouping a large number of tools into multiple tool ranges, the screening range can be pre-narrowed using tool characteristics, increasing the difficulty of primary screening. Then, a fine screening is performed using a screening strategy matching the tool range, and different screening strategies are formulated for the tool characteristics of the tools within different tool ranges, which can improve the flexibility and pertinence of screening, thereby improving the interaction efficiency and the personalized needs of users.
[0024] Figure 1 FIG. schematically shows an exemplary system architecture to which an artificial intelligence-based interaction method and apparatus according to an embodiment of the present disclosure can be applied.
[0025] It should be noted that Figure 1 The example shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which an artificial intelligence-based interaction method and apparatus can be applied may include a terminal device, but the terminal device can implement the artificial intelligence-based interaction method and apparatus provided by embodiments of the present disclosure without interacting with a server.
[0026] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0028] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.
[0029] The server 105 may be a server providing various services, such as a background management server that supports the content browsed by users using the terminal devices 101, 102, 103 (only as an example). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0030] It should be noted that the artificial intelligence-based interaction method provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103.
[0031] Alternatively, the artificial intelligence-based interaction method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can generally be set in the server 105. The artificial intelligence-based interaction method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the artificial intelligence-based interaction device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0032] It should be understood,Figure 1 The number of terminal devices, networks, and servers in
[0033] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc., of the user's personal information all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good customs.
[0034] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.
[0035] It should be noted that the sequence numbers of the various operations in the following methods are only used as representations of the operations for description and should not be regarded as indicating the execution order of the various operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0036] Figure 2 A flowchart of an interaction method based on artificial intelligence according to an embodiment of the present disclosure is schematically shown.
[0037] As Figure 2 shown, the method includes operations S210 to S240.
[0038] In operation S210, based on the intent information represented by the input content, determine the tool scope of the tool to be called.
[0039] In operation S220, use a screening strategy that matches the tool scope to determine a target tool for processing the input content from multiple tools belonging to the tool scope.
[0040] In operation S230, process the input content by calling the target tool to obtain an intermediate result.
[0041] In operation S240, input the intermediate result and the input content into the interactive large model to obtain a feedback result for the input content.
[0042] The input content, which can also be referred to as a query, can be input by the user through a human-computer interaction interface. There is no limitation on the content type of the input content. For example, it can include one or more of text, image, document, video, and voice.
[0043] For example, input content A includes the text "Please search for the vehicle in this image" and the vehicle image.
[0044] It is possible to determine the intent information represented by the input content in response to the received input content. The intent information can include at least one of the task to be executed and the expected result.
[0045] For example, the intent information represented by input content A may include tasks to be executed such as "analyze image" and "search for knowledge".
[0046] The tool scope can refer to a set of tools. Multiple tool scopes can be pre-divided, and each tool scope can include multiple tools capable of executing tasks. The type of tools is not limited. For example, it can include search engines, deep learning models, computing functions, knowledge graphs, etc., as long as they are tools that can be called.
[0047] For example, the tool scope can include an internal tool library and an external tool library. The internal tool library can include at least one tool deployed locally. The external tool library can include at least one tool communicating through network protocols.
[0048] Based on at least one of the tasks to be executed and the expected results included in the intent information, the tool scope of the tool to be called can be determined to complete the preliminary screening.
[0049] For example, the tool scope matching "analyze image" is the external tool library. The tool scope matching "search for knowledge" is the internal tool library.
[0050] Each tool scope has a different screening strategy. According to the screening strategy matching the tool scope, the target tool for processing the input content can be determined from the multiple tools belonging to the tool scope.
[0051] For example, the tool types of the tools belonging to the internal tool library are fewer, and the number of tools under each tool type is also fewer. The screening strategy can adopt a simplified screening strategy.
[0052] Also for example, the tool types of the tools belonging to the external tool library are more, and the number of tools under each tool type is also more. The screening strategy can adopt a complex screening strategy.
[0053] After determining the target tool, the target tool can be called to process the input content to obtain an intermediate result.
[0054] For example, by calling an optical character recognition tool, it is determined that the vehicle image includes a license plate number "***" and a vehicle model number "+++". By calling an object recognition tool, attribute information such as the color of the vehicle is recognized. The vehicle image can also be processed by calling a knowledge search tool to obtain text information for describing the vehicle. The license plate number, vehicle model number, attribute information, and text information are all used as intermediate results.
[0055] The intermediate result and the input content are input into an interactive large model to use the interactive large model to process the intermediate result based on the intent represented by the input content to obtain a feedback result for the input content.
[0056] For example, the feedback result includes "The vehicle in the image is produced by Company AA and is rated as a star product of Company AA. The model of the vehicle is +++, which belongs to the off-road style and is suitable for driving on relatively rough roads. The driving experience..."
[0057] The interactive large model is applied in the human-computer interaction scenario to process the input content and intermediate results, thereby obtaining the feedback result for the user. The interactive large model can execute multiple tasks of different types, and the tasks can vary according to the different input content and intermediate results. For example, the input content includes an image and the text "Please introduce the car in the image", and the intermediate result can include the introduction content of the car. The interactive large model can execute the artificial intelligence generation task according to the image, text, and the introduction content of the car, integrate the introduction content, and obtain the relevant description for describing the car as the feedback result. Also for example, the input content includes the text "Please draw an image according to the song of singer AA", and the intermediate result can include the song of singer AA. The interactive large model can execute the text-to-image task according to the text and the song of singer AA, and obtain an image as the feedback result.
[0058] By using the interactive method provided in the embodiments of the present disclosure, the input content is processed by calling the target tool to obtain the intermediate result, so that the feedback result obtained by the interactive large model based on the intermediate result and the input content is accurate and effective, improving the interactive intelligence and user experience. In addition, a large number of tools are grouped into multiple tool ranges, which can pre-narrow the screening range by using the tool characteristics and increase the difficulty of the primary screening. Then, a fine screening is performed by using the screening strategy matching the tool range, and different screening strategies are formulated for the tool characteristics of the tools within different tool ranges, which can improve the flexibility and pertinence of the screening, and thus improve the interactive efficiency.
[0059] It should be noted that: The differences between the interactive large model mentioned above, the information generation large model, the rewriting large model, the expanding large model, etc. mentioned below lie in the different input data to be processed and the output results respectively. The similarities are: They are all models with a large number of parameters in the structure, and the order of the number of parameters is generally in the tens of millions, hundreds of millions or more, and may reach billions or tens of billions. In terms of the network structure of the large model, for example, network structures such as UFO (Unified Feature Optimization) can be adopted. By using the large model to execute tasks, the relatively powerful feature understanding ability and expression ability of the large model can be exerted, and efficient processing can be achieved. For example, the large model can include one or a combination of multiple of the large language model (LLM), the large vision model (LVM), or the multimodal large model (MLM).
[0060] The above provides a general description of the AI-based interaction method, and the following will elaborate on each operation.
[0061] For the operation S210 as Figure 2 shown, based on the intent information represented by the input content, determining the tool scope of the tool to be invoked may include: determining the task to be executed based on the intent information. Based on the task to be executed and the mapping relationship, determining the tool scope.
[0062] The input content can be subjected to intent recognition to obtain intent information. An encoder-decoder can be used to process the input content for intent recognition, but it is not limited to this. An intent recognition large model can also be used, as long as it is a deep learning model capable of performing intent recognition.
[0063] The intent information may include the task to be executed. However, it is not limited to this. The task to be executed included in the intent information and the predefined tasks associated with the intent information can also be combined and used as the task to be executed together.
[0064] A task table can be pre-generated, which includes multiple tasks that have a corresponding relationship with the intent information. Based on the intent information represented by the input content and the task table, determining the task to be executed.
[0065] For example, if the intent information includes the task of "querying the weather", the predefined tasks associated with the intent information may include "querying scenic spots", "querying food", etc. The tasks of "querying the weather", "querying scenic spots", and "querying food" can be combined into a task set, establishing a corresponding relationship with the intent information of "querying the weather", and generating a task table.
[0066] In the case where the input content includes "querying the weather in City AA in March", if the intent information is determined to include "querying the weather", then based on the task table and the intent information, the tasks to be executed are determined to include "querying the weather", "querying scenic spots", and "querying food".
[0067] Before dividing the tool scope, it is possible to pre-determine the tasks that each of the multiple tools within each tool scope can execute, obtaining a mapping relationship. The mapping relationship represents the mapping relationship between the tool scope and the task to be executed.
[0068] For example, if the intent information represents searching for images, based on the intent information and the task table, the tasks to be executed that match the image search can be determined to include image analysis and knowledge search. By querying the mapping relationship, it is determined that the tool scope corresponding to the execution of the image analysis task is the external tool library, and the tool scope corresponding to the execution of the knowledge search task is the internal tool library. Both the external tool library and the internal tool library are used as the tool scope of the tool to be invoked.
[0069] By using the method of determining the tool scope based on the task to be executed, the tool scope can be accurately matched through the clarity of the task to be executed. In addition, through the mapping relationship, the screening difficulty can be reduced, thereby improving the effectiveness and efficiency of the primary screening, and providing a basis for the fine screening of the target tool in the subsequent stage.
[0070] Exemplarily, based on the task to be executed and the mapping relationship that match the intent information, determining the tool scope may further include: determining the initial tool scope based on the task to be executed and the mapping relationship that match the intent information. In the case where multiple initial tool scopes are included, it is determined whether the multiple tasks to be executed corresponding to the multiple initial tool scopes are the same. In the case where they are the same, based on the tool characteristics of the initial tool scope, such as tool attribute information, one is determined from the multiple initial tool scopes as the tool scope of the tool to be called.
[0071] For example, in the case where the initial tool scope includes an internal tool library and an external tool library, it is determined whether the task to be executed corresponding to the external tool library is the same as the task to be executed corresponding to the internal tool library. In the case where they are the same, the communication duration can be determined based on the tool characteristics of the initial tool scope, such as the deployment method, and then the calling difficulty of the tool can be quantified through the communication duration. Then, based on the communication duration, the internal tool library can be determined as the tool scope of the tool to be called. This simplifies the subsequent screening while reducing the tool calling difficulty and improving the interaction efficiency.
[0072] In the case where they are not the same, multiple initial tool scopes can be simultaneously used as the screening objects for subsequent fine screening.
[0073] The following will be through the attached Figure 3 An illustration of how to determine the tool scope will be given.
[0074] Figure 3 A flowchart showing the process of determining the tool scope according to an embodiment of the present disclosure is schematically illustrated.
[0075] As Figure 3 shown, determining the tool scope includes operations S310 to S350.
[0076] In operation S310, based on the intent information, the task to be executed is determined.
[0077] In operation S320, based on the task to be executed and the mapping relationship, it is determined whether there are multiple initial tool scopes corresponding to a single task to be executed. In the case where there are multiple, operation S330 is executed. In the case where there is one, operation S350 is executed.
[0078] In operation S330, based on the tool attribute information of the tools belonging to the initial tool scope, the recognition result for characterizing the calling difficulty is determined.
[0079] In operation S340, based on the recognition result, the initial tool range with the lowest call difficulty is taken as the tool range of the tool to be called.
[0080] In operation S350, the initial tool range is determined as the tool range of the tool to be called.
[0081] The above has described in detail the primary screening for how to determine the screening range, and the following will describe in detail the fine screening of the tool.
[0082] Exemplarily, the tool range may include an internal tool library and an external tool library. Different screening strategies can be set according to the deployment methods and tool characteristics of different tool ranges.
[0083] For Figure 2 the operation S220 as shown, using a screening strategy matching the tool range, a target tool for processing the input content is determined from multiple tools belonging to the tool range, which may include: when the tool range includes an internal tool library, it is determined that the screening strategy includes a simplified screening strategy. When the tool range includes an external tool library, it is determined that the screening strategy includes a complex screening strategy.
[0084] Specifically, the internal tool library may include text processing tools, knowledge search tools, chess-playing tools, etc. The multiple tools in the internal tool library are quite different from each other. For example, the tool types, names, data types to be processed, and obtained results are all different. A simplified screening strategy can be adopted. The multiple tools in the external tool library are less different from each other and are of a wide variety. For example, image processing tools may include object recognition tools, character recognition tools, image correction tools, and texture generation tools, etc. In order to screen out accurate and effective tools from the external tool library, a complex screening strategy can be adopted.
[0085] Compared with the complex screening strategy, the simplified screening strategy has a lower processing difficulty and higher processing efficiency. Different screening strategies are adopted according to different tool ranges, improving the flexibility and effectiveness of tool screening.
[0086] The following will describe in detail the simplified screening strategy and the complex screening strategy respectively. First, the simplified screening strategy is introduced.
[0087] According to an embodiment of the present disclosure, when the tool range includes an internal tool library, a simplified screening strategy can be adopted. The simplified screening strategy may include: determining a target tool from multiple tools based on the similarity between the input content and the reference input content of the tool.
[0088] The internal tool library may include tools deployed locally. Being deployed locally can be understood as being deployed on an execution entity such as a server.
[0089] The reference input content can include the reference content for invoking the tool. The reference content can include the historical input content for invoking the tool, and can also include the content obtained by transforming the historical input content. Whether to invoke the tool can be determined based on the similarity between the input content and the reference input content. In the case of determining to invoke the tool, the tool is used as the target tool.
[0090] The similarity between the input content and the reference input content of the tool can include at least one of the following: format similarity, keyword similarity, semantic similarity, and type similarity.
[0091] The format similarity can represent whether the content types are the same, with 1 for the same and 0 for different. The keyword similarity can represent the same proportion or quantity of keywords. The semantic similarity can refer to the semantic matching degree.
[0092] Combining the tool features of multiple tools in the internal tool library and using a simplified screening strategy to screen tools can ensure the screening accuracy while improving the processing efficiency.
[0093] According to the embodiments of the present disclosure, the number of target tools can be not limited. For example, the target tool can include at least one.
[0094] Optionally, the target tool can include multiple ones. For example, in the process of screening the target tool from multiple tools, if the similarity 1 between the input content and the reference input content of tool 1, the similarity 2 between the input content and the reference input content of tool 2, and the similarity 3 between the input content and the reference input content of tool 3 are all higher than the predetermined similarity threshold, then tools 1, 2, and 3 with similarities higher than the predetermined similarity threshold can be used as the target tools.
[0095] In the case where the target tool includes multiple ones, since the tasks performed by the multiple target tools for processing the input content and the expected results obtained are basically the same, when multiple target tools are invoked simultaneously, it will cause waste of resources and redundancy of information. Therefore, in the internal tool library, the number of target tools for performing a single task can be limited to one. Exemplarily, the following method can be used to ensure that the target tool for performing a single task is unique.
[0096] For example, based on the format similarity between the content type of the input content and the content type of the reference input content, the initial target tool is determined from multiple tools. In the case where the number of initial target tools includes multiple ones, based on the semantic similarity between the semantic information of the input content and the semantic information of the reference input content of the initial target tool, the target tool is determined from the multiple initial target tools. In the case where the number of initial target tools includes one, the initial target tool is used as the target tool.
[0097] The content type can include text, images, videos, etc.
[0098] The matching of the content type is simple and effective. When the types of tools in the internal tool library are single and the number of tools is small, the screening efficiency and effectiveness are high.
[0099] When it is determined that there are multiple initial target tools, screening can be performed again based on semantic similarity, thereby ensuring the uniqueness of the target tool.
[0100] For example, the semantic information of the input content includes image recognition. The semantic information of the reference input content of the initial target tool A includes image correction, the semantic information of the reference input content of the initial target tool B includes texture repair, and the semantic information of the reference input content of the initial target tool C includes object recognition. Then, the initial target tool C with the highest semantic similarity can be used as the target tool.
[0101] Using the screening strategy of the target tool in the internal tool library provided by the embodiments of the present disclosure to screen the tools, thereby improving the screening effectiveness while ensuring the screening efficiency, and avoiding the problems of redundant intermediate results and high noise content.
[0102] The above has described the simplified screening strategy in detail, and the following will describe the method for obtaining the reference input content.
[0103] According to an embodiment of the present disclosure, before performing the operation S220 as Figure 2 shown, the interaction method based on artificial intelligence may further include: obtaining reference input content.
[0104] Obtaining the reference input content may include: collecting the historical input content of the invoked tool, and collecting the tendency evaluation information of the feedback result for the historical input content. The feedback result is generated from the intermediate result obtained by invoking the tool. When the tendency evaluation information represents a positive attitude of the user, the historical input content of the tool is used as the reference input content.
[0105] Obtaining the reference input content may also include: determining the historical input content of the tool. Using a rewriting large model to rewrite the historical input content to obtain the reference input content.
[0106] Optionally, when the tendency evaluation information represents a positive attitude of the user, the historical input content of the tool can be obtained. The historical input content is input into the rewriting large model to rewrite the historical input content using the rewriting large model to obtain the reference input content.
[0107] The rewriting may include: synonym replacement, content change with unchanged semantics, language transformation, etc.
[0108] By rewriting the historical input content, the resulting reference input content can be a set of content rather than just one. This can make the keywords used in the simplified screening strategy more accurately matched and reduce the workload and cost of collecting historical input content.
[0109] Next, the simplified screening strategy will be described by Figure 4 illustrating the simplified screening strategy.
[0110] Figure 4 FIG. schematically shows a schematic diagram of tool screening using the simplified screening strategy according to an embodiment of the present disclosure.
[0111] As Figure 4 shown, based on the format similarity 410 between the content type of the input content and the content type of the reference input content, a plurality of initial target tools 430 are determined from a plurality of tools in the internal tool library 420. Based on the semantic similarity 440 between the semantic information of the input content and the semantic information of the reference input content of the initial target tool, a target tool 450 is determined from the plurality of initial target tools. For example, the initial target tool with the highest semantic similarity is used as the target tool.
[0112] In the case where the number of initial target tools includes one, the initial target tool is used as the target tool.
[0113] The above has described the simplified screening strategy in detail. Next, the complex screening strategy will be described in detail.
[0114] In the case where the tool scope includes an external tool library, a complex screening strategy can be adopted. The complex screening strategy can include: determining a target tool from a plurality of tools based on the semantic similarity between the intent information and the tool description information of the tool.
[0115] The tool description information can be information used to characterize the tool features. For example, the tool description information can include information such as the function, identification, communication method, call parameters, etc. of the tool.
[0116] Based on the semantic similarity between the intent information and the tool description information, the tool with the highest semantic similarity can be determined as the target tool.
[0117] For example, if the intent information includes "analyze image", both the optical character recognition tool and the object recognition tool with image analysis functions are used as target tools. Tools such as texture generation tools and image correction tools in the image processing tool are not called.
[0118] Screening target tools from an external tool library can leverage the characteristics of a large number of tools and rich tool types in the external tool library, and use tool description information with sufficient tool feature information as reference data to improve the effectiveness of screening. In addition, compared with using keyword similarity, using semantic similarity as a quantification index can lower the screening criteria using semantic similarity and expand the quantity and richness of the screened target tools.
[0119] Optionally, in the case of multiple target tools, multiple target tools can be called simultaneously. For example, when calling tools in an external tool library, through a multi-threaded parallel call method, while increasing the richness and diversity of intermediate results, it avoids resource consumption problems caused by running tools on a local processor.
[0120] The following will be described by Figure 5 illustrate the complex screening strategy.
[0121] Figure 5 FIG. schematically shows a tool screening diagram using a complex screening strategy according to an embodiment of the present disclosure.
[0122] As Figure 5 shown, based on the semantic similarity 530 between the intent information 510 and the tool description information 520, the target tool 550, such as target tool 1, target tool 2,... target tool n, can be determined from multiple tools in the external tool library 540.
[0123] In the case of multiple target tools, multiple target tools are called in parallel using multiple threads, and the results each target tool feeds back are used as intermediate results.
[0124] The above has described the complex screening strategy in detail. The following will describe how to obtain tool description information.
[0125] According to an embodiment of the present disclosure, before performing the operation S220 as Figure 2 shown, the artificial intelligence-based interaction method may further include an operation: generating tool description information for each tool in the external tool library.
[0126] Optionally, generating tool description information for each tool in the external tool library may include: extracting target tool content from the development specification document of the tool according to a prompt information template. Filling the target tool content into the prompt information template to obtain prompt information. Inputting the prompt information into an information generation large model to obtain tool description information.
[0127] The development specification document of the tool can be referred to as a tool specification document or a design specification document. It may include content such as requirement specifications, design specifications, and function specifications.
[0128] The development instruction document of the tool is the basis for developing the tool. Therefore, in the case where a third party making the tool cannot provide tool description information, the development instruction document can be provided. Thus, by using the development instruction document to generate tool description information, while ensuring the richness and effectiveness of the tool description information, the difficulty of data acquisition can be reduced. The information generation large model can be used to perform intelligent generation tasks. Specifically, it can perform information generation processing on the target tool content in the prompt information to obtain tool description information. However, it is not limited to this. The development instruction document can also be directly input into the information generation large model to obtain tool description information.
[0129] Compared with the method of using the information generation large model to process the development instruction document to generate tool description information, the method of using the information generation large model to process the target tool content to generate tool description information can simplify the data processing volume of the information generation large model. In addition, it can also remove redundant information and improve the conciseness and effectiveness of the data input into the information generation large model.
[0130] Optionally, the function information of the tool can be directly used as the tool description information. However, the amount of information in the function information is small and the richness is low. Using the function information as the tool description information will affect the accuracy and effectiveness of the fine screening of the tool.
[0131] Next, an explanation of how to obtain the prompt information will be given.
[0132] According to an embodiment of the present disclosure, filling the target tool content into the prompt information template to obtain the prompt information may include: identifying the target tool content to obtain an expansion recognition result. In the case where the expansion recognition result indicates that the target tool content is to be expanded, input the target tool content into the expansion large model to use the expansion large model to expand the target tool content to obtain the expanded content. Fill the expanded content into the prompt information template to obtain the prompt information.
[0133] The target tool content can be extracted from the development instruction document by using the keyword matching method. However, it is not limited to this. The target tool content can also be extracted from the development instruction document by using the rule matching method.
[0134] A binary classification model can be used to identify the target tool content to obtain an expansion recognition result indicating whether to expand. In the case where the expansion recognition result indicates that the target tool content is to be expanded, use the expansion large model to expand the target tool content and fill the expanded content into the prompt information template. Otherwise, directly fill the target tool content into the prompt information template.
[0135] There is no limitation on the model structure of the binary classification model, as long as it is a deep learning model that can classify the target tool content to determine whether rewriting is required, such as a support vector machine, a random forest decision tree, etc.
[0136] Optionally, identifying the target tool content may further include: identifying the content richness of the target tool content to obtain a richness degree. The content richness identification may refer to the identification of the types of attribute content. Correspondingly, the richness degree can be used to evaluate the richness of the target tool content.
[0137] For example, the target tool content may include different categories of attribute content such as tool identifiers, tool functions, tool call methods, call parameters, etc. Identifying the target tool content may be the identification of the types of attribute content involved in the target tool content. If the number of types involved is large, the richness degree is high. Otherwise, the richness degree is low.
[0138] An expansion recognition result can be generated based on the richness degree. For example, when the richness degree is less than the richness degree threshold, the expansion recognition result indicates that the target tool content needs to be expanded. Otherwise, the expansion recognition result indicates that there is no need to expand the target tool content.
[0139] When the expansion recognition result indicates that the target tool content needs to be expanded, it means that the amount of information covered by the target tool content is not substantial, and there may be problems such as missing tool call methods or tool functions. Generating tool description information directly using the target tool content with insufficient information will reduce the effectiveness of the tool description information. When the amount of information covered by the target tool content is not substantial, the target tool content is input into the expansion large model to use the expansion large model to expand the target tool content, and tool description information is generated using the expanded content, thereby improving the effectiveness of the tool description information.
[0140] The information prompt template may refer to a template with a fixed pattern. The prompt information is obtained by filling the information prompt template and is used to indicate the operations expected to be performed by the large model and the results expected to be obtained. The expansion large model can be used to expand the content to obtain an output result with more information than the input content. Optionally, the expansion large model can be used to expand the target tool content to obtain expanded content with more information.
[0141] Exemplarily, the information prompt template may include the following content:
[0142] # Task description: You are a tool description generator, and your task is to generate a tool description information that meets the requirements based on the input information.
[0143] 1. When generating tool description information, it is necessary to refer to the <tool function>, <usage scenario>, <limiting conditions> provided by the user
[0144] 2. Just output the final description directly
[0145] # Content to be filled in:
[0146] <Tool function>
[0147] <Content to be filled in>
[0148] <Usage scenario>
[0149] <Content to be filled in>
[0150] <Limiting conditions>
[0151] <Content to be filled in>
[0152] Optionally, in addition to the information to be filled in such as tool functions, usage scenarios, and limiting conditions, the information prompt template can also include the following examples, for example, example tool description information recommended to be generated and example tool description information not recommended to be generated.
[0153] By expanding the information richness of the information prompt template, it can better guide the information generation large model, thereby improving the generation effect of the tool description information.
[0154] The above has described in detail how to determine the fine screening of the target tool. The following will describe in detail how to call the target tool to obtain intermediate results.
[0155] By calling the target tool to process the input content to obtain intermediate results, it can include: calling the target tool according to the call method matching the target tool to utilize the target tool to process the input content and obtain intermediate results.
[0156] The call methods of the internal tool library and the external tool library can be different according to different deployment methods. For example, the internal tool library can directly transmit information through wired communication to achieve the purpose of calling the target tool. Also for example, the external tool can communicate through network protocols to achieve the purpose of calling the target tool.
[0157] In addition, the call method can also include content transformation of the input content based on the input parameters of the target tool. For example, if the input content includes text and images, the input content can be content-transformed according to the content type of the input parameters of the target tool to meet the input parameter requirements of the target tool.
[0158] By flexibly adjusting the call method of the target tool, the stability and success rate of the target tool call can be improved.
[0159] The above has described in detail how to call the target tool to obtain the intermediate result. The following will describe how to obtain the feedback result by using the intermediate result.
[0160] According to an embodiment of the present disclosure, for the operation S240 as Figure 2 shown, inputting the intermediate result and the input content into the interactive large model to obtain the feedback result for the input content may include: performing noise reduction processing on the intermediate result to obtain the denoised content. Based on the denoised content and the input content, inputting them into the interactive large model for content generation processing to obtain the feedback result. However, it is not limited thereto. It may also include: by calling the interactive large model, performing multitask processing on the intermediate result and the input content to obtain the feedback result for the input content. The multitask processing may include: performing noise reduction processing on the intermediate result to obtain the denoised content. Based on the denoised content and the input content, performing content generation processing to obtain the feedback result.
[0161] The noise reduction processing may include deleting redundant content. By performing noise reduction processing, the processing accuracy of the content generation processing is improved.
[0162] The content generation processing may refer to the processing of Artificial Intelligence Generated Content. For example, using the interactive large model to perform generative processing of summaries on the intermediate result and the input content, or performing generative processing of songs.
[0163] Compared with the method of performing noise reduction processing on the intermediate result by using the interactive large model, pre-performing noise reduction processing on the intermediate result before inputting it into the interactive large model can reduce the noise content, thereby improving the processing accuracy of the interactive large model while reducing the operating cost of the interactive large model.
[0164] Optionally, the feedback result may be displayed on the human-computer interaction interface to complete a single-round interaction with the user.
[0165] Figure 6 Schematically shows a block diagram of an artificial intelligence-based interaction device according to an embodiment of the present disclosure.
[0166] As Figure 6 shown, the artificial intelligence-based interaction device 600 includes: a primary screening module 610, a fine screening module 620, an intermediate processing module 630, and a feedback module 640.
[0167] The primary screening module 610 is used to determine the screening range of the tool to be called based on the intention information represented by the input content.
[0168] The fine screening module 620 is configured to determine a target tool for processing the input content from multiple tools belonging to the tool scope by using a screening strategy matching the tool scope.
[0169] The intermediate processing module 630 is configured to obtain an intermediate result by calling the target tool to process the input content.
[0170] The feedback module 640 is configured to input the intermediate result and the input content into an interactive large model to obtain a feedback result for the input content.
[0171] According to an embodiment of the present disclosure, the fine screening module includes: a first screening sub-module.
[0172] The first screening sub-module is configured to, when the tool scope includes an internal tool library, determine a target tool from multiple tools based on the similarity between the input content and the reference input content of the tool, where the internal tool library includes tools deployed locally, and the reference input content includes reference content for calling the tool.
[0173] According to an embodiment of the present disclosure, the first screening sub-module includes: a first screening unit and a second screening unit.
[0174] The first screening unit is configured to determine an initial target tool from multiple tools based on the type similarity between the content type of the input content and the content type of the reference input content.
[0175] The second screening unit is configured to, when the number of initial target tools is multiple, determine a target tool from multiple initial target tools based on the semantic similarity between the semantic information of the input content and the semantic information of the reference input content of the initial target tool.
[0176] According to an embodiment of the present disclosure, the interactive device based on artificial intelligence further includes: a historical determination module and a rewriting module.
[0177] The historical determination module is configured to determine the historical input content of the tool.
[0178] The rewriting module is configured to rewrite the historical input content by using a rewriting large model to obtain the reference input content.
[0179] According to an embodiment of the present disclosure, the fine screening module includes: a second screening sub-module.
[0180] The second screening sub-module is configured to, when the tool scope includes an external tool library, determine a target tool from multiple tools based on the semantic similarity between the intent information and the tool description information of the tool, where the external tool library includes tools communicating through a network protocol, and the tool description information represents the tool characteristics of the tool.
[0181] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device further includes: an extraction module, a filling module, and an input module.
[0182] The extraction module is configured to extract target tool content from the development specification document of the tool according to the prompt information template.
[0183] The filling module is configured to fill the target tool content into the prompt information template to obtain prompt information.
[0184] The input module is configured to input the prompt information into the information generation large model to obtain tool description information.
[0185] According to an embodiment of the present disclosure, the filling module includes: an identification sub-module, an expansion sub-module, and a filling sub-module.
[0186] The identification sub-module is configured to identify the target tool content to obtain an expansion identification result.
[0187] The expansion sub-module is configured to, when the expansion identification result indicates that the target tool content needs to be expanded, input the target tool content into the expansion large model to expand the target tool content using the expansion large model to obtain expanded content.
[0188] The expansion sub-module is configured to fill the expanded content into the prompt information template to obtain prompt information.
[0189] According to an embodiment of the present disclosure, the primary screening module includes: a task determination sub-module and a scope determination sub-module.
[0190] The task determination sub-module is configured to determine a task to be executed based on the intent information.
[0191] The scope determination sub-module is configured to determine a tool scope based on the task to be executed and the mapping relationship, where the mapping relationship represents the corresponding relationship between the task to be executed and the tool scope.
[0192] According to an embodiment of the present disclosure, the feedback module includes: a noise reduction sub-module and a feedback sub-module.
[0193] The noise reduction sub-module is configured to perform noise reduction processing on the intermediate result to obtain noise-reduced content.
[0194] The feedback sub-module is configured to input the noise-reduced content and the input content into the interaction large model for content generation processing to obtain a feedback result.
[0195] Figure 7 A block diagram of an artificial intelligence-based intelligent agent according to an embodiment of the present disclosure is schematically shown.
[0196] AsFigure 7 As shown, the artificial intelligence-based agent 700 may include an input module 710, a processing module 720, and an output module 730.
[0197] The input module 710 is configured to receive input content.
[0198] The processing module 720 is configured to obtain a feedback result by executing the above-described artificial intelligence-based interaction method based on the input content received by the input module.
[0199] The output module 730 is configured to output the feedback result obtained by the processing module.
[0200] According to an embodiment of the present disclosure, the input module 710 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the agent 700 can understand and process. The input module 710 is the primary link for the agent 700 to interact with the outside world, enabling the agent 700 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0201] In an example, the input module 710 may input the input content described above.
[0202] In an example, the processing module 720 is the core support for the agent 700's ability to handle complex tasks. The processing module 720 may execute the artificial intelligence-based interaction method described above.
[0203] In an example, the performance of the processing module 720 may be closely related to the interaction large model on which the agent 700 is based. To fully utilize the capabilities of the interaction large model, the internal structure of the processing module 720 may be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.
[0204] In an example, after the agent 700 obtains the input content, the processing module 720 may determine the tool scope of the tool to be called based on the intent information represented by the input content. Using a screening strategy that matches the tool scope, the target tool for processing the input content is determined from multiple tools belonging to the tool scope. The processing module 720 may process the input content by calling the target tool to obtain an intermediate result. The processing module 720 may call the interaction large model based on the intermediate result and the input content to obtain a feedback result for the input content, and transmit the feedback result to the output module 730.
[0205] It can be understood that although large models have excellent language understanding and generation capabilities, like humans, they can perform only a limited number of tasks without any tools. After the agent 700 is given the ability to call tools, it can perform tasks such as completing mathematical operations with the help of a calculator, performing data analysis with the help of the Python language, and obtaining weather forecasts with the help of a search engine.
[0206] In the example, the output module 730 can output the feedback results described above.
[0207] The agent 700 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and enhance flexibility and versatility.
[0208] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0209] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0210] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described above.
[0211] According to an embodiment of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0212] Figure 8 FIG. shows a schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0213] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0214] Multiple components in the device 800 are connected to the input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0215] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as an AI-based interaction method. For example, in some embodiments, the AI-based interaction method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the AI-based interaction method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the AI-based interaction method in any other appropriate manner (e.g., by means of firmware).
[0216] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0217] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0218] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0220] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0221] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0222] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0223] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An interactive method based on artificial intelligence, comprising: Determine the tool scope of the tool to be called based on the intention information represented by the input content; Determine a target tool for processing the input content from a plurality of tools belonging to the tool range by using a screening strategy matching the tool range; Processing the input content by calling the target tool to obtain an intermediate result; as well as The intermediate result and the input content are input into the interactive large model to obtain a feedback result for the input content.
2. The method according to claim 1, wherein: The method of using a screening strategy that matches the tool range to determine a target tool for processing the input content from a plurality of tools belonging to the tool range includes: In a case where the tool scope includes an internal tool library, the target tool is determined from a plurality of the tools based on a similarity between the input content and a reference input content of the tool, wherein the internal tool library includes locally deployed tools, and the reference input content includes reference content for calling the tool.
3. The method according to claim 2, wherein: The determining the target tool from the plurality of tools based on the similarity between the input content and the reference input content of the tool comprises: determining an initial target tool from a plurality of the tools based on a type similarity between a content type of the input content and a content type of the reference input content; and In the case where the number of the initial target tools includes a plurality, the target tool is determined from the plurality of the initial target tools based on the semantic similarity between the semantic information of the input content and the semantic information of the reference input content of the initial target tool.
4. The method according to claim 2 or 3, further comprising: determining historical inputs to said tool; as well as The historical input content is rewritten using the rewriting large model to obtain the reference input content.
5. The method according to any one of claims 1 to 4, wherein: The method of using a screening strategy that matches the tool range to determine a target tool for processing the input content from a plurality of tools belonging to the tool range includes: In a case where the tool range includes an external tool library, the target tool is determined from a plurality of the tools based on the semantic similarity between the intent information and the tool description information of the tool, wherein the external tool library includes tools that communicate via a network protocol, and the tool description information characterizes tool features of the tool.
6. The method according to claim 5, further comprising: Extracting target tool content from the development description document of the tool according to the prompt information template; Filling the target tool content into the prompt information template to obtain prompt information; as well as The prompt information is input into the information generation model to obtain the tool description information.
7. The method according to claim 6, wherein: The step of filling the target tool content into the prompt information template to obtain the prompt information includes: Identify the target tool content to obtain an expanded identification result; In the case where the expansion recognition result indicates that the target tool content is expanded, the target tool content is input into the expansion large model, so as to expand the target tool content using the expansion large model to obtain the expanded content; and The expanded content is filled into the prompt information template to obtain the prompt information.
8. The method according to any one of claims 1 to 7, wherein: The step of determining the tool range of the tool to be called based on the intention information represented by the input content includes: Based on the intention information, determining a task to be performed; and The tool range is determined based on the tasks to be performed and the mapping relationship, wherein the mapping relationship represents the corresponding relationship between the tasks to be performed and the tool range.
9. The method according to any one of claims 1 to 8, wherein: The step of inputting the intermediate result and the input content into the interactive macro model to obtain a feedback result for the input content includes: Performing noise reduction processing on the intermediate result to obtain noise-reduced content; and Based on the denoised content and the input content, the content is input into the interactive large model, and content generation processing is performed to obtain the feedback result.
10. An interactive device based on artificial intelligence, comprising: A primary screening module, used to determine the screening scope of the tool to be called based on the intent information represented by the input content; A fine screening module, used for determining a target tool for processing the input content from a plurality of tools belonging to the tool range by using a screening strategy matching the tool range; An intermediate processing module, used to process the input content by calling the target tool to obtain an intermediate result; as well as A feedback module is used to input the intermediate result and the input content into the interactive large model to obtain a feedback result for the input content.
11. An artificial intelligence-based agent, wherein: The agent is configured to perform the method according to any one of claims 1 to 9.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Cited By
Question and answer data processing method, training method and training system
CN121278053A