Interaction method and device based on large model, intelligent agent and storage medium

By adjusting the execution logic in the large language model based on object-related features and intent understanding, the problem of low response content matching in large language model interactions is solved, and efficient and accurate user interaction is achieved.

CN120706575AActive Publication Date: 2025-09-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510896743.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-26
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Large language models find it difficult to flexibly output responses that match user quality requirements when interacting with users, resulting in high complexity of interactive operations and high learning costs.

Method used

By understanding the intent of the input information based on object-related features, the target intent is determined, and the big model is used to perform semantic understanding tasks based on the execution logic that matches the target intent, generate reply information, and automatically adjust the thinking mode and task execution mode of the big model to match user needs.

Benefits of technology

It achieves precise matching of reply information and user needs, reduces the user's learning cost and complexity of interactive operations, and improves interaction efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706575A_ABST
    Figure CN120706575A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device based on a large model, an intelligent agent and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, intelligent search, intelligent customer service, intelligent medical treatment and the like. The interaction method based on the large model comprises the following steps: receiving input information of a target object; performing intention understanding on the input information based on the object related features of the target object to obtain a target intention, the target intention representing a content quality demand for the reply information; executing a semantic understanding task on the input information by utilizing the large model based on execution logic matched with the target intention to obtain reply information; and pushing reply information to the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as deep learning, intelligent search, smart customer service, and smart medical care. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) can be used to semantically understand the voice, text and other information input by users to generate response content that meets their needs. Summary of the Invention

[0003] The present disclosure provides a large model-based interaction method, device, intelligent agent, electronic device and storage medium.

[0004] According to one aspect of the present disclosure, an interaction method based on a big model is provided, including: receiving input information of a target object; understanding the intent of the input information based on object-related features of the target object to obtain a target intent, wherein the target intent represents the content quality requirements for the reply information; using the big model to perform a semantic understanding task on the input information based on an execution logic matching the target intent to obtain reply information; and pushing the reply information to the target object.

[0005] According to another aspect of the present disclosure, an interaction device based on a large model is provided, including: a receiving module for receiving input information of a target object; a first obtaining module for understanding the intent of the input information based on object-related features of the target object to obtain a target intent, where the target intent represents the content quality requirements for the reply information; a second obtaining module for using the large model to perform a semantic understanding task on the input information based on an execution logic matching the target intent to obtain reply information; and a pushing module for pushing the reply information to the target object.

[0006] According to another aspect of the present disclosure, an artificial intelligence agent is provided, including: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and executing the method provided by the embodiment of the present disclosure by calling the large model; and an output module for outputting output information obtained by the processing module.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to an embodiment of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided according to an embodiment of the present disclosure.

[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided according to the embodiment of the present disclosure when executed by a processor.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 Schematically illustrates an exemplary system architecture to which a large model-based interaction method and apparatus according to an embodiment of the present disclosure can be applied;

[0013] Figure 2 The flowchart of the interaction method based on the large model according to the embodiment of the present disclosure is schematically shown;

[0014] Figure 3 The following schematically illustrates a principle diagram of an interaction method based on a large model according to an embodiment of the present disclosure;

[0015] Figure 4 Schematically illustrates a principle diagram of an interaction method based on a large model provided according to another embodiment of the present disclosure;

[0016] Figure 5 The following schematically illustrates an application scenario diagram of an interaction method based on a large model according to an embodiment of the present disclosure;

[0017] Figure 6 A block diagram schematically illustrates an interactive device for a large model according to an embodiment of the present disclosure;

[0018] Figure 7 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0019] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the large model-based interaction method of an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0022] The inventors discovered that the large language model can process user input to perform semantic understanding and generate responses that match the user's needs based on the large language model's relatively powerful semantic understanding and generation capabilities. At the same time, the relevant platform can provide users with options such as "deep thinking" based on their actual needs to control the large language model to output high-quality responses. However, during user interaction, the large language model finds it difficult to flexibly output responses that match the user's quality requirements based on the user's input, resulting in users having to control the quality of the response content at a high learning cost, which leads to high complexity and learning costs in interactive operations.

[0023] The embodiments of the present disclosure provide a large-scale model-based interaction method, apparatus, intelligent agent, electronic device, and storage medium. The large-scale model-based interaction method includes: receiving input information from a target object; performing intent understanding on the input information based on object-related features of the target object to obtain a target intent, wherein the target intent represents the content quality requirements for the reply information; using the large-scale model to perform a semantic understanding task on the input information based on execution logic matching the target intent to obtain a reply information; and pushing the reply information to the target object.

[0024] According to the embodiments of the present disclosure, the target intent is determined by understanding the intent of the input information of the target object through object-related features, so that the target intent can more accurately represent the target object's requirements for the content quality level of the reply information output by the large model after processing the input information. By controlling the large model to perform semantic understanding tasks based on the execution logic that matches the target intent, it is possible to automatically adjust the execution logic of the large model for performing semantic understanding tasks on the input information according to the preferences or habits of the target object, and flexibly adjust the thinking mode and task execution mode of the large model, thereby automatically adjusting the content quality of the reply information to match the intent expressed by the input information of the target object, avoiding a large deviation between the content quality of the reply information and the requirements of the target object, or the execution logic of the large model being too long, which affects the user experience, and improves interaction efficiency and reply accuracy.

[0025] Figure 1 An exemplary system architecture to which the large model-based interaction method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0026] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not imply that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the large-model-based interaction method and apparatus may be applied may include a terminal device, but the terminal device may implement the large-model-based interaction method and apparatus provided in the embodiments of the present disclosure without interacting with a server.

[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0029] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0030] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0031] It should be noted that the large model-based interaction method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the large model-based interaction device provided in the embodiment of the present disclosure can also be set in the terminal device 101, 102, or 103.

[0032] Alternatively, the interaction method based on the big model provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the interaction device based on the big model provided in the embodiment of the present disclosure may generally be set in the server 105. The interaction method based on the big model provided in the embodiment of the present disclosure may also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the interaction device based on the big model provided in the embodiment of the present disclosure may also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0034] It should be noted that the information obtained in any embodiment of this disclosure, including but not limited to object-related characteristics, is obtained only with the authorization of the relevant user or organization. The purpose of the information obtained is disclosed in advance, and necessary encryption or desensitization measures are adopted, in accordance with relevant laws and regulations and without violating public order and good morals.

[0035] Figure 2 The flowchart of the interaction method based on the large model according to the embodiment of the present disclosure is schematically shown.

[0036] like Figure 2 As shown, the large model-based interaction method includes operations S210 to S240.

[0037] In operation S210 , input information of a target object is received.

[0038] In operation S220 , the input information is subjected to intention understanding based on the object-related features of the target object to obtain the target intention.

[0039] In operation S230 , the semantic understanding task is performed on the input information using the large model based on the execution logic matching the target intent to obtain reply information.

[0040] In operation S240 , the reply information is pushed to the target object.

[0041] According to embodiments of the present disclosure, the input information can be data in any modality, such as text or voice, input by the target subject. For example, the target subject can use voice input to input a request such as "Help me check today's weather." Voice recognition can be performed on the request to obtain text or a vector corresponding to the request as input information.

[0042] According to the embodiments of the present disclosure, a large model, or large language model, can be understood as a model built based on deep learning algorithms that has the ability to process natural language and generate response information. Large models can have billions or even hundreds of billions of parameters. These parameters enable the model to capture and understand the complexity and diversity of language and output generative content in any modality, such as text and images, based on the large-scale model parameters.

[0043] According to an embodiment of the present disclosure, object-related features may represent features related to the target object's own attributes, interactive behaviors, etc. For example, object-related features may be the target object's interest preference features, resource content features that have performed interactive behavior operations, etc.

[0044] According to embodiments of the present disclosure, understanding the intent of input information based on object-related features of a target object may include processing the input information and object-related features using a neural network algorithm to obtain the target intent. For example, the object-related features and input information may be processed using a convolutional neural network algorithm to obtain the target intent. However, this is not limited to this, and other types of algorithms may be used to process the object-related features and input information to obtain the target intent. The embodiments of the present disclosure do not limit the specific algorithm type used to process the object-related features and input information.

[0045] According to an embodiment of the present disclosure, the target intent represents the content quality requirements for the reply information. The content quality requirements can represent any requirement level related to content quality, such as the richness of the reply information, the number of knowledge points, and logical coherence. For example, the content quality requirement can be the quality requirements for the logic and vividness of the press release for the input information "A press release about the sports meeting needs to be written."

[0046] In some embodiments, the content quality requirements of the response information represented by the target intent can be used to control the task execution logic and thinking mode of the large model. By using the large model to adjust the execution logic and thinking mode of the semantic task based on the content quality requirements represented by the target intent, the large model can output high-quality response information through the execution logic corresponding to deep thinking, or quickly generate response information that matches the content quality requirements through the execution logic corresponding to non-deep thinking.

[0047] For example, when the target intent indicates that the content quality requirements meet the preset quality requirement conditions, the large model can perform semantic understanding tasks based on the deep thinking pattern corresponding to the target intent to output high-quality response information.

[0048] For example, when the target intent indicates that the content quality requirement does not meet the preset quality requirement conditions, the large model can perform semantic understanding tasks based on the non-deep thinking mode corresponding to the target requirement intent to quickly output reply information.

[0049] In some embodiments, a first execution duration of the semantic understanding task performed by the large model based on the deep thinking mode is longer than a second execution duration of the semantic understanding task performed by the large model based on the deep thinking mode. For example, the first execution duration for determining a reply message is one second longer than the second execution duration for determining a reply message.

[0050] According to the embodiments of the present disclosure, by detecting the intent of the input information based on object-related features and controlling the execution logic of the large model according to the content quality requirements of the reply information represented by the target intent, it is possible to automatically adjust the task execution logic and thinking mode of the large model according to the input information and object-related features of the target object without the target object's perception, generate reply information that matches the actual content quality requirements of the target object, avoid the target object controlling the thinking mode of the large model by performing interactive operations, reduce the target object's learning cost and interactive operation execution cost, and improve the interactive experience.

[0051] In some embodiments, the reply content may include information in any data format, such as text, images, tables, etc. The embodiments of the present disclosure do not limit the specific information format of the reply content.

[0052] In some embodiments, the object-related features include interactive behavior features of the target object with respect to at least one of the following target information: historical reply information; and target knowledge information related to the knowledge domain of the input information.

[0053] Historical replies can be replies pushed to a target object during a historical time period. The target object's interactive behavior with respect to historical replies can include actions such as "like" and "favorite" responses, or input requests for historical replies. For example, in response to the historical reply "Yesterday's temperature was 18°C," the target object might input a request like "What will the temperature be tomorrow?"

[0054] According to an embodiment of the present disclosure, the interactive behavior characteristics for historical reply information can represent the contextual information content of the input information for the target object. By understanding the intention of the input information based on the interactive behavior characteristics for historical reply information, the target intention can more accurately represent the actual content reply intention of the current input information input by the target object.

[0055] In some embodiments, the target knowledge information related to the knowledge domain of the input information may include information such as papers, news, and video content that the target subject has viewed over a historical period. For example, if the input information is "How is the performance of vehicle A?", the target knowledge information related to the knowledge domain of the input information may include vehicle review videos, vehicle promotional advertisements, and so on. Interaction behavior features for the target knowledge information may, for example, represent characteristic data of interaction behaviors such as browsing time, comment content, and forwarding behavior.

[0056] According to the embodiments of the present disclosure, by semantically understanding the input information based on the interactive behavior characteristics executed for the target knowledge information related to the knowledge field of the input information, the target object's understanding of the knowledge field related to the input information can be more accurately learned to detect the target object's content quality requirements for the reply information, thereby avoiding outputting reply information with a lot of basic knowledge introduction and repetition for the knowledge field that the target object has a deep understanding of, and avoiding outputting relatively simple reply information for the knowledge field that the target object is not familiar with. In this way, based on the target intention, the large model can be controlled to perform the semantic understanding task based on the matching execution logic to output reply information that matches the actual needs of the target object.

[0057] In some embodiments, object interaction features may include historical response information; target knowledge information related to the knowledge domain of the input information; and the input information may be subjected to intent detection based on the historical response information, the target knowledge information, and other object-related features to obtain the target intent. The embodiments of the present disclosure will not be further described herein.

[0058] In some embodiments, understanding the intent of the input information based on the object-related features of the target object to obtain the target intent may also include: understanding the intent of the object-related features, the input information and the performance description information related to multiple candidate large models to obtain the target intent.

[0059] According to an embodiment of the present disclosure, the target intent also indicates that a large model among multiple candidate large models is used to perform semantic understanding tasks. The candidate large model can be a fine-tuned model for performing semantic understanding tasks in a specific field, or a model for performing semantic understanding tasks with specific demand attributes. For example, the multiple candidate large models are a first candidate large model for the automotive knowledge field and a second candidate large model for the ship knowledge field. The performance description information can be used to describe the knowledge field of the first candidate large model, the response information quality evaluation result, the average execution time of the semantic understanding task, and other information.

[0060] According to embodiments of the present disclosure, a deep learning model can be used to process object-related features, input information, and multiple performance descriptions to output a target intent. For example, a long short-term memory network algorithm can be used to process object-related features, input information, and multiple performance descriptions to capture temporal semantic features and output the target intent. The target intent can also include an indicator of a large model suitable for performing a semantic understanding task on the input information, so that the large model matching the target intent can be used to perform the semantic understanding task based on the corresponding execution logic to obtain a response message.

[0061] In one example, the target intent may specify the copywriting model among the medical domain model, the mechanical domain model, and the copywriting model as the designated model. The copywriting model is then used to execute the execution logic corresponding to the deep thinking mode to process the input message "Please write an essay about spring." The copywriting model can process the input message based on the execution logic corresponding to the deep thinking mode and output a prose copy as the response message.

[0062] In some embodiments, understanding the intent of object-related features, input information, and performance description information related to multiple candidate large models may include: performing feature fusion on object-related features, input information, and performance description information related to multiple candidate large models based on an attention mechanism to obtain target fusion features; and performing intent understanding based on the target fusion features to obtain target intent.

[0063] In one embodiment, feature fusion of object-related features, input information, and performance description information related to multiple candidate large models can be performed based on the Transformer model, so as to learn the degree of match between the input information and the preferences and historical interaction habits of the target object, and the performance description information of the candidate large models based on the attention mechanism, so that the large model indicated in the target intent can be adapted to perform semantic understanding tasks on the input information, and can perform semantic understanding tasks based on execution logic matching the target intent, so that the output reply information can more accurately match the actual content quality requirements and task execution time requirements of the target object, and further improve the interaction efficiency and reduce the execution efficiency of the semantic understanding task by automatically scheduling the large model and controlling the execution logic of the large model.

[0064] In one embodiment, a trained scheduling model can be used to process object-related features, input information, and performance description information related to multiple candidate large models using an attention mechanism to understand intent and obtain a target intent. The trained scheduling model can be trained based on labeled sample data, which can include sample object-related features, sample input information, and multiple performance descriptions of a sample object. The sample data can be labeled with a selection label for a sample large model from the candidate large models, and the sample large model can have a thinking mode label indicating "deep thinking mode" or "non-deep thinking mode." By using labels to perform diversified fitting annotation on the sample data, the trained scheduling model trained based on the sample data can simulate the target object's requirements for content quality and the large model's execution logic. The trained scheduling model can thus accurately and comprehensively determine the content quality requirements represented by the input information and the execution logic requirements of the large model. This allows for automated model scheduling and execution logic control based on the preferences represented by the target object's object-related features and the semantics of the input information, improving the match between the reply information and the target object.

[0065] In some embodiments, understanding the intent of the input information based on the object-related features of the target object to obtain the target intent may include: using multiple trained classification models to process the input information and object-related features respectively to obtain multiple initial model call intentions; and determining the model call intent in the target intent based on the multiple initial model call intentions, and the model call intent is used to determine the large model from multiple candidate large models.

[0066] In one embodiment, a trained classification model can be used to determine an initial model call intent by processing input information and object-related features. The initial model call intent can, for example, be the scores corresponding to multiple candidate large models. By fusing the scores corresponding to the multiple initial model call intents, the resulting target score is used as the model call intent. The candidate large model with the highest target score is then determined as the large model designated by the target intent.

[0067] In some embodiments, some of the candidate large models may be deep thinking models capable of performing semantic understanding tasks based on the execution logic of a deep thinking mode, while another portion of the candidate large models may be deep thinking models capable of performing semantic understanding tasks based on the execution logic of a non-deep thinking mode. Thus, by determining the large model for the semantic understanding task from the candidate large models, the execution logic of the large model corresponding to the content quality requirement expressed by the target intent can be determined, and the large model can be controlled to perform the semantic understanding task.

[0068] It should be noted that the classification model can be constructed based on any type of neural network algorithm, for example, the classification model can be constructed based on an attention network algorithm, a long short-term memory network algorithm, etc. The embodiments of the present disclosure do not limit the specific algorithm type for constructing the classification model.

[0069] Figure 3 The schematic diagram schematically shows the principle of the interaction method based on the large model according to the embodiment of the present disclosure.

[0070] like Figure 3 As shown, the multiple candidate large models may include the first candidate large model M301, the second candidate large model M302... to the nth candidate large model M30n. The input information 301 and the object-related features 302 of the target object, as well as the performance description information 303 related to each of the multiple candidate large models are input into multiple trained classifiers. For example, the input information 301, the object-related features 302 and the performance description information 303 can be input into the first classifier, the second classifier and the third classifier respectively. The first classifier, the second classifier and the third classifier each output the corresponding scores of the first candidate large model M301, the second candidate large model M302... to the nth candidate large model M30n as the initial model call intention. By fusing multiple initial model call intentions, the model call intention is obtained. The model call intention can indicate the first candidate large model M301 as the designated large model for performing the semantic understanding task. Wherein, n is an integer greater than 1.

[0071] By feeding the input information into the trained intent understanding model, a first mode intent within the target intent is output. This first mode intent can indicate that the content quality requirements of the reply information meet the preset quality requirements. The first candidate large model M301 performs a semantic understanding task on the input information 301 based on the execution logic representing the deep thinking mode corresponding to the first mode intent, and outputs a reply information 304.

[0072] According to an embodiment of the present disclosure, using a large model to perform a semantic understanding task on input information based on an execution logic that matches the target intent may include: based on a first mode intent, using a large model to perform semantic understanding on the input information to obtain a semantic understanding result; and using the large model to perform a semantic understanding task based on multiple demand information.

[0073] According to an embodiment of the present disclosure, the target intent may include a first mode intent, which indicates that the content quality requirements of the reply information meet preset quality requirements. For example, the first mode intent may indicate that the quality evaluation indicators of the reply content, such as word count and logic, are greater than or equal to a preset quality evaluation indicator threshold.

[0074] According to an embodiment of the present disclosure, the first mode intent may instruct the large model to perform a semantic understanding task on the input information based on a deep learning model. For example, the large model may perform semantic understanding on the input information to obtain a semantic understanding result including multiple pieces of requirement information. The multiple pieces of requirement information may match the requirement intent represented by the input information. For example, the multiple pieces of requirement information may be multiple sub-questions broken down after understanding the question text represented by the input information.

[0075] For example, if the input information is "What's the weather like today?", the multiple requirements in the semantic understanding results could include "What city is the target person located in?", "Which authorized or open query interface is used to query the city?", "What factors affect temperature, rainfall, or snowfall during multiple time periods today and in the future?", etc. This multiple requirement information can be used to represent the requirements required to execute the multiple steps of the semantic understanding task used to generate the response information, as output by the large model after semantic understanding of the input information.

[0076] According to the embodiments of the present disclosure, by utilizing a large model to perform semantic understanding tasks based on multiple demand information, the large model can be enabled to perform semantic understanding tasks according to a deep thinking mode for multiple demand information mined from input information and object-related features. Thus, when the target intent obtained through identification is to indicate the need to provide reply information with higher content quality, the first mode intention instructs the large model to decompose or mine the demand intention of the input information through the object preference and intention represented by the object-related features, thereby achieving deep thinking on the demand represented by the input information of the target object, and utilizing the large model to perform semantic understanding tasks based on multiple demand information that matches the input information to achieve the execution of semantic understanding tasks based on a deep thinking mode, so that the reply information can more accurately meet the actual needs of the user, thereby improving the interaction efficiency and interaction experience of the target object.

[0077] In some embodiments, using a large model to perform semantic understanding on input information to obtain multiple demand information that matches the demand intention represented by the input information can also include: using a large model to perform semantic understanding on input information and object-related features to obtain multiple demand prompt words for multiple subtasks in the semantic understanding task.

[0078] According to an embodiment of the present disclosure, a subtask within a semantic understanding task for input information may include utilizing a large model to semantically understand the demand information and outputting a response to the demand information. For example, a subtask may be a process for semantically understanding and generating a response to the multiple demand information enclosed by ", such as "where is the city where the target object is located," "which authorized or open query interface is used to query the city," "what factors affect the temperature, rainfall, or snowfall during multiple time periods today and in the future," and so on, including "where is the city where the target object is located."

[0079] According to an embodiment of the present disclosure, the demand prompt words can be used to control the subtasks in the semantic understanding task to be executed according to the demand conditions that match the demand intentions of the target object. For example, for the demand text content "where is the city where the target object is located" in the demand information, the demand prompt words can be "the city location is represented by the longitude and latitude coordinate range based on open authorization". This enables the large model to more accurately execute the subtasks by processing the information content and demand prompt words of the demand information, so as to output the execution results of the subtasks that can meet the demand conditions of the deep thinking mode, and then generate reply information based on the execution results of multiple subtasks, so that the reply information can meet the actual needs of the target object.

[0080] In some embodiments, using a large model to perform a semantic understanding task based on multiple requirement information also includes: using a gating network of the large model to process the multiple requirement information so as to determine multiple target expert networks related to multiple subtasks from multiple expert networks.

[0081] According to an embodiment of the present disclosure, a large model may include a gating network and multiple expert networks. The gating network and the multiple expert networks may be model structure layers constructed based on a deep learning algorithm. The gating network is used to determine a target expert network that can execute subtasks for the demand information by semantically understanding multiple demand information. The target expert network can match the demand conditions or demand intentions such as the output file format, professional field, number of tokens, etc. represented by the demand information, thereby activating the model parameters of the target expert network in the multiple expert networks in the large model through the gating network to execute subtasks, so as to reduce the model parameter requirements required for the large model to execute multiple subtasks and reduce the computational overhead of the computing device to execute semantic understanding tasks. At the same time, by selecting the activated expert network to execute subtasks based on the demand information obtained by intention mining of the input information through the gating network, the large model can execute multiple subtasks more accurately through multiple expert network models that match the demand information in the deep thinking mode, so as to improve the accuracy and content quality of the reply information determined by the execution results of the task subtasks.

[0082] In one embodiment, the gated network of the large model can also be used to process the information content and demand prompt words in multiple demand information to determine multiple target expert networks. Therefore, based on the demand conditions represented by the demand prompt words, the target expert network that meets the demand intention represented by the demand information can be more accurately activated to improve the accuracy of the subtask execution results and further match the content quality of the reply information with the actual needs of the target object.

[0083] In some embodiments, a large model is used to perform a semantic understanding task based on multiple demand information, including: using the first target expert network of the large model to perform a first subtask based on the first demand information to obtain a first execution result; using the second target expert network of the large model to perform a second subtask based on the second demand information and the first execution result to obtain a second execution result, wherein the reply information is determined based on the first execution result and the second execution result.

[0084] In one embodiment, the first target expert network processes the demand information content in the first demand information "where is the city where the target object is located" and the demand prompt word "the city location is represented by the longitude and latitude coordinate range based on the open authorization", and outputs the first execution result of the first execution subtask, which is the authorized longitude and latitude coordinate area of ​​the city where the target object is located. The second target expert network is used to process the demand information content in the first demand information "which authorized or open query interface is used to execute the query in the city" and the first execution result to determine the query interface address as the second execution result. By utilizing the target expert network to execute the subtask to be executed later according to the demand information based on the execution results of the subtask that has been executed, it is possible to control multiple target expert networks based on multiple demand information to execute semantic understanding tasks according to the execution logic of multiple subtasks, so as to realize that the large model executes the semantic understanding task based on the execution logic corresponding to the deep thinking mode, thereby improving the information content quality of the reply information.

[0085] In some embodiments, the semantic understanding task executed based on the execution logic corresponding to the first mode intention can display multiple requirement information and the execution results corresponding to the multiple requirement information in the display interface to show the deep thinking process of the large model for the input information, so as to prompt the target object's reply information to the input information as the generated result of the large model executed through the deep thinking mode, so that the target object can clearly understand the execution process of the semantic understanding task.

[0086] Figure 4 A schematic diagram of the principle of an interaction method based on a large model provided according to another embodiment of the present disclosure is schematically shown.

[0087] like Figure 4As shown, the large model for performing the semantic understanding task can include a gating network M411 and multiple expert networks M420. The multiple expert networks M420 include a first expert network M421, a second expert network M422, and so on to an nth expert network M2n. The gating network M411 is used to process the first requirement information 401 and the second requirement information 402 among the multiple requirement information, and outputs two target expert networks from the multiple expert networks M420 for processing the first requirement information 401 and the second requirement information 402, respectively. The two target expert networks are the first expert network M421 and the second expert network M422. The gating network M411 can also transmit the first requirement information 401 and the second requirement information to the first expert network M421 and the second expert network M422, so that the multiple target expert networks can perform multiple subtasks in the semantic understanding task based on the multiple requirement information and output multiple execution results. By fusing the multiple execution results, a reply message 403 can be output.

[0088] In some embodiments, using the large model to perform a semantic understanding task on the input information based on the execution logic that matches the target intent may include: based on the second mode intent, performing a semantic understanding task on the input information according to the model parameters of the expert network in the large model to obtain reply information.

[0089] According to an embodiment of the present disclosure, the second mode intentionally indicates that the content quality requirement of the reply message does not meet the preset quality requirement condition. For example, the second mode intention may indicate that the quality evaluation indicators of the reply content, such as word count and logic, are less than the preset quality evaluation indicator threshold.

[0090] According to an embodiment of the present disclosure, the second mode intention can instruct the large model to perform a semantic understanding task on the input information based on a non-deep learning mode. For example, the large model can perform semantic understanding on the input information "1+1 equals how many" and output the response information "1+1=2".

[0091] In one embodiment, the input information "1+1 equals something" can be processed based on a target expert network within a large model that has mathematical calculation capabilities. This allows the target expert network's model parameters to be activated to perform a semantic understanding task on the input information "1+1 equals something," outputting the response information "1+1=2." This allows the semantic understanding task to be performed by activating only the model parameters of the target expert network within the large model, reducing the data size of the model parameters involved in performing the semantic understanding task and the computational overhead of the computing device. Furthermore, the response information can be generated based on the response information's content quality requirements, improving the degree of match between the response quality and the target's intent, and enhancing response accuracy and interaction accuracy.

[0092] Figure 5The application scenario diagram of the large model-based interaction method according to an embodiment of the present disclosure is schematically shown.

[0093] like Figure 5 As shown, display interface 500 may display the target subject's first input message 501 as "Medication for relieving frequent headaches." By understanding the intent of first input message 501 based on the target subject's object-related features, it can be determined that the target subject's first input message 501 is intended to relieve headaches. The determined target intent may include a first mode intent instructing the target subject to quickly purchase the relevant medication. The large model can process first input message 501 based on the execution logic corresponding to the deep thinking mode that matches the first mode intent, resulting in a first response message 510.

[0094] After browsing the first reply information 510, the target object can input the second input information 502 "What are the main ingredients of drug A?". Based on the target object's deep understanding of the effects of various compounds in the relevant drug field knowledge represented in the object interaction feature, the second input information 502 can be understood based on the object interaction feature to determine that the target intention can represent the second mode intention. The second input information 502 is subjected to the semantic understanding task by performing the execution logic of the non-deep thinking mode that matches the second mode intention through the large model, and the second reply information 520 can be output relatively quickly. This allows the target object to quickly understand that the main ingredients of drug A are compound A and compound B based on knowledge accumulation, avoids information redundancy in outputting explanation content for the compounds, and ensures that the reply information meets the actual content quality requirements of the target object.

[0095] According to an embodiment of the present disclosure, the interaction method based on the large model may further include: when the target intent also includes a retrieval intent representing the need to perform a retrieval task, calling a retrieval resource to perform a retrieval based on the input information to obtain a retrieval result.

[0096] According to an embodiment of the present disclosure, the search resource may include an authorized search resource interface, and the search result may be obtained by calling the search resource interface to perform information search according to input information.

[0097] In one embodiment, a search resource interface can be called to perform a search based on the input information "What is the price of vehicle A?" The search results obtained may include the sales prices of various models of vehicle A, discount information for vehicle A, and the time periods corresponding to the discount information. By utilizing a large model to process the search results and input information, the discounted sales prices for vehicle A in various time periods can be output as response information.

[0098] According to an embodiment of the present disclosure, reply information is determined by performing a semantic understanding task on input information based on retrieval results and target intent using a large model.

[0099] For example, the large model can be used to represent the first mode intention based on the target intention, and the retrieval results and input information can be processed through the execution logic corresponding to the deep thinking mode to obtain the reply information.

[0100] For another example, the large model can be used to represent the second mode intention based on the target intention, and the retrieval results and input information can be processed through the execution logic corresponding to the non-deep thinking mode to obtain reply information.

[0101] In some embodiments, calling the retrieval resource to perform data retrieval based on the input information may also include: rewriting the input information using the large model based on the prompt words determined based on the target intent to obtain updated input information; calling the retrieval resource to perform data retrieval based on the updated input information.

[0102] According to embodiments of the present disclosure, the prompt words determined based on the target intent can include prompt words that match the content quality requirements indicated by the first or second mode intent. This controls the large model to rewrite the input information so that the updated input information avoids defects such as unclear semantics and an overly broad search scope. Data retrieval can then be performed based on the updated input information, resulting in retrieval results that meet the semantic understanding task.

[0103] For example, if the input information is "price of car A," the macro model can be used to rewrite the input information based on the first pattern intent to obtain "the corresponding price of car A1 and car A2." This allows a search based on "the corresponding price of car A1 and car A2" to obtain search results.

[0104] Figure 6 A block diagram of an interaction device of a large model according to an embodiment of the present disclosure is schematically shown.

[0105] like Figure 6 As shown, the large model-based interaction device 600 includes: a receiving module 610 , a first obtaining module 620 , a second obtaining module 630 and a pushing module 640 .

[0106] The receiving module 610 is configured to receive input information of a target object.

[0107] The first obtaining module 620 is used to understand the intention of the input information based on the object-related features of the target object to obtain the target intention, which represents the content quality requirements for the reply information.

[0108] The second acquisition module 630 is used to use the large model to perform a semantic understanding task on the input information based on the execution logic that matches the target intention to obtain reply information.

[0109] The push module 640 is used to push the reply information to the target object.

[0110] According to an embodiment of the present disclosure, the second obtaining module includes: a first obtaining unit and a second obtaining unit.

[0111] The first acquisition unit is used to perform semantic understanding on the input information based on the first mode intention using a large model to obtain a semantic understanding result, wherein the semantic understanding result includes multiple demand information that matches the demand intention represented by the input information, and the first mode intention indicates that the content quality requirement of the reply information meets the preset quality requirement conditions.

[0112] The second obtaining unit is used to use the large model to perform semantic understanding tasks based on multiple demand information to obtain response information.

[0113] According to an embodiment of the present disclosure, a plurality of demand information is used for a plurality of subtasks in a semantic understanding task; wherein the second obtaining unit includes: a first obtaining subunit and a second obtaining subunit.

[0114] The first obtaining subunit is used to use the first target expert network of the large model to execute the first subtask based on the first requirement information to obtain a first execution result.

[0115] The second obtaining subunit is used to use the second target expert network of the large model to perform the second subtask based on the second demand information and the first execution result to obtain the second execution result, wherein the reply information is determined based on the first execution result and the second execution result.

[0116] According to an embodiment of the present disclosure, the second obtaining module includes a third obtaining unit.

[0117] The third obtaining unit is used to perform a semantic understanding task on the input information based on the second mode intention and the model parameters of the expert network in the large model to obtain reply information, wherein the second mode intention indicates that the content quality requirement of the reply information does not meet the preset quality requirement conditions.

[0118] According to an embodiment of the present disclosure, the second obtaining unit further includes a first determining subunit.

[0119] The first determination subunit is used to process multiple demand information using the gating network of the large model, so as to determine multiple target expert networks related to multiple subtasks from multiple expert networks.

[0120] According to an embodiment of the present disclosure, the second obtaining module includes a fourth obtaining unit.

[0121] The fourth obtaining unit is used to use the large model to perform semantic understanding on the input information and object-related features, and obtain multiple requirement prompt words for multiple subtasks in the semantic understanding task.

[0122] According to an embodiment of the present disclosure, the first obtaining module includes a target intention obtaining unit.

[0123] The target intention acquisition unit is used to understand the intention of object-related features, input information and performance description information related to multiple candidate large models to obtain the target intention. The target intention also indicates the large model among the multiple candidate large models for performing semantic understanding tasks.

[0124] According to an embodiment of the present disclosure, the target intention obtaining unit includes: a target fusion feature obtaining subunit and a target intention obtaining subunit.

[0125] The target fusion feature acquisition subunit is used to fuse object-related features, input information and performance description information related to multiple candidate large models based on the attention mechanism to obtain target fusion features.

[0126] The target intention acquisition subunit is used to understand the intention based on the target fusion features and obtain the target intention.

[0127] According to an embodiment of the present disclosure, the first obtaining module includes: an initial model calling intention obtaining unit and a large model determining unit.

[0128] The initial model calling intention obtaining unit is used to use multiple trained classification models to process input information and object-related features respectively to obtain multiple initial model calling intentions.

[0129] The large model determination unit is used to determine the model call intention in the target intention based on multiple initial model call intentions, and the model call intention is used to determine the large model from multiple candidate large models.

[0130] According to an embodiment of the present disclosure, the object-related features include interactive behavior features of the target object with respect to at least one of the following target information: historical reply information; target knowledge information related to the knowledge field of the input information.

[0131] According to an embodiment of the present disclosure, the large model-based interaction device further includes a retrieval result obtaining module.

[0132] The retrieval result acquisition module is used to call the retrieval resources to perform retrieval based on the input information and obtain the retrieval results when the target intent also includes the retrieval intent that represents the need to perform the retrieval task. The reply information is determined by using the large model to perform the semantic understanding task on the input information based on the retrieval results and the target intent.

[0133] According to an embodiment of the present disclosure, the search result obtaining module includes: an updating unit and a calling unit.

[0134] The updating unit is used to rewrite the input information using the large model based on the prompt word determined by the target intention to obtain updated input information.

[0135] The calling unit is used to call the retrieval resource to perform data retrieval based on the updated input information.

[0136] Figure 7 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0137] In the embodiments of the present disclosure, Figure 7 As shown, the AI ​​agent 700 may include an input module 710 , a processing module 720 and an output module 730 .

[0138] Input module 710, for receiving input information;

[0139] A processing module 720 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the large model-based interaction method provided according to an embodiment of the present disclosure;

[0140] The output module 730 is used to output the output information obtained by the processing module.

[0141] According to an embodiment of the present disclosure, the input module 710 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., a user or the external environment) and converting it into a format that can be understood and processed by the AI ​​agent 700. The input module 710 is the primary link for the AI ​​agent 700 to interact with the outside world, enabling the AI ​​agent 700 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.

[0142] In an example, the input module 710 may input the input information, object-related features, etc. described above.

[0143] In the example, the processing module 720 is the core support for the AI ​​agent 700 to handle complex tasks. The processing module 720 can execute the large model-based interaction method described above.

[0144] In this example, the performance of processing module 720 may be closely related to the large model underlying AI agent 700. To fully leverage the capabilities of the large model, the internal structure of processing module 720 may be designed to be highly configurable and extensible to handle a variety of different types of tasks and requirements in real-world scenarios.

[0145] In the example, after the AI ​​agent 700 obtains the required voice, the processing module 720 can use the large model to perform semantic understanding tasks on the input information, obtain reply information, and pass the reply information to the output module 730.

[0146] Understandably, while the large model possesses excellent language understanding and generation capabilities, like humans, it can only perform limited tasks without tools. However, once AI Agent 700 is empowered with tool-based capabilities, it can perform tasks such as mathematical calculations using a calculator, data analysis using Python, and weather forecasting using search engines.

[0147] In an example, the output module 730 may output the reply information described above.

[0148] The AI ​​agent 700 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0149] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0150] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0151] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.

[0152] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.

[0153] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the large model-based interaction method of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0154] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0155] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the large model-based interaction method. For example, in some embodiments, the large model-based interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the large model-based interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the large model-based interaction method in any other suitable manner (e.g., via firmware).

[0157] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0162] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0163] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0164] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An interactive method based on a large model, comprising: Receive input information from the target object; Performing intent understanding on the input information based on object-related features of the target object to obtain a target intent, wherein the target intent represents a content quality requirement for the reply information; Using the large model to perform a semantic understanding task on the input information based on execution logic matching the target intent, to obtain the reply information; Push the reply information to the target object.

2. The method according to claim 1, wherein The using of the large model to perform a semantic understanding task on the input information based on execution logic matching the target intent includes: Based on the first mode intention, using the large model to perform semantic understanding on the input information to obtain a semantic understanding result, wherein the semantic understanding result includes multiple demand information that matches the demand intention represented by the input information, and the first mode intention indicates that the content quality requirement of the reply information meets a preset quality requirement condition; and The semantic understanding task is performed based on the plurality of demand information using the large model.

3. The method according to claim 2, wherein: The plurality of requirement information are used for a plurality of subtasks in the semantic understanding task; The step of using the large model to perform the semantic understanding task based on the plurality of demand information includes: Utilizing the first target expert network of the large model to execute a first subtask based on first requirement information, and obtaining a first execution result; The second target expert network of the large model is used to execute a second subtask based on second requirement information and the first execution result to obtain a second execution result, wherein the reply information is determined based on the first execution result and the second execution result.

4. The method according to claim 3, wherein: The performing of the semantic understanding task based on the plurality of demand information using the large model further includes: The gating network of the large model is used to process the plurality of requirement information so as to determine a plurality of target expert networks related to the plurality of subtasks from a plurality of expert networks.

5. The method according to claim 2, wherein: The use of the large model to perform semantic understanding on the input information to obtain multiple demand information that matches the demand intent represented by the input information includes: The large model is used to perform semantic understanding on the input information and the object-related features to obtain multiple requirement prompt words for multiple subtasks in the semantic understanding task.

6. The method according to claim 1, wherein The using of the large model to perform a semantic understanding task on the input information based on execution logic matching the target intent includes: Based on the second mode intention, a semantic understanding task is performed on the input information according to the model parameters of the expert network in the large model to obtain the reply information, wherein the second mode intention indicates that the content quality requirement of the reply information does not meet the preset quality requirement conditions.

7. The method according to claim 1, wherein The performing intention understanding on the input information based on the object-related features of the target object to obtain the target intention includes: The object-related features, the input information and the performance description information related to multiple candidate large models are used to understand the intention and obtain the target intention. The target intention also indicates that the large model among the multiple candidate large models is used to perform the semantic understanding task.

8. The method according to claim 7, wherein: The performing intention understanding on the object-related features, the input information, and the performance description information related to the plurality of candidate large models includes: Performing feature fusion on the object-related features, the input information, and performance description information related to multiple candidate large models based on an attention mechanism to obtain target fusion features; The target intention is understood based on the target fusion features to obtain the target intention.

9. The method according to claim 1, wherein The performing intention understanding on the input information based on the object-related features of the target object to obtain the target intention includes: Using multiple trained classification models to process the input information and the object-related features respectively to obtain multiple initial model call intentions; and The model calling intention in the target intention is determined based on the multiple initial model calling intentions, and the model calling intention is used to determine the large model from multiple candidate large models.

10. The method according to claim 1, wherein The object-related features include interactive behavior features of the target object with respect to at least one of the following target information: Historical reply information; Target knowledge information related to the knowledge domain of the input information.

11. The method according to claim 1, wherein The method further comprises: In the case where the target intent also includes a retrieval intent that represents the need to perform a retrieval task, the retrieval resource is called to perform a retrieval based on the input information to obtain a retrieval result, wherein the reply information is determined by performing a semantic understanding task on the input information based on the retrieval result and the target intent using a large model.

12. The method according to claim 11, wherein The calling of the retrieval resource to perform data retrieval based on the input information includes: Based on the prompt word determined by the target intention, the input information is rewritten using the large model to obtain updated input information; The retrieval resource is called to perform data retrieval based on the updated input information.

13. An interactive device based on a large model, comprising: A receiving module, used for receiving input information of a target object; A first obtaining module is configured to understand the intent of the input information based on the object-related features of the target object to obtain a target intent, wherein the target intent represents a content quality requirement for the reply information; A second acquisition module is configured to use the large model to perform a semantic understanding task on the input information based on an execution logic that matches the target intent, thereby obtaining the reply information; A push module is used to push the reply information to the target object.

14. The device according to claim 13, wherein The second obtaining module includes: a first obtaining unit, configured to perform semantic understanding of the input information using the large model based on a first pattern intent, to obtain a semantic understanding result, wherein the semantic understanding result includes a plurality of demand information that matches the demand intent represented by the input information, and the first pattern intent indicates that the content quality requirement of the reply information satisfies a preset quality requirement condition; and The second obtaining unit is used to use the large model to perform the semantic understanding task based on the multiple demand information to obtain the reply information.

15. The device according to claim 14, wherein The plurality of requirement information are used for a plurality of subtasks in the semantic understanding task; Wherein, the second obtaining unit includes: A first obtaining subunit is configured to use the first target expert network of the large model to execute a first subtask based on first requirement information to obtain a first execution result; The second obtaining subunit is used to use the second target expert network of the large model to perform a second subtask based on second demand information and the first execution result to obtain a second execution result, wherein the reply information is determined based on the first execution result and the second execution result.

16. The device according to claim 13, wherein The second obtaining module includes: The third obtaining unit is used to perform a semantic understanding task on the input information based on the second mode intention and the model parameters of the expert network in the large model to obtain the reply information, wherein the second mode intention indicates that the content quality requirement of the reply information does not meet the preset quality requirement conditions.

17. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method according to any one of claims 1 to 12 by calling the large model; An output module is used to output the output information obtained by the processing module.

18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 12.

20. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Interaction method and related device

    CN113703883A

  • Human-computer interaction method, device and equipment and storage medium

    CN117421398A

  • Universal terminal perception interaction processing method, control device and storage medium

    CN118034637A

  • Interaction method and device based on large model, electronic equipment and storage medium

    CN118897925A

  • Information interaction method and device, electronic equipment and storage medium

    CN119129646A