Large model-based interaction method and device, intelligent agent and storage medium
By using object-related features and intent understanding as a basis, and by adjusting the execution logic of the large model, the problem of content quality matching in large language model interaction is solved, and efficient and accurate user interaction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-03-17
AI Technical Summary
Large language models struggle to flexibly output responses that match users' quality needs during interactions, resulting in high complexity and learning costs in interactive operations.
By receiving input information from the target object, the system performs intent understanding based on object-related features, utilizes a large model to execute semantic understanding tasks based on execution logic that matches the target intent, generates response information, and automatically adjusts the large model's thinking mode and task execution mode to match user needs.
It achieves accurate matching of response information with user needs, reduces the user's learning cost and the complexity of interaction, and improves interaction efficiency and experience.
Smart Images

Figure CN120706575B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning, intelligent search, smart customer service, and smart healthcare. Background Technology
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) can be used to perform semantic understanding on user input such as voice and text, and generate response content that meets their needs. Summary of the Invention
[0003] This disclosure provides a large-model-based interaction method, device, intelligent agent, electronic device, and storage medium.
[0004] According to one aspect of this disclosure, a large-model-based interaction method is provided, comprising: receiving input information from a target object; performing intent understanding on the input information based on object-related features of the target object to obtain a target intent, wherein the target intent represents the content quality requirements for response information; using a large-model to perform a semantic understanding task on the input information based on execution logic matching the target intent to obtain response information; and pushing the response information to the target object.
[0005] According to another aspect of this disclosure, a large-model-based interactive device is provided, comprising: a receiving module for receiving input information from a target object; a first obtaining module for performing intent understanding on the input information based on object-related features of the target object to obtain a target intent, wherein the target intent represents the content quality requirements for response information; a second obtaining module for performing a semantic understanding task on the input information using a large model based on execution logic matching the target intent to obtain response information; and a pushing module for pushing the response information to the target object.
[0006] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and executing the method provided in the embodiments of this disclosure by calling the large model; and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to an embodiment of this disclosure.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform a method provided according to embodiments of this disclosure.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 The illustration schematically shows an exemplary system architecture for applying large-model-based interaction methods and apparatus according to embodiments of the present disclosure;
[0013] Figure 2 A flowchart illustrating a large-model-based interaction method according to an embodiment of the present disclosure is shown schematically.
[0014] Figure 3 The schematic diagram illustrates the principle of a large-model-based interaction method according to an embodiment of the present disclosure;
[0015] Figure 4 The illustration shows a schematic diagram of the principle of a large-model-based interaction method provided according to another embodiment of the present disclosure;
[0016] Figure 5 The diagram illustrates an application scenario of the interaction method based on a large model according to an embodiment of the present disclosure.
[0017] Figure 6 A block diagram of an interactive device for a large model according to an embodiment of the present disclosure is shown schematically;
[0018] Figure 7 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and
[0019] Figure 8 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure based on a large model interaction method is shown. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0022] The inventors discovered that large language models can perform semantic understanding by processing user input and generate responses that match the user's needs based on their powerful semantic understanding and generation capabilities. Furthermore, related platforms can provide users with options such as "deep thinking" to control the quality of the responses output by the large language model. However, during actual user interaction, large language models struggle to flexibly match the quality of user input with their desired response quality. This forces users to invest significant learning effort in controlling response quality, resulting in high complexity and learning costs in the interaction process.
[0023] Embodiments of this disclosure provide a large-model-based interaction method, device, intelligent agent, electronic device, and storage medium. The large-model-based interaction method includes: receiving input information from a target object; performing intent understanding on the input information based on the object-related features of the target object to obtain a target intent, whereby the target intent represents the content quality requirements for response information; using a large model to perform a semantic understanding task on the input information based on execution logic matching the target intent to obtain response information; and pushing the response information to the target object.
[0024] According to embodiments of this disclosure, by understanding the target object's input information through object-related features to determine the target intent, the target intent can more accurately represent the target object's requirements for the content quality level of the response information output by the large model after processing the input information. By controlling the large model to perform semantic understanding tasks based on execution logic that matches the target intent, the execution logic of the large model for performing semantic understanding tasks on the input information is automatically adjusted according to the target object's preferences or habits. This allows for flexible adjustment of the large model's thinking mode and task execution mode, thereby automatically adjusting the content quality of the response information to match the intent expressed by the target object's input information. This avoids large deviations between the content quality of the response information and the target object's needs, or excessively long execution logic of the large model that could negatively impact user experience, thus improving interaction efficiency and response accuracy.
[0025] Figure 1 The illustration schematically depicts an exemplary system architecture for applying large-model-based interaction methods and apparatus according to embodiments of the present disclosure.
[0026] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture for applying the large-model-based interaction method and apparatus may include a terminal device, but the terminal device may implement the large-model-based interaction method and apparatus provided by embodiments of this disclosure without interacting with a server.
[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0029] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0031] It should be noted that the large-model-based interaction method provided in this disclosure can generally be executed by terminal devices 101, 102, or 103. Accordingly, the large-model-based interaction device provided in this disclosure can also be disposed in terminal devices 101, 102, or 103.
[0032] Alternatively, the large-model-based interaction method provided in this disclosure can generally be executed by server 105. Correspondingly, the large-model-based interaction device provided in this disclosure can generally be located in server 105. The large-model-based interaction method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the large-model-based interaction device provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0034] It should be noted that the information obtained in any embodiment of this disclosure, including but not limited to object-related characteristics, was obtained under the authorization of the relevant users or organizations. Furthermore, the intended use of the obtained information was disclosed in advance, and necessary encryption or de-identification measures were adopted, complying with relevant laws and regulations and not violating public order and good morals.
[0035] Figure 2 A flowchart illustrating a large-model-based interaction method according to an embodiment of the present disclosure is shown schematically.
[0036] like Figure 2 As shown, the interaction method based on the large model includes operations S210~S240.
[0037] In operation S210, input information from the target object is received.
[0038] In operation S220, the input information is understood based on the object-related features of the target object to obtain the target intent.
[0039] In operation S230, a large model is used to perform a semantic understanding task on the input information based on the execution logic that matches the target intent, and the response information is obtained.
[0040] In operation S240, a response message is pushed to the target object.
[0041] According to embodiments of this disclosure, the input information can be data of any modality, such as text or speech, input by the target object. For example, the target object can input a request, such as "Please check today's weather," via voice input. The input information can be obtained by performing speech recognition on the requested speech to obtain the corresponding text or vector.
[0042] According to embodiments of this disclosure, a large model, or large language model, can be understood as a model built based on deep learning algorithms that possesses natural language processing capabilities and response information generation capabilities. A large model can have billions or even hundreds of billions of parameters. These parameters enable the model to capture and understand the complexity and diversity of language, and to output generative content of arbitrary modalities such as text and images based on a large number of model parameters.
[0043] According to embodiments of this disclosure, object-related features can represent features related to the target object's own attributes, interactive behaviors, etc. For example, object-related features can be the target object's interest and preference features, resource content features that have undergone interactive operations, etc.
[0044] According to embodiments of this disclosure, intent understanding of input information based on object-related features of a target object can include processing the input information and object-related features using a neural network algorithm to obtain the target intent. For example, a convolutional neural network algorithm can be used to process object-related features and input information to obtain the target intent. However, this is not limited to this; other types of algorithms can also be used to process object-related features and input information to obtain the target intent. Embodiments of this disclosure do not limit the specific type of algorithm used to process object-related features and input information.
[0045] According to embodiments of this disclosure, the target intent represents the content quality requirements for the response information. Content quality requirements can represent any level of requirement related to content quality, such as the richness of the content, the number of knowledge points, and the logical coherence of the response information. For example, content quality requirements can be the level of quality requirement for the logic and vividness of a press release for the input information "A press release about the sports meet needs to be written."
[0046] In some embodiments, the content quality requirements of the response information representing the target intent can be used to control the task execution logic and thinking mode of the large model. The execution logic and thinking mode of the semantic task can be adjusted by leveraging the large model based on the content quality requirements of the target intent representation. This allows the large model to output high-quality response information through execution logic corresponding to deep thinking, or to quickly generate response information matching the content quality requirements through execution logic corresponding to non-deep thinking.
[0047] For example, if the target intent indicates that the content quality requirements meet the preset quality requirements, the large model can perform semantic understanding tasks based on the deep thinking patterns corresponding to the target intent, so as to output high-quality response information.
[0048] For example, if the target intent indicates that the content quality requirements do not meet the preset quality requirements, the large model can perform semantic understanding tasks based on a non-deep thinking mode corresponding to the target intent, so as to quickly output response information.
[0049] In some embodiments, the first execution time of the large model performing the semantic understanding task based on deep thinking patterns is longer than the second execution time of the large model performing the semantic understanding task based on deep thinking patterns. For example, the first execution time for determining the response information is 1 second longer than the second execution time for determining the response information.
[0050] According to embodiments of this disclosure, by performing intent detection on input information based on object-related features and controlling the execution logic of the large model according to the content quality requirements of the response information represented by the target intent, it is possible to automatically adjust the task execution logic and thinking mode of the large model based on the input information and object-related features of the target object without the target object's awareness, and generate response information that matches the actual content quality requirements of the target object. This avoids the target object controlling the thinking mode of the large model through interactive operations, reduces the learning cost and interactive operation execution cost of the target object, and improves the interactive experience.
[0051] In some embodiments, the response content may include information in any data format such as text, images, and tables. The embodiments of this disclosure do not limit the specific information format of the response content.
[0052] In some embodiments, object-related features include interactive behavior features of the target object in response to at least one of the following target information: historical response information; target knowledge information related to the knowledge domain of the input information.
[0053] Historical response information can be response information pushed to the target audience within a historical time period. The target audience's interactive behavior characteristics in response to historical response information can represent interactive behaviors such as liking or saving historical response information, or it can also represent the request information entered in response to historical response information, such as the historical request information "What will the temperature be tomorrow?" in response to the historical response information "Yesterday's temperature was 18℃".
[0054] According to embodiments of this disclosure, the interactive behavior characteristics of historical response information can represent the contextual information content of the input information for the target object. By understanding the intent of the input information based on the interactive behavior characteristics of historical response information, the target intent can more accurately represent the actual content response intent of the target object when inputting the current input information.
[0055] In some embodiments, the target knowledge information related to the knowledge domain of the input information may include information such as papers, news, and video content browsed by the target object during a historical period. For example, if the input information is "How is the performance of vehicle A?", the target knowledge information related to the knowledge domain of the input information could be vehicle review videos, vehicle advertisements, etc. Interactive behavior characteristics targeting the target knowledge information may represent feature data such as browsing time, comment content, and forwarding behavior.
[0056] According to embodiments of this disclosure, semantic understanding of input information is achieved by performing interactive behavior features based on target knowledge information related to the knowledge domain of the input information. This allows for a more accurate learning of the target object's understanding of the knowledge domain related to the input information, thereby detecting the target object's content quality requirements for the response information. This avoids outputting repetitive, basic information for knowledge domains the target object has in-depth knowledge of, and also avoids outputting superficial response information for knowledge domains the target object is unfamiliar with. Therefore, the large model can be controlled based on the target intent to perform semantic understanding tasks according to matching execution logic, outputting response information that matches the actual needs of the target object.
[0057] In some embodiments, object interaction features may include historical response information; target knowledge information related to the knowledge domain of the input information can be used to detect intent in the input information based on historical response information, target knowledge information, and other object-related features to obtain the target intent. Embodiments of this disclosure will not be described in detail here.
[0058] In some embodiments, the intention understanding of the input information based on the object-related features of the target object to obtain the target intention may further include: performing intention understanding on the object-related features, the input information, and the performance description information related to multiple candidate large models to obtain the target intention.
[0059] According to embodiments of this disclosure, the target intent further indicates a large model among multiple candidate large models for performing semantic understanding tasks. The candidate large model may be a fine-tuned model for performing a domain-specific semantic understanding task, or a model for performing a semantic understanding task with specific requirement attributes. For example, the multiple candidate large models may be a first candidate large model for the automotive knowledge domain and a second candidate large model for the naval knowledge domain. Performance description information may be used to describe the knowledge domain, response information quality assessment results, average execution time of the semantic understanding task, and other information of the first candidate large model.
[0060] According to embodiments of this disclosure, a deep learning model can be used to process object-related features, input information, and multiple performance descriptions to output a target intent. For example, a Long Short-Term Memory (LSTM) network algorithm can be used to process object-related features, input information, and multiple performance descriptions to capture temporal semantic features and output a target intent. This target intent may also include an indicator of a large model suitable for performing semantic understanding tasks on the input information, so as to utilize a large model matching the target intent to perform semantic understanding tasks based on corresponding execution logic to obtain response information.
[0061] In one example, the target intent could specify the copywriting model as the designated large model from the medical field, mechanical field, and copywriting model. The copywriting model would then be used to process the input information "Please write an essay about spring" based on the execution logic corresponding to the deep thinking mode. The copywriting model could then process the input information based on the execution logic corresponding to the deep thinking mode, outputting the essay as the response information.
[0062] In some embodiments, intent understanding of object-related features, input information, and performance description information related to multiple candidate large models may include: performing feature fusion on object-related features, input information, and performance description information related to multiple candidate large models based on an attention mechanism to obtain target fused features; and performing intent understanding based on the target fused features to obtain target intent.
[0063] In one embodiment, a Transformer model can be used to fuse object-related features, input information, and performance descriptions associated with multiple candidate large models. This allows for the learning of the matching degree between the input information and the target object's preferences and historical interaction habits, as well as the performance descriptions of the candidate large models, based on an attention mechanism. This enables the large model indicated in the target intent to be adapted to perform semantic understanding tasks related to the input information. Furthermore, the semantic understanding task can be executed based on execution logic that matches the target intent, allowing the output response information to more accurately match the target object's actual content quality requirements and task execution time requirements. Additionally, by automating the scheduling of large models and controlling their execution logic, interaction efficiency can be improved, while reducing the execution efficiency of the semantic understanding task.
[0064] In one embodiment, a trained scheduling model can be used to process attention-based attention mechanisms to understand the target intent by analyzing object-related features, input information, and performance descriptions associated with multiple candidate large models. The trained scheduling model can be trained on labeled sample data, which may include object-related features of the sample object, sample input information, and multiple performance descriptions. The labels on the sample data can be selection labels for the candidate large models, which may have labels representing "deep thinking mode" or "non-deep thinking mode." This diverse labeling approach allows the trained scheduling model to simulate the target object's needs regarding content quality and the execution logic of the large model. Thus, the trained scheduling model can accurately and comprehensively determine the content quality needs represented by the input information and the execution logic needs of the large model. This enables automated model scheduling and execution logic control based on the target object's preferences for object-related features and the semantics of the input information, improving the matching degree between the response information and the target object.
[0065] In some embodiments, understanding the input information based on the object-related features of the target object to obtain the target intent may include: processing the input information and object-related features using multiple trained classification models to obtain multiple initial model invocation intents; and determining the model invocation intent in the target intent based on the multiple initial model invocation intents, wherein the model invocation intent is used to determine the large model from multiple candidate large models.
[0066] In one embodiment, a trained classification model can be used to determine an initial model invocation intent by processing input information and object-related features. The initial model invocation intent could be, for example, the scores corresponding to multiple candidate large models. By fusing the scores corresponding to multiple initial model invocation intents, the resulting target score is used as the model invocation intent. The candidate large model with the highest target score is then determined as the large model specified by the target intent.
[0067] In some embodiments, some of the candidate large models may be deep thinking models capable of performing semantic understanding tasks based on execution logic of a deep thinking mode, while others may be deep thinking models capable of performing semantic understanding tasks based on execution logic of a non-deep thinking mode. Thus, by determining the large model used for the semantic understanding task from the candidate large models, the execution logic of the large model corresponding to the content quality requirements of the target intent representation can be determined, and the large model can be controlled to perform the semantic understanding task.
[0068] It should be noted that the classification model can be constructed based on any type of neural network algorithm, such as attention network algorithm, long short-term memory network algorithm, etc. The embodiments of this disclosure do not limit the specific type of algorithm used to construct the classification model.
[0069] Figure 3 The illustration shows a schematic diagram of the principle of a large-model-based interaction method according to an embodiment of the present disclosure.
[0070] like Figure 3 As shown, multiple candidate large models can include candidate large model M301, candidate large model M302, ... up to candidate large model M30n. The input information 301 of the target object, object-related features 302, and performance description information 303 associated with each of the multiple candidate large models are input into multiple trained classifiers. For example, the input information 301, object-related features 302, and performance description information 303 can be input into a first classifier, a second classifier, and a third classifier, respectively. The first classifier, the second classifier, and the third classifier each output a score corresponding to candidate large model M301, candidate large model M302, ... up to candidate large model M30n as an initial model invocation intent. By fusing multiple initial model invocation intents, a model invocation intent is obtained. The model invocation intent can instruct candidate large model M301 to be used as the specified large model for performing the semantic understanding task. Here, n is an integer greater than 1.
[0071] By feeding the input information into a trained intent understanding model, the first pattern intent in the target intent is output. The first pattern intent indicates that the content quality requirement of the response information meets the preset quality requirement conditions. The first candidate large model M301 is used to perform a semantic understanding task on the input information 301 based on the execution logic corresponding to the deep thinking pattern of the first pattern intent, and the response information 304 is output.
[0072] According to embodiments of this disclosure, performing a semantic understanding task on input information using a large model based on execution logic matching the target intent may include: performing semantic understanding on the input information using a large model based on a first pattern intent to obtain a semantic understanding result; and performing a semantic understanding task using a large model based on multiple requirement information.
[0073] According to embodiments of this disclosure, the target intent may include a first mode intent, which indicates that the content quality requirements of the response information meet preset quality requirement conditions. For example, the first mode intent may indicate that the quality evaluation indicators such as the number of words and logical consistency of the response content are greater than or equal to preset quality evaluation indicator thresholds.
[0074] According to embodiments of this disclosure, the first mode intent can instruct a large model to perform a semantic understanding task on the input information based on a deep learning mode. For example, by performing semantic understanding on the input information using a large model, a semantic understanding result including multiple demand information can be obtained. The multiple demand information is matched with the demand intent represented by the input information; for example, the multiple demand information can be multiple sub-questions decomposed from the question text represented by the input information after understanding it.
[0075] For example, if the input is "What's the weather like today?", the semantic understanding result could contain multiple pieces of information such as "What city is the target object located in?", "Which authorized or open query interface does the target object use to execute the query?", "What factors affect the temperature, rainfall, or snowfall at various times today and in the future?", and so on. Thus, these multiple pieces of information can represent the requirements of the large model after performing semantic understanding on the input information, representing the various execution steps needed to perform the semantic understanding task of generating the response information.
[0076] According to embodiments of this disclosure, by utilizing a large model to perform semantic understanding tasks based on multiple demand information, the large model can perform semantic understanding tasks according to a deep thinking mode, targeting multiple demand information mined from input information and object-related features. Thus, when the identified target intent indicates a need for high-quality response information, the large model, guided by a first mode intent, decomposes or mines the demand intent of the input information based on object preferences and intents represented by object-related features. This enables deep thinking about the demands represented by the input information of the target object, and by utilizing the large model to perform semantic understanding tasks based on multiple demand information matching the input information, the deep thinking mode is achieved, ensuring that the response information accurately meets the user's actual needs and improves the interaction efficiency and experience of the target object.
[0077] In some embodiments, using a large model to perform semantic understanding on the input information to obtain multiple demand information that matches the demand intent represented by the input information may further include: using a large model to perform semantic understanding on the input information and object-related features to obtain multiple demand prompt words for multiple subtasks in the semantic understanding task.
[0078] According to embodiments of this disclosure, a subtask in the semantic understanding task for input information may include using a large model to perform semantic understanding on the demand information to output response content for the demand information. For example, the subtask may be a process of semantic understanding and response content generation for the "where is the target object located" option within a set of multiple demand information enclosed in quotation marks, such as "where is the target object located?", "which authorized or open query interface is used to execute the query for the city?", "what factors affect the temperature, rainfall, or snowfall for multiple time periods today and in the future?", etc.
[0079] According to embodiments of this disclosure, requirement prompts can be used to control subtasks in a semantic understanding task to execute according to requirement conditions that match the target object's intent. For example, for the requirement text content "Where is the target object located?" in the requirement information, the requirement prompt could be "The city location is represented by a range of latitude and longitude coordinates under open licensing." This allows the large model to process the information content and requirement prompts of the requirement information to execute subtasks more accurately, outputting execution results that meet the requirements of the deep thinking mode. Furthermore, response information can be generated based on the execution results of multiple subtasks, ensuring that the response information meets the actual needs of the target object.
[0080] In some embodiments, performing a semantic understanding task based on multiple requirements information using a large model further includes: using a gating network of the large model to process the multiple requirements information in order to identify multiple target expert networks related to multiple sub-tasks from multiple expert networks.
[0081] According to embodiments of this disclosure, a large model may include a gating network and multiple expert networks. The gating network and the multiple expert networks can be model structure layers constructed based on deep learning algorithms. The gating network is used to determine, through semantic understanding of multiple demand information, a target expert network capable of performing sub-tasks based on the demand information. The target expert network can be matched with demand conditions or intents such as the output file format, professional domain, and number of tokens represented by the demand information. Thus, the gating network can activate the model parameters of the target expert network among the multiple expert networks in the large model to execute sub-tasks, reducing the amount of model parameters required for the large model to execute multiple sub-tasks and reducing the computational overhead of the computing device performing semantic understanding tasks. Simultaneously, by selecting the activated expert network to execute sub-tasks based on the demand information obtained through intent mining of the input information through the gating network, the large model can more accurately execute multiple sub-tasks through multiple expert network models matching the demand information in a deep thinking mode, thereby improving the accuracy and content quality of the response information determined by the execution results of the sub-tasks.
[0082] In one embodiment, a large model gating network can be used to process the information content and requirement prompts of multiple requirement information to determine multiple target expert networks. This allows for more accurate activation of target expert networks that meet the requirement intent represented by the requirement prompts, based on the requirement conditions indicated by the requirement prompts. This improves the accuracy of subtask execution results and further matches the content quality of the response information with the actual needs of the target object.
[0083] In some embodiments, a semantic understanding task is performed using a large model based on multiple requirement information, including: using a first target expert network of the large model to perform a first sub-task based on the first requirement information to obtain a first execution result; using a second target expert network of the large model to perform a second sub-task based on the second requirement information and the first execution result to obtain a second execution result, wherein the response information is determined based on the first execution result and the second execution result.
[0084] In one embodiment, the first target expert network, based on the demand information content "where is the city where the target object is located" and the demand prompt "the city location is represented by a range of open authorized latitude and longitude coordinates" in the first demand information, outputs the authorized latitude and longitude coordinate area of the city where the target object is located as the first execution result of the first execution subtask. The second target expert network processes the demand information content "which authorized or open query interface is used to perform the query on the city" in the first demand information and the first execution result to determine the query interface address as the second execution result. By using the target expert network to execute subsequent subtasks based on the execution results of already executed subtasks according to the demand information, it is possible to control multiple target expert networks to execute semantic understanding tasks according to the execution logic of multiple subtasks based on multiple demand information. This allows the large model to execute semantic understanding tasks based on the execution logic corresponding to the deep thinking mode, improving the information content quality of the response information.
[0085] In some embodiments, the semantic understanding task executed based on the execution logic corresponding to the first mode intent can be demonstrated by displaying multiple requirement information and the execution results corresponding to the multiple requirement information in the display interface. This shows the deep thinking process of the large model in response to the input information, so as to indicate to the target object that the response information to the input information is the result generated by the large model through the deep thinking mode, and the target object can clearly understand the execution process of the semantic understanding task.
[0086] Figure 4 The illustration shows a schematic diagram of the principle of a large-model-based interaction method provided according to another embodiment of the present disclosure.
[0087] like Figure 4As shown, the large model used to perform semantic understanding tasks can include a gating network M411 and multiple expert networks M420. The multiple expert networks M420 include a first expert network M421, a second expert network M422, ..., an nth expert network M2n. The gating network M411 processes the first requirement information 401 and the second requirement information 402 from multiple requirement information, outputting two target expert networks from the multiple expert networks M420, respectively, for processing the first requirement information 401 and the second requirement information 402. The two target expert networks are the first expert network M421 and the second expert network M422. The gating network M411 can also transmit the first requirement information 401 and the second requirement information to the first expert network M421 and the second expert network M422, so that the multiple target expert networks can perform multiple sub-tasks in the semantic understanding task based on the multiple requirement information, outputting multiple execution results. By fusing the multiple execution results, response information 403 can be output.
[0088] In some embodiments, performing a semantic understanding task on the input information using a large model based on execution logic that matches the target intent may include: performing a semantic understanding task on the input information based on the second mode intent and according to the model parameters of the expert network in the large model to obtain response information.
[0089] According to embodiments of this disclosure, the second mode is intended to indicate that the content quality requirements of the response information do not meet preset quality requirement conditions. For example, the second mode may indicate that the quality evaluation indicators such as the number of words or logical consistency of the response content are less than preset quality evaluation indicator thresholds.
[0090] According to embodiments of this disclosure, the second mode intent can instruct a large model to perform a semantic understanding task on the input information based on a non-deep learning mode. For example, the large model can perform semantic understanding on the input information "1+1 equals what", and output the response information "1+1=2".
[0091] In one embodiment, the input information "1+1 equals what?" can be processed using a target expert network with mathematical calculation capabilities within a large model. This allows the model parameters of the target expert network to be activated to perform a semantic understanding task on the input information "1+1 equals what?", outputting the response information "1+1=2". This approach reduces the data size of the model parameters involved in the semantic understanding task by activating only the model parameters of the target expert network within the large model, lowering the computational overhead of the computing device. Furthermore, by generating response information based on the content quality requirements of the target object's input information, the response quality and the degree of matching with the target object's intent are improved, thereby enhancing response accuracy and interaction precision.
[0092] Figure 5The diagram illustrates an application scenario of the interaction method based on a large model according to an embodiment of the present disclosure.
[0093] like Figure 5 As shown, the display interface 500 can show the first input information 501 input by the target object as "medication to relieve frequent headache symptoms". By understanding the intent of the first input information 501 based on the object-related features of the target object, it can be determined that the intent of the first input information 501 is to relieve headaches. The determined target intent may include a first pattern intent to instruct the target object to quickly purchase the relevant medication. The large model can process the first input information 501 based on the execution logic corresponding to the deep thinking pattern that matches the first pattern intent, and obtain the first response information 510.
[0094] After viewing the first response message 510, the target user can input a second message 502: "What are the main components of drug A?". Based on the target user's deep understanding of the functions of various compounds in the relevant drug domain, as indicated by the object interaction features, the intent of the second input message 502 can be understood using these features, confirming that the target intent represents a second-mode intent. By using a large model to perform semantic understanding on the second input message 502 based on the execution logic representing a non-deep thinking mode that matches the second-mode intent, the second response message 520 can be output relatively quickly. This allows the target user to quickly understand, based on their accumulated knowledge, that the main components of drug A are compound A and compound B, avoiding information redundancy caused by outputting explanations of the compounds, and ensuring that the response message meets the target user's actual content quality requirements.
[0095] According to embodiments of this disclosure, the interaction method based on a large model may further include: when the target intent also includes a retrieval intent that represents the need to perform a retrieval task, invoking retrieval resources to perform a retrieval based on the input information, and obtaining retrieval results.
[0096] According to embodiments of this disclosure, the search resources may include an authorized search resource interface. By calling the search resource interface to perform information retrieval based on input information, search results can be obtained.
[0097] In one embodiment, a search can be performed based on the input information "What is the price of vehicle A?" by calling a retrieval resource interface. The search results can include the sales prices of various models of vehicle A, discount information for vehicle A, and the time periods corresponding to the discount information. By using a large model to process the search results and input information, the discounted sales prices of vehicle A for various time periods can be output as response information.
[0098] According to embodiments of this disclosure, the response information is determined by performing a semantic understanding task on the input information using a large model based on the retrieval results and the target intent.
[0099] For example, a large model can be used to represent the first mode intent based on the target intent, and the retrieval results and input information can be processed through the execution logic corresponding to the deep thinking mode to obtain the response information.
[0100] For example, a large model can be used to represent the second mode intent based on the target intent, and the retrieval results and input information can be processed through the execution logic corresponding to the non-deep thinking mode to obtain the response information.
[0101] In some embodiments, invoking retrieval resources to perform data retrieval based on input information may further include: rewriting the input information using a large model based on prompts determined by the target intent to obtain updated input information; and invoking retrieval resources to perform data retrieval based on the updated input information.
[0102] According to embodiments of this disclosure, the prompts determined based on the target intent may include prompts that match the content quality requirements indicated by the first or second mode intent. This controls the large model to rewrite the input information, ensuring that the updated input information avoids semantic ambiguity, excessively broad search scope, and other defects. Consequently, data retrieval can be performed based on the updated input information to obtain retrieval results that satisfy the semantic understanding task.
[0103] For example, if the input information is "price of car A", it can be rewritten based on the first pattern intent using a large model to obtain "the corresponding prices of car A1 and A2 models". Therefore, a search can be performed based on "the corresponding prices of car A1 and A2 models" to obtain search results.
[0104] Figure 6 A block diagram of an interactive device for a large model according to an embodiment of the present disclosure is shown schematically.
[0105] like Figure 6 As shown, the interactive device 600 based on a large model includes: a receiving module 610, a first obtaining module 620, a second obtaining module 630, and a push module 640.
[0106] The receiving module 610 is used to receive input information from the target object.
[0107] The first acquisition module 620 is used to understand the intent of the input information based on the object-related features of the target object, and obtain the target intent, which represents the content quality requirements for the response information.
[0108] The second acquisition module 630 is used to perform a semantic understanding task on the input information using a large model based on execution logic that matches the target intent, and obtain response information.
[0109] The push module 640 is used to push reply information to the target object.
[0110] According to embodiments of this disclosure, the second obtaining module includes: a first obtaining unit and a second obtaining unit.
[0111] The first obtaining unit is used to perform semantic understanding on the input information based on the first mode intent using a large model to obtain a semantic understanding result. The semantic understanding result includes multiple demand information that matches the demand intent represented by the input information. The first mode intent indicates that the content quality requirements of the response information meet the preset quality requirement conditions.
[0112] The second acquisition unit is used to perform semantic understanding tasks based on multiple demand information using a large model to obtain response information.
[0113] According to embodiments of this disclosure, multiple requirements are used for multiple sub-tasks in a semantic understanding task; wherein, the second obtaining unit includes: a first obtaining sub-unit and a second obtaining sub-unit.
[0114] The first obtaining subunit is used to execute the first subtask based on the first requirement information using the first objective expert network of the large model, and obtain the first execution result.
[0115] The second obtaining subunit is used to perform a second subtask based on the second requirement information and the first execution result using the second objective expert network of the large model, and obtain a second execution result, wherein the response information is determined based on the first execution result and the second execution result.
[0116] According to embodiments of this disclosure, the second obtaining module includes a third obtaining unit.
[0117] The third obtaining unit is used to perform semantic understanding of the input information based on the second mode intent and the model parameters of the expert network in the large model to obtain the response information. The second mode intent indicates that the content quality requirements of the response information do not meet the preset quality requirement conditions.
[0118] According to embodiments of this disclosure, the second obtaining unit further includes a first determining subunit.
[0119] The first determining subunit is used to process multiple requirement information using a gated network of a large model, so as to determine multiple target expert networks related to multiple sub-tasks from multiple expert networks.
[0120] According to embodiments of this disclosure, the second obtaining module includes a fourth obtaining unit.
[0121] The fourth acquisition unit is used to perform semantic understanding on the input information and object-related features using a large model, and to obtain multiple requirement prompts for multiple subtasks in the semantic understanding task.
[0122] According to embodiments of this disclosure, the first obtaining module includes a target intent obtaining unit.
[0123] The target intent acquisition unit is used to understand the intent of object-related features, input information and performance description information related to multiple candidate large models, and obtain the target intent. The target intent also instructs the large model among the multiple candidate large models to perform semantic understanding tasks.
[0124] According to embodiments of this disclosure, the target intent acquisition unit includes: a target fusion feature acquisition subunit and a target intent acquisition subunit.
[0125] The target fusion feature acquisition subunit is used to fuse object-related features, input information, and performance description information related to multiple candidate large models based on the attention mechanism to obtain target fusion features.
[0126] The target intent acquisition subunit is used to understand intent based on target fusion features and obtain the target intent.
[0127] According to embodiments of this disclosure, the first obtaining module includes: an initial model invocation intent obtaining unit and a large model determining unit.
[0128] The initial model invocation intent acquisition unit is used to process input information and object-related features using multiple trained classification models to obtain multiple initial model invocation intents.
[0129] The large model determination unit is used to determine the model invocation intent in the target intent based on multiple initial model invocation intents. The model invocation intent is used to determine the large model from multiple candidate large models.
[0130] According to embodiments of this disclosure, object-related features include interactive behavior features of the target object in response to at least one of the following target information: historical response information; target knowledge information related to the knowledge domain of the input information.
[0131] According to embodiments of this disclosure, the interactive device based on a large model further includes a search result acquisition module.
[0132] The retrieval results acquisition module is used to retrieve retrieval results by invoking retrieval resources based on the input information when the target intent also includes a retrieval intent that represents the need to perform a retrieval task. The response information is determined by performing a semantic understanding task on the input information based on the retrieval results and the target intent using a large model.
[0133] According to embodiments of this disclosure, the retrieval result obtaining module includes an update unit and a calling unit.
[0134] The update unit is used to rewrite the input information based on the prompt words determined by the target intent, using the large model to obtain updated input information.
[0135] The invocation unit is used to invoke retrieval resources to perform data retrieval based on the updated input information.
[0136] Figure 7 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0137] In embodiments of this disclosure, such as Figure 7 As shown, the AI agent 700 may include an input module 710, a processing module 720, and an output module 730.
[0138] Input module 710 is used to receive input information;
[0139] The processing module 720 is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and obtain output information by calling the large model to execute the interaction method based on the large model according to the embodiments of this disclosure.
[0140] Output module 730 is used to output the output information obtained by the processing module.
[0141] According to embodiments of this disclosure, the input module 710 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 700 can understand and process. The input module 710 is the primary link for the AI agent 700 to interact with the outside world, enabling the AI agent 700 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0142] In the example, the input module 710 can input the input information described above, object-related features, etc.
[0143] In the example, processing module 720 is the core support for the AI agent 700's ability to handle complex tasks. Processing module 720 can execute the large model-based interaction methods described above.
[0144] In the example, the performance of the processing module 720 is closely related to the large model on which the AI agent 700 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 720 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0145] In the example, after the AI agent 700 acquires the request voice, the processing module 720 can use a large model to perform a semantic understanding task on the input information, obtain the response information, and pass the response information to the output module 730.
[0146] Understandably, while large models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. Once the AI agent 700 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.
[0147] In the example, output module 730 can output the response information described above.
[0148] The AI agent 700 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0149] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0150] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0151] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.
[0152] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0153] Figure 8 A schematic block diagram of an example electronic device for implementing a large-model-based interaction method of embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0154] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0155] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0156] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the large model-based interaction method. For example, in some embodiments, the large model-based interaction method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the large model-based interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the large model-based interaction method by any other suitable means (e.g., by means of firmware).
[0157] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0158] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0159] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0162] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0163] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0164] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A large model-based interaction method, comprising: receiving input information of a target object; performing intent understanding on the input information based on object-related features of the target object to obtain a target intent, the target intent representing a content quality requirement for reply information; performing a semantic understanding task on the input information based on execution logic matched with the target intent by using a large model to obtain the reply information, the target intent including a second mode intent, the second mode intent indicating that the content quality requirement of the reply information does not meet a preset quality requirement condition, and based on the second mode intent, performing the semantic understanding task on the input information according to model parameters of an expert network in the large model; pushing the reply information to the target object.
2. The method of claim 1, wherein, The performing of the semantic understanding task on the input information based on the execution logic matched with the target intent by using the large model comprises: performing semantic understanding on the input information by using the large model based on a first mode intent to obtain a semantic understanding result, wherein the semantic understanding result includes a plurality of requirement information matched with a requirement intent represented by the input information, and the first mode intent indicates that the content quality requirement of the reply information meets the preset quality requirement condition; and performing the semantic understanding task by using the large model based on the plurality of requirement information.
3. The method of claim 2, wherein, The plurality of requirement information is used for a plurality of subtasks in the semantic understanding task. The performing of the semantic understanding task by using the large model based on the plurality of requirement information comprises: performing a first subtask based on first requirement information by using a first target expert network of the large model to obtain a first execution result; performing a second subtask based on second requirement information and the first execution result by using a second target expert network of the large model to obtain a second execution result, wherein the reply information is determined based on the first execution result and the second execution result.
4. The method of claim 3, wherein, The performing of the semantic understanding task by using the large model based on the plurality of requirement information further comprises: processing the plurality of requirement information by using a gating network of the large model to determine a plurality of target expert networks related to the plurality of subtasks from a plurality of expert networks.
5. The method of claim 2, wherein, The performing of the semantic understanding on the input information by using the large model to obtain the plurality of requirement information matched with the requirement intent represented by the input information comprises: performing semantic understanding on the input information and the object-related features by using the large model to obtain a plurality of requirement trigger words for a plurality of subtasks in the semantic understanding task.
6. The method of claim 1, wherein, The performing of the intent understanding on the input information based on the object-related features of the target object to obtain the target intent comprises: performing intent understanding on the object-related features, the input information, and performance description information related to a plurality of candidate large models to obtain the target intent, the target intent further indicating that the large model in the plurality of candidate large models is used to perform the semantic understanding task.
7. The method of claim 6, wherein, The performing of the intent understanding on the object-related features, the input information, and the performance description information related to a plurality of candidate large models comprises: The object-related features of the target object, the input information, and performance description information related to a plurality of candidate large models are fused based on an attention mechanism to obtain target fusion features; An intent understanding is performed based on the target fusion features to obtain a target intent.
8. The method of claim 1, wherein, The intent understanding of the input information based on the object-related features of the target object obtains a target intent, including: The input information and the object-related features are processed by a plurality of trained classification models to obtain a plurality of initial model invocation intents; and A model invocation intent in the target intent is determined according to a plurality of the initial model invocation intents, and the model invocation intent is used to determine the large model from a plurality of candidate large models.
9. The method of claim 1, wherein, The object-related features include interactive behavior features of the target object for at least one of the following target information: historical reply information; target knowledge information related to the knowledge field of the input information.
10. The method of claim 1, wherein, The method further includes: In the case that the target intent further includes a retrieval intent representing a need to perform a retrieval task, a retrieval resource is invoked to perform retrieval based on the input information to obtain a retrieval result, wherein the reply information is determined by a large model based on the retrieval result and the target intent to perform a semantic understanding task on the input information.
11. The method of claim 10, wherein, The invocation of the retrieval resource to perform data retrieval based on the input information includes: updating the input information by rewriting the input information based on a prompt word determined based on the target intent using the large model to obtain updated input information; invoking a retrieval resource to perform data retrieval based on the updated input information.
12. A large model-based interaction device, comprising: a receiving module configured to receive input information of a target object; a first obtaining module configured to perform intent understanding on the input information based on object-related features of the target object to obtain a target intent, the target intent representing a content quality requirement for reply information; a second obtaining module configured to perform a semantic understanding task on the input information based on an execution logic matched with the target intent using a large model to obtain the reply information; a pushing module configured to push the reply information to the target object; wherein the second obtaining module includes: a third obtaining unit configured to perform a semantic understanding task on the input information based on a second mode intent according to model parameters of an expert network in the large model to obtain the reply information, wherein the second mode intent indicates that the content quality requirement of the reply information does not meet a preset quality requirement condition.
13. The apparatus of claim 12, wherein, The second obtaining module includes: a first obtaining unit configured to perform semantic understanding on the input information using the large model based on a first mode intent to obtain a semantic understanding result, wherein the semantic understanding result includes a plurality of requirement information matched with a requirement intent represented by the input information, and the first mode intent indicates that the content quality requirement of the reply information meets a preset quality requirement condition; and a second obtaining unit configured to perform the semantic understanding task based on a plurality of the requirement information using the large model to obtain the reply information.
14. The apparatus of claim 13, wherein, The plurality of requirement information is used for a plurality of sub-tasks in the semantic understanding task. The second obtaining unit comprises: A first obtaining sub-unit configured to execute a first sub-task based on first requirement information by using a first target expert network of the large model to obtain a first execution result; A second obtaining sub-unit configured to execute a second sub-task based on second requirement information and the first execution result by using a second target expert network of the large model to obtain a second execution result, wherein the reply information is determined based on the first execution result and the second execution result.
15. An artificial intelligence agent product, comprising: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method of any one of claims 1 to 11 by calling the large model; an output module configured to output the output information obtained by the processing module.
16. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 11.
17. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to execute the method of any one of claims 1 to 11.
18. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 11. The computer instructions are used to enable the computer to execute the method of any one of claims 1 to 11.
18. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 11.
Citation Information
Patent Citations
Universal terminal perception interaction processing method, control device and storage medium
CN118034637A
User question answering method and device, equipment and medium
CN119336867A