Interaction method and device based on multi-model cooperation, intelligent agent and electronic equipment
Through the interactive method of multi-model collaboration, the first and second models are used to understand the intent and semantics of demand information, and generate reply information that matches user needs, which solves the problem of low resource content matching on the Internet and improves the accuracy of information push and user experience.
Patent Information
- Application Number
- CN202510970335.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-14
AI Technical Summary
It is difficult for users to quickly find resource content that matches their needs and interests on the Internet. Existing search technologies are unable to accurately understand user needs, resulting in inaccurate information push and increased interaction frequency and complexity.
An interactive method based on multi-model collaboration is adopted. The first model is used to understand the intention of demand information and resource-related features, and generate demand intention description text. The second model is used to semantically understand the demand intention description text and retrieval results, and generate reply information to improve the matching degree.
Through multi-model collaboration, the accuracy of reply information is improved, the frequency and complexity of interactions are reduced, the user experience is improved, the user interaction experience is improved, the accuracy of information and efficiency are achieved, the technical problems that have not been effectively solved in the existing technology are solved, and the user experience is improved.
Smart Images

Figure CN120805926A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to the technical field of intelligent reply, intelligent search, resource recommendation, and intelligent customer service. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, a user can conveniently browse resource information such as news and videos through a terminal device such as a smart phone. Alternatively, the user can also input search information through the terminal device to search for information content, so as to meet information acquisition needs such as travel planning and knowledge learning. SUMMARY
[0003] The present disclosure provides an interaction method and device based on multi-model cooperation, an intelligent agent, and a storage medium.
[0004] According to an aspect of the present disclosure, an interaction method based on multi-model cooperation is provided, which includes: receiving demand information input by a target object; performing intent understanding on the demand information and resource-related features by using a first large model to obtain demand intent description text, wherein the resource-related features are related to resources browsed by the target object, and the demand intent description text represents a demand degree of the target object for resource content based on a natural language form; performing semantic understanding on the demand intent description text and a search result determined based on the demand information by using a second large model to obtain reply information; and pushing the reply information to the target object.
[0005] According to another aspect of the present disclosure, an interaction device based on multi-model cooperation is provided, which includes: a receiving module configured to receive demand information input by a target object; a first obtaining module configured to perform intent understanding on the demand information and resource-related features by using a first large model to obtain demand intent description text, wherein the resource-related features are related to resources browsed by the target object, and the demand intent description text represents a demand degree of the target object for resource content based on a natural language form; a second obtaining module configured to perform semantic understanding on the demand intent description text and a search result determined based on the demand information by using a second large model to obtain reply information; and a pushing module configured to push the reply information to the target object.
[0006] According to another aspect of the present disclosure, an artificial intelligence intelligent agent is provided, which includes: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, execute a method according to an embodiment of the present disclosure by calling the large model, and obtain output information; and an output module configured to output the output information obtained by the processing module.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the embodiments of the present disclosure.
[0008] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method provided by the embodiments of the present disclosure.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method provided by the embodiments of the present disclosure.
[0010] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:
[0012] Figure 1 An exemplary system architecture to which the interaction method and device based on multi-model cooperation according to the embodiments of the present disclosure can be applied is schematically shown;
[0013] Figure 2 A flowchart of the interaction method based on multi-model cooperation according to the embodiments of the present disclosure is schematically shown;
[0014] Figure 3 A principle schematic diagram of the interaction method based on multi-model cooperation provided by the embodiments of the present disclosure is schematically shown;
[0015] Figure 4 A principle schematic diagram of the interaction method based on multi-model cooperation provided by another embodiment of the present disclosure is schematically shown;
[0016] Figure 5 A block diagram of the interaction device based on multi-model cooperation according to the embodiments of the present disclosure is schematically shown;
[0017] Figure 6 A structural block diagram of an artificial intelligence agent according to the embodiments of the present disclosure is schematically shown; and
[0018] Figure 7A schematic block diagram of an example electronic device that can be used to implement the multi-model collaboration based interaction method of embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0019] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which show various details of embodiments of the present disclosure to assist in that understanding, and should be considered in that light. Thus, those of ordinary skill in the art will recognize the embodiments described herein can vary as a function of, inter alia, the applicable audio encoding / decoding standards and / or communication standards. Wherever possible, the description herein has been presented to include, without limitation, details of the preferred embodiments; however, it is contemplated that the disclosed embodiments can be practiced without these specific details. In addition, it should be appreciated that the description presented is not necessarily oriented in time and / or order. For example, operations can be performed in various recited orders and / or concurrently. Moreover, the description given with reference to the figures is illustrative only, and it will thus be recognized that modifications and / or additions can be made to the preferred embodiments that fall within the scope and spirit of the present disclosure. Likewise, every control and / or procedure illustrated has associated latency, and, so where appropriate, additional controls and / or procedures can be added to the application in order to permit increased system responsiveness.
[0020] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations, necessary security measures are taken, and the public order and good customs are not violated.
[0021] The inventors have found that, with the rapid development of Internet technology, the massive resource data generated on the Internet makes it difficult for users to quickly find the required resource content, and it is easy to make the searched information content irrelevant to the user's needs, or the searched information content is too different from the user's actual interests and demand scenarios, resulting in difficulty in meeting the user's needs.
[0022] The present disclosure provides a multi-model collaboration based interaction method, device, agent and storage medium. The multi-model collaboration based interaction method includes: receiving demand information input by a target object; using a first large model to understand the intention of the demand information and resource related features to obtain a demand intention description text, wherein the resource related features are related to resources browsed by the target object, and the demand intention description text represents the demand degree of the target object for resource content based on natural language form; using a second large model to understand the semantics of the demand intention description text and the retrieval result determined based on the demand information to obtain reply information; and pushing the reply information to the target object.
[0023] According to an embodiment of the present disclosure, by utilizing the first large model to perform intent understanding on the demand information and the resource-related features to obtain the demand description text, the demand degree of the target object on the resource content based on the natural language description of the target object is realized, so as to capture the interest degree of the target object on the resource content browsed in the resource. Thus, based on the second large model, the semantic understanding on the demand intent description text and the retrieval result is performed, so that the second large model can accurately understand the demand degree and the interest change of the target object on the resource content through the natural language attributes such as the syntax structure of the natural language form of the demand description text and the demand degree description word, and the matching degree between the information content of the retrieval result based on the demand information, and then the generated reply information can be matched with the demand degree of the target object on the resource content, so as to improve the matching degree between the reply information and the actual demand of the user, thereby improving the reply information pushing accuracy, reducing the interaction frequency and the interaction complexity, and improving the user experience.
[0024] Figure 1 An exemplary system architecture to which the multi-model cooperation based interaction method and device according to an embodiment of the present disclosure can be applied is schematically shown.
[0025] It should be noted that, Figure 1 The system architecture shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the multi-model cooperation based interaction method and device can be applied can include terminal devices, but the terminal devices can not need to interact with the server to implement the multi-model cooperation based interaction method and device provided by the embodiments of the present disclosure.
[0026] As Figure 1 shown, the system architecture 100 according to this embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.
[0027] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software, etc. (only as examples).
[0028] The terminal devices 101, 102, and 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and the like.
[0029] The server 105 can be a server providing various services, such as a background management server providing support for content browsed by a user using the terminal devices 101, 102, and 103 (only as an example). The background management server can perform analysis and the like on received user requests and the like, and feed back the processing results (such as a webpage, information, or data generated or obtained according to a user request) to the terminal devices.
[0030] The server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in a cloud computing service system, and solves the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server 105 can also be a server of a distributed system, or a server combined with a blockchain.
[0031] It should be noted that the interaction method based on multi-model cooperation provided in the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Accordingly, the interaction apparatus based on multi-model cooperation provided in the embodiments of the present disclosure can also be arranged in the terminal devices 101, 102, or 103.
[0032] Alternatively, the interaction method based on multi-model cooperation provided in the embodiments of the present disclosure can also be generally executed by the server 105. Accordingly, the interaction apparatus based on multi-model cooperation provided in the embodiments of the present disclosure can generally be arranged in the server 105. The interaction method based on multi-model cooperation provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105. Accordingly, the interaction apparatus based on multi-model cooperation provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105.
[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system 100 is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers.
[0034] Figure 2 An illustrative flowchart of the interaction method based on multi-model cooperation according to the embodiments of the present disclosure is shown.
[0035] As shown in Figure 2 The interaction method based on multi-model cooperation includes operations S210-S240.
[0036] In operation S210, demand information input by a target object is received.
[0037] In operation S220, a first large model is used to perform intent understanding on the demand information and resource-related features to obtain demand intent description text.
[0038] In operation S230, a second large model is used to perform semantic understanding on the demand intent description text and search results determined based on the demand information to obtain reply information.
[0039] In operation S240, the reply information is pushed to the target object.
[0040] According to an embodiment of the present disclosure, the demand information input by the target object can include data in any modality such as text, voice, image, etc. The target object can input the demand information through a terminal device such as a smartphone.
[0041] According to an embodiment of the present disclosure, the first large model and the second large model can be language large models with generative capabilities based on deep learning algorithms. The language large models have a large number of parameters, for example, hundreds of millions to tens of billions of model parameters, to implement semantic understanding of information represented by natural language and perform generative data processing tasks based on a large number of parameters, generating structured or unstructured data such as text, tables, charts, etc. The first large model and the second large model are large models suitable for performing data processing tasks in different task requirement conditions, and the first large model and the second large model can have different models adapted to different scenarios.
[0042] In some embodiments, the resource-related features are related to resources browsed by the target object, which can include video resources, news resources, commodity information resources, etc. The resource-related features can be related to any resource content such as the theme content of the resource, the video screenshot content, the comment content, the text content, etc.
[0043] According to an embodiment of the present disclosure, the demand intent description text represents the demand degree of the target object for the resource content in the form of natural language, which can include attention degree, preference degree, etc. The demand intent description text can describe the demand degree of the target object for the resource content based on adjectives, words representing time changes, phrases representing probabilities, etc.
[0044] According to an embodiment of the present disclosure, the search results can be any type of information such as page text, image, video, etc. obtained by searching based on the demand information.
[0045] In some embodiments, the demand intention description text can also represent the change of the demand degree of the target object for the resource content, so that the second large model can more accurately understand the change of the demand intention of the target object to process the retrieval result, and generate reply information matched with the change of the demand intention, so as to improve the matching degree of the reply information and the actual demand of the target object, and improve the accuracy and timeliness of the reply.
[0046] According to the embodiments of the present disclosure, the demand intention description text represents the demand degree of the target object for the resource content based on the natural language form, so that the second large model can perform semantic understanding based on the demand degree of the target object for the resource content expressed by the natural language form of the demand intention description text, and learn the preference degree or attention degree of the target object for different resource contents. Thereby, the second large model can perform semantic understanding on the demand intention description text and the retrieval result determined based on the demand information, to generate reply information matched with the demand degree of the target object for the resource content.
[0047] In one example, the demand information can be "where to travel on the weekend?". The demand intention description text is "the user is interested in city A and city B, for example, has a possible experience demand for the rail transit of city A, has a photo experience demand for B scenic spot of city B, but has a very high interest in the cruise project of city B in the past month". It should be understood that city A, city B, B scenic spot, and city B cruise project can be resource contents in the resources browsed by the target object. The words such as "has a possible experience demand" and "has a very high interest" in the demand intention description text can be words representing the demand degree.
[0048] The second large model can generate reply information as: "According to your interest preferences, you can refer to the following itinerary planning content: B City famous city river cruise project, ticket price is 110 yuan, duration of 1 hour. You can reserve tickets through the following link: XXXY.COM. In addition, you can also visit B scenic spot after the cruise project, and reserve B scenic spot tickets through the following link: XyyyyY.COM". The demand intention description text can be expressed in natural language form to express the demand degree of the target object for the resource content and the change of the demand degree. To avoid the second large model from being difficult to understand the demand information according to the discretized resource content, it is realized that the demand intention description text can more clearly and accurately understand the demand degree of the target object for the resource content based on the demand intention description text, and process the information related to each resource content in the retrieval result according to the demand degree for the resource content, and generate reply information matched with the actual intention of the target object. Therefore, the first large model and the second large model can be used to process resource-related features and demand information to generate reply information that more accurately represents the actual demand intention of the user.
[0049] In one embodiment, the first large model and the second large model are instructed by different types of role attribute instructions to perform data processing tasks of different task requirement conditions based on the respective role attributes corresponding to the role attribute instructions. The role attributes represented by the role attribute instructions can include demand preference analysis expert attributes and reply content planning expert attributes. Therefore, the first large model and the second large model can be instructed to perform tasks by inputting role attribute instructions representing demand preference analysis expert attributes and reply content planning expert attributes to the first large model and the second large model respectively, to generate demand intention description text and reply information.
[0050] In some embodiments, the first large model is used to understand the intention of the demand information and the resource-related features to obtain the demand intention description text can include: based on the intention understanding prompt information, using the first large model to execute the intention understanding task according to the demand information and the resource-related features.
[0051] According to an embodiment of the present disclosure, the intention understanding prompt information is used to prompt the first large model to perform an intention understanding task based on the resource content-related features to obtain the demand intention description text. The resource content-related features can represent the interaction behavior of the target object to the resource content, such as the browsing duration, the browsing frequency, the like behavior, and the like. The intention understanding prompt information can be represented in any format of data such as a symbol, a text, a script code, and the like. The intention understanding prompt information can be used as input data of the first large model to prompt the first large model to understand the demand degree of the target object to different resource content by the interaction behavior of the target object to the resource content-related features and the demand intention represented by the input demand information, so that the demand intention description text can accurately describe the demand degree of the target object to different resource content of the input demand information, and thus the second large model can output more accurate reply information by processing the search result and the demand intention description text.
[0052] In some embodiments, the resource content-related features can include at least one of a browsing duration feature, a browsing frequency feature, a like behavior feature, and a comment behavior feature. The resource content-related features can be obtained by feature extraction or feature encoding of the interaction behavior of the target object to the resource content such as the browsing duration, the browsing frequency, the like behavior, and the comment behavior.
[0053] In some embodiments, performing the intention understanding task by the first large model according to the demand information and the resource-related features based on the intention understanding prompt information can include performing the intention understanding task by the first large model according to the demand information and the associated resource feature in the resource-related features based on the intention understanding prompt information.
[0054] The associated resource feature satisfies a preset semantic similarity condition with the demand information. The associated resource feature can be related to an associated resource in the resource browsed by the target object. The associated resource can include a resource related to the theme semantics represented by the demand information. For example, the demand information is “where to go on the weekend?”, and the associated resource feature can be related to a video, a news, and the like resource related to tourism browsed by the target object.
[0055] According to an embodiment of the present disclosure, by controlling the first large model to perform the intention understanding task according to the demand information and the associated resource feature based on the intention understanding prompt information, the resource-related features unrelated to the intention represented by the demand information in the first large model can be removed, and the noise interference on the first large model can be reduced, so that the demand intention description text can more accurately and finely describe the demand degree or the attention degree of the demand information to different resource content, and thus the matching degree between the reply information output by the second large model and the actual intention of the target object can be improved.
[0056] In some embodiments, the first large model can analyze the multiple resource contents that the target object needs to pay attention to by inputting the current demand information and the demand degree of the resource contents corresponding to the resource contents based on the prompt function of the intent understanding prompt information, the demand information, the historical search information of the target object, and the resource content features of the resources searched based on the historical search information that meet the semantic similarity condition with the demand information. Then, the demand degree description accuracy of the demand intent description text for different resource contents of the target object is improved. Thus, the demand degree fine-grained description manner of the resource contents by the intent description text is realized to improve the accuracy of the subsequent reply information.
[0057] In some embodiments, the retrieval result can include a basic retrieval result retrieved based on the basic demand information, and an extended retrieval result retrieved based on the extended demand information. The basic demand information meets the semantic similarity condition with the basic intent of the semantic representation of the demand information. The extended demand information related to the extended intent of the target object meets the semantic difference condition with the basic demand information. The second large model can generate the reply information by processing the demand information, the basic retrieval result, and the extended retrieval result, so as to meet the potential demand of the target object in the process of inputting the demand information and improve the reply accuracy and the reply quality.
[0058] In some embodiments, the retrieval result is determined based on the following operation: updating the demand information based on the resource-related features and the object environment attributes related to the demand information to obtain target demand information; and retrieving based on the target demand information to obtain the detection result.
[0059] According to embodiments of the present disclosure, the object environment attribute is related to the environment around the target object in the process of inputting the demand information. For example, it can be related to the weather, time, geographical location, and the like around the target object. By updating the demand information based on the resource-related features and the object environment attribute, the target demand information obtained can be related to the demand intent of the target object in a specific environment condition, and the target demand information can be more accurately retrieved to the resources related to the theme of the resource-related features, so that the target demand information can more accurately represent the actual information retrieval intent of the target object inputting the demand information. Thus, the retrieval result can be retrieved by the target demand information, so that the second large model can more completely and comprehensively process the information related to the resource content in the retrieval result according to the demand degree of the target object for the resource content based on the demand intent description text and the retrieval result retrieved based on the target demand information, and improve the matching degree between the reply information and the actual demand of the target object.
[0060] In some embodiments, the object environment attribute includes at least one of the following attributes related to the process of inputting the demand information by the target object: time attribute, location attribute, and weather condition attribute.
[0061] In some embodiments, the time attribute can include time period data, such as a field or identifier representing a time period "9:00-11:00 am" or the like. Or the time attribute can also include data representing a specific scenario time period attribute such as a holiday, lunch, etc.
[0062] In some embodiments, the location attribute can include data representing the geographic location of the target object. The weather condition attribute can represent the weather condition near the target object during the input of the demand information, and the weather condition attribute can represent information related to weather conditions such as rainfall, typhoon warning, snowfall, temperature, humidity, etc.
[0063] By updating the demand information based on the resource-related features and the object environment attributes related to the demand information, the target demand information can accurately represent the current search intent and demand intent of the target object, and adapt to specific conditions such as the current weather, geographic location, time requirement, etc. of the target object, for example, it can avoid determining the target demand information as a search information of a tourist attraction with a recommended travel duration of more than 4 days when the target object is on a workday and there is no holiday. Thus, the second large model can generate reply information that accurately represents the actual demand of the user by processing the search results retrieved based on the target demand information and the demand intent description text, thereby improving the accuracy of information reply.
[0064] In some embodiments, the target demand information can include the basic demand information and the extended demand information output by the third large model by processing the resource-related features and the object environment attributes. The basic demand intent can have a high semantic similarity with the demand information, and can change the semantic ambiguity, logical disorder or information missing of the demand information, so that the basic demand information can accurately represent the demand intent represented by the demand information, and avoid the second large model processing the demand information to produce hallucinations, resulting in the reply information not matching the actual demand intent.
[0065] Thus, by retrieving the target demand information that meets the time condition, location condition, environment condition, etc. represented by the object environment attributes, the basic search results and the extended search results can be obtained, so that the second large model can process the basic search results and the extended search results according to the demand degree of the resource content represented by the demand intent description text, so that the reply information can more accurately represent the basic demand intent and the extended demand intent of the target object, reduce the interactive process of the target object continuing to follow up the extended demand intent to input demand information, reduce the complexity and operation step learning cost of the interaction, improve the efficiency of information interaction and the quality of reply information pushing.
[0066] It should be noted that any data involved in any embodiment of the present disclosure, including but not limited to resource-related features, object environment attributes, etc., is obtained under the condition of obtaining authorization from the relevant user or institution, and the purpose of obtaining the data is informed to the relevant user or institution before obtaining the data, and necessary encryption or desensitization measures are taken, in compliance with relevant regulations and without violating public order and good customs.
[0067] In some embodiments, the retrieval based on the target demand information to obtain the detection result can include: retrieving based on the target demand information to obtain a plurality of initial detection results; and determining the detection result from the plurality of initial detection results based on a semantic matching degree between the initial detection results and the target demand information.
[0068] In one example, a plurality of initial retrieval results retrieved based on the basic demand information can be subjected to semantic relevance detection with the basic demand information. The semantic matching degree between the initial retrieval results and the basic demand information is obtained, so that the retrieval results related to the theme of the basic demand information can be determined from the initial retrieval results according to the semantic matching degree, to avoid the large model processing resources irrelevant to the basic demand intent to obtain reply information, and to improve the quality of the reply information.
[0069] In one example, a plurality of initial retrieval results retrieved based on the extended demand information can be subjected to semantic relevance detection with the extended demand information. The semantic matching degree between the initial retrieval results and the extended demand information is obtained, so that the retrieval results related to the theme of the extended demand information can be determined from the initial retrieval results according to the semantic matching degree, to avoid the large model processing resources irrelevant to the extended demand intent to obtain reply information, and to improve the quality of the reply information.
[0070] In one example, the plurality of initial detection results can be retrieved based on the basic demand information and the extended demand information respectively, the semantics of the initial retrieval results can be subjected to semantic relevance detection with the respective basic demand information or extended demand information, and the respective retrieval results of each basic demand information or extended demand information can be determined from the plurality of initial retrieval results corresponding to each basic demand information or extended demand information according to the semantic matching degree. Thus, the second large model can be used to process the retrieval results corresponding to the plurality of basic demand information and extended demand information to obtain the reply information, so that the second large model can output reply information satisfying the basic demand intent and the extended demand intent in combination with the demand degree of the target object for resource content, to further improve the quality of information reply.
[0071] In some embodiments, updating the demand information based on the resource-related features and the object environment attributes related to the demand information can include: performing at least one subtask in a demand information updating task by using the third large model, the subtask can include at least one of a first subtask and a second subtask.
[0072] The first subtask can include updating the demand information based on semantic features of the demand information to obtain basic demand information. By using the third large model to process the demand information to perform the first subtask, the output basic demand information can correct or supplement the content of the demand information that is not clearly expressed in semantics, so that the basic demand information can accurately represent the demand intention represented by the demand information.
[0073] In some embodiments, based on the task prompt information for the first subtask, the third large model is prompted to perform semantic understanding by processing the demand information, a preset knowledge base, and context content related to the demand information of the target object to update the demand information, and output rewritten basic demand information. The knowledge base can contain knowledge content related to various fields, and by combining the context content and the knowledge base to understand and rewrite the content of the demand information, the third large model can accurately understand the demand intention of the demand information, and by outputting the updated basic demand intention, it can supplement or correct the language expression defects in the demand information, so that the basic retrieval result retrieved based on the basic demand information can meet the actual demand intention of the target object. Thus, the quality of the reply information can be improved by processing the basic retrieval result and the demand intention description information based on the second large model.
[0074] The second subtask can include detecting resource browsing habits of the target object based on search resource features in the resource-related features to obtain resource browsing habit attributes, and detecting extended demand based on the resource browsing habit attributes and the object environment attributes to obtain extended demand information representing an extended intention of the target object.
[0075] According to embodiments of the present disclosure, the browsing habit attribute can represent the browsing preferences of the target object for resource length, resource playback time, and resource publisher (such as a specific video producer) of resource content. By using the third large model to process the resource browsing habit attribute and the object environment attribute to detect the extended demand, the extended demand information can be matched with the browsing habits of the target object for the resource and the current environment around the target object, so as to improve the matching of the extended demand information with the extended intention representing the potential demand of the target object.
[0076] For example, the third model can be used to process the target environment attributes "temperature 36 degrees, humidity 60" and the reading habit attribute "prefer blogger A's travel guides" to understand extended intent. The resulting extended demand information can include "search for blogger A's travel guide posts recommending summer escapes." This allows the extended search results to include page resources related to blogger A's travel guides on the theme of "summer escapes." By using the second model to process the extended search results and the demand intent description text, travel planning information on the theme of "summer escapes" can be generated as a response message. The travel planning information can include blogger A's recommended locations and planned routes to meet the target audience's potential needs.
[0077] In some embodiments, the third model can obtain basic demand information and extended demand information by executing the first subtask and the second subtask, and the embodiments of the present disclosure will not be described in detail here.
[0078] In some embodiments, the third model can be prompted to perform the first and second subtasks based on the respective prompt information for the first and second subtasks. The task prompt information may include any prompt words such as the output text word limit and the input information type, and the embodiments of the present disclosure are not further described here.
[0079] It should be understood that the target demand information includes at least one of basic demand information and extended demand information. In some embodiments, the retrieval results include at least one of basic retrieval results and extended retrieval results. The basic retrieval results are retrieved based on the basic demand information; the extended retrieval results are retrieved based on the extended demand information. The second largest model can generate reply information that can meet the diverse demand intentions of the target object by semantically understanding the demand intention description text with at least one of the basic retrieval results and the extended retrieval results, so as to improve the diversity, accuracy and timeliness of the reply information content, improve the quality of the reply and reduce the complexity of the interactive operation.
[0080] In some embodiments, updating the demand information based on resource-related characteristics and object environment attributes related to the demand information to obtain target demand information can also include: using the third model to perform the demand information update task by processing resource-related characteristics, object environment attributes and object preference description text to obtain target demand information.
[0081] According to embodiments of the present disclosure, the object preference description text is determined based on resource-related features. For example, the object preference description text can be obtained by processing resource-related features such as resource content, historical search information, and historical comment data of resources browsed by the target object in a preset historical period based on a trained large language model.
[0082] In some embodiments, the object preference description text describes at least one preference attribute of the target object in a structured manner. For example, the object preference description text can include a preference attribute name and description text corresponding to the preference attribute. For example, the object preference description text can be represented based on the content in Table 1 as follows.
[0083] Table 1
[0084]
[0085] By describing one or more ticket number attributes of the target object through the structured object preference description text, the resource-related features of the target object based on interaction operations in a historical period can be summarized and compressed, and noise interference caused by the target object mistakenly clicking to browse resource content can be removed. Thus, between processing the demand information of the target object by the first large model and the second large model through cooperation, the preference attributes of the target object can be summarized in a fine-grained manner based on the object preference description text, and the actual preference intentions such as attention and preference degree of multiple preference attributes can be described in a natural language form as a long-term user portrait of the target object, so that the third large model can clearly understand the preference intentions of the target object based on the structured natural language text, reduce the interference of noise data in the resource-related features, and reduce the problem of excessive computational overhead caused by processing the resource-related features of the long-term interaction behavior of the target object, improve the rewriting accuracy of the target demand information, and improve the matching degree between the retrieval results and the basic demand intentions or extended demand intentions of the target object. Then, based on the second large model, the quality of the reply information is improved by cooperating to process the retrieval results and the demand intention description text.
[0086] In some embodiments, the object preference description text can be determined based on the following operation: based on the preference understanding prompt information related to the multiple preference attributes, the expert large model is used to perform semantic understanding on the resource-related features and the object attribute features of the target object to obtain the object preference description text.
[0087] It should be noted that the object preference description text includes description text corresponding to the preference attribute, for example, can include description text corresponding to the preference attribute name.
[0088] According to an embodiment of the present disclosure, the expert large model is used to understand and extract the semantics of the object preference attributes in the resource-related features of the target object, and then output the object preference description text. Based on the preference understanding prompt information related to the multiple preference attributes, the preference attribute name based on the multiple preference attributes can be used to prompt the semantic understanding and preference attribute description text extraction of the resource-related features and the object attribute features, and then generate description text corresponding to each of the multiple preference attributes.
[0089] It should be noted that the expert large model can be a large model trained to understand and extract semantics representing object preference attributes in resource-related features. For example, the preference understanding prompt information can be used as a prompt to output structured object preference description text by processing resource-related features representing resource content, target object browsing duration features for resource content, and interaction behavior attribute features in resource-related features.
[0090] For example, the preference understanding prompt information can be: "The input resource-related features need to be understood and extracted according to the following six preference attributes to obtain the description text of the target object for each preference attribute name. The preference attribute names can be interest preference attribute, consumption preference attribute, lifestyle preference attribute, travel decision preference attribute, cultural type preference attribute, and emotional demand preference attribute. The output description text is structured and represented by the preference attribute name."
[0091] In some embodiments, the object preference description text is determined by semantic fusion of resource-related features and object attribute features of the target object. For example, an expert large model can be used to process resource-related features, object attribute features, and preference understanding prompt information to output object preference description text.
[0092] Figure 3 An illustrative schematic diagram of an interaction method based on multi-model cooperation according to an embodiment of the present disclosure is shown.
[0093] As shown in Figure 3 The first input data set 310 composed of demand information and other input data input by the target object is processed based on cooperation of the first large model, the second large model, and the third large model having different functions to obtain reply information 321. The first large model performs an intent understanding task on the demand information and resource-related features in the first input data set 310 based on intent understanding prompt information to obtain demand intent description text capable of representing the demand degree of the target object for resource content. The third large model performs a demand information updating task by processing resource-related features, object environment attributes, demand information, and object preference description text in the first input data set 310 to obtain basic demand information and extended demand information. The object preference description text is determined by processing resource-related features in a preset historical period using an expert large model, and the object preference description text represents the preference attributes of the target object based on a structured description method.
[0094] The networking search is performed based on the basic demand information and the extended demand information respectively, to obtain the basic search result and the extended search result corresponding to the basic demand information and the extended demand information respectively. The second large model is used to process the object preference description text, the demand information and the object environment attribute, the demand description text, the basic search result, the extended search result, and the basic search information and the extended demand information in the first input data set 310, to obtain the reply information 321 capable of meeting the actual demand intention of the target object.
[0095] In some embodiments, the second large model is used to perform semantic understanding on the demand intention description text and the search result determined based on the demand information, to obtain the reply information. The second large model is also used to perform semantic understanding on the demand intention description text, the search result, and the object preference description text, to obtain the reply information.
[0096] According to embodiments of the present disclosure, the object preference description text describes at least one preference attribute of the target object in a structured manner, and the demand intention description text can describe the demand degree and the demand degree change of the target object for the resource content in a fine-grained manner. Thus, the preference of the target object and the demand degree for the resource content can be more accurately captured by using the second large model to process the demand intention description text and the object preference description text, so as to effectively summarize and screen the resource content in the search result, make the output reply information more accurately meet the actual preference of the target object, and describe and plan the resource content with a higher demand degree of the target object, thereby improving the matching degree between the reply information and the actual demand intention of the target object.
[0097] In one embodiment, the second large model is used to perform semantic understanding on the demand intention description text, the basic search result, and the object preference description text, so that the reply information can more accurately meet the basic demand intention expressed by the demand information input by the target object, and the frequency of interactive operations caused by the target object repeatedly inputting new demand information to obtain detailed reply information is avoided, thereby improving the user experience.
[0098] In one embodiment, the second large model is used to perform semantic understanding on the demand intention description text, the basic search result, the extended search result, and the object preference description text, so that the reply information can more accurately meet the basic demand intention expressed by the demand information input by the target object, and the potential demand of the target object expressed by inputting the demand information, and the frequency of interactive operations caused by the target object repeatedly inputting new demand information to obtain reply information related to the extended demand intention is avoided, thereby improving the user experience.
[0099] In some embodiments, the semantic understanding of the demand intention description text, the retrieval result, and the object preference description text by the second large model can further include: based on the semantic understanding of the demand intention description text, the retrieval result, the object preference description text, and the object environment attribute related to the demand information by the second large model.
[0100] For example, the demand intention description text, the basic retrieval result, the extended retrieval result, the object preference description text, and the object environment attribute related to the demand information can be input data of the second large model, so that the second large model can perform semantic understanding and fusion on the basic retrieval result and the extended retrieval result under the condition of fully understanding the demand degree of the target object for the resource content, the actual preference content of the target object, and the environment of the target object, so that the reply information can fuse the resource information in the retrieval result according to the demand degree of the target object for the resource content and the preference attribute of the target object, and the reply information can meet the requirements of the target object for the current time stage, geographical location, weather condition, and other environmental conditions, improve the matching degree of the reply information and the actual demand and potential demand of the target object, and generate the final personalized reply information to improve the interactive experience.
[0101] In one embodiment, the demand intention description text, the basic demand information, the extended demand information, the basic retrieval result, the extended retrieval result, the object preference description text, and the object environment attribute related to the demand information can be input data of the second large model, so that the reply information can meet the actual demand of the target object.
[0102] In some embodiments, the reply information includes travel planning information described in a structured form.
[0103] According to embodiments of the present disclosure, the travel planning information includes: a travel item matched with the resource content preferred by the target object, and travel planning content corresponding to the travel item.
[0104] The travel item can be a resource content such as a scenic spot name or a restaurant name, and the travel planning content corresponding to the travel item can represent planning content for describing a ticket booking method, an opening time period, and transportation planning content.
[0105] In one embodiment, the order of the plurality of travel items can be arranged according to the demand degree of the target object for the resource content, and the travel item and the travel planning content can be related to the basic demand intention and the extended demand intention represented by the demand information input by the target object, so that the target object can output the personalized reply information that can meet the diversified demand of the target object by inputting the demand information, thereby improving the interactive experience.
[0106] Figure 4A schematic diagram of the principles of an interaction method based on multi-model collaboration provided according to another embodiment of the present disclosure is schematically shown.
[0107] like Figure 4 As shown, the first, second, and third models with different functions collaborate to process the second input data set 410 consisting of the target object's demand information and other input data to obtain the response information 321. The demand information can be, for example, the text "Where to go this weekend?" entered by the target object.
[0108] The first large model performs the intention understanding task based on the intention understanding prompt information for the demand information in the second input data set 410 and the associated resource features related to the demand information in the resource-related features, and obtains the demand intention description text 421 that can represent the degree of demand of the target object for the resource content.
[0109] The demand intention description text 421 may be, for example: "Users may want to know about tourist attractions in different cities, such as the unique light rail in City A, which can be photographed and checked in, and they especially hope to take a boat tour of the scenic river in City A. In addition, some users may visit the famous temple A in City B." In the demand intention description text 421, "the unique light rail in City A", "taking a boat tour", "the scenic river", and "the famous temple A" may be resource content. The demand intention description text 421 may use vocabulary and syntactic structures such as "especially hope" and "maybe" to express the degree of demand for resource content by the target object. It should be noted that the degree of demand for resource content in the demand intention description text can be expressed by any natural language expression method, and the embodiments of the present disclosure do not limit the specific form of the natural language expression method.
[0110] The third model performs the demand information update task by processing the resource-related features, object environment attributes, demand information, and object preference description text in the second input dataset 410, obtaining basic demand information 431 and extended demand information 432. Basic demand information 431 could be, for example, "Good weekend destinations in City A," while extended demand information 432 could be, for example, "Recommended cruises and other attractions." Searches are performed using basic demand information 431 and extended demand information 432, respectively, to obtain basic and extended search results.
[0111] The second large model processes the object preference description text, requirement information, object environment attributes, requirement description text, basic search results, extended search results, basic search information, and extended requirement information in the second input dataset 410 to obtain itinerary planning information 441 that can meet the actual requirements of the target object as response information. The itinerary planning information describes the itinerary items and itinerary plan content in a structured form.
[0112] The itinerary items can be, for example, "1, A City Cruise + Famous Restaurant", "2, A City Riverside Park Family Tour", and "3, A City Zoo". The itinerary planning content can be specific planning content corresponding to the itinerary items. By using the second large model to perform semantic understanding on the basis search result, the extended search result, and the demand intention description text, the object preference description text, and the object environment attribute, and generating the itinerary planning information 441 described in a structured manner, the actual demand of the user can be effectively met, the problem of high interaction operation complexity caused by multiple rounds of interaction to obtain reply information can be reduced, and the user experience can be further improved.
[0113] Figure 5 A block diagram of an interaction device based on multi-model cooperation according to an embodiment of the present disclosure is schematically shown.
[0114] As shown in Figure 5 The interaction device 500 based on multi-model cooperation includes a receiving module 510, a first obtaining module 520, a second obtaining module 530, and a pushing module 540.
[0115] The receiving module 510 is configured to receive demand information input by a target object.
[0116] The first obtaining module 520 is configured to perform intention understanding on the demand information and resource-related features by using a first large model to obtain demand intention description text, wherein the resource-related features are related to resources browsed by the target object, and the demand intention description text represents a demand degree of the target object for resource content based on a natural language form.
[0117] The second obtaining module 530 is configured to perform semantic understanding on the demand intention description text and a search result determined based on the demand information by using a second large model to obtain reply information.
[0118] The pushing module 540 is configured to push the reply information to the target object.
[0119] According to an embodiment of the present disclosure, the second obtaining module includes a first obtaining unit.
[0120] The first obtaining unit is configured to perform semantic understanding on the demand intention description text, the search result, and an object preference description text by using the second large model to obtain the reply information, wherein the object preference description text is determined by performing semantic fusion on resource-related features and object attribute features of the target object, and the object preference description text describes at least one preference attribute of the target object based on a structured manner.
[0121] According to an embodiment of the present disclosure, the first obtaining unit includes a semantic understanding subunit.
[0122] The semantic understanding subunit is configured to perform semantic understanding on the demand intention description text, the search result, the object preference description text, and the object environment attribute related to the demand information by using a second large model, where the object environment attribute is related to the environment around the target object in the process of inputting the demand information.
[0123] According to an embodiment of the present disclosure, the search result is determined based on the following operation: updating the demand information based on the resource-related feature and the object environment attribute related to the demand information to obtain target demand information, where the object environment attribute is related to the environment around the target object in the process of inputting the demand information; and performing search based on the target demand information to obtain the search result.
[0124] According to an embodiment of the present disclosure, updating the demand information based on the resource-related feature and the object environment attribute related to the demand information to obtain target demand information comprises: performing a demand information updating task by processing the resource-related feature, the object environment attribute, and the object preference description text by using a third large model, to obtain the target demand information, where the object preference description text is determined according to the resource-related feature, and the object preference description text describes at least one preference attribute of the target object in a structured manner.
[0125] According to an embodiment of the present disclosure, updating the demand information based on the resource-related feature and the object environment attribute related to the demand information comprises: performing at least one of the following subtasks by using a third large model: updating the demand information based on the semantic feature of the demand information to obtain basic demand information; detecting a resource browsing habit attribute of the target object based on a search resource feature in the resource-related feature, and performing extended demand detection based on the resource browsing habit attribute and the object environment attribute to obtain extended demand information representing an extended intention of the target object; and wherein the target demand information comprises at least one of the basic demand information and the extended demand information.
[0126] According to an embodiment of the present disclosure, the search result comprises at least one of the following: a basic search result obtained based on the basic demand information; and an extended search result obtained based on the extended demand information.
[0127] According to an embodiment of the present disclosure, the object preference description text is determined based on the following operation: performing semantic understanding on the resource-related feature and the object attribute feature of the target object by using an expert large model based on preference understanding prompt information related to a plurality of preference attributes, to obtain the object preference description text, where the object preference description text comprises description text corresponding to the preference attribute.
[0128] According to an embodiment of the present disclosure, the retrieving based on the target demand information obtains a detection result, including: retrieving based on the target demand information to obtain a plurality of initial detection results; and determining the detection result from the plurality of initial detection results based on a semantic matching degree between the initial detection result and the target demand information.
[0129] According to an embodiment of the present disclosure, the object environment attribute includes at least one attribute related to the following in the process of inputting the target demand information by the target object: a time attribute, a location attribute, and a weather condition attribute.
[0130] According to an embodiment of the present disclosure, the reply information includes trip planning information described in a structured form, and the trip planning information includes: a trip item matched with the resource content preferred by the target object, and trip planning content corresponding to the trip item.
[0131] According to an embodiment of the present disclosure, the first obtaining module includes a first execution unit.
[0132] The first execution unit is configured to perform an intent understanding task based on the demand information and resource-related features by using the first large model based on intent understanding prompt information, wherein the intent understanding prompt information is used to prompt the first large model to perform the intent understanding task based on at least one feature related to resource content, including: a browsing duration feature, a browsing frequency feature, a like behavior feature, and a comment behavior feature.
[0133] According to an embodiment of the present disclosure, the first execution unit includes a first execution subunit.
[0134] The first execution subunit is configured to perform an intent understanding task based on the demand information and associated resource features in the resource-related features by using the first large model based on the intent understanding prompt information, wherein the associated resource features satisfy a preset semantic similarity condition with the demand information.
[0135] Figure 6 An illustrative structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0136] In an embodiment of the present disclosure, as shown in Figure 6 The AI agent 600 can include an input module 610, a processing module 620, and an output module 630.
[0137] The input module 610 is configured to receive input information.
[0138] The processing module 620 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling a large language model to perform an interaction method based on multi-model cooperation according to an embodiment of the present disclosure.
[0139] The output module 630 is configured to output the output information obtained by the processing module.
[0140] According to an embodiment of the present disclosure, the input module 610 is responsible for receiving or perceiving information such as queries, requests, instructions, signals or data from the outside world (e.g., a user or an external environment), and converting the information into a format that can be understood and processed by the AI agent 600. The input module 610 is the first step for the AI agent 600 to interact with the outside world, and it enables the AI agent 600 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to the information.
[0141] In an example, the input module 610 can input the demand information and resource-related features described above.
[0142] In an example, the processing module 620 is the core support for the AI agent 600 to process complex tasks. The processing module 620 can perform the interaction method based on multi-model collaboration described above.
[0143] In an example, the performance of the processing module 620 can be closely related to the large model on which the AI agent 600 is based. In order to fully utilize the capabilities of the large model, the internal structure of the processing module 620 can be designed to be highly configurable and scalable in order to cope with various types of tasks and demands in real scenarios.
[0144] In an example, after obtaining the demand information, the processing module 620 can process the demand information and the resource-related features using the first large model to obtain a demand intent description text, process the demand intent description text and the retrieval results determined based on the demand information using the second large model to obtain reply information, and pass the reply information to the output module 630.
[0145] It can be understood that although the large language model has excellent language understanding and generation capabilities, it, like a human, can only solve a limited number of tasks without the aid of any tools. When the AI agent 600 is given the ability to call tools, it can achieve tasks such as completing mathematical operations with a calculator, completing data analysis with Python, and completing weather forecasts with a search engine.
[0146] In an example, the output module 630 can output the reply information described above.
[0147] The AI agent 600 according to an embodiment of the present disclosure can simply and effectively improve the degree of intelligence and improve flexibility and versatility.
[0148] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0149] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0150] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0151] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0152] Figure 7 A schematic block diagram of an example electronic device that can be used to implement the multi-model collaboration-based interaction method of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0153] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0154] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0155] The computing unit 701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the multi-model collaboration based interaction method. For example, in some embodiments, the multi-model collaboration based interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the computing unit 701, one or more steps of the multi-model collaboration based interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the multi-model collaboration based interaction method by any other suitable means, such as by means of firmware.
[0156] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0157] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0158] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0160] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0161] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0162] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0163] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An interactive method based on multi-model collaboration, comprising: Receive demand information input by the target object; Using the first model to understand the intention of the demand information and resource-related features, a demand intention description text is obtained, wherein the resource-related features are related to the resources browsed by the target object, and the demand intention description text represents the degree of the target object's demand for resource content in a natural language form; Using the second model to perform semantic understanding on the demand intention description text and the search results determined based on the demand information to obtain reply information; and Push the reply information to the target object.
2. The method according to claim 1, wherein The search results are determined based on the following operations: updating the demand information based on the resource-related characteristics and the object environment attributes related to the demand information to obtain target demand information, wherein the object environment attributes are related to the environment surrounding the target object during the process of inputting the demand information; A search is performed based on the target demand information to obtain the detection result.
3. The method according to claim 2, wherein: The updating of the demand information based on the resource-related characteristics and the object environment attributes related to the demand information to obtain target demand information includes: The target demand information is obtained by performing the demand information update task by processing the resource-related characteristics, the object environment attributes and the object preference description text using the third model, wherein the object preference description text is determined according to the resource-related characteristics, and the object preference description text describes at least one preference attribute of the target object in a structured manner.
4. The method according to claim 2, wherein: The updating of the demand information based on the resource-related characteristics and the object environment attributes related to the demand information includes: Using the third model, perform at least one of the following subtasks: updating the demand information based on the semantic features of the demand information to obtain basic demand information; Performing resource browsing habit detection on the target object based on the search resource feature in the resource-related feature to obtain resource browsing habit attributes, and performing extension demand detection based on the resource browsing habit attributes and the object environment attributes to obtain extension demand information representing the extension intention of the target object; The target demand information includes at least one of the basic demand information and the extended demand information.
5. The method according to claim 4, wherein The search results include at least one of the following: Basic search results retrieved based on the basic demand information; The extended search results retrieved based on the extended requirement information.
6. The method according to claim 3, wherein: The object preference description text is determined based on the following operations: Based on the preference understanding prompt information related to multiple preference attributes, the expert big model is used to perform semantic understanding on the resource-related features and the object attribute features of the target object to obtain the object preference description text, which includes the description text corresponding to the preference attributes.
7. The method according to any one of claims 2 to 6, wherein The searching based on the target demand information to obtain the detection result includes: Performing a search based on the target demand information to obtain a plurality of initial detection results; and The detection result is determined from a plurality of the initial detection results based on a semantic matching degree between the initial detection result and the target requirement information.
8. The method according to claim 1, wherein The second model is used to perform semantic understanding on the demand intention description text and the search results determined based on the demand information to obtain reply information, including: The second model is used to perform semantic understanding on the demand intention description text, the retrieval results and the object preference description text to obtain the reply information, wherein the object preference description text is determined by semantically fusing the resource-related features and the object attribute features of the target object, and the object preference description text describes at least one preference attribute of the target object in a structured manner.
9. The method according to claim 8, wherein The using the second model to perform semantic understanding on the demand intention description text, the search results, and the object preference description text includes: Based on the use of the second largest model, semantic understanding is performed on the demand intention description text, the retrieval results, the object preference description text and the object environment attributes related to the demand information, wherein the object environment attributes are related to the environment surrounding the target object during the process of inputting the demand information.
10. The method according to claim 9, wherein: The object environment attribute includes at least one of the following attributes related to the process of the target object inputting the requirement information: Time attributes, location attributes, and weather condition attributes.
11. The method according to claim 1, wherein The reply information includes itinerary planning information based on a structured description, and the itinerary planning information includes itinerary items that match the resource content preferred by the target object, and itinerary planning content corresponding to the itinerary items.
12. The method according to claim 1, wherein The first model is used to understand the intention of the demand information and resource-related features to obtain a demand intention description text, including: Based on the intent understanding prompt information, the first large model is used to perform the intent understanding task according to the requirement information and the resource-related characteristics, wherein the intent understanding prompt information is used to prompt the first large model to perform the intent understanding task based on at least one of the following characteristics related to the resource content: Browsing time characteristics, browsing frequency characteristics, like behavior characteristics, and comment behavior characteristics.
13. The method according to claim 12, wherein: The intention understanding prompt information is based on the first model, and the intention understanding task is performed according to the demand information and the resource-related characteristics, including: Based on the intention understanding prompt information, the first large model is used to perform the intention understanding task according to the demand information and the associated resource features in the resource-related features, and the associated resource features and the demand information meet the preset semantic similarity conditions.
14. An interactive device based on multi-model collaboration, comprising: A receiving module is used to receive demand information input by the target object; A first acquisition module is configured to use the first large model to perform intention understanding on the demand information and resource-related features to obtain a demand intention description text, wherein the resource-related features are related to the resources browsed by the target object, and the demand intention description text represents the degree of demand of the target object for the resource content in a natural language form; a second obtaining module, configured to use a second large model to perform semantic understanding on the demand intention description text and the search results determined based on the demand information to obtain reply information; and A push module is used to push the reply information to the target object.
15. The device according to claim 14, wherein The second obtaining module includes: The first acquisition unit is used to use the second large model to perform semantic understanding on the demand intention description text, the retrieval results and the object preference description text to obtain the reply information, wherein the object preference description text is determined by semantically fusing the resource-related features and the object attribute features of the target object, and the object preference description text describes at least one preference attribute of the target object in a structured manner.
16. The device according to claim 15, wherein The first obtaining unit includes: A semantic understanding sub-unit is used to perform semantic understanding on the demand intention description text, the retrieval results, the object preference description text and the object environment attributes related to the demand information based on the use of the second large model, wherein the object environment attributes are related to the environment surrounding the target object during the process of inputting the demand information.
17. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by executing the method according to any one of claims 1 to 13 by calling the large model; An output module is used to output the output information obtained by the processing module.
18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.
20. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Interaction method of artificial intelligence screen, artificial intelligence screen and storage medium
CN117649469A
Information retrieval method and device and electronic equipment
CN118916443A
Response method and device based on large model, electronic equipment and storage medium
CN119202225A
Memory retrieval method based on large language model and related device
CN119903125A
Interaction method and device based on large model, intelligent agent and storage medium
CN120216804A