Interaction method and device based on large model, intelligent agent and storage medium
By combining the large model and image subject information in human-computer interaction to understand the required text in semantics, the problem of inefficiency caused by semantic fuzzy user input text is solved, and more accurate and efficient push of reply information is achieved.
Patent Information
- Application Number
- CN202510678069.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-27
AI Technical Summary
During human-computer interaction, when the user requests the big model to process the target image by inputting text, the semantic expression of the input text may be blurred, resulting in multiple rounds of dialogue and multiple jump scenes, reducing the response efficiency and accuracy.
By receiving the target image and demand text, using the big model to semantic understanding of the demand text, identifying the target subject based on the image subject information, and then determining the reply information matching the processing requirements and pushing it to the target object.
It improves the timeliness and accuracy of reply information, reduces the complexity of interaction, and improves the user experience.
Smart Images

Figure CN120216804A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as deep learning and large models, and can be applied to application scenarios such as knowledge search, intelligent question answering, intelligent customer service, intelligent e-commerce, and AI (Artificial Intelligence) medical treatment. Background Art
[0002] With the rapid development of artificial intelligence technology, in the process of human-computer interaction, multi-modal data such as text, images, and voices input by users can be processed based on Artificial Intelligence Generated Content (AIGC) technology to generate reply content that meets the user's needs. Summary of the Invention
[0003] The present disclosure provides an interaction method, device, intelligent agent, and storage medium based on a large model.
[0004] According to one aspect of the present disclosure, there is provided an interaction method based on a large model, including: receiving a target image and a requirement text input by a target object, where the requirement text represents the processing requirement of the target object for the target image; based on the main body information of the image, using the large model to perform semantic understanding on the requirement text to obtain an intention understanding result, where the main body information of the image is determined by performing main body recognition on the target image; and determining a reply information that matches the processing requirement based on the intention understanding result, and pushing the reply information to the target object.
[0005] According to another aspect of the present disclosure, there is provided an interaction device based on a large model, including: a receiving module, configured to receive a target image and a requirement text input by a target object, where the requirement text represents the processing requirement of the target object for the target image; a first obtaining module, configured to perform semantic understanding on the requirement text using the large model based on the main body information of the image to obtain an intention understanding result, where the main body information of the image is determined by performing main body recognition on the target image; and a pushing module, configured to determine a reply information that matches the processing requirement based on the intention understanding result, and push the reply information to the target object.
[0006] According to another aspect of the present disclosure, there is provided an intelligent agent of artificial intelligence, including: an input module, configured to receive input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the interaction method based on the large model provided in the embodiments of the present disclosure by calling the large model to obtain output information; and an output module, configured to output the output information obtained by the processing module.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a large model-based interaction method provided according to an embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute a large model-based interaction method provided according to an embodiment of the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor, implements a large model-based interaction method provided according to an embodiment of the present disclosure.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which a large model-based interaction method and apparatus according to an embodiment of the present disclosure can be applied;
[0013] Figure 2 Schematically shows a flowchart of a large model-based interaction method according to an embodiment of the present disclosure;
[0014] Figure 3 Schematically shows a flowchart of determining reply information matching a processing requirement based on an intention understanding result according to an embodiment of the present disclosure;
[0015] Figure 4 Schematically shows an application scenario diagram of a large model-based interaction method according to an embodiment of the present disclosure;
[0016] Figure 5 Schematically shows a block diagram of a large model-based interaction apparatus according to an embodiment of the present disclosure;
[0017] Figure 6 Schematically shows a block diagram of the structure of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure; and
[0018] Figure 7FIG. 0 shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the large model-based interaction method of the present disclosure. DETAILED IMPLEMENTATION MANNER
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0020] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0021] The inventors found that users can request service resources integrating large models to process a target image by inputting text, so as to quickly obtain the image processing result. However, the input text may have vague semantic expressions and requires multi-hop scenarios such as multi-round dialogue methods to determine the actual intention of the user, resulting in low efficiency and poor accuracy in replying to the user.
[0022] Embodiments of the present disclosure provide a large model-based interaction method, apparatus, intelligent agent, and storage medium. The large model-based interaction method includes: receiving a target image and a requirement text input by a target object, where the requirement text represents the processing requirement of the target object for the target image; based on the image main body information, using a large model to perform semantic understanding on the requirement text to obtain an intention understanding result, and the image main body information is determined by performing main body recognition on the target image; and determining a reply information matching the processing requirement based on the intention understanding result, and pushing the reply information to the target object.
[0023] According to the embodiments of the present disclosure, for the requirement text input by the target object for the target image, by performing main body recognition on the target image to obtain the image main body information representing the target main body in the target image, and using the large model to perform semantic understanding on the requirement text according to the image main body information, the intention understanding result can more clearly represent the processing intention of the requirement text for the target image, avoiding semantic defects such as vague semantic expressions and unclear requirement main bodies in the requirement text. By determining the reply information according to the intention understanding result, the reply information can more accurately match the actual requirement of the target object for the target image, avoiding the reduction of interaction efficiency caused by determining the user's intention through multiple interactions, thereby improving the timeliness of pushing the reply information, reducing the complexity of interaction operations, and improving the user experience.
[0024] Figure 1Schematically shows an exemplary system architecture to which the interaction method and apparatus based on a large model according to an embodiment of the present disclosure can be applied.
[0025] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the interaction method and apparatus based on a large model can be applied may include terminal devices, but the terminal devices may be able to implement the interaction method and apparatus based on a large model provided by the embodiments of the present disclosure without interacting with the server.
[0026] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only examples).
[0028] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0029] The server 105 may be a server providing various services, such as a background management server that provides support for the content browsed by users using the terminal devices 101, 102, 103 (only an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0030] The server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server 105 can also be a server of a distributed system, or a server combined with a blockchain.
[0031] It should be noted that the interaction method based on the large model provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Correspondingly, the interaction device based on the large model provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103.
[0032] Alternatively, the interaction method based on the large model provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the interaction device based on the large model provided by the embodiments of the present disclosure can generally be set in the server 105. The interaction method based on the large model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the interaction device based on the large model provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0033] For example, when the user inputs a demand text while browsing a news resource containing a vehicle image, the terminal devices 101, 102, 103 can receive the vehicle image and the demand text that the user is browsing, and then send the obtained vehicle image and demand text to the server 105. The server 105 performs semantic understanding on the demand text by using the large model according to the vehicle main body information in the vehicle image to obtain an intention understanding result; determines a reply message according to the intention understanding result, and pushes the reply message to the terminal devices 101, 102, 103. Or a server or a server cluster capable of communicating with the terminal devices 101, 102, 103 and / or the server 105 performs semantic understanding on the demand text by using the large model according to the vehicle main body information in the vehicle image to obtain an intention understanding result, and pushes the reply message to the terminal devices 101, 102, 103 and / or the server 105.
[0034] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0035] Figure 2The flowchart of the interaction method based on the large model according to an embodiment of the present disclosure is schematically shown.
[0036] As Figure 2 shown, the interaction method based on the large model includes operations S210 to S230.
[0037] In operation S210, receive the target image and the requirement text input by the target object.
[0038] In operation S220, based on the image main body information, use the large model to perform semantic understanding on the requirement text to obtain the intention understanding result.
[0039] In operation S230, determine the reply information matching the processing requirement based on the intention understanding result, and push the reply information to the target object.
[0040] According to an embodiment of the present disclosure, the target image may be an image associated with the requirement text. For example, it may be an image input by the target object, or may also be an image browsed by the target object during a specified period or on a specified page.
[0041] According to an embodiment of the present disclosure, the requirement text represents the processing requirement of the target object for the target image. For example, the requirement text may represent an identification requirement, a query requirement, an editing requirement, etc. for the target image. The specific type of the processing requirement is not limited in the embodiments of the present disclosure.
[0042] In some embodiments, the requirement text may also represent the requirement that the target object needs to ask or query in combination with the target image. For example, if the target image is an image representing the A Tower building in City A, for the target image A, the requirement text may be "How tall is the building in the picture?". The requirement text represents the requirement of asking about the height of the building in combination with the target image A.
[0043] It should be noted that the requirement text can be determined based on any type of interaction operation such as the text input operation, text copy operation, voice input operation, etc. of the target object, as long as it can meet the requirement of being user input. The specific data format of the requirement text is not limited in the embodiments of the present disclosure.
[0044] According to an embodiment of the present disclosure, the image main body information is determined by performing main body recognition on the target image. The image main body information represents the target main body in the target image, and the target main body may be a person, an animal, a building, a vehicle, etc. The image main body information may represent attribute information such as the attribute type, shape, name, etc. of the target main body.
[0045] It should be noted that the image main body information may be determined by performing main body recognition on the target image before receiving the requirement text, or may also be determined by performing main body recognition on the target image when receiving the requirement text or after receiving the requirement text.
[0046] In some embodiments, the target image can be processed based on a target detection algorithm to determine the image main body information. For example, a target detection algorithm constructed based on the ResNet algorithm can be used to perform main body recognition on the target image to obtain the size and category of target main bodies such as vehicles in the target image.
[0047] In some embodiments, a trained vision large model can also be used to process the target image for main body recognition to obtain the image main body information. For example, a vision large model constructed based on the CogVLM (CogVision-Language Model) model can be used to process the target image for main body recognition.
[0048] According to the embodiments of the present disclosure, the image main body information characterizes the attribute information of the target main body in the target image. By using the image main body information as a requirement prompt to control the processing of the requirement text by the large model, the large model can more accurately and fully understand the requirement intention expressed by the requirement text according to the requirement prompt effect of the image main body information, avoid generating understanding hallucinations or understanding deviations for the requirement text, enable the intention understanding result to more accurately characterize the requirement intention of the target object of the input requirement text, so that the reply information matching the processing requirement can be determined based on the intention understanding result to meet the actual requirement intention of the target object for the target image, improve the quality of the reply information, so that the pushed reply information can more accurately meet the actual needs of the user, and avoid the user from performing multi-round interaction operations to meet the actual processing requirements for the target image, so as to improve the user experience of the target object.
[0049] It should be noted that the acquisition of information involved in the embodiments of the present disclosure, including but not limited to the target image, requirement text, etc., is obtained after obtaining the authorization of the relevant user or institution. And before obtaining the information, it is clearly informed to the user or institution that its purpose is to push accurate reply information, and at the same time, necessary encryption or desensitization measures are taken for the acquired information to avoid information leakage, which complies with relevant regulations and does not violate public order and good customs.
[0050] According to the embodiments of the present disclosure, the image main body information is determined based on the following operations: based on the preset prompt information, using the large model to perform semantic understanding on the target image and the requirement text to obtain a semantic understanding result; from at least one preset tool, calling a target tool that matches the functional attribute represented by the semantic understanding result to perform main body recognition on the target image to obtain the image main body information.
[0051] According to an embodiment of the present disclosure, the preset prompt information represents the functional attributes of a preset tool. The preset tool can be software resources or hardware resources such as a program or device for processing images. The functional attributes of the preset tool can include attribute information such as the type of processing method for image processing by the preset tool and the processing performance, which are used to describe the functions of the preset tool.
[0052] For example, the preset tool can be a text recognition tool, and the functional attributes of the text recognition tool can include the type of text language that can be recognized, etc.
[0053] Again, for example, the preset tool can be a target subject classification tool, and the functional attributes of the target subject classification tool can include the type of subject attribute of the target subject that can be recognized. The type of subject attribute can represent the classification of the subject, and the classification of the subject can include cultural relics, buildings, etc. Invoking the preset tool corresponding to the type of subject attribute can more accurately identify the target subject in the target image, so as to avoid misidentifying the target subject as a subject in other subject attribute types, resulting in logical contradictions or hallucinations in the reply information, and improving the matching degree between the reply information and the demand intention of the target object.
[0054] According to an embodiment of the present disclosure, the large model can be a multimodal large model for processing multimodal data. The multimodal large model can process and understand information in multiple modalities such as text, images, audio, and video. By integrating data in different modalities, the multimodal large model can more comprehensively understand and generate information, thereby better simulating human cognitive abilities. By using the multimodal large model to understand the target image and the demand text under the control condition of the preset prompt information, the large model can understand and match the functional attributes of the preset tool with the subject attributes of the target subject in the target image and the demand semantics of the demand text, so that the functional attributes represented by the semantic understanding result can be adapted to the subject recognition of the target image according to the intention represented by the demand text, improving the recognition accuracy of the image subject information, and further improving the accuracy of the intention understanding result for representing the demand intention. Thus, by using the large model to process the target image and the demand text, the target tool adapted to the subject recognition of the target image can be invoked to improve the quality of the reply information.
[0055] In one example, the target tool can be software resources. For example, the target tool can be an application program resource for identifying vehicle models. The application program resource can be invoked through the interface of the application program resource to process the target image and obtain the image subject information.
[0056] In one example, the target tool may be a hardware resource. For example, it may be a specified chip for text recognition of specified language text. The target image can be sent to the cache area for the specified chip, and a control instruction can be sent to control the specified chip to process the target image to obtain the image body information.
[0057] According to an embodiment of the present disclosure, the functional attribute of the preset tool characterizes that the preset tool is used for subject recognition of an image body with specified subject attributes. The specified subject attributes may represent the type of the subject attributes of the image body. For example, the specified subject attributes may be the subject categories for classifying the image body. The specified subject attributes may include at least one of animal attributes, plant attributes, vehicle attributes, and building attributes.
[0058] In some embodiments, the animal attributes may include bird attributes, fish animal attributes, tropical animal attributes, first-class protected animal attributes, and so on. Those skilled in the art can determine the specified animal attributes according to actual needs. The plant attributes may include flower attributes, tree attributes, algae attributes, crop attributes, etc.
[0059] In some embodiments, the vehicle attributes may include car attributes, airplane attributes, etc., and the building attributes may include bridge attributes, ancient building attributes, etc.
[0060] In one embodiment, the prompt message A is "The preset tool A is used for recognizing sedan-class subjects", and the prompt message A can be used to describe the functional attribute of the preset tool A for recognizing the sedan body in the image. The prompt message B is "The preset tool B is used for recognizing ship-class subjects", and the prompt message B can be used to describe the functional attribute of the preset tool B for recognizing the ship body in the image. When the target image is a promotional image of a new sedan, by using the large model to process the target image and the requirement text "How much is this car" according to the prompt message A and the prompt message B, a semantic understanding result including a control command representing "calling the preset tool with the function of recognizing sedan-class subjects to perform subject recognition on the target image" can be obtained. By calling the preset tool A that matches the functional attribute of recognizing the sedan body according to the control command to process the target image, the image body information including the model, color, brand, etc. of the vehicle in the target image can be determined.
[0061] The functional attribute can characterize that the corresponding preset tool is applicable to identifying an image subject with any one of animal attributes, plant attributes, vehicle attributes, and building attributes. By means of the preset prompt information characterizing the functional attribute, to prompt the large model to match the image semantics and text semantics of the target image with the functional attribute on the condition of understanding the image semantics of the target image and the text semantics of the demand text, it can make the functional attribute represented by the semantic understanding result match the specified subject attribute of the target subject to be recognized in the target image, so that the corresponding target tool in multiple preset tools can be called according to the functional attribute represented by the semantic understanding result to process the target image, and relatively accurate image subject information can be obtained.
[0062] In some embodiments, an initial model can be trained through sample images and labels corresponding to the specified subject attribute to obtain a preset tool with the functional attribute of subject recognition for images of image subjects with the specified subject attribute.
[0063] In some embodiments, prompting the large model to perform semantic understanding on the demand text and the target image through the preset prompt information may further include using the large model to retrieve associated images with a high similarity to the target image from a specified retrieval space, so that the large model can strengthen the degree of semantic understanding of the demand text and the target image through the retrieved associated images, avoid inaccurate selection of the target tool caused by unclear image subjects in the target image or overly vague semantic expressions in the demand text, make the semantic understanding result more accurately represent the subject attribute of the target subject in the target image, and improve the adaptability between the functional attribute of the target tool called by the semantic understanding result and the specified subject attribute of the target subject. Thereby, the accuracy of the target subject information can be improved, and it is ensured that the subsequent determined reply information matches the actual demand intention of the target object.
[0064] According to an embodiment of the present disclosure, the target tool includes a retrieval tool and a specified subject recognition tool, and the specified recognition tool is used to recognize an image subject with the specified subject attribute.
[0065] For example, the demand text is "What are the prices of these kinds of ships", and the ship subject in the target image is relatively small, making it difficult to determine the specific model and brand of the ship subject. By prompting the large model to perform semantic understanding on the demand text and the target image through the preset prompt information, it can be determined that the functional attribute represented by the semantic understanding result may include a retrieval attribute and a ship class recognition attribute. The ship class recognition attribute, as the functional attribute, indicates a relatively high recognition accuracy for subject recognition of image subjects with ship class attributes in the image, and further enables the image subject information to more accurately determine the brand and model of the ship.
[0066] In one embodiment, a target tool that matches the functional attributes represented by the semantic understanding result is called to perform subject recognition on the target image, and the obtained image subject information includes: calling a retrieval tool to obtain associated images that meet the preset similarity condition with the target image; calling a specified recognition tool to perform subject recognition on at least one of the target image and the associated images to obtain the image subject information.
[0067] For example, a retrieval tool with an image retrieval function retrieves associated images in a specified retrieval space whose similarity with the target image meets a preset similarity threshold. The associated images can represent associated image subjects similar to the target subject of the target image, and the associated image subjects can have specified subject attributes with the target subject. For example, the associated image subjects and the target subject can both have a yacht theme attribute. By calling a target tool whose functional attribute is to perform subject recognition on the image subjects with specified subject attributes to perform subject recognition on at least one of the associated images and the target image, more accurate image subject information can be determined. The image subject information can, for example, include yacht models, yacht names, etc. Thus, the large model can be used to process information such as yacht models and yacht names in the image subject information to prompt the large model to accurately understand the demand text and obtain a more accurate intention understanding result, so as to improve the matching degree between the reply information and the actual demand intention of the target object.
[0068] Figure 3 Schematically shows a flowchart for determining reply information that matches the processing demand based on the intention understanding result according to an embodiment of the present disclosure.
[0069] As Figure 3 shown, reply information that matches the processing demand can be determined based on operations S301 to S304.
[0070] After the start operation is executed, operation S301 is executed to determine whether the intention understanding result represents the need for retrieval. If the judgment result of operation S301 is yes, it can be understood that the intention understanding result includes a first intention representing the need for retrieval, and operations S302 and S303 are executed.
[0071] In operation S302, when the intention understanding result includes a first intention representing the need for retrieval, a retrieval tool is called to perform a retrieval based on at least one of the intention understanding result and the demand text to obtain retrieval information.
[0072] In operation S303, a first reply information is determined based on the retrieval information.
[0073] If the judgment result of operation S301 is no, it can be understood that the intention understanding result includes a second intention representing the need not to perform retrieval, and operation S304 is executed.
[0074] In operation S304, when the intention understanding result includes a second intention indicating that retrieval is not required, based on at least one of the image subject information and the target image, a second response message is determined.
[0075] In one example, the target image only includes a ship body, the image subject information represents a "ship of brand A1", and the demand text is "What is the price of a car of brand A and model A?" The intention understanding result may include "Whether to retrieve: Yes" indicating the first intention. Since the subject information of the demand text is "a car of brand A and model A" and the subject information of the demand text is clear, retrieval can be directly performed based on the demand text, and the retrieved information obtained includes "130,000 yuan". By using the trained large language model to process the retrieved information "130,000 yuan" and the demand text, the first response message can be determined as "The price of a car of brand A and model A is 130,000 yuan".
[0076] In one example, the intention understanding result further includes a target demand text, which is determined by using a large model to update the demand text based on the image subject information. For example, it can be obtained by using the large model to perform semantic understanding on the demand text based on the image subject information to obtain the target demand text and the first intention. Among them, operation S302 may include calling a retrieval tool to perform retrieval according to the target demand text in the intention understanding result to obtain the retrieved information.
[0077] For example, the target image is a car of brand A, the image subject information represents a "car of brand A and model B", and the demand text is "What is this price?" The intention understanding result may include "Whether to retrieve: Yes" indicating the first intention, and the intention understanding result may further include the target demand text "What is the price of a car of brand A and model B". The target demand text rewrites the demand text with unclear subject information to obtain a target demand text that can more clearly represent the demand intention of the target object for the target image.
[0078] At the same time, since the actual intention represented by the demand text is to query the price of a "car of brand A and model B", it is necessary to call a retrieval tool to perform retrieval with the target demand text "What is the price of a car of brand A and model B" containing the image subject information "car of brand A and model B" as the query text. The retrieved information obtained includes "130,000 yuan". By generating the first response message "a car of brand A and model B" with the retrieved information "130,000 yuan" and the image subject information, and pushing the first response message to the target object.
[0079] In one embodiment, the intention understanding result further includes a target requirement text for supplementing and explaining the requirement text. For example, if the requirement text is "What is the price of Model A car of Brand A", since the image main body information of the target image includes "Model A car of Brand A" and the iconic building "Tower B" of Country B, the target requirement text includes "Retrieve in the currency unit of Country B". Thus, based on the requirement text "What is the price of Model A car of Brand A" and the target requirement text "Retrieve in the currency unit of Country B", a retrieval tool can be called for retrieval, and the retrieved information "The price of Model B car of Brand A in Country B is 1.3 million (currency unit of Country B)" can be obtained. And the retrieved information is pushed to the target object as the first reply information.
[0080] In one embodiment, based on at least one of the image main body information and the target image, the determination of the second reply information is represented as follows in the following embodiments.
[0081] The target image is a press release file, the requirement text is "Please identify the text in the file", and the intention understanding result includes "Whether retrieval is required: No" as the second intention. The text characters identified according to the image main body information can be used as the second reply information to be pushed to the target object.
[0082] In one embodiment, based on at least one of the image main body information and the target image, the determination of the second reply information may further include: using a large model to process the target image and the requirement text to obtain the second reply information, where the requirement main body keyword in the requirement text matches the image main body information.
[0083] For example, 8 cats and 3 dogs in the target image are used as the target main body, and the requirement text is "What colors are the cats in the image and how many are there". The image main body information includes the classifications of specifying the main body attributes "long-haired cats", "Persian cats", "wolf dogs". Thus, it can be determined that the requirement main body keyword "cats" in the requirement text matches "long-haired cats" and "Persian cats" in the image main body information. The intention understanding result includes "Whether retrieval is required: No", and thus the large language model can be used to process the requirement text "What colors are the cats in the image and how many are there" and the target image to obtain the second reply information "There are 8 cats in the image, and they are all white".
[0084] In some embodiments, different target images and requirement texts, and the content included in the corresponding intention understanding results can be represented based on Table 1 below.
[0085] Table 1
[0086]
[0087] According to an embodiment of the present disclosure, it is determined whether it is necessary to call a retrieval tool for retrieval through the first intention or the second intention included in the intention understanding result. The main body information of the image and the requirement text can be processed by a large model to accurately detect whether retrieval is required for the response content of the target object. Thereby, it is possible to dynamically determine whether to call the retrieval interface resource according to the actual requirements of the target object, avoiding the redundant consumption of computing power resources and response delay caused by frequent calls to the retrieval interface resource. At the same time, it is also possible to avoid generating response information that does not match the actual requirement intention of the target object by regularly calling the retrieval tool to obtain retrieval information, improving the quality and efficiency of the response information.
[0088] In addition, in the case where the intention understanding result includes the first intention, the target requirement text included in the intention understanding result can be used to rewrite or supplement the requirement text with ambiguous semantic expression by using the main body information of the image. Thus, the retrieval tool can be called to perform retrieval according to at least one of the requirement text and the target requirement text, improving the matching degree between the retrieved retrieval information and the actual intention of the target object, and further enabling the response information to more accurately represent the actual intention of the target object, improving the accuracy of the response information.
[0089] Figure 4 The application scenario diagram of the interaction method based on the large model according to the embodiment of the present disclosure is schematically shown.
[0090] As Figure 4 shown, in this application scenario, the target object can input the target image P401 and the requirement text T401 "What is the price of this?". The tool call component is constructed based on a trained multi-modal large model. According to the preset prompt information, the multi-modal large model is used to process the target image P401 and the requirement text T401 to obtain the semantic understanding result. The preset prompt information can represent the functional attributes of multiple preset tools in the image processing tool library. For example, the preset prompt information can respectively represent the "car brand and model identification" functional attribute of the car identification tool, the "ship model, displacement and ship type identification" functional attribute of the ship identification tool, and the "tropical plant species identification" functional attribute of the plant identification tool. The car identification tool can be called to perform main body identification on the target image T401 through the functional attribute represented by the semantic understanding result, and the main body information of the image "Car of Brand A and Model A" can be obtained.
[0091] Use the reply information planning component to process the image main body information "A-brand A-model car", as well as the target image P401 and the requirement text T401. The reply information planning component can be constructed based on a fine-tuned multi-modal large model, and use the multi-modal large model to perform semantic understanding on the image main body information "A-brand A-model car" against the image P401 and the requirement text T401, and obtain the intention recognition result including the first intention "Whether retrieval is needed: Yes"; the rewritten requirement intention "Whether rewriting is needed: Yes", and the target requirement text represents the updated text of the requirement text T401 "What is the real-time price of the A-brand A-model car?".
[0092] By calling the retrieval resource interface of the retrieval engine, realize using the retrieval engine as a retrieval tool to retrieve according to "What is the real-time price of the A-brand A-model car?", and obtain the retrieval information "The price of the A-brand A-model car was just adjusted on January 1st, and the latest preferential price is 120,000 yuan".
[0093] By using the reply generation component constructed based on the trained large language model to process the retrieval information "The price of the A-brand A-model car was just adjusted on January 1st, and the latest preferential price is 120,000 yuan" and the requirement text T401, the reply information T402 "The selling price of the A-brand A-model car on January 1st is 120,000 yuan" can be generated
[0094] Figure 5 Schematically shows a block diagram of an interaction device based on a large model according to an embodiment of the present disclosure.
[0095] As Figure 5 shown, the interaction device 500 based on the large model includes: a receiving module 510, a first obtaining module 520, and a pushing module 530.
[0096] The receiving module 510 is configured to receive the target image and the requirement text input by the target object, and the requirement text represents the processing requirement of the target object for the target image.
[0097] The first obtaining module 520 is configured to perform semantic understanding on the requirement text by using the large model based on the image main body information, and obtain the intention understanding result, and the image main body information is determined by performing main body recognition on the target image.
[0098] The pushing module 530 is configured to determine the reply information matching the processing requirement based on the intention understanding result, and push the reply information to the target object.
[0099] According to an embodiment of the present disclosure, the pushing module 530 includes: a first determining sub-module and a second determining sub-module.
[0100] The first determination sub-module is configured to, when the intent understanding result includes a first intent indicating that retrieval is required, call a retrieval tool to perform retrieval based on at least one of the intent understanding result and the requirement text, obtain retrieval information, and determine a first reply message according to the retrieval information.
[0101] The second determination sub-module is configured to, when the intent understanding result includes a second intent indicating that retrieval is not required, determine a second reply message based on at least one of the image subject information and the target image.
[0102] According to an embodiment of the present disclosure, the intent understanding result further includes a target requirement text, and the target requirement text is determined by updating the requirement text using a large model based on the image subject information.
[0103] The first determination sub-module includes a retrieval unit.
[0104] The retrieval unit is configured to call a retrieval tool to perform retrieval according to the target requirement text and obtain retrieval information.
[0105] According to an embodiment of the present disclosure, the second determination sub-module includes a second reply message obtaining unit.
[0106] The second reply message obtaining unit is configured to process the target image and the requirement text using a large model to obtain a second reply message, where the requirement subject keyword in the requirement text matches the image subject information.
[0107] According to an embodiment of the present disclosure, the image subject information is determined based on the following operations: based on preset prompt information, using a large model to perform semantic understanding on the target image and the requirement text to obtain a semantic understanding result, where the preset prompt information represents the functional attributes of a preset tool; calling a target tool that matches the functional attributes represented by the semantic understanding result from at least one preset tool to perform subject recognition on the target image to obtain the image subject information.
[0108] According to an embodiment of the present disclosure, the target tool includes a retrieval tool and a specified subject recognition tool, and the specified recognition tool is configured to recognize the image subject with specified subject attributes.
[0109] Calling a target tool that matches the functional attributes represented by the semantic understanding result to perform subject recognition on the target image to obtain the image subject information includes: calling a retrieval tool to obtain associated images that meet the preset similarity condition with the target image; calling the specified recognition tool to perform subject recognition on at least one of the target image and the associated images to obtain the image subject information.
[0110] According to an embodiment of the present disclosure, the functional attribute of the preset tool characterizes that the preset tool is used for subject recognition of an image subject with specified subject attributes, and the specified subject attributes include at least one of the following: animal attribute, plant attribute, vehicle attribute, building attribute.
[0111] Figure 6 Schematically shows a structural block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure.
[0112] In an embodiment of the present disclosure, as Figure 6 shown, the AI intelligent agent 600 may include an input module 610, a processing module 620, and an output module 630.
[0113] The input module 610 is configured to receive input information;
[0114] The processing module 620 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by invoking the large model to execute the interaction method based on the large model provided according to an embodiment of the present disclosure;
[0115] The output module 630 is configured to output the output information obtained by the processing module.
[0116] According to an embodiment of the present disclosure, the input module 610 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (such as users or the external environment), and converting it into a format that the AI intelligent agent 600 can understand and process. The input module 610 is the primary link for the AI intelligent agent 600 to interact with the outside world, enabling the AI intelligent agent 600 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0117] In an example, the input module 610 may input the target image and the requirement text described above.
[0118] In an example, the processing module 620 is the core support for the AI intelligent agent 600 to handle complex tasks. The processing module 620 may execute the interaction method based on the large model described above.
[0119] In an example, the performance of the processing module 620 may be closely related to the large model on which the AI intelligent agent 600 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 620 can be designed to be highly configurable and extensible to cope with various different types of tasks and requirements in real scenarios.
[0120] In the example, after the AI agent 600 receives the requirement text and the target image, the processing module 620 can use the large model to process the requirement text and the image body information for the target image, obtain the intention understanding result, use the large model to process the intention understanding result to obtain the reply information, and transmit the reply information to the output module 630.
[0121] It can be understood that although the large language model has excellent language understanding and generation capabilities, like humans, without any tools, the tasks it can solve are very limited. When the AI agent 600 is given the ability to call tools, it can perform tasks such as completing mathematical operations with the help of a calculator, performing data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine.
[0122] In the example, the output module 630 can output the reply information described above.
[0123] The AI agent 600 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.
[0124] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0125] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0126] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0127] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0128] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 for a large model-based interaction method that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0129] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0130] Multiple components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0131] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the interaction method based on a large model. For example, in some embodiments, the interaction method based on a large model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the interaction method based on a large model described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the interaction method based on a large model in any other suitable manner (e.g., by means of firmware).
[0132] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0136] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0137] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0138] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0139] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An interaction method based on a large model, comprising: Receiving a target image and a requirement text input by a target object, where the requirement text represents the processing requirement of the target object for the target image; Based on the image main body information, using the large model to perform semantic understanding on the requirement text to obtain an intention understanding result, where the image main body information is determined by performing main body recognition on the target image; And Based on the intention understanding result, determining a reply information that matches the processing requirement, and pushing the reply information to the target object.
2. The method according to claim 1, wherein, The determining the reply information that matches the processing requirement based on the intention understanding result includes: In the case where the intention understanding result includes a first intention indicating that retrieval is required, calling a retrieval tool to perform retrieval according to at least one of the intention understanding result and the requirement text to obtain retrieval information, and determining a first reply information according to the retrieval information; and In the case where the intention understanding result includes a second intention indicating that retrieval is not required, determining a second reply information based on at least one of the image main body information and the target image.
3. The method according to claim 2, wherein The intention understanding result further includes a target requirement text, and the target requirement text is determined by using the large model to update the requirement text according to the image main body information; The calling the retrieval tool to perform retrieval according to the intention understanding result to obtain retrieval information includes: Calling the retrieval tool to perform retrieval according to the target requirement text to obtain the retrieval information.
4. The method according to claim 2, wherein, The determining the second reply information based on at least one of the image main body information and the target image includes: Using the large model to process the target image and the requirement text to obtain the second reply information, where the requirement main body keyword in the requirement text matches the image main body information.
5. The method according to claim 1, wherein, The image main body information is determined based on the following operations: Based on preset prompt information, using the large model to perform semantic understanding on the target image and the requirement text to obtain a semantic understanding result, where the preset prompt information represents the functional attributes of a preset tool; From at least one of the preset tools, calling a target tool that matches the functional attributes represented by the semantic understanding result to perform main body recognition on the target image to obtain the image main body information.
6. The method according to claim 5, wherein The target tools include a retrieval tool and a specified main body recognition tool, and the specified recognition tool is used to recognize the main body of an image with specified main body attributes; Calling a target tool that matches the functional attributes represented by the semantic understanding result to perform main body recognition on the target image to obtain the image main body information includes: Calling the retrieval tool to obtain associated images that meet the preset similarity condition with the target image; Calling the specified main body recognition tool to perform main body recognition on at least one of the target image and the associated images to obtain the image main body information.
7. The method according to claim 5, wherein, The functional attributes of the preset tool represent that the preset tool is used to perform main body recognition on the main body of an image with specified main body attributes, and the specified main body attributes include at least one of the following: Animal attribute, plant attribute, transportation vehicle attribute, building attribute.
8. An interaction device based on a large model, comprising: A receiving module, configured to receive a target image and a requirement text input by a target object, where the requirement text represents the processing requirement of the target object for the target image; A first obtaining module, configured to perform semantic understanding on the requirement text by using a large model based on image main body information, and obtain an intention understanding result, where the image main body information is determined by performing main body recognition on the target image; And A pushing module, configured to determine a reply message matching the processing requirement based on the intention understanding result, and push the reply message to the target object.
9. An intelligent agent of artificial intelligence, comprising: An input module, configured to receive input information; A processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method according to any one of claims 1 to 7 by invoking the large model to obtain output information; An output module, configured to output the output information obtained by the processing module.
10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 7.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
12. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 - 4.
Citation Information
Cited By
Interaction method and device based on multi-model cooperation, intelligent agent and electronic equipment
CN120805926A
Interaction method and device based on multi-model cooperation, agent and electronic device
CN120805926B