Information interaction method and apparatus, and computer device and computer-readable storage medium

By leveraging the collaborative efforts of large language models, visual tools, and multimodal recognizers, task instructions are optimized to generate target images that meet user needs. This addresses the inaccuracy of visual tasks in traditional methods and enhances the visual task processing capabilities of artificial intelligence systems.

WO2025246751A1PCT designated stage Publication Date: 2025-12-04BEIJING SMARTMORE INTELLIGENT TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/091148
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2025-04-25
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Traditional information interaction methods suffer from inaccurate task processing when handling visual tasks.

Method used

Task instructions are generated using a large language model, and the task is executed using visual tools. A multimodal recognizer generates visual feedback information, and the task instructions are optimized using the large language model. This process is repeated iteratively until a target image that meets the user's needs is obtained.

Benefits of technology

It improves the accuracy and adaptability of visual task processing, and significantly enhances the performance and accuracy of artificial intelligence systems when processing visual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025091148_04122025_PF_FP_ABST
    Figure CN2025091148_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to an information interaction method and apparatus, and a computer device and a computer-readable storage medium. The method comprises: in response to a user requirement, by means of a large language model, generating a task instruction corresponding to the user requirement (S202); by means of a visual tool, executing a target task corresponding to the task instruction, so as to obtain a candidate image (S204); by means of a multi-modal recognizer, generating visual feedback information of the candidate image, and sending the visual feedback information to the large language model (S206); and by means of the large language model, optimizing the task instruction on the basis of the visual feedback information, so as to obtain an optimized task instruction, and performing cyclic iteration processing on the basis of the optimized task instruction until a target image is obtained (S208).
Need to check novelty before this filing date? Find Prior Art

Description

Information exchange methods, devices, computer equipment and computer-readable storage media

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on May 31, 2024, with application number 2024107075784, entitled "Information Interaction Method, Apparatus, Computer Equipment and Computer-Readable Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, and in particular to an information interaction method, apparatus, computer device, and computer-readable storage medium. Background Technology

[0004] With the development of cognitive science and computer science, artificial intelligence models such as large language models have emerged. These models can handle visual tasks, such as image generation, reasoning segmentation, and reasoning detection. Traditional information exchange methods use large language models to send visual tasks to other tools to complete them.

[0005] However, traditional information exchange methods suffer from inaccurate task processing. Summary of the Invention

[0006] This application provides an information interaction method, apparatus, computer device, computer-readable storage medium, and computer program product.

[0007] In a first aspect, this application provides an information exchange method, executed by a computer device, comprising:

[0008] Responding to user needs, the system generates task instructions corresponding to user needs through a large language model.

[0009] Candidate images are obtained by executing the target task corresponding to the task instruction using visual tools;

[0010] Visual feedback information for candidate images is generated using a multimodal recognizer, and this visual feedback information is then sent to a large language model;

[0011] The task instructions are optimized based on visual feedback information using a large language model to obtain optimized task instructions. The optimized task instructions are then processed iteratively until the target image is obtained; the target image meets the user's needs.

[0012] Secondly, this application provides an information interaction device, comprising:

[0013] The model analysis module is used to respond to user needs and generate task instructions corresponding to user needs through a large language model.

[0014] The task execution module is used to execute the target task corresponding to the task instruction through visual tools and obtain candidate images;

[0015] The feedback information generation module is used to generate visual feedback information for candidate images through a multimodal recognizer and send the visual feedback information to the large language model; and

[0016] The model analysis module is also used to optimize task instructions based on visual feedback information using a large language model, to obtain optimized task instructions, and to perform iterative processing based on the optimized task instructions until the target image is obtained; the target image meets the user's needs.

[0017] Thirdly, this application provides a computer device including a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps in the above-described method.

[0018] Fourthly, this application provides a computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, implement the steps in the method described above.

[0019] Fifthly, this application provides a computer program product including computer-readable instructions that, when executed by a processor, implement the steps in the method described above. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the disclosed drawings without creative effort.

[0021] Figure 1 is an application environment diagram of an information interaction method provided in an embodiment of this application;

[0022] Figure 2 is a flowchart illustrating an information interaction method provided in an embodiment of this application;

[0023] Figure 3 is a comparative diagram of an information interaction method provided in an embodiment of this application and a traditional technology;

[0024] Figure 4 is a flowchart illustrating another information interaction method provided in an embodiment of this application;

[0025] Figure 5 is a comparative diagram of another information interaction method provided in the embodiments of this application and traditional technology;

[0026] Figure 6 is a schematic diagram of an information interaction method provided in an embodiment of this application;

[0027] Figure 7 is a schematic diagram comparing the image processing of the first ECO mechanism and the LISA model provided in the embodiments of this application;

[0028] Figure 8 is a schematic diagram comparing the image processing of the second ECO mechanism and the LISA model provided in the embodiments of this application;

[0029] Figure 9 is a schematic diagram comparing the image processing of the third ECO mechanism and the LISA model provided in the embodiments of this application;

[0030] Figure 10 is a schematic diagram comparing the image processing of the fourth ECO mechanism and the LISA model provided in the embodiments of this application;

[0031] Figure 11 is a schematic diagram comparing the first ECO mechanism provided in the embodiments of this application and image processing using only visual tools;

[0032] Figure 12 is a schematic diagram comparing the second ECO mechanism provided in the embodiments of this application with image processing using only visual tools;

[0033] Figure 13 is a schematic diagram comparing the third ECO mechanism provided in the embodiments of this application with image processing using only visual tools;

[0034] Figure 14 is a schematic diagram comparing the fourth ECO mechanism provided in the embodiments of this application with image processing using only visual tools;

[0035] Figure 15 is a comparative diagram of the fifth ECO mechanism provided in the embodiments of this application and image processing using only visual tools;

[0036] Figure 16 is a schematic diagram comparing the sixth ECO mechanism provided in the embodiments of this application with image processing using only visual tools;

[0037] Figure 17 is a structural block diagram of an information interaction device provided in an embodiment of this application;

[0038] Figure 18 is an internal structural diagram of a computer device provided in an embodiment of this application;

[0039] Figure 19 is an internal structural diagram of another computer device provided in an embodiment of this application;

[0040] Figure 20 is an internal structure diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] The information interaction method provided in this application embodiment is applied to the application environment shown in Figure 1. The computer equipment includes at least one of a terminal and a server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc.; the server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0043] Both the terminal and the server can be used independently for the information interaction method provided in the embodiments of this application.

[0044] For example, in response to user requests, the terminal generates task instructions corresponding to those requests using a large language model. Next, the terminal executes the target task corresponding to the task instructions using visual tools, obtaining candidate images. Then, the terminal generates visual feedback information for the candidate images using a multimodal recognizer and sends this visual feedback information to the large language model. Finally, the terminal optimizes the task instructions based on the visual feedback information using the large language model, obtaining optimized task instructions, and iteratively processes these optimized instructions until the target image is obtained; the target image satisfies the user's request.

[0045] For example, in response to a user's request, the server generates a task instruction corresponding to the user's request using a large language model. Next, the server executes the target task corresponding to the task instruction using a visual tool, obtaining candidate images. Then, the server generates visual feedback information for the candidate images using a multimodal recognizer and sends this visual feedback information to the large language model. Finally, the server optimizes the task instruction based on the visual feedback information using the large language model, obtaining an optimized task instruction, and iteratively processes this optimized instruction until the target image is obtained; the target image satisfies the user's request.

[0046] Terminals and servers can also work together to execute the information interaction methods provided in the embodiments of this application.

[0047] For example, in response to a user's request, the terminal generates a task instruction corresponding to the user's request using a large language model. Next, the terminal executes the target task corresponding to the task instruction using visual tools, obtaining candidate images. Then, the server retrieves the candidate images from the terminal. The server then generates visual feedback information for the candidate images using a multimodal recognizer and sends this visual feedback information to the large language model. Finally, the server optimizes the task instruction based on the visual feedback information using the large language model, obtaining an optimized task instruction, and iteratively processes this optimized instruction until the target image is obtained; the target image satisfies the user's request.

[0048] The aforementioned information interaction method, apparatus, computer equipment, computer-readable storage medium, and computer program product generate task instructions corresponding to user needs through a large language model; execute the target task corresponding to the task instructions through a visual tool to obtain candidate images; generate visual feedback information of the candidate images through a multimodal recognizer and send the visual feedback information to the large language model. That is, the multimodal recognizer acts as the "eye" to provide visual feedback to the large language model. The large language model can obtain the task execution status of the candidate images from the visual feedback information to optimize the task instructions, obtain more accurate optimized task instructions, and perform iterative processing based on the optimized task instructions until the target image that meets the user's needs is obtained. In other words, through the information interaction between the large language model, visual tools, and multimodal recognizer, a target image that better meets the user's needs can be obtained, improving the accuracy of image task processing.

[0049] As shown in Figure 2, this application provides an information interaction method, which will be described using an application to a computer device as an example. The information interaction method includes the following steps:

[0050] S202. Responding to user needs, generate task instructions corresponding to user needs through a large language model.

[0051] The computer equipment includes an information interaction system, which comprises a large language model (Brain LLM), visual tools, and a multimodal recognizer.

[0052] In some embodiments, the computer device includes an information interaction system that includes a Brain LLM, visual tools, a multimodal recognizer, and a user interface.

[0053] Information interaction systems simulate the eye-hand coordination ability of human infants during the learning process through the ECO (Evolutionary Eye-Hand Coordination) mechanism. By using visual feedback information, they optimize the interaction between large language models (LLMs) and visual tools, thereby better adapting to complex and ever-changing task requirements.

[0054] In this system, the large language model acts as the "brain" of the information interaction system, responsible for processing user input commands, invoking visual tools, and analyzing visual feedback from the multimodal recognizer to determine whether to continue optimization or terminate the task. Examples of large language models include high-performance models such as GPT-4-Turbo.

[0055] Visual tools, acting as the "hands" of an information interaction system, are used to execute specific image processing tasks, such as image generation, image segmentation, or image detection, based on the task instructions of a large language model. For example, a visual tool is at least one image processing model such as DallE-3 or Stable Diffusion XL.

[0056] As the "eyes" of an information interaction system, a multimodal recognizer can understand and interpret image content, transforming candidate images into visual feedback information that a large language model can understand, and is responsible for providing visual feedback information about the candidate images output by visual tools. For example, a multimodal recognizer is a multimodal model such as LLaVA-1.5, which has image description and logical reasoning capabilities.

[0057] The selection of visual tools and multimodal recognizers offers a degree of flexibility. Computer devices can choose different visual tools and multimodal recognizers based on the specific requirements of the target task and available resources. For example, when the target task requires high detail in image generation, GPT-4V is selected as the visual feedback provider. Furthermore, when the target task requires high logical reasoning ability, LLaVA-1.5 is selected as the multimodal recognizer.

[0058] The user interface is used to interact with users, and it can receive user input and display target images.

[0059] The computer device acquires user requests, responds to user requests by sending the user requests to a large language model, and generates task instructions corresponding to the user requests through the large language model.

[0060] The computer device obtains user input requirements through the user interface and sends the user requirements to the large language model, which then generates task instructions corresponding to the user requirements.

[0061] User requirements include at least one of image generation, image segmentation, or image detection, while task instructions refer to instructions for calling visual tools.

[0062] Specifically, a large language model is used to generate task instructions corresponding to user needs, and these task instructions are sent to a vision tool to instruct the vision tool to perform the target task.

[0063] S204. Execute the target task corresponding to the task instruction using visual tools to obtain candidate images.

[0064] The target task is an image processing task that matches the user's needs, such as at least one of image generation, image segmentation, or image detection. Candidate images are the images output by the visual tool after performing the target task.

[0065] For example, when the user's requirement is image generation, the vision tool executes the target task of image generation according to the task instructions corresponding to the user's requirement, and generates a candidate image.

[0066] For example, when the user's requirement is image segmentation, the vision tool executes the target task of image segmentation according to the task instructions corresponding to the user's requirement, and obtains candidate images after image segmentation, such as regions of interest such as human figures, animals, and buildings.

[0067] For example, when the user's requirement is image detection, the vision tool executes the target task of image detection according to the task instructions corresponding to the user's requirement, and obtains candidate images after image detection, such as face regions, action regions, etc.

[0068] Specifically, the target task corresponding to the task instruction is executed through a visual tool to obtain candidate images, which are then sent to the multimodal recognizer.

[0069] S206. Generate visual feedback information for candidate images through a multimodal recognizer and send the visual feedback information to the large language model.

[0070] Among them, visual feedback information is used to characterize relevant information of candidate images, including information such as image features of candidate images.

[0071] In this process, a multimodal recognizer is used to analyze candidate images to obtain visual feedback information about them.

[0072] S208. The task instructions are optimized based on visual feedback information using a large language model to obtain optimized task instructions. The optimized task instructions are then processed iteratively until the target image is obtained. The target image meets the user's requirements.

[0073] The process involves obtaining the target image when the iteration reaches the termination condition. The termination condition can be set as needed. Iteration termination conditions include at least one of the following: the candidate image meets user requirements, the number of iterations reaches a threshold, etc. The number of iterations can be set as needed.

[0074] The process involves iterative processing based on optimized task instructions. When a candidate image is determined to meet the user's requirements based on visual feedback information, that candidate image is used as the target image.

[0075] Specifically, the target image is sent to the user interface through a large language model, and the target image is then presented to the user through the user interface.

[0076] As can be seen, in this embodiment, a large language model generates task instructions corresponding to user needs; a visual tool executes the target task corresponding to the task instructions to obtain candidate images; a multimodal recognizer generates visual feedback information of the candidate images and sends the visual feedback information to the large language model. That is, the multimodal recognizer acts as the "eyes" providing visual feedback to the large language model. The large language model can obtain the task execution status of the candidate images from the visual feedback information to optimize the task instructions, obtaining more accurate optimized task instructions. Based on the optimized task instructions, iterative processing is performed until a target image that meets user needs is obtained. In other words, through the information interaction between the large language model, the visual tool, and the multimodal recognizer, a target image that better meets user needs can be obtained, improving the accuracy of image task processing. Furthermore, the large language model can obtain direct visual information about the output of the visual tool after executing the target task through visual feedback information, thereby better understanding the task requirements and adjusting the task instructions of the visual tool. This visual feedback mechanism significantly improves the adaptability and decision-making quality of the large language model in visual tasks. By simulating the human eye-hand coordination learning process, the embodiments of this application are expected to significantly improve the performance and accuracy of artificial intelligence systems when processing visual tasks, thereby promoting the application and development of artificial intelligence technology in multiple fields.

[0077] In some embodiments, as shown in Figure 3, user requirements include "highlighting the cue ball to strike other balls in the image" and "generating an image of a horse riding an astronaut." Traditional techniques, directly using visual tools for image processing, cannot generate accurate images. However, by employing an information interaction system, a large language model generates task instructions corresponding to the user's requirements. Visual tools then execute the target task based on these instructions, obtaining candidate images. A multimodal recognizer generates visual feedback information for the candidate images and sends this feedback to the large language model. The large language model can then obtain the task execution status of the candidate images from the visual feedback, optimizing the task instructions to obtain more accurate ones. This optimized task instructions are then iteratively processed until a target image that meets the user's requirements is obtained, resulting in a more accurate, consistent, and aesthetically pleasing target image.

[0078] In some embodiments, iterative processing is performed based on optimized task instructions until the target image is obtained, including:

[0079] During the current iteration, the target task corresponding to the optimized task instruction is executed using a visual tool to obtain new candidate images;

[0080] The visual feedback information of the new candidate images is generated by the multimodal recognizer and then sent to the large language model.

[0081] When a new candidate image is determined to meet the user's needs based on visual feedback information from the large language model, it is used as the target image.

[0082] Specifically, based on the visual feedback information of the new candidate image, the large language model determines whether the new candidate image meets the user's needs. When the new candidate image meets the user's needs, it is used as the target image.

[0083] As can be seen, in this embodiment, during the current iteration, the target task corresponding to the optimized task instruction is executed through a visual tool to obtain a new candidate image, as well as visual feedback information for generating the new candidate image through a multimodal recognizer. Therefore, when the large language model determines that the new candidate image meets the user's needs based on the visual feedback information of the new candidate image, it can accurately obtain the target image that meets the user's needs, thus improving the accuracy of task processing.

[0084] In some embodiments, after sending the visual feedback information of the new candidate image to the large language model, the method further includes:

[0085] When the large language model determines that the new candidate image does not meet the user's needs based on the visual feedback information of the new candidate image, the optimized task instructions are further optimized based on the visual feedback information of the new candidate image, and the process enters the next iterative processing cycle until the target image is obtained.

[0086] Understandably, when the large language model determines that a new candidate image does not meet the user's needs based on the visual feedback information of the new candidate image, the large language model can continue to optimize the optimized task instructions, thereby avoiding the output of images that do not meet the user's needs and improving the accuracy of task processing.

[0087] In some embodiments, task instructions are optimized based on visual feedback information using a large language model to obtain optimized task instructions, including:

[0088] Based on visual feedback information, a large language model is used to determine the task reinforcement instructions corresponding to the visual tools.

[0089] The task instructions are optimized based on the task enhancement instructions to obtain the optimized task instructions.

[0090] The task reinforcement instruction is used to instruct the visual tool to strengthen the execution of the target task. For example, the user requirement is "to generate an image of a horse riding on an astronaut," while the visual feedback information indicates that in the candidate images generated by the visual tool, "the horse is not riding on the astronaut." Therefore, based on the visual feedback information, the large language model determines that the corresponding task reinforcement instruction for the visual tool is "the horse must be riding on the astronaut." Then, based on this task reinforcement instruction, the task instruction is optimized to obtain the optimized task instruction, which instructs the visual tool to generate a candidate image of "the horse riding on the astronaut."

[0091] Among them, the optimized task instructions include task enhancement instructions.

[0092] As can be seen, in this embodiment, by using a large language model based on visual feedback information to determine the task enhancement instructions corresponding to the visual tool, the task instructions can be optimized more accurately based on these task enhancement instructions, resulting in more accurate optimized task instructions, and thus more accurately instructing the visual tool to perform the target task.

[0093] In some embodiments, the task instruction includes at least two sub-task instructions, and the candidate image includes a sub-candidate image corresponding to each sub-task instruction; the task instruction is optimized based on visual feedback information using a large language model to obtain an optimized task instruction, including:

[0094] The intermediate image is determined from each sub-candidate image by using a large language model based on the visual feedback information of each sub-candidate image;

[0095] The subtask instructions corresponding to the intermediate image are optimized based on the visual feedback information of the intermediate image to obtain the optimized task instructions.

[0096] The task instructions include at least two different sub-task instructions, which are used to instruct the vision tool to obtain at least two different versions of the sub-candidate images.

[0097] Specifically, a large language model is used to generate task instructions corresponding to user needs; for each sub-task instruction, a visual tool is used to execute the target task corresponding to the sub-task instruction to obtain sub-candidate images; a multimodal recognizer is used to generate visual feedback information of the sub-candidate images, and the visual feedback information of the sub-candidate images is sent to the large language model.

[0098] In some embodiments, a large language model is used to determine the intermediate image as the sub-candidate image whose sharpness meets a preset condition based on the visual feedback information of each sub-candidate image. For example, the sub-candidate image with the highest sharpness is determined as the intermediate image.

[0099] In some embodiments, the most colorful sub-candidate image is determined as the intermediate image from among the sub-candidate images based on the visual feedback information of each sub-candidate image using a large language model.

[0100] In some embodiments, the large language model can also use other methods to determine the intermediate image, such as randomly determining the intermediate image or selecting the last generated sub-candidate image as the intermediate image, etc., which are not limited here.

[0101] As can be seen, in this embodiment, the intermediate image is determined from each sub-candidate image by using the visual feedback information of each sub-candidate image through the large language model. That is, an evolutionary strategy is adopted to optimize the collaborative work between the large language model and the vision tool. The process of natural selection is simulated for at least two sub-candidate images to determine the intermediate image. Based on the intermediate image, further iteration is performed to optimize the sub-task instructions corresponding to the intermediate image to obtain more accurate optimized task instructions.

[0102] In some embodiments, an intermediate image is determined from the sub-candidate images based on visual feedback information of each sub-candidate image using a large language model, including:

[0103] Based on the visual feedback information of each sub-candidate image, the degree of matching between each sub-candidate image and the user's needs is determined by using a large language model;

[0104] The intermediate image is determined from each sub-candidate image based on the degree of matching between each sub-candidate image and the user's needs.

[0105] Understandably, the higher the degree of matching between the sub-candidate image and the user's needs, the better the sub-candidate image meets the user's needs.

[0106] In this process, the sub-candidate image with the highest matching degree is determined from each sub-candidate image using a large language model, and then used as the intermediate image.

[0107] In this process, the second-highest matching sub-candidate image is determined from each sub-candidate image using a large language model and used as the intermediate image.

[0108] In some embodiments, the large language model may also use other methods to determine intermediate images, which are not limited here.

[0109] As can be seen, in this embodiment, the large language model determines the degree of matching between each sub-candidate image and the user's needs based on the visual feedback information of each sub-candidate image. Based on the degree of matching between each sub-candidate image and the user's needs, the intermediate image can be more accurately determined from each sub-candidate image.

[0110] By generating multiple different versions of subtask instructions in each iteration (called "evolution"), a wider solution space is explored, allowing the information interaction system to consider multiple possibilities (diversity and exploration) at the same time. The subtask instructions corresponding to the best candidate images are selected for optimization based on the visual feedback information of the multimodal recognizer. That is, by simulating the process of natural selection, the best-performing subtask instructions are retained and further iterative improvements are made on this basis.

[0111] Since multiple versions of subtask instructions are generated in each iteration, this embodiment can evaluate multiple different optimization directions in a shorter time, thereby accelerating the process of finding the global optimal solution and improving the speed and efficiency of iterative optimization, especially when dealing with complex or high-dimensional problems.

[0112] In this embodiment, by exploring multiple solutions in parallel, the risk of getting stuck in local optima is reduced. Even if the performance of a certain version of the subtask instruction is poor, the information interaction system can still try other versions of the subtask instruction, thus having the opportunity to find a better solution to obtain a target image that better meets the user's needs.

[0113] In some embodiments, as shown in Figure 4, the user inputs "a horse is riding an astronaut," and the large language model generates a task instruction corresponding to the user's needs. The task instruction includes four sub-task instructions: Sub-task instruction 0: a photo of an astronaut standing on the lunar surface with a horse humorously placed on his back, as if the horse is riding on him; Sub-task instruction 1: an illustration of a comical scene, with a horse sitting on the astronaut's shoulders; Sub-task instruction 2: a cartoon of a horse playing; Sub-task instruction 3: rendering an astronaut in a spacesuit and a horse lying on the ground.

[0114] The system executes the target task corresponding to each sub-task instruction using visual tools, resulting in four sub-candidate images. A multimodal recognizer generates visual feedback information for each sub-candidate image and sends this information to a large language model. Based on the visual feedback information, the large language model determines the middle image as the second sub-candidate image. The sub-task instruction corresponding to the second sub-candidate image is then optimized based on its visual feedback information, resulting in an optimized task instruction. This optimized task instruction is then iteratively processed until the obtained sub-candidate image matches the user's requirements (meets user needs). The sub-candidate image that matches the user's requirements is then used as the target image and output.

[0115] In some embodiments, as shown in Figure 5, the user input generates an image of a horse riding an astronaut. Traditional techniques cannot generate an image that meets the user's needs. When the information interaction system adopts an evolutionary strategy, it executes the target task corresponding to each sub-task instruction through visual tools, obtaining at least two sub-candidate images. An intermediate image is then determined from these at least two sub-candidate images, and the sub-task instructions are further optimized based on this intermediate image. This iterative optimization can generate a target image that meets the user's needs. This evolutionary strategy has significant advantages in improving system performance, accelerating optimization speed, and avoiding local optima. Especially in complex environments where large language models and visual tools work together, this strategy can better leverage its advantages, thereby achieving more efficient and accurate task execution and significantly improving the performance and adaptability of the information interaction system when handling complex visual tasks.

[0116] In some embodiments, after iterative processing based on optimized task instructions until the target image is obtained, the method further includes:

[0117] In response to new user requirements for the target image, based on the new user requirements, return to the steps of generating task instructions corresponding to the user requirements through the large language model until a new target image is generated; the new target image satisfies the new user requirements.

[0118] In this process, after the target image is presented through the large language model, the computer device obtains new user requirements for the target image. In response to these new user requirements, the computer sends the new user requirements to the large language model and generates task instructions corresponding to these new user requirements through a loop iteration until a new target image is generated.

[0119] In this process, the computer device obtains new user requirements for the target image through the user interface, sends the new user requirements to the large language model, and generates task instructions corresponding to the new user requirements through the large language model, which are then iterated in a loop until a new target image is generated.

[0120] Specifically, a new target image is sent to the user interface through a large language model, and the user interface then presents the new target image to the user.

[0121] For example, after the target image is presented through the user interface, if the user is not satisfied with the target image, they can input new user requirements in the user interface. The new user requirements are then sent to the large language model through the user interface. The large language model generates task instructions corresponding to the new user requirements and iterates in a loop until a new target image is generated.

[0122] As can be seen, in this embodiment, in response to new user requirements for the target image, the large language model is further adjusted and optimized based on the new user requirements, and can more accurately generate new target images that meet the new user requirements.

[0123] In some embodiments, the method further includes:

[0124] When the target task is image segmentation and the visual feedback indicates that the user's needs are unreasonable, a prompt message is generated based on the visual feedback information using a large language model and presented in a perceptible manner.

[0125] Here, "perceptible mode" refers to a mode that the user can perceive. For example, perceptible modes include at least one of the following: images, videos, and voice.

[0126] Specifically, when the target task is image segmentation, if the segmentation object required by the user is not found in the candidate image by the multimodal recognizer, visual feedback information indicating that the user's requirement is unreasonable is generated and sent to the large language model.

[0127] As shown in Figure 6, when a user inputs an image and the user's request is "Please cut the banana into segments in this image", the ECO mechanism can effectively identify this unreasonable request and avoid generating an incorrect image.

[0128] As can be seen, in this embodiment, when the target task is image segmentation and the visual feedback information indicates that the user's needs are unreasonable, that is, the multimodal recognizer can provide accurate visual feedback, which helps to reduce the illusion phenomenon caused by unreasonable user input. Furthermore, by generating prompt information based on the visual feedback information through the large language model, incorrect image segmentation results are avoided, thus improving the accuracy of information interaction.

[0129] In some embodiments, another information interaction method is also provided, applied to a computer device, the information interaction method comprising the following steps:

[0130] Step A1: In response to user needs, generate task instructions corresponding to user needs through the large language model; the task instructions include at least two sub-task instructions.

[0131] Step A2: For each subtask instruction, execute the target task corresponding to the subtask instruction using a visual tool to obtain the sub-candidate image;

[0132] Step A3: Generate visual feedback information for each sub-candidate image using a multimodal recognizer, and send the visual feedback information to the large language model;

[0133] Among them, when the target task is image segmentation and the visual feedback information indicates that the user's needs are unreasonable, the large language model generates prompt information based on the visual feedback information and presents the prompt information in a perceptible manner.

[0134] Step A4: Based on the visual feedback information of each sub-candidate image, determine the degree of matching between each sub-candidate image and the user's needs using a large language model; and determine the intermediate image from each sub-candidate image based on the degree of matching between each sub-candidate image and the user's needs.

[0135] Step A5: Based on the visual feedback information of the intermediate image, determine the task enhancement instructions corresponding to the visual tool using the large language model; optimize the sub-task instructions corresponding to the intermediate image based on the task enhancement instructions to obtain the optimized task instructions.

[0136] Step A6: In the current iteration process, the target task corresponding to the optimized task instruction is executed through a vision tool to obtain a new candidate image;

[0137] Step A7: Generate visual feedback information for new candidate images using a multimodal recognizer, and send the visual feedback information for new candidate images to the large language model;

[0138] Step A8: When the large language model determines that the new candidate image meets the user's needs based on the visual feedback information of the new candidate image, the new candidate image is used as the target image; when the large language model determines that the new candidate image does not meet the user's needs based on the visual feedback information of the new candidate image, the optimized task instructions are optimized based on the visual feedback information of the new candidate image, and the process enters the next iterative processing cycle until the target image is obtained; the target image meets the user's needs.

[0139] Step A9: In response to new user requirements for the target image, based on the new user requirements, return to the step of generating task instructions corresponding to the user requirements through the large language model until a new target image is generated; the new target image satisfies the new user requirements.

[0140] Referring to Table 1 and Figures 7 to 10, it can be seen that in this embodiment, the performance of the inference segmentation task can be improved. In the inference segmentation task, after adopting the Evolutionary Eye-Hand Coordination (EHCO) mechanism, the gIoU (Generalized Intersection over Union) index increased from 53.3% to 59.2%, and the cIoU (Centered Intersection over Union) index increased from 60.8% to 61.2%. This is a significant performance improvement, demonstrating the effectiveness of the ECO mechanism in improving image segmentation accuracy.

[0141] Table 1:

[0142] X-decoder, SEEM, and LISA are all models (visual tools) used for inference segmentation tasks, capable of segmenting objects in images based on textual cues. "LISA in Evolutionary Eye-Hand Coordination w SLS" indicates that the visual tool LISA employs the aforementioned ECO mechanism (without an evolutionary strategy), while "LISA in Evolutionary Eye-Hand Coordination w ES" indicates that the visual tool LISA employs the aforementioned ECO mechanism (with an evolutionary strategy).

[0143] Referring to Table 2 and Figures 11 to 16, it can be seen that the consistency of the reference image generation task is enhanced:

[0144] In image generation tasks, user research results show that adopting the ECO mechanism significantly improves the consistency between generated images and user needs. For example, when dealing with a scene of "an astronaut riding a horse," the ECO mechanism is able to better understand the user's creative needs and generate an image that is more consistent with those needs.

[0145] Table 2:

[0146] Among them, DallE-3 and GPT-4 (visual tools) are text-to-image generation models that can generate corresponding images based on user-input text descriptions. DallE-3 in ECO w SLS indicates that the visual tool DallE-3 uses the aforementioned ECO mechanism (without an evolutionary strategy), while DallE-3 in ECO w ES indicates that the visual tool DallE-3 uses the aforementioned ECO mechanism (with an evolutionary strategy). As shown in Table 2, DallE-3 in ECO w SLS and DallE-3 in ECO w ES both scored highly in terms of consistency with user needs and aesthetics, indicating that the image processing of DallE-3 in ECO w SLS and DallE-3 in ECO w ES better meets user needs and is more aesthetically pleasing.

[0147] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0148] Based on the same inventive concept, this application also provides an information interaction device. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more information interaction device embodiments provided below can be found in the limitations of the information interaction method above, and will not be repeated here.

[0149] As shown in Figure 17, this application embodiment provides an information interaction device 1700, including:

[0150] The model analysis module 1702 is used to respond to user needs and generate task instructions corresponding to user needs through a large language model.

[0151] The task execution module 1704 is used to execute the target task corresponding to the task instruction through a visual tool to obtain candidate images;

[0152] The feedback information generation module 1706 is used to generate visual feedback information of candidate images through a multimodal recognizer and send the visual feedback information to the large language model;

[0153] The model analysis module 1702 is also used to optimize the task instructions based on visual feedback information through a large language model, obtain optimized task instructions, and perform iterative processing based on the optimized task instructions until the target image is obtained; the target image meets the user's needs.

[0154] In some embodiments, regarding the iterative processing based on optimized task instructions until the target image is obtained, the model analysis module 1702 is specifically used for:

[0155] During the current iterative processing, the target task corresponding to the optimized task instruction is executed through visual tools to obtain new candidate images;

[0156] The visual feedback information of the new candidate images is generated by the multimodal recognizer and then sent to the large language model.

[0157] When a new candidate image is determined to meet the user's needs based on visual feedback information from the large language model, it is used as the target image.

[0158] In some embodiments, the model analysis module 1702 is further configured to optimize the optimized task instructions based on the visual feedback information of the new candidate image through the large language model, and enter the next iterative processing process until the target image is obtained when the new candidate image does not meet the user's needs.

[0159] In some embodiments, in optimizing task instructions based on visual feedback information using a large language model to obtain optimized task instructions, the model analysis module 1702 is specifically used for:

[0160] Based on visual feedback information, a large language model is used to determine the task reinforcement instructions corresponding to the visual tools.

[0161] The task instructions are optimized based on the task enhancement instructions to obtain the optimized task instructions.

[0162] In some embodiments, the task instruction includes at least two sub-task instructions, and the candidate image includes a sub-candidate image corresponding to each sub-task instruction; in optimizing the task instruction based on visual feedback information using a large language model to obtain an optimized task instruction, the model analysis module 1702 is specifically used for:

[0163] The intermediate image is determined from each sub-candidate image by using a large language model based on the visual feedback information of each sub-candidate image;

[0164] The subtask instructions corresponding to the intermediate image are optimized based on the visual feedback information of the intermediate image to obtain the optimized task instructions.

[0165] In some embodiments, in determining the intermediate image from the various sub-candidate images based on the visual feedback information of each sub-candidate image using a large language model, the model analysis module 1702 is specifically used for:

[0166] Based on the visual feedback information of each sub-candidate image, the degree of matching between each sub-candidate image and the user's needs is determined by using a large language model;

[0167] The intermediate image is determined from each sub-candidate image based on the degree of matching between each sub-candidate image and the user's needs.

[0168] In some embodiments, the model analysis module 1702 is further configured to, in response to new user requirements for the target image, return to the step of executing the task instruction corresponding to the user requirements through the large language model, until a new target image is generated; the new target image satisfies the new user requirements.

[0169] In some embodiments, the model analysis module 1702 is further configured to generate prompt information based on the visual feedback information through a large language model when the target task is an image segmentation task and the visual feedback information indicates that the user's needs are unreasonable, and to present the prompt information in a perceptible manner.

[0170] Each module in the aforementioned information interaction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0171] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in Figure 18. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The database of the computer device stores at least one of user requirements, candidate images, and target images. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement the steps in the above-described information interaction method.

[0172] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as shown in Figure 19. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer-readable instructions are executed by the processor, they implement the steps in the aforementioned information interaction method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen; the input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs or touchpads set on the casing of the computer device, or external keyboards, touchpads or mice, etc.

[0173] Those skilled in the art will understand that the structures shown in Figure 18 or Figure 19 are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0174] In some embodiments, a computer device is provided, the computer device including a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps in the above method embodiments.

[0175] In some embodiments, as shown in FIG20, an internal structure diagram of a computer-readable storage medium is provided. The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps in the above-described method embodiments.

[0176] In some embodiments, a computer program product is provided, the computer program product including computer-readable instructions that, when executed by a processor, implement the steps in the above method embodiments.

[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0180] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An information interaction method applied to a computer device, comprising: generating, by a large language model, a task instruction corresponding to a user demand in response to the user demand; executing, by a visual tool, a target task corresponding to the task instruction to obtain a candidate image; generating, by a multi-modal recognizer, visual feedback information of the candidate image and sending the visual feedback information to the large language model; and optimizing, by the large language model, the task instruction based on the visual feedback information to obtain an optimized task instruction, and performing cyclic iteration processing based on the optimized task instruction until a target image is obtained; the target image meets the user demand.

2. The method of claim 1, wherein, The cyclic iteration processing based on the optimized task instruction until the target image is obtained comprises: in a current cyclic iteration processing process, executing, by the visual tool, a target task corresponding to the optimized task instruction to obtain a new candidate image; generating, by the multi-modal recognizer, visual feedback information of the new candidate image and sending the visual feedback information of the new candidate image to the large language model; and when it is determined by the large language model based on the visual feedback information of the new candidate image that the new candidate image meets the user demand, the new candidate image is taken as the target image.

3. The method of claim 2, wherein, After the visual feedback information of the new candidate image is sent to the large language model, the method further comprises: when it is determined by the large language model based on the visual feedback information of the new candidate image that the new candidate image does not meet the user demand, optimizing, by the large language model, the optimized task instruction based on the visual feedback information of the new candidate image, and entering a next cyclic iteration processing process until the target image is obtained.

4. The method of claim 1, wherein, The optimization of the task instruction by the large language model based on the visual feedback information to obtain the optimized task instruction comprises: determining, by the large language model based on the visual feedback information, a task reinforcement instruction corresponding to the visual tool; and optimizing the task instruction based on the task reinforcement instruction to obtain the optimized task instruction.

5. The method of claim 1, wherein, The task instruction comprises at least two sub-task instructions, and the candidate image comprises a sub-candidate image corresponding to each sub-task instruction; the optimization of the task instruction by the large language model based on the visual feedback information to obtain the optimized task instruction comprises: determining, by the large language model based on the visual feedback information of each sub-candidate image, an intermediate image from each sub-candidate image; and optimizing the sub-task instruction corresponding to the intermediate image based on the visual feedback information of the intermediate image to obtain the optimized task instruction.

6. The method of claim 5, wherein, The determination of the intermediate image from each sub-candidate image by the large language model based on the visual feedback information of each sub-candidate image comprises: determining, by the large language model based on the visual feedback information of each sub-candidate image, a matching degree between each sub-candidate image and the user demand; and determine an intermediate image from the respective sub-candidate images based on the degree of matching between each sub-candidate image and the user demand.

7. The method of claim 5, wherein, Before the step of determining an intermediate image from the respective sub-candidate images based on the visual feedback information of each sub-candidate image by the large language model, the method further comprises: for each sub-task instruction included in the task instruction, executing a target task corresponding to the sub-task instruction by the visual tool to obtain the sub-candidate image; and generating visual feedback information of the sub-candidate image by the multi-modal recognizer, and sending the visual feedback information of the sub-candidate image to the large language model.

8. The method of claim 5, wherein, The step of determining an intermediate image from the respective sub-candidate images based on the visual feedback information of each sub-candidate image by the large language model comprises: determining, by the large language model, a sub-candidate image whose clarity meets a preset condition as an intermediate image from the respective sub-candidate images based on the visual feedback information of each sub-candidate image.

9. The method of claim 1, wherein, After the step of performing cyclic iteration processing based on the optimized task instruction until a target image is obtained, the method further comprises: in response to a new user demand for the target image, based on the new user demand, returning to the step of generating a task instruction corresponding to the user demand by the large language model until a new target image is generated; the new target image meets the new user demand.

10. The method of claim 1, wherein, The method further comprises: when the target task is an image segmentation task and the visual feedback information indicates that the user demand is unreasonable, generating prompt information based on the visual feedback information by the large language model, and presenting the prompt information in a perceptible manner.

11. An information interaction device, characterized in that, comprises: a model analysis module configured to generate a task instruction corresponding to a user demand by a large language model in response to the user demand; a task execution module configured to execute a target task corresponding to the task instruction by a visual tool to obtain a candidate image; a feedback information generation module configured to generate visual feedback information of the candidate image by a multi-modal recognizer, and send the visual feedback information to the large language model; and the model analysis module is further configured to optimize the task instruction based on the visual feedback information by the large language model to obtain an optimized task instruction, and perform cyclic iteration processing based on the optimized task instruction until a target image is obtained; the target image meets the user demand.

12. The apparatus of claim 11, wherein, In the aspect of performing cyclic iteration processing based on the optimized task instruction until a target image is obtained, the model analysis module is specifically configured to: in a current cyclic iteration processing process, execute a target task corresponding to the optimized task instruction by the visual tool to obtain a new candidate image; generate visual feedback information of the new candidate image by the multi-modal recognizer, and send the visual feedback information of the new candidate image to the large language model; and when it is determined that the new candidate image meets the user demand based on the visual feedback information of the new candidate image by the large language model, take the new candidate image as a target image.

13. The apparatus of claim 12, wherein, After the visual feedback information of the new candidate image is sent to the large language model, the model analysis module is specifically configured to: When it is determined by the large language model that the new candidate image does not meet the user demand based on the visual feedback information of the new candidate image, the large language model is used to optimize the optimized task instruction based on the visual feedback information of the new candidate image, and a next loop iteration process is entered until a target image is obtained.

14. The apparatus according to claim 11, characterized in that, In the aspect of optimizing the task instruction based on the visual feedback information by the large language model to obtain an optimized task instruction, the model analysis module is specifically configured to: determine a task reinforcement instruction corresponding to the visual tool based on the visual feedback information by the large language model; and optimize the task instruction based on the task reinforcement instruction to obtain an optimized task instruction.

15. The apparatus of claim 11, wherein, The task instruction includes at least two sub-task instructions, and the candidate image includes a sub-candidate image corresponding to each sub-task instruction; in the aspect of optimizing the task instruction based on the visual feedback information by the large language model to obtain an optimized task instruction, the model analysis module is specifically configured to: determine an intermediate image from each sub-candidate image based on the visual feedback information of each sub-candidate image by the large language model; and optimize the sub-task instruction corresponding to the intermediate image based on the visual feedback information of the intermediate image to obtain an optimized task instruction.

16. The apparatus of claim 15, wherein, In the aspect of determining an intermediate image from each sub-candidate image based on the visual feedback information of each sub-candidate image by the large language model, the model analysis module is specifically configured to: determine a matching degree between each sub-candidate image and the user demand based on the visual feedback information of each sub-candidate image by the large language model; and determine an intermediate image from each sub-candidate image based on the matching degree between each sub-candidate image and the user demand.

17. The apparatus of claim 11, wherein, After the loop iteration process based on the optimized task instruction is performed until a target image is obtained, the model analysis module is further configured to: in response to a new user demand for a target image, based on the new user demand, return to execute the step of generating a task instruction corresponding to the user demand by the large language model until a new target image is generated; and the new target image meets the new user demand.

18. The apparatus of claim 11, wherein, When the target task is an image segmentation task and the visual feedback information indicates that the user demand is unreasonable, the model analysis module is further configured to generate prompt information based on the visual feedback information by the large language model, and present the prompt information in a perceptible manner. 19.A computer device, comprising a memory and a processor, wherein the memory stores computer readable instructions, and the computer device is characterized in that, The processor executes the computer readable instructions to implement the steps of the method of any one of claims 1-10.

20. A computer-readable storage medium having stored thereon computer-readable instructions, wherein, The computer readable instructions are executed by the processor to implement the steps of the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN116843795A

  • Image generation method and data processing method for image generation

    CN117409109A

  • Information interaction method and device, electronic equipment and storage medium

    CN117690002A

  • Information interaction method and device, computer equipment and computer readable storage medium

    CN118447121A