Information interaction method, device, electronic device and storage medium

Through a large language model, the requirements description text are processed, visual task attributes are generated and image processing is automatically executed, which solves the problem of inefficient image processing requirements in the prior art, realizes efficient and accurate image processing and improves user experience.

CN117690002BActive Publication Date: 2025-05-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311694607.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-05-23
Estimated Expiration
2043-12-11

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process image processing requirements, especially in the field of computer vision, where users need to manually select diverse visual processing resources, resulting in complex operations and inefficient efficiency.

Method used

By using a large language model to process the requirements description text, visual task attributes matching the image processing intention are generated, so that the image processing results are automatically determined and feedback information is displayed on the interactive interface.

Benefits of technology

It improves the efficiency and accuracy of image processing, simplifies user operation processes, reduces dependence on diversified visual processing resources, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117690002B_ABST
    Figure CN117690002B_ABST
Patent Text Reader

Abstract

The present disclosure provides an information interaction method, device, electronic device and storage medium, which relate to the field of artificial intelligence technology, specifically computer vision, deep learning, large models and other technical fields, and can be applied to artificial intelligence content generation, human-computer interaction and other scenarios. The specific implementation scheme is: in response to obtaining a demand description text, using a large language model to process the demand description text, and obtain visual task attributes that match the image processing intent represented by the demand description text, wherein the demand description text is associated with the image to be processed; according to the visual task attributes, determine the image processing result related to the image to be processed; generate feedback information according to the image processing result; and display the feedback information on the interactive interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, deep learning, large models and other technical fields, and can be applied to scenarios such as artificial intelligence content generation and human-computer interaction. Background Art

[0002] With the rapid development of computer vision technology, it is possible to process images such as photos and videos through computer vision technology. For example, images can be processed based on computer vision functions such as target detection and image classification. Computer vision technology is widely used in scenes such as film and television product production and intelligent security. Summary of the invention

[0003] The present disclosure provides an information interaction method, device, electronic device and storage medium.

[0004] According to one aspect of the present disclosure, an information interaction method is provided, including: in response to obtaining a requirement description text, using a large language model to process the requirement description text to obtain visual task attributes that match the image processing intent represented by the requirement description text, wherein the requirement description text is associated with an image to be processed; determining an image processing result related to the image to be processed based on the visual task attributes; generating feedback information based on the image processing result; and displaying the feedback information on an interactive interface.

[0005] According to another aspect of the present disclosure, an information interaction device is provided, including: a visual task attribute acquisition module, which is used to process the demand description text using a large language model in response to obtaining a demand description text, and obtain visual task attributes that match the image processing intent represented by the demand description text, wherein the demand description text is associated with an image to be processed; an image processing result determination module, which is used to determine an image processing result related to the image to be processed based on the visual task attributes; a feedback information generation module, which is used to generate feedback information based on the image processing result; and a display module, which is used to display the feedback information on an interactive interface.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided according to an embodiment of the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method provided according to an embodiment of the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method provided according to the embodiment of the present disclosure is implemented.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0011] Figure 1 An exemplary system architecture to which the information interaction method and device according to an embodiment of the present disclosure can be applied is schematically shown;

[0012] Figure 2 A flowchart of an information interaction method according to an embodiment of the present disclosure is schematically shown;

[0013] Figure 3 The schematic diagram shows a principle diagram of an information interaction method according to an embodiment of the present disclosure;

[0014] Figure 4 The application scenario diagram of the information interaction method according to the embodiment of the present disclosure is schematically shown;

[0015] Figure 5 A diagram schematically shows an application scenario of an information interaction method according to another embodiment of the present disclosure;

[0016] Figure 6 A block diagram schematically shows an information interaction device according to an embodiment of the present disclosure; and

[0017] Figure 7 A block diagram of an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0018] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] In the technical solution disclosed in the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and do not violate public order and good morals.

[0020] Embodiments of the present disclosure provide an information interaction method, apparatus, electronic device, and storage medium. The information interaction method includes: in response to obtaining a requirement description text, using a large language model to process the requirement description text to obtain visual task attributes that match the image processing intent represented by the requirement description text, wherein the requirement description text is associated with an image to be processed; determining an image processing result related to the image to be processed based on the visual task attributes; generating feedback information based on the image processing result; and displaying the feedback information on an interactive interface.

[0021] According to the embodiments of the present disclosure, by obtaining the demand description text and using a large language model to understand the semantics of the natural language representation in the demand description text, the obtained visual task attributes can be matched with the demand intent represented by the demand description text, and then the corresponding visual task can be performed through the visual task attributes, and the image processing results can be conveniently determined, avoiding the target object from obtaining the image processing results corresponding to the demand intent through diverse and complex visual processing resources, saving the complex operation process generated by selecting visual processing resources, thereby improving the image processing efficiency, and generating and displaying feedback information based on the image processing results can improve the timeliness of the target object obtaining the processing results, thereby improving the user experience.

[0022] A large language model (LLM) may include a deep learning model trained using a large amount of text data, which can be used to understand the meaning of language text and generate natural language text. The large language model can handle a variety of natural language tasks, such as text classification, question and answer, conversation, etc. Since large language models usually contain billions of parameters, large-scale parameters can help large language models learn complex patterns in natural language data, thereby having more outstanding performance in natural language processing (NLP) tasks. In addition, the disclosed embodiments can develop extended plug-in tools (service resources) based on the large language model to combine the large language model with computer vision service resources, so that the large language model can perform computer vision tasks such as detection, classification, and understanding of images and videos.

[0023] Figure 1 An exemplary system architecture to which the information interaction method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0024] It should be noted that Figure 1The examples shown are only examples of system architectures to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the information interaction method and apparatus can be applied may include a terminal device, but the terminal device may implement the information interaction method and apparatus provided by the embodiments of the present disclosure without interacting with the server.

[0025] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0026] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software, etc. (only as examples).

[0027] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0028] The server 105 may be a server that provides various services, such as a background management server (only an example) that provides support for the content browsed by the user using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0029] Alternatively, the server can also be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server can also be a server for a distributed system, or a server combined with blockchain.

[0030] It should be noted that the information interaction method provided in the embodiment of the present disclosure can also be generally performed by the server 105. Accordingly, the information interaction device provided in the embodiment of the present disclosure can generally be set in the server 105. The information interaction method provided in the embodiment of the present disclosure can also be performed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the information interaction device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0031] Alternatively, the information interaction method provided in the embodiment of the present disclosure may also be generally executed by the terminal device 101, 102, or 103. Accordingly, the information interaction apparatus provided in the embodiment of the present disclosure may also be arranged in the terminal device 101, 102, or 103.

[0032] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0033] Figure 2 The flowchart of the information interaction method according to the embodiment of the present disclosure is schematically shown.

[0034] like Figure 2 As shown, the information interaction method includes operations S210 to S240.

[0035] In operation S210, in response to acquiring the requirement description text, the requirement description text is processed using a large language model to obtain visual task attributes that match the image processing intent represented by the requirement description text, wherein the requirement description text is associated with the image to be processed.

[0036] In operation S220, an image processing result related to the image to be processed is determined according to the visual task attribute.

[0037] In operation S230, feedback information is generated according to the image processing result.

[0038] In operation S240, feedback information is displayed on the interactive interface.

[0039] According to the embodiments of the present disclosure, the demand description text may be natural language information used to characterize the image processing requirements of the target object, and the demand description text may be obtained based on the input operation of the target object to the information input box in the interactive interface. However, it is not limited to this, and the demand description text may also be obtained by other means, for example, the sound information expressing the demand description text of the target object may be collected based on a sound collection device, and the demand description text represented by the sound information may be obtained by recognizing the sound information through a speech recognition algorithm. The embodiments of the present disclosure do not limit the specific method of obtaining the demand description text.

[0040] According to an embodiment of the present disclosure, the image to be processed may include an image that needs to be processed that describes the text representation of the requirement. The image to be processed may be obtained through an image input operation of the target object, or may also include being queried from an associated storage device or storage means based on the semantics of the text representation of the requirement. The embodiment of the present disclosure does not limit the specific method of obtaining the image to be processed, as long as the image that needs to be processed that describes the text representation of the requirement can be obtained.

[0041] According to an embodiment of the present disclosure, a large language model (LLM) may include a pre-trained model, which may identify the image processing intent of the text representation of the demand description based on the natural language understanding capability of the large language model, and then generate visual task attributes that match the image processing intent of the text representation of the demand description, thereby avoiding the time-consuming operation caused by the target object performing selection operations on a large number of image processing service resources, saving operation time and complexity, and improving the overall image processing efficiency.

[0042] According to an embodiment of the present disclosure, the visual task attributes may include attribute information related to the execution of an image processing task (or visual task), for example, may include task configuration parameters of the visual task, the visual task type, and the like.

[0043] According to an embodiment of the present disclosure, determining an image processing result related to an image to be processed according to visual task attributes may include executing a visual task generated based on the visual task attributes, thereby achieving image processing for the image to be processed and obtaining an image processing result.

[0044] According to the embodiments of the present disclosure, the image processing result may include any type of information, for example, it may include text information describing the image to be processed, or it may also include an image or image block obtained after performing any visual task such as target detection, image classification, image clipping, etc. on the image to be processed. The embodiments of the present disclosure do not limit the specific information type of the image processing result.

[0045] According to the embodiments of the present disclosure, the feedback information can be used to characterize the image processing result. For example, the feedback information can include any type of information such as text, logo, image block, etc. that characterizes the image processing result. The embodiments of the present disclosure do not limit the specific information type of the feedback information, and those skilled in the art can select it according to actual needs.

[0046] According to an embodiment of the present disclosure, the interactive interface may include an interface for the target object to browse information, such as a display screen of a smart phone. Displaying the feedback information on the interactive interface may include displaying the feedback information by rendering the feedback information on the interactive interface, or may also include playing the feedback information in an audio format through an audio playback device associated with the interactive interface, so that the feedback information in the audio format can assist the target object to conveniently obtain the image processing result for the image to be processed.

[0047] It should be noted that the information processing operations in any embodiment of the present disclosure, including but not limited to information acquisition operations and image processing operations, are all performed after obtaining the authorization of the relevant user. And after obtaining the information, necessary encryption measures are adopted to protect the security of the information and avoid information leakage. The data obtained according to the method provided in the embodiment of the present disclosure, including but not limited to image processing results and feedback information, are all reviewed in accordance with relevant laws and regulations or specifications, and displayed after passing the review.

[0048] According to an embodiment of the present disclosure, the information interaction method may further include: in response to receiving an image to be processed input by a target object, performing image processing intent detection on the image to be processed to obtain a requirement description text.

[0049] According to the embodiments of the present disclosure, target detection can be performed on the image to be processed based on the target detection algorithm, and image processing intent detection can be implemented on the image to be processed based on the obtained detection result to obtain the requirement description text.

[0050] For example, the image to be processed may contain Bridge A, and the detection result obtained by performing target detection on the image to be processed may be "Bridge A". According to the detection result, the requirement description text representing the image processing intention may be determined to be "building recognition".

[0051] According to an embodiment of the present disclosure, image processing intent detection may be performed on the image to be processed based on other types of methods. For example, text in the image to be processed may be identified based on OCR technology, and the identified text may be used as demand description text.

[0052] According to the embodiments of the present disclosure, by performing image processing intention detection on the image to be processed input by the target object and then obtaining the requirement description text, the image processing intention represented by the image to be processed can be automatically and preliminarily identified without the target object performing the input operation of inputting the requirement description text, thereby saving the image processing operation steps and improving the subsequent image processing efficiency and information interaction efficiency.

[0053] According to an embodiment of the present disclosure, the information interaction method may further include: updating the received requirement description text according to a preset requirement prompt template to obtain a new requirement description text.

[0054] According to an embodiment of the present disclosure, a requirement prompt template (or prompt template) can be used to help a large language model understand the intent semantic attributes of the text representation of the requirement description, and control the large language model to accurately predict a visual task attribute prompt tag sequence that matches the image processing intent. The prompt tag sequence can include any type of prompt tags such as characters, fields, and text.

[0055] According to an embodiment of the present disclosure, updating a received requirement description text according to a preset requirement prompt template may include adding keywords related to the image processing intent in the requirement description text to the requirement prompt template, and the resulting new requirement description text may include a prompt tag sequence for controlling the large language model to perform precise predictions, thereby enabling the large language model to more accurately understand the image processing intent represented by the requirement description text, and control the large language model to accurately predict visual attribute tasks that match the image processing intent, and further, by accurately predicting the visual task attributes, the image processing results that match the requirement description text can be determined, thereby achieving accurate and efficient processing of the image to be processed.

[0056] According to an embodiment of the present disclosure, determining an image processing result related to an image to be processed based on visual task attributes may include: generating a visual task based on the visual task attributes and the image to be processed; executing the visual task based on at least one service resource associated with the visual task attributes to obtain a task execution result; and obtaining an image processing result based on the task execution result.

[0057] According to an embodiment of the present disclosure, a visual task may include a visual task type and visual task configuration parameters for performing image processing on an image to be processed. By calling service resources associated with the visual task attributes to execute the visual task, it is possible to execute the visual task on the image to be processed according to the image processing intention represented by the demand description text, thereby avoiding problems such as too long operation steps and service resource call errors caused by the target object calling service resources by performing selection operations or calling the service resource interface by manually sending control instructions to execute the visual task. This saves the learning cost required for the target object to perform image processing on the image to be processed, reduces the complexity of the operation process, and improves image processing efficiency and accuracy.

[0058] According to an embodiment of the present disclosure, a visual task attribute includes multiple subtask attributes and execution dependencies between the multiple subtask attributes, a visual task includes sub-visual tasks corresponding to the subtask attributes, and service resources are associated with the subtask attributes.

[0059] According to the embodiments of the present disclosure, subtask attributes can characterize the attributes of subtasks such as task configuration parameters of subtasks of visual tasks, subtask objects, visual processing types of subtasks, and service resource identifiers of service resources that need to be called to execute subtasks. Multiple subtasks are characterized by execution dependencies to characterize the logical relationship of executing visual tasks. By using a large language model to process the demand description text, and then obtaining visual task attributes containing multiple subtask attributes and execution dependencies, the execution of the visual task can be finely disassembled to clearly characterize the execution process of the visual task. By calling the service resources associated with the subtask attributes to execute sub-visual tasks, fine-grained execution of visual tasks can be achieved, thereby improving the execution accuracy of visual tasks.

[0060] According to an embodiment of the present disclosure, executing a visual task according to at least one service resource associated with a visual task attribute to obtain a task execution result includes: according to an execution dependency, calling a kth service resource associated with a kth subtask attribute to execute a kth sub-visual task to obtain a kth sub-task execution result, wherein k>1 and k is an integer.

[0061] According to an embodiment of the present disclosure, the kth sub-visual task may include a task execution result determined according to the kth sub-task attribute, and may include one or more sub-task execution results. The kth sub-task attribute may include a kth service resource identifier of a service resource suitable for executing the kth sub-visual task, and the kth service resource may be called by the kth service resource identifier, and the kth sub-visual task may be executed based on the configuration parameters of the kth sub-task attribute, so as to obtain a kth sub-task execution result that matches the image processing intent represented by the requirement description text.

[0062] For example, when the kth subtask attribute indicates an image segmentation task and indicates the image region parameter for image segmentation of the processed image, the kth visual subtask can be generated based on the image region parameter and the subtask type of the visual subtask. The kth service resource is called through the kth service resource identifier to execute the kth visual subtask, thereby obtaining the execution result of the kth subtask.

[0063] According to an embodiment of the present disclosure, the kth visual sub-task may also be determined according to the k-1th sub-task execution result and the kth sub-task attribute.

[0064] For example, the kth subtask attribute can be used to represent the identification of the traffic light state of the traffic light in the image to be processed, and the execution result of the k-1th subtask can be the traffic light image block obtained after image segmentation of the traffic light image area in the image to be processed (image segmentation subtask execution result). The traffic light state recognition subtask (kth sub-visual task) can be generated based on the traffic light image block obtained after segmentation and the subtask attribute parameters represented by the kth subtask attribute.

[0065] According to an embodiment of the present disclosure, the sub-vision task includes at least one of the following: a target detection sub-task, an image clipping sub-task, and an image description sub-task.

[0066] According to an embodiment of the present disclosure, the target detection subtask may include a subtask for detecting people, buildings, traffic lights and other targets in the processed image, for example, it may include a traffic light status detection subtask, a building name detection subtask, etc.

[0067] It should be noted that the subtask execution result of the target detection subtask may include the recognition result of the attributes such as the type of the target to be detected, and may also include the description information related to the target to be detected. For example, in the target detection subtask of detecting the building in the image block, the subtask execution result of the target detection subtask may include the name of the building and the construction process introduction information related to the building and other related information.

[0068] According to an embodiment of the present disclosure, the image editing subtask may include a subtask for editing a plurality of continuous video frame images, or may also include a subtask for performing editing processes such as cropping and contrast adjustment on any frame of an image.

[0069] In one embodiment of the present disclosure, the image to be processed may be a continuous multi-frame video frame image, and the image clipping subtask may be a subtask for clipping the continuous multi-frame video frame image. The subtask execution result of the image clipping subtask may include the clipped multi-frame video frame image, or may also include description information of the clipping operation process.

[0070] According to an embodiment of the present disclosure, the image description subtask may include a subtask of describing text in the image to be processed, or may also include a subtask of describing the layout of objects contained in the image to be processed.

[0071] In one embodiment of the present disclosure, the image to be processed may include a scanned image of a bill, and the image description subtask may include a subtask of describing the bill type of the bill.

[0072] According to an embodiment of the present disclosure, the service resources suitable for executing the sub-vision task may be software service resources such as plug-ins with vision sub-task service capabilities, or may also include hardware service resources such as chips with image processing capabilities.

[0073] According to an embodiment of the present disclosure, obtaining an image processing result according to a task execution result may include: fusing multiple subtask execution results based on an execution dependency relationship to obtain an image processing result.

[0074] According to the embodiments of the present disclosure, the execution dependency relationship can characterize the logical order of execution of multiple visual subtasks, and part or all of the multiple subtask execution results can be arranged based on the dependency relationship to achieve the fusion of multiple subtask execution results based on the execution dependency relationship, and the obtained arranged multiple subtask execution results can be used as image processing results.

[0075] According to an embodiment of the present disclosure, multiple subtask execution results can be fused based on execution dependency relationships, and multiple subtask execution results and the execution dependency relationships between multiple subtask execution results can also be processed based on a neural network model, thereby achieving full fusion of multiple subtask execution results according to the execution dependency relationships. For example, the neural network model can fuse the image blocks, image identifiers, and building names contained in the multiple subtask execution results according to the execution dependency relationships, thereby generating an image identifier and a building name corresponding to the image identifier in the image to be processed, thereby obtaining a fused image processing result, so that the target object can quickly understand the image processing result corresponding to the image to be processed based on the feedback information generated by the fused image processing result.

[0076] Figure 3 The schematic diagram shows a principle diagram of an information interaction method according to an embodiment of the present disclosure.

[0077] like Figure 3As shown, input information 301 can be pushed to the service module 300, and the input information 301 can include a demand description text: "the clothing brand of the first person on the right", and an image to be processed 3011. The service module 300 may include a large language model 310 and multiple service resources. The demand description text is input into the large language model 310, and a visual task attribute can be output. The visual task attribute may include multiple subtask attributes. The multiple subtask attributes are target detection subtask attribute 311, image block determination subtask attribute 312, image segmentation subtask attribute 313, and clothing recognition subtask attribute 314. The execution dependency relationship between the multiple subtask attributes can be characterized based on the arrangement order of the target detection subtask attribute 311, the image block determination subtask attribute 312, the image segmentation subtask attribute 313, and the clothing recognition subtask attribute 314.

[0078] like Figure 3 As shown, for the first visual sub-task, a target detection sub-task R311 (first visual sub-task) can be generated based on the target detection sub-task attribute 311 and the image to be processed 3011. Based on the target detection sub-task attribute 311, a service resource (target detection resource 321) associated with the target detection sub-task attribute 311 can be called to execute the target detection sub-task R311 to obtain the first sub-task execution result. The first sub-task execution result can be, for example, performing target detection on the image to be processed to obtain a detection box corresponding to each task.

[0079] like Figure 3 As shown, for the second sub-visual task, the image block determination sub-task R312 (second sub-visual task) can be generated based on the image block determination sub-task attribute 312 and the first sub-task execution result. Based on the image block determination sub-task attribute 312, the service resource (image block determination resource 322) associated with the image block determination sub-task attribute 312 can be called to execute the image block determination sub-task R312 to obtain the second sub-task execution result. The second sub-task execution result can be, for example, determining that the detection box corresponding to the rightmost person in the first sub-task execution result is the detection box that needs to be segmented.

[0080] like Figure 3 As shown, for the third visual subtask, an image segmentation subtask R313 (third visual subtask) can be generated based on the image segmentation subtask attribute 313 and the second subtask execution result. Based on the image segmentation subtask attribute 313, the service resource (image segmentation resource 323) associated with the image segmentation subtask attribute 313 can be called to execute the image segmentation subtask R313 to obtain the third subtask execution result. The third subtask execution result can be, for example, image segmentation according to the detection frame determined as requiring image segmentation in the second subtask execution result, thereby obtaining the rightmost character image block in the image to be processed 3011.

[0081] like Figure 3 As shown, for the fourth visual sub-task, a clothing recognition sub-task R314 (the fourth visual sub-task) can be generated based on the clothing recognition sub-task attribute 314 and the execution result of the third sub-task. Based on the clothing recognition sub-task attribute 314, a service resource (clothing recognition resource 324) associated with the clothing recognition sub-task attribute 314 can be called to execute the clothing recognition sub-task R314 to obtain the fourth sub-task execution result. The fourth sub-task execution result can be, for example, clothing detection of the rightmost character image block obtained after segmentation according to the execution result of the third sub-task to obtain information related to the clothing brand.

[0082] It should be noted that the information processing operations in the embodiments of the present disclosure, including but not limited to information acquisition operations and image processing operations, are all performed after obtaining the authorization of the relevant users. And after obtaining the information, necessary encryption measures are adopted to protect the security of the information and avoid information leakage. The data obtained according to the method provided in the embodiments of the present disclosure, including but not limited to image processing results and feedback information, are all reviewed in accordance with relevant laws and regulations or specifications, and are displayed after passing the review.

[0083] According to an embodiment of the present disclosure, generating feedback information based on an image processing result may include: processing the image processing result using a large language model to obtain a description text of the processing result; and generating feedback information based on the result description text.

[0084] According to the embodiments of the present disclosure, by using a large language model to process image processing results, a result description text for describing the image processing results can be generated based on the analysis ability and text prediction ability of the large language model, so that the feedback information generated according to the result description text can characterize the image processing results through natural language, thereby reducing the difficulty of understanding the image processing results and improving the browsing efficiency of the target object.

[0085] According to an embodiment of the present disclosure, using a large language model to process an image processing result to obtain a description text of the processing result may include: updating a preset feedback prompt template based on the image processing result to obtain feedback prompt information; and using the large language model to process the feedback prompt information to obtain a description text of the processing result.

[0086] According to an embodiment of the present disclosure, a feedback prompt template may include a feedback prompt tag sequence, which can be used to control a large language model to accurately predict the natural language used to describe the image processing results, thereby generating a processing result description text that describes the image processing results in natural language.

[0087] In one embodiment of the present disclosure, the feedback prompt information may include the following paragraphs surrounded by “ / / ”:

[0088] / / I try to describe the execution results of the following multiple sub-visual tasks.

[0089] The current sub-vision task execution results are:

[0090] "{picture1}", "{picture2}";

[0091] Please describe the above execution sub-results separately, and merge the description contents to obtain the description text of the image processing result. / /

[0092] It should be noted that "{picture1}" and "{picture2}" in the feedback prompt information can be the subtask execution results, and the feedback prompt template can be a feedback prompt mark sequence composed of other fields and characters in the feedback prompt information except "{picture1}" and "{picture2}".

[0093] According to an embodiment of the present disclosure, the requirement description text is obtained according to the input operation of the target object on the interactive interface.

[0094] According to an embodiment of the present disclosure, the information interaction method may further include: generating a requirement description floating window representing the requirement description text on the interactive interface.

[0095] According to an embodiment of the present disclosure, the target object can generate a demand description text by inputting text, voice and other information into the demand description floating window, thereby facilitating the target object to perform image processing on the image to be processed through chat interaction, avoiding performing visual tasks on the image to be processed by selecting a visual processing module, and saving operation steps.

[0096] According to an embodiment of the present disclosure, displaying feedback information in an interactive interface may include: generating a feedback information floating window suitable for displaying feedback information at a second position in the interactive interface that is within a preset distance range from a first position of a requirement description floating window.

[0097] Figure 4 The application scenario diagram of the information interaction method according to the embodiment of the present disclosure is schematically shown.

[0098] like Figure 4As shown, the interactive interface 400 may include a requirement description floating window 410 representing the requirement description text, and the requirement description floating window 410 may include the requirement description text "the clothing brand of the first person on the right", and an image to be processed 411 associated with the requirement description text. The information interaction method provided by the embodiment of the present disclosure can be executed according to the requirement description text, and feedback information corresponding to the requirement description text can be obtained. Feedback information can be displayed in the feedback information floating window 420, and the feedback information may include result description text: "The first person on the right in the figure below is the person in the dotted box in the figure below, and the clothing brand is XXX", and may also include image processing results 421. The image processing result 421 can mark the objects to be detected that need to be identified in the requirement description text based on the dotted box to facilitate browsing and viewing of the target object.

[0099] like Figure 4 As shown, the interactive interface 400 may further include an input box 430, and the target object may input a requirement description text by inputting text in the input box 430, or may upload the image to be processed by dragging the image to be processed to the input box 430. It should be understood that the target object may request a visual task service based on natural language by inputting the requirement description text and the image to be processed in the input box 430.

[0100] The information interaction method provided by the embodiment of the present disclosure can integrate various computer vision functions and generate instructions for controlling computer vision service resources based on dialogue interaction, so as to realize fine-grained classification of visual tasks of images by fine-grained control of service resources, and effectively expand the capability range of large language models.

[0101] According to an embodiment of the present disclosure, determining an image processing result related to an image to be processed according to visual task attributes may also include: generating a service call request according to the visual task attributes; and sending the service call request to a cloud service end, wherein the cloud service end is configured to call cloud service resources corresponding to the visual task attributes according to the service call request; and processing the image to be processed according to the called cloud service resources to obtain an image processing result.

[0102] According to an embodiment of the present disclosure, a service call request may include task attribute parameters of a visual task and a service resource identifier of a service resource required to perform the visual task. By sending a service call request to the cloud, the cloud service end can generate a visual task according to the service call request, and call the cloud service resource corresponding to the visual task to perform the visual task, thereby performing image processing on the image to be processed and obtaining an image processing result. Then, by obtaining the image processing result from the cloud service end, cloud deployment of the visual task can be realized, reducing the computational overhead of local service resources used to perform the visual task, and improving the execution efficiency of the visual task.

[0103] According to an embodiment of the present disclosure, a visual task attribute may include multiple subtask attributes and execution dependencies between the multiple subtask attributes, and a service call request may also correspondingly include multiple subtask attributes and execution dependencies between the subtask attributes. After receiving the service call request, the cloud service end may generate multiple sub-visual tasks based on the multiple subtask attributes, and call cloud service resources corresponding to the multiple subtask attributes. In addition, the cloud service end may execute the sub-visual tasks according to the execution dependencies, so that the cloud service end may generate image processing results according to the method provided in the embodiment of the present disclosure.

[0104] According to an embodiment of the present disclosure, a service call request may also include some sub-task attributes in the visual task attributes, so that the visual task can be jointly executed by calling the cloud service resources of the cloud service end and calling the service resources of the local service end, thereby improving the scope of application of the information interaction method.

[0105] According to an embodiment of the present disclosure, the information interaction method may further include: sending the image to be processed to the cloud service end; and receiving the image processing result sent from the cloud service end.

[0106] According to an embodiment of the present disclosure, sending the image to be processed to the cloud server may include generating an image message based on the image to be processed, and asynchronously sending the image message to the cloud server through a message queue to achieve sending the image to be processed to the cloud server.

[0107] According to an embodiment of the present disclosure, receiving the image processing result sent from the cloud server may include asynchronously acquiring the image processing result from the cloud server through a message queue.

[0108] Figure 5 An application scenario diagram of an information interaction method according to another embodiment of the present disclosure is schematically shown.

[0109] like Figure 5 As shown, the application scenario 500 may include a client 510, a local server 520, and a cloud server 530. The target object may input a requirement description text 501 and an image to be processed 502 through the client 510. The client 510 may send the requirement description text 501 to the server 520, and send the image to be processed 502 to the cloud server 530.

[0110] The local server 520 may include a large language model service module and a plug-in service module. The large language model service module may be constructed based on a pre-trained large language model. When the local server 520 obtains the demand description text 501, the large language model service module will process the demand description text 501 based on the large language model and generate visual task attributes. The visual task attributes may include multiple subtask attributes and execution dependencies between multiple subtask attributes, and the subtask attributes may include the cloud service resource identifier of the called cloud service resource. Multiple subtask attributes and the execution dependencies between multiple subtask attributes may be stored in the visual task list to facilitate structured storage of subtask attributes and execution dependencies. The plug-in service module will generate a service call request 521 based on the execution dependencies between multiple subtask attributes and multiple subtask attributes generated by the large language model service module. The local server 520 sends the service call request 521 to the cloud service 530. The cloud service end 530 receives the service call request 521, and can generate multiple sub-visual tasks based on the multiple sub-task attributes in the service call request 521 and the execution dependency relationship between the multiple sub-task attributes, and call the cloud service resources corresponding to the sub-task attributes according to the execution dependency relationship to sequentially execute the multiple sub-visual tasks to obtain multiple sub-task execution results. The cloud service end 530 can generate an image processing result message 531 based on the multiple sub-task execution results. The local service end 520 can push the multiple sub-task execution results in the image processing result message 531 to the plug-in service module, and the plug-in service module merges the multiple sub-task execution results according to the execution dependency relationship, and pushes the merged image processing results to the large language model service module. The large language model service module can process the image processing results based on the large language model and output the result description text. The local service end 520 can generate feedback information 522 based on the image processing results and the result description text, and send the feedback information 522 to the client 510 so as to display the feedback information 522 on the interactive interface of the client 510. According to the information interaction method provided by the embodiment of the present disclosure, the target object can request to perform a visual task through a natural language interaction method similar to chatting. There is no need to select a visual service from a large list of visual functions, nor is there any need to access several complex visual service interfaces, thereby improving the overall efficiency of visual task execution.

[0111] In another embodiment of the present disclosure, the local server can generate service call requests corresponding to the subtask attributes one by one for the multiple subtask attributes, and send the multiple service call requests to the cloud server. The cloud server can also send the multiple subtask execution results to the local server respectively.

[0112] It should be noted that Figure 5The local server shown in the figure can be a server or a server cluster, and the large language model service module and the plug-in service module can be deployed in any server or server cluster of the local server.

[0113] In another embodiment of the present disclosure, the cloud server may include an object server corresponding to the permissions of the target object, and the object server may deploy private service resources related to the target object, thereby facilitating the target object to quickly perform visual tasks by calling the private service resources, thereby realizing a personalized processing process for the image to be processed.

[0114] According to an embodiment of the present disclosure, the information interaction method may further include: performing authority authentication on a target object related to the input requirement description text to obtain an authority authentication result; and determining an object server corresponding to the authority of the target object based on the authority authentication result, wherein the cloud service includes the object server.

[0115] According to the embodiments of the present disclosure, the target object can be authenticated according to the permission token related to the target object, so that the permission authentication result can be obtained. By determining the object server corresponding to the permission authentication result, a service call request can be sent to the object server related to the permission of the target object, and the image processing result can be obtained from the object server. In this way, the object server can be adapted according to the permission attributes of the target object, so that the target object can perform visual tasks by calling the cloud service resources in the object server, avoiding the distribution of the image to be processed to other cloud servers, resulting in information leakage of the image to be processed, and ensuring the information security of the target object.

[0116] Figure 6 A block diagram of an information interaction device according to an embodiment of the present disclosure is schematically shown.

[0117] like Figure 6 As shown, the information interaction device 600 includes: a visual task attribute acquisition module 610, an image processing result determination module 620, a feedback information generation module 630 and a display module 640.

[0118] The visual task attribute acquisition module 610 is used to process the requirement description text using a large language model in response to obtaining the requirement description text, and obtain visual task attributes that match the image processing intent represented by the requirement description text, wherein the requirement description text is associated with the image to be processed.

[0119] The image processing result determination module 620 is used to determine the image processing result related to the image to be processed according to the visual task attributes.

[0120] The feedback information generating module 630 is used to generate feedback information according to the image processing result.

[0121] The display module 640 is used to display the feedback information on the interactive interface.

[0122] According to an embodiment of the present disclosure, the image processing result determination module includes: a visual task generation submodule, a task execution result determination submodule and an image processing result acquisition submodule.

[0123] The visual task generation submodule is used to generate visual tasks according to visual task attributes and images to be processed.

[0124] The task execution result determination submodule is used to execute the visual task according to at least one service resource associated with the visual task attribute to obtain the task execution result.

[0125] The image processing result acquisition submodule is used to obtain the image processing result according to the task execution result.

[0126] According to an embodiment of the present disclosure, a visual task attribute includes multiple subtask attributes and execution dependencies between the multiple subtask attributes, a visual task includes sub-visual tasks corresponding to the subtask attributes, and service resources are associated with the subtask attributes.

[0127] According to an embodiment of the present disclosure, the task execution result determination submodule includes a sub-visual task execution unit.

[0128] A sub-visual task execution unit is used to call the kth service resource associated with the kth sub-task attribute to execute the kth sub-visual task according to the execution dependency, and obtain the kth sub-task execution result, wherein k>1 and k is an integer, the kth sub-visual task is determined according to the kth sub-task attribute, and the task execution result includes the sub-task execution result.

[0129] According to an embodiment of the present disclosure, the image processing result obtaining submodule includes an image processing result obtaining unit.

[0130] The image processing result obtaining unit is used to fuse multiple subtask execution results based on the execution dependency relationship to obtain the image processing result.

[0131] According to an embodiment of the present disclosure, the sub-vision task includes at least one of the following: a target detection sub-task, an image clipping sub-task, and an image description sub-task.

[0132] According to an embodiment of the present disclosure, the feedback information generating module includes: a processing result description text obtaining submodule and a feedback information generating submodule.

[0133] The processing result description text acquisition submodule is used to process the image processing results using the large language model to obtain the processing result description text.

[0134] The feedback information generation submodule is used to generate feedback information according to the result description text.

[0135] According to an embodiment of the present disclosure, the processing result description text obtaining submodule includes: a feedback prompt information obtaining unit and a processing result description text obtaining unit.

[0136] The feedback prompt information obtaining unit is used to update a preset feedback prompt template based on the image processing result to obtain feedback prompt information.

[0137] The processing result description text obtaining unit is used to use the large language model to process the feedback prompt information to obtain the processing result description text.

[0138] According to an embodiment of the present disclosure, the requirement description text is obtained according to the input operation of the target object on the interactive interface.

[0139] The information interaction device also includes: a requirement description floating window generation module.

[0140] The requirement description floating window generation module is used to generate a requirement description floating window representing the requirement description text in the interactive interface.

[0141] According to an embodiment of the present disclosure, the display module includes a feedback information floating window generation submodule.

[0142] The feedback information floating window generation submodule is used to generate a feedback information floating window suitable for displaying feedback information at a second position within a preset distance range from the first position of the requirement description floating window in the interactive interface.

[0143] According to an embodiment of the present disclosure, the information interaction device also includes a requirement description script determination module.

[0144] The requirement description text determination module is used to detect the image processing intention of the image to be processed in response to receiving the image to be processed input by the target object, and obtain the requirement description text.

[0145] According to an embodiment of the present disclosure, the information interaction device also includes an updating module.

[0146] The updating module is used to update the received requirement description text according to a preset requirement prompt template to obtain a new requirement description text.

[0147] According to an embodiment of the present disclosure, the image processing result determination module includes: a request generation module and a first sending module.

[0148] The request generation module is used to generate service call requests based on visual task attributes.

[0149] The first sending module is used to send a service call request to the cloud server, wherein the cloud server is configured to call the cloud service resources corresponding to the visual task attributes according to the service call request; and process the image to be processed according to the called cloud service resources to obtain the image processing result.

[0150] According to an embodiment of the present disclosure, the information interaction device further includes: a second sending module and a receiving module.

[0151] The second sending module is used to send the image to be processed to the cloud server.

[0152] The receiving module is used to receive the image processing results sent from the cloud server.

[0153] According to an embodiment of the present disclosure, the information interaction device further includes: an authority authentication result obtaining module and an object server determining module.

[0154] The permission authentication result acquisition module is used to perform permission authentication on the target object related to the input requirement description text and obtain the permission authentication result.

[0155] The object server determination module is used to determine the object server corresponding to the authority of the target object according to the authority authentication result, wherein the cloud server includes the object server.

[0156] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0157] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0158] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0159] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method as described above.

[0160] Figure 7A block diagram of an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0161] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0162] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0163] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the information interaction method. For example, in some embodiments, the information interaction method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the information interaction method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the information interaction method in any other appropriate manner (e.g., by means of firmware).

[0164] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0165] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0166] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0167] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0168] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0169] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.

[0170] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0171] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. An information interaction method, include: In response to acquiring the requirement description text, updating the received requirement description text according to a preset requirement prompt template to obtain a processed requirement description text; Processing the processed requirement description text using a large language model to obtain visual task attributes that match the image processing intent represented by the requirement description text and include multiple subtask attributes, wherein the requirement description text is associated with the image to be processed; Generating a plurality of sub-vision tasks corresponding to the plurality of sub-task attributes according to the image to be processed and the vision task attributes, wherein the vision task attributes further include an execution dependency relationship between the plurality of sub-task attributes, and the execution dependency relationship represents an execution logic order of the plurality of sub-vision tasks; According to the execution dependency, calling the service resource associated with the subtask attribute to execute the sub-visual task, and obtaining the sub-task execution result; Obtaining an image processing result according to the execution results of the plurality of subtasks; generating feedback information according to the image processing result; The feedback information is displayed on the interactive interface.

2. The method according to claim 1, in, Obtaining the image processing result according to the execution results of the plurality of subtasks comprises: The image processing result is obtained by fusing the execution results of the multiple subtasks based on the execution dependency relationship.

3. The method according to claim 1, in, The sub-visual tasks include at least one of the following: Object detection subtask, image clipping subtask, image description subtask.

4. The method according to claim 1, in, The generating feedback information according to the image processing result comprises: Processing the image processing result using the large language model to obtain a processing result description text; and The feedback information is generated according to the result description text.

5. The method according to claim 4, in, The using of the large language model to process the image processing result to obtain a processing result description text comprises: Updating a preset feedback prompt template based on the image processing result to obtain feedback prompt information; and The feedback prompt information is processed using the large language model to obtain the processing result description text.

6. The method according to claim 4, in, The requirement description text is obtained according to the input operation of the target object on the interactive interface; The method further comprises: Generating a requirement description floating window representing the requirement description text on the interactive interface; The displaying of the feedback information on the interactive interface includes: A feedback information floating window suitable for displaying the feedback information is generated at a second position in the interactive interface that is within a preset distance range from the first position of the requirement description floating window.

7. The method according to claim 1, further comprising: include: In response to receiving the image to be processed input by the target object, image processing intention detection is performed on the image to be processed to obtain the requirement description text.

8. The method according to claim 1, in, Determining the image processing result related to the image to be processed according to the visual task attribute includes: Generate a service call request according to the visual task attributes; and The service call request is sent to a cloud service end, wherein the cloud service end is configured to call a cloud service resource corresponding to the visual task attribute according to the service call request; and process the image to be processed according to the called cloud service resource to obtain the image processing result.

9. The method according to claim 8, further comprising: include: Sending the image to be processed to the cloud service end; as well as Receive the image processing result sent from the cloud service end.

10. The method according to claim 8, further comprising: include: Performing authority authentication on the target object related to the input requirement description text to obtain an authority authentication result; as well as According to the permission authentication result, an object server corresponding to the permission of the target object is determined, wherein the cloud server includes the object server.

11. An information interaction device, include: A visual task attribute acquisition module is used to update the received demand description text according to a preset demand prompt template in response to obtaining the demand description text, and obtain a processed demand description text; use a large language model to process the processed demand description text to obtain visual task attributes that match the image processing intent represented by the demand description text and include multiple subtask attributes, and the demand description text is associated with the image to be processed; An image processing result determination module, used to determine an image processing result related to the image to be processed according to the visual task attribute; A feedback information generating module, used for generating feedback information according to the image processing result; as well as A display module, used to display the feedback information on an interactive interface; Wherein, the image processing result determination module is configured as follows: Generate a plurality of visual sub-tasks according to the image to be processed and the visual task attributes, wherein the visual task attributes further include an execution dependency relationship between the plurality of sub-task attributes, and the execution dependency relationship represents an execution logic order of the plurality of visual sub-tasks; According to the execution dependency, calling the service resource associated with the subtask attribute to execute the sub-visual task, and obtaining the sub-task execution result; The image processing result is obtained according to the execution results of the multiple subtasks.

12. The device according to claim 11, in, The image processing result determination module comprises: The image processing result obtaining unit is used to fuse the execution results of the plurality of subtasks based on the execution dependency relationship to obtain the image processing result.

13. The device according to claim 11, in, The sub-visual tasks include at least one of the following: Object detection subtask, image clipping subtask, image description subtask.

14. The device according to claim 11, in, The feedback information generating module comprises: a processing result description text obtaining submodule, used to process the image processing result using the large language model to obtain a processing result description text; and The feedback information generating submodule is used to generate the feedback information according to the result description text.

15. The device according to claim 14, in, The processing result description text obtaining submodule includes: A feedback prompt information obtaining unit, used to update a preset feedback prompt template based on the image processing result to obtain feedback prompt information; and The processing result description text obtaining unit is used to use the large language model to process the feedback prompt information to obtain the processing result description text.

16. The device according to claim 14, in, The requirement description text is obtained according to the input operation of the target object on the interactive interface; The device also includes: A requirement description floating window generating module, used to generate a requirement description floating window representing the requirement description text on the interactive interface; Wherein, the display module includes: The feedback information floating window generation submodule is used to generate a feedback information floating window suitable for displaying the feedback information at a second position within a preset distance range from the first position of the requirement description floating window in the interactive interface.

17. The device according to claim 11, further comprising: include: The requirement description text determination module is used to perform image processing intention detection on the image to be processed in response to receiving the image to be processed input by the target object, so as to obtain the requirement description text.

18. The device according to claim 11, in, The image processing result determination module comprises: A request generation module, used to generate a service call request according to the visual task attributes; and The first sending module is used to send the service call request to the cloud service end, wherein the cloud service end is configured to call the cloud service resources corresponding to the visual task attributes according to the service call request; and process the image to be processed according to the called cloud service resources to obtain the image processing result.

19. The device according to claim 18, further comprising: include: A second sending module, used for sending the image to be processed to the cloud service end; as well as A receiving module is used to receive the image processing result sent from the cloud service end.

20. The device according to claim 18, further comprising: include: The authorization authentication result obtaining module is used to perform authorization authentication on the target object related to the input requirement description text to obtain the authorization authentication result; as well as The object server determination module is used to determine the object server corresponding to the authority of the target object according to the authority authentication result, wherein the cloud server includes the object server.

21. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

23. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Searching method and device, electronic equipment and storage medium

    CN116628327A

  • Multi-model cooperation method based on large-scale language model

    CN116976306A

  • Multi-modal machine learning architectures integrating language models and computer vision systems

    US11803710B1