Image processing method and device based on intelligent agent

By using an agent-based image processing method, image elements are extracted into image blocks and layered for editing, which solves the problems of convenience and intuitiveness in image processing, and enables personalized adjustment of image elements and ensures image integrity.

CN121414906APending Publication Date: 2026-01-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511509075.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently achieve convenient and intuitive image processing, especially in adjusting and personalizing specific elements within an image, failing to meet user needs.

Method used

By using an agent-based image processing method, image elements are extracted into image blocks, layers are created and edited, and image filling and position adaptation are combined. Through the collaborative work of the agent image component and the server, the synchronous display and editing of image elements are achieved.

Benefits of technology

It improves the convenience and intuitiveness of image processing, meets users' personalized processing needs for specific elements in images, and ensures the integrity and consistency of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414906A_ABST
    Figure CN121414906A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an agent-based image processing method and device.The agent-based image processing method comprises the steps that in the process that a user carries out image processing based on an agent, according to image elements corresponding to a user instruction input by the user in an agent image assembly, the image elements corresponding to the user instruction input by the user in the agent image assembly; the method comprises the following steps: extracting an element image block of an image element, creating a layer of the element image block, editing the element image block in the layer according to an editing instruction, then carrying out image filling processing on an element extraction area in an image of the extracted image element, and carrying out position adaptation processing on the element image block according to an editing position of the element image block. Therefore, image processing of the image elements is realized based on the intelligent agent image component.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of image processing technology, and in particular to an image processing method and apparatus based on intelligent agents. Background Technology

[0002] With the advancement of digital image processing technology, image editing tools have been widely deployed and used in various terminal devices and application platforms. As users' demands for personalized image processing increase, image processing targeting image elements within images has received more and more attention. For example, adjusting specific elements in an image has become a common requirement in the image processing process. Furthermore, with the continuous development of artificial intelligence technology and the increasing demands of users for the intuitiveness, convenience, and responsiveness of image processing tools, how to better achieve image processing has become a key focus for all parties. Summary of the Invention

[0003] This specification provides one or more embodiments of an intelligent agent-based image processing method, comprising: extracting element image blocks of image elements corresponding to user commands input by a user in an intelligent agent image component; creating a layer of the element image blocks and performing editing processing on the element image blocks in the layer according to editing commands; performing image filling processing on the element extraction region in the image from which the image elements are extracted, and performing position adaptation processing on the element image blocks according to the editing position of the element image blocks.

[0004] This specification provides one or more embodiments of another agent-based image processing method, including: acquiring a user instruction input by a user in an agent image component and submitting it to a server; acquiring editing processing data of the element image block corresponding to the user instruction in a created layer from the server and synchronously displaying the edited data; acquiring the user-inputted editing instruction and submitting it to the server; receiving the image filling processing result returned by the server for the element extraction region in the image from which the image element is extracted and synchronously displaying the result; and receiving the adaptation processing result returned by the server for the position adaptation processing of the element image block and synchronously displaying the result.

[0005] This specification provides one or more embodiments of an agent-based image processing apparatus, comprising: an image block extraction module configured to extract element image blocks of an image element corresponding to a user instruction input by a user into an agent image component; an image block editing module configured to create a layer of the element image blocks and perform editing processing on the element image blocks in the layer according to editing instructions; and a position adaptation module configured to perform image filling processing on the element extraction region in the image from which the image elements are extracted, and perform position adaptation processing on the element image blocks according to the edited position of the element image blocks.

[0006] This specification provides one or more embodiments of another agent-based image processing apparatus, comprising: an instruction acquisition module configured to acquire user instructions input by a user in an agent image component and submit them to a server; an editing display module configured to acquire editing processing data of the element image blocks corresponding to the user instructions in a created layer from the server and synchronously display the edited data; an instruction submission module configured to acquire the editing instructions input by the user and submit them to the server; and a result display module configured to receive and synchronously display the image filling processing results returned by the server for the element extraction regions in the image from which the image elements are extracted, and to receive and synchronously display the adaptation processing results returned by the server for the position adaptation processing of the element image blocks.

[0007] This specification provides one or more embodiments of an agent-based image processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: extract element image blocks of an image element corresponding to a user instruction input by a user in an agent image component; create a layer of the element image blocks and perform editing processing on the element image blocks in the layer according to editing instructions; perform image filling processing on the element extraction region in the image from which the image element was extracted, and perform position adaptation processing on the element image blocks according to the editing position of the element image blocks.

[0008] This specification provides one or more embodiments of another agent-based image processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: acquire user instructions input by a user in an agent image component and submit them to a server; acquire editing processing data of element image blocks corresponding to the user instructions in a created layer from the server and synchronously display the edited data; acquire the user-inputted editing instructions and submit them to the server; receive and synchronously display image filling processing results returned by the server for element extraction regions in the image from which the image elements are extracted, and receive and synchronously display adaptation processing results returned by the server for position adaptation processing of the element image blocks.

[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions that, when executed, perform the following process: extracting element image blocks from image elements corresponding to user instructions input by a user in an intelligent agent image component; creating a layer of the element image blocks and performing editing processing on the element image blocks in the layer according to editing instructions; performing image filling processing on the element extraction region in the image from which the image elements are extracted, and performing position adaptation processing on the element image blocks according to the edited position of the element image blocks.

[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: acquiring user instructions input by a user in an intelligent agent image component and submitting them to a server; acquiring editing processing data of the element image block corresponding to the user instructions in a created layer from the server and synchronously displaying the edited data; acquiring the user-inputted editing instructions and submitting them to the server; receiving and synchronously displaying the image filling result returned by the server for the element extraction region in the image from which the image elements are extracted, and receiving and synchronously displaying the adaptation processing result returned by the server for the position adaptation processing of the element image block. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1A schematic diagram illustrating an implementation environment for one or more embodiments of an agent-based image processing method provided in this specification; Figure 2 A flowchart of an image processing method based on an intelligent agent, provided for one or more embodiments of this specification; Figure 3 A schematic diagram of a first image processing page provided for one or more embodiments of this specification; Figure 4 A schematic diagram of a second image processing page provided for one or more embodiments of this specification; Figure 5 A schematic diagram of a third image processing page provided for one or more embodiments of this specification; Figure 6 A schematic diagram of a fourth image processing page provided for one or more embodiments of this specification; Figure 7 A timing diagram illustrating the processing of an agent-based image processing method applied to an image processing scenario, provided in one or more embodiments of this specification. Figure 8 A flowchart of an image processing method based on an intelligent agent, provided for one or more embodiments of this specification; Figure 9 A schematic diagram illustrating an embodiment of an agent-based image processing apparatus provided in one or more embodiments of this specification; Figure 10 A schematic diagram of another embodiment of an agent-based image processing apparatus provided in one or more embodiments of this specification; Figure 11 A schematic diagram of the structure of an agent-based image processing device provided in one or more embodiments of this specification; Figure 12 This is a schematic diagram of another agent-based image processing device provided in one or more embodiments of this specification. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0013] The agent-based image processing method provided in one or more embodiments of this specification is applicable to the image processing implementation environment. (Refer to...) Figure 1 The implementation environment includes at least: Server 101, client 102, and intelligent agent image component 103 deployed on client 102; The server 101 is used to obtain user commands submitted by the client 102 and inputted by the user in the intelligent agent image component 103, extract element image blocks of image elements based on the user commands, edit the element image blocks, fill the extracted element regions and perform position adaptation processing on the element image blocks, and return the filling and adaptation processing results to the client 102. The server 101 can be deployed on a server, which can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The client 102 is used to submit user instructions and / or editing instructions input by the user in the intelligent agent image component 103 to the server 101, and to receive the filling and adaptation results obtained after image processing of the image elements returned by the server 101 and display them synchronously. The client 102 can be deployed on a terminal device, which can be a mobile phone, personal computer, tablet computer, e-book reader, wearable device, device for information interaction based on AR (Augmented Reality) / VR (Virtual Reality), and laptop computer, etc.

[0014] In this implementation environment, client 102 first obtains the user command input by the user in the intelligent agent image component 103 and submits it to server 101. Server 101 extracts the element image block of the image element according to the image element corresponding to the user command and creates a layer of the element image block. Based on this, client 102 obtains the editing processing data of the element image block in the created layer from server 101 and performs editing synchronous display. It also obtains the editing command input by the user and submits it to server 101. Server 101 performs editing processing of the element image block in the layer according to the editing command, performs image filling processing on the element extraction area in the image of the extracted image element, and performs position adaptation processing on the element image block according to the editing position of the element image block. It returns the filling processing result and the adaptation processing result to client 102. Client 102 receives the filling processing result and the adaptation processing result returned by server 101 and displays them synchronously. In this way, image processing of image elements is realized based on the intelligent agent image component.

[0015] This specification provides one or more embodiments of an agent-based image processing method as follows: Reference Figure 2The image processing method based on intelligent agents provided in this embodiment specifically includes steps S202 to S206.

[0016] Step S202: Extract the element image block of the image element according to the image element corresponding to the user instruction input by the user in the intelligent agent image component.

[0017] The intelligent agent image component described in this embodiment refers to a component or plugin used for image processing. Specifically, it can be an image processing component integrated with an intelligent agent, or an interactive plugin integrated with an intelligent agent, such as an interactive plugin that interacts with the user to process images. Alternatively, the intelligent agent image component can also be an image editing tool integrated with an intelligent agent, including tools for cropping, repairing, removing objects, and / or filling objects in images. The intelligent agent image component integrates an intelligent agent, which includes an entity capable of autonomously performing tasks, making decisions, and learning and adjusting according to environmental changes. Optionally, the intelligent agent image component includes an image editing tool integrated with an intelligent agent.

[0018] The user instruction refers to the instruction input by the user through the intelligent agent's image component or the intelligent agent's dialogue interface. Specifically, it can be an instruction used by the user to express the intention to edit or manipulate the image content. The user instruction can be any form of instruction. Specifically, the user instruction can be a text instruction, such as natural text language entered by the user in the intelligent agent's dialog box or input box to express the image processing intention. The user instruction can also be a semantic instruction, such as an instruction obtained by semantic recognition of the image processing instruction input by the user's voice. The user instruction can also be an action instruction, such as a specific operation behavior performed by the user through touch interaction. Optionally, user instructions may include text or semantic instructions entered into the agent's image component or the agent's dialogue interface, or action instructions entered into the agent's image component.

[0019] The image element refers to an identifiable visual object or image unit in the image, specifically the target object that the user intends to manipulate through user commands; the element image block refers to a sub-image data block extracted from the image that corresponds to a certain image element; the element image block is the direct object of operation processed during image processing; optionally, the element image block includes element images generated by the intelligent agent.

[0020] In practice, when a user is processing an image through an intelligent agent image component, the user inputs a user command expressing the intention to process the image. The user's client submits the user command to the server. Accordingly, the server determines the image elements in the image to be processed based on the user command submitted by the client, that is, it determines the image elements corresponding to the user command and extracts the element image blocks of the image elements.

[0021] Here, the server can be a server applied to the intelligent agent, and correspondingly, the client can be the client of the intelligent agent, or the user's access client or user terminal; in addition, the server can also be the server of the intelligent agent image component. In this case, the execution entity that interacts with the server can be the intelligent agent image component. Based on this, the client in this embodiment can also be replaced by the intelligent agent image component.

[0022] In the specific execution process, during image processing, the user-submitted instructions can be action instructions or trigger instructions. Therefore, in determining image elements, the image elements can be determined based on the user-submitted action instructions or trigger instructions. In one optional implementation of this embodiment, the image elements corresponding to the user instructions are determined in the following way: The image corresponding to the user command is segmented to obtain each image element, and the image element corresponding to the image position carried by the user command is determined as the target element.

[0023] Specifically, after obtaining the user's input command into the intelligent agent image component and determining the image corresponding to the user's command, the image is first segmented to obtain each image element. For example, an image segmentation model can be called to segment the image to identify the boundaries of different objects in the image, thereby decomposing the image into multiple image elements. On this basis, the target element is further determined in each image element based on the image position carried by the user command, thereby obtaining the target element.

[0024] Furthermore, when the user-submitted command is a text command or a voice command, the image element can also be determined based on the text command or voice command; in another optional implementation provided in this embodiment, the image element corresponding to the user command is determined in the following way: The image corresponding to the user command is subjected to image semantic recognition to obtain each image semantic element, and the element keywords carried by the user command are matched with each image semantic element to obtain the target element.

[0025] For example, if a user inputs the command: "Delete the blue vehicle on the right side of the image from the image," the system first performs image semantic recognition on the image corresponding to the user's command to obtain the semantic elements of each image. Then, it extracts the keyword elements from the user's command to obtain keyword elements such as "right side," "blue," and "vehicle." Finally, it performs image semantic matching between the keyword elements and the semantic elements of each image to obtain the target element.

[0026] Here, in the process of determining image elements, the image elements can be image elements in images uploaded by the user, or they can be image elements generated by the intelligent agent.

[0027] Step S204: Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions.

[0028] As described above, based on the extraction of image elements into image blocks, in order to achieve image processing of image elements in the image while ensuring image integrity, a layer of image elements into image blocks is created, and image processing is performed on the layer.

[0029] In practice, based on the extracted image elements, a layer of the image elements is created. Then, the editing data of the image elements in the created layer can be returned to the client so that the client can edit and display the data synchronously. When the client receives the user's edit instructions, the image elements in the layer are edited according to the instructions.

[0030] The editing instructions refer to user-inputted instructions for manipulating element image blocks, such as displacement instructions, scaling instructions, rotation instructions, and / or deletion instructions. Editing instructions can be any form of instruction; specifically, editing instructions can be text instructions, semantic instructions, action instructions, or trigger instructions. Optionally, editing instructions include text instructions or semantic instructions input into the agent image component or the agent's dialogue interface, or action instructions input into the agent image component.

[0031] It should be noted that the editing command submitted by the user and the user command input above can be the same command or different commands. For example, if the editing command and the user command are the same command, and the user long-presses an image element in the image and drags it, the image element corresponding to the user command can be determined based on the user's long-press action, and the editing command can be determined based on the user's dragging action to be a displacement operation on the element image block of the image element. In this case, when the editing command and the user command are the same command, the editing of the element image block on the layer according to the editing command can be replaced by: the editing of the element image block on the layer according to the user command.

[0032] In the specific execution process, during the editing of element image blocks, the editing of element image blocks can be based on user action commands or trigger commands for image elements; in an optional implementation method provided in this embodiment, the editing of element image blocks on the layer according to editing commands includes: Based on the position or trajectory information carried by the displacement command, the element image block is displaced and rendered on the layer, and the displacement rendering data is synchronized to the client for synchronous display.

[0033] Specifically, when a user submits a displacement command through a trigger operation, the position information or trajectory information carried by the displacement command is determined based on the user's displacement operation on the element image. Then, the displacement processing and displacement rendering of the element image block on the layer are performed based on the position information or trajectory information. After that, the displacement rendering data is synchronized to the client for displacement synchronization display.

[0034] For example, a user submits a displacement command by dragging an image element to a target location. The position information is determined based on the coordinate information of the target location carried in the displacement command submitted by the user, or the trajectory information is determined based on the coordinate sequence continuously generated during the user's dragging process. Subsequently, the spatial coordinate parameters of the element image block in the layer are updated, and the display position of the element image block in the image is adjusted synchronously to achieve displacement processing. The movement process of the element image block is drawn in real time according to the trajectory information to achieve displacement rendering.

[0035] Furthermore, editing keywords can be extracted based on user text or voice commands for editing image elements, and the element image blocks can be edited based on these editing keywords. In another optional implementation of this embodiment, the element image blocks are edited on a layer according to editing commands, including: The displacement position of the element image block is determined based on the displacement keyword carried by the displacement command, and the displacement processing of the element image block is performed on the layer according to the displacement position. The displacement processing result is then synchronized to the client.

[0036] Specifically, the displacement command is first determined based on the user's input text or voice command. Then, the displacement position of the element image block is determined based on the displacement keywords carried in the displacement command to indicate the target location. The element image block is then displaced on the layer according to its displacement position. Finally, the displacement processing result is synchronized to the client for synchronized display. Here, the displacement command and the aforementioned user command can be the same command.

[0037] Step S206: Perform image filling processing on the element extraction region in the image from which the image elements are extracted, and perform position adaptation processing on the element image block according to the editing position of the element image block.

[0038] In practice, to avoid obvious editing marks in the processed image, or to prevent damage to the overall composition of the image due to lighting imbalance or spatial misalignment, the element extraction area is first determined in the image where the image elements have been extracted, and image filling processing is performed on the element extraction area; and, when the element image block is moved to the editing position, the position adaptation processing of the element image block is performed according to the editing position of the element image block.

[0039] Among them, the element extraction region refers to the area left in the image to be repaired after the image element is identified and segmented and extracted; the position adaptation processing refers to the process of adjusting the appearance attributes of the element image block according to the visual environment characteristics of the editing position after the element image block is moved to the new editing position, so as to make the element image block consistent with the context of the editing position; for example, the position adaptation processing of the element image block can be performed based on the lighting, shadow, perspective and / or blur degree of the editing position.

[0040] Optionally, images of extracted image elements and / or images corresponding to edited locations are generated by the agent.

[0041] In practical applications, users may not only move or modify element image blocks within the original image, but may also add element image blocks from one image to another image for compositing. In this case, in addition to editing the element image blocks, before performing image filling on the extracted element area, it is also possible to determine whether the original image from which the element image block was extracted is the same image as the image corresponding to the current editing position of the element image block. In one optional implementation of this embodiment, after editing the element image block on the layer according to the editing instructions, the method further includes: Determine whether the image of the extracted element image block is the same image as the image corresponding to the edit position; If so, perform image filling processing on the element extraction regions in the image from which the image elements are extracted; If not, perform fusion processing between the element image block and the image corresponding to the edit position based on the edit position of the element image block.

[0042] Specifically, based on the editing of the element image block on the layer according to the editing instructions, the image of the extracted element image block is judged. Specifically, it is judged whether the image of the extracted element image block is the same image as the image corresponding to the editing position. If so, it indicates that the editing operation is an internal structural adjustment of the image, and the element extraction area in the image of the extracted image element is filled. If not, it indicates that the editing operation is a cross-image fusion editing operation, and the element image block and the image corresponding to the editing position are merged according to the editing position of the element image block.

[0043] For example, if a user wants to... Figure 3 A car 301, identified from other images, is added to the image. After extracting car 301 from other images, it is determined whether the image of extracted car 301 is the same image as the image corresponding to the edit position. If the result is negative, then the image of car 301 is fused with the image corresponding to the edit position based on the edit position of car 301. For example, ... Figure 3 The size of the China Automotive 301 Figure 3 To adapt the size of other image elements in the image, or for example, to... Figure 3 The image style of the CR301 has been modified to match... Figure 3 The overall image style is compatible with the style of the image. Based on this, the image obtained after fusion processing is as follows: Figure 4 As shown.

[0044] Here, if the result of determining whether the image of the extracted element image block and the image corresponding to the editing position are the same image is negative, that is, if the image of the extracted element image block and the image corresponding to the editing position are not the same image, the image of the extracted element image block can be called the first image, and the image corresponding to the editing position can be called the second image; that is, the image of the extracted element image block can be replaced with the first image, and the image corresponding to the editing position can be replaced with the second image; in this case, step S206 can also be replaced with: performing image filling processing on the element extraction area in the first image, and performing position adaptation processing on the element image block according to the editing position of the second image.

[0045] Furthermore, if the determination result of whether the image of the extracted element image block and the image corresponding to the editing position are the same image is negative, in addition to the above-mentioned fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block, the following can also be performed: predict the displacement trajectory of the element image block based on the displacement position to obtain the predicted trajectory, and determine the positional relationship between the predicted trajectory and the corresponding image; determine the editing type of the element image block based on the positional relationship, and perform image editing adaptation of the element image block and the corresponding image according to the editing type; this embodiment does not limit this.

[0046] In specific implementation, during the image filling process of the element extraction region, in order to ensure the continuity and integrity of the image being processed, and to ensure the user's interactive experience in image processing, image filling processing can be performed on the element extraction region that is temporarily in a "vacant" state after the image elements are extracted. In one optional implementation of this embodiment, image filling processing of the element extraction region in the image from which image elements are extracted includes: If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, the image after extracting the image elements is input into the image filling model for image region filling processing to obtain the filled image, and synchronized to the client for image editing and synchronous display.

[0047] The preset displacement filling condition refers to the preset logical condition used to determine whether to trigger image filling processing. The preset displacement filling condition can be set based on the displacement distance or the user's operation state. For example, the displacement distance can be set to be greater than or equal to a displacement threshold, or the current user's operation state can be set to "release", or the element image block can be set to enter the target area. An image filling model is a model used for image region filling processing. The input of an image filling model is the image after extracting image elements, and the output is the filled image obtained after image region filling processing. Here, the filled image refers to the image in which the content of the extracted element regions in the image has been supplemented.

[0048] In the specific execution process, the real-time position of the element image block and the displacement data of the real-time position of the element image block relative to the element extraction area can be detected. If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, the image after extracting the image elements is input into the image filling model so that the image filling model can perform image area filling processing and output the filled image. After that, the filled image output by the image filling model is obtained and synchronized with the client so that the client can perform image editing and synchronous display.

[0049] In this context, the image region filling process performed by the image filling model can be based on image association features obtained through image association parsing. In one optional implementation of this embodiment, the image region filling process includes: Identify associated elements and / or associated image regions in the image after extracting image elements that are associated with the element image blocks; Image association analysis is performed on element image blocks and associated elements and / or associated image regions to obtain image association features; The element extraction region is filled with pixels according to the pixel filling parameters corresponding to the image association features.

[0050] The associated elements refer to image elements that have a logical relationship with the image element blocks in terms of semantics or space, such as people and their shadows, or cars and their shadows; the associated image regions refer to image regions that are associated with the image element blocks in terms of spatial layout, lighting environment, or structural continuity, such as the ground under a person's feet, or adjacent regions in the same lighting direction.

[0051] Specifically, in the process of image region filling, the process begins by identifying associated elements and / or associated image regions in the image after extracting image elements, based on the original image and the image after extracting image elements. Image association analysis is then performed on the image blocks and associated elements and / or associated image regions, such as analyzing light and shadow relationships, perspective relationships, and / or occlusion relationships, to obtain image association features. Then, pixel filling processing is performed on the element extraction region according to the pixel filling parameters corresponding to the image association features, thereby obtaining the filled image.

[0052] In addition, during the image region filling process, besides processing the element extraction region based on the associated elements and / or associated image regions of the element image block, the associated elements and / or associated image regions can also be processed. In one optional implementation of this embodiment, the image region filling process further includes: Determine pixel adjustment parameters for associated elements and / or associated image regions based on image association features; The filled image is obtained by adjusting the pixels of associated elements and / or associated image regions according to the pixel adjustment parameters.

[0053] Specifically, during the image region filling process, there may be situations where the displacement of element image blocks affects the rationality of associated elements and / or associated image regions. In such cases, the associated elements and / or associated image regions can be processed. Specifically, the pixel adjustment parameters of the associated elements and / or associated image regions can be determined first based on the image association features, and then the pixel adjustment parameters can be used to adjust the pixels of the associated elements and / or associated image regions to obtain the filled image.

[0054] It should be noted that the two image region filling methods provided above can be executed by the image filling model during the actual execution process; in addition, image region filling can also be performed without calling the image filling model. In this case, image region filling can be performed based on the image filling module of the intelligent agent without calling the image filling model. Based on this, image filling processing is performed on the element extraction region in the image from which image elements are extracted, including: if the real-time position of the element image block and the displacement data of the element extraction region meet the preset displacement filling conditions, determining the associated elements and / or associated image regions in the image after extracting image elements that are associated with the element image block; performing image association analysis on the element image block and the associated elements and / or associated image regions to obtain image association features; performing pixel filling processing on the element extraction region according to the pixel filling parameters corresponding to the image association features to obtain the filled image, and synchronizing it with the client for image editing and synchronous display; Alternatively, image filling processing can be performed on the element extraction region in the image from which image elements are extracted, including: if the real-time position of the element image block and the displacement data of the element extraction region meet the preset displacement filling conditions, determining the pixel adjustment parameters of the associated elements and / or associated image regions based on the image association features; adjusting the pixels of the associated elements and / or associated image regions according to the pixel adjustment parameters to obtain the filled image, and synchronizing it with the client for synchronized image editing and display.

[0055] In practice, based on obtaining the element image block and determining its editing position, in order to achieve the fusion and visual coordination between the processed element image block and the image environment, the implementation method for position adaptation processing of the element image block can be determined based on different editing intentions and context scenarios. The following provides four implementation methods for position adaptation processing of element image blocks, and each of the four implementation methods is explained in detail.

[0056] (1) Implementation method one In the specific execution process, after merging the element image blocks to the editing position, position adaptation processing of the element image blocks can be performed based on spatial relationship recognition; in an optional implementation method provided in this embodiment, the position adaptation processing of the element image blocks according to the editing position of the element image blocks includes: Merge the element image blocks into the image corresponding to the editing position to obtain the merged image; Spatial relationships between image elements and element image blocks in a merged image are identified, and image adjustment processing is performed on the element image blocks in the merged image based on the obtained image space.

[0057] Specifically, in the process of position adaptation of element image blocks, the element image blocks are first merged into the image corresponding to the editing position to obtain a merged image. For example, the element image blocks can be embedded into the underlying image according to the target coordinates, scaling ratio and / or rotation angle according to the transformation parameters. Subsequently, since there are certain physical spatial relationships between different objects in the real visual scene, such as occlusion, projection, perspective consistency, relative size ratio and / or lighting direction, the spatial relationship between the image elements and element image blocks in the merged image is identified, and the element image blocks in the merged image are adjusted based on the obtained image space.

[0058] For example, in the process of spatial relationship recognition of image elements and element image blocks, spatial relationship recognition can be performed based on a scene understanding model. Subsequently, in the process of image adjustment processing of element image blocks based on image space, the image adjustment processing can be multi-dimensional. For example, if spatial relationship recognition indicates that the element image block should be in a relatively far position, then the element image block is appropriately shrunk; if spatial relationship recognition indicates that the element image block should be occluded by other objects, then a corresponding mask is generated on the element image block to hide the occluded part; if spatial relationship recognition indicates the direction of the light source, then a shadow layer with a matching angle is added to the element image block, and the brightness and darkness distribution is adjusted.

[0059] For example, Figure 4 The vehicle in the middle moved from the corresponding position 401 to Figure 5 When the vehicle is positioned at location 501, it is first merged into the image corresponding to 501 to obtain a merged image. Considering that the vehicle is traveling horizontally when at position 401, but on a downward slope when at position 501, the vehicle angle needs to be adjusted appropriately to maintain perspective consistency. Based on this, the spatial relationship between the vehicle and the image elements surrounding it in the merged image is identified, and the vehicle image is adjusted based on the obtained image space. The adjusted image is shown below. Figure 5 As shown.

[0060] (2) Implementation Method Two In the specific execution process, the position adaptation processing of element image blocks can also be performed based on the image fusion dimension determined according to the fusion keywords. Specifically, the user's editing intent can be identified by parsing the fusion keywords in the user's instructions or editing instructions, and then the image fusion dimension can be determined to perform position adaptation processing on the element image blocks. In an optional implementation method provided in this embodiment, the position adaptation processing of element image blocks is performed according to the editing position of the element image blocks, including: Determine the image fusion dimensions based on the fusion keywords carried in user instructions or editing instructions; The element image blocks are merged into the image corresponding to the editing position to obtain the merged image, and the element image blocks and / or image elements in the merged image are adjusted according to the image fusion parameters corresponding to the image fusion dimension to obtain the target image.

[0061] Among them, image fusion dimension refers to the visual attribute dimension that affects the image fusion effect, such as lighting dimension, color dimension, texture dimension, edge blending, perspective dimension and / or shadow dimension.

[0062] Specifically, semantic parsing is performed on user commands or editing commands to determine the fusion keywords carried by the user commands or editing commands, and the image fusion dimension is further determined based on the fusion keywords; based on merging the element image blocks into the image corresponding to the editing position to obtain the merged image, the element image blocks and / or image elements in the merged image are processed according to the image fusion parameters corresponding to the image fusion dimension to obtain the target image.

[0063] (3) Implementation method three In the specific execution process, the element image blocks can also be adjusted and merged based on the image block adjustment parameters determined according to the fusion keywords, so as to achieve the fusion processing of the element image blocks and the images corresponding to the editing positions; in an optional implementation method provided in this embodiment, the fusion processing of the element image blocks and the images corresponding to the editing positions is performed according to the editing positions of the element image blocks, including: The image block adjustment parameters are determined based on the fusion keywords carried by the user instructions or editing instructions, and the element image blocks are adjusted according to the image block adjustment parameters; The adjusted element image blocks are merged into the image corresponding to the editing position to obtain the target image. Among them, image block adjustment parameters refer to a set of parameters used to control the appearance attributes of element image blocks; for example, image block adjustment parameters can be brightness adjustment parameters, contrast adjustment parameters, color shift parameters, blur parameters and / or transparency parameters.

[0064] Specifically, in the process of fusing an element image block with the image corresponding to the editing position based on the editing position of the element image block, firstly, fusion keywords are extracted based on user instructions or editing instructions, and image block adjustment parameters are determined based on the fusion keywords. Then, the element image block is adjusted according to the image block adjustment parameters. For example, the corresponding image processing module can be called to adjust the element image block according to the image block adjustment parameters. Finally, the adjusted element image block is merged into the image corresponding to the editing position to obtain the target image.

[0065] (4) Implementation method four In the specific execution process, during the fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block, a large language model can also be introduced, and the fusion processing of the element image block and the image corresponding to the editing position can be realized through the large language model; in an optional implementation method provided in this embodiment, the fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block includes: The adjusted element image blocks are merged into the image corresponding to the edit position to obtain the merged image; The merged images are input into a large language model for image semantic understanding and image element adjustment to obtain the target image.

[0066] Large Language Model (LLM) refers to a pre-trained natural language model. LLM can be based on foundation models or pre-trained models. The architecture of LLM can be a neural network architecture with a large number of parameters, a Transform architecture, or other architectures. Specifically, LLM can directly use foundation models or pre-trained models. Alternatively, based on foundation models or pre-trained models, the foundation models or pre-trained models can be fine-tuned for specific tasks such as image semantic understanding and image element adjustment to obtain a large language model capable of performing specific tasks such as image semantic understanding and image element adjustment.

[0067] Specifically, in the process of fusing the element image block with the image corresponding to the editing position based on the editing position of the element image block, the adjusted element image block can first be merged into the image corresponding to the editing position to obtain a merged image. Then, the merged image is input into the large language model so that the large language model can perform image semantic understanding and image element adjustment based on the merged image and output the target image. Based on this, the target image is obtained.

[0068] It should be noted that, in addition to the four implementation methods for adapting the position of element image blocks based on their editing positions provided above, new implementation methods can be obtained by combining the above-mentioned methods. Alternatively, one or more operations of the above-mentioned implementation methods can be combined to obtain new implementation methods. For example, based on the fusion keywords carried by user instructions or editing instructions, the image fusion dimension can be determined; the element image blocks can be merged into the image corresponding to the editing position to obtain a merged image; and the merged image and the image fusion parameters corresponding to the image fusion dimension can be input into a large language model for image adjustment processing to obtain the target image.

[0069] In practical applications, during the process of adapting the position of an element image block to its edit position, the user's operational intention can be predicted based on the user's operational trend, and the position can be adapted accordingly, thereby reducing the number of manual adjustment steps required by the user. In one optional implementation of this embodiment, the position adaptation processing of the element image block based on its editing position includes: Based on the displacement position, the displacement trajectory of the element image block is predicted to obtain the predicted trajectory, and the positional relationship between the predicted trajectory and the corresponding image is determined. The editing type of the element image block is determined based on its positional relationship, and the image editing of the element image block is adapted to the corresponding image according to the editing type.

[0070] Specifically, the displacement trajectory of the element image block can be predicted based on the displacement position to obtain the predicted trajectory. Specifically, the trajectory prediction model can be called to predict the displacement trajectory, and then the positional relationship between the predicted trajectory and the corresponding image can be determined. Based on this, the editing type of the element image block can be determined according to the positional relationship, such as perspective scaling or partial occlusion. Then, image editing adaptation can be performed on the element image block and the corresponding image according to the determined editing type.

[0071] In one optional implementation of this embodiment, during the image editing adaptation process, image editing adaptation of element image blocks and corresponding images is performed according to the editing type, including: If the edit type is deletion, the element deletion rendering data is generated based on the deletion rendering parameters and element image blocks corresponding to the deletion type, and the element deletion rendering data is synchronized to the client for synchronous display of deletion.

[0072] Specifically, if the user's intention is determined to delete an element image block based on user operation or displacement trajectory analysis, and the editing type is determined to be deletion, then the deletion rendering parameters corresponding to the deletion type are determined, and the element deletion rendering data is generated based on the deletion rendering parameters and the element image block. The element deletion rendering data is then synchronized to the client for synchronized deletion display.

[0073] For example, when a user wants to Figure 4 The car at position 401 is deleted from the image. Users can drag the car at position 401, for example, by moving it towards... Figure 4 The image shown is dragged along its right edge. When the user drags the car outside the image boundary, the system determines the user's intention to delete the car based on preset rules (such as the car leaving the image range by more than a certain threshold). The edit type is then set to delete. Based on this, further deletion rendering data is generated according to the deletion rendering parameters corresponding to the deletion type and the car to be deleted. Deletion rendering parameters can include the car's location and size, the type of deletion animation (such as fading or disappearing directly), and / or the method of background restoration after deletion. This element deletion rendering data is synchronized to the client. After receiving the element deletion rendering data, the client can then... Figure 6 The deletion is shown in the synchronized display.

[0074] Here, in the process of position adaptation of element image blocks based on their editing positions, in addition to determining the editing type of the element image block based on its positional relationship and performing position adaptation based on the editing type, based on the positional relationship determined above, it is also possible to directly detect whether the positional relationship meets the element deletion condition. If the positional relationship meets the element deletion condition, position adaptation is performed. Based on this, the above-mentioned position adaptation of element image blocks based on their editing positions includes: predicting the displacement trajectory of the element image block based on its displacement position to obtain a predicted trajectory, and determining the positional relationship between the predicted trajectory and the corresponding image; if the positional relationship meets the element deletion condition, generating element deletion rendering data based on the deletion rendering parameters and the element image block, and synchronizing the element deletion rendering data to the access terminal of the intelligent agent for deletion synchronization display.

[0075] It should be added that each optional implementation method and each feasible execution method in steps S202 to S206 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted. This embodiment does not specifically limit the subsequent operations after the execution conditions.

[0076] In summary, the image processing method based on intelligent agents provided in this embodiment allows users to first extract image blocks corresponding to the user commands input into the intelligent agent image component, and create layers of these image blocks. Then, based on editing commands, the image blocks are further edited within the layers. Image filling is performed on the extracted areas of the extracted image elements, and positional adaptation is performed based on the edited positions of the image blocks. The filling and adaptation results are then returned to the client for synchronous display. This enhances the convenience and efficiency of image processing based on intelligent agent image components, improving the user's interactive experience in image processing.

[0077] The following example uses the application of an agent-based image processing method provided in this embodiment in an image processing scenario, combined with... Figure 7 The image processing method based on intelligent agents provided in this embodiment will be further described below. Figure 7 An agent-based image processing method applied to image processing scenarios includes the following steps.

[0078] Step S704: Perform element segmentation processing on the image corresponding to the user instruction to obtain each image element, and determine the image element corresponding to the image position carried by the user instruction as the target element.

[0079] Step S706: Extract the element image block of the target element and create a layer of the element image block.

[0080] Step S708: Return to the client the editing data of the image block of the image element corresponding to the user instruction in the created layer.

[0081] Step S714: Based on the position information or trajectory information carried by the displacement command, perform displacement processing and displacement rendering of the element image block in the layer.

[0082] Step S716, and synchronize the displacement rendering data to the client for displacement synchronization display.

[0083] Step S720: Determine whether the image of the extracted element image block is the same image as the image corresponding to the editing position. If so, proceed to step S722.

[0084] Step S722: If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, determine the associated elements and associated image areas in the image after extracting the image elements that are associated with the element image block.

[0085] Step S724: Perform image association analysis on the element image block and associated elements and associated image regions to obtain image association features.

[0086] Step S726: Perform pixel filling processing on the element extraction region according to the pixel filling parameters corresponding to the image association features to obtain the filling processing result.

[0087] Step S728: Merge the element image blocks into the image corresponding to the editing position to obtain the merged image.

[0088] Step S730: Spatial relationship recognition is performed on the image elements and element image blocks in the merged image, and image adjustment processing is performed on the element image blocks in the merged image based on the obtained image space to obtain the adaptation processing result.

[0089] Step S732: Return to the client the image filling result of the element extraction region in the image from which the image elements are extracted, and the adaptation result of the position adaptation processing of the element image block.

[0090] It should be noted that any one or more steps in steps S704 to S708, S714 to S716, and S720 to S732 can be combined with any one or more steps in steps S202 to S206 to form a new implementation method according to the needs of implementation and deployment; in addition, any one of steps in steps S704 to S708, S714 to S716, and S720 to S732 can be selected according to the actual deployment needs. One or more technical features can be combined with one or more technical features provided in steps S202 to S206 to form a new implementation method; or, any one or more technical features in steps S704 to S708, S714 to S716, and S720 to S732 can be replaced with any one or more technical features provided in steps S202 to S206 to form a new implementation method according to the actual deployment needs, which will not be elaborated here.

[0091] Furthermore, it should be noted that steps S704 to S708, S714 to S716, and S720 to S732 provided in this embodiment can be executed by the server. It should be noted that the steps S704 to S708, S714 to S716, and S720 to S732 executed by the server can cooperate with steps S702, S710 to S712, S718, and S734 executed by the client in the following embodiment during execution. Therefore, when reading this embodiment, please refer to the corresponding content of steps S702, S710 to S712, S718, and S734 provided in the following method embodiment. When reading the following method embodiment, please refer to the corresponding content of steps S704 to S708, S714 to S716, and S720 to S732 provided in this embodiment.

[0092] One or more embodiments of another agent-based image processing method provided in this specification are as follows: Reference Figure 8 The image processing method based on intelligent agents provided in this embodiment specifically includes steps S802 to S808.

[0093] Step S802: Obtain the user command input by the user in the intelligent agent image component and submit it to the server.

[0094] The intelligent agent image component described in this embodiment refers to a component or plugin used for image processing. Specifically, it can be an image processing component integrated with an intelligent agent, or an interactive plugin integrated with an intelligent agent, such as an interactive plugin that interacts with the user to process images. Alternatively, the intelligent agent image component can also be an image editing tool integrated with an intelligent agent, including tools for cropping, repairing, removing objects, and / or filling objects in images. The intelligent agent image component integrates an intelligent agent, which includes an entity capable of autonomously performing tasks, making decisions, and learning and adjusting according to environmental changes. Optionally, the intelligent agent image component includes an image editing tool integrated with an intelligent agent.

[0095] The user instruction refers to the instruction input by the user through the intelligent agent's image component or the intelligent agent's dialogue interface. Specifically, it can be an instruction used by the user to express the intention to edit or manipulate the image content. The user instruction can be any form of instruction. Specifically, the user instruction can be a text instruction, such as natural text language entered by the user in the intelligent agent's dialog box or input box to express the image processing intention. The user instruction can also be a semantic instruction, such as an instruction obtained by semantic recognition of the image processing instruction input by the user's voice. The user instruction can also be an action instruction, such as a specific operation behavior performed by the user through touch interaction. Optionally, user instructions may include text or semantic instructions entered into the agent's image component or the agent's dialogue interface, or action instructions entered into the agent's image component.

[0096] In practice, when a user accesses the intelligent agent image component to perform image processing, the user can input user commands expressing the intention of image processing through the intelligent agent image component. Correspondingly, the user commands input by the user in the intelligent agent image component are obtained and submitted to the server so that the server can perform subsequent image processing operations based on the user commands.

[0097] Step S804: Obtain the editing data of the image block of the image element corresponding to the user instruction in the created layer from the server and display the edit synchronously.

[0098] As described above, after obtaining the user's input command in the intelligent agent image component and submitting it to the server, the server obtains the user command and determines the image elements in the image to be processed based on the user command submitted by the client. That is, the server determines the image element corresponding to the user command and extracts the element image block of the image element.

[0099] The image element refers to an identifiable visual object or image unit in the image, specifically the target object that the user intends to manipulate through user commands; the element image block refers to a sub-image data block extracted from the image that corresponds to a certain image element; the element image block is the direct object of operation processed during image processing; optionally, the element image block includes element images generated by the intelligent agent.

[0100] Here, the server can be a server applied to the intelligent agent, and correspondingly, the client can be the client of the intelligent agent, or the user's access client or user terminal; in addition, the server can also be the server of the intelligent agent image component. In this case, the execution entity that interacts with the server can be the intelligent agent image component. Based on this, the client in this embodiment can also be replaced by the intelligent agent image component.

[0101] In the specific execution process, during image processing, the user-submitted instructions can be action instructions or trigger instructions. Based on this, the server can determine image elements based on the user-submitted action instructions or trigger instructions. In an optional implementation provided in this embodiment, the image elements corresponding to the user instructions are determined in the following way: The image corresponding to the user command is segmented to obtain each image element, and the image element corresponding to the image position carried by the user command is determined as the target element.

[0102] Specifically, after the server obtains the user's input command in the intelligent agent image component and determines the image corresponding to the user's command, the server first performs element segmentation processing on the image to obtain each image element. For example, it can call an image segmentation model to perform element segmentation processing on the image to identify the boundaries of different objects in the image, thereby decomposing the image into multiple image elements. On this basis, the server further determines the target element in each image element based on the image position carried by the user command, thereby obtaining the target element.

[0103] Furthermore, when the user-submitted command is a text command or a voice command, the server can also determine the image element based on the text command or voice command; in another optional implementation provided in this embodiment, the image element corresponding to the user command is determined in the following way: The image corresponding to the user command is subjected to image semantic recognition to obtain each image semantic element, and the element keywords carried by the user command are matched with each image semantic element to obtain the target element.

[0104] For example, if a user inputs a command to delete the blue vehicle on the right side of the image, the server first performs image semantic recognition on the image corresponding to the user's command to obtain the semantic elements of each image. Then, it extracts element keywords from the user's command to obtain element keywords such as "right side", "blue", and "vehicle". Finally, it performs image semantic matching between the element keywords and each image semantic element to obtain the target element.

[0105] Here, during the process of determining image elements, the image elements can be image elements in images uploaded by the user, or image elements generated by the intelligent agent.

[0106] In practice, the server extracts the image elements corresponding to the user's input command in the intelligent agent image component, creates a layer of image elements based on the image elements of the image elements, and then returns the editing data of the image elements in the created layer to the client. Correspondingly, here, the editing data of the image elements corresponding to the user's command in the created layer is obtained from the server, and the editing data is edited and displayed synchronously.

[0107] Here, the editing data of the image block corresponding to the user command in the created layer is obtained from the server and the editing is displayed synchronously; it can also be replaced by: receiving the editing data of the image block corresponding to the user command in the created layer synchronized by the server and the editing is displayed synchronously.

[0108] Step S806: Obtain the editing instructions input by the user and submit them to the server.

[0109] In practice, when users access the intelligent agent image component to process images, they can also submit editing instructions. Based on this, the system obtains the editing instructions input by the user and submits them to the server, so that the server can perform editing processing of the element image blocks in the layer according to the editing instructions.

[0110] The editing instructions refer to user-inputted instructions for manipulating element image blocks, such as displacement instructions, scaling instructions, rotation instructions, and / or deletion instructions. Editing instructions can be any form of instruction; specifically, editing instructions can be text instructions, semantic instructions, action instructions, or trigger instructions. Optionally, editing instructions include text instructions or semantic instructions input into the agent image component or the agent's dialogue interface, or action instructions input into the agent image component.

[0111] It should be noted that the editing command submitted by the user and the user input command mentioned above can be the same command or different commands. For example, if the editing command and the user command are the same command, and the user long-presses and drags an image element in the image, the image element corresponding to the user command can be determined based on the user's long-press action, and the editing command can be determined to be a displacement operation on the element image block based on the user's drag operation. In this case, when the editing command and the user command are the same command, the server's editing of the element image block on the layer according to the editing command can be replaced by: editing the element image block on the layer according to the user command.

[0112] In the specific execution process, during the editing of element image blocks on the server side, the server can edit the element image blocks based on the user's action commands or trigger commands for the image elements; in one optional implementation of this embodiment, the editing of element image blocks on the layer according to editing commands includes: Based on the position or trajectory information carried by the displacement command, the element image block is displaced and rendered on the layer, and the displacement rendering data is synchronized to the client for synchronous display.

[0113] Specifically, when a user submits a displacement command through a trigger operation, the server determines the position or trajectory information carried by the displacement command based on the user's displacement operation on the element image, and performs displacement processing and displacement rendering of the element image block on the layer based on the position or trajectory information. After that, the server synchronizes the displacement rendering data to the client. Correspondingly, here, the server receives the displacement rendering data synchronized by the server and performs displacement synchronization display, that is, it obtains the displacement rendering data processed and rendered by the server and performs displacement synchronization display.

[0114] For example, a user submits a displacement command by dragging an image element to a target location. The server determines the location information based on the coordinate information of the target location carried in the displacement command submitted by the user, or determines the trajectory information based on the coordinate sequence continuously generated during the user's dragging process. Subsequently, the server updates the spatial coordinate parameters of the element image block in the layer and synchronously adjusts the display position of the element image block in the image to achieve displacement processing. The server also draws the movement process of the element image block in real time according to the trajectory information to achieve displacement rendering.

[0115] In addition, the server can extract editing keywords based on the user's text or voice commands for editing image elements, and perform editing processing on the element image blocks based on the editing keywords; in another optional implementation provided in this embodiment, the editing processing of element image blocks on the layer according to editing commands includes: The displacement position of the element image block is determined based on the displacement keyword carried by the displacement command, and the displacement processing of the element image block is performed on the layer according to the displacement position. The displacement processing result is then synchronized to the client.

[0116] Specifically, the server first determines the displacement command based on the user's input text or voice command. Then, based on the displacement keywords carried in the displacement command indicating the target location, it determines the displacement position of the element image block and performs displacement processing on the layer according to the displacement position. Afterward, the server synchronizes the displacement processing result to the client. Correspondingly, here, the server receives the synchronized displacement rendering data and displays it synchronously. Here, the displacement command and the aforementioned user command can be the same command.

[0117] Step S808: Receive the image filling processing result returned by the server for the element extraction region in the image from which the image elements are extracted, and display it synchronously; and receive the adaptation processing result returned by the server for the position adaptation processing of the element image block, and display it synchronously.

[0118] In practical applications, to avoid obvious editing marks in the processed image, or to prevent damage to the overall composition of the image due to lighting imbalance or spatial misalignment, the server can determine the element extraction area in the image where the image elements have been extracted, and perform image filling processing on the element extraction area; in addition, the server can also perform position adaptation processing on the element image block according to the editing position of the element image block when the element image block is moved to the editing position.

[0119] Among them, the element extraction region refers to the area left in the image to be repaired after the image element is identified and segmented and extracted; the position adaptation processing refers to the process by which the server adjusts the appearance attributes of the element image block according to the visual environment characteristics of the editing position after the element image block is moved to the new editing position, so as to make the element image block consistent with the context of the editing position; for example, the server can perform position adaptation processing on the element image block based on the lighting, shadow, perspective and / or blur degree of the editing position.

[0120] Optionally, images of extracted image elements and / or images corresponding to edited locations are generated by the agent.

[0121] In the actual execution process, users may not only move or modify element image blocks within the original image, but may also add element image blocks from one image to another image for compositing. In this case, on the basis of the server's editing of the element image blocks, and before the image filling process of the element extraction area, the server can also determine whether the original image from which the element image block is extracted is the same image as the image corresponding to the current editing position of the element image block. In one optional implementation of this embodiment, after editing the element image block on the layer according to the editing instructions, the method further includes: Determine whether the image of the extracted element image block is the same image as the image corresponding to the edit position; If so, perform image filling processing on the element extraction regions in the image from which the image elements are extracted; If not, perform fusion processing between the element image block and the image corresponding to the edit position based on the edit position of the element image block.

[0122] Specifically, in addition to editing the element image blocks on the layer according to the editing instructions, the server can also judge the image of the extracted element image block. Specifically, it can judge whether the image of the extracted element image block is the same image as the image corresponding to the editing position. If so, it indicates that the editing operation is an internal structural adjustment of the image, and the server performs image filling processing on the element extraction area in the image of the extracted image element. If not, it indicates that the editing operation is a cross-image fusion editing operation, and the server performs fusion processing on the element image block and the image corresponding to the editing position according to the editing position of the element image block.

[0123] Furthermore, if the server determines that the image of the extracted element image block and the image corresponding to the editing position are the same image and the result is negative, in addition to the above-mentioned fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block, the server can also perform: predict the displacement trajectory of the element image block based on the displacement position to obtain the predicted trajectory, and determine the positional relationship between the predicted trajectory and the corresponding image; determine the editing type of the element image block based on the positional relationship, and perform image editing adaptation of the element image block and the corresponding image according to the editing type; this embodiment does not limit this.

[0124] In the specific execution process, during the image filling process of the element extraction region, in order to ensure the continuity and integrity of the image being processed and to guarantee the user's interactive experience, the server can perform image filling processing on the element extraction region that is temporarily in a "gap" state after extracting the image elements. In one optional implementation of this embodiment, image filling processing on the element extraction region in the image from which image elements are extracted includes: If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, the image after extracting the image elements is input into the image filling model for image region filling processing to obtain the filled image, and synchronized to the client for image editing and synchronous display.

[0125] The preset displacement filling condition refers to the preset logical condition used to determine whether to trigger image filling processing. The preset displacement filling condition can be set based on the displacement distance or the user's operation state. For example, the displacement distance can be set to be greater than or equal to a displacement threshold, or the current user's operation state can be set to "release", or the element image block can be set to enter the target area. An image filling model is a model used for image region filling processing. The input of an image filling model is the image after extracting image elements, and the output is the filled image obtained after image region filling processing. Here, the filled image refers to the image in which the content of the extracted element regions in the image has been supplemented.

[0126] During the specific execution process, the server can detect the real-time position of the element image block and the displacement data of the real-time position of the element image block relative to the element extraction area. If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, the server will input the image after extracting the image elements into the image filling model so that the image filling model can perform image area filling processing and output the filled image. After that, the server obtains the filled image output by the image filling model and synchronizes it with the client. Correspondingly, here, the filled image returned by the server is obtained and the image is edited and displayed synchronously.

[0127] In this context, the image region filling process performed by the image filling model can be based on image association features obtained through image association parsing. In one optional implementation of this embodiment, the image region filling process includes: Identify associated elements and / or associated image regions in the image after extracting image elements that are associated with the element image blocks; Image association analysis is performed on element image blocks and associated elements and / or associated image regions to obtain image association features; The element extraction region is filled with pixels according to the pixel filling parameters corresponding to the image association features.

[0128] The associated elements refer to image elements that have a logical relationship with the image element blocks in terms of semantics or space, such as people and their shadows, or cars and their shadows; the associated image regions refer to image regions that are associated with the image element blocks in terms of spatial layout, lighting environment, or structural continuity, such as the ground under a person's feet, or adjacent regions in the same lighting direction.

[0129] Specifically, in the process of image region filling, the process begins by identifying associated elements and / or associated image regions in the image after extracting image elements, based on the original image and the image after extracting image elements. Image association analysis is then performed on the image blocks and associated elements and / or associated image regions, such as analyzing light and shadow relationships, perspective relationships, and / or occlusion relationships, to obtain image association features. Then, pixel filling processing is performed on the element extraction region according to the pixel filling parameters corresponding to the image association features, thereby obtaining the filled image.

[0130] In addition, during the image region filling process, besides processing the element extraction region based on the associated elements and / or associated image regions of the element image block, the associated elements and / or associated image regions can also be processed. In one optional implementation of this embodiment, the image region filling process further includes: Determine pixel adjustment parameters for associated elements and / or associated image regions based on image association features; The filled image is obtained by adjusting the pixels of associated elements and / or associated image regions according to the pixel adjustment parameters.

[0131] Specifically, during the image region filling process, there may be situations where the displacement of element image blocks affects the rationality of associated elements and / or associated image regions. In such cases, the associated elements and / or associated image regions can be processed. Specifically, the pixel adjustment parameters of the associated elements and / or associated image regions can be determined first based on the image association features, and then the pixel adjustment parameters can be used to adjust the pixels of the associated elements and / or associated image regions to obtain the filled image.

[0132] It should be noted that the two image region filling methods provided above can be executed by the image filling model during the actual execution process. Furthermore, the server may not call the image filling model for image region filling. In this case, without calling the image filling model, the server can perform image region filling based on the agent's image filling module. Based on this, image filling processing is performed on the element extraction region in the image from which image elements are extracted, including: if the real-time position of the element image block and the displacement data of the element extraction region meet the preset displacement filling conditions, determining the associated elements and / or associated image regions in the image after image element extraction that are associated with the element image block; performing image association analysis on the element image block and the associated elements and / or associated image regions to obtain image association features; performing pixel filling processing on the element extraction region according to the pixel filling parameters corresponding to the image association features to obtain a filled image, and synchronizing it with the client for synchronized image editing and display. Alternatively, image filling processing can be performed on the element extraction region in the image from which image elements are extracted, including: if the real-time position of the element image block and the displacement data of the element extraction region meet the preset displacement filling conditions, determining the pixel adjustment parameters of the associated elements and / or associated image regions based on the image association features; adjusting the pixels of the associated elements and / or associated image regions according to the pixel adjustment parameters to obtain the filled image, and synchronizing it with the client for synchronized image editing and display.

[0133] In practice, after obtaining the element image block and determining its editing position, the server can also determine the implementation method for position adaptation processing of the element image block based on different editing intentions and context scenarios in order to achieve fusion and visual coordination between the processed element image block and the image environment. The following provides four implementation methods for position adaptation processing of element image blocks, and explains each of the four implementation methods in detail.

[0134] (1) Implementation method one In the specific execution process, after merging the element image blocks to the editing position, the server can perform position adaptation processing on the element image blocks based on spatial relationship recognition. In one optional implementation of this embodiment, the position adaptation processing of the element image blocks according to the editing position of the element image blocks includes: Merge the element image blocks into the image corresponding to the editing position to obtain the merged image; Spatial relationships between image elements and element image blocks in a merged image are identified, and image adjustment processing is performed on the element image blocks in the merged image based on the obtained image space.

[0135] Specifically, during the server-side position adaptation process for element image blocks, the server first merges the element image blocks into the image corresponding to the editing position to obtain a merged image. For example, the element image blocks can be embedded into the underlying image according to the target coordinates, scaling ratio and / or rotation angle. Subsequently, since there are certain physical spatial relationships between different objects in the real visual scene, such as occlusion, projection, perspective consistency, relative size ratio and / or lighting direction, the server can identify the spatial relationship between the image elements and element image blocks in the merged image, and perform image adjustment processing on the element image blocks in the merged image based on the obtained image space.

[0136] For example, during the process of spatial relationship recognition of image elements and element image blocks on the server side, the server can perform spatial relationship recognition based on a scene understanding model. Subsequently, during the image adjustment processing of element image blocks based on image space, the image adjustment processing performed by the server can be multi-dimensional. For example, if spatial relationship recognition indicates that the element image block should be in a relatively far position, then the element image block is appropriately shrunk; if spatial relationship recognition indicates that the element image block should be occluded by other objects, then a corresponding mask is generated on the element image block to hide the occluded part; if spatial relationship recognition indicates the direction of the light source, then a shadow layer with a matching angle is added to the element image block, and the brightness and darkness distribution is adjusted.

[0137] (2) Implementation Method Two In the specific execution process, the server can also perform position adaptation processing on the element image blocks based on the image fusion dimension determined according to the fusion keywords. Specifically, this can be achieved by parsing the fusion keywords in user commands or editing commands to identify the user's editing intent, and then determining the image fusion dimension to perform position adaptation processing on the element image blocks. In one optional implementation of this embodiment, the position adaptation processing of the element image blocks based on the editing position of the element image blocks includes: Determine the image fusion dimensions based on the fusion keywords carried in user instructions or editing instructions; The element image blocks are merged into the image corresponding to the editing position to obtain the merged image, and the element image blocks and / or image elements in the merged image are adjusted according to the image fusion parameters corresponding to the image fusion dimension to obtain the target image.

[0138] Among them, image fusion dimension refers to the visual attribute dimension that affects the image fusion effect, such as lighting dimension, color dimension, texture dimension, edge blending, perspective dimension and / or shadow dimension.

[0139] Specifically, the server can perform semantic parsing on user commands or editing commands to determine the fusion keywords carried by the user commands or editing commands, and further determine the image fusion dimension based on the fusion keywords; then, based on the merged image obtained by merging the element image blocks into the image corresponding to the editing position on the server, the element image blocks and / or image elements in the merged image are processed according to the image fusion parameters corresponding to the image fusion dimension to obtain the target image.

[0140] (3) Implementation method three In the specific execution process, the server can also adjust and merge the element image blocks based on the image block adjustment parameters determined according to the fusion keywords, so as to realize the fusion processing of the element image blocks and the images corresponding to the editing positions; in an optional implementation method provided in this embodiment, the fusion processing of the element image blocks and the images corresponding to the editing positions according to the editing positions of the element image blocks includes: The image block adjustment parameters are determined based on the fusion keywords carried by the user instructions or editing instructions, and the element image blocks are adjusted according to the image block adjustment parameters; The adjusted element image blocks are merged into the image corresponding to the editing position to obtain the target image. Among them, image block adjustment parameters refer to a set of parameters used to control the appearance attributes of element image blocks; for example, image block adjustment parameters can be brightness adjustment parameters, contrast adjustment parameters, color shift parameters, blur parameters and / or transparency parameters.

[0141] Specifically, in the process of fusing the element image block with the image corresponding to the editing position based on the editing position of the element image block, the server first extracts fusion keywords based on user instructions or editing instructions, determines the image block adjustment parameters based on the fusion keywords, and then adjusts the element image block according to the image block adjustment parameters. For example, the corresponding image processing module can be called to adjust the element image block according to the image block adjustment parameters. Finally, the adjusted element image block is merged into the image corresponding to the editing position to obtain the target image.

[0142] (4) Implementation method four In the specific execution process, during the fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block on the server side, a large language model can also be introduced. Based on this, the server side can call the large language model to realize the fusion processing of the element image block and the image corresponding to the editing position. In an optional implementation method provided in this embodiment, the fusion processing of the element image block and the image corresponding to the editing position based on the editing position of the element image block includes: The adjusted element image blocks are merged into the image corresponding to the edit position to obtain the merged image; The merged images are input into a large language model for image semantic understanding and image element adjustment to obtain the target image.

[0143] Large Language Model (LLM) refers to a pre-trained natural language model. LLM can be based on foundation models or pre-trained models. The architecture of LLM can be a neural network architecture with a large number of parameters, a Transform architecture, or other architectures. Specifically, LLM can directly use foundation models or pre-trained models. Alternatively, based on foundation models or pre-trained models, the foundation models or pre-trained models can be fine-tuned for specific tasks such as image semantic understanding and image element adjustment to obtain a large language model capable of performing specific tasks such as image semantic understanding and image element adjustment.

[0144] Specifically, during the process of fusing the element image block with the image corresponding to the editing position on the server side, the server can first merge the adjusted element image block into the image corresponding to the editing position to obtain a merged image. Then, the merged image is input into the large language model so that the large language model can perform image semantic understanding and image element adjustment based on the merged image and output the target image. Based on this, the target image is obtained.

[0145] It should be noted that, in addition to the four server-side implementation methods for adapting the position of element image blocks based on their editing positions, as provided above, new implementation methods can be obtained by combining these methods. Alternatively, one or more operations from the aforementioned implementation methods can be combined to obtain new implementation methods. For example, based on the fusion keywords carried by user instructions or editing instructions, the image fusion dimension can be determined; the element image blocks can be merged into the image corresponding to the editing position to obtain a merged image; and the merged image and the image fusion parameters corresponding to the image fusion dimension can be input into a large language model for image adjustment processing to obtain the target image.

[0146] In practical applications, during the process of adapting the position of the element image block to the edit position on the server side, the user's operation intention can be predicted based on the user's operation trend, and the position can be adapted accordingly, thereby reducing the steps of manual adjustment by the user. In one optional implementation of this embodiment, the position adaptation processing of the element image block based on its editing position includes: Based on the displacement position, the displacement trajectory of the element image block is predicted to obtain the predicted trajectory, and the positional relationship between the predicted trajectory and the corresponding image is determined. The editing type of the element image block is determined based on its positional relationship, and the image editing of the element image block is adapted to the corresponding image according to the editing type.

[0147] Specifically, the server can predict the displacement trajectory of the element image block based on the displacement position to obtain the predicted trajectory. Specifically, it can call the trajectory prediction model to predict the displacement trajectory and then determine the positional relationship between the predicted trajectory and the corresponding image. Based on this, the server determines the editing type of the element image block according to the positional relationship, such as perspective scaling or partial occlusion, and then performs image editing adaptation on the element image block and the corresponding image according to the determined editing type.

[0148] In one optional implementation of this embodiment, during the image editing adaptation process, image editing adaptation of element image blocks and corresponding images is performed according to the editing type, including: If the edit type is deletion, the element deletion rendering data is generated based on the deletion rendering parameters and element image blocks corresponding to the deletion type, and the element deletion rendering data is synchronized to the client for synchronous display of deletion.

[0149] Specifically, if the user's intention is determined to delete an element image block based on user operation or displacement trajectory analysis, and the editing type is determined to be deletion, then the deletion rendering parameters corresponding to the deletion type are determined, and the element deletion rendering data is generated based on the deletion rendering parameters and the element image block. The element deletion rendering data is then synchronized to the client for synchronized deletion display.

[0150] Here, in the process of position adaptation of element image blocks based on their editing positions, in addition to determining the editing type of the element image block based on its positional relationship and performing position adaptation based on the editing type, based on the positional relationship determined above, it is also possible to directly detect whether the positional relationship meets the element deletion condition. If the positional relationship meets the element deletion condition, position adaptation is performed. Based on this, the above-mentioned position adaptation of element image blocks based on their editing positions includes: predicting the displacement trajectory of the element image block based on its displacement position to obtain a predicted trajectory, and determining the positional relationship between the predicted trajectory and the corresponding image; if the positional relationship meets the element deletion condition, generating element deletion rendering data based on the deletion rendering parameters and the element image block, and synchronizing the element deletion rendering data to the access terminal of the intelligent agent for deletion synchronization display.

[0151] In specific implementation, the server performs image filling processing on the element extraction area of ​​the image containing the extracted image elements, and performs position adaptation processing on the element image block according to the editing position of the element image block. The server can then return the filling processing result and the adaptation processing result to the client. Correspondingly, here, the server receives and synchronously displays the image filling processing result of the element extraction area of ​​the image containing the extracted image elements, and also receives and synchronously displays the adaptation processing result of the element image block position adaptation processing.

[0152] Here, step S808 can also be replaced by: if the server returns (synchronously) the image filling processing result of the element extraction region in the image of the extracted image elements, the image filling processing result is displayed synchronously; or, it can also be replaced by: if the server returns (synchronously) the adaptation processing result of the element image block position adaptation processing, the adaptation processing result is displayed synchronously; this embodiment does not limit this.

[0153] It should be added that each optional implementation method and each feasible execution method in steps S802 to S808 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted. This embodiment does not specifically limit the subsequent operations after the execution conditions.

[0154] In summary, the image processing method based on intelligent agents provided in this embodiment allows users to first obtain user commands input into the intelligent agent image component and submit them to the server. The server then extracts image blocks of the corresponding image elements based on the user commands, creates layers of these image blocks, and retrieves the editing data of the image blocks in the created layers from the server, displaying the edited data synchronously. Next, the server retrieves user-inputted editing commands and submits them to the server, enabling the server to edit the image blocks in the layers according to the commands. Furthermore, the server performs image filling processing on the extracted areas of the extracted image elements and performs position adaptation processing on the image blocks based on their edited positions. The server then returns the filling and adaptation results to the client. Correspondingly, the server receives and displays the image filling results for the extracted areas of the extracted image elements, and also receives and displays the position adaptation results for the image blocks. This improves the convenience and efficiency of image processing based on intelligent agent image components, enhancing the user's image processing interaction experience.

[0155] The following example uses the application of an agent-based image processing method provided in this embodiment in an image processing scenario, combined with... Figure 7 The image processing method based on intelligent agents provided in this embodiment will be further described below. Figure 7 An agent-based image processing method applied to image processing scenarios includes the following steps.

[0156] Step S702: Obtain the user command input by the user in the intelligent agent image component and submit it to the server.

[0157] Step S710: Obtain the editing data of the image block of the image element corresponding to the user instruction in the created layer from the server and display the edit synchronously.

[0158] Step S712: Obtain the editing instructions input by the user and submit them to the server.

[0159] Step S718: Obtain displacement rendering data from the server and perform displacement synchronization display.

[0160] Step S734: Receive the image filling processing result returned by the server for the element extraction region in the image from which the image elements are extracted, and display it synchronously; and receive the adaptation processing result returned by the server for the position adaptation processing of the element image block, and display it synchronously.

[0161] It should be noted that any one or more of steps S702, S710 to S712, S718, and S734 can be combined with any one or more of steps S802 to S808 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features can be selected from steps S702, S710 to S712, S718, and S734 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features from steps S702, S710 to S712, S718, and S734 can be replaced with any one or more of the technical features provided in steps S802 to S808 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0162] This specification provides an embodiment of an image processing device based on an intelligent agent as follows: In the above embodiments, an image processing method based on intelligent agents is provided, and correspondingly, an image processing device based on intelligent agents is also provided, which will be described below with reference to the accompanying drawings.

[0163] Reference Figure 9 This illustration shows a schematic diagram of an embodiment of an image processing device based on an intelligent agent provided in this embodiment.

[0164] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.

[0165] This embodiment provides an image processing device based on an intelligent agent, the device comprising: The image block extraction module 902 is configured to extract element image blocks of the image elements according to the image elements corresponding to the user commands input by the user in the intelligent agent image component; Image block editing module 904 is configured to create a layer of the element image block and perform editing processing of the element image block on the layer according to editing instructions; The position adaptation module 906 is configured to perform image filling processing on the element extraction region in the image from which the image elements are extracted, and to perform position adaptation processing on the element image block according to the editing position of the element image block.

[0166] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0167] Another embodiment of the agent-based image processing device provided in this specification is as follows: In the above embodiments, another agent-based image processing method is provided, and correspondingly, another agent-based image processing apparatus is also provided, which will be described below with reference to the accompanying drawings.

[0168] Reference Figure 10 This illustration shows a schematic diagram of an embodiment of an image processing device based on an intelligent agent provided in this embodiment.

[0169] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.

[0170] This embodiment provides an image processing device based on an intelligent agent, the device comprising: The instruction acquisition module 1002 is configured to acquire user instructions input by the user in the intelligent agent image component and submit them to the server. The editing and display module 1004 is configured to obtain the editing processing data of the image block of the image element corresponding to the user instruction in the created layer from the server and perform synchronous editing and display. The instruction submission module 1006 is configured to obtain the editing instruction input by the user and submit it to the server; The result display module 1008 is configured to receive and synchronously display the image filling processing result returned by the server for the element extraction region in the image from which the image elements are extracted, and to receive and synchronously display the adaptation processing result returned by the server for the position adaptation processing of the element image block.

[0171] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0172] This specification provides an embodiment of an intelligent agent-based image processing device as follows: Corresponding to the agent-based image processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide an agent-based image processing device for executing the agent-based image processing method described above. Figure 11 This is a schematic diagram of the structure of an intelligent agent-based image processing device provided for one or more embodiments of this specification.

[0173] This embodiment provides an image processing device based on an intelligent agent, comprising: like Figure 11As shown, device 1100 mainly consists of a communication interface 1102, a user interface 1104, a processor 1106, and a data storage 1108. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1110. The communication interface 1102 enables device 1100 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1102 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1102 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1102 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1102 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1104 includes receiving user input and providing output to the user. Therefore, user interface 1104 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1104 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1104 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1100 may support remote access from other devices via communication interface 1102 or another physical interface (not shown). User interface 1104 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1104 may also be configured as a display device for rendering or displaying text fragments.

[0174] Processor 1106 may include one or more general-purpose processors and / or dedicated processors. Data storage 1108 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1106. Data storage 1108 may include removable and non-removable components.

[0175] Processor 1106 is capable of executing program instructions 1118 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1108 to perform the various functions described herein. Data store 1108 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1100, enable device 1100 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1118 by processor 1106 may result in processor 1106 using data 1112. For example, program instructions 1118 may include an operating system 1122 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1100 and one or more application programs 1120 (e.g., a browser, social application, or game application). Similarly, data 1112 may include operating system data 1116 and application data 1114. Operating system data 1116 is primarily accessible to operating system 1122, while application data 1114 is primarily accessible to one or more application programs 1120. Application data 1114 may reside in a file system visible or hidden to the user of device 1100. Application 1120 may communicate with operating system 1122 via one or more application programming interfaces (APIs). These APIs facilitate application 1120 reading and / or writing application data 1114, transmitting or receiving information via communication interface 1102, receiving or displaying information on user interface 1104, etc. In some terms, application 1120 may be simply referred to as an "app". Furthermore, application 1120 may be downloaded to device 1100 through one or more online app stores or app markets. However, applications may also be installed on device 1100 in other ways, such as through a web browser or a physical interface on device 1100 (e.g., a USB port).

[0176] In one specific embodiment, the agent-based image processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the agent-based image processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Based on the image elements corresponding to the user commands input by the user in the intelligent agent image component, extract the element image blocks of the image elements; Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions; The image filling process is performed on the element extraction region in the image from which the image elements are extracted, and the position adaptation process of the element image block is performed according to the editing position of the element image block.

[0177] Another embodiment of an agent-based image processing device provided in this specification is as follows: Corresponding to the other agent-based image processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide another agent-based image processing apparatus for executing the other agent-based image processing method provided above. Figure 12 This is a schematic diagram of another agent-based image processing device provided in one or more embodiments of this specification.

[0178] This embodiment provides an image processing device based on an intelligent agent, comprising: like Figure 12As shown, device 1200 mainly consists of a communication interface 1202, a user interface 1204, a processor 1206, and a data storage 1208. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1210. The communication interface 1202 enables device 1200 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1202 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1202 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1202 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1202 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1204 includes receiving user input and providing output to the user. Therefore, user interface 1204 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1204 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1204 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1200 may support remote access from other devices via communication interface 1202 or another physical interface (not shown). User interface 1204 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1204 may also be configured as a display device for rendering or displaying text fragments.

[0179] Processor 1206 may include one or more general-purpose processors and / or dedicated processors. Data storage 1208 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1206. Data storage 1208 may include removable and non-removable components.

[0180] Processor 1206 is capable of executing program instructions 1218 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1208 to perform the various functions described herein. Data storage 1208 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1200, enable device 1200 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1218 by processor 1206 may result in processor 1206 using data 1212. For example, program instructions 1218 may include an operating system 1222 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1200 and one or more application programs 1220 (e.g., a browser, social application, or game application). Similarly, data 1212 may include operating system data 1216 and application data 1214. Operating system data 1216 is primarily accessible to operating system 1222, while application data 1214 is primarily accessible to one or more application programs 1220. Application data 1214 may reside in a file system visible or hidden from the user of device 1200. Application 1220 may communicate with operating system 1222 via one or more application programming interfaces (APIs). These APIs facilitate application 1220 reading and / or writing application data 1214, transmitting or receiving information via communication interface 1202, receiving or displaying information on user interface 1204, etc. In some terms, application 1220 may be simply referred to as an "app". Furthermore, application 1220 may be downloaded to device 1200 through one or more online app stores or app markets. However, applications may also be installed on device 1200 in other ways, such as through a web browser or a physical interface on device 1200 (e.g., a USB port).

[0181] In one specific embodiment, the agent-based image processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the agent-based image processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Obtain user commands input into the intelligent agent image component and submit them to the server; The server retrieves the editing data of the image block corresponding to the user instruction in the created layer, and displays the edited data synchronously. Obtain the editing instructions input by the user and submit them to the server; The system receives and synchronously displays the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and also receives and synchronously displays the adaptation results returned by the server for the position adaptation of the element image blocks.

[0182] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the above-described agent-based image processing method, and based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.

[0183] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Based on the image elements corresponding to the user commands input by the user in the intelligent agent image component, extract the element image blocks of the image elements; Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions; The image filling process is performed on the element extraction region in the image from which the image elements are extracted, and the position adaptation process of the element image block is performed according to the editing position of the element image block.

[0184] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of an image processing method based on an intelligent agent described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0185] Another embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the other agent-based image processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.

[0186] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Obtain user commands input into the intelligent agent image component and submit them to the server; The server retrieves the editing data of the image block corresponding to the user instruction in the created layer, and displays the edited data synchronously. Obtain the editing instructions input by the user and submit them to the server; The system receives and synchronously displays the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and also receives and synchronously displays the adaptation results returned by the server for the position adaptation of the element image blocks.

[0187] It should be noted that the embodiments of another computer-readable storage medium described in this specification and the embodiments of another intelligent agent-based image processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0188] This specification provides an example of a computer program product as follows: Corresponding to the above-described agent-based image processing method, and based on the same technical concept, one or more embodiments of this specification also provide a computer program product.

[0189] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Based on the image elements corresponding to the user commands input by the user in the intelligent agent image component, extract the element image blocks of the image elements; Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions; The image filling process is performed on the element extraction region in the image from which the image elements are extracted, and the position adaptation process of the element image block is performed according to the editing position of the element image block.

[0190] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of an image processing method based on an intelligent agent described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0191] Another example of a computer program product provided in this specification is as follows: Corresponding to the other agent-based image processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide another computer program product.

[0192] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Obtain user commands input into the intelligent agent image component and submit them to the server; The server retrieves the editing data of the image block corresponding to the user instruction in the created layer, and displays the edited data synchronously. Obtain the editing instructions input by the user and submit them to the server; The system receives and synchronously displays the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and also receives and synchronously displays the adaptation results returned by the server for the position adaptation of the element image blocks.

[0193] It should be noted that the embodiment of another computer program product in this specification and the embodiment of another image processing method based on intelligent agents in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0194] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment and computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. When reading the relevant content of the device embodiment, equipment embodiment and computer-readable storage medium embodiment, please refer to the description of the method embodiment.

[0195] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps, and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims. This specification uses specific terms to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0196] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0197] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0198] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0199] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0200] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0201] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0202] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0203] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0204] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0205] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0206] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0207] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0208] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising at least one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0209] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0210] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.

Claims

1. An agent-based image processing method, comprising: Based on the image elements corresponding to the user commands input by the user in the intelligent agent image component, extract the element image blocks of the image elements; Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions; The image filling process is performed on the element extraction region in the image from which the image elements are extracted, and the position adaptation process of the element image block is performed according to the editing position of the element image block.

2. The image processing method based on intelligent agents according to claim 1, wherein the image elements corresponding to the user instruction are determined in the following manner: The image corresponding to the user instruction is segmented to obtain each image element, and the image element corresponding to the image position carried by the user instruction is determined as the target element. or, Image semantic recognition is performed on the image corresponding to the user instruction to obtain each image semantic element, and the element keywords carried by the user instruction are matched with each image semantic element to obtain the target element.

3. The image processing method based on intelligent agents according to claim 1, wherein the step of editing the element image block in the layer according to the editing instruction includes: Based on the position or trajectory information carried by the displacement command, the element image block is subjected to displacement processing and displacement rendering in the layer, and the displacement rendering data is synchronized to the client for displacement synchronization display. or, The displacement position of the element image block is determined according to the displacement keyword carried by the displacement command, and the displacement processing of the element image block is performed on the layer according to the displacement position. The displacement processing result is then synchronized to the client.

4. The image processing method based on intelligent agents according to claim 1, wherein the image filling process for the element extraction region in the image from which the image elements are extracted includes: If the real-time position of the element image block and the displacement data of the element extraction area meet the preset displacement filling conditions, the image after extracting the image element is input into the image filling model for image area filling processing to obtain a filled image, and then synchronized to the client for image editing and synchronous display.

5. The image processing method based on intelligent agents according to claim 4, wherein the image region filling process includes: Identify associated elements and / or associated image regions in the image after extracting the image elements that are associated with the image blocks of the elements; Image association analysis is performed on the element image block and the associated element and / or associated image region to obtain image association features; The element extraction region is filled with pixels according to the pixel filling parameters corresponding to the image association features.

6. The image processing method based on intelligent agents according to claim 5, wherein the image region filling process further includes: The pixel adjustment parameters of the associated elements and / or the associated image regions are determined based on the image association features; The filled image is obtained by adjusting the pixels of the associated elements and / or the associated image regions according to the pixel adjustment parameters.

7. The image processing method based on an intelligent agent according to claim 1, wherein the step of performing position adaptation processing of the element image block according to the edit position of the element image block includes: The element image blocks are merged into the image corresponding to the edit position to obtain a merged image; Spatial relationship recognition is performed between image elements and element image blocks in the merged image, and image adjustment processing is performed on element image blocks in the merged image based on the obtained image space; or, The image fusion dimension is determined based on the fusion keywords carried by the user instruction or the editing instruction; The element image blocks are merged into the image corresponding to the editing position to obtain a merged image, and the element image blocks and / or image elements in the merged image are adjusted according to the image fusion parameters corresponding to the image fusion dimension to obtain the target image.

8. The image processing method based on intelligent agents according to claim 1, wherein the step of performing position adaptation processing of the element image block according to the edit position of the element image block includes: Based on the displacement position, the displacement trajectory of the element image block is predicted to obtain the predicted trajectory, and the positional relationship between the predicted trajectory and the corresponding image is determined. The editing type of the element image block is determined based on the positional relationship, and the image editing adaptation of the element image block and the corresponding image is performed according to the editing type.

9. The image processing method based on intelligent agents according to claim 8, wherein the step of performing image editing adaptation of the element image block and the corresponding image according to the editing type includes: If the editing type is deletion, element deletion rendering data is generated based on the deletion rendering parameters corresponding to the deletion type and the element image block, and the element deletion rendering data is synchronized to the client for deletion synchronization display.

10. The image processing method based on an intelligent agent according to claim 8, wherein the step of performing the editing operation on the element image block according to the editing instruction on the layer is executed after the editing operation is executed, and the step of performing the image filling operation on the element extraction region in the image from which the image elements are extracted is executed before the editing operation is executed, further comprising: Determine whether the image from which the element image block is extracted is the same image as the image corresponding to the edit position; If so, perform the image filling operation on the element extraction region in the image from which the image elements are extracted; If not, perform a fusion process between the element image block and the image corresponding to the edit position based on the edit position of the element image block.

11. The image processing method based on an intelligent agent according to claim 10, wherein the step of fusing the element image block with the image corresponding to the edit position based on the edit position of the element image block includes: The image block adjustment parameters are determined according to the fusion keywords carried by the user instruction or the editing instruction, and the element image block is adjusted according to the image block adjustment parameters; The adjusted element image blocks are merged into the image corresponding to the edited position to obtain the target image; or, The adjusted element image blocks are merged into the image corresponding to the edited position to obtain the merged image; The merged image is input into a large language model for image semantic understanding and image element adjustment to obtain the target image.

12. The image processing method based on an intelligent agent according to claim 1, wherein the image of the extracted image element, and / or the image corresponding to the edit position, is generated by the intelligent agent; Alternatively, the element image block may include element images generated by the agent.

13. The agent-based image processing method according to claim 12, wherein the agent image component includes an image editing tool integrated by the agent; The user instructions and / or the editing instructions include text instructions or semantic instructions input into the agent image component or the agent's dialogue interface, or action instructions input into the agent image component.

14. An agent-based image processing method, comprising: Obtain user commands input into the intelligent agent image component and submit them to the server; The server retrieves the editing data of the image block corresponding to the user instruction in the created layer, and displays the edited data synchronously. Obtain the editing instructions input by the user and submit them to the server; The system receives and synchronously displays the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and also receives and synchronously displays the adaptation results returned by the server for the position adaptation of the element image blocks.

15. The image processing method based on intelligent agents according to claim 14, wherein receiving the image filling processing result of the extracted element region in the image from which the image elements are extracted, returned by the server, and synchronously displaying it, includes: Receive the filled image returned by the server, and update the display of the image corresponding to the user command based on the filled image; The step of receiving and synchronously displaying the adaptation results returned by the server for the position adaptation of the element image block includes: Receive the target image returned by the server, update the image corresponding to the editing instruction based on the target image, and display it.

16. An agent-based image processing apparatus, comprising: The image block extraction module is configured to extract element image blocks of the image elements corresponding to the user commands input by the user in the intelligent agent image component; The image block editing module is configured to create a layer of the element image block and perform editing processing on the element image block in the layer according to editing instructions; The position adaptation module is configured to perform image filling processing on the element extraction region in the image from which the image elements are extracted, and to perform position adaptation processing on the element image block according to the editing position of the element image block.

17. An agent-based image processing apparatus, comprising: The instruction acquisition module is configured to acquire user instructions input by the user in the intelligent agent image component and submit them to the server. The editing and display module is configured to obtain the editing processing data of the image block of the image element corresponding to the user instruction in the created layer from the server and display the edited data synchronously. The instruction submission module is configured to obtain the editing instructions input by the user and submit them to the server; The result display module is configured to receive and synchronously display the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and to receive and synchronously display the adaptation results returned by the server for the position adaptation of the element image blocks.

18. An agent-based image processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Based on the image elements corresponding to the user commands input by the user in the intelligent agent image component, extract the element image blocks of the image elements; Create a layer for the element image block, and perform editing processing on the element image block in the layer according to the editing instructions; The image filling process is performed on the element extraction region in the image from which the image elements are extracted, and the position adaptation process of the element image block is performed according to the editing position of the element image block.

19. An agent-based image processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Obtain user commands input into the intelligent agent image component and submit them to the server; The server retrieves the editing data of the image block corresponding to the user instruction in the created layer, and displays the edited data synchronously. Obtain the editing instructions input by the user and submit them to the server; The system receives and synchronously displays the image filling results returned by the server for the element extraction regions in the image from which the image elements are extracted, and also receives and synchronously displays the adaptation results returned by the server for the position adaptation of the element image blocks.

20. A computer-readable storage medium for storing computer-executable instructions that, when executed, implement the steps of the method of claim 1 or claim 14.