Image processing method and apparatus, device, and storage medium
Through user interaction with digital assistants and machine learning models, precise processing of local image regions is achieved, solving the problems of uncontrollable and inefficient generation results in traditional image processing techniques, and improving the accuracy and efficiency of image processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2026-04-02
AI Technical Summary
Traditional image processing techniques suffer from difficulties in controlling the generated results and low efficiency in local erasure and local re-editing.
By interacting with the digital assistant, the first region in the image and adjustment information are determined. Then, a machine learning model is used to generate an adjusted second image, achieving precise processing of local areas.
It improves the accuracy and efficiency of image processing, and can precisely adjust the visual effect of local areas according to user needs.
Smart Images

Figure CN2025094389_02042026_PF_FP_ABST
Abstract
Description
Method, device, equipment and storage medium for image processing
[0001] The present application claims priority to the Chinese patent application No. 202411377799.6, filed on September 29, 2024, entitled “Method, device, equipment and storage medium for image processing”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, equipment and computer readable storage medium for image processing. BACKGROUND
[0003] In the field of computer vision (CV), various image processing techniques have been significantly developed and have wide applications. For example, it is desirable to generate and use target images in many application scenarios such as social, game, image editing, etc. For example, machine learning based image processing techniques can be used in such application scenarios to improve the experience of users. SUMMARY
[0004] In a first aspect of the present disclosure, a method for image processing is provided. The method comprises: presenting a first image in an interaction of a user with a digital assistant; determining a first region in the first image and adjustment information for the first region based on an interaction operation from the user, the adjustment information indicating an adjustment of a visual effect on the first region; and presenting a second image as an adjustment result of the first image in response to receiving an adjustment indication for the first region, a second region corresponding to the first region in the second image having the adjusted visual effect.
[0005] In a second aspect of the present disclosure, a device for image processing is provided. The device comprises: a presentation module configured to present a first image in an interaction of a user with a digital assistant; a receiving module configured to determine a first region in the first image and adjustment information for the first region based on an interaction operation from the user, the adjustment information indicating an adjustment of a visual effect on the first region; and a processing module configured to present a second image as an adjustment result of the first image in response to receiving an adjustment indication for the first region, a second region corresponding to the first region in the second image having the adjusted visual effect.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure can be implemented;
[0012] FIGS. 2A to 2G show schematic diagrams of example interfaces for local regeneration according to some embodiments of the present disclosure;
[0013] FIGS. 3A to 3E show schematic diagrams of example interfaces for local update according to some embodiments of the present disclosure;
[0014] FIGS. 4A to 4C show schematic diagrams of further example interfaces for local update according to some embodiments of the present disclosure;
[0015] FIG. 5 shows a flowchart of an example process of image processing according to some embodiments of the present disclosure;
[0016] FIG. 6 shows a schematic structural block diagram of an example apparatus for image processing according to some embodiments of the present disclosure; and
[0017] FIG. 7 shows a block diagram of an electronic device capable of implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the range of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0019] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.
[0020] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in a text manner in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0021] It can be understood that the above notification and acquisition of user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0022] It can be understood that the data (including but not limited to the data itself, acquisition or use of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, on the contrary, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0024] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.
[0025] In this document, unless explicitly stated, performing a step “in response to A” does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0026] In the description of embodiments of the disclosure, the term "includes" and its conjugations are to be understood as open-ended terms that mean "comprises but not limited to." The term "based on" is to be understood as "based, at least in part, on." The term "one embodiment" or "the embodiment" is to be understood as "at least one embodiment." The term "some embodiments" is to be understood as "at least some embodiments." Other explicit or implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.
[0027] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. In this article, "model" can also be referred to as "machine learning model", "machine learning network" or "network", which are used interchangeably in this article. A model can also include different types of processing units or networks.
[0028] As briefly mentioned above, image processing technology has been applied to a variety of application scenarios. Traditional image processing technology has certain defects in local erasing and local re-editing, resulting in problems such as difficulty in controlling the generated results, low image processing efficiency, etc. in the process of image processing.
[0029] Embodiments of the present disclosure propose a scheme for image processing. According to embodiments of the present disclosure, in the interaction between a user and a digital assistant, a first image is presented. Based on an interaction operation from the user, a first region in the first image and adjustment information for the first region are determined, and the adjustment information indicates an adjustment of a visual effect of the first region. If an adjustment indication for the first region is received, a second image is obtained and presented as an adjustment result of the first image. In the second image, a second region corresponding to the first region has an adjusted visual effect. According to various embodiments of the present disclosure, in the interaction between a user and a digital assistant, a local region of an image can be adjusted based on the interaction operation of the user. In this way, the local region of the image can be accurately processed, and the image processing efficiency can be improved.
[0030] Various example implementations of the scheme are described in detail below in further conjunction with the accompanying drawings.
[0031] Example environment
[0032] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, a terminal device 110 has an application 120 installed therein. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.
[0033] In some embodiments, the application 120 can be any suitable application, e.g., an application that can provide a query service. In the environment 100 of FIG. 1, the terminal device 110 can render a page 150 of the application 120 if the application 120 is in an active state. The page 150 can include various types of pages that the application 120 can provide, such as an information interaction page, a query page, a search page, a search result rendering page, etc.
[0034] In some embodiments, the terminal device 110 communicates with a server 130 to implement the provision of a service of the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.). The server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc.
[0035] It should be understood that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0036] Some example embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. It should be understood that the pages shown in the accompanying drawings are merely examples and various page designs can actually exist. Various graphical elements in the pages can have different arrangements and different visual representations, one or more elements among them can be omitted or replaced, and one or more other elements can also exist. Embodiments of the present disclosure are not limited in this respect. Furthermore, in the following, example embodiments will be mainly described with respect to the terminal device 110. It should be understood that the actions described with respect to the terminal device 110 can be performed by the application 120 on the terminal device 110 or can be performed by the application 120 in coordination with its server (e.g., the server 130).
[0037] Example interactions
[0038] In the interaction between the user 140 and the digital assistant, the user 140 can be helped to achieve the adjustment of the local region of the image, so as to obtain the image as expected by the user. In the interaction between the user 140 and the digital assistant, the terminal device 110 can present a first image. The first image can be regarded as an initial image. In some embodiments, the first image can be an image specified by the user 140 in any suitable manner, for example, an image uploaded by the user 140. In some embodiments, the first image can be an image generated by using a machine learning model, for example, generated according to user input during the interaction between the user 140 and the digital assistant. The terminal device 110 can determine a first region in the first image and adjustment information for the first region based on the interactive operation from the user 140. The adjustment information can indicate the adjustment of the visual effect of the first region. That is, based on the interactive operation of the user 140, the local region to be adjusted in the first image and the information about how to adjust the region can be determined. In embodiments of the present disclosure, the visual effect can include a visual element or visual content (for example, a mountain, a tree, a person, and the like) in the image, and can also include an attribute (for example, color, brightness, chroma, definition, and the like) of the visual element or visual content.
[0039] In some embodiments, the user 140 can be provided with interface elements to facilitate the user 140 to autonomously select a region and give adjustment information. For example, if a local adjustment request for the first image is received, the terminal device 110 can present a region selection control and a prompt input portal, as will be described below with reference to FIGS. 2A to 2G. The terminal device 110 can receive, via the region selection control, a selection of a first region in the first image. The terminal device 110 can receive, via the prompt input portal, user input including adjustment information. Then, if an adjustment instruction for the first region is received, the terminal device 110 can present a second image as a result of the adjustment of the first image, a second region corresponding to the first region in the second image having an adjusted visual effect. For example, the terminal device 110 can send the adjustment information, information of the first image, and information of the first region to a server, and the server can update the first image to the second image by using a machine learning model.
[0040] In embodiments of the present disclosure, the adjustment on the local region can include various appropriate forms of adjustment. In some embodiments, the adjustment on the visual effect of the first region can include local redrawing, and the second region can include visual content generated based on the adjustment information. In other words, the local adjustment can indicate the regeneration of the local image. One example is described below with reference to FIGS. 2A-2G. FIGS. 2A-2G show schematic diagrams of example interfaces 200A-200G for local regeneration, according to some embodiments of the present disclosure. The interfaces 200A-200G can be provided by the terminal device 110 shown in FIG. 1, for example.
[0041] As shown in the interface 200A, the terminal device 110 can present the first image 210 in the interface. The interface 200A can also include a “modify” control 215, a “download” control 220, a “share” control 225, and an “add as work” control 230, for example. In some examples, the terminal device 110 can save the first image 210 to the local or a target address if it detects that the user 140 clicks the “download” control 220. In some examples, the terminal device 110 can share the first image 210 to a target session or a target address if it detects that the user 140 clicks the “share” control 225. In some examples, the terminal device 110 can save the first image 210 to a work collection if it detects that the user 140 clicks the “add as work” control 230. It can be appreciated that the controls shown in FIG. 2A are merely exemplary and are not limiting.
[0042] If the terminal device 110 detects that the user 140 clicks the “modify” control 215, it can present the interface 200B as shown in FIG. 2B. The terminal device 110 presents a panel 240 for selecting a processing operation after receiving the user 140 clicking the “modify” control 215. In the example of FIG. 2B, the panel 240 can include options for multiple processing operations, which can include high-definition, expansion, local redrawing, and local erasing, for example. In some examples, the terminal device 110 can present the interface 200C if it detects that the user 140 clicks any position in the region corresponding to the local redrawing in the panel 240.
[0043] As shown in interface 200C of FIG. 2C, terminal device 110 presents control 250, interface element 255, and prompt input entry 260 upon detecting that user 140 clicks any position in the region corresponding to the partial redrawing in panel 240. In an example, user 140 can adjust the size of the region selection control by sliding the circular region in control 250. Interface element 255 can be used to show the numerical size of the region selection control, and can slide together with the circular region in control 250. Terminal device 110 can receive, via the region selection control, a selection of the first region in the first image. In some examples, the region selection control can be presented in any reasonable graphical form, such as a circle, an arrow, etc. Terminal device 110 can receive, via prompt input entry 260, a user input including adjustment information. In some examples, prompt input entry 260 can include a corresponding prompt word (e.g., "describe the redrawing content or directly generate…") and "send" control 265. Terminal device 110 can present interface 200D if it receives, via the region selection control, a selection of the first region in the first image by user 140.
[0044] The above-described manner of determining the first region (i.e., determining the region to be adjusted) is merely exemplary and is not intended to be limiting. In embodiments of the present disclosure, the first region can be determined in various suitable manners. In some embodiments, terminal device 110 can receive a selection of an initial region in the first image and description information for adjusting the first image. For example, the initial region can be a region corresponding to a selection of the first image received by terminal device 110 for the first time. For example, the description information can be "retain the buildings and the sky in the first image." For another example, the description information can be "retain the person in the first image." Then, terminal device 110 can update the initial region to the first region based on the description information. For example, terminal device 110 can expand, shrink, move, etc. the initial region based on the description information to determine the first region. In some examples, assuming the description information indicates "retain the buildings and the grassland," terminal device 110 can expand the initial region to determine, in the first image, the region other than the buildings and the grassland as the first region based on the description information. In such embodiments, the first region is determined by combining the selection operation and the description information.
[0045] In some embodiments, the first region can be determined based on at least one of a natural language input indicating a region to be adjusted, a predetermined type of operation on a region in the first image, or an interactive control for selecting a region to be adjusted. The natural language input can include a text input or a voice input from the user, and can describe a feature of a region to be adjusted in the first image. The predetermined type of operation on a region in the first image can be provided for the user to specify a region to be adjusted. In some examples, the predetermined type of operation on a region in the first image can include a smearing operation, a circling operation, etc. The interactive control for selecting a region to be adjusted can provide the user with a suitable region selection control, such as a marquee control, etc.
[0046] As shown in interface 200D of FIG. 2D, the first region 270 of the first image 210 selected by the user 140 is highlighted in a specific visual style to visually distinguish from the regions that are not selected.
[0047] As shown in interface 200E of FIG. 2E, the terminal device 110 presents a prompt input entry 275 containing the user input if it receives the user input of the user 140 at the prompt input entry 260, which can include the adjustment information. In some examples, the input content of the user input can be “replace the selected region with sky”. In some examples, the input content of the user input can be “partial redraw: replace the selected region with a green grassland”. The terminal device 110 can provide the user input to the digital assistant and present interface 200F if it detects a click of the “send” control 265 by the user 140.
[0048] As shown in interface 200F of FIG. 2F, the terminal device 110 can present a user input display box 280 and a picture processing display interface 285 if it provides the user input to the digital assistant. The user input display box 280 can include a reference picture of the first image and the input content of the user input. The picture processing display interface 285 can include a “picture generating” icon 286 and a generation degree prompt word 287 (e.g., “picture generating 78%”). The interface 200F further includes a “regenerate” region 288, and the terminal device 110 can perform the image generation operation again if it receives a click of the user 140 at any position in the “regenerate” region 288.
[0049] In some embodiments, the first image is determined based on an interaction message issued by the user to the digital assistant. For example, the interaction message can be "draw an image with a little girl walking on a street in the rain, and the little girl's right hand is holding a white puppy." In such embodiments, the terminal device 110 can determine the adjustment target for the first region based on the adjustment information and the interaction message previously used to generate the first image. Then, the terminal device 110 can update the first image to the second image according to the adjustment target. In some examples, the adjustment information can include "replace the puppy in the image with a kitten." Accordingly, in the second image, the puppy is replaced with a kitten.
[0050] As shown in interface 200G of FIG. 2G, the terminal device 110 presents the second image 290 as the result of the adjustment to the first image 210 after detecting that the image generation is completed. The second region 295 corresponding to the first region 270 in the second image 290 has an adjusted visual effect. In some examples, if the input content input by the user is "replace the selected region with a big tree", the visual content presented by the second region 295 includes a big tree. In some examples, if the input content input by the user is "replace the selected region with a starry sky", the visual content presented by the second region 295 includes a starry sky.
[0051] In some embodiments, the adjustment to the visual effect of the first region can include a partial effect update. The partial effect update can also be referred to as partial modification or partial erasure. The partial effect update can include an update to part of the content in the partial region, an update to the attributes of the presented content, and the like.
[0052] In some embodiments, the second region can include visual content that is partially the same as the first region. That is, the second region obtained by the partial effect update to the first region can include visual content that is partially the same (the part that is not subject to the partial effect update) and partially different (the part that is subject to the partial effect update) as the first region. An example is described below with reference to FIGS. 3A-3E. FIGS. 3A-3E show schematic diagrams of example interfaces 300A-300E for partial update, according to some embodiments of the present disclosure. The interfaces 300A-300E can be provided by the terminal device 110 shown in FIG. 1, for example.
[0053] As shown in the example interfaces 200A and 300A of FIG. 2A and FIG. 3A, upon the user 140 tapping the "modify" control 215 in the interface 200A, the terminal device 110 can present the interface 300A. As shown in FIG. 3A, the interface 300A presents an operation editing region 310. In some examples, the operation editing region 310 can include the first image 210, a processing operation selection region, and a prompt input portal. In some examples, the processing operation selection region can include a "high definition" control, an "expand" control, a "partial redraw" control, and a "partial erase" control. In some examples, the prompt input portal can be presented with a prompt word, such as "describe how you want to edit the picture." Upon detecting the user 140 tapping any location of the "partial erase" control in the processing operation selection region, the terminal device 110 can present the interface 300B.
[0054] As shown in the interface 300B of FIG. 3B, the terminal device 110 can receive a first region 320 of the first image 210 selected by the user 140 via the region selection control. In the interface 300B, the first region 320 selected by the user 140 is highlighted in a specific visual style to visually distinguish from the unselected regions. The interface 300B can further include an operation prompt box 325. In some examples, the operation prompt box 325 can include a thumbnail of the first image and an editing region 326 (e.g., "partial erase"). In addition, the terminal device 110 can receive user input via the editing region 326 in the operation prompt box 325. In some examples, the terminal device 110 can display the name of the selected processing operation, such as "partial erase," in the editing region 326 in the absence of receiving user input.
[0055] As shown in the interface 300C of FIG. 3C, upon receiving user input, the terminal device 110 can present an operation prompt box 330. In some examples, the editing region 335 of the operation prompt box 330 can present the input content of the user input, such as "erase XXX but not YYY." Upon detecting the user 140 tapping the "send" control 335 in the editing region 326, the terminal device 110 can present the interface 300D.
[0056] As shown in the interface 300D of FIG. 3D, in the interactive interface with the digital assistant, upon providing the operation prompt box containing the user input to the digital assistant, the terminal device 110 can present a picture processing demonstration interface 340 in the interactive interface. In some examples, the picture processing demonstration interface 340 can include a "picture generating" icon 386 and a generation progress prompt word 387 (e.g., "picture generating 75%").
[0057] As shown in interface 300E of FIG. 3E, upon detecting completion of the image generation, the terminal device 110 presents a second image 350 as the result of the adjustment to the first image 210. Illustratively, the second image 350 has an adjusted visual effect in a second region 355 corresponding to the first region 320. Continuing the example above, if the user input is “erase XXX but do not erase YYY”, then the second region 355 can include the visual content of YYY but no longer include the visual content of XXX.
[0058] In some embodiments, the second region can have at least one image property of a different dimension than the first region. In some examples, such image property can be any suitable property, such as but not limited to one of color, transparency, brightness, sharpness, etc. One example is described below with reference to FIGS. 4A-4C. FIGS. 4A-4C show schematic diagrams of further example interfaces 400A-400C for partial update, in accordance with some embodiments of the present disclosure. The interfaces 400A-400C can be provided, for example, by the terminal device 110 shown in FIG. 1.
[0059] As shown in example interfaces 300B and 400A of FIGS. 3B and 4A, upon detecting the user input of the user 140 in the edit region 326 of the interface 300B, the terminal device 110 can present the interface 400A.
[0060] As shown in interface 400A of FIG. 4A, the terminal device 110 can present an operation prompt box 430. Illustratively, the operation prompt box 430 can include a thumbnail of the first image and an edit region containing the user input of “lightly erase but do not change the original”. Further, upon detecting a click of the “send” control 335 of the edit region 435 by the user 140, the terminal device 110 presents the interface 400B.
[0061] As shown in FIG. 4B, upon providing the operation prompt box containing the user input to the digital assistant in the interaction interface 400B with the digital assistant, the terminal device 110 can present a picture processing demonstration interface 445 in the interaction interface.
[0062] As shown in interface 400C of FIG. 4C, upon detecting completion of the image generation, the terminal device 110 can present a second image 440 as the result of the adjustment to the first image 210. Illustratively, the second image 440 has an adjusted visual effect in a second region 455 corresponding to the first region 320. Continuing the example above, in the case where the user input is “lightly erase but do not change the original”, the visual content in the second region 455 itself remains unchanged relative to the first region, but the color is faded. It should be appreciated that this is merely illustrative, and other than color, the brightness, chroma, sharpness, etc. of the region can also be changed, for example.
[0063] Some example embodiments of the present disclosure will be further described below. In some embodiments described above, the local region to be adjusted and the adjustment information are selected or provided by the user 140 through interface elements. Alternatively or additionally, in some embodiments, the local region to be adjusted and / or the adjustment information for the local region can be determined based on the interactive messages between the user and the digital assistant.
[0064] In some embodiments, the terminal device 110 can determine the first region based on one or more first interactive messages received from the user 140 to the digital assistant. For example, the one or more first interactive messages can indicate an image adjustment requirement. In some examples, the first interactive message can be "select all the clouds in the image", and the corresponding first region is composed of all the clouds in the first image. In some examples, the first interactive message can be "select the little girl in the image", and the corresponding first region is the area where the little girl in the first image is located. In more examples, the first interactive message can also be "I want to select the lower right corner of the image". In this way, the user 140 can express the selection requirement for the first region in the first image through natural language interaction with the digital assistant.
[0065] In some embodiments, the terminal device 110 can select a candidate region in the first image based on the element or location in the first image indicated by the one or more interactive messages. In some examples, the element in the first image can indicate a girl, a cup, a sun, etc. In some examples, the location in the first image can indicate the center, the lower right corner, the upper left corner, etc. The terminal device 110 can present the indication information about the candidate region. For example, the terminal device 110 can determine a candidate region according to the interactive messages between the user 140 and the digital assistant, and present the candidate region to the user 140. The terminal device 110 can determine the candidate region as the first region after receiving a positive indication for the candidate region. In some examples, the user 140 can further adjust the candidate region before confirming the candidate region.
[0066] In some embodiments, the terminal device 110 can receive one or more second interactive messages from the user 140 to the digital assistant, and the one or more second interactive messages describe image visual requirements. The terminal device 110 can determine the adjustment information based on the one or more second interactive messages. In some examples, the user 140 can express the image processing requirement to the digital assistant through voice input after confirming the first region. For example, the user 140 inputs the following content "modify the little boy in the image to a little girl" through voice input. In this way, the user 140 can express the image processing requirement through interaction with the digital assistant, rather than necessarily through the input box.
[0067] In some embodiments, the second image can be generated in the following manner. Illustratively, the terminal device 110 can generate a local image using a machine learning model based on the first region and the adjustment information. The terminal device 110 determines the second image based on the local image and a region of the first image other than the first region. Further, the terminal device 110 can take the second image as a processing result of the first image.
[0068] In this way, the embodiments of the present disclosure can accurately process a local part of the first image through user input, and can also improve the image processing efficiency.
[0069] Example processes, apparatuses, and devices
[0070] FIG. 5 illustrates a flowchart of an example process 500 of image processing according to some embodiments of the present disclosure. The process 500 can be implemented at the terminal device 110. The process 500 is described below with reference to FIG. 1.
[0071] As shown in FIG. 5, at block 510, the terminal device 110 presents a first image in a user interaction with a digital assistant.
[0072] At block 520, the terminal device 110 determines a first region in the first image and adjustment information for the first region based on an interaction operation from the user, the adjustment information indicating an adjustment of a visual effect of the first region.
[0073] At block 530, the terminal device 110 presents a second image as a result of the adjustment of the first image in response to receiving the adjustment indication for the first region, a second region corresponding to the first region in the second image having the adjusted visual effect.
[0074] In some embodiments, determining the first region and the adjustment information includes: presenting a region selection control and a prompt input portal in response to receiving a local adjustment request for the first image; receiving a selection of the first region in the first image via the region selection control; and receiving user input including the adjustment information via the prompt input portal.
[0075] In some embodiments, determining the first region includes: receiving a selection of an initial region in the first image and description information for adjusting the first image; and updating the initial region to the first region based on the description information.
[0076] In some embodiments, the first image is determined based on an interaction message issued by the user to the digital assistant, and the second image is generated in the following manner: determining an adjustment target for the first region based on the adjustment information and the interaction message; and updating the first image to the second image according to the adjustment target.
[0077] In some embodiments, the first region is determined based on at least one of: a natural language input indicating a region to be adjusted, a predetermined type of operation on a region in the first image, or an interactive control for selecting a region to be adjusted.
[0078] In some embodiments, determining the first region in the first image comprises: receiving one or more first interaction messages issued by the user to the digital assistant, the one or more first interaction messages indicating image adjustment requirements; and determining the first region based on the one or more first interaction messages.
[0079] In some embodiments, determining the first region based on the one or more first interaction messages comprises: selecting a candidate region in the first image based on an element or a location in the first image indicated by the one or more interaction messages; presenting indication information about the candidate region; and in response to receiving a positive indication for the candidate region, determining the candidate region as the first region.
[0080] In some embodiments, determining the adjustment information comprises: receiving one or more second interaction messages issued by the user to the digital assistant, the one or more second interaction messages describing image visual requirements; and determining the adjustment information based on the one or more second interaction messages.
[0081] In some embodiments, the adjustment of the visual effect of the first region comprises partial redrawing, and the second region comprises visual content generated based on the adjustment information.
[0082] In some embodiments, the adjustment of the visual effect of the first region comprises partial effect update, and the second region comprises visual content that is partially the same as the first region or has at least one dimension of image attribute that is different from the first region.
[0083] In some embodiments, the second image is generated by: generating a partial image based on the first region and the adjustment information using a machine learning model; and determining the second image based on the partial image and a region in the first image other than the first region.
[0084] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 6 shows a schematic structural block diagram of an example apparatus 600 for image processing according to certain embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0085] As shown in FIG. 6, the apparatus 600 includes a presentation module 610 configured to present a first image in an interaction of a user with a digital assistant; a determination module 620 configured to determine, based on an interaction operation from the user, a first region in the first image and adjustment information for the first region, the adjustment information indicating an adjustment of a visual effect on the first region; and a receiving module 630 configured to present a second image as an adjustment result of the first image in response to receiving an adjustment indication for the first region, a second region corresponding to the first region in the second image having the adjusted visual effect.
[0086] In some embodiments, the determination module 620 is further configured to present a region selection control and a prompt input portal in response to receiving a local adjustment request for the first image; receive, via the region selection control, a selection of the first region in the first image; and receive, via the prompt input portal, a user input including the adjustment information.
[0087] In some embodiments, the determination module 620 is further configured to receive a selection of an initial region in the first image and description information for adjusting the first image; and update the initial region to the first region based on the description information.
[0088] In some embodiments, the first image is determined based on an interaction message issued by the user to the digital assistant, and the second image is generated by: determining an adjustment target for the first region based on the adjustment information and the interaction message; and updating the first image to the second image according to the adjustment target.
[0089] In some embodiments, the first region is determined based on at least one of: a natural language input indicating a region to be adjusted, a predetermined type of operation on a region in the first image, or an interaction control for selecting a region to be adjusted.
[0090] In some embodiments, the determination module 620 is further configured to receive one or more first interaction messages issued by the user to the digital assistant, the one or more first interaction messages indicating an image adjustment requirement; and determine the first region based on the one or more first interaction messages.
[0091] In some embodiments, the determination module 620 is further configured to select a candidate region in the first image based on an element or a position in the first image indicated by the one or more interaction messages; present indication information about the candidate region; and determine the candidate region as the first region in response to receiving a positive indication for the candidate region.
[0092] In some embodiments, the determining module 620 is further configured to receive one or more second interaction messages issued by the user for the digital assistant, the one or more second interaction messages describing image visual requirements; and determine the adjustment information based on the one or more second interaction messages.
[0093] In some embodiments, the adjustment of the visual effect of the first region comprises a partial redraw, and the second region comprises visual content generated based on the adjustment information.
[0094] In some embodiments, the adjustment of the visual effect of the first region comprises a partial effect update, and the second region comprises visual content that is partially the same as the first region or has at least one dimension of image properties that is different from the first region.
[0095] In some embodiments, the second image is generated by: generating a partial image based on the first region and the adjustment information using a machine learning model; and determining the second image based on the partial image and a region of the first image other than the first region.
[0096] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can be used to implement the terminal device 110 of FIG. 1.
[0097] As shown in FIG. 7, the electronic device 700 is in the form of a general electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processing units or processors 710, memory 720, storage 730, one or more communication units 740, one or more input devices 770, and one or more output devices 760. The processor 710 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processors perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0098] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media that is locally and / or remotely accessible by the electronic device 700, including volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., random access memory (RAM), cache memory, etc.), a non-volatile memory (e.g., read-only memory (ROM), electrically programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.), or some combination thereof. The storage device 730 can be a removable storage and / or non-removable storage including, for example, a machine-readable medium that can be used to store information and / or data, which can be accessed by the electronic device 700.
[0099] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive for reading from and / or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from and / or writing to a removable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 727 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0100] The communication unit 740 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0101] The input device 770 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 740, as needed, one or more devices that enable a user to interact with the electronic device 700, or any device (e.g., a network card, a modem, etc.) that enables the electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0102] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0103] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0104] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk drive, or any other suitable non-transitory computer readable medium can store the computer program product.
[0107] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and deletions can be made to the disclosed implementations without materially departing from the novel teachings of the disclosure set forth herein. Accordingly, embodiments of the present disclosure shall not be limited by what has been particularly shown and described hereinabove. Rather, the scope of the present disclosure includes both combinations and sub-combinations of the various features described hereinabove, as well as variations and modifications thereof that would be apparent to one of ordinary skill in the art upon reading the foregoing description. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
Claims
1. A method for image processing, comprising: presenting a first image in an interaction between a user and a digital assistant; determining a first region in the first image and adjustment information for the first region based on an interaction operation from the user, the adjustment information indicating an adjustment of a visual effect on the first region; and in response to receiving an indication of the adjustment for the first region, presenting a second image as a result of the adjustment on the first image, a second region in the second image corresponding to the first region having the adjusted visual effect.
2. The method of claim 1, wherein determining the first region and the adjustment information comprises: in response to receiving a request for a local adjustment on the first image, presenting a region selection control and a prompt input portal; receiving, via the region selection control, a selection of the first region in the first image; and receiving, via the prompt input portal, a user input comprising at least a portion of the adjustment information.
3. The method of claim 1, wherein determining the first region comprises: receiving a selection of an initial region in the first image and description information for adjusting the first image; and updating the initial region to the first region based on the description information.
4. The method of claim 1, wherein the first image is determined based on an interaction message issued by the user to the digital assistant, and the second image is generated by: determining an adjustment target for the first region based on the adjustment information and the interaction message; and updating the first image to the second image according to the adjustment target.
5. The method of claim 1, wherein the first region is determined based on at least one of: a natural language input indicating a region to be adjusted, a predetermined type of operation on a region in the first image, or an interaction control for selecting a region to be adjusted.
6. The method of claim 1, wherein determining the first region in the first image comprises: receiving one or more first interaction messages issued by the user to the digital assistant, the one or more first interaction messages indicating an image adjustment requirement; and determining the first region based on the one or more first interaction messages.
7. The method of claim 6, wherein determining the first region based on the one or more first interaction messages comprises: selecting a candidate region in the first image based on an element or a location in the first image indicated by the one or more interaction messages; presenting indication information about the candidate region; and in response to receiving a positive indication for the candidate region, determining the candidate region as the first region.
8. The method of claim 1, wherein determining the adjustment information comprises: receiving one or more second interaction messages issued by the user to the digital assistant, the one or more second interaction messages describing an image visual requirement; and determining the adjustment information based on the one or more second interaction messages. 9. The method of claim 1, wherein the adjustment of the visual effect of the first region comprises partial redrawing, and the second region comprises visual content generated based on the adjustment information.
10. The method of claim 1, wherein the adjustment of the visual effect of the first region comprises partial effect update, and the second region comprises visual content that is partially same as the first region, or has image properties that are different from the first region in at least one dimension.
11. The method of claim 1, wherein the second image is generated by: generating, based on the first region and the adjustment information, a partial image using a machine learning model; and determining the second image based on the partial image and a region of the first image other than the first region.
12. An apparatus for image processing, comprising: a presentation module configured to present a first image in an interaction of a user with a digital assistant; a receiving module configured to determine, based on an interaction operation from the user, a first region in the first image and adjustment information for the first region, the adjustment information indicating an adjustment of a visual effect of the first region; and a processing module configured to present, in response to receiving the adjustment indication for the first region, a second image as a result of the adjustment of the first image, the second image having an adjusted visual effect in a second region corresponding to the first region.
13. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method of any of claims 1-11.
14. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method of any of claims 1-11.
15. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any of claims 1-11.
Citation Information
Patent Citations
Image processing method, model training method and related device
CN114943789A
Image processing method and device, equipment and storage medium
CN115100359A
Image generation method and device, electronic equipment and computer readable storage medium
CN116580127A
Image processing method and device, equipment and storage medium
CN118172236A
Image processing method and device, equipment and storage medium
CN119415002A