Image editing method and apparatus, and device, medium and program product

By detecting image editing operations and generating recommended instruction sets, and leveraging computer vision and natural language processing technologies, we can simplify user interactions in image editing applications, solving the complex interaction issues in existing technologies and achieving an efficient and convenient image editing experience.

WO2025209146A1PCT designated stage Publication Date: 2025-10-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/082415
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-03-13
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing image editing applications have a high interaction threshold, and users need to manually select editing operations, resulting in a poor experience.

Method used

By detecting interactive operations on the image to be edited, a recommended instruction set is generated, including encapsulated instructions for image editing operations, simplifying the user interaction process. Computer vision and natural language processing technologies are used to predict user intentions and provide personalized image editing suggestions.

Benefits of technology

It lowers the threshold for using image editing, improves the convenience of interaction and the efficiency of image editing, and generates high-quality editing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082415_09102025_PF_FP_ABST
    Figure CN2025082415_09102025_PF_FP_ABST
Patent Text Reader

Abstract

On the basis of the embodiments of the present disclosure, provided are an image editing method and apparatus, and a device, a medium and a program product. The method comprises: detecting an interactive operation executed on a target image to be edited; on the basis of the interactive operation, determining a recommended instruction set for interacting with a target model, wherein the recommended instruction set comprises at least one instruction, each instruction being defined as corresponding to at least one image editing operation; presenting the recommended instruction set to a user; and in response to detecting the selection of the user in respect of a first instruction in the recommended instruction set, providing the first instruction to the target model, so that the target model determines, on the basis of the first instruction, at least one image editing operation to be applied to the target image. Thus, the image editing efficiency can be improved, and edited high-quality images can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device, medium and program product for image editing

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment, media and program products for image editing” and application number 202410390208.2, filed on April 1, 2024, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, computer-readable storage media, and computer program products for image editing. Background Art

[0003] With the development of Internet technology, people have an increasing demand for expressing and sharing content on Internet platforms, especially for displaying images. Therefore, related image editing applications have emerged to assist people in editing images. Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for image editing is provided. The method includes: detecting an interactive operation performed on a target image to be edited; determining a recommended instruction set for interacting with a target model based on the interactive operation, the recommended instruction set including at least one instruction, each instruction defined as corresponding to at least one image editing operation; causing the recommended instruction set to be presented to a user; and in response to detecting a user selection of a first instruction in the recommended instruction set, providing the first instruction to the target model, so that the target model determines, based on the first instruction, at least one image editing operation to be applied to the target image.

[0005] In a second aspect of the present disclosure, a device for image editing is provided. The device includes: a detection module configured to detect an interactive operation performed on a target image to be edited; a determination module configured to determine a recommended instruction set for interacting with a target model based on the interactive operation, the recommended instruction set including at least one instruction, each instruction being defined as corresponding to at least one image editing operation; a presentation module configured to cause the recommended instruction set to be presented to a user; and a providing module configured to, in response to detecting a user selecting a first instruction in the recommended instruction set, provide the first instruction to the target model, so that the target model determines, based on the first instruction, at least one image editing operation to be applied to the target image.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, cause the device to perform the method of the first aspect.

[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a flowchart of a process for image editing according to some embodiments of the present disclosure;

[0013] FIG3 shows an example block diagram of an architecture 300 for image editing according to some embodiments of the present disclosure;

[0014] FIG4 shows a flowchart of a process of image editing according to some embodiments of the present disclosure;

[0015] FIG5 shows a block diagram of an apparatus for image editing according to some embodiments of the present disclosure; and

[0016] FIG6 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0019] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably in this article. A model can also include different types of processing units or networks.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] Most current image editing applications usually require users to manually select editing operations, and the interaction threshold is generally high, which affects the user experience. People expect image editing applications to be able to edit images more conveniently and quickly, and obtain high-quality images.

[0028] FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an image editing application 120 is installed on a terminal device 110. In some embodiments, image editing application 120 may be an application for analyzing, repairing, beautifying, compositing, or performing other processing on images. In some embodiments, image editing application 120 may also be any other suitable application capable of editing images.

[0029] In some embodiments, user 140 can interact with image editing application 120 via terminal device 110 and / or a device attached to terminal device 110. In some examples, user 140 can transmit a target image 145 to be edited to image editing application 120 via terminal device 110. Image editing application 120 edits target image 145 to generate an edited image 150. Edited image 150 can be saved on terminal device 110 or the user can select to perform other operations. User 140 can directly use the saved edited image 150 on terminal device 110 or transmit the edited image 150 to other devices / systems via terminal device 110.

[0030] In some embodiments, the terminal device 110 communicates with the server device 130 to provide services for the image editing application 120. In the environment 100, the target model 135 communicates with the server device 130. In some examples, the server device 130 can provide services for the image editing application 120 installed in the terminal device 110 by calling the target model 135. The target model 135 can run on a device / system other than the server device 130 and the terminal device 110. In some embodiments of the present disclosure, the target model 135 can parse content such as text or instructions to determine one or more image editing operations to be applied to the target image 145. Although only a single target model 135 is shown in FIG1 , it is understood that, depending on the specific application needs, there may be more models, and different models may be configured to determine different image editing operations in different ways, or to provide other functions in the image editing application 120, etc.

[0031] In some embodiments, the target model 135 may have content generation capabilities. In some embodiments, the target model 135 may include a language model, such as a large language model. The target model 135 may be pre-trained from a large amount of data so that it can understand the semantic information of text modalities and other modalities. A suitable machine learning model may be selected based on actual needs. In some embodiments, based on the target model 135, an interactive interface with dialogue capabilities may be provided, in which user input may be received, and the semantics of the user input may be understood with the help of the target model 135, and a response to the user input may be determined. In an image editing scenario, the target model 135 may be used to understand the user's text input related to image editing, and the image editing operation that the user desires to perform may be determined from the text input, thereby realizing conversational image editing or conversational image beautification.

[0032] The target model 135 can run locally on the terminal device 110 or the server device 130 or be deployed on a remote device, a cloud environment, etc. In the case of local operation, the terminal device 110 or the server device 130 can directly provide the model input to the locally installed target model 135 and obtain the model output generated by the target model 135. In the case of remote operation, the terminal device 110 or the server device 130 provides data to other devices through a communication connection with other devices. The other devices determine the model input based on the acquired data and provide the model input to the target model 135. After obtaining the model output of the target model 135, the other electronic devices provide the model output through a communication connection with the terminal device 110 or the server device 130.

[0033] In some embodiments, the terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server device 130 can be various types of computing systems / servers that can provide computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0034] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0035] By introducing a model with conversational capabilities, we can support conversational image editing applications. While this approach simplifies the traditional manual image editing process, this text-only approach still has a high barrier to entry. This is because users often have diverse language preferences, and translation software is often required to support foreign languages. Furthermore, users are required to provide very specific text descriptions to interact with the application, all of which make interaction difficult.

[0036] Therefore, people expect image editing applications to simplify the interaction process and lower the usage threshold, so that they can edit images more conveniently and quickly and obtain high-quality images.

[0037] According to an embodiment of the present disclosure, an improved scheme for image editing is proposed. According to the scheme of the embodiment of the present disclosure, an interactive operation performed on a target image to be edited is detected. Based on the interactive operation, a recommended instruction set for interacting with a target model is determined, the recommended instruction set including at least one instruction, each instruction being defined as corresponding to at least one image editing operation. The recommended instruction set is then presented to the user. In response to detecting a user's selection of a first instruction in the recommended instruction set, the first instruction is provided to the target model, so that the target model determines at least one image editing operation to be applied to the target image based on the first instruction. Thus, in this way, people can interact more simply and friendly when using image editing applications, effectively lowering the threshold for using the application, so that image editing can be performed more conveniently and quickly, and high-quality images can be obtained.

[0038] The following will continue to describe some example embodiments of the present disclosure with reference to the accompanying drawings. Figure 2 shows a flowchart of a process 200 for image editing according to some embodiments of the present disclosure. For ease of discussion, these embodiments will be described with reference to the environment 100 of Figure 1. These embodiments can be implemented in the server device 130 of Figure 1. In other embodiments, these embodiments can also be implemented locally on the client, that is, in the terminal device 110, or through the collaboration of the terminal device 110 and the server device 130. The following specific embodiments are implemented in the server device 130 as an example.

[0039] At block 210 , the server device 130 detects an interaction operation performed on the target image 145 to be edited. As will be described in detail below, in some embodiments, the interaction operation considered may include the user 140 selecting to upload the target image 145 and / or a subsequent user 140 modifying the target image 145.

[0040] FIG3 shows an example block diagram of an architecture 300 for image editing according to some embodiments of the present disclosure. The architecture 300 may be implemented in the environment 100 of FIG1 . The architecture 300 may be shown as the architecture employed by the steps of the process 200.

[0041] As shown in FIG3 , architecture 300 may include server device 130, terminal device 110, target model 135, and instruction phrase library 350. Instruction phrase library 350 may communicate with server device 130. Server device 130 may include instruction recommendation module 310 and text processing module 320. Terminal device 110 may include user interaction module 330 and local rendering module 340.

[0042] In some embodiments, the user interaction module 330 can be used for interaction between the user 140 and the terminal device 110. The user's selection of the target image 145 can be implemented in the user interaction module 330 of the terminal device 110. Specifically, through the user interaction module 330, the user 140 uploads the target image 145 to be edited in the image editing application 120. The user interaction module 330 can then upload the selected target image 145 to the server device 130. The instruction recommendation module 310 in the server device 130 can detect the target image 145 and perform analysis and processing.

[0043] Through the user interaction module 330, the user 140 can also retouch the target image in the image editing application 120. The user interaction module 330 can upload the updated image to the server device 130, and the instruction recommendation module 310 in the server device 130 can detect the updated image and perform analysis and processing.

[0044] Continuing with reference to Figure 2, in box 220, the server device 130 determines a recommended instruction set for interacting with the target model 135 based on the interactive operation. The recommended instruction set includes at least one instruction, and each instruction is defined as corresponding to at least one image editing operation. In some embodiments, the instruction may be a photo editing instruction phrase for indicating an image editing operation. The image editing operation may have atomic capabilities for image beautification, such as adjusting the brightness of an image, whitening the skin of a person in an image, and so on. By encapsulating the image editing operation in the form of an instruction for interacting with the target model 135 and recommending it to the user, the user's interaction complexity can be simplified, allowing the user to select the desired image editing operation more conveniently and quickly. In some embodiments, a combination of multiple image editing operations can be encapsulated into one instruction, which can further improve editing efficiency.

[0045] 3 , pre-configured instructions may be stored in an instruction phrase library 350. The instruction phrase library 350 may be run on the server device 130 or on another device / system different from the server device 130. In some embodiments, the instructions in the instruction phrase library 350 may be configured based on the image editing operations supported by the image editing application 120.

[0046] In some embodiments, in response to detecting that the interaction operation is for the user to select the target image 145 for editing, the server device 130 extracts visual features of the target image 145 and determines a recommended instruction set associated with the visual features.

[0047] As shown in FIG3 , after the user 140 uploads the target image 145 to be edited in the image editing application 120 , the instruction recommendation module 310 of the server device 130 can obtain the target image 145 via the user interaction module 330 of the terminal device 110 .

[0048] In some embodiments, the visual features of the target image 145 extracted by the server device 130 indicate at least one of the following: whether the target image contains a human portrait, whether the target image does not contain a human portrait, the number of human portraits contained in the target image, and the type of object contained in the target image.

[0049] Among user 140's image editing needs, portrait editing is often of particular interest. Furthermore, portraits are more difficult to edit than other object types, such as scenery and animals. Therefore, in this embodiment, using portraits as a visual feature distinction criterion can more quickly meet the customer's image editing needs.

[0050] In some embodiments, when determining the recommended instruction set associated with the visual feature, server device 130 may provide different recommended instruction sets depending on whether target image 145 contains a human portrait or not. It may also provide different recommended instruction sets depending on whether target image 145 contains a single portrait or multiple portraits. It may also provide different recommended instruction sets depending on whether the target image 145 contains a human portrait, an animal, a natural scene, or other object type.

[0051] Specifically, for images containing human figures, the personalized features of the figures can be mapped to multiple tags and scores. The tags are then prioritized based on the scores, and command phrases from the corresponding tag clusters are selected based on the priorities. For images without human figures, command phrases for common scenarios can be randomly extracted.

[0052] Therefore, providing different recommended instruction sets based on these different conditions is conducive to providing user 140 with fast and high-quality image editing services.

[0053] In some embodiments, in response to the interactive operation being an image editing operation of the target image 145 by the user 140 , the server device 130 determines the target image 145 updated by the image editing operation and determines a recommended instruction set matching the updated target image 145 .

[0054] As shown in Figure 3, after user 140 performs editing operations such as photo editing interactions on a target image in image editing application 120, an updated target image 145 is generated. Via user interaction module 330 of terminal device 110, instruction recommendation module 310 of server device 130 can obtain updated target image 145 and determine that the content of target image 145 has changed. Based on this, instruction recommendation module 310 can provide a recommended instruction set that matches the updated target image 145.

[0055] For example, if the image editing operation is image magnification or image cropping, the instruction recommendation module 310 may recommend a recommended instruction set suitable for editing the image area based on the image area magnified or cropped by the user 140 .

[0056] In some embodiments, when determining a recommended instruction set that matches the updated target image 145 , the server device 130 may determine the user's focus area in the target image 145 based on the image editing operation, and determine the recommended instruction set based on the focus area.

[0057] Specifically, the updated content of the target image 145 can be input into the visual algorithm in real time, and the visual algorithm detects the key points of the position change and predicts the user's center of attention. The visual algorithm can then recommend instruction phrases to the user 140 under the tag cluster with the visual center semantics as the label.

[0058] For example, if the image editing operation is to enlarge the eyes of a portrait, the server device 130 can recognize that the user 140's area of ​​interest is the eyes, and then provide a recommended instruction set related to the eyes.

[0059] In such an embodiment, the server device 130 can quickly and accurately understand the image editing needs of the user 140 and recommend an appropriate instruction set, which effectively improves the interaction experience between the user 140 and the image editing application 120.

[0060] 2 , at block 220 , server device 130 causes the recommended instruction set to be presented to the user, for example, via a user interface of terminal device 110 . In some embodiments, when causing the recommended instruction set to be presented to the user, server device 130 may present the recommended instruction set through an interactive window with user 140 in image editing application 120 .

[0061] Specifically, in the image editing application 120, an interaction window can be provided between the user 140 and the digital assistant, in which a recommended instruction set is displayed so that the user 140 can select a recommended instruction from the recommended instruction set to conveniently determine the image editing operation to be performed, thereby simplifying the interaction process.

[0062] Continuing with reference to FIG2 , in box 220 , in response to detecting the user 140 selecting a first instruction in the recommended instruction set, the server device 130 provides the first instruction to the target model 135 , so that the target model 135 determines at least one image editing operation to be applied to the target image 145 based on the first instruction.

[0063] As shown in FIG3 , the selection of the first instruction can be implemented by the user interaction module 330 of the terminal device 110. After the user 140 selects the first instruction from the recommended instruction set, the selection result (i.e., the selected first instruction) can be sent to the target model 135 via the user interaction module 330. The target model 135 determines one or more image editing operations to be applied to the target image 145 based on the first instruction, thereby realizing the image editing requirements of the user 140.

[0064] In the above embodiment, by packaging the image editing function into a photo editing instruction phrase, the photo editing instruction phrase can be between natural language and beautification atomic capability, short and easy to understand, thereby determining the specific type and degree of beautification atomic capability through the instruction phrase.

[0065] In addition, after the user uploads the image, the personalized features of the image are intelligently identified by using computer vision algorithms. The user's photo editing intention can be predicted before the traditional interactive dialogue mode begins, thereby providing the user with a recommended instruction set.

[0066] In addition, during the user's interactive photo editing process, it can adaptively identify and detect the user's attention center, predict the user's photo editing intention, and update the recommendation results in real time, thereby improving the efficiency and quality of image editing.

[0067] In some embodiments, the server device 130 receives text input from the user 140 via the interactive window between the image editing application 120 and the user 140. If it is determined that the semantics of the text input is incomplete, another recommended instruction set matching the text input is determined and presented to the user 140.

[0068] Specifically, when the user 140 inputs a partial text, the recommendation algorithm can retrieve matching instruction phrases using the text keywords. Therefore, the user can use incomplete text with keywords to input his or her image editing intention.

[0069] For example, assuming that user 140 enters "brightness" in the interactive window, that is, does not enter a complete sentence, such as "adjust the brightness", the server device 130 can determine the matching recommended instruction set based on the keyword "brightness" and present it to user 140.

[0070] In some embodiments, if it is determined that the semantics of the text input are ambiguous, the server device 130 determines another recommended instruction set based on predetermined rules, and causes the recommended instruction set to be presented to the user 140 .

[0071] Specifically, the fuzzy text detection model can be used to filter out fuzzy prompt words, and then a recommendation algorithm can be used to recommend instruction phrases with similar semantics as tag clusters to the user 140. The specific content of the predetermined rules can be set according to actual needs.

[0072] In some embodiments, in response to detecting user 140 selecting a second instruction from another set of recommended instructions, the second instruction is provided to target model 135 so that target model 135 determines at least one image editing operation to be applied to target image 145 based on the second instruction.

[0073] As shown in FIG3 , the selection of the second instruction can be implemented by user interaction module 330 of terminal device 110. After user 140 selects the second instruction from the other recommended instruction set, user interaction module 330 can send the selection result to target model 135. Target model 135 determines the image editing operation to be applied to target image 145 based on the second instruction, thereby implementing the image editing requirements of user 140.

[0074] In some embodiments, if the text input is determined to be semantically complete and semantically unambiguous, the text input is provided to target model 135 so that target model 135 can determine at least one image editing operation to be applied to target image 145 based on the text input.

[0075] The text processing module 320 of the server device 130 can determine the semantic completeness and ambiguity of the input text. If the text input is semantically complete and unambiguous, the text processing module 320 can provide the text to the target model 135. When the target model 135 is a large language model, the large language model can deeply understand the meaning of the text and efficiently process natural language tasks. Therefore, the target model 135 can directly determine the image editing operation to be applied to the target image 145 based on the complete and unambiguous text, without the need for recommendation instructions.

[0076] Therefore, compared with traditional text interaction links, this solution adds a recommendation function for incomplete and ambiguous user input, recommending a set of semantically matching instruction phrases, thereby speeding up image editing and effectively improving the user experience.

[0077] FIG4 shows a flow chart of a process 400 for image editing according to some embodiments of the present disclosure. Process 400 may be implemented in environment 100 of FIG1 . Process 400 may be shown as a specific embodiment of the steps of process 200. Process 400 may be implemented using architecture 300 of FIG3 .

[0078] As shown in FIG4 , the user 140 selects ( 410 ) a target image 145 to be edited in the image editing application 120 , and the target image 145 can be uploaded ( 430 ) to the instruction recommendation module 310 of the server device 130 through the user interaction module 330 of the terminal device 110 . Based on the original target image 145 , the instruction recommendation module 310 can use a visual algorithm to detect whether the target image 145 contains a human portrait ( 432 ). If a human portrait is contained, the algorithm can be used to detect the human body personalized features of the human portrait ( 434 ), and then the recommended instruction phrase ( 440 ) can be obtained through the recommendation algorithm. If a human portrait is not contained ( 436 ), the recommended instruction phrase ( 440 ) can be directly obtained through the recommendation algorithm. These instruction phrases can be grouped into a recommended instruction set 490 . The recommended instruction set 490 can be presented ( 492 ) to the user 140 through the user interaction module 330 .

[0079] At the terminal device 110, after the user 140 selects the instruction (416), the user interaction module 330 provides (494) the instruction to the target model 130. The target model 130 can parse (480) the instruction, for example, to parse out the image editing operation, sequence number, and operation item corresponding to the instruction, such as "add" or "undo" the image editing operation. The parsed instruction is then provided to the local rendering module 340 of the terminal device 110, which performs the corresponding image editing operation on the target image 145 and renders the edited image.

[0080] In some embodiments, the image editing operations generated by the target model 130 from user text input can also be packaged into one or more instructions and stored in the instruction phrase library 350. That is, the instructions in the instruction phrase library 350 can be derived from textual interactions with the user. For example, when a user uses the target model to perform image editing interactions, they often need to indicate the execution of a certain type of image editing operation through text input. In this case, such an image editing operation can be packaged into a single instruction. The newly packaged instruction can be recommended to the user when the user has similar image editing intentions during subsequent use.

[0081] In FIG4 , after user 140 edits a target image 145 in the image editing application 120 ( 412 ), the user interaction module 330 of the terminal device 110 can detect whether an image editing interaction has occurred on the updated target image 145 ( 420 ). If it is determined that the image content has changed ( 422 ), an attention mechanism recognition algorithm 424 can be used to identify the user 140's area of ​​interest, and then a recommendation algorithm can be used to recommend instruction phrases that match the area of ​​interest ( 440 ). These instruction phrases can be aggregated into a recommended instruction set 490 . Subsequently, user 140 can select instructions based on the recommended instruction set 490 and perform corresponding image editing on the target image 145 via the target model 130 and the local rendering module 340 .

[0082] In addition, user 140 may also enter text in image editing application 120 (414). After the text processing module 320 of server-side device 130 detects the entered text (470), it determines whether the text is complete (472). If the text is incomplete, instruction recommendation module 310 may use keywords in the text to recommend matching instruction phrases (450). These instruction phrases may be aggregated into a recommended instruction set 490. Subsequently, user 140 may select an instruction based on the recommended instruction set 490 and perform corresponding image editing on target image 145 via target model 130 and local rendering module 340.

[0083] If the text is complete, the text processing module 320 can perform a recognition judgment (474) on the complete text to determine whether the image editing function mentioned in the text is not supported. The recognition judgment is mainly used to determine whether the target model 135 can process and respond to the currently input text. If the input text passes the recognition judgment, that is, it is determined that the image editing function mentioned in the text is supported, the text is then subjected to a semantically fuzzy text judgment (476). If it is determined that the text is semantically fuzzy, the instruction recommendation module 310 can recommend matching instruction phrases (460) based on predetermined rules, such as based on the approximate words of the fuzzy prompt words found. These instruction phrases can be grouped into a recommended instruction set 490. Subsequent users 140 can select instructions based on the recommended instruction set 490 and perform corresponding image editing on the target image 145 via the target model 130 and the local rendering module 340.

[0084] If the recognition judgment fails (478), the user interaction module 330 may send a relevant prompt to the user 140 to remind the user 140 that the photo editing function is not supported.

[0085] In addition, if the result that the text is not ambiguous is obtained based on the fuzzy text judgment (476), it is determined that the text is semantically complete and not ambiguous, and the text is directly provided to the target model 130, so that the target model 130 determines one or more image editing operations to be applied to the target image 145 based on the text input.

[0086] If the user 140 is not satisfied with the recommendation result, the user may repeat the steps of image editing ( 412 ) and text editing ( 414 ) in FIG. 4 to further process the image until a satisfactory editing result is obtained.

[0087] Through process 400, the interaction process between 140 and the image editing application 120 is simplified. Before entering a text conversation, the image content can be quickly understood through the image editing application 120, and personalized instruction recommendations can be given in advance. It can also adaptively perceive changes in image content under user interaction behavior, infer the user's next action, and recommend instructions in advance. In addition, in the traditional text conversation mode, users who enter semantically ambiguous text will be directly rejected for photo editing, while this solution can recommend related functions to users, making the user experience better. Therefore, this solution effectively improves the user experience.

[0088] 5 shows a schematic structural block diagram of an apparatus 500 for image editing according to certain embodiments of the present disclosure. Apparatus 500 may be implemented as or included in server device 130. Each module / component in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0089] As shown in the figure, the apparatus 500 includes a detection module 510 configured to detect an interactive operation performed on a target image to be edited.

[0090] The apparatus 500 further includes a determining module 520 configured to determine a recommended instruction set for interacting with the target model based on the interaction operation, wherein the recommended instruction set includes at least one instruction, and each instruction is defined as corresponding to at least one image editing operation.

[0091] The apparatus 500 further includes a presentation module 530 configured to enable the recommended instruction set to be presented to the user.

[0092] The apparatus 500 further includes a providing module 540 configured to provide the first instruction to the target model in response to detecting a user selection of the first instruction in the recommended instruction set, so that the target model determines at least one image editing operation to be applied to the target image based on the first instruction.

[0093] In some embodiments, the determination module 520 is further configured to extract visual features of the target image in response to the user selecting a target image for editing through an interactive operation; and determine a recommended instruction set associated with the visual features.

[0094] In some embodiments, the visual feature indicates at least one of the following: the target image includes a human portrait, the target image does not include a human portrait, the number of human portraits included in the target image, and the type of object included in the target image.

[0095] In some embodiments, the determination module 520 is further configured to determine the target image updated by the image editing operation in response to the interactive operation being an image editing operation performed by the user on the target image; and determine a recommended instruction set matching the updated target image.

[0096] In some embodiments, the determination module 520 includes a focus module configured to determine a focus area of ​​the user in the target image based on the image editing operation; and determine a recommended instruction set based on the focus area.

[0097] In some embodiments, the presentation module 530 is further configured to present the recommended instruction set through an interactive window with the user in the image editing application.

[0098] In some embodiments, the device 500 also includes a text interaction module, which is configured to receive user text input via an interaction window with the user in an image editing application; if it is determined that the semantics of the text input is incomplete, determine another recommended instruction set that matches the text of the text input; if it is determined that the semantics of the text input is ambiguous, determine another recommended instruction set based on predetermined rules; cause the other recommended instruction set to be presented to the user; and in response to detecting the user's selection of a second instruction in the other recommended instruction set, provide the second instruction to the target model, so that the target model determines at least one image editing operation to be applied to the target image based on the second instruction.

[0099] In some embodiments, the device 500 also includes a model application module that is configured to provide the text input to the target model if it is determined that the text input is semantically complete and semantically unambiguous, so that the target model determines at least one image editing operation to be applied to the target image based on the text input.

[0100] The units and / or modules included in the device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 500 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0101] FIG6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 may be used to implement the server device 130 of FIG1 or the apparatus 500 of FIG5 .

[0102] As shown in FIG6 , electronic device 600 is a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.

[0103] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 600.

[0104] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0105] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0106] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0107] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0108] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0110] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0111] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0112] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for image editing, comprising: detecting interactive operations performed on a target image to be edited; determining a recommended instruction set for interacting with the target model based on the interaction operation, the recommended instruction set comprising at least one instruction, each instruction being defined as corresponding to at least one image editing operation; causing the recommended instruction set to be presented to a user; as well as In response to detecting a selection by the user of a first instruction from the set of recommended instructions, the first instruction is provided to the target model so that the target model determines at least one image editing operation to be applied to the target image based on the first instruction.

2. The method according to claim 1, wherein determining a recommended instruction set for interacting with the target model based on the interaction operation comprises: extracting visual features of the target image in response to the user selecting the target image for editing through the interactive operation; as well as A recommended instruction set associated with the visual feature is determined.

3. The method of claim 2, wherein the visual feature indicates at least one of: The target image includes a portrait, The target image does not contain a human portrait, The number of human figures contained in the target image, The type of object contained in the target image.

4. The method according to any one of claims 1 to 3, wherein determining a recommended instruction set for interacting with a target model based on the interaction operation comprises: In response to the interactive operation being an image editing operation performed by the user on the target image, determining the target image updated by the image editing operation; as well as A recommended instruction set matching the updated target image is determined.

5. The method of claim 4 , wherein determining a recommended instruction set matching the updated target image comprises: Based on the image editing operation, determining a region of interest of the user in the target image; as well as Based on the area of ​​interest, the recommended instruction set is determined.

6. The method of any one of claims 1 to 5, wherein causing the recommended set of instructions to be presented to the user comprises: The recommended instruction set is presented through an interactive window with the user in an image editing application.

7. The method according to any one of claims 1 to 6, further comprising: receiving text input from the user via an interactive window with the user in the image editing application; If it is determined that the semantics of the text input is incomplete, determining another recommended instruction set that matches the text of the text input; If it is determined that the semantics of the text input are ambiguous, determining another recommended instruction set based on predetermined rules; causing the another recommended instruction set to be presented to the user; as well as In response to detecting a selection by the user of a second instruction from the other recommended instruction set, the second instruction is provided to the target model so that the target model determines at least one image editing operation to be applied to the target image based on the second instruction.

8. The method according to claim 7, further comprising: If it is determined that the text input is semantically complete and semantically unambiguous, the text input is provided to the target model so that the target model determines at least one image editing operation to be applied to the target image based on the text input.

9. An apparatus for image editing, comprising: a detection module configured to detect an interactive operation performed on a target image to be edited; a determination module configured to determine a recommended instruction set for interacting with the target model based on the interaction operation, the recommended instruction set comprising at least one instruction, each instruction being defined as corresponding to at least one image editing operation; a presentation module configured to cause the recommended instruction set to be presented to a user; as well as A providing module is configured to provide the first instruction to the target model in response to detecting the user's selection of the first instruction in the recommended instruction set, so that the target model determines at least one image editing operation to be applied to the target image based on the first instruction.

10. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 8.

12. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Fuzzy instruction interaction method suitable for smart home

    CN108710310A

  • Image editing method and device, electronic equipment and storage medium

    CN111429551A

  • Image editing method and device, electronic equipment, storage medium and program product

    CN116168119A

  • Recommendation method and device based on interaction, computer equipment and storage medium

    CN117370404A

  • Method and device for generating 2D virtual human video, storage medium and equipment

    CN117574963A