Image editing method and electronic device

CN122593673APending Publication Date: 2026-08-18HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510175792.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而,采用目前的图像编辑方法仍然会出现背景编辑效果不好的问题

Benefits of technology

[0040]应当理解的是,本申请的第二方面至第六方面与本申请的第一方面的技术方案相对应,各方面及对应的可行实施方式所取得的有益效果相似,不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593673A_ABST
    Figure CN122593673A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image editing method and an electronic device, and relate to the technical field of terminals. The method comprises: obtaining a first editing instruction input for a first image, the first editing instruction being used to instruct adjustment of a background of the first image; generating first indication information according to the first image and the first editing instruction, the first indication information comprising the first image and the first editing instruction, the first indication information being used to instruct rewriting of the first editing instruction based on the first image; inputting the first indication information to a first model to obtain a second editing instruction output by the first model, the second editing instruction being used to instruct adjustment of the background of the first image; and generating a second image according to the second editing instruction, the first image and a second model, the second image being different from the background of the first image. In this way, the background of the first image is edited based on the information of the text content and the first image, and the effect of the image background editing method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to an image editing method and an electronic device. Background Technology

[0002] With the development of artificial intelligence, electronic devices can support the ability to personalize image backgrounds, aiming to meet the diverse needs of users.

[0003] In current implementations, some algorithms can edit image backgrounds based on user-provided text. In this process, users can input simple keywords to indicate the desired image background. The algorithm can then adjust the original image based on these keywords to achieve the desired background editing. However, current image editing methods still suffer from poor background editing results. Summary of the Invention

[0004] This application provides an image editing method and electronic device, applicable to the field of terminal technology, to improve the effect of image background editing methods.

[0005] Firstly, embodiments of this application propose an image editing method. The method includes:

[0006] Obtain a first editing instruction input for the first image, the first editing instruction being used to instruct adjustments to the background of the first image.

[0007] Based on the first image and the first editing instruction, a first instruction information is generated. The first instruction information includes the first image and the first editing instruction, and is used to instruct the first editing instruction to be rewritten based on the first image.

[0008] The first instruction information is input into the first model to obtain the second editing instruction output by the first mode. The second editing instruction is used to instruct the background of the first image to be adjusted.

[0009] A second image is generated based on the second editing instructions, the first image, and the second model. The background of the second image is different from that of the first image.

[0010] In this embodiment, first instruction information can be generated based on the first image and the first editing instruction to guide the first model to generate reasonable rewriting instructions. These rewriting instructions then guide the second model to generate a second image that matches the foreground target with the background image. In this process, existing first and second models can be applied without the need for joint training or complex adjustments to the architecture and training process of existing models. This allows image editing functionality to be achieved while ensuring the generation of reasonable images, reducing the complexity of the image editing method and improving its adaptability and effectiveness.

[0011] In one possible implementation, first instruction information is generated based on the first image and the first editing instruction, including:

[0012] Image segmentation processing is performed on the first image to obtain the foreground image of the first image.

[0013] First instruction information is generated based on the foreground image of the first image and the first editing instruction.

[0014] The first instruction information is used to instruct the foreground image based on the first image to rewrite the first editing instruction.

[0015] In this implementation, image segmentation separates the foreground from the background, allowing for more precise background adjustments without affecting the foreground object. This facilitates maintaining the integrity and quality of the foreground during subsequent editing. Furthermore, rewriting editing instructions using information from the foreground image makes the instructions more targeted and refined, thereby improving the accuracy of the editing results.

[0016] In one possible implementation, the first instruction information is further used to indicate a first condition for rewriting the first editing instruction based on the foreground image of the first image, the first condition being used to restrict the background of the image to be adjusted by the rewritten editing instruction to be compatible with the foreground image of the first image.

[0017] In this implementation, by using a first condition to ensure that the background of the image to be adjusted matches the foreground image of the first image, the adjusted background and foreground target are well-matched, thus maintaining the overall visual consistency and harmonious aesthetics of the image. The first condition prevents the background and foreground from clashing in color, style, or subject matter, thereby avoiding discordant editing results. This method makes the edited image look natural and harmonious. This is crucial for generating natural and realistic images.

[0018] In one possible implementation, after performing image segmentation on the first image, the method further includes obtaining a foreground mask of the first image.

[0019] In this implementation, the foreground mask can precisely identify which parts of the image belong to the foreground, facilitating specific operations on the foreground or background in subsequent processing steps without affecting other parts. By using a mask, editing operations can be quickly applied to specific areas without manual selection or adjustment, improving editing efficiency. At the same time, the foreground mask reduces the risk of accidental manipulation of the foreground.

[0020] In one possible implementation, a second image is generated based on a second editing instruction, a first image, and a second model. The background of the second image differs from that of the first image, including:

[0021] The second editing command, the foreground image of the first image, and the foreground mask of the first image are input into the second model to obtain the second image output by the second model.

[0022] In one possible implementation, the background of the second image is consistent with the background of the image adjusted as instructed by the second editing command.

[0023] In this implementation, by using a foreground mask, the foreground portion is not subjected to unnecessary modifications during processing, ensuring that its quality and detail are preserved. The second model can flexibly generate background effects according to a second editing command to meet the user's needs.

[0024] In one possible implementation, the first model is a multimodal large language model, and the second model is an image editing model.

[0025] In this implementation, the multimodal large language model can understand and process the first editing instruction and the first image, and generate a second editing instruction based on the first instruction information that more accurately describes the relationship between the foreground target and the content described in the first editing instruction. The image editing model can perform specific image processing tasks based on the second editing instruction, the foreground image, and the foreground mask, generating high-quality image results by efficiently applying the second editing instruction, the foreground image, and the foreground mask.

[0026] In one possible implementation, obtaining the first editing instruction input for the first image includes:

[0027] In response to user actions on the first interface of the first application, an input box is displayed, and a first image is displayed on the first interface.

[0028] Based on the text input operation performed in the input box, obtain the first editing instruction input for the first image.

[0029] In this implementation, users can directly see the first image on the first interface and input editing commands through the input box. This intuitive interaction reduces operational complexity, making it easier for users to understand and use. Furthermore, by detecting the input box, the first editing command can be obtained in a timely manner to perform image editing processing on the first image.

[0030] Secondly, embodiments of this application provide an image editing device, which can be an electronic device, or a chip or chip system within an electronic device. The image editing device may include a display unit and a processing unit.

[0031] When the image editing device is an electronic device, the display unit therein can be a display screen. The display unit is used to perform the display steps so that the electronic device implements an image editing method described in the first aspect or any possible implementation of the first aspect.

[0032] When the image editing apparatus is an electronic device, the processing unit may be a processor. The image editing apparatus may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement an image editing method described in the first aspect or any possible implementation thereof.

[0033] When the image editing device is a chip or chip system within an electronic device, the processing unit can be a processor. The processing unit executes instructions stored in the storage unit to cause the electronic device to implement an image editing method described in the first aspect or any possible implementation of the first aspect. The storage unit can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).

[0034] For example, a processing unit is used to process a first editing instruction and a first image, and to process the first editing instruction and the first image based on an image editing device. A display unit is used to display an operation interface and an image.

[0035] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory for storing code instructions, and the processor for running the code instructions to perform the methods described in the first aspect or any possible implementation of the first aspect.

[0036] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0037] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0038] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.

[0039] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).

[0040] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0041] Figure 1 This is a schematic diagram illustrating an application scenario of the image editing method provided in the embodiments of this application;

[0042] Figure 2 A flowchart illustrating an image editing method provided in an embodiment of this application;

[0043] Figure 3 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of the software structure of a terminal device provided in an embodiment of this application;

[0045] Figure 5 A schematic flowchart illustrating the image editing method provided in an embodiment of this application;

[0046] Figure 6A schematic diagram of image segmentation results provided in an embodiment of this application;

[0047] Figure 7 A schematic diagram illustrating the generation of a second editing instruction provided in an embodiment of this application;

[0048] Figure 8 A schematic diagram of the image background editing results provided in the embodiments of this application;

[0049] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0050] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0051] 1. Multimodal large language model

[0052] A multimodal large model is an artificial intelligence model capable of simultaneously processing and understanding multiple types of data (such as text, images, audio, and video). It can not only understand information from a single modality but also perform information conversion between modalities, such as generating images from text, generating descriptions from images, and converting speech to text.

[0053] 2. Other terms

[0054] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0055] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0056] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.

[0057] 3. Electronic equipment

[0058] The electronic devices in this application embodiment may include handheld devices with image processing functions, vehicle-mounted devices, etc. For example, some electronic devices include: mobile phones, tablets, PDAs, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, terminal devices in 5G networks, or future evolution of public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.

[0059] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0060] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0061] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.

[0062] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0063] To better understand the technical solution of this application, the relevant technologies involved in this application will be further described in detail below.

[0064] With the continuous development of artificial intelligence technology, electronic devices can now support image background editing to meet users' personalized needs. This image editing technology allows users to easily modify image backgrounds according to their specific requirements. At the same time, this technology significantly improves the efficiency and flexibility of background editing, further satisfying users' growing demand for personalized and customized image content.

[0065] The following is combined Figure 1 The application scenarios of image editing methods are introduced. Figure 1 This is a schematic diagram illustrating an application scenario of the image editing method provided in the embodiments of this application.

[0066] like Figure 1 As shown in (a), this can be understood as the gallery interface of an electronic device. The gallery interface can display multiple functional controls, such as Photos, Albums, Moments, and Discover, as shown in the figure. Users can click on any of these functional controls to display the corresponding content. (See reference...) Figure 1 In (a), after the user opens the Gallery app, the content corresponding to the Photos control can be displayed by default, which can show multiple images stored at different shooting times. For example... Figure 1 As shown in (a), 5 images are displayed during this time period today, and 9 images can be displayed during this time period yesterday.

[0067] Users can click on any image displayed in the gallery interface. (Reference) Figure 1 In (a), the user clicks on image 101. In response to the user's click on image 101, the system can redirect to display the following: Figure 1 The image details page shown in (b) of the image. Figure 1 In illustration (b), image 101 and its capture time (5:10 PM on June 5, 2020, as shown in the image) can be displayed. Image 101 depicts a person meditating on a mat. Furthermore, the image details page can display multiple function controls, such as share, favorite, edit, delete, and more (102) as shown in the image. Users can click on any of these function controls to display the corresponding content. (See reference...) Figure 1 In (b), the user can click on the more functional control 102. In response to the user's click on the more functional control 102, the following can be displayed: Figure 1 The interface shown in (c) is shown in the image.

[0068] Figure 1(c) indicates that a more functional window 104 can be displayed above the image details interface 103. This window can display multiple function controls, such as "Set as Wallpaper," "Add to," "Cast," "Print," and "Background Editing" 105, as shown in the image. Users can click on any of these function controls to display their corresponding content. (See reference...) Figure 1 In (c), the user can click on the background editing control 105. In response to the user's click on the background editing control 105, the following can be displayed: Figure 1 The first interface is shown in (d) in the diagram.

[0069] like Figure 1 As shown in (d), the first interface can display a first image 101, an input box 106, a submit button, and a cancel button, etc. The user can click on the input box 106 to bring up a virtual keyboard. Furthermore, by clicking on the virtual keyboard, the user can input a first editing command into the input box. The first editing command is used to instruct adjustments to the background of the first image 101. (Refer to...) Figure 1 In (d), the user clicks input box 106. Responding to the user's click action on the input box, the following can be displayed: Figure 1 The interface shown in (e) is shown in the image.

[0070] like Figure 1 As shown in (e), a virtual keyboard is displayed below the cancel button. Users can enter the first editing command (the white cloud shown in the image) in the input box by clicking on some of the virtual keys on the virtual keyboard. After completing the text input in the input box, the user can click the submit button to issue an image background editing command to the electronic device, or click the cancel button to abandon the background editing operation for the first image.

[0071] Reference Figure 1 In (e), the user enters the first editing instruction "white clouds" in the input box. Furthermore, by clicking the submit button, the electronic device can be instructed to perform background editing on the first image according to the first editing instruction. After the electronic device completes the background editing operation on the first image, it can jump to display as shown below. Figure 1 The interface shown in (f) is shown in the figure. Figure 1 (f) in the diagram illustrates the second image 107, which is an image after background editing of the first image according to the first editing instruction. Figure 1In the illustration, the target person in the original image is suspended in the sky, surrounded by floating clouds. Since this scene of a person suspended in the sky is not realistic, the second image 107 generated based on the first editing instruction of "clouds" can be considered unreasonable. In this embodiment, images whose content does not conform to reality can be defined as unreasonable images. Afterwards, the user can save the second image 107 to the gallery application by clicking the save button, or the user can cancel the saving operation for the second image 107 by clicking the cancel button.

[0072] Based on the application scenarios described above, it can be determined that users can use simple text input to describe the changes they expect to make to the background of the original image, thus instructing the user to edit the background of the original image. However, editing the background of the original image solely based on the user's text input may result in situations where the image illustrated in the above scenarios is unreasonable.

[0073] To address this issue, one approach involves employing a multimodal large language model and an image editing model to implement image editing functionality. This implementation integrates the multimodal large language model and the image editing model, and uses a constructed loss function to jointly train both models, thereby adjusting their parameters. The following section will combine... Figure 2 This paper introduces an existing implementation of an image editing method. Figure 2 This is a flowchart illustrating an image editing method provided in an embodiment of this application.

[0074] like Figure 2 As shown, user input commands can be fed into the embedding layer to convert the text commands into a representation usable for image editing, i.e., to obtain the text features corresponding to the user input commands. Additionally, the original image can be input into the adapter to obtain the image features corresponding to the original image. Then, the text and image features can be input into a multimodal large language model. The multimodal large language model can understand and process the image features of the original image and the text features of the user input commands, and output optimized editing commands through a language head. The optimized editing commands can provide a more accurate description of the original image and the user input commands.

[0075] After receiving the optimized editing instructions, the editing head processes them to generate corresponding image editing instructions. This converts the optimized instructions into a text format more suitable for the visual editing model, ensuring accurate understanding and execution. Image editing instructions allow for a closer integration of text and image information, guaranteeing a high degree of match between verbal descriptions and visual editing operations, thus avoiding discrepancies. The original image and editing instructions can then be input into the image editing model to guide the editing process. Finally, the image editing model outputs the image with the background edited.

[0076] For example, suppose there is an original image with the content "There is a small cabin in the forest," and the user inputs the instruction "Make the cabin look like it's in the desert." Based on the above, the optimized editing instruction could be, for example, "Change the surrounding environment to make it more like a dry desert; remove the green trees and add sand dunes and cacti." After the editing head processes the optimized editing instruction, it can generate image editing instructions such as: "Remove the forest background and replace it with a desert landscape; add specific elements, such as cacti and sand dunes; adjust the color style to make the image drier and more desolate." The background-edited image output by the image editing model can include, for example, the small cabin, the desert background, and elements such as cacti and sand dunes.

[0077] When training this ensemble model, two loss functions need to be constructed to guide the parameter tuning of the multimodal large language model and the image editing model, respectively. For example, a target loss function for text can be constructed to optimize the understanding and expression of text instructions, thereby improving the quality of text information. In one implementation, the target loss function for text can be constructed by calculating the difference between the output text of the multimodal large language model and the target text, thus guiding the parameter tuning of the multimodal large language model.

[0078] Furthermore, a target loss function can be constructed for the image to optimize the generation of the image after background editing. In one implementation, the target loss function can be constructed by calculating the difference between the output image of the image editing model and the target image, thereby guiding the image editing model to adjust its parameters.

[0079] refer to Figure 2 It can be determined that this image editing method requires treating the embedding layer, adapter, language head, editing head, multimodal large language model, and image editing model as a whole, and based on... Figure 2The architecture shown is meticulously designed. Training the model requires joint training of these modules based on this architecture. However, this increases the complexity and cost of model training. Furthermore, the need to construct two loss functions—one for text and one for images—during training also increases computational load, potentially prolonging training time.

[0080] To address the issues described above, this application proposes the following technical concept: First, user input instructions and the original image can be input into a multimodal large language model, which can output edited instructions. The edited instructions can accurately indicate the relationship between the content of the user input instructions and the foreground objects in the original image, thus accurately guiding the background editing of the original image. Then, the edited instructions and the first image can be input into an image editing model. The image editing model can edit the background of the original image according to the edited instructions to output an edited image. Notably, the multimodal large language model and the image editing model are decoupled, meaning that they do not need to be jointly trained during the training phase.

[0081] The image editing method of this application embodiment can be executed by an electronic device equipped with image processing, or by a chip, chip system, or processor that supports the implementation of the image editing method by the electronic device, or by a logic module or software that can implement all or part of the functions of the electronic device. This application does not impose specific limitations in this regard. The image editing method of this application embodiment will be described in detail below using an electronic device as the execution subject.

[0082] Electronic devices can be, for example, terminal devices. The following section will first combine... Figure 3 and Figure 4 A brief introduction to the terminal equipment.

[0083] For example, Figure 3 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application.

[0084] Figure 3This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0085] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device. In other embodiments of this application, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0086] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. In one implementation, for example, the image editing method provided in this application may be executed by processor 110.

[0087] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel.

[0088] Camera 193 is used to capture still images or videos.

[0089] The software system of a terminal device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This application uses the layered architecture Android system as an example to illustrate the software structure of the terminal device.

[0090] For example, Figure 4 This is a schematic diagram of the software structure of a terminal device provided in an embodiment of this application.

[0091] like Figure 4 As shown, the layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the system may include an application layer, an application framework layer, an Android runtime and system libraries, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that this application uses the Android system as an example; however, the solution can also be implemented in other operating systems (such as HarmonyOS, iOS, etc.) as long as the functions implemented by each module are similar to those in the embodiments of this application.

[0092] The application layer can include a series of application packages.

[0093] like Figure 4 As shown, the application package may include applications such as gallery, camera, calendar, phone, map, music, settings, email, video, and social media. Of course, the application layer may also include other application packages, such as third-party applications like payment apps, shopping apps, banking apps, and social media apps; this application is not limited to these. In this embodiment, corresponding operations can be performed in the gallery application to achieve image editing functionality.

[0094] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0095] like Figure 4 As shown, the application framework layer may include a window manager, content provider, resource manager, view system, notification manager, etc.

[0096] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0097] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0098] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0099] The system library can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0100] The HAL layer is a wrapper around Linux kernel drivers, providing interfaces to the upper layers and shielding them from the implementation details of the lower-level hardware.

[0101] The HAL layer may include a Wi-Fi HAL, an audio HAL, a Camera HALServer unit, and software code libraries. In this embodiment, a multimodal large language model and an image editing model can be deployed in the HAL layer. The multimodal large language model can, for example, process the original image and user input instructions to obtain instructions that accurately indicate the relationship between foreground objects in the original image and the content described in the user input instructions, thereby guiding the adjustment of the background of the original image. The image editing model can, for example, adjust the background of the original image according to the instructions output by the multimodal large language model to obtain an edited image.

[0102] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0103] The technical solutions of the embodiments of this application and how the technical solutions of the embodiments of this application solve the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific examples. The following specific embodiments can be implemented independently or in combination with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0104] The following describes the workflow of the image editing method with specific examples. Figures 5 to 8 A detailed introduction will be given, including, Figure 5 This is a schematic flowchart of the image editing method provided in the embodiments of this application. Figure 6 This is a schematic diagram of the image segmentation results provided in an embodiment of this application. Figure 7 This is a schematic diagram illustrating the generation of a second editing instruction provided in an embodiment of this application. Figure 8 This is a schematic diagram of the image background editing result provided in an embodiment of this application. The image editing method may include, for example, the following steps:

[0105] S501. Process the first image to obtain the foreground image and foreground mask of the first image.

[0106] A first image may contain numerous pixels, including background pixels and foreground pixels (pixels of interest). When editing the background of a first image, it's crucial to clearly distinguish between the foreground and background to ensure the integrity and quality of the foreground objects during subsequent editing. Therefore, before background editing, the first image can be segmented to obtain a foreground image and a foreground mask. The foreground mask can be used to quickly locate foreground objects in the foreground image and ensure their integrity.

[0107] In one implementation, an image segmentation model can be used to obtain the foreground image and foreground mask of the first image. (See reference...) Figure 5 To understand this, after the first image is input into the image segmentation model, the image segmentation model can output the foreground image and foreground mask of the first image.

[0108] The following will use specific examples, combined with Figure 6 The results of image segmentation are presented. For example... Figure 6 As shown, suppose there is a first image 601. In the first image 601, a car is driving on the road, sunlight casts soft shadows, trees stand around the road, and mountains stand in the distance. After inputting the first image 601 into the image segmentation model, the image segmentation model can output the foreground image 602 and the foreground mask 603 of the first image.

[0109] Reference Figure 6 In the foreground image 602, pixels corresponding to the target pixel region in the first image (i.e., the car shown in the figure) can be retained, while pixels in the first image other than the target pixel region can be discarded, i.e., their pixel values ​​can be set to 0 (black). In the foreground mask 603, the pixel values ​​corresponding to the target pixel region in the first image can be set to 255 (white), while the pixel values ​​in the first image other than the target pixel region can be set to 0. Thus, by using the foreground mask, the pixel region containing the car object can be quickly located in the foreground image.

[0110] S502. Input the first editing instruction and the foreground image of the first image into the first model to obtain the second editing instruction.

[0111] It's understandable that in real life, multiple objects within a space will have certain spatial relationships. Therefore, in an image, the multiple objects contained within the image need to satisfy certain spatial relationships to make the image's content plausible. For example, people, cars, and other objects cannot be suspended in the sky, etc.

[0112] It can also be understood that when instructing the background editing of an image via text commands, only when the text information can accurately describe the spatial relationship between multiple objects in the expected image can the image editing model be guided to generate an image with a reasonable spatial relationship based on the text information.

[0113] However, in practical applications, the first editing instruction input by the user often fails to adequately represent the spatial relationship between the expected image background and the foreground object in the first image. Therefore, before performing background editing on the original image based on the image editing model, the first editing instruction can be expanded to obtain a second editing instruction based on the spatial relationship between the background information described by the first editing instruction and the foreground object in the first image. The second editing instruction is used to instruct adjustments to the background of the first image, and it can accurately describe the spatial relationship between the background content described by the first editing instruction and the foreground object in the first image.

[0114] To ensure the rationality and harmony between background content and foreground objects in the desired image, certain conditions can be set to restrict their relationship. For example, these conditions might include ensuring the background color coordinates with the main color tone of the foreground, and maintaining consistent lighting conditions between the background and foreground. Consistent lighting conditions between the background and foreground can be understood as maintaining the same light source direction, light intensity, and shadow effects between the two.

[0115] In this embodiment, firstly, first instruction information can be generated based on the first editing instruction and the first image. This first instruction information is used to instruct the first editing instruction to be rewritten based on the first image. Then, the first editing instruction can be expanded using a first model. (See also...) Figure 5 By understanding this, the first editing instruction, the foreground image of the first image, and the first instruction information can be input into the first model, and the first model can output the second editing instruction.

[0116] In this embodiment, the first model is used to process the first editing instruction and the foreground image based on the first instruction information to obtain the second editing instruction. In one implementation, the first model can be a multimodal large language model. This embodiment can use any existing model as long as it can achieve the goal of generating the second editing instruction; therefore, the selection of the first model is not limited.

[0117] In one implementation, the first instruction information may include a first image and a first editing instruction. Furthermore, the first instruction information may also be used to indicate a first condition for rewriting the first editing instruction based on the foreground image of the first image. This first condition restricts the background of the image to be adjusted according to the foreground image of the first image. For example, when images containing objects such as people or vehicles suspended in the sky, or the sun appearing in the night sky, which do not conform to reality, it indicates that the adjusted image background is not compatible with the foreground image of the first image.

[0118] In one implementation, the adaptation of the image background to the foreground image of the first image can be manifested in the following ways: the spatial positional relationship between objects in the background and foreground objects in the image is reasonable; the color of the background is coordinated with the main color of the foreground; and the lighting conditions of the background and foreground are consistent.

[0119] In one implementation, the first instruction information can be set to generate a reasonable image editing instruction based on the foreground target image and the new background content described in the first editing instruction.

[0120] The following will use specific examples, combined with Figure 7 The results of generating the second editing command will be described. For example... Figure 7 As shown, the foreground image 602 of the first image is input into the first model. Based on the first instruction information, the first model is instructed to generate a reasonable image editing instruction according to the foreground target image and the new background content described in the first editing instruction. Assume the current first editing instruction is "desert".

[0121] Based on the above description, it can be determined that the first model can output the second editing instruction, namely, a car parked in a desert landscape, with a clear blue sky and sand dunes as the background. The vehicle is located on flat sand, with sunlight casting soft shadows. The desert scene should be vast, with several sand dunes in the distance. Based on the description of the second editing instruction, it can be determined that the spatial position of the car in the desert landscape is reasonable, that is, the vehicle is located on flat sand as described in the text. Furthermore, the second editing instruction specifies the lighting scene for the foreground object car, namely, the sunlight casting soft shadows as described in the text. In addition, the second editing instruction also provides a specific description of the elements contained in the desert background, namely, the background is a clear blue sky and sand dunes, the desert scene should be vast, and several sand dunes in the distance.

[0122] Based on the first image 601, different first editing instructions can be input, for example. Let's assume the current first editing instruction is "ancient building." By inputting the foreground image 602 of the first image, the first instruction information, and the first editing instruction into the first model, a second editing instruction can be obtained. For example, the second editing instruction could be: A car is parked in front of an ancient building with intricate stone carvings and towering pillars. The car's sleek design contrasts sharply with the historical background, creating a striking effect. The scene is bathed in warm sunlight, highlighting the details of the car and the building.

[0123] Based on the description in the second editing instruction, it can be determined that the spatial relationship between the car and the ancient building is reasonable, i.e., the car is parked in front of the ancient building as described in the text. Furthermore, the second editing instruction specifies the lighting scene for the foreground object (car) and the background object (ancient building), i.e., the scene described in the text is bathed in warm sunlight, highlighting the details of both the car and the building.

[0124] In the two examples above, the target object is a car. Below, we can further illustrate the result of the second editing instruction generated when the target object is a person. Suppose we have an image showing three men squatting on the floor of an office. And suppose the first editing instruction is blue sky and white clouds. By inputting the foreground image of the first image (the three men), the first instruction information, and the first editing instruction into the first model, we can obtain the second editing instruction.

[0125] As an example, the second editing instruction could be: Three young men dressed casually are sitting on the ground against a backdrop of blue sky and white clouds. The first man is wearing a light blue T-shirt and dark pants, the second is wearing a gray hoodie and black pants, and the third is wearing a navy camouflage shirt and blue pants. All three are wearing glasses and looking to the side. Based on the description of the second editing instruction, it can be determined that the spatial relationship between the target object (three men) and the background object (blue sky and white clouds) in the first image is reasonable; that is, the text describes three young men dressed casually sitting on the ground against a backdrop of blue sky and white clouds. Furthermore, the second editing instruction provides a detailed description of the target object and the background object.

[0126] S503. Input the second editing command, the foreground image of the first image, and the foreground mask into the second model to obtain the second image.

[0127] After receiving the second editing instruction, the foreground image of the first image, and the foreground mask, image processing techniques can be used to generate the background image of the first image based on the information in the second editing instruction. Furthermore, image fusion and other methods can be used to fuse the foreground and background images of the first image to obtain the second image. The content presented in the second image corresponds to the content described in the second editing instruction.

[0128] In this embodiment, a second image can be generated using a second model based on second editing instructions. See also... Figure 5 After understanding the process, the second editing command, the foreground image, and the foreground mask are input into the second model, and the second model can output the second image.

[0129] In this embodiment, the second model is used to perform background editing on the original image according to the second editing instructions. In one implementation, the second model can be an image editing model. This embodiment can use any existing model as long as it can achieve the image editing goal; therefore, the choice of the second model is not limited.

[0130] The following will use specific examples, combined with Figure 8 The results of image background editing are described below. For example... Figure 8 As shown, the second editing command, the foreground image of the first image, and the foreground mask are input into the second model, and the second model can output the second image 801. In the illustration of the second image 801, the target object car in the first image is located on flat sand. Furthermore, the background of the second image may include image elements such as a clear blue sky, sand dunes, and the car's shadow cast by sunlight on the sand, all generated by the second model.

[0131] In this embodiment, the foreground image and foreground mask of the first image are obtained through image segmentation to determine the target pixel region (foreground target) and non-target pixel region (background) in the first image, which facilitates maintaining the integrity and quality of the foreground during subsequent editing. Simultaneously, rewriting editing instructions using information from the foreground image makes the instructions more targeted and refined, thereby improving the accuracy of the editing effect.

[0132] When rewriting the first editing instruction input by the user, prompt information is generated based on the first editing instruction and the foreground image. The first condition restricts the background of the image to be adjusted according to the first image to be compatible with the foreground image of the first image. This ensures that the image content described by the second editing instruction output by the first model is reasonable as a whole, and thus makes the edited image look natural and harmonious.

[0133] Next, the second image is obtained by inputting the second editing command, the foreground image, and the foreground mask into the second model. The use of the foreground image and the foreground mask ensures that the foreground part is not subjected to unnecessary modifications during processing, thus preserving its quality and details.

[0134] In this way, image editing functions can be implemented while ensuring the generation of reasonable images, without the need for joint training of the existing first and second models, or complex adjustments to the architecture and training process of the existing models. This reduces the complexity of image editing methods and improves their adaptability and effectiveness.

[0135] Based on the above introduction, the actual effect of the image editing method provided in this application embodiment can be verified through multiple test images. The foreground targets in the test images include objects, single people, and multiple people. In the specific implementation process, the second image generated by the image editing method can be subjectively evaluated from three dimensions: image material, image quality, and visual aesthetics. Regarding the evaluation of image material, the standard mainly focuses on whether the generated material matches the content requirements (including subject and object, relationship, scene, etc.) and whether the generated material elements (such as subject, environment, etc.) meet the specific needs of the content. Regarding the evaluation of image quality, the standard mainly focuses on the detail representation of the image, the presence of noise, and the presence of traces of manual editing. Furthermore, regarding the evaluation of visual aesthetics, the standard mainly focuses on whether the image's color tone is balanced, whether the composition is harmonious, and whether there is obvious distortion in the image.

[0136] In the final experimental results, for test images with objects as the foreground target, the score without instruction rewriting was 4.09, while the score using the method described in this solution was 4.55. For test images with a single person as the foreground target, the score without instruction rewriting was 3.87, while the score using the method described in this solution was 4.13. Similarly, for test images with multiple people as the foreground target, the score without instruction rewriting was 3.86, while the score using the method described in this solution was 4.16. Statistically, the average score without instruction rewriting was 3.94, while the average score using the method described in this solution was 4.28. Based on the above statistical data, it can be determined that the image editing method proposed in this application demonstrates good editing results.

[0137] It should be noted that the module names involved in the embodiments of this application can all be defined as other names, as long as they can achieve the function of each module, and no specific restrictions are placed on the module names.

[0138] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0139] The image editing method according to the embodiments of this application has been described above. The apparatus for performing the above method provided in the embodiments of this application is described below. Those skilled in the art will understand that the methods and apparatus can be combined with and referenced by each other, and the related apparatus provided in the embodiments of this application can perform the steps in the above image editing method.

[0140] The image editing method provided in this application can be applied to electronic devices with image processing capabilities. Electronic devices include terminal devices, and the specific device form of the terminal device can be referred to the above-described related information, which will not be repeated here.

[0141] In one implementation, this application provides an electronic device. Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application.

[0142] like Figure 9 As shown, the electronic device 900 includes: a processor 901 and a memory 902; the memory 902 stores computer execution instructions; the processor 901 executes the computer execution instructions stored in the memory 902, causing the electronic device 900 to perform the above-described method.

[0143] When the memory 902 is set up independently, the electronic device also includes a bus 903 for connecting the memory 902 and the processor 901.

[0144] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.

[0145] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0146] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0147] This application provides a computer program product, which includes a computer program that, when run, causes a computer to perform the above-described method.

[0148] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0149] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.

Claims

1. An image editing method, characterized in that, include: Obtain a first editing instruction input for the first image, the first editing instruction being used to instruct the background of the first image to be adjusted; Based on the first image and the first editing instruction, first instruction information is generated, the first instruction information including the first image and the first editing instruction, the first instruction information being used to instruct the first editing instruction to be rewritten based on the first image; The first instruction information is input into the first model to obtain the second editing instruction output by the first model. The second editing instruction is used to instruct the background of the first image to be adjusted. A second image is generated based on the second editing instruction, the first image, and the second model, wherein the background of the second image is different from that of the first image.

2. The method according to claim 1, characterized in that, The step of generating first instruction information based on the first image and the first editing instruction includes: The first image is segmented to obtain the foreground image of the first image; The first instruction information is generated based on the foreground image of the first image and the first editing instruction; The first instruction information is used to instruct the first editing instruction to be rewritten based on the foreground image of the first image.

3. The method according to claim 1 or 2, characterized in that, The first indication information is further used to indicate a first condition for rewriting the first editing instruction based on the foreground image of the first image, the first condition being used to restrict the background of the image to be adjusted by the rewritten editing instruction to be adapted to the foreground image of the first image.

4. The method according to claim 2, characterized in that, After performing the image segmentation process on the first image, the method further includes obtaining a foreground mask of the first image.

5. The method according to claim 4, characterized in that, The step of generating a second image based on the second editing instruction, the first image, and the second model, wherein the background of the second image is different from that of the first image, includes: The second editing command, the foreground image of the first image, and the foreground mask of the first image are input into the second model to obtain the second image output by the second model.

6. The method according to claim 5, characterized in that, The background of the second image is consistent with the background of the image adjusted as indicated by the second editing instruction.

7. The method according to any one of claims 1-6, characterized in that, The first model is a multimodal large language model, and the second model is an image editing model.

8. The method according to any one of claims 1-7, characterized in that, The step of obtaining the first editing instruction input for the first image includes: In response to a user operation applied to the first interface of the first application, an input box is displayed, and the first image is displayed on the first interface; Based on the text input operation performed in the input box, the first editing instruction input for the first image is obtained.

9. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 8.

10. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 8.