An image editing method, device, electronic device, storage medium and program product

By determining the embedded feature vector and key point vector of the image to be edited, and combining attribute control encoding to generate the edited image, the geometric position control and integrity problems of image editing in the prior art are solved, and highly accurate multimodal image editing is achieved.

CN122156370APending Publication Date: 2026-06-05SHENZHEN DESAY SV AUTOMOTIVE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN DESAY SV AUTOMOTIVE CO LTD
Filing Date
2026-02-11
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing image editing techniques cannot accurately control geometric positions, resulting in easily shifted images. Furthermore, local area editing cannot ensure integrity, and multimodal control and image support are insufficient.

Method used

By determining the embedded feature vector and key point vector of the image to be edited, and combining attribute control encoding to generate the edited image, the identity features of the target object are extracted and the position is controlled by the diffusion inversion generation operation, thus ensuring the integrity of the image.

Benefits of technology

It improves the accuracy of image editing, reduces the occurrence of geometric misalignment, and realizes image editing under multimodal control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156370A_ABST
    Figure CN122156370A_ABST
Patent Text Reader

Abstract

An image editing method and device, electronic equipment, storage medium and program product are disclosed. The specific implementation scheme comprises: determining a to-be-edited image and an embedding feature vector corresponding to the to-be-edited image; extracting a key point vector of the to-be-edited image; determining control information and an attribute control code corresponding to the control information; and generating an edited image based on the to-be-edited image, the embedding feature vector, the key point vector and the attribute control code. By determining the embedding feature vector of the to-be-edited image, the identity feature of the target object is extracted, the position of the target object is controlled through the key point vector, the occurrence of geometric misplacement is reduced, the multi-modal control is realized through the control information, the editing of the to-be-edited image is completed under the condition of ensuring the integrity of the image, and the accuracy of image editing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an image editing method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] Currently, image editing techniques for faces mainly rely on generative models (such as diffusion models or generative adversarial networks). Representative solutions include: using multimodal large models to understand text prompts for image generation, which can support text commands such as "add a hat" and "change hairstyle"; or local region editing based on diffusion models + image masks.

[0003] However, methods using large multimodal models often fail to accurately control geometric positions (hats frequently shift or obscure facial features), while local region editing cannot ensure integrity, resulting in easily shifted images. Furthermore, existing methods can only accept text-based input, and the edited images can only be visible light images, exhibiting significant shortcomings in multimodal control and image support. Summary of the Invention

[0004] This invention provides an image editing method, apparatus, electronic device, storage medium, and program product to complete image editing through control information of multiple modalities, thereby improving the accuracy of image editing.

[0005] According to one aspect of the present invention, an image editing method is provided, comprising: Determine the image to be edited and the embedded feature vector corresponding to the image to be edited, wherein the image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object; Extract the key point vectors of the image to be edited, wherein the key point vectors include information about the face of the target object in the image to be edited; Determine control information and the attribute control code corresponding to the control information. The control information includes information for editing the face of the target object in the image to be edited. The control information includes text and / or information generated from reference images. Based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding, an edited image is generated. The edited image includes an image obtained by editing the face of the target object in the image to be edited.

[0006] According to another aspect of the present invention, an image editing apparatus is provided, comprising: The first determining module is used to determine the image to be edited and the embedded feature vector corresponding to the image to be edited, wherein the image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object; The extraction module is used to extract key point vectors from the image to be edited, wherein the key point vectors include facial information of the target object in the image to be edited; The second determining module is used to determine control information and the attribute control code corresponding to the control information. The control information includes information set for editing the face of the target object in the image to be edited. The control information includes text and / or information generated from reference images. The generation module is used to generate an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding. The edited image includes an image obtained by editing the face of the target object in the image to be edited.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image editing method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image editing method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the image editing method described in any embodiment of the present invention.

[0010] The technical solution of this invention involves determining the image to be edited and the embedded feature vector corresponding to the image to be edited; extracting the key point vector of the image to be edited; determining control information and the attribute control code corresponding to the control information; and generating an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control code. By determining the embedded feature vector of the image to be edited, the identity features of the target object are extracted; the position control of the target object is achieved through the key point vector, reducing the occurrence of geometric misalignment; and the control information achieves multimodal control. This process completes the editing of the image to be edited while ensuring image integrity, thus improving the accuracy of image editing.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of an image editing method provided according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of an image generation method provided according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of an image editing device according to Embodiment 3 of the present invention; Figure 4 This is a block diagram of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart of an image editing method according to Embodiment 1 of the present invention. This embodiment is applicable to situations involving image editing. The method can be executed by an image editing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Determine the image to be edited and the embedded feature vector corresponding to the image to be edited.

[0017] The image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object.

[0018] In this embodiment, the image to be edited can be understood as an image containing a target object. The image to be edited can be an unprocessed visible light image or an infrared image. The embedded feature vector can be understood as a vector indicating the identity characteristics of the target object. The embedded feature vector can be used to ensure the consistency of the identity characteristics of the target object during the editing process of the image to be edited. The target object can be understood as an object contained in the image to be edited, and the image to be edited may contain the face of the target object.

[0019] Specifically, in specific scenarios (such as in-vehicle facial recognition or security applications), a face extraction model can be trained to extract features from the image to be edited. For the image to be edited, the face extraction model can extract the identity features of the target object. For example, it can extract identifiers indicating the target object's identity and quantize these identifiers into an embedded feature vector using the face extraction model.

[0020] For example, the image to be edited is represented as (I_{src}), and the embedded feature vector (E_{id}) of the image to be edited is extracted using a face extraction model trained in a vertical scenario.

[0021] S120. Extract the key point vector of the image to be edited.

[0022] The key point vector includes information about the face of the target object in the image to be edited.

[0023] In this embodiment, the key point vector can be understood as information indicating the key points of the face of the target object in the image to be edited. The key point vector can be used to determine the accurate position and geometric alignment of each element of the target object's face.

[0024] Specifically, keypoint vectors can be extracted from the image to be edited using a facial extraction model. For example, keypoint vectors can be obtained by extracting information about each keypoint on the face of the target object in the image to be edited. This can be achieved by dividing the target object's face into a triangle-based polygonal mesh, where the vertices of the polygonal mesh are the keypoints.

[0025] For example, the key points of the target object's face can include key points of various parts such as the head, hair, and eyes. The key point vector of the image to be edited is represented as (K_{face}), and the key point vector can be the prior geometric information of the target object's face.

[0026] S130. Determine the control information and the attribute control code corresponding to the control information.

[0027] The control information includes settings for editing the face of the target object in the image to be edited, and the control information includes information generated from text and / or reference images.

[0028] In this embodiment, control information can be understood as input information for editing a target object in the image to be edited. Control information can be for editing the face of the target object, such as adding different elements to the face. Attribute control encoding can be understood as a sequence generated by an encoder based on the control information; attribute control encoding can indicate the characteristics of the control information.

[0029] Specifically, it can receive control information input as an image to be edited. This control information can be text-based or reference image-based. For different types of control information, different types of encoders can be used to encode the control information into attribute control codes.

[0030] For example, text-based control information could include: adding a baseball cap to the target object, changing the target object's hairstyle to curly hair, etc. Image-based control information could include: image templates containing information such as hats, hairstyles, and glasses. Attribute control encoding can be represented as (C_{attr}).

[0031] S140. Based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding, generate the edited image.

[0032] The edited image includes an image obtained by editing the face of the target object in the image to be edited.

[0033] In this embodiment, the edited image can be understood as the image obtained by editing the face of the target object in the image to be edited. For example, the edited image can be the image after adding a hat to the face of the target object in the image to be edited. The edited image is a virtual image generated based on the image to be edited.

[0034] Specifically, a mask can first be generated based on keypoint vectors and attribute control encoding. This mask can be used to segment the parts of the target object's face that need to be edited, such as the head region, hair region, or eye region. Finally, the image to be edited, the embedded feature vectors, keypoint vectors, attribute control encoding, and the mask are used as conditions for the diffusion inversion generation operation to output the edited image.

[0035] The technical solution of this invention involves determining the image to be edited and the embedded feature vector corresponding to the image to be edited; extracting the key point vector of the image to be edited; determining control information and the attribute control code corresponding to the control information; and generating an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control code. By determining the embedded feature vector of the image to be edited, the identity features of the target object are extracted; the position control of the target object is achieved through the key point vector, reducing the occurrence of geometric misalignment; and the control information achieves multimodal control. This process completes the editing of the image to be edited while ensuring image integrity, thus improving the accuracy of image editing.

[0036] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.

[0037] In one embodiment, determining the control information and the attribute control code corresponding to the control information includes: In response to input operations on the input interface, determine control information; The control information is encoded to obtain the attribute control code corresponding to the control information.

[0038] In this embodiment, the input interface can be understood as a page for inputting control information. On the input interface, control information can be entered in text or as reference images. The input operation can be understood as the operation of inputting control information.

[0039] Specifically, in response to input operations on the input interface, control information is determined, and the type of control information is identified. When the control information is text, a text encoder can be used to encode the control information, extract the text encoding vector, and use this text encoding vector as the attribute control encoding. When the control information is a reference image, an image encoder can be used to encode the control information, extract the image encoding vector, and use this image encoding vector as the attribute control encoding.

[0040] In one embodiment, determining the image to be edited and the embedded feature vector corresponding to the image to be edited includes: Determine the image to be edited, and the image type of the image to be edited; Based on the image type, extract the initial embedding feature vector of the image to be edited; Determine the feature space difference between the initial embedded feature vector and the image to be edited, wherein the feature space difference is used to constrain the information in the initial embedded feature vector related to the identity features of the target object; If the difference in the feature space is less than a set difference threshold, the initial embedding feature vector is determined as the embedding feature vector of the image to be edited.

[0041] In this embodiment, the image type can be understood as the type of the image to be edited, which may include visible light image type and infrared image type. The initial embedded feature vector can be understood as the feature vector of the image to be edited extracted by the face extraction model. The feature space difference can be understood as constraining the values ​​of information related to the identity features of the target object in the initial embedded feature vector. The feature space difference can be used to ensure the consistency between the identity features of the target object in the initial embedded feature vector and the target object in the image to be edited. The set difference threshold can be understood as a threshold set to determine the consistency between the target object in the initial embedded feature vector and the target object in the image to be edited.

[0042] Specifically, the image to be edited and its image type are determined. Image type can include visible light images or infrared images. Based on different image types, a face extraction model can be used to extract the initial embedding feature vector of the image to be edited. If the image to be edited is a visible light image, the initial embedding feature vector can be extracted directly. If it is an infrared image, the image to be edited needs to be mapped before extracting the initial embedding feature vector. The feature space difference between the initial embedding feature vector and the image to be edited is calculated. The larger the feature space difference value, the greater the difference between the initial embedding feature vector and the image to be edited. If the feature space difference is less than a set difference threshold, the initial embedding feature vector is determined as the embedding feature vector of the image to be edited.

[0043] For example, the difference threshold can be set to 0.35 cosine distance. Therefore, if the difference in the feature space is less than 0.35 cosine distance, the initial embedded feature vector is determined as the embedded feature vector of the image to be edited.

[0044] Optionally, extracting the initial embedding feature vector of the image to be edited based on the image type includes: If the image type indicates that the image to be edited is a visible light image, an initial embedding feature vector of the image to be edited is extracted based on a face extraction model; If the image type indicates that the image to be edited is an infrared image, the image to be edited is mapped to obtain a mapped image, and the initial embedding feature vector of the mapped image is extracted based on the face extraction model.

[0045] In this embodiment, the mapped image can be understood as the image obtained by mapping an infrared image when the image to be edited is an infrared image.

[0046] Specifically, when the image type indicates that the image to be edited is a visible light image, the initial embedding feature vector of the image to be edited is extracted based on the face extraction model. When the image type indicates that the image to be edited is an infrared image, the image to be edited is mapped, which can map the infrared image and the visible light image to a unified embedding space to obtain the mapped image, and the initial embedding feature vector of the mapped image is extracted based on the face extraction model.

[0047] For example, the loss functions of the face extraction model are as follows: Identity consistency loss L_{ID} = 1 - \cos(E_{id}(I_{src}), E_{id}(I_{gen})), Geometric consistency loss L_{geom} = |K_{face}(I_{src}) - K_{face}(I_{gen})|, and visual naturalness loss (L_{adv}) can be learned based on the discriminator to achieve a natural blending effect. The total loss function L = L_{recon} + \lambda_1 L_{ID} + \lambda_2 L_{geom} + \lambda_3 L_{adv}.

[0048] In one embodiment, extracting the key point vectors of the image to be edited includes: Identify at least one key point of the face of the target object in the image to be edited; For each key point, extract the key point information corresponding to the key point. The key point information includes the key point coordinates, the pose angle of the part corresponding to the key point in the image to be edited, and the situation where the key point is occluded in the image to be edited. The key point information corresponding to each key point is quantized into a key point vector of the image to be edited.

[0049] In this embodiment, key point information can be understood as the information of each key point in the image to be edited. Key point information may include the coordinates of the key point, the pose angle of the part corresponding to the key point in the image to be edited, and the situation where the key point is occluded in the image to be edited.

[0050] Specifically, at least one keypoint of the face of the target object in the image to be edited is identified. This keypoint can include keypoints of various parts such as the head, hair, and eyes. Therefore, the keypoint vector can be used to determine the accurate position and geometric alignment of elements such as hats, glasses, and hairstyles. For each keypoint, the keypoint coordinates, the pose angle of the corresponding part in the image to be edited, and the occlusion status of the keypoint in the image to be edited can be extracted using a facial extraction model. The extracted information is used as keypoint information. Finally, the keypoint information corresponding to each keypoint is quantized into a keypoint vector of the image to be edited.

[0051] Example 2 Figure 2 This is a flowchart of an image generation method according to Embodiment 2 of the present invention. This embodiment focuses on the method for generating edited images as described in the above embodiments. Figure 2 As shown, the method includes: S210. Determine the image to be edited and the embedded feature vector corresponding to the image to be edited.

[0052] S220. Extract the key point vector of the image to be edited.

[0053] S230. Determine the control information and the attribute control code corresponding to the control information.

[0054] S240. Based on the key point vector and the attribute control encoding, generate an edit mask.

[0055] The edit mask is associated with the location of the target object indicated by the attribute control encoding.

[0056] In this embodiment, the edit mask can be understood as the binarized information of the part of the target object indicated by the attribute control encoding.

[0057] For example, based on keypoint vectors and attribute control encoding, an edit mask (M_{edit}) can be automatically generated by a generator in a generative adversarial network. For instance, when the control information is to add a baseball cap to a target object, the attribute control encoding indicates the head as the part of the target object; therefore, the head is visible in the binarized image corresponding to the generated edit mask.

[0058] Optionally, generating the edit mask based on the keypoint vector and the attribute control encoding includes: Determine the set area indicated by the attribute control code, the set area including the set area of ​​the target object's face; Determine at least one key point corresponding to the key point vector; For each key point, if the key point is located in the area corresponding to the set part, the pixel value corresponding to the key point is set to a first value; otherwise, the pixel value corresponding to the key point is set to a second value. The set of pixel values ​​corresponding to each key point is determined as the edit mask.

[0059] In this embodiment, the defined location can be understood as the facial region of the target object indicated by the attribute control code. A pixel can be understood as the smallest physical point on an image. A pixel may contain color information and is the basic unit constituting an image. Each pixel has a specific position in the image, which can be determined by coordinates. The first value can be understood as the pixel value corresponding to the defined key point, indicating that the pixel value is visible. The second value can be understood as the pixel value corresponding to the defined key point, indicating that the pixel value is invisible.

[0060] Specifically, the designated area indicated by the attribute control code is determined, such as the head region, hair region, or eye region. At least one keypoint corresponding to each keypoint vector is identified. For each keypoint, if it lies within the region corresponding to the designated area, it indicates that the keypoint is the area indicated by the control information, and therefore the pixel value corresponding to the keypoint is set to a first value; otherwise, it indicates that the keypoint is not the area indicated by the control information, and therefore the pixel value corresponding to the keypoint is set to a second value. Finally, the set of pixel values ​​corresponding to all keypoints is determined as the edit mask.

[0061] For example, for each keypoint, if the keypoint is located in the area corresponding to the set part, the pixel value corresponding to the keypoint is set to a first value of 1; otherwise, the pixel value corresponding to the keypoint is set to a second value of 0. Therefore, the edit mask is a set of 0s and 1s. The binary image formed by the edit mask is a special kind of digital image, characterized in that each pixel in the image has only two values, such as 0 and 1. When displayed, the binary image can be represented as an image composed of black and white.

[0062] S250. Perform diffusion inversion generation operation on the image to be edited, the embedded feature vector, the key point vector, the attribute control code, and the editing mask to obtain the edited image.

[0063] For example, using the image to be edited (I_{src}), the embedded feature vector (E_{id}), the key point vector (K_{face}), the attribute control encoding (C_{attr}), and the editing mask (M_{edit}) as conditions, a diffusion inversion generation operation is performed to output the edited image (I_{gen}).

[0064] The technical solution of this invention generates an editing mask based on the keypoint vector and the attribute control code; a diffusion inversion generation operation is performed on the image to be edited, the embedded feature vector, the keypoint vector, the attribute control code, and the editing mask to obtain the edited image. By determining the embedded feature vector and keypoint vector of the image to be edited, accurate segmentation of the parts of the target object to be edited is achieved, reducing the occurrence of geometric misalignment, and completing the editing of the image to be edited while ensuring image integrity, thus improving the accuracy of image editing.

[0065] Example 3 Figure 3 This is a schematic diagram of the structure of an image editing device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes: The first determining module 310 is used to determine the image to be edited and the embedded feature vector corresponding to the image to be edited, wherein the image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object; Extraction module 320 is used to extract key point vectors from the image to be edited, the key point vectors including facial information of the target object in the image to be edited; The second determining module 330 is used to determine control information and the attribute control code corresponding to the control information. The control information includes information set for editing the face of the target object in the image to be edited. The control information includes text and / or information generated from reference images. The generation module 340 is used to generate an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding. The edited image includes an image obtained by editing the face of the target object in the image to be edited.

[0066] The image editing apparatus provided in this embodiment of the invention determines the image to be edited and the embedded feature vector corresponding to the image to be edited through a first determining module; extracts the key point vector of the image to be edited through an extraction module; determines control information and the attribute control code corresponding to the control information through a second determining module; and generates an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control code through a generating module. Through the cooperation between these modules, the embedded feature vector of the image to be edited is determined, enabling the extraction of the identity features of the target object; the key point vector enables positional control of the target object, reducing geometric misalignment; and the control information enables multimodal control. This completes the editing of the image to be edited while ensuring image integrity, thus improving the accuracy of image editing.

[0067] In one embodiment, the generation module 340 includes: A generation unit is configured to generate an edit mask based on the key point vector and the attribute control code, wherein the edit mask is related to the part of the target object indicated by the attribute control code; The inversion unit is used to perform a diffusion inversion generation operation on the image to be edited, the embedded feature vector, the key point vector, the attribute control code, and the editing mask to obtain the edited image.

[0068] In one embodiment, the generating unit is specifically used for: Determine the set area indicated by the attribute control code, the set area including the set area of ​​the target object's face; Determine at least one key point corresponding to the key point vector; For each key point, if the key point is located in the area corresponding to the set part, the pixel value corresponding to the key point is set to a first value; otherwise, the pixel value corresponding to the key point is set to a second value. The set of pixel values ​​corresponding to each key point is determined as the edit mask.

[0069] In one embodiment, the second determining module 330 is specifically used for: In response to input operations on the input interface, determine control information; The control information is encoded to obtain the attribute control code corresponding to the control information.

[0070] In one embodiment, the first determining module 310 includes: The first determining unit is used to determine the image to be edited and the image type of the image to be edited; An extraction unit is used to extract the initial embedding feature vector of the image to be edited based on the image type; The second determining unit is used to determine the feature space difference between the initial embedded feature vector and the image to be edited, wherein the feature space difference is used to constrain the information in the initial embedded feature vector related to the identity features of the target object; The third determining unit is used to determine the initial embedded feature vector as the embedded feature vector of the image to be edited when the feature space difference is less than a set difference threshold.

[0071] In one embodiment, the extraction unit is specifically used for: If the image type indicates that the image to be edited is a visible light image, an initial embedding feature vector of the image to be edited is extracted based on a face extraction model; If the image type indicates that the image to be edited is an infrared image, the image to be edited is mapped to obtain a mapped image, and the initial embedding feature vector of the mapped image is extracted based on the face extraction model.

[0072] In one embodiment, the extraction module 320 is specifically used for: Identify at least one key point of the face of the target object in the image to be edited; For each key point, extract the key point information corresponding to the key point. The key point information includes the key point coordinates, the pose angle of the part corresponding to the key point in the image to be edited, and the situation where the key point is occluded in the image to be edited. The key point information corresponding to each key point is quantized into a key point vector of the image to be edited.

[0073] The image editing device provided in this embodiment of the invention can execute the image editing method provided in any embodiment of the invention. Through the cooperation and coordination between the modules, the image editing is completed, and it has the corresponding functional modules and beneficial effects of the execution method.

[0074] Example 4 According to embodiments of the present invention, the present invention also provides an electronic device, a computer-readable storage medium, and a computer program product.

[0075] Figure 4 This is a block diagram of an electronic device according to Embodiment 4 of the present invention, which implements the image editing method described in the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0076] like Figure 4 As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412 or a random access memory (RAM) 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0077] Multiple components in the electronic device are connected to the I / O interface 415, including: an input unit 416, such as a keyboard, mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a disk, optical disk, etc.; and a communication unit 419, such as a network card, modem, wireless transceiver, etc. The communication unit 419 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0078] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as image editing methods.

[0079] In some embodiments, the image editing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the image editing method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to perform the image editing method by any other suitable means (e.g., by means of firmware).

[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0085] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0086] In some embodiments, the computer program product includes a computer program that, when executed by a processor, implements the image editing method provided in the embodiments of the present invention.

[0087] The technical solution of this invention provides an image editing method, apparatus, electronic device, storage medium, and program product. It involves determining the image to be edited and its corresponding embedded feature vector; extracting keypoint vectors from the image to be edited; determining control information and its corresponding attribute control code; and generating an edited image based on the image to be edited, the embedded feature vector, the keypoint vector, and the attribute control code. By determining the embedded feature vector of the image to be edited, the identity features of the target object are extracted; the position of the target object is controlled through the keypoint vector, reducing geometric misalignment; and the control information enables multimodal control. This process completes the editing of the image to be edited while ensuring image integrity, thus improving the accuracy of image editing.

[0088] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.

[0089] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image editing method, characterized in that, include: Determine the image to be edited and the embedded feature vector corresponding to the image to be edited, wherein the image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object; Extract the key point vectors of the image to be edited, wherein the key point vectors include information about the face of the target object in the image to be edited; Determine control information and the attribute control code corresponding to the control information. The control information includes information for editing the face of the target object in the image to be edited. The control information includes text and / or information generated from reference images. Based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding, an edited image is generated. The edited image includes an image obtained by editing the face of the target object in the image to be edited.

2. The method according to claim 1, characterized in that, The process of generating the edited image based on the embedded feature vector, the key point vector, and the attribute control encoding includes: Based on the key point vector and the attribute control code, an edit mask is generated, the edit mask being related to the part of the target object indicated by the attribute control code; The image to be edited, the embedded feature vector, the key point vector, the attribute control encoding, and the editing mask are subjected to diffusion inversion generation to obtain the edited image.

3. The method according to claim 2, characterized in that, The process of generating an edit mask based on the keypoint vector and the attribute control encoding includes: Determine the set area indicated by the attribute control code, the set area including the set area of ​​the target object's face; Determine at least one key point corresponding to the key point vector; For each key point, if the key point is located in the area corresponding to the set part, the pixel value corresponding to the key point is set to a first value; otherwise, the pixel value corresponding to the key point is set to a second value. The set of pixel values ​​corresponding to each key point is determined as the edit mask.

4. The method according to claim 1, characterized in that, The determination of control information and the attribute control code corresponding to the control information includes: In response to input operations on the input interface, determine control information; The control information is encoded to obtain the attribute control code corresponding to the control information.

5. The method according to claim 1, characterized in that, The process of determining the image to be edited and the embedded feature vector corresponding to the image to be edited includes: Determine the image to be edited, and the image type of the image to be edited; Based on the image type, extract the initial embedding feature vector of the image to be edited; Determine the feature space difference between the initial embedded feature vector and the image to be edited, wherein the feature space difference is used to constrain the information in the initial embedded feature vector related to the identity features of the target object; If the difference in the feature space is less than a set difference threshold, the initial embedding feature vector is determined as the embedding feature vector of the image to be edited.

6. The method according to claim 5, characterized in that, The step of extracting the initial embedding feature vector of the image to be edited based on the image type includes: If the image type indicates that the image to be edited is a visible light image, an initial embedding feature vector of the image to be edited is extracted based on a face extraction model; If the image type indicates that the image to be edited is an infrared image, the image to be edited is mapped to obtain a mapped image, and the initial embedding feature vector of the mapped image is extracted based on the face extraction model.

7. The method according to claim 1, characterized in that, The extraction of the key point vector of the image to be edited includes: Identify at least one key point of the face of the target object in the image to be edited; For each key point, extract the key point information corresponding to the key point. The key point information includes the key point coordinates, the pose angle of the part corresponding to the key point in the image to be edited, and the situation where the key point is occluded in the image to be edited. The key point information corresponding to each key point is quantized into a key point vector of the image to be edited.

8. An image editing device, characterized in that, include: The first determining module is used to determine the image to be edited and the embedded feature vector corresponding to the image to be edited, wherein the image to be edited includes an image containing a target object, and the embedded feature vector indicates the identity features of the target object; The extraction module is used to extract key point vectors from the image to be edited, wherein the key point vectors include facial information of the target object in the image to be edited; The second determining module is used to determine control information and the attribute control code corresponding to the control information. The control information includes information set for editing the face of the target object in the image to be edited. The control information includes text and / or information generated from reference images. The generation module is used to generate an edited image based on the image to be edited, the embedded feature vector, the key point vector, and the attribute control encoding. The edited image includes an image obtained by editing the face of the target object in the image to be edited.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image editing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image editing method according to any one of claims 1-7.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the image editing method according to any one of claims 1-7.