Image processing method and device and computer program product

By determining the target area based on the prompt text in image processing and expanding the mask according to the texture complexity, the problem of poor edge processing in image restoration is solved, and better restoration effect and efficiency are achieved.

CN120689247APending Publication Date: 2025-09-23BEIJING AUTONAVI YUNMAP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510757980.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the process of removing and repairing specific objects in images, there is a problem of poor edge processing effect in the repair area, and inaccurate positioning of specific objects leads to low processing efficiency.

Method used

The target area is determined from the image to be processed based on the hint text of the target object, an initial mask is generated, and the texture complexity of the area corresponding to the initial mask is expanded to obtain the target mask, and finally the repair processing is performed based on the target mask.

Benefits of technology

The repair quality of the mask edge is improved, the overall consistency and naturalness of the image are maintained, the problems of jagged edges and residual content in the repair area in the existing technology are solved, and the processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689247A_ABST
    Figure CN120689247A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and device and a computer program product, and relates to the field of image processing. According to the specific implementation scheme, the method comprises the steps of determining a target area where a target object is located from a to-be-processed image based on a prompt text for the target object; performing segmentation processing on the to-be-processed image based on the target area to generate an initial mask of the target object; based on the texture complexity of the corresponding region of the initial mask, performing expansion processing on the initial mask to obtain a target mask; and repairing the to-be-processed image based on the target mask to obtain a repaired image. In the scheme, the initial mask is expanded based on the texture complexity of the area corresponding to the initial mask, so that the initial mask can be expanded to a reasonable amplitude, and the repairing quality of the edge of the mask can be improved based on the target mask obtained after the expansion, so that a relatively good repairing effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, device, and computer program product. Background Art

[0002] With the rapid development of digital image processing technology, the demand for image content editing is becoming increasingly diverse. Removing and repairing specific objects in images has also become an important application direction. For example, it is used in scenarios such as removing irrelevant logos or distracting objects in images.

[0003] The inventors of the present application have discovered that when removing and repairing specific objects in an image, there is a problem of poor edge processing effect on the repaired area. Summary of the Invention

[0004] The present application provides an image processing method, apparatus, and computer program product.

[0005] This application provides the following solutions:

[0006] According to a first aspect, there is provided an image processing method, the method comprising:

[0007] Based on the prompt text for the target object, determining the target area where the target object is located from the image to be processed;

[0008] Segment the image to be processed based on the target area to generate an initial mask of the target object;

[0009] Based on the texture complexity of the area corresponding to the initial mask, the initial mask is expanded to obtain the target mask;

[0010] The image to be processed is repaired based on the target mask to obtain a repaired image.

[0011] According to a second aspect, there is provided an image processing apparatus, the apparatus comprising:

[0012] The target detection module is used to determine the target area where the target object is located from the image to be processed based on the prompt text for the target object;

[0013] The target segmentation module is used to segment the image to be processed based on the target area and generate an initial mask of the target object;

[0014] A mask expansion module is used to expand the initial mask based on the texture complexity of the area corresponding to the initial mask to obtain the target mask;

[0015] The image restoration module is used to restore the image to be processed based on the target mask to obtain a restored image.

[0016] According to a third aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned image processing method when executed by a processor.

[0017] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0018] In this application, the target region of the target object is determined from the image to be processed based on the prompt text for the target object; the image to be processed is segmented based on the target region to generate an initial mask of the target object; the initial mask is expanded based on the texture complexity of the region corresponding to the initial mask to obtain a target mask; and the image to be processed is then inpainted based on the target mask to obtain a restored image. In this solution, by inflating the initial mask based on the texture complexity of the region corresponding to the initial mask, the initial mask is expanded to a reasonable extent. Inpainting based on the target mask obtained by the expansion process can improve the quality of the mask edge restoration, thereby achieving a better image restoration effect.

[0019] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;

[0022] Figure 2 A flowchart of an image processing method provided in an embodiment of the present application;

[0023] Figure 3 A flowchart of a specific implementation of the image processing method provided in an embodiment of the present application;

[0024] Figure 4 This is a schematic diagram of the structure of a target detector provided in an embodiment of the present application;

[0025] Figure 5 Schematic diagram of the structure of the cross-modal alignment module in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of the overall flow of the image processing method provided in an embodiment of the present application;

[0027] Figure 7 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;

[0028] Figure 8 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0030] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0031] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0032] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0033] The inventors of this application discovered that when removing and repairing specific objects in an image, there is a problem with poor edge processing of the repaired area. Specifically, during the removal and repair process of specific objects in an image, the area of ​​the specific object needs to be masked. If the masking area is too small, it is easy to cause edge artifacts. If the masking area is too large, it is easy to cause the repair area to be too large, resulting in low processing efficiency.

[0034] In addition, the removal and restoration of specific objects in an image depends on the positioning of the specific objects. Therefore, how to improve the accuracy of positioning specific objects in an image has become an important technical issue when removing and restoring specific objects.

[0035] In view of this, this application provides a new idea. In order to facilitate the understanding of this application, the system architecture on which this application is based is first described. Figure 1 The following is an exemplary system architecture to which the embodiments of the present application can be applied: Figure 1 As shown in , the system architecture may include: a server side and a user side running on a user terminal.

[0036] The server and client are the two main components of an application service. The server, with the server as the primary hardware infrastructure, can include one or more software service modules. The client can be a client running on the user's terminal, a mini-program, or a web application running through a browser.

[0037] User terminals where the user end resides may include, but are not limited to, smart mobile terminals, wearable devices, and PCs (Personal Computers). Smart mobile devices may include mobile phones, tablets, PDAs (Personal Digital Assistants), and internet-connected car terminals. Wearable devices may include smart watches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and mixed reality devices (i.e., devices that support both VR and AR).

[0038] A server can be a standalone server, a server cluster, or even a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0039] As an optional method, a user can submit a prompt text and an image to be processed through the user terminal, which then sends the prompt text and the image to be processed to the server terminal. Upon receiving the prompt text and the image to be processed, the server terminal can use the image processing method provided in the embodiments of the present application to repair the image to be processed, generate a repaired image, and return the generated repaired image to the user terminal. This will be described in detail in subsequent embodiments.

[0040] As another optional method, the user can submit a prompt text and an image to be processed through the user terminal, and the user terminal can use the image processing method provided in the embodiment of the application to repair the image to be processed and generate a repaired image. The details will be detailed in subsequent embodiments.

[0041] It should be understood that Figure 1 The numbers of the server ends, user ends, and user terminals in the embodiment are merely illustrative. Any number of server ends, user ends, and user terminals may be provided according to implementation requirements.

[0042] Figure 2 This is a flow chart of the image processing method provided in the embodiment of the present application. The method can be performed by Figure 1 The system shown in the figure is executed on the user side or server side. Figure 2 As shown in , the method may include the following steps:

[0043] Step S210: Based on the prompt text for the target object, determine the target area where the target object is located from the image to be processed.

[0044] Step S220: Segment the image to be processed based on the target area to generate an initial mask of the target object.

[0045] Step S230: Based on the texture complexity of the area corresponding to the initial mask, the initial mask is expanded to obtain a target mask.

[0046] Step S240: performing restoration processing on the image to be processed based on the target mask to obtain a restored image.

[0047] As can be seen from the above process, this application determines the target area of ​​the target object from the image to be processed based on the prompt text for the target object; segments the image to be processed based on the target area to generate an initial mask of the target object; dilates the initial mask based on the texture complexity of the area corresponding to the initial mask to obtain the target mask; and then performs repair processing on the image to be processed based on the target mask to obtain a repaired image. In this solution, by dilating the initial mask based on the texture complexity of the area corresponding to the initial mask, the initial mask can be expanded to a reasonable extent. Repair processing based on the target mask obtained after the dilation process can improve the quality of repair of the mask edge, thereby achieving a better repair effect.

[0048] The following describes in detail the steps in the above process and the effects that can be further produced in conjunction with the embodiments.

[0049] First, the above step S210, namely "determining the target area where the target object is located from the image to be processed based on the prompt text for the target object" is described in detail in conjunction with the embodiment.

[0050] The target objects may include but are not limited to various irrelevant identifiers (such as image identifiers, text identifiers, etc.) or interference objects (such as watermarks, obstructions, etc.) in the image.

[0051] The hint text is text information describing the target object. As an example, the hint text can be the type or name of the target object.

[0052] As an example, Figure 3 A flowchart of a specific implementation of the image processing method provided in an embodiment of the present application.

[0053] Reference Figure 3 In the example, the target object is the panda in the image to be processed, and the prompt text can be "panda".

[0054] The prompt text can provide semantic information about the target object, and based on the semantics of the target object, the target area where the target object is located can be accurately located.

[0055] The above step S220 , namely “segmenting the image to be processed based on the target area to generate an initial mask of the target object”, is described in detail below with reference to an embodiment.

[0056] The target region can serve as a hint for segmenting the target object, guiding fine segmentation of the target region, thereby reducing interference from irrelevant areas and accurately segmenting the initial mask of the target object from the image to be processed. The initial mask can provide precise regional guidance for subsequent image restoration.

[0057] The above step S230 , namely “dilation processing of the initial mask based on the texture complexity of the area corresponding to the initial mask to obtain the target mask”, is described in detail below with reference to an embodiment.

[0058] The initial mask corresponding area is the area covered by the initial mask in the image to be processed.

[0059] By dilating the initial mask, the uncertain area at the edge of the initial mask can be effectively covered, so that the edge of the initial mask can transition smoothly, avoiding sharp edges, and providing richer contextual information for the repair process, which helps to improve the repair quality.

[0060] In this case, contextual information refers to the image content surrounding the masked area, which can be used to infer missing content. If image inpainting is performed directly based on the initial mask, the area corresponding to the initial mask cannot provide effective contextual information and cannot serve as a valid reference for image inpainting. In this solution, by dilating the initial mask, the area corresponding to the target mask obtained by dilation contains more contextual information, thus providing a valid reference for image inpainting.

[0061] Reference Figure 3In the example shown in , after the initial mask of the target object panda is expanded, the target mask obtained is as follows Figure 3 As shown in c in , the target mask corresponds to the panda pattern and a portion of the edge band surrounding the outline of the panda pattern. The edge band includes a portion of the leaf pattern that serves as the background of the panda. When performing image restoration, the color gradient and texture continuation of the edge band can be used as a reference to generate more natural restoration content, which provides more contextual information as a reference for image restoration.

[0062] When dilating the initial mask, if the resulting target mask is too small, edge artifacts may occur. If the resulting target mask is too large, the inpainted area may be too large, resulting in low processing efficiency. Therefore, it is necessary to reasonably determine the dilation amplitude of the initial mask.

[0063] The texture complexity in the area corresponding to the initial mask can characterize the image details contained therein. A high texture complexity indicates that the area corresponding to the initial mask contains more image details, and when performing the restoration process, it is necessary to rely on more contextual information in the image, that is, a larger dilation amplitude is required. Correspondingly, a low texture complexity indicates that the area corresponding to the initial mask contains fewer image details, and when performing the restoration process, it is not necessary to rely on too much contextual information in the image, that is, a smaller dilation amplitude is required. It can be seen that the texture complexity in the area corresponding to the initial mask has a certain correlation with the dilation amplitude of the initial mask. Therefore, the initial mask can be dilated based on the texture complexity of the area covered by the mask, and the dilation amplitude of the initial mask can be reasonably determined, thereby avoiding edge artifacts caused by a target mask that is too small, ensuring a good restoration effect on the mask edge, and avoiding an excessively large restoration area due to an excessively large target mask, thereby improving processing efficiency.

[0064] The above step S240 , namely “performing restoration processing on the image to be processed based on the target mask to obtain a restored image”, is described in detail below with reference to an embodiment.

[0065] The object to be processed is repaired based on the target mask, and filling content consistent with the original image background can be generated for the area of ​​the target mask.

[0066] The target mask can provide the contextual information in the image required for the restoration process. Restoration based on the target mask can improve the restoration quality of the mask edge, maintain the overall consistency and naturalness of the image, reduce restoration traces, and thus achieve better restoration effects.

[0067] Reference Figure 3 For example, Figure 3As shown in a, the target object is the panda in the image to be processed, and the prompt text can be "panda". After performing target detection on the image to be processed based on the prompt text, the target area can be determined. Figure 3 As shown in b, the part of the image area corresponding to the panda in the image to be processed will be detected as the target area, and the target area is represented by a detection frame. After determining the target area, the image to be processed can be segmented based on the target area to obtain an initial mask, and then the initial mask is expanded to obtain the target mask. The target mask is as follows Figure 3 As shown in c, the target mask corresponds to the panda pattern and a portion of the edge band surrounding the panda pattern outline. This edge band can provide reference information for color gradient and texture continuity for image restoration, which helps to generate a more natural restored image. Image restoration is performed on the image to be processed based on the target mask, and the restored image is obtained as shown in Figure 3 As shown in (d), the area corresponding to the target mask will be filled with new content to ensure a smooth transition with the background and ensure the overall consistency of the restored image.

[0068] In an optional manner of the present application, the initial mask is expanded based on the texture complexity of the area corresponding to the initial mask, including:

[0069] The dilation factor is determined based on the texture complexity of the area corresponding to the initial mask. The texture complexity is positively correlated with the dilation factor.

[0070] Dilate the initial mask based on the dilation factor.

[0071] The dilation factor is a coefficient used to expand the degree of an object in an image during a morphological operation. For example, the dilation factor can be the size of a morphological kernel or the number of iterations of a dilation operation.

[0072] In an embodiment of the present application, a preset calculation rule can be provided to calculate the expansion factor based on the texture complexity of the area corresponding to the initial mask. The preset calculation rule can indicate that there is a positive correlation between the texture complexity and the expansion factor, that is, the higher the texture complexity, the higher the expansion factor, and the lower the texture complexity, the lower the expansion factor.

[0073] As an example, the preset calculation rule may be a linear mapping relationship between texture complexity and dilation factor.

[0074] In an embodiment of the present application, the expansion factor is dynamically adjusted based on the texture complexity of the area corresponding to the initial mask, so that the initial mask can be expanded to a reasonable amplitude, thereby avoiding the situation where the expansion amplitude of the initial mask does not match the actual texture complexity due to providing a fixed value of the expansion factor, thereby affecting the image restoration quality.

[0075] In an optional manner of the present application, the dilation factor is determined based on the texture complexity of the area corresponding to the initial mask, including:

[0076] Determine pixel gradient values ​​between boundary pixels of the area corresponding to the initial mask and at least some of the pixels adjacent to the boundary pixels;

[0077] The texture complexity of the area corresponding to the initial mask is determined based on the pixel gradient value.

[0078] The boundary pixels are pixels located at the boundary of the mask coverage area, and the adjacent pixels are pixels adjacent to the boundary pixels.

[0079] The pixel gradient value between the boundary pixel point and the adjacent pixel point can reflect the pixel change at the boundary position of the mask coverage area. The texture complexity of the mask coverage area can be determined based on the pixel gradient value.

[0080] As an example, adjacent pixels are pixels directly adjacent to boundary pixels in four directions (up, down, left, and right). It is understood that pixel gradient values ​​for boundary pixels can be calculated with all of the aforementioned adjacent pixels to accurately capture pixel changes in all directions. Alternatively, pixel gradient values ​​can be calculated with a subset of the adjacent pixels, such as selecting adjacent pixels in a specified direction to focus on pixel changes in the specified direction and reduce the overall computational effort.

[0081] As an example, the pixel gradient values ​​between the boundary pixel point and each adjacent pixel point can be calculated separately, and the average of the pixel gradient values ​​between the boundary pixel point and each adjacent pixel point can be used as the overall pixel gradient value of the boundary pixel point. Then, the average of the overall pixel gradient values ​​of all boundary pixels can be normalized as the texture complexity of the mask coverage area.

[0082] It is understood that in addition to using pixel gradient values ​​to calculate texture complexity, other methods can also be used to calculate texture complexity. For example, the variance of the local binary pattern (LBP) in the neighborhood of the boundary pixel can be calculated, and the texture complexity can be calculated based on the variance of the local binary pattern. For another example, the information entropy of the grayscale value in the neighborhood of the boundary pixel can be calculated, and the texture complexity can be calculated based on the information entropy.

[0083] In an optional embodiment of the present application, the prompt text includes detection prompt text for prompting the large language model to detect the target object. Based on the prompt text for the target object, determining the target area where the target object is located from the image to be processed includes:

[0084] Input the detection prompt text and the image to be processed into the large language model, and extract the text prompt features of the target object from the response text output by the large language model;

[0085] Perform feature extraction on the image to be processed to obtain image features, and perform feature extraction on the detection prompt text to obtain text features;

[0086] Perform feature enhancement processing on image features based on text prompt features to obtain enhanced image features;

[0087] Determine the target image features from the enhanced image features, and the similarity between the target image features and the text prompt features meets the preset conditions;

[0088] The target area where the target object is located is determined based on the target image features.

[0089] The detection prompt text prompts the large language model to detect the target object. This prompt text allows for a more abstract description of the target object. Leveraging the large language model's comprehension capabilities, it can better understand the semantic information contained in the prompt text. For example, the prompt text could be "Detect objects associated with people in the image" or "Detect objects related to sports in the image."

[0090] After the detection prompt text and the image to be processed are input into the large language model, the large model will output a response text. Part of the text can be extracted from the response text as a text prompt feature. This text prompt feature integrates the semantic information of the target object and the location information of the target object in the image to be processed.

[0091] As an example, the response text may contain non-semantic anchor tags (such as <det>), the positioning mark can be used to locate the text prompt feature, so as to extract the text prompt feature from the response text based on the positioning mark.

[0092] In this example, the large language model can be specifically a multimodal large language model (MLLM). By fine-tuning the multimodal large language model, it is possible to introduce non-explicit semantics into the vocabulary of the multimodal large language model. <det>Tag, embed the hidden part of the tag into the response text, and make a positioning prompt of the text prompt feature.

[0093] The image content and the detection prompt text can be respectively extracted through the corresponding feature extraction network to obtain image features and text features.

[0094] The text prompt features and image features are semantically aligned across modalities, and the aligned target image features are output. The target image features are more focused on the target area where the target object is located. Target area detection based on the target image features can achieve higher detection accuracy.

[0095] Specifically, image features can be enhanced based on the textual hint features to obtain enhanced image features. From these enhanced image features, image features with high similarity to the textual features can be selected as target image features, thereby achieving semantic alignment with the image. Finally, target detection can be performed based on the target image features to determine the target region where the target object is located.

[0096] As an example, the preset condition may be that the similarity between the target image feature and the text feature is greater than a preset similarity threshold.

[0097] Related technologies generally require prompt text to have a clear directionality (such as a clear name or category), cannot support ambiguous prompt text, and lack semantic generalization capabilities. In this example, the reasoning ability of the large language model can better understand the semantic information in the detection prompt text, thereby supporting ambiguous prompt text and having semantic generalization capabilities, realizing the recognition of multiple possible item categories. For example, if the prompt text is "Detect sports-related items", the large language model can recognize multiple possible categories of items in the image, such as "basketball", "running shoes", "yoga mat", etc.

[0098] As an example, Figure 4 The following is a schematic diagram of the structure of an object detector provided in an embodiment of the present application. The object detector is used to align the detection prompt text with the image to be processed and detect the target area.

[0099] The object detector may specifically include a multimodal large language model, a detection encoder, a cross-modal alignment module, and a detection decoder.

[0100] Among them, the image to be processed and the detection prompt text can be input into the multimodal large language model, and the text prompt features can be extracted from the response text output by the multimodal large language model.

[0101] In this example, the image content of the image to be processed can be as follows Figure 4 As shown in , the detection prompt text may be "Please detect something used for contacting other people in this image", which instructs the multimodal large model to detect items used to contact people from the image to be processed.

[0102] The detection encoder is used to extract image features from the image to be processed and text features from the detection prompt text.

[0103] The cross-modal alignment module is used to perform cross-modal semantic alignment of text prompt features and image features, and output the aligned target image features.

[0104] The detection decoder is used to predict the target area based on the target image features. Figure 4 As shown in , the target area can be represented by a detection box. In this example, the target area is actually the area where the phone is located in the image to be processed.

[0105] In an optional manner of the present application, feature enhancement processing is performed on image features based on text prompt features to obtain enhanced image features, including:

[0106] Determine key features and value features based on text prompt features;

[0107] determining query features based on image features;

[0108] Cross-attention processing is performed based on query features, key features, and value features to obtain enhanced image features.

[0109] The text hint features can be linearly transformed to generate key features and value features, while the image features can be linearly transformed to generate query features. Cross-attention processing is then performed on the query, key, and value features to generate enhanced image features. By linearly transforming the text hint features as key and value features and performing cross-attention processing on the image features, relevant features of the target region within the image features are activated, thereby enhancing the image features.

[0110] As an example, continue Figure 4 The example shown in Figure 5 Schematic diagram of the structure of the cross-modal alignment module in an embodiment of the present application.

[0111] like Figure 5 As shown in , after cross-attention processing of image features based on text prompt features, enhanced image features can be obtained. Then, target image features with high similarity to text features can be screened from the enhanced image features.

[0112] In an optional manner of the present application, determining the target area where the target object is located from the image to be processed based on the prompt text for the target object includes:

[0113] Using the target detection model, based on the prompt text of the target object, the target area where the target object is located is determined from the image to be processed;

[0114] Object detection models include any of the following:

[0115] Marrying DINO with Grounded Pre-Training for Open-Set Object Detection (Grounding-DINO), a language-guided self-distillation unlabeled learning model.

[0116] Grounded Language-Image Pre-training (GLIP) is a language-based vision-language pre-training model.

[0117] In the embodiment of the present application, the target region detection can be implemented based on the aforementioned target detector or a known target detection model. The target detection model uses the prompt text as a semantic guide to perform target detection on the processed image and determine the target region.

[0118] In an embodiment of the present application, the target detection model may be Grounding-DINO or GLIP.

[0119] In the processing flow implemented by Grounding-DINO, multimodal feature extraction is first performed, specifically feature extraction for the prompt text and the image to be processed. The prompt text can be encoded into high-dimensional semantic vectors as text features using a pre-trained language model to capture the contextual information of the keywords. Multi-scale visual features are extracted from the image to be processed using a convolutional neural network to preserve spatial structure and texture information. Next, feature alignment between the text features and the image features is performed using a cross-modal attention mechanism. Specifically, the cross-modal attention mechanism establishes semantic associations between the prompt text and the image to be processed. By calculating the spatial similarity between the text features and the image features, an attention weight matrix is ​​generated to highlight regions in the image related to the prompt text. This mechanism effectively addresses the semantic ambiguity problem faced by traditional single-modal detectors in complex scenarios. Finally, the cross-modal attention weights are fused with image features to predict the bounding box of the target region. By introducing a multi-scale feature pyramid, the system can simultaneously process large-scale objects and small-scale details, significantly improving localization accuracy. Furthermore, the non-maximum suppression (NMS) algorithm can be used to filter out redundant boxes, ensuring the uniqueness and accuracy of the output results. Object detection based on Grounding-DINO can effectively avoid the semantic ambiguity problem of traditional single-modal detectors in complex scenarios.

[0120] In an optional embodiment of the present invention, segmenting the image to be processed based on the target region to generate an initial mask of the target object includes:

[0121] Through the target segmentation model, the image to be processed is segmented based on the target area to generate an initial mask of the target object;

[0122] Object segmentation models include any of the following:

[0123] Segment Anything Model 2 (SAM2) for image and video segmentation;

[0124] Edge Detection with Transformer (EDTER) model based on transformer.

[0125] In the embodiment of the present application, the target segmentation model may adopt SAM2 or EDTER.

[0126] SAM2 is based on a Transformer architecture and can handle objects of arbitrary shapes and scales. Its encoder extracts global contextual information from the image to be processed, while the decoder recovers spatial details through layer-by-layer upsampling, ultimately outputting a high-resolution mask.

[0127] In an optional manner of the present application, performing restoration processing on the image to be processed based on the target mask to obtain a restored image includes:

[0128] Through the target restoration model, the image to be processed is restored based on the target mask to obtain the restored image;

[0129] Target repair models include any of the following:

[0130] Large Mask Inpainting Model LAMA;

[0131] Mask-Aware Transformer for LargeHole Image Inpainting (MAT).

[0132] In an embodiment of the present application, the target repair model may be LAMA or MAT.

[0133] LAMA is based on an improved Residual Network (ResNet) architecture that incorporates Fast Fourier Convolution (FFC) technology, enabling it to model image features in both the frequency and spatial domains. Its encoder extracts local texture and global structure information through multi-level convolution, while the decoder gradually restores the image content through deconvolution.

[0134] Traditional convolution operations are limited by their local receptive field and struggle to model long-range dependencies. FFC, however, effectively captures the overall structural information of an image by extracting global features in the frequency domain. Furthermore, spatial convolution preserves local details, and the frequency and spatial domains work together to achieve high-quality restoration results.

[0135] Furthermore, during the restoration process, LAMA integrates feature information from different levels through a multi-scale feature pyramid fusion algorithm. Bottom-level features preserve local texture, mid-level features align structural information, and top-level features model global semantics. A dynamic weight allocation mechanism adaptively balances the contributions of features at different levels, ensuring a natural transition between the restored area and surrounding content.

[0136] After the restoration is complete, LAMA can further optimize the restoration results through the post-processing module. This module combines image smoothing algorithms and edge enhancement technology to eliminate minor defects in the repaired area and improve the overall visual effect.

[0137] As an example, Figure 6 A schematic diagram of the overall flow of the image processing method provided in an embodiment of the present application.

[0138] like Figure 6 As shown in [1], we first perform target detection on the image to generate the target region based on the prompt text. We then segment the image based on the target region to generate an initial mask. We then dilate the initial mask to generate the target mask. Finally, we perform inpainting on the image based on the target mask to obtain the inpainted image.

[0139] In this example, by dynamically aligning the prompt text with image features, the recognition of abstract concepts and the accuracy of target object detection are significantly improved compared to fixed encoding methods. By combining target region prediction with segmentation mask generation, the problems of jagged edges and residual content are effectively addressed. By introducing a dynamic expansion factor and adaptively adjusting the expansion factor, the restoration effect on boundaries is guaranteed, maintaining overall image consistency when restoring images with complex backgrounds, and improving restoration speed.

[0140] The above method provided in the embodiment of the present application can be applied to a variety of application scenarios, including but not limited to: removing and repairing target objects in images in the fields of e-commerce, medical security, etc.

[0141] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0142] According to an embodiment of another aspect, an image processing apparatus is provided. Figure 7 A schematic block diagram of an image processing device according to an embodiment is shown. Figure 1 The user side or server side in the architecture shown. Figure 7 As shown, the image processing device 700 includes:

[0143] The target detection module 710 is configured to determine a target area where a target object is located from the image to be processed based on a prompt text for the target object;

[0144] The target segmentation module 720 is used to segment the image to be processed based on the target area and generate an initial mask of the target object;

[0145] A mask expansion module 730 is configured to expand the initial mask based on the texture complexity of the area corresponding to the initial mask to obtain a target mask;

[0146] The image restoration module 740 is configured to perform restoration processing on the image to be processed based on the target mask to obtain a restored image.

[0147] As an optional manner, when the mask expansion module 730 performs expansion processing on the initial mask based on the texture complexity of the area corresponding to the initial mask, it is specifically configured to:

[0148] The dilation factor is determined based on the texture complexity of the area corresponding to the initial mask. The texture complexity is positively correlated with the dilation factor.

[0149] Dilate the initial mask based on the dilation factor.

[0150] As an optional manner, when determining the dilation factor based on the texture complexity of the area corresponding to the initial mask, the mask dilation module 730 is specifically configured to:

[0151] Determine pixel gradient values ​​between boundary pixels of the area corresponding to the initial mask and at least some of the pixels adjacent to the boundary pixels;

[0152] The texture complexity of the area corresponding to the initial mask is determined based on the pixel gradient value.

[0153] As an optional manner, the prompt text includes detection prompt text for prompting the large language model to detect the target object. The target detection module 710 is specifically configured to:

[0154] Input the detection prompt text and the image to be processed into the large language model, and extract the text prompt features of the target object from the response text output by the large language model;

[0155] Perform feature extraction on the image to be processed to obtain image features, and perform feature extraction on the detection prompt text to obtain text features;

[0156] Perform feature enhancement processing on image features based on text prompt features to obtain enhanced image features;

[0157] Determine the target image features from the enhanced image features, and the similarity between the target image features and the text features meets a preset condition;

[0158] The target area where the target object is located is determined based on the target image features.

[0159] As an optional manner, when the target detection module 710 performs feature enhancement processing on the image features based on the text prompt features to obtain the enhanced image features, it is specifically used to:

[0160] Determine key features and value features based on text prompt features;

[0161] determining query features based on image features;

[0162] Cross-attention processing is performed based on query features, key features, and value features to obtain enhanced image features.

[0163] As an optional approach, the target detection module 710 is specifically configured to:

[0164] Using the target detection model, based on the prompt text of the target object, the target area where the target object is located is determined from the image to be processed;

[0165] Object detection models include any of the following:

[0166] Grounding-DINO;

[0167] GLIP.

[0168] As an optional approach, the object segmentation module 720 is specifically configured to:

[0169] Through the target segmentation model, the image to be processed is segmented based on the target area to generate an initial mask of the target object;

[0170] Object segmentation models include any of the following:

[0171] SAM;

[0172] EDTER.

[0173] As an optional approach, the image restoration module 740 is specifically configured to:

[0174] Through the target restoration model, the image to be processed is restored based on the target mask to obtain the restored image;

[0175] Target repair models include any of the following:

[0176] LAMA;

[0177] MAT.

[0178] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.

[0179] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0180] And an electronic device comprising:

[0181] one or more processors; and

[0182] A memory associated with one or more processors, the memory being used to store program instructions, which, when read and executed by one or more processors, execute the steps of any one of the method embodiments described above.

[0183] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the method embodiments described above when executed by a processor.

[0184] in, Figure 8 The electronic device architecture is shown as an example, and may include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, the video display adapter 811, the disk drive 812, the input / output interface 813, the network interface 814, and the memory 820 may be communicatively connected via a communication bus 830.

[0185] Among them, the processor 810 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.

[0186] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store an operating system 821 for controlling the operation of the electronic device 800, and a basic input and output system (BIOS) 822 for controlling the low-level operations of the electronic device 800. In addition, a web browser 823, a data storage management system 824, and an image processing device 825, etc. can also be stored. The above-mentioned image processing device 825 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.

[0187] The input / output interface 813 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0188] The network interface 814 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0189] The bus 830 comprises a pathway for transmitting information between the various components of the device (eg, the processor 810 , the video display adapter 811 , the disk drive 812 , the input / output interface 813 , the network interface 814 , and the memory 820 ).

[0190] It should be noted that although the above device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0191] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0192] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.< / det> < / det>

Claims

1. An image processing method, characterized in that: include: Based on the prompt text for the target object, determining a target area where the target object is located from the image to be processed; Segmenting the image to be processed based on the target area to generate an initial mask of the target object; Based on the texture complexity of the area corresponding to the initial mask, the initial mask is expanded to obtain a target mask; The image to be processed is repaired based on the target mask to obtain a repaired image.

2. The method according to claim 1, characterized in that The dilation processing of the initial mask based on the texture complexity of the area corresponding to the initial mask includes: Determining a dilation factor based on texture complexity of an area corresponding to the initial mask, wherein the texture complexity is positively correlated with the dilation factor; The initial mask is expanded based on the expansion factor.

3. The method according to claim 2, characterized in that The step of determining the dilation factor based on the texture complexity of the area corresponding to the initial mask includes: Determine pixel gradient values ​​between boundary pixels of an area corresponding to the initial mask and at least some of the adjacent pixels of the boundary pixels; The texture complexity of the area corresponding to the initial mask is determined based on the pixel gradient value.

4. The method according to claim 1, wherein The prompt text includes detection prompt text for prompting the large language model to detect the target object. The determining of the target area where the target object is located from the image to be processed based on the prompt text for the target object includes: Inputting the detection prompt text and the image to be processed into a large language model, and extracting the text prompt features of the target object from the response text output by the large language model; Performing feature extraction on the image to be processed to obtain image features, and performing feature extraction on the detection prompt text to obtain text features; Performing feature enhancement processing on the image feature based on the text prompt feature to obtain enhanced image features; Determining a target image feature from the enhanced image features, wherein the similarity between the target image feature and the text feature satisfies a preset condition; A target area where the target object is located is determined based on the target image features.

5. The method according to claim 4, characterized in that The performing feature enhancement processing on the image feature based on the text prompt feature to obtain enhanced image features includes: Determining a key feature and a value feature based on the text prompt feature; determining a query feature based on the image feature; Cross-attention processing is performed based on the query feature, the key feature, and the value feature to obtain enhanced image features.

6. The method according to claim 1, characterized in that The step of determining the target area where the target object is located from the image to be processed based on the prompt text for the target object includes: Determining the target area where the target object is located from the image to be processed based on the prompt text of the target object by using the target detection model; The target detection model includes any of the following: Grounding-DINO, a language-guided self-distillation unlabeled learning model; Language-based vision-language pre-training model GLIP.

7. The method according to claim 1, characterized in that The segmenting of the image to be processed based on the target area to generate an initial mask of the target object includes: Segmenting the image to be processed based on the target area using a target segmentation model to generate an initial mask of the target object; The target segmentation model includes any of the following: Segmentation Everything Model (SAM) in images and videos; Transformer-based edge detection model EDTER.

8. The method according to claim 1, characterized in that The repairing process is performed on the image to be processed based on the target mask to obtain a repaired image, including: Performing restoration processing on the image to be processed based on the target mask using the target restoration model to obtain a restored image; The target repair model includes any of the following: Large Mask Inpainting Model LAMA; Mask-aware transformer large hole image restoration model MAT.

9. An image processing device, characterized in that: include: A target detection module is used to determine a target area where a target object is located from an image to be processed based on a prompt text for the target object; A target segmentation module, configured to segment the image to be processed based on the target area and generate an initial mask of the target object; A mask expansion module, configured to expand the initial mask based on the texture complexity of the area corresponding to the initial mask to obtain a target mask; The image restoration module is used to perform restoration processing on the image to be processed based on the target mask to obtain a restored image.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.