Visual data processing method, image erasing method, computing device and storage medium

By extracting the correlation information of the area to be processed during the image erasure process and adjusting it using a diffusion model, the problems of smearing and blurring after image erasure are solved, thus improving image quality.

CN121746211APending Publication Date: 2026-03-27ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing image erasing techniques tend to result in strong smearing, blurring, and high distortion when erasing specific parts of an image, leading to a decrease in image quality.

Method used

By identifying the region to be processed in the visual data, performing content processing on the region, extracting the correlation information of the intermediate visual data, such as content, edge and depth information, and adjusting it using a diffusion model, the target visual data is obtained.

Benefits of technology

It improves the blur and distortion of intermediate visual data, enhances image quality, avoids the smearing effect of secondary generation, and ensures the clarity and realism of the erased image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746211A_ABST
    Figure CN121746211A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of image processing, in particular to a visual data processing method, an image erasing method, computing equipment and a storage medium, and the visual data processing method comprises the following steps: determining to-be-processed visual data and a to-be-processed area included in the to-be-processed visual data; the to-be-processed area is subjected to area content processing, intermediate visual data corresponding to the to-be-processed visual data are obtained, and the intermediate visual data comprise the to-be-processed area after area content processing; according to the mesopic vision data, at least two kinds of associated information of the mesopic vision data are determined, and the at least two kinds of associated information are attribute information of a to-be-processed area after the area content is processed; and adjusting the mesopic vision data according to the at least two kinds of associated information to obtain target vision data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of image processing technology, and in particular to visual data processing methods, image erasure methods, computing devices, and storage media. Background Technology

[0002] With the widespread use of smartphones, taking photos has become an important way for people to record moments in their daily lives. This has led to a dramatic increase in the number of photos for personal and commercial use. These photos sometimes contain unnecessary elements or sensitive information, such as road signs, billboards, pedestrians, or even personally identifiable information, which may need to be removed from the images to protect privacy or improve image quality. Therefore, image erasing technology has emerged and become a key function in image editing software.

[0003] Image erasing techniques are typically based on computer vision algorithms, allowing users to select and remove specific portions of an image, then fill in the gaps with the surrounding environment or other intelligently generated content. However, when erasing specific parts of an image, the erased areas may appear heavily smeared, resulting in a blurred and distorted image, ultimately degrading image quality. Therefore, an effective technical solution is urgently needed to address these issues. Summary of the Invention

[0004] In view of the above, embodiments of this specification provide a visual data processing method. One or more embodiments of this specification also relate to a visual data processing apparatus, an image erasing method, an image erasing device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a visual data processing method is provided, comprising: Determine the visual data to be processed and the regions to be processed included in the visual data to be processed; The region to be processed is subjected to region content processing to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing. Based on the intermediate visual data, at least two types of related information are determined, wherein the at least two types of related information are attribute information of the region to be processed after the region content processing; Based on the at least two types of related information, the intermediate visual data is adjusted to obtain the target visual data.

[0006] According to a second aspect of the embodiments of this specification, a visual data processing apparatus is provided, comprising: The first determining module is configured to determine the visual data to be processed and the region to be processed included in the visual data to be processed; The first processing module is configured to perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after the region content processing. The second determining module is configured to determine at least two types of related information of the intermediate visual data based on the intermediate visual data, wherein the at least two types of related information are attribute information of the region to be processed after the region content processing; The second processing module is configured to adjust the intermediate visual data based on the at least two types of association information to obtain the target visual data.

[0007] According to a third aspect of the embodiments of this specification, an image erasing method is provided, comprising: Determine the image to be processed and the region to be processed included in the image to be processed; The area to be processed is erased to obtain an intermediate image corresponding to the image to be processed, wherein the intermediate image includes the erased area to be processed; Based on the intermediate image, at least two types of associated information are determined, wherein the at least two types of associated information are attribute information of the erased area to be processed; The intermediate image is adjusted based on the at least two types of related information to obtain the target image.

[0008] According to a fourth aspect of the embodiments of this specification, an image erasing apparatus is provided, comprising: The first determining module is configured to determine the image to be processed and the region to be processed included in the image to be processed; The first processing module is configured to erase the area to be processed to obtain an intermediate image corresponding to the image to be processed, wherein the intermediate image includes the erased area to be processed; The second determining module is configured to determine at least two types of associated information related to the intermediate image based on the intermediate image, wherein the at least two types of associated information are attribute information of the erased area to be processed; The second processing module is configured to adjust the intermediate image based on the at least two types of association information to obtain the target image.

[0009] According to a fifth aspect of the embodiments of this specification, an image erasing system is provided, including an edge device and a cloud-side device, wherein... The edge device is configured to send the visual data to be processed to the cloud-side device; The cloud-side device is configured to receive visual data to be processed and determine the region to be processed included in the visual data to be processed; perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing; determine at least two kinds of association information of the intermediate visual data based on the intermediate visual data, wherein the at least two kinds of association information are attribute information of the region to be processed after region content processing; adjust the intermediate visual data based on the at least two kinds of association information to obtain target visual data, and send the target visual data to the edge device.

[0010] According to a sixth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.

[0011] According to a seventh aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0012] According to an eighth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0013] One embodiment of this specification provides a visual data processing method. By determining the visual data to be processed and the region to be processed included in the visual data to be processed, the method first performs initial region content processing on the region to be processed to obtain intermediate visual data. Then, based on the intermediate visual data, it determines at least two kinds of association information associated with the intermediate visual data. The at least two kinds of association information are used as guiding information to guide the adjustment of the intermediate visual data to obtain target visual data. This achieves the adjustment of the region to be processed obtained from the initial region content processing and the region content processing, improving the problems of blurriness, high distortion, and smearing that may exist in the intermediate visual data, thereby ensuring the quality of the target visual data obtained after adjustment. Attached Figure Description

[0014] Figure 1 This is a schematic diagram illustrating an application scenario of a visual data processing method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a visual data processing method provided in one embodiment of this specification; Figure 3 This is an architecture diagram of a visual data processing method provided in one embodiment of this specification; Figure 4 This is a flowchart illustrating the processing procedure of a visual data processing method provided in one embodiment of this specification. Figure 5 This is a schematic diagram of the structure of a visual data processing device provided in one embodiment of this specification; Figure 6 This is a flowchart of an image erasing method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the structure of an image erasing device provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0015] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0016] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0017] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0018] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0019] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0020] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0021] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0022] Image erasure: refers to erasing the content of a specified area in an image and filling the area with its surrounding background.

[0023] Generative erasure: refers to erasing images using a diffusion model.

[0024] Secondary generation: This refers to the process where, instead of filling the area that needs to be erased with a background, new foreground content, such as people or vehicles, is generated.

[0025] Diffusion models are a type of latent variable model, which is a Markov chain trained using variational estimation. Diffusion models learn the latent structure of a dataset by modeling how data points spread throughout the latent space. After training, randomly sampled noise can be fed into the diffusion model, and the model learns a denoising process to generate new data.

[0026] GAN (Generative Adversarial Network) is an unsupervised learning method that learns by having two neural networks compete against each other. A GAN consists of a generator network and a discriminator network. The generator network randomly samples data from the latent space as input, and its output should mimic real samples in the training set as closely as possible. The discriminator network takes either real samples or the output of the generator network as input, and its goal is to distinguish the generator network's output from real samples as much as possible. After training, the generator network can be used to generate data.

[0027] blip model: A pre-trained bidirectional encoder model for image processing.

[0028] WD14 model: A deep learning-based technical model that learns how to convert images into simple text descriptions by training on a large dataset.

[0029] The Sobel operator is a gradient-based edge detection operator. It detects edges by calculating the gradients of image grayscale values ​​in the horizontal and vertical directions. The Sobel operator uses two 3x3 matrices as convolution kernels to estimate the gradients of the image in the x-axis and y-axis directions, respectively.

[0030] Canny edge detection is a multi-stage edge detection algorithm designed to find edge points that are most likely to represent true edges under given noise conditions. Canny edge detection includes the following steps: (1) Gaussian filtering: First, Gaussian filtering is applied to the input image to remove noise; (2) Calculation of gradient strength and direction: The gradient strength and direction of each pixel in the image are calculated using the Sobel operator or a similar difference operator; (3) Non-maximum suppression: Each pixel is examined along the gradient direction, and only those whose gradient values ​​are locally maximum are considered edges; (4) Double threshold detection and edge connection: Edges are connected using two thresholds, high and low. Edges below the low threshold are suppressed, edges above the high threshold are preserved, and edges in between are preserved only if both ends are connected to the high threshold edge.

[0031] The Laplacian operator is a second-order differential operator used to detect second-order changes in grayscale values ​​in an image, that is, to detect locations where grayscale values ​​change abruptly. The Laplacian operator is commonly used to highlight details in an image and can also be used for edge detection.

[0032] Lineart: An image processed by an algorithm that retains only the most prominent edges and contours of the original image.

[0033] In practical applications, during the erasure of specific parts of an image, secondary generation often occurs. That is, instead of filling the background in the area to be erased, new foreground content is generated. This results in a smeared, blurry, and distorted image with poor quality. Therefore, an effective technical solution is urgently needed to address these problems.

[0034] This specification provides a visual data processing method, and also relates to a visual data processing apparatus, an image processing method, an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0035] See Figure 1 , Figure 1 The illustration shows an application scenario diagram of a visual data processing method according to an embodiment of this specification. The visual data processing method includes: determining visual data to be processed and a region to be processed included in the visual data to be processed; performing region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing; determining at least two types of association information based on the intermediate visual data, wherein the at least two types of association information are attribute information of the region to be processed after region content processing; and adjusting the intermediate visual data based on the at least two types of association information to obtain target visual data.

[0036] like Figure 1 As shown, Figure 1 It includes end-side device 102 and cloud-side device 104.

[0037] In practice, an enterprise can provide image processing services to users by developing an application. Users can use the application through the edge device 102. In the application, users can upload visual data to be processed and the regions to be processed included in the visual data. The edge device 102 can send the uploaded visual data and regions to be processed to the cloud device 104 of the application. The cloud device 104 can perform regional content processing on the regions to be processed to obtain intermediate visual data and determine at least two kinds of association information associated with the intermediate visual data. Based on the at least two kinds of association information, the intermediate visual data is adjusted to obtain target visual data, and the target visual data is sent to the edge device 102 for display through the display interface of the edge device 102.

[0038] In practical applications, intermediate visual data can be visual data after the area to be processed has been pre-erased, and target visual data can be visual data after secondary adjustment of the intermediate visual data.

[0039] In addition, a company can also provide image processing services to users by developing web pages. For example, the company can deploy a question-and-answer model on the web page. This model can obtain the visual data to be processed and the area to be processed uploaded by the user in the form of a question-and-answer dialogue, execute the above-mentioned visual data processing methods, and return the processed target visual data to the user.

[0040] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0041] Cloud-side device 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Cloud-side device 104 can also be a server for a distributed system, or a server integrated with blockchain. Cloud-side device 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0042] It is worth noting that the visual data processing method provided in the embodiments of this specification can be executed by the cloud-side device 104. In other embodiments of this specification, the model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the cloud-side device 104, thereby executing the visual data processing method provided in the embodiments of this specification. In other embodiments, the visual data processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the cloud-side device 104.

[0043] See Figure 2 , Figure 2 A flowchart of a visual data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0044] Step 202: Determine the visual data to be processed and the regions to be processed included in the visual data to be processed.

[0045] In this context, "visual data to be processed" can be understood as the visual data that needs to be processed. Visual data can be understood as data acquired through various imaging devices (such as cameras, scanners, satellites, etc.) and primarily used for visual perception. Visual data can include images, visual sequences, etc. The "region to be processed" within the visual data to be processed can be understood as the area within the visual data that needs to be erased.

[0046] For ease of understanding, this specification uses images as visual data as an example for illustration.

[0047] Based on this, we can determine the image that needs to be processed, and the area in the image that needs to be erased.

[0048] For example, if the visual data to be processed is an image of person A, and the background of person A in the image also includes pedestrians, then the area to be processed in the visual data to be processed can be the location of the pedestrians in the image.

[0049] In practical applications, determining the visual data to be processed and the region to be processed included in the visual data to be processed includes: Receive visual data to be processed sent by the client; In response to a processing instruction sent by the client for the visual data to be processed, the region to be processed included in the visual data to be processed is determined.

[0050] The processing instructions for the visual data to be processed can include user selection and erasing instructions on the visual data displayed on the client's interface. The area to be processed can be understood as the area in the visual data that needs to be erased.

[0051] Based on this, the server can receive visual data to be processed sent by the user through the client. Furthermore, the user can perform point operations or smear operations on the visual data to be processed displayed on the client's interface. The client can send the user's point operation or smear operation command to the server. The server responds to the point operation or smear operation command and determines the area to be processed that the user wants to erase in the visual data to be processed.

[0052] Using the previous example, a user can send an image of person A to the server via the client. The user can then erase or scribble on the image at the location of pedestrians in the background of person A. The client sends the erase or scribble instruction to the server, and the server responds to the instruction by determining the area to be erased in the image. This area is the part the user wants to erase.

[0053] In addition, users can directly upload visual data containing the area to be processed through the client, or users can send command text to the server through the client via text input or verbal description, and the server can determine the area to be processed in the visual data based on the command text. This specification does not limit this approach.

[0054] In summary, by interacting with the client, the visual data to be processed and the area to be processed can be determined, which facilitates the subsequent determination of the parts of the visual data to be processed that need to be erased.

[0055] Step 204: Perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing.

[0056] The process of processing the region to be processed can be understood as performing pre-erasing of the content within that region. Intermediate visual data can be understood as the visual data after the region to be processed within the visual data to be processed has undergone pre-erasing. Understandably, intermediate visual data may exhibit strong smearing, blurriness, and high distortion.

[0057] Based on this, the content of the region to be processed in the visual data can be pre-erased to obtain intermediate visual data.

[0058] Using the previous example, we can perform pre-erasing on the location of the pedestrian in the image to obtain the erased image.

[0059] In specific implementation, the step of performing region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed includes: Using a visual data processing model, the region to be processed is erased to obtain intermediate visual data corresponding to the visual data to be processed.

[0060] The visual data processing model can be understood as a model used to erase content from visual data.

[0061] Based on this, the visual data to be processed and the regions to be processed included in the visual data to be processed can be input into the visual data processing model. In the visual data processing model, the regions to be processed in the visual data to be processed are pre-erased according to the regions to be processed, and intermediate visual data corresponding to the visual data to be processed is obtained. The intermediate visual data is the pre-erasing result.

[0062] In practical applications, the visual data processing model can be a GAN model, and the region to be processed can be input into the visual data processing model in the form of a mask of the region to be processed in the visual data.

[0063] Using the previous example, the image and the area to be processed where the pedestrian is located in the image can be input into the visual data processing model to obtain the erased image.

[0064] In summary, by using the GAN model to pre-erase the input visual data based on the mask of the region to be processed, the foreground information in the visual data to be processed is effectively removed, avoiding the secondary generation of the region to be processed. Furthermore, the intermediate visual data, as the pre-erase result, can serve as a guide for the pre-erase result of the subsequent diffusion model to complete the data.

[0065] Step 206: Based on the intermediate visual data, determine at least two types of related information of the intermediate visual data, wherein the at least two types of related information are attribute information of the region to be processed after the region content processing.

[0066] Among them, at least two types of related information associated with intermediate visual data can serve as information guidance for subsequent diffusion model completion.

[0067] In one embodiment of this specification, the at least two types of associated information include content information, edge information, and / or depth information of the region to be processed after the region content has been processed.

[0068] The content information of the region to be processed after region content processing can be understood as the visual content information contained in the intermediate visual data. For example, it can describe the environment and scene contained in the intermediate visual data, such as parks, playgrounds, and grasslands. The edge information of the region to be processed after region content processing can be understood as the parts of the intermediate visual data with significant brightness changes. This typically corresponds to the boundaries between different objects or abrupt changes in the internal structure of an object. Edge information can mark changes in the outline or surface properties of an object. The depth information of the region to be processed after region content processing can be understood as the distance of each point in the scene of the intermediate visual data from the observer.

[0069] Understandably, at least two types of related information can be the content information and edge information of the region to be processed after the region content has been processed, or the content information and depth information of the region to be processed after the region content has been processed, or the edge information and depth information of the region to be processed after the region content has been processed, or the content information, edge information and depth information of the region to be processed after the region content has been processed.

[0070] In summary, by determining the content information, edge information, and / or depth information of the region to be processed after content processing, it is easier to guide the subsequent diffusion model completion based on this information, thus ensuring the adjustment of the pre-erasure results.

[0071] In specific implementation, determining at least two types of related information about the intermediate visual data based on the intermediate visual data includes: Content extraction is performed on the intermediate visual data to obtain the content information of the region to be processed after the region content processing; Edge detection is performed on the intermediate visual data to obtain the edge information of the region to be processed after the region content processing; Depth detection is performed on the intermediate visual data to obtain the depth information of the region to be processed after the region content is processed.

[0072] In practical applications, see Figure 3 , Figure 3 An architecture diagram of a visual data processing method according to an embodiment of this specification is shown, such as... Figure 3 As shown, this visual data processing method can be executed by a visual data processing system 300. The visual data processing system 300 may include an erasure guidance unit 302 and a diffusion completion model (i.e., a diffusion model) 304. Specifically, the visual data to be processed and the region to be processed included in the visual data to be processed can be input into the erasure guidance unit 302. In the erasure guidance unit 302, the region to be processed included in the visual data to be processed is pre-erased to obtain intermediate visual data (i.e., pre-erasing result). Content is extracted from the intermediate visual data to obtain content information. Edge detection is performed on the intermediate visual data to obtain edge information. Depth detection is performed on the intermediate visual data to obtain depth information. The pre-erasing result, content information, edge information, and depth information are input into the diffusion completion model 304. Through multiple denoising iterations, the final erasure result (i.e., target visual data) is obtained.

[0073] In summary, by determining content information, edge information, and depth information based on intermediate visual data, three types of guiding information are identified, which facilitates subsequent adjustments to the pre-erasing results based on these three types of guiding information.

[0074] Further, the step of extracting content from the intermediate visual data to obtain content information of the region to be processed after content processing includes: By using a content extraction model, the intermediate visual data is extracted to obtain the text content corresponding to the intermediate visual data; The text content is determined as the content information of the area to be processed after the area content has been processed.

[0075] The text content corresponding to the intermediate visual data can be understood as the text content describing the background information included in the intermediate visual data. The content extraction model can be understood as a model used to extract the visual content from the visual data.

[0076] Based on this, intermediate visual data can be input into the content extraction model. In the content extraction model, content extraction is performed on the intermediate visual data to obtain text content that describes the background content included in the intermediate visual data. This text content is then identified as the content information of the area to be processed after erasure.

[0077] In practical applications, content extraction models can be blip models, WD14 models, etc. Specifically, the WD14 model can be used to label intermediate visual data to obtain textual descriptions (i.e., content information) of the visual content of the intermediate visual data, and these textual descriptions can be used as positive textual guides for the diffusion model.

[0078] In summary, since the text description of the tag is extracted from the pre-erasure result, it will not describe the foreground information that has already been erased again. By using this text description as a guide for content information, the secondary generation of foreground information can be effectively avoided.

[0079] Further, the step of performing edge detection on the intermediate visual data to obtain edge information of the region to be processed after the region content processing includes: According to the edge detection algorithm, edge detection is performed on the intermediate visual data to obtain the edge information of the region to be processed after the region content processing.

[0080] Edge detection algorithms can be understood as algorithms used for image edge detection. Edge detection algorithms can include Sobel operator, Canny edge detection, Laplacian operator, lineart, etc.

[0081] Based on this, edge detection algorithms can be used to perform edge detection on intermediate visual data to obtain edge information.

[0082] In practice, edge detection algorithms can be used to perform noise reduction, gradient calculation, non-maximum suppression, and hysteresis thresholding on intermediate visual data to obtain edge information.

[0083] In practical applications, the Canny edge detection algorithm can be used to calculate the edge information of the pre-erasing result. Optionally, the region to be processed in the pre-erasing result can be a blank region without edge information. Using blank edge information as a guide can suppress the generation of the foreground. Based on this, blank edge information can be used as the edge information of the region to be processed after content processing, and as the edge information guide for subsequent diffusion model completion.

[0084] In summary, edge detection is performed on the pre-erasing results to obtain edge information, which facilitates subsequent adjustments to the pre-erasing results under the guidance of the edge information.

[0085] Further, the step of performing depth detection on the intermediate visual data to obtain depth information of the region to be processed after the region content processing includes: According to the depth detection algorithm, depth detection is performed on the intermediate visual data to obtain the depth information corresponding to the intermediate visual data; The depth information is determined as the depth information of the region to be processed after the region content is processed.

[0086] Specifically, a depth detection model can be used to perform depth detection on intermediate visual data to obtain the depth information of the intermediate visual data, and this depth information can be used as the depth information of the region to be processed after the region content is processed.

[0087] In practical applications, the depth information of the region to be processed after regional content processing can be consistent with the depth information of the background region in the intermediate visual data.

[0088] In summary, by using depth information consistent with the background as the depth information guide for subsequent diffusion model completion, the generation of foreground can be effectively suppressed, thereby avoiding secondary image generation.

[0089] Step 208: Adjust the intermediate visual data according to the at least two types of association information to obtain the target visual data.

[0090] The target visual data can be understood as the final erasure result.

[0091] Specifically, the pre-erasing result (i.e., intermediate visual data) can be adjusted based on at least two related pieces of information to obtain the final erasing result (i.e., target visual data).

[0092] Using the previous example, the target visual data could be an image after all pedestrians except person A have been erased from the image.

[0093] In one embodiment of this specification, adjusting the intermediate visual data based on the at least two types of association information to obtain target visual data includes: Feature extraction is performed on the at least two types of related information to obtain the related information features of the at least two types of related information; Feature extraction is performed on the intermediate visual data to obtain the intermediate visual data features; The associated information features and the intermediate visual data features are combined and calculated to obtain the target visual data.

[0094] Specifically, when the associated information is content information, the associated information feature is the content information feature (i.e., text feature vector); when the associated information is edge information, the associated information feature is the edge information feature (i.e., image feature vector); and when the associated information is depth information, the associated information feature is the depth information feature (i.e., image feature vector).

[0095] In practical applications, when at least two types of associated information include content information, edge information, and depth information, the content information in text form can be encoded to obtain the content information features corresponding to the content information. The edge information and depth information in image form can be encoded to obtain the edge information features corresponding to the edge information and the depth information features corresponding to the depth information. The intermediate visual data can be extracted using an image encoding network to obtain the intermediate visual data features. The content information features and the intermediate visual data features can be combined for calculation. The edge information features, depth information features, and intermediate visual data features can be combined for calculation to obtain the target visual data.

[0096] In summary, by using three types of guidance—content information guidance, edge information guidance, and depth information guidance—the final erasure result is ensured to be as consistent as possible with the pre-erasure result in terms of text description, edge information, and depth information, effectively avoiding the secondary generation of pre-erasure result completion.

[0097] In another embodiment of this specification, adjusting the intermediate visual data based on the at least two types of association information to obtain the target visual data includes: Input the at least two types of related information and the intermediate visual data into the diffusion model; In the diffusion model, the intermediate visual data is denoised iteratively based on the at least two types of association information to obtain the target visual data.

[0098] The diffusion model can be understood as a model used to complete the pre-erasing results. Diffusion models can include denoising networks, text encoding networks, and image encoding networks.

[0099] Based on this, when at least two types of associated information include content information, edge information, and depth information, the content information, edge information, depth information, and intermediate visual data can be input into a diffusion model. In the diffusion model, the intermediate visual data is input into an image encoding network to obtain the intermediate visual data features output by the image encoding network. The content information in text form is input into a text encoding network, which encodes the content information into a text feature vector to obtain the content information features. The content information features and the intermediate visual data features are then combined to participate in the network calculation. The edge information and depth information in image form are input into an image encoding network, which encodes the edge information and depth information into image feature vectors respectively to obtain the edge information features and depth information features. The edge information features and depth information features are then combined with the intermediate visual data features to participate in the network calculation.

[0100] In practical applications, intermediate image features, content information features, edge information features, and depth information features can be input into the denoising network of the diffusion model. Gaussian noise is input into the latent space, and the denoising network of the diffusion model predicts the latent space noise. The difference between the current latent space noise feature and the predicted latent space noise is calculated, and the above steps are iterated multiple times. When the difference meets a preset difference threshold, the final generated result (i.e., the final predicted latent space noise) is obtained, realizing the denoising iteration of the diffusion model. In the diffusion model, the intermediate image features, content information features, edge information features, and depth information features are concatenated with the final predicted latent space noise to achieve a merged calculation, thereby obtaining the target concatenated feature. This target concatenated feature is then decoded to obtain the target image.

[0101] In summary, by using three types of guidance—content information guidance, edge information guidance, and depth information guidance—the output of the diffusion erasure model is made as consistent as possible with the pre-erasure result in terms of text description, edge information, and depth information. Furthermore, the completion result of the diffusion model has low inconsistency and high realism, resulting in high-quality erasure completion results.

[0102] In addition, after obtaining the target visual data, the process also includes: The target visual data is sent to the client and displayed through the client's display interface.

[0103] Specifically, the final erasure result can be sent to the client, and the final erasure result can be displayed to the user through the client's display interface.

[0104] In addition, users can send adjustment commands through the client, and the server can respond to the adjustment commands to adjust the target visual data to meet the user's image processing needs.

[0105] One embodiment of this specification provides a visual data processing method. By determining the visual data to be processed and the region to be processed included in the visual data to be processed, the method first performs initial region content processing on the region to be processed to obtain intermediate visual data. Then, based on the intermediate visual data, it determines at least two kinds of association information associated with the intermediate visual data. The at least two kinds of association information are used as guiding information to guide the adjustment of the intermediate visual data to obtain target visual data. This achieves the adjustment of the region to be processed obtained from the initial region content processing and the region content processing, improving the problems of blurriness, high distortion, and smearing that may exist in the intermediate visual data, thereby ensuring the quality of the target visual data obtained after adjustment.

[0106] The following is in conjunction with the appendix Figure 4Taking the application of the visual data processing method provided in this specification in image erasure as an example, the visual data processing method will be further explained. Figure 4 A flowchart illustrating the processing steps of a visual data processing method according to an embodiment of this specification is shown, specifically including the following steps.

[0107] Step 402: Determine the image to be processed and the region to be processed within that image.

[0108] Step 404: Input the image to be processed and the region to be processed into the visual data processing model, and perform erasure processing on the image to be processed according to the region to be processed to obtain an intermediate image.

[0109] Step 406: Input the intermediate image into the content extraction model to obtain content information.

[0110] Step 408: Input the intermediate image into the edge detection model to obtain edge information.

[0111] Step 410: Input the intermediate image into the depth detection model to obtain depth information.

[0112] Step 412: Input the intermediate image into the image encoding network to obtain the intermediate image features.

[0113] Step 414: Input the content information into the text encoding network in the diffusion model to obtain the content information features.

[0114] Step 416: Input the edge information and depth information into the image coding network in the diffusion model to obtain edge information features and depth information features.

[0115] Step 418: Input the intermediate image features, content information features, edge information features, and depth information features into the denoising network in the diffusion model, and use the denoising network to merge and calculate the intermediate image features, content information features, edge information features, and depth information features to obtain the target image.

[0116] Specifically, when merging and calculating intermediate image features, content information features, edge information features, and depth information features, these features can be input into the denoising network of the diffusion model. Gaussian noise is input into the latent space, and the denoising network of the diffusion model predicts the latent space noise. The difference between the current latent space noise feature and the predicted latent space noise is calculated, and the above steps are iterated multiple times. When the difference meets a preset difference threshold, the final generated result (i.e., the final predicted latent space noise) is obtained, thus realizing the denoising iteration of the diffusion model. In the diffusion model, the intermediate image features, content information features, edge information features, and depth information features are concatenated with the final predicted latent space noise to achieve merging and calculation, thereby obtaining the target concatenated features. These target concatenated features are then decoded to obtain the target image.

[0117] In summary, by identifying the visual data to be processed and the regions to be processed within it, initial content processing is performed on these regions to obtain intermediate visual data. Then, based on this intermediate visual data, at least two types of associated information are identified. This information serves as guiding information to adjust the intermediate visual data, ultimately yielding the target visual data. This process adjusts the regions obtained from the initial content processing, improving issues such as blurriness, high distortion, and a smeared appearance in the intermediate visual data, thereby ensuring the quality of the adjusted target visual data.

[0118] Corresponding to the above method embodiments, this specification also provides embodiments of a visual data processing device. Figure 5 A schematic diagram of the structure of a visual data processing apparatus according to one embodiment of this specification is shown. Figure 5 As shown, the device includes: The first determining module 502 is configured to determine the visual data to be processed and the region to be processed included in the visual data to be processed; The first processing module 504 is configured to perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after the region content processing. The second determining module 506 is configured to determine at least two kinds of related information of the intermediate visual data based on the intermediate visual data, wherein the at least two kinds of related information are attribute information of the region to be processed after the region content processing. The second processing module 508 is configured to adjust the intermediate visual data according to the at least two types of association information to obtain target visual data.

[0119] In an optional embodiment, the second processing module 508 is further configured to: Feature extraction is performed on the at least two types of related information to obtain the related information features of the at least two types of related information; Feature extraction is performed on the intermediate visual data to obtain the intermediate visual data features; The associated information features and the intermediate visual data features are combined and calculated to obtain the target visual data.

[0120] In an optional embodiment, the second processing module 508 is further configured to: Input the at least two types of related information and the intermediate visual data into the diffusion model; In the diffusion model, the intermediate visual data is denoised iteratively based on the at least two types of association information to obtain the target visual data.

[0121] In one optional embodiment, the at least two types of associated information include content information, edge information, and / or depth information of the region to be processed after the region content has been processed.

[0122] In an optional embodiment, the second determining module 506 is further configured to: Content extraction is performed on the intermediate visual data to obtain the content information of the region to be processed after the region content processing; Edge detection is performed on the intermediate visual data to obtain the edge information of the region to be processed after the region content processing; Depth detection is performed on the intermediate visual data to obtain the depth information of the region to be processed after the region content is processed.

[0123] In an optional embodiment, the second determining module 506 is further configured to: By using a content extraction model, the intermediate visual data is extracted to obtain the text content corresponding to the intermediate visual data; The text content is determined as the content information of the area to be processed after the area content has been processed.

[0124] In an optional embodiment, the second determining module 506 is further configured to: According to the edge detection algorithm, edge detection is performed on the intermediate visual data to obtain the edge information of the region to be processed after the region content processing.

[0125] In an optional embodiment, the second determining module 506 is further configured to: According to the depth detection algorithm, depth detection is performed on the intermediate visual data to obtain the depth information corresponding to the intermediate visual data; The depth information is determined as the depth information of the region to be processed after the region content is processed.

[0126] In an optional embodiment, the first processing module 504 is further configured to: Using a visual data processing model, the region to be processed is erased to obtain intermediate visual data corresponding to the visual data to be processed.

[0127] In an optional embodiment, the first determining module 502 is further configured to: Receive visual data to be processed sent by the client; In response to the processing instruction sent by the client for the visual data to be processed, the region to be processed included in the visual data to be processed is determined; The target visual data is sent to the client and displayed through the client's display interface.

[0128] In summary, by identifying the visual data to be processed and the regions to be processed within it, initial content processing is performed on these regions to obtain intermediate visual data. Then, based on this intermediate visual data, at least two types of associated information are identified. This information serves as guiding information to adjust the intermediate visual data, ultimately yielding the target visual data. This process adjusts the regions obtained from the initial content processing, improving issues such as blurriness, high distortion, and a smeared appearance in the intermediate visual data, thereby ensuring the quality of the adjusted target visual data.

[0129] The above is an illustrative scheme of a visual data processing device according to this embodiment. It should be noted that the technical solution of this visual data processing device and the technical solution of the above-described visual data processing method belong to the same concept. For details not described in detail in the technical solution of the visual data processing device, please refer to the description of the technical solution of the above-described visual data processing method.

[0130] See Figure 6 , Figure 6 A flowchart of an image erasing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0131] Step 602: Determine the image to be processed and the region to be processed included in the image to be processed; Step 604: Erasure the area to be processed to obtain an intermediate image corresponding to the image to be processed, wherein the intermediate image includes the erased area to be processed; Step 606: Based on the intermediate image, determine at least two types of associated information related to the intermediate image, wherein the at least two types of associated information are attribute information of the erased area to be processed; Step 608: Adjust the intermediate image according to the at least two types of association information to obtain the target image.

[0132] The at least two types of associated information include content information, edge information, and / or depth information of the region to be processed after the region content has been processed.

[0133] Specifically, by determining the visual data to be processed and the regions to be processed included in the visual data, the regions to be processed are first subjected to initial content processing to obtain intermediate visual data. Then, based on the intermediate visual data, at least two kinds of related information are determined. These at least two kinds of related information are used as guiding information to guide the adjustment of the intermediate visual data to obtain the target visual data. This achieves the adjustment of the regions to be processed obtained from the initial content processing and the regions to be processed after content processing, improving the problems of blurriness, high distortion, and smearing that may exist in the intermediate visual data, thereby ensuring the quality of the target visual data obtained after adjustment.

[0134] The above is an illustrative scheme of an image erasing method according to this embodiment. It should be noted that the technical solution of this image erasing method belongs to the same concept as the technical solution of the above-described visual data processing method. For details not described in detail in the technical solution of the image erasing method, please refer to the description of the technical solution of the above-described visual data processing method.

[0135] Corresponding to the above method embodiments, this specification also provides embodiments of an image erasing device. Figure 7 A schematic diagram of an image erasing device according to one embodiment of this specification is shown. Figure 7 As shown, the device includes: The first determining module 702 is configured to determine the image to be processed and the region to be processed included in the image to be processed; The first processing module 704 is configured to perform an erasure process on the area to be processed to obtain an intermediate image corresponding to the image to be processed, wherein the intermediate image includes the erased area to be processed. The second determining module 706 is configured to determine at least two types of associated information related to the intermediate image based on the intermediate image, wherein the at least two types of associated information are attribute information of the erased area to be processed; The second processing module 708 is configured to adjust the intermediate image based on the at least two types of association information to obtain the target image.

[0136] Specifically, by determining the visual data to be processed and the regions to be processed included in the visual data, the regions to be processed are first subjected to initial content processing to obtain intermediate visual data. Then, based on the intermediate visual data, at least two kinds of related information are determined. These at least two kinds of related information are used as guiding information to guide the adjustment of the intermediate visual data to obtain the target visual data. This achieves the adjustment of the regions to be processed obtained from the initial content processing and the regions to be processed after content processing, improving the problems of blurriness, high distortion, and smearing that may exist in the intermediate visual data, thereby ensuring the quality of the target visual data obtained after adjustment.

[0137] The above is a schematic representation of an image erasing device according to this embodiment. It should be noted that the technical solution of this image erasing device and the technical solution of the aforementioned visual data processing method belong to the same concept. Details not described in detail in the technical solution of the image erasing device can be found in the description of the technical solution of the aforementioned visual data processing method.

[0138] Corresponding to the above method embodiments, this specification also provides an image erasing system, including an edge device and a cloud-side device, wherein, The edge device is configured to send the visual data to be processed to the cloud-side device; The cloud-side device is configured to receive visual data to be processed and determine the region to be processed included in the visual data to be processed; perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing; determine at least two kinds of association information of the intermediate visual data based on the intermediate visual data, wherein the at least two kinds of association information are attribute information of the region to be processed after region content processing; adjust the intermediate visual data based on the at least two kinds of association information to obtain target visual data, and send the target visual data to the edge device.

[0139] Specifically, by determining the visual data to be processed and the regions to be processed included in the visual data, the regions to be processed are first subjected to initial content processing to obtain intermediate visual data. Then, based on the intermediate visual data, at least two kinds of related information are determined. These at least two kinds of related information are used as guiding information to guide the adjustment of the intermediate visual data to obtain the target visual data. This achieves the adjustment of the regions to be processed obtained from the initial content processing and the regions to be processed after content processing, improving the problems of blurriness, high distortion, and smearing that may exist in the intermediate visual data, thereby ensuring the quality of the target visual data obtained after adjustment.

[0140] The above is an illustrative scheme of an image erasing system according to this embodiment. It should be noted that the technical solution of this image erasing system and the technical solution of the above-described visual data processing method belong to the same concept. For details not described in detail in the technical solution of the image erasing system, please refer to the description of the technical solution of the above-described visual data processing method.

[0141] Figure 8 A structural block diagram of a computing device 800 according to one embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.

[0142] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0143] In one embodiment of this application, the aforementioned components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0144] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.

[0145] The processor 820 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above method.

[0146] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0147] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0148] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0149] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0150] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.

[0151] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0152] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0153] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0154] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0155] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A visual data processing method, comprising: Determine the visual data to be processed and the regions to be processed included in the visual data to be processed; The region to be processed is subjected to region content processing to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing. Based on the intermediate visual data, at least two types of related information are determined, wherein the at least two types of related information are attribute information of the region to be processed after the region content processing; Based on the at least two types of related information, the intermediate visual data is adjusted to obtain the target visual data.

2. The method according to claim 1, wherein adjusting the intermediate visual data based on the at least two types of association information to obtain target visual data includes: Feature extraction is performed on the at least two types of related information to obtain the related information features of the at least two types of related information; Feature extraction is performed on the intermediate visual data to obtain the intermediate visual data features; The associated information features and the intermediate visual data features are combined and calculated to obtain the target visual data.

3. The method according to claim 1, wherein adjusting the intermediate visual data based on the at least two types of association information to obtain target visual data includes: Input the at least two types of related information and the intermediate visual data into the diffusion model; In the diffusion model, the intermediate visual data is denoised iteratively based on the at least two types of association information to obtain the target visual data.

4. The method according to claim 1, wherein the at least two types of associated information include content information, edge information, and / or depth information of the region to be processed after the region content processing.

5. The method according to claim 4, wherein determining at least two types of association information of the intermediate visual data based on the intermediate visual data includes: Content extraction is performed on the intermediate visual data to obtain the content information of the region to be processed after the region content processing; and / or Edge detection is performed on the intermediate visual data to obtain the edge information of the region to be processed after the region content processing; and / or Depth detection is performed on the intermediate visual data to obtain the depth information of the region to be processed after the region content is processed.

6. The method according to claim 5, wherein extracting content from the intermediate visual data to obtain content information of the region to be processed after content processing includes: By using a content extraction model, the intermediate visual data is extracted to obtain the text content corresponding to the intermediate visual data; The text content is determined as the content information of the area to be processed after the area content has been processed.

7. The method according to claim 5, wherein performing depth detection on the intermediate visual data to obtain depth information of the region to be processed after content processing includes: According to the depth detection algorithm, depth detection is performed on the intermediate visual data to obtain the depth information corresponding to the intermediate visual data; The depth information is determined as the depth information of the region to be processed after the region content is processed.

8. The method according to any one of claims 1-7, wherein performing region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed includes: Using a visual data processing model, the region to be processed is erased to obtain intermediate visual data corresponding to the visual data to be processed.

9. The method according to any one of claims 1-7, wherein determining the visual data to be processed and the region to be processed included in the visual data to be processed comprises: Receive visual data to be processed sent by the client; In response to the processing instruction sent by the client for the visual data to be processed, the region to be processed included in the visual data to be processed is determined; After obtaining the target visual data, the process also includes: The target visual data is sent to the client and displayed through the client's display interface.

10. An image erasing method, comprising: Determine the image to be processed and the region to be processed included in the image to be processed; The area to be processed is erased to obtain an intermediate image corresponding to the image to be processed, wherein the intermediate image includes the erased area to be processed; Based on the intermediate image, at least two types of associated information about the intermediate image are determined, wherein the at least two types of associated information are attribute information of the erased area to be processed; The intermediate image is adjusted based on the at least two types of related information to obtain the target image.

11. The method according to claim 10, wherein the at least two types of associated information include content information, edge information, and / or depth information of the region to be processed after the region content processing.

12. An image erasing system, comprising an edge device and a cloud-side device, wherein, The edge device is configured to send the visual data to be processed to the cloud-side device; The cloud-side device is configured to receive visual data to be processed and determine the region to be processed included in the visual data to be processed; perform region content processing on the region to be processed to obtain intermediate visual data corresponding to the visual data to be processed, wherein the intermediate visual data includes the region to be processed after region content processing; determine at least two kinds of association information of the intermediate visual data based on the intermediate visual data, wherein the at least two kinds of association information are attribute information of the region to be processed after region content processing; adjust the intermediate visual data based on the at least two kinds of association information to obtain target visual data, and send the target visual data to the edge device.

13. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.