Intelligent matting method and device, medium, equipment and product

By combining a target matting model and an edge inpainting network, the problems of long processing time and poor adaptability of existing matting techniques are solved, achieving efficient and accurate image and video matting, which is suitable for applications in multiple fields.

CN121010621APending Publication Date: 2025-11-25HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511544820.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing image matting technologies are time-consuming and rely on manual labor, which cannot meet the needs of batch processing. They are particularly inefficient when dealing with hair and transparent materials, and have poor adaptability to complex scenes, making it difficult to achieve real-time image matting and high precision.

Method used

A pre-trained target matting model is used to process the object to be processed, generate mask data, and repair it through an edge repair network. The matting results are optimized by combining user operation commands to adapt to different target types and scene features.

Benefits of technology

It achieves efficient and accurate image and video cutout, improves batch processing efficiency, reduces manual intervention, supports real-time cutout and maintains high accuracy in complex scenes, and reduces user operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010621A_ABST
    Figure CN121010621A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and particularly provides an intelligent matting method and device, a medium, equipment and a product, and the method can comprise the steps: obtaining a to-be-processed object after preprocessing; wherein the object to be processed is an image or a video; processing the to-be-processed object by using a target matting model to generate mask data; wherein the target matting model comprises a feature positioning module, a feature reconstruction module and a prediction module; and repairing the mask data, and generating and outputting a matting result matched with a preset parameter. According to the embodiment of the invention, the longitude and efficiency of matting can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a method, apparatus, medium, device, and product for intelligent image matting. Background Technology

[0002] Background removal and cutout technology is one of the core applications in the field of image processing. This technology is widely used in e-commerce design (changing product backgrounds), film and television post-production (green screen cutout), live streaming interaction (virtual background replacement), UI / UX design (element extraction), and other scenarios.

[0003] Currently, traditional image cutout techniques primarily rely on manual methods. Specifically, this involves manually adjusting edges pixel by pixel using tools such as the Pen Tool and Quick Selection Tool in Photoshop. Especially when dealing with hair or transparent materials (such as wedding dresses or glass), a single image can take several minutes, making it unsuitable for batch processing. Furthermore, the cutout effect is entirely dependent on human experience.

[0004] Therefore, how to provide an efficient and accurate intelligent image matting method has become an urgent technical problem to be solved. Summary of the Invention

[0005] The purpose of some embodiments of this application is to provide a method, apparatus, medium, device and product for intelligent image cutout. The technical solutions of the embodiments of this application can realize intelligent image cutout, improve the efficiency and accuracy of image cutout, and have high practicality.

[0006] In a first aspect, some embodiments of this application provide an intelligent image matting method, comprising: obtaining a preprocessed object to be processed; wherein the object to be processed is an image or a video; processing the object to be processed using a target matting model to generate mask data; wherein the target matting model includes a feature localization module, a feature reconstruction module, and a prediction module; repairing the mask data to generate and output a matting result that matches preset parameters; wherein the matting result is a matted image or a matted video.

[0007] Some embodiments of this application process the object to be processed using a target matting model to generate mask data, and then repair the mask data to generate the matting result. Some embodiments of this application can achieve intelligent matting, enabling batch processing of image or video matting tasks, improving matting efficiency and accuracy, and demonstrating high practicality.

[0008] In some embodiments, obtaining the preprocessed object to be processed includes: acquiring three-channel data of the original visual data; preprocessing the three-channel data to obtain the object to be processed; wherein the type of preprocessing includes noise reduction, resolution adjustment, and color normalization.

[0009] Some embodiments of this application preprocess the three-channel data of the original visual data to obtain the object to be processed, providing data support for subsequent accurate image matting.

[0010] In some embodiments, the step of processing the object to be processed using the target matting model to generate mask data includes: using the feature localization module to locate the object to be processed and outputting multi-scale features; using the feature reconstruction module to iteratively optimize the multi-scale features and output enhanced features; and using the prediction module to integrate the enhanced features and output the mask data.

[0011] Some embodiments of this application use multiple modules in the target matting model to perform different processing on the object to be processed and output mask data, thereby accurately identifying the target type and scene features in the object to be processed, and providing effective data support for subsequent matting.

[0012] In some embodiments, repairing the mask data and generating and outputting a matting result that matches preset parameters includes: inputting the mask data and original visual data into an edge repair network and outputting edge-processed data; performing region repair on the edge-processed data to obtain the matting result; wherein the region includes a vulnerability region and a redundant region.

[0013] Some embodiments of this application can perform edge repair on mask data through an edge repair network, followed by region repair, to solve the problems of edge discontinuity and background residue.

[0014] In some embodiments, performing region repair on the edge-processed data to obtain the matting result includes: identifying the defective regions and the redundant regions within the edge-processed data; filling the defective regions and deleting the redundant regions to obtain the matting result.

[0015] Some embodiments of this application can ensure the integrity of the foreground in an image or video by identifying and processing defective and redundant regions within the image or video frame.

[0016] In some embodiments, after generating and outputting a cutout result that matches preset parameters, the method further includes: generating an adjusted cutout result in response to a user's operation instruction; wherein the operation instruction is used to adjust the cutout result; and displaying the adjusted cutout result.

[0017] Some embodiments of this application optimize the cutout results through user operation commands to obtain adjusted cutout results, thereby obtaining high-precision adjusted cutout results.

[0018] Secondly, some embodiments of this application provide an intelligent image matting apparatus, comprising: an acquisition module for acquiring a preprocessed object to be processed; wherein the object to be processed is an image or video; a generation module for processing the object to be processed using a target matting model to generate mask data; wherein the target matting model includes a feature localization module, a feature reconstruction module, and a prediction module; and a repair module for repairing the mask data to generate and output a matting result matching preset parameters; wherein the matting result is a matted image or a matted video.

[0019] Thirdly, some embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.

[0020] Fourthly, some embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method as described in any embodiment of the first aspect.

[0021] Fifthly, some embodiments of this application provide a computer program product, the computer program product including a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of some embodiments of this application, the accompanying drawings used in some embodiments of this application will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 System diagrams for intelligent image matting provided for some embodiments of this application; Figure 2 One of the flowcharts of the intelligent image matting method provided for some embodiments of this application; Figure 3 The second flowchart of the intelligent image matting method provided for some embodiments of this application; Figure 4 Block diagrams of an intelligent image matting apparatus provided for some embodiments of this application; Figure 5 A schematic diagram of an electronic device provided for some embodiments of this application. Detailed Implementation

[0024] The technical solutions of some embodiments of this application will now be described with reference to the accompanying drawings.

[0025] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0026] With the increasing demand for digital content creation, users have raised higher requirements for the efficiency, accuracy, and scene adaptability of image matting. However, existing image matting technologies still have significant limitations. While existing AI image matting solutions based on semantic segmentation (such as Mask R-CNN and DeepLab series) can achieve automation, they still have the following problems: low accuracy in recognizing the boundary area between "hair background" and "transparent object background", easily resulting in "edge breaks" (such as missing hair strands) or "background residue" (such as background reflections on glass that have not been removed); weak scene adaptability: in non-standard scenes (such as dynamic backgrounds without green screens, low-light environments, and multiple overlapping objects), it is easy to misidentify the background as the foreground (such as identifying a swaying curtain as the edge of a person) or miss foreground details. Moreover, dynamic videos, live broadcasts, and other scenes require "real-time image matting" (frame rate ≥ 30fps), but existing models usually have a large number of parameters (such as deep learning models with hundreds of millions of parameters) in pursuit of accuracy, resulting in slow inference speed on ordinary terminal devices (such as laptops and mobile phones). If the number of parameters is reduced to improve speed, it will lead to a significant decrease in accuracy. Existing tools require users to manually set parameters (such as edge feathering values ​​and cutout range), and cannot automatically adapt strategies based on target type (such as "people," "products," and "animals") and scene characteristics (such as "static images," "dynamic videos," and "low-light conditions"). For example, processing "animal fur" and "hard goods" requires different edge optimization logic, but existing tools do not implement differentiated processing.

[0027] In view of this, some embodiments of this application provide an intelligent image matting method. This method can process the object to be processed using a pre-trained target matting model to generate mask data; then, the mask data is repaired to output the final matting result. Embodiments of this application can accurately process complex areas such as hair, transparent materials, and fine edges in the object to be processed without manual intervention, reducing background residue and foreground omissions; real-time matting of dynamic videos and live streaming scenes can be achieved on terminal devices (i.e., frame rate ≥ 30fps, single-frame inference time ≤ 100ms), ensuring consistency between frame results (no jitter, no jumps); by automatically identifying target type (people / products / animals) and scene features (static / dynamic / low light), it adapts to the corresponding matting strategy, eliminating the need for users to manually adjust parameters and improving scene adaptability; it also provides a "one-click" matting workflow, supporting batch processing of static images and real-time processing of dynamic videos, and is compatible with mainstream design tools and platforms, optimizing user experience and reducing user operating costs.

[0028] The following is in conjunction with the appendix Figure 1 The overall structure of the intelligent image matting system provided by some embodiments of this application is illustrated by way of example.

[0029] like Figure 1 As shown, some embodiments of this application provide an intelligent image matting system. This intelligent image matting system includes a terminal 100. A user 200 can start the matting system on the terminal 100 and upload raw visual data (such as single / batch static images, or local / real-time dynamic videos). The terminal 100 can automatically identify the type of the input raw visual data and preprocess it to obtain the object to be processed. Then, a pre-trained target matting model can process the object to be processed to obtain mask data. Finally, the mask data is repaired to generate and output a matted image or matted video with preset parameters.

[0030] In some embodiments of this application, the terminal 100 can be a mobile terminal or a non-portable computer terminal; users can launch the image matting system through a web interface (browser), client (Windows / Mac), or mobile terminal (iOS / Android); this application does not impose specific limitations on the embodiments. Furthermore, the target image matting model in this application embodiment can be pre-trained and deployed on the terminal 100.

[0031] The image cutout system in this application provides a RESTful API, supporting POST requests to upload data and retrieve cutout results, with an interface response time of ≤500ms (cloud processing). Based on this API, Photoshop and Figma plugins can be developed, allowing users to directly call the system's cutout function within their design tools without switching software. The system can also interface with e-commerce platforms and live streaming software to automate the entire process from "product upload → automatic cutout → details page generation" to "live stream startup → real-time cutout → virtual background replacement."

[0032] The following is in conjunction with the appendix Figure 2 The present application provides an exemplary embodiment of the intelligent image matting process performed by terminal 100.

[0033] Please see the appendix Figure 2 , Figure 2 A flowchart of an intelligent image matting method is provided for some embodiments of this application. The intelligent image matting method may include: S210, Obtain the preprocessed object to be processed; wherein, the object to be processed is an image or video.

[0034] Specifically, S210 may include: acquiring three-channel data of raw visual data; preprocessing the three-channel data to obtain the object to be processed; wherein the types of preprocessing include noise reduction, resolution adjustment and color normalization.

[0035] For example, in a specific embodiment of this application, the original visual data can be a single or batch of static images (JPG / PNG), or it can be a local dynamic video or a live video stream. After the user uploads the original visual data, the system can prompt the user to select the corresponding "application scenario" (e.g., e-commerce, live streaming, film and television, design, etc.) through a pop-up window; if the user does not select, the system automatically determines the scenario type to which the original visual data belongs through a lightweight classification model. The inference time of this classification model is ≤10ms, which is highly efficient.

[0036] For static images, the RGB three-channel data of the image is used. If the input is in PNG format and includes an alpha channel, the original alpha channel is automatically ignored (based on the result generated by this system). For dynamic video, RGB three-channel data of consecutive frames and inter-frame motion vectors are acquired synchronously to ensure inter-frame consistency. During processing by terminal 100, data can be transmitted in memory. If terminal 100 is a cloud device, it is transmitted encrypted using the HTTPS protocol to ensure data security.

[0037] The three channels of data are denoised, for example, to eliminate noise in low light conditions; resolution adaptation (i.e., resolution adjustment) is performed to adjust the input resolution according to the device performance, balancing speed and accuracy; color normalization is performed, for example, to unify brightness and contrast, and improve the stability of model recognition.

[0038] Specifically, for noise reduction: In low-light scenes, Non-Local Means (NLM) is used to eliminate Gaussian noise; in strong-light scenes, Adaptive Histogram Equalization (CLAHE) is used to avoid overexposure and loss of target features. Resolution adaptation: The input resolution is automatically adjusted according to device performance; for example, for mobile phones or laptops, the default input resolution is 1280×720px, while high-performance PCs / servers support 4K (3840×2160px) input; users can manually set the resolution (e.g., range: 320×240px~4096×4096px). Color normalization: RGB channel pixel values ​​are normalized from "0-255" to "0-1" and the mean of the ImageNet dataset (R=0.485, G=0.456, B=0.406) is subtracted to improve the recognition stability of the AI ​​semantic understanding model (as a specific example of a target matting model).

[0039] S220, The object to be processed is processed using a target matting model to generate mask data; wherein, the target matting model includes a feature localization module, a feature reconstruction module, and a prediction module.

[0040] For example, in a specific embodiment of this application, an AI semantic understanding model is used as input to the preprocessed object to be processed, i.e., video frames in an image or video, to complete semantic segmentation and mask generation. Specifically, the AI ​​semantic understanding model combines the target type and scene features in the image or video frame to generate an initial binary mask (1 for foreground, 0 for background, as a specific example of mask data). Compared with traditional semantic segmentation models, it can more accurately identify the blurred boundaries between "transparent object background" and "hair background". It is understandable that when the object to be processed is a video, it is necessary to process the video frames frame by frame, which is equivalent to processing images, but the video frames can be processed in batches, which is more efficient.

[0041] The AI ​​semantic understanding model comprises a Localization Module, a Reconstruction Module, and a Predictor Module. The model is trained using a constructed training dataset. This dataset includes over 200,000 labeled images, covering three main categories: people (different hairstyles, clothing), goods (hard, transparent, plush, etc.), and animals (densely furred, such as cats and dogs). Each category contains 150,000 static images and 50,000 dynamic video frames, with pixel-level annotation accuracy. The training strategy employs a "pre-training + LoRA fine-tuning" approach. Semantic understanding capabilities are first pre-trained on a large-scale general image dataset, then fine-tuned using LoRA on a dedicated matting dataset (with only 1.2 million new parameters), balancing model accuracy and lightweight design. The model's training accuracy metrics are: Intersection over Union (IoU) ≥ 92% for static image masks and IoU ≥ 90% for dynamic video frames (used to measure inter-frame consistency).

[0042] In practical applications, the object to be processed is first converted into a 512×512×4 feature map. The key regions of the target in the feature map are located through the feature localization module, and multi-scale features are output. The semantic differences between the target and the background (such as the texture features of hair and the light reflection features of transparent objects) are extracted from the multi-scale features through the feature reconstruction module to obtain the enhanced features. Finally, the prediction module combines the multi-degree fusion of the target features and the enhanced features to generate an initial binary mask of 1024×1024 (1 for foreground and 0 for background).

[0043] Specifically, in some embodiments of this application, S220 may include: S221, The feature localization module is used to locate the object to be processed and output multi-scale features.

[0044] For example, in a specific embodiment of this application, the localization module (i.e., the feature localization module) serves as the "vanguard" for semantic extraction, consisting of multiple layers of Transformer Blocks and auxiliary structures. After inputting the preprocessed image I (as a specific example of the object to be processed), the Transformer architecture captures global semantic information, quickly locating the key regions of the target in the image. For example, in complex scene images, it can efficiently identify the approximate location of subjects such as animals and objects, providing direction for subsequent fine processing. Simultaneously, features are optimized using components such as ASPP (Spatial Pyramid Pooling with Hollows), and the initial distinction between the target and the background is strengthened through classification supervision (LCE loss), outputting multi-scale features {F1l, F2l, F3l, Fe}, providing basic semantic clues for the reconstruction module. Alternatively, the preprocessed image can be converted into a 512×512×4 feature map, allowing the localization module to quickly locate the key regions of the target.

[0045] S222, The feature reconstruction module is used to iteratively optimize the multi-scale features and output the enhanced features.

[0046] For example, in a specific embodiment of this application, the reconstruction module (i.e., the feature reconstruction module) plays a "refining" role, using the BiRef Block as the core unit to iteratively optimize the multi-scale features output by the localization module. An internal and external reference mechanism is constructed through the BiRef Block: on the one hand, adaptive cropping patches from high-resolution images are used as "internal references" to preserve the original resolution and detailed features; on the other hand, gradient maps are introduced as "external references" to assist the model in capturing subtle structures such as hair strands and transparent materials. The multi-scale features are processed sequentially by the BiRef Block to generate enhanced features {F1d+, F2d+, F3d+}, enabling in-depth mining of semantic differences between the target and background, accurately distinguishing hair texture from the surrounding environment, and light reflection from transparent objects from occluded backgrounds.

[0047] S223, The enhanced features are integrated using the prediction module, and the mask data is output.

[0048] For example, in a specific embodiment of this application, the prediction module, as the final "output terminal," integrates the enhanced features from the reconstruction module to generate a semantic mask (as a specific example of mask data). Based on multi-scale feature fusion, the feature representation is further optimized through the Reconstruction Block, combining "Inward Reference" and "Outward Reference" mechanisms, ultimately outputting a 1024×1024 initial binary mask M (1 foreground, 0 for background). Under a 1280×720px input and GPU environment, the inference time can be controlled within ⩽50ms, meeting the efficiency requirements. Simultaneously, supervised training using loss functions such as BCE (Binary Cross-Entropy) and IoU (Intersection over Union) (LBCE, LIoU) ensures the accuracy of mask generation, making the foreground and background segmentation more closely match the real boundaries.

[0049] S230, the mask data is repaired to generate and output a cutout result that matches the preset parameters; wherein, the cutout result is a cutout image or a cutout video.

[0050] For example, in some embodiments of this application, edge defects and other details existing in the generated initial binary mask M are repaired, and a cutout result conforming to preset parameters is output. These preset parameters can be automatically output based on user needs, such as outputting PNG format (with an alpha channel) for static images. For another example, static images are output in PNG format (with an alpha channel) by default, while also supporting user selection of JPG (white or transparent background) or WebP (adjustable compression ratio), with the resolution consistent with the input or scaled according to user needs; for batch-processed images or videos, compressed packages (ZIP / RAR) are supported, automatically named with "original filename + cutout" for easy user organization.

[0051] In some embodiments of this application, S230 may include: S231, the mask data and the original visual data are input into the edge repair network, and the edge-processed data is output.

[0052] For example, in a specific embodiment of this application, an edge refinement network (as a specific example of an edge repair network, based on the U-Net++ lightweight architecture) is activated to perform pixel-level optimization on the edges of hair, hair strands, and transparent materials. Through "edge detection → pixel completion → transparency gradient adjustment," the problems of "edge discontinuity" and "background residue" are solved (such as accurately removing background pixels between hair strands and retaining a semi-transparent effect at glass edges). The following describes the processing procedure of the edge refinement network on the initial binary mask M and the object to be processed: Edge detection: Edge pixels (edge ​​width ≤ 3 pixels) of the initial binary mask are extracted using the Canny edge detection algorithm, and the area to be optimized is marked; Pixel-level optimization: Differentiated strategies are adopted for different edge types; Hair edges: Single hair strands are identified using the hair strand detection algorithm to supplement missing hair strand pixels in the initial mask, while removing residual background pixels between hair strands; Transparent edges: The Alpha channel value is adjusted according to the RGB brightness gradient of the original image (i.e., the original visual data) (e.g., the Alpha value of glass edges is set to 128-192 to achieve a semi-transparent effect); Smoothing: Gaussian blur (radius 1-2 pixels) is used to smooth the edges to avoid "jagged edges".

[0053] S232, perform region repair on the edge-processed data to obtain the matting result; wherein, the region includes a vulnerability region and a redundant region.

[0054] For example, in a specific embodiment of this application, gaps (such as small foreground omissions) or redundancies (such as background misjudged as foreground) in the mask of the edge-processed data are identified, and the corresponding pixels are automatically filled or deleted to ensure the integrity of the foreground.

[0055] In some embodiments of this application, S232 may include: identifying the vulnerability region and the redundant region within the edge-processed data; filling the vulnerability region and deleting the redundant region to obtain the matting result.

[0056] For example, in a specific embodiment of this application, small-area gaps (area < 100 pixels) in the mask are identified by “connected region analysis”, and the gaps are automatically filled according to the color and texture of the surrounding foreground pixels; redundant areas in the mask that are not semantically related to the foreground (such as misjudging small objects in the background as foreground) are identified and deleted after confirmation by “semantic matching” (i.e., comparing the semantic features of the redundant areas with those of the foreground).

[0057] In some embodiments of this application, after executing S230, the intelligent image cutout method may further include: generating an adjusted image cutout result in response to a user's operation instruction; wherein the operation instruction is used to adjust the image cutout result; and displaying the adjusted image cutout result.

[0058] For example, in a specific embodiment of this application, terminal 100 provides user 200 with lightweight web or client-side editing tools. Users can use these tools to adjust the cutout result and obtain the adjusted cutout result. For instance, the brush tool supports foreground brushes (marking the background as the foreground) and background brushes (marking the foreground as the background), with adjustable brush size (1-20 pixels); edge feathering supports 0-20 pixel edge feathering to soften the transition between the foreground and the new background; simultaneously, it saves the last 5 adjustment operations based on user commands, supports undoing or redoing, adjusts the cutout result, and provides a real-time preview of the adjusted cutout result, with a delay of ≤50ms.

[0059] During the image matting process in this application, the system can also monitor the CPU / GPU usage of terminal 100 in real time. If the usage of terminal 100 is >80%, the number of model parameters is automatically reduced (e.g., switching to "lightweight mode" reduces the number of model parameters by 50%) to ensure inference speed. If the usage is <50%, "high-precision mode" is automatically enabled. Mobile devices can enable "low-power mode" to reduce battery consumption by reducing the number of model inferences (e.g., inference once every 2 frames in a live streaming scene, with the mask generated by interpolation algorithm in the middle frames). During the image matting process, all data for each matting step is automatically recorded, such as input data features, such as target type, scene parameters (lighting or dynamic range), and input resolution; output result data, such as mask IoU (automatically calculated by the system), user fine-tuning records (fine-tuning area, parameter adjustment), and output format; system performance data, such as inference time, device type, and CPU / GPU usage. All data is timestamped and stored locally or in the cloud (after user authorization). The data retention period is set to a default number of days (e.g., 30 days, which can be customized).

[0060] Furthermore, recording data during the matting process allows for iterative optimization of the relevant models. For example, low-precision cases (e.g., IoU < 85%) and frequently fine-tuned cases by users can be selected from historical matting data as training samples for model iteration. A reinforcement learning algorithm is used, with the reward function being the fewest user fine-tuning attempts and the highest IoU. Every 10,000 valid data points accumulated, the target matting model and edge refinement network undergo LoRA fine-tuning (fine-tuning time ≤ 2 hours, cloud GPU environment). After each iteration, the IoU of the target scene can improve by 1-2%, and the user fine-tuning rate can decrease by 5-8%.

[0061] The following is in conjunction with the appendix Figure 3 The specific process of intelligent image matting provided by some embodiments of this application is illustrated by way of example.

[0062] Please see the appendix Figure 3 , Figure 3 A flowchart of an intelligent image matting method provided for some embodiments of this application.

[0063] The above process is illustrated below by example.

[0064] S310 acquires three channels of raw visual data.

[0065] S320 preprocesses the three-channel data to obtain the object to be processed.

[0066] S330 uses a target matting model to process the object to be processed and generate mask data.

[0067] S340 inputs mask data and raw visual data into the edge inpainting network and outputs edge-processed data.

[0068] S350 identifies vulnerable and redundant areas within the data after edge processing, fills the vulnerable areas and deletes the redundant areas to obtain the matting result.

[0069] S360 responds to user commands, generating and displaying the adjusted cutout results.

[0070] It is understood that the specific implementation process of S310~S360 can be referred to the method implementation examples provided above. To avoid repetition, it will not be described again here.

[0071] It should be noted that the aforementioned intelligent image cutout method can be applied in the e-commerce field. For example, e-commerce merchants need to generate background-free images in batches for products such as clothing, accessories, and electronic products for use in product detail page design and multi-scene poster composition (such as changing the background of clothing in different scenes). The intelligent image cutout process described in this application can automatically identify the product type (such as processing the plush edges of "cotton clothing" and retaining the transparent texture of "glass water cups"). Batch processing 100 product images takes only 3 minutes, and the edge precision meets the high-definition display requirements of e-commerce platforms. It supports integration with the "Merchant Material Center" of e-commerce platforms, and the image cutout results are directly synchronized to the backend without manual uploading.

[0072] The aforementioned intelligent background removal method can be applied to dynamic background removal scenes in film and television post-production. For example, in film and television production, it is necessary to remove the background of "people or props filmed without a green screen" and replace them with special effects backgrounds (such as science fiction scenes or historical scenes); the 4K resolution video processing precision of this application meets the needs of professional post-production, and the edge details of people's hair and clothing wrinkles are complete; it supports integration with film and television software such as Premiere and After Effects, and the background removal results can be directly imported into the timeline, reducing the format conversion process.

[0073] The aforementioned intelligent image cutout method can also be applied to element extraction scenarios in the design field. For example, UI designers need to extract icons and illustrations from reference images for app interface design; this application automatically identifies the boundary between the "icon or illustration" and the background, preserves the alpha channel in the extraction results, and can be directly dragged into design tools such as Figma and Sketch without manually adjusting the transparency; it supports the simultaneous extraction of multiple elements (such as extracting multiple icons from one image at the same time), improving efficiency many times over.

[0074] The aforementioned intelligent background removal method can also be applied to MR e-commerce shopping guide scenarios. For example, when users wear MR glasses to select clothing, home appliances, and other products in physical stores or warehouses, the real background of the products needs to be removed in real time, and a virtual usage scene (such as clothing worn on a virtual model or the placement effect of home appliances in the user's home) needs to be superimposed to assist in purchasing decisions; the product background removal can be completed within 1 second using TOF depth data, with the spatial offset between clothing and virtual models being ≤1 pixel, and the relative position of the product and the virtual scene remaining consistent when the user turns their head; it supports fine-tuning of product size with gestures, and the results are directly synchronized to the company's MR e-commerce platform without manual operation, improving shopping guide efficiency by more than 3 times.

[0075] The aforementioned intelligent image matting method can also be applied to MR remote collaboration scenarios. For example, when factory workers wear MR glasses to troubleshoot equipment malfunctions, the operating area (such as equipment parts) needs to be matted in real time for remote experts to overlay virtual repair annotations (such as disassembly animations) to avoid the annotations being obscured by the real background. By automatically distinguishing between the real operating area and virtual annotations, only the real background is removed, and the matted area moves synchronously with the parts, ensuring that the annotations always fit the surface of the parts. It supports integration with the company's MR collaboration software, improving remote collaboration efficiency by 50% and reducing troubleshooting time by 40%.

[0076] As can be seen from the above embodiments of this application, the fully automatic image cutout process of this application shortens the processing time of a single static image from "several minutes" to "≤2 seconds", and improves the batch processing efficiency by more than 100 times; the combination of AI semantic understanding and edge refinement network improves the cutout accuracy of hair and transparent materials to more than 95%, reduces edge breakage and background residue problems by 80%, and eliminates the need for manual secondary correction; by automatically identifying target types and scene features, it adapts to differentiated cutout strategies, and maintains stable accuracy (30% improvement compared to traditional AI cutout) in complex scenes such as "no green screen dynamic background", "low light", and "multiple target overlap", with strong scene adaptability; at the same time, one-click operation and lightweight fine-tuning tools reduce the usage threshold for non-professional users; it supports multi-format output and integration with external tools, adapting to the needs of e-commerce, film and television, design and other fields; it reduces the labor cost of manual cutout and secondary correction, and the terminal device adaptation capability reduces the dependence on high-configuration hardware, allowing small and medium-sized businesses and individual users to use professional-grade cutout functions at low cost.

[0077] Please refer to Figure 4 , Figure 4 The diagram shows a block diagram of an intelligent image matting apparatus provided in some embodiments of this application. It should be understood that this intelligent image matting apparatus corresponds to the method embodiments described above and is capable of performing the various steps involved in the method embodiments described above. The specific functions of this intelligent image matting apparatus can be found in the description above, and detailed descriptions are appropriately omitted here to avoid repetition.

[0078] Figure 4 The intelligent matting device includes at least one software function module that can be stored in a memory or embedded in the intelligent matting device in the form of software or firmware. The intelligent matting device includes: an acquisition module 410, used to acquire a preprocessed object to be processed; wherein the object to be processed is an image or video; a generation module 420, used to process the object to be processed using a target matting model to generate mask data; wherein the target matting model includes a feature localization module, a feature reconstruction module, and a prediction module; and a repair module 430, used to repair the mask data, generate and output a matting result that matches preset parameters; wherein the matting result is a matted image or a matted video.

[0079] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0080] Some embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can perform the operation of any of the methods corresponding to the methods provided in the above embodiments.

[0081] Some embodiments of this application also provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operation of any of the methods corresponding to the above embodiments provided in the above embodiments.

[0082] like Figure 5 As shown, some embodiments of this application provide an electronic device 500, which includes a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520. When the processor 520 reads the program from the memory 510 via a bus 530 and executes the program, it can implement the methods of any of the above embodiments.

[0083] Processor 520 can process digital signals and can include various computing architectures. For example, it can be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 520 can be a microprocessor.

[0084] The memory 510 can be used to store instructions executed by the processor 520 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of this application. The processor 520 of this disclosure embodiment can be used to execute the instructions in the memory 510 to implement the methods shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.

[0085] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0087] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method of intelligent matting, the method comprising: The method comprises the following steps: obtaining a pre-processed object to be processed; wherein the object to be processed is an image or a video; processing the object to be processed by using a target matting model to generate mask data; wherein the target matting model comprises a feature positioning module, a feature reconstruction module and a prediction module; repairing the mask data to generate and output a matting result matching a preset parameter; wherein the matting result is a matting image or a matting video.

2. The method of claim 1, wherein, The method comprises the following steps: collecting three-channel data of original visual data; pre-processing the three-channel data to obtain the object to be processed; wherein the types of pre-processing include denoising, resolution adjustment and color normalization.

3. The method of claim 1 or 2, wherein, The method comprises the following steps: using the feature positioning module to perform target positioning on the object to be processed to output multi-scale features; using the feature reconstruction module to iteratively optimize the multi-scale features to output enhanced features; using the prediction module to integrate the enhanced features to output the mask data.

4. The method of claim 1 or 2, wherein, The method comprises the following steps: inputting the mask data and the original visual data into an edge repair network to output edge-processed data; performing regional repair on the edge-processed data to obtain the matting result; wherein the regions include a void region and a redundant region.

5. The method of claim 4, wherein, The method comprises the following steps: identifying the void region and the redundant region in the edge-processed data; filling the void region and deleting the redundant region to obtain the matting result.

6. The method of claim 1 or 2, wherein, After the generation and output of the matting result matching the preset parameter, the method further comprises the following steps: in response to a user's operation instruction, generating an adjusted matting result; wherein the operation instruction is used to adjust the matting result; displaying the adjusted matting result.

7. An apparatus for intelligent matting, the apparatus comprising: The method comprises the following steps: an obtaining module for obtaining a pre-processed object to be processed; wherein the object to be processed is an image or a video; a generating module for processing the object to be processed by using a target matting model to generate mask data; wherein the target matting model comprises a feature positioning module, a feature reconstruction module and a prediction module; a repairing module for repairing the mask data to generate and output a matting result matching a preset parameter; wherein the matting result is a matting image or a matting video.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is run by the processor to execute the method of any one of claims 1-6.

9. An electronic device, comprising: The computer readable storage medium stores a computer program, wherein the computer program is run by the processor to execute the method of any one of claims 1-6. The computer readable storage medium stores a computer program, wherein the computer program is run by the processor to execute the method of any one of claims 1-6.

10. A computer program product, characterised in that, The computer program product comprises a computer program, wherein the computer program, when executed by a processor, performs the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Video restoration method, device and equipment and storage medium

    CN112102212A

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN113658197A

  • Image restoration method, figure image restoration method, electronic equipment and storage medium

    CN118840295A

  • General matting method and system based on artificial intelligence

    CN119672058A

  • Image segmentation method, computing device, electronic device and storage medium

    CN120259349A

Cited By

  • Image processing system

    CN121685570A