Vision large model and neural network cooperative graph data processing method and system
Patent Information
- Application Number
- CN202611283644.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-25
AI Technical Summary
该方式在精度上具有一定优势,但高度依赖人工操作,处理效率低,不适用于海量商业图像的批量化处理
[0044]采用上述的技术方案,本发明与现有技术相比,其具有的有益效果为:本方案突破了复杂边缘与材质的抠图瓶颈(兼顾全局与局部),通过两级网络的解耦与协同,第一级视觉基础模型作为“指南针”确保不丢失全局语义主体并避免背景误判,第二级精修网络则专注于局部高频细节的深度挖掘。该架构使得系统能够精准处理发丝、透明材质、反光物体等传统单端网络难以攻克的高难场景,输出边缘过渡自然、无光晕伪影的高质量结果。
Smart Images

Figure CN122821143A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer vision technology, image processing technology, neural network inference technology, and image matting technology, and particularly to a graph data processing method and system based on the collaboration of large visual models and neural networks. Background Technology
[0002] Image matting is an important technique in computer vision and image editing. Its purpose is to estimate the alpha value (transparency) of each pixel in an input image relative to the foreground object, thereby separating the foreground object from the background. In e-commerce, film and television post-production, advertising design, and digital content generation, it is often necessary to perform batch matting on a large number of product images, people images, or object images, requiring high edge naturalness, fidelity of transparent areas, and high-quality image synthesis.
[0003] Existing image matting methods typically include manual interactive matting methods and automated neural network matting methods. Manual interactive matting methods usually require users to manually provide a ternary map, which is used to label and determine the foreground region, the background region, and the unknown region. This method has certain advantages in accuracy, but it is highly dependent on manual operation, has low processing efficiency, and is not suitable for batch processing of massive amounts of commercial images.
[0004] While automated neural network matting methods can reduce manual intervention, they still have significant shortcomings in complex scenes. On one hand, single-stage end-to-end networks often struggle to simultaneously achieve global semantic accuracy and local edge detail. When the input image contains complex objects such as hair strands, animal fur, transparent glass, semi-transparent gauze, or reflective materials, the model is prone to issues like missing foreground structures, edge artifacts, color bleeding, or inaccurate estimation of transparent regions. On the other hand, while some two-stage matting methods incorporate coarse segmentation results and refinement networks, they typically rely on fixed-radius spatial operations such as morphological dilation and erosion when generating unknown regions. Fixed spatial operations cannot adaptively adjust based on image resolution, local boundary complexity, and model prediction confidence, easily leading to unknown regions that are too large or too small. When the unknown region is too large, the refinement network introduces more background noise and increases computational overhead; when the unknown region is too small, it may miss true edge details, reducing the accuracy of alpha transparency prediction.
[0005] Furthermore, existing general-purpose visual models suffer from insufficient domain generalization capabilities in specific vertical business scenarios. For example, under specific lighting conditions, with specific product materials, specific shooting backgrounds, or images with special transparent materials, general-purpose models are prone to domain shifts, making it difficult to consistently meet the stability, accuracy, and consistency requirements of commercial delivery. Existing technologies typically lack a closed-loop mechanism to continuously feed enterprise-specific business data back into the model training and update process, making it difficult for the system to continuously improve its processing performance as business data accumulates.
[0006] Therefore, it is necessary to provide a new graph data processing method and system that can combine the global semantic understanding capability of large visual models and the local detail refinement capability of high-resolution neural networks without the need for manual provision of ternary graphs, and guide the refinement calculation area through an adaptive spatial probability field, thereby taking into account automated processing efficiency, global semantic stability, local edge accuracy, and adaptability to specific business scenarios. Summary of the Invention
[0007] In view of this, the purpose of this invention is to propose a graph data processing method and system based on the collaboration of a large visual model and a neural network, which is reliable in implementation, flexible in application, and can take into account processing efficiency, global semantic stability, local edge accuracy and adaptability to specific business scenarios.
[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows: A graph data processing method based on the collaboration of large visual models and neural networks, comprising: S1. Obtain the original image to be processed, perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor suitable for neural network calculation; determine the semantic processing scale according to the resolution of the original image to be processed, and generate a semantic processing tensor for input to the first-level visual large model, while retaining the high-resolution image tensor corresponding to the resolution of the original image to be processed. S2. Input the semantic processing tensor into the first-level visual large model, extract global semantic features through the first-level visual large model, and output a foreground probability map representing the probability that each pixel belongs to the target foreground, and use the foreground probability map as a global semantic mask. S3. Based on the foreground probability map, the semantic processing scale, and the local gradient information of the foreground probability map, calculate the classification uncertainty and boundary uncertainty of each pixel, and generate an adaptive spatial probability field according to the classification uncertainty and boundary uncertainty. The adaptive spatial probability field is used to divide the image region into a foreground region, a background region, and a region with unknown dynamic width. S4. Spatially align the adaptive spatial probability field with the high-resolution image tensor and fuse them in the channel dimension to obtain a joint feature tensor; S5. Input the joint feature tensor into the second-level high-resolution refinement neural network, generate a focus mask based on the dynamic width unknown region, and perform multi-scale feature fusion, attention inference and high-frequency detail prediction within the spatial range corresponding to the focus mask to obtain a high-precision Alpha channel prediction map. S6. Combine the high-precision Alpha channel prediction image with the original image to be processed to output the image data processing result including the transparency channel.
[0009] As one possible implementation, further, in step S1 of this solution, the semantic processing scale is determined in the following manner: Let the tensor of the original image to be processed be... Its size is The longest edge threshold is The pixel area threshold is Then determine the scaling factor:
[0010] The working resolution is obtained based on the scale factor:
[0011]
[0012] in, and These represent the height and width of the original image, respectively. and These represent the working resolution height and width, respectively. Indicates the scaling factor. This indicates a round-down operation.
[0013] In step S1, when the original image to be processed meets the ultra-high resolution processing conditions, the tensor of the original image is divided into blocks with a side length of... The overlap width is Multiple image patches, with a step size of for adjacent image patches. The model outputs of each image patch are normalized and fused according to the overlap weights to obtain a full-image prediction result that is aligned with the spatial position of the original image.
[0014] As a possible implementation, further, in step S2 of this scheme, the first-level visual large model includes a general pre-trained visual backbone network and a segmentation-decoding network; during the model training phase, the training samples include RGB images and their corresponding binary masks, and the binary masks are normalized to obtain the mask ground truth. ,in:
[0015] In the formula, This represents a label mask with pixel values ranging from 0 to 255.
[0016] In step S2, the first-level visual large model outputs the logits tensor. And obtain the foreground probability map using the Sigmoid function:
[0017] The training loss of the first-level visual large model includes at least binary cross-entropy loss, soft Dice loss, and edge-weighted binary cross-entropy loss, where the soft Dice loss is:
[0018] in, Indicates the first Foreground prediction probability of 1 pixel, Indicates the first The true value of the mask in pixels. To prevent smoothing factors with a denominator of zero.
[0019] As one possible implementation, further, in step S3 of this solution, the classification uncertainty is calculated in the following manner: Based on foreground probability map Calculate the marginal confidence level:
[0020] Calculate the binary Shannon entropy:
[0021] And calculate the normalized classification uncertainty:
[0022] in, Indicates the marginal confidence level. Represents the binary Shannon entropy. This indicates uncertainty in normalized classification.
[0023] In step S3, the boundary uncertainty is calculated as follows: Foreground probability map PPerform horizontal and vertical Sobel convolutions respectively to obtain the horizontal gradient. and vertical gradient And calculate the probability gradient magnitude:
[0024] Further, we obtain boundary uncertainties:
[0025] in, Indicates the magnitude of the probability gradient. This represents the gradient normalization scaling parameter. This indicates uncertainty at the boundary.
[0026] As a preferred implementation method, preferably, in step S3 of this solution, the classification uncertainty is determined... and boundary uncertainty Calculate the overall uncertainty :
[0027] or
[0028] in, and Let represent the weights of classification uncertainty and boundary uncertainty, respectively. ; Based on the comprehensive uncertainty And the semantic processing scale determines the pixel-level unknown bandwidth. and make the pixel-level unknown bandwidth It increases with the increase of the overall uncertainty; Pixels that meet the high confidence condition for the foreground are marked as defined foreground regions, and pixels that meet the high confidence condition for the background are marked as defined background regions. Pixels located in the transition range between the foreground and background and with corresponding pixel-level unknown bandwidth are also marked. The covered pixels are marked as regions of unknown dynamic width, thereby generating the adaptive spatial probability field.
[0029] As one possible implementation, further, in step S4 of this solution, let the high-resolution image tensor be... Its spatial dimensions are The number of channels is The adaptive spatial probability field is Its spatial dimensions are The number of channels is Then, it is first obtained through spatial alignment mapping:
[0030] in, This represents the bilinear resampling space alignment mapping function. This represents a probabilistic prior tensor with the same tensor space size as the high-resolution image. Then, by splicing the channels, we get:
[0031] in, Let the joint feature tensor have the following number of channels: ; This is a channel splicing function.
[0032] and through Convolution performs learnable channel projection on the joint feature tensor to obtain the projected joint features. :
[0033] in, for Convolution operation function.
[0034] As a possible implementation, further, in step S5 of this solution, the focusing mask includes a hard focusing mask and a context focusing mask.
[0035] In step S5 of this scheme, a hard focusing mask is generated based on the dynamically unknown region in the adaptive spatial probability field. ,in, Represents pixels It belongs to the region with an unknown dynamic width. Represents pixels It does not belong to the region with unknown dynamic width; For the hard focusing mask Perform a radius of The extended processing yields the context focus mask:
[0036] in, Indicates the context expansion radius, This indicates an extension operation.
[0037] In step S5 of this scheme, the second-level high-resolution refinement neural network performs gating processing on intermediate features based on the context focusing mask:
[0038] in, Indicates the features after gating. This represents element-wise multiplication. This indicates that the mask will be expanded to match the number of feature channels.
[0039] As one possible implementation, further, in step S5 of this solution, the final Alpha channel is synthesized in the following manner:
[0040] in, Represents pixels The final alpha transparency value, Represents pixels Context focus mask, This represents the refinement transparency value output by the second-level high-resolution refinement neural network within the context-focused region. This represents a rough transparency value obtained from the output of determining the foreground region, determining the background region, or the first-level visual large model.
[0041] Step S6 of this solution also includes: updating the parameters of the first-level visual large model and the second-level high-resolution refined neural network based on the labeled business samples.
[0042] Based on the above, this solution also proposes a graph data processing system based on the collaboration of a large visual model and a neural network. This system applies the aforementioned graph data processing method based on the collaboration of a large visual model and a neural network, and includes: The image preprocessing module is used to acquire the original image to be processed, perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor, and generate a semantic processing tensor for input to the first-level visual large model and a high-resolution image tensor corresponding to the resolution of the original image to be processed. The feature extraction module is used to input the semantic processing tensor into the first-level visual large model to extract global semantic features and output a foreground probability map; The spatial probability field generation module is used to calculate classification uncertainty and boundary uncertainty based on the foreground probability map, semantic processing scale and local gradient information, and generate an adaptive spatial probability field containing a foreground region, a background region and a region with unknown dynamic width. The feature alignment and fusion module is used to spatially align and channel-fuse the adaptive spatial probability field with the high-resolution image tensor to obtain a joint feature tensor. The high-precision alpha prediction module is used to input the joint feature tensor into the second-level high-resolution refinement neural network and perform local high-frequency detail prediction based on the unknown dynamic width region to obtain a high-precision alpha channel prediction map. The image synthesis module is used to synthesize the high-precision Alpha channel prediction image with the original image to be processed, and output the image data processing result including the transparency channel; The closed-loop fine-tuning module is used to update the parameters of the first-level visual large model and the second-level high-resolution refined neural network based on labeled business samples.
[0043] As a preferred implementation method, the high-precision alpha prediction module of this solution preferably includes: The focus region determination unit is used to determine the dynamically unknown region based on the adaptive spatial probability field, and to generate a hard focus mask and a context focus mask; The computation scheduling unit is used to perform loss domain masking, feature gating, attention masking or region pruning scheduling based on the hard focusing mask or context focusing mask, so that the computation focus of the second-level high-resolution refinement neural network is concentrated on the region with unknown dynamic width. The Alpha synthesis unit is used to synthesize the refined transparency value output by the second-level high-resolution refinement neural network with the coarse transparency value corresponding to a defined region to obtain the full-image Alpha channel.
[0044] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: This solution breaks through the bottleneck of complex edge and material matting (taking into account both global and local aspects). Through the decoupling and collaboration of two-level networks, the first-level visual basic model acts as a "compass" to ensure that the global semantic subject is not lost and to avoid background misjudgment, while the second-level refinement network focuses on the in-depth mining of local high-frequency details. This architecture enables the system to accurately handle high-difficulty scenes such as hair strands, transparent materials, and reflective objects, which are difficult for traditional single-end networks to handle, and outputs high-quality results with natural edge transitions and no halo artifacts.
[0045] In addition, this scheme eliminates the spatial errors and computational waste caused by a fixed receptive field. Through an adaptive spatial probability map generation mechanism, it dynamically adjusts the width of the unknown region based on the image working resolution and the local network prediction confidence, overcoming the problems of "under-containment (loss of details)" or "over-containment (introduction of a large amount of background noise)" that are easily caused by traditional fixed-size morphological operations in images of different resolutions or regions with abrupt feature changes. This scheme not only significantly improves the feature focusing efficiency in the refinement stage but also significantly enhances the final prediction accuracy.
[0046] In its implementation, this approach eliminates the need for manual interaction in providing Trimaps or prior annotations, enabling fully automated batch processing and significantly reducing labor costs in commercial applications. Furthermore, thanks to the strong constraint of dynamically adaptive unknown regions, the computationally intensive second-level refinement network only needs to perform dense computations within a very small dynamic candidate region, thus balancing computational resource consumption with the performance requirements of high-resolution output.
[0047] Unlike rigid, open-source, general-purpose models that cannot be flexibly iterated upon, this solution's system architecture natively supports closed-loop, continuous fine-tuning training on enterprise-level proprietary business datasets. This allows the system model to self-evolve as data from vertical business domains (such as specific product categories) accumulates, effectively solving the long-tail feature recognition problem in specific business scenarios. This makes the system's generalization ability in specific domains superior to general-purpose models, and also builds a competitive technological barrier and data flywheel effect for application enterprises. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a schematic diagram illustrating the implementation process of the graph data processing method based on the collaboration of large visual models and neural networks in this solution; Figure 2 This is one of the schematic diagrams of auto parts obtained by the processing method of this solution; Figure 3 This is the second schematic diagram of the auto parts obtained by the processing method of this solution; Figure 4 This is a schematic diagram of the unit module connections of the graph data processing system based on the collaboration of a large visual model and a neural network in this solution. Detailed Implementation
[0050] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] like Figure 1 As shown in the figure, this embodiment proposes a graph data processing method based on the collaboration of a large visual model and a neural network, which includes: S1. Obtain the original image to be processed, perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor suitable for neural network calculation; determine the semantic processing scale according to the resolution of the original image to be processed, and generate a semantic processing tensor for input to the first-level visual large model, while retaining the high-resolution image tensor corresponding to the resolution of the original image to be processed. S2. Input the semantic processing tensor into the first-level visual large model, extract global semantic features through the first-level visual large model, and output a foreground probability map representing the probability that each pixel belongs to the target foreground, and use the foreground probability map as a global semantic mask. S3. Based on the foreground probability map, the semantic processing scale, and the local gradient information of the foreground probability map, calculate the classification uncertainty and boundary uncertainty of each pixel, and generate an adaptive spatial probability field according to the classification uncertainty and boundary uncertainty. The adaptive spatial probability field is used to divide the image region into a foreground region, a background region, and a region with unknown dynamic width. S4. Spatially align the adaptive spatial probability field with the high-resolution image tensor and fuse them in the channel dimension to obtain a joint feature tensor; S5. Input the joint feature tensor into the second-level high-resolution refinement neural network, generate a focus mask based on the dynamic width unknown region, and perform multi-scale feature fusion, attention inference and high-frequency detail prediction within the spatial range corresponding to the focus mask to obtain a high-precision Alpha channel prediction map. S6. Combine the high-precision Alpha channel prediction image with the original image to be processed to output the image data processing result including the transparency channel.
[0052] This embodiment's solution, based on a graph data processing method that combines a large visual model with a neural network, can be used to automatically perform image matting on input images and output an RGBA image with an opacity channel.
[0053] As a specific explanation of the above steps S1-S6, this embodiment may include the following technical details: S1, Image Acquisition and Preprocessing; The system receives the original image (original high-resolution image) to be processed from the system front-end, image processing platform, or API interface. The original image to be processed can be an RGB image, and its content can include people, goods, animals, transparent objects, semi-transparent objects, or reflective objects.
[0054] Let the tensor of the original image corresponding to the original image to be processed be: ;in, Indicates the image height. 3 represents the image width, and 3 represents the three RGB color channels.
[0055] The original image tensor is subjected to size transformation, color space conversion, and channel normalization to transform it into a tensor format suitable for neural network computation.
[0056] Channel normalization can be expressed as:
[0057] in, Indicates the first The pixel values of each color channel. Indicates the first The mean of each color channel, Indicates the first The standard deviation of each color channel This represents the normalized channel value.
[0058] Since direct full-resolution forward processing of the original image is costly, and simply scaling it down would weaken high-frequency boundary priors such as hair strands, this scheme employs adaptive downsampling or block-based Tiling; while the second stage still uses the original resolution. (Or alignment features) are processed in a way that carries high frequencies.
[0059] For ultra-high resolution images, to avoid excessive memory consumption and decreased inference efficiency caused by directly inputting full-resolution data into the first-level visual model, this scheme determines the semantic processing scale based on the image size. The longest edge threshold is set as follows: The pixel area threshold is , Indicates the image height. The scale factor represents the image width. It can be determined as follows:
[0060] According to the scaling factor Get the working resolution:
[0061]
[0062] in, Indicates the working resolution height. Indicates the working resolution width. This represents the floor operation. The semantic processing tensor is obtained after proportional scaling. The semantic processing tensor is used as input to the first-level visual large model.
[0063] In this scheme, bilinear resampling can be written as a four-neighbor convex combination (with coefficients summing to 1), and the semantic processing tensor... At the target pixel • Neighboring pixels + … + • Neighboring pixels ( , (for the fractional part), its effect is equivalent to varying with the scale factor. The low-pass filter is reduced and enhanced, resulting in slightly blurred details but preserving the main semantics.
[0064] In another implementation, when the local details of the input image are dense or the aspect ratio is not suitable for direct proportional scaling, an overlapping block processing strategy can be used. Let the block side length be... The overlap width is Then the step size between adjacent image blocks is: .
[0065] Crop the original image into multiple sizes. The image patches are processed and input into the model for forward inference. For overlapping regions, the output of each image patch is fused using weight normalization.
[0066]
[0067] in, Indicates the first The model output for each image patch This represents the spatial weight of the image patch during the fusion process. Represents the sum of overlapping weights. This represents the predicted result after fusion, aligned with the spatial position of the original image. Overlapping block processing reduces block boundary discontinuities and lowers the inference load per run while maintaining high-resolution details.
[0068] In this step, the semantic processing tensor is used to extract global semantic information in the first-level visual large model, while the high-resolution image tensor is reserved for the subsequent second-level refinement network to predict local details.
[0069] S2, global semantic mask extraction based on the first-level visual large model; The semantic processing tensor obtained in step S1 is input into the first-level visual large model. The first-level visual large model may include a pre-trained Transformer visual backbone network, a deep convolutional neural network backbone network, or a combination of the two to form a visual segmentation network, and is connected to a segmentation decoding structure to output pixel-level foreground probabilities. That is, the global semantic features of the image are extracted using the deep receptive field of the first-level visual large model.
[0070] In this step, the first-level visual large model predicts and outputs single-channel or multi-channel probability maps that define the foreground and background regions, namely the "Global Semantic Mask", so that it can provide reliable semantic priors in complex lighting and cluttered backgrounds.
[0071] The first-level visual large model outputs the logits tensor. The foreground probability map is obtained through the Sigmoid function and is defined as follows:
[0072] in, or This represents the probability that each pixel belongs to the target foreground. This represents the Sigmoid function.
[0073] If the output resolution of the first-level visual model differs from the resolution of the original image, then bilinear interpolation is used to adjust the foreground probability map. Align the image to the original spatial dimensions. The aligned foreground probability map serves as a global semantic mask, used to identify the approximate region of the target foreground in the input image and to provide semantic priors for subsequent adaptive spatial probability field generation.
[0074] During the model training phase, paired training samples can be used, with each training sample consisting of an RGB image and a corresponding binary labeled mask. Let the original labeled mask be... If the pixel value ranges from 0 to 255, then the true value of the mask Y can be expressed as: ;in, .
[0075] During data training, the pixel values of the training samples are first normalized to... Then, the mean and variance of the pre-trained model are used for standardization to align the data distribution with the distribution during model pre-training.
[0076] Channel normalization can be expressed as:
[0077] in, Indicates the first The pixel values of each color channel (0-255). Indicates the first The mean of each color channel, Indicates the first The standard deviation of each color channel This represents the normalized channel value.
[0078] To improve the model's robustness to foreground targets, reduce background noise interference, and enhance generalization ability, this scheme also performs data augmentation through random background augmentation, the formula of which is defined as:
[0079] in, For the augmented output image, For future goals, The background image is a randomly sampled image. The mask for the foreground target (which is synonymous with the target region, takes the value 0 or 1, and can also be a soft mask). This is pixel-by-pixel multiplication.
[0080] In this step of data augmentation, the foreground target is fused with the random background through a mask to generate diverse training samples.
[0081] For geometric transformations (such as rotation, flipping, scaling, cropping, etc.), the image must be modified accordingly. and tags Execution is synchronized to ensure a one-to-one correspondence between spatial locations and to avoid label misalignment.
[0082] The training loss for a first-level visual large model can include binary cross-entropy loss, soft Dice loss, and edge-weighted binary cross-entropy loss.
[0083] Binary cross-entropy loss (BCE With Logits loss) can be determined by directly receiving the model's raw output (logits).
[0084] The soft Dice loss can be expressed as:
[0085] in, Indicates the first Foreground prediction probability of 1 pixel, Indicates the first The true value of the mask in pixels. It is a smoothing factor used to avoid the denominator being zero.
[0086] Edge-weighted binary cross-entropy loss can be applied based on the labeled mask. Edge map generation. Specifically, for Sobel edge detection is performed to obtain the edge weight map. Then reconstruct the weights, which are defined as follows:
[0087] in, This represents the edge weighting coefficient.
[0088] This solution will use a weighted graph. Multiplying the BCE loss pixel by pixel achieves weighted amplification of the loss in edge regions. At the same time, the edge-weighted loss enables the first-level visual model to output a more stable probability distribution in the target boundary region.
[0089] In this scheme, the system as a whole adopts a weighted multi-loss fusion strategy, and the total loss is the weighted sum of the losses of each branch, which is defined as:
[0090] in, For the model number k The output of each branch (such as the prediction results at different levels). For the corresponding tags, These are the weighting coefficients for the loss of each branch. It is a single-branch loss function, which consists of binary cross-entropy loss, soft Dice loss, and edge-weighted binary cross-entropy loss.
[0091] S3, Adaptive Spatial Probability Field Generation; Foreground probability map based on the output of step S2 An adaptive spatial probability field is generated. This adaptive spatial probability field replaces the traditional fixed-radius ternary image generation method and dynamically determines the width of the unknown region based on the prediction confidence of each pixel, the boundary complexity, and the image scale.
[0092] First, calculate the foreground probability map. The marginal confidence of each pixel is defined as follows:
[0093] in, The larger the value of C, the more certain it is that the pixel belongs to the foreground or background; The smaller the value, the less certain the classification result of that pixel is.
[0094] Then, calculate the binary Shannon entropy: And further, the normalized classification uncertainty is obtained:
[0095]
[0096] in, Represents the binary Shannon entropy. This represents the uncertainty of the normalized classification. The larger the value of U, the more unstable the foreground / background classification at that pixel, or the lower the confidence of the classification prediction at that pixel.
[0097] Secondly, for the foreground probability map Perform horizontal and vertical Sobel convolutions respectively to obtain the horizontal gradient. and vertical gradient And calculate the probability gradient magnitude, which is defined as:
[0098] in, This indicates the intensity of local changes in the foreground probability map. When the foreground probability changes drastically in the local space, it usually indicates that the region is located near the target boundary, i.e., Large The neighborhood changes drastically, which constitutes a potential semantic boundary.
[0099] Furthermore, the boundary uncertainty is calculated, and it is defined as follows:
[0100] in, This represents the gradient normalization scaling parameter. This indicates uncertainty at the boundary.
[0101] Based on classification uncertainty and boundary uncertainty The overall uncertainty is obtained Defined as:
[0102] in, Represents the classification uncertainty weights. Represents the boundary uncertainty weights, and .
[0103] Based on comprehensive uncertainty Generate pixel-level unknown bandwidth at the image working scale .
[0104] In one implementation, pixel-level unknown bandwidth It can be represented as:
[0105] in, Indicates the minimum unknown bandwidth. Indicates the maximum unknown bandwidth. This represents the scaling factor determined by the image resolution or working scale. The larger the overall uncertainty, or the higher the image resolution, the larger the corresponding unknown bandwidth; the smaller the overall uncertainty, or the smoother the local boundaries, the smaller the corresponding unknown bandwidth.
[0106] Subsequently, based on the foreground probability diagram Comprehensive uncertainty and pixel-level unknown bandwidth The image is divided into a defined foreground region, a defined background region, and a region with unknown dynamic width.
[0107] Specifically, pixels that satisfy high foreground probability and low uncertainty are identified as foreground regions, pixels that satisfy low foreground probability and low uncertainty are identified as background regions, and pixels that are within the transition range between foreground and background and are covered by pixel-level unknown bandwidth are identified as dynamic width unknown regions.
[0108] The adaptive spatial probability field generated in the above manner can dynamically adjust the range of unknown regions according to different image resolutions and different local boundary complexities, thereby providing more accurate computational region constraints for subsequent high-resolution refinement networks.
[0109] S4, Feature Alignment and Fusion; Spatial alignment and channel fusion are performed between the adaptive spatial probability field output in step S3 and the high-resolution image tensor retained in step S1.
[0110] Let the high-resolution image tensor be... Its size is ,in, This indicates the number of image channels or the number of shallow feature channels.
[0111] Let the adaptive spatial probability field be... Its size is ,in, This represents the number of channels in the probability field.
[0112] when Space dimensions and When there is inconsistency, Spatial alignment is defined as follows:
[0113] in, Spatial alignment mapping can be achieved using bilinear resampling; This represents a probabilistic prior tensor with the same tensor space size as the high-resolution image.
[0114] Then, the high-resolution image tensor With probability prior tensor The concatenation is performed along the channel dimension, and is defined as follows:
[0115] in, This is a channel splicing function. Denotes the joint characteristic tensor, whose size is This joint feature tensor simultaneously contains high-resolution texture information from the original image and semantic probability priors from the output of the first-level visual large model.
[0116] The pixel-wise vector form is defined as: That is, the same location Up Dimensional color / features and dimensional probability prior is listed as The operator is a parameterless information aggregation, and the first layer of the second-level convolution can be linearly mixed with it.
[0117] In a further embodiment, it can be achieved through... Convolution performs learnable channel projection, which is defined as:
[0118] in, for Convolution function, This represents the joint features after projection. Through learnable channel projection, the second-level high-resolution refinement neural network can adaptively adjust the weight relationship between image texture features and spatial probability priors, that is, to... Channel projection is This enables learnable channel reweighting.
[0119] The tensor shapes before and after fusion can be seen in the table below:
[0120] S5, Alpha prediction based on a second-level high-resolution refined neural network; The joint feature tensor obtained in step S4 is input into the second-level high-resolution refinement neural network. The second-level high-resolution refinement neural network is used to make fine predictions of local high-frequency edges, transparent transition regions, and fine structures within regions of unknown dynamic width.
[0121] Regarding the discrete unknown band indication, it is assumed that the output of step S3 assigns a tri-state or equivalent label to each pixel: foreground FG determined, background BG determined, and unknown UNK; or outputs continuous unknown probabilities. ∈ [0,1].
[0122] First, a hard focusing mask is generated based on the adaptive spatial probability field. When pixels When the region is a dynamically unknown area, let: When pixel When it does not belong to the region with unknown dynamic width, let: .
[0123] Since there can still be weights with high centers and low edges within the unknown band, for example, a smooth weight mask can be used. ( (for low-pass kernels), or for Directly cut and normalize to obtain ∈ [0,1], used to smooth gradient discontinuities caused by hard boundaries.
[0124] To enable the second-stage high-resolution refining neural network to obtain the necessary contextual information around unknown regions, a hard focusing mask can be applied. Perform extended processing (for) Make radius Morphological dilatation), which is defined as:
[0125] in, Indicates a context-focused mask. Indicates an extension operation. This refers to the context spread radius. The context spread radius can be set based on image resolution, unknown bandwidth, or network receptive field size.
[0126] During the inference phase, a context-focusing mask can be used to gate intermediate features, defined as follows:
[0127] in, Indicates the features after gating. This represents element-wise multiplication. This means expanding the mask to match the number of feature channels. Through feature gating, the network can suppress feature responses in non-interested regions, allowing the main computation to focus on regions with unknown dynamic widths and their contextual neighborhood.
[0128] During the training phase, region weights can be applied to the Alpha supervision loss. Let the pixel-level Alpha supervision loss be... The region-weighted loss can then be expressed as:
[0129] in, The region weight map can be represented by either a hard focus mask or a soft focus weight. Indicates the normalization factor; This represents the weak constraint loss for a defined region, used to keep the Alpha value of the defined foreground region close to 1 and the Alpha value of the defined background region close to 0. This indicates a weak constraint weight.
[0130] In another implementation, it can be based on a hard focusing mask. The minimum bounding rectangle is used to generate the Region of Interest (ROI), and the main forward computation of the second-level high-resolution refinement neural network is performed within the ROI. The alpha value of the region outside the ROI can be directly obtained from the coarse alpha value generated by determining the foreground region, determining the background region, or the output of the first level. This method can further reduce the invalid computation of the high-resolution refinement network.
[0131] The second-level high-resolution refinement neural network outputs refinement transparency values. The final alpha channel can be synthesized as follows:
[0132] in, Represents pixels The final alpha transparency value, This represents the refinement transparency value output by the second-level high-resolution refinement neural network. This represents a rough transparency value obtained from the output of determining the foreground region, determining the background region, or the first-level visual large model.
[0133] Through this synthesis method, regions with unknown dynamic widths are precisely predicted by a second-level high-resolution refinement neural network, while regions with known widths are stably supplemented by coarse results, thus ensuring that every pixel in the entire image has a definite Alpha value.
[0134] S6, image post-processing, RGBA output and closed-loop fine-tuning; The final Alpha channel obtained in step S5 Compared with the original input image The image is composited to obtain an RGBA image that includes the alpha channel. Its definition is:
[0135] in, These represent the original input image in pixels. The red, green, and blue channel values at that location. This indicates the corresponding transparency value.
[0136] Before output, edge smoothing, isolated noise removal, and transparency range cropping can be performed on the Alpha channel to make the edge transition of the output image more natural and reduce halos, jagged edges, and background color penetration.
[0137] in, Figure 2 , Figure 3 A schematic diagram of the auto parts obtained by the processing method of this solution is shown.
[0138] To enhance the system's adaptability to specific business scenarios, image processing results confirmed by users, manual correction results, and semi-automatic annotation results are continuously collected to construct a proprietary business fine-tuning dataset. This dataset may include product images with specific lighting styles, transparent material product images, reflective material product images, images of people with complex backgrounds, and other business-representative samples.
[0139] Based on the proprietary business fine-tuning dataset, low-rank or full-parameter fine-tuning can be performed on the first-level large-scale visual model and the second-level high-resolution refined neural network. In the low-rank fine-tuning approach, only a small number of adaptation parameters in the model are updated to reduce training costs and minimize disruption to the original general capabilities. In the full-parameter fine-tuning approach, all parameters of the model can be updated based on a large-scale business sample to enhance the model's adaptability to specific business scenarios.
[0140] Through a closed-loop fine-tuning mechanism, the system can continuously absorb new samples from business scenarios, reduce the impact of domain offset on the stability of model output, and continuously improve the processing quality of complex products, complex materials, and complex edge images.
[0141] This application's image data processing solution achieves a balance between high precision and full automation, eliminating the need for manual provision of trimmaps and enabling end-to-end automated high-quality image matting to meet the batch processing needs of massive commercial images. This solution addresses the challenge of matting complex materials and edges by employing a two-level "coarse-to-fine" network collaborative architecture. This resolves artifacts and structural loss issues that arise when handling delicate edges such as transparency, semi-transparency, and hair-like details in existing methods, balancing global semantic integrity with high precision in local details. Furthermore, this solution eliminates the computational and precision losses associated with fixed spatial operations by innovatively introducing an "adaptive spatial probability field generation mechanism," replacing traditional fixed-radius morphological operations. This allows the model to dynamically and adaptively generate unknown regions based on image resolution and local boundary uncertainties, improving feature focusing efficiency and final accuracy.
[0142] This solution also overcomes the domain offset problem of general large models by building a mechanism that supports continuous fine-tuning and optimization of proprietary business data. This enables the system to continuously iterate and improve the robustness of image matting in specific business scenarios, thereby enhancing the generalization ability of specific domains and forming a data closed loop.
[0143] Combination Figure 4 As shown, based on the above, this solution also proposes a graph data processing system based on the collaboration of a large visual model and a neural network. This system applies the aforementioned graph data processing method based on the collaboration of a large visual model and a neural network, and includes: The image preprocessing module is used to acquire the original image to be processed, perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor, and generate a semantic processing tensor for input to the first-level visual large model and a high-resolution image tensor corresponding to the resolution of the original image to be processed. The feature extraction module is used to input the semantic processing tensor into the first-level visual large model to extract global semantic features and output a foreground probability map; The spatial probability field generation module is used to calculate classification uncertainty and boundary uncertainty based on the foreground probability map, semantic processing scale and local gradient information, and generate an adaptive spatial probability field containing a foreground region, a background region and a region with unknown dynamic width. The feature alignment and fusion module is used to spatially align and channel-fuse the adaptive spatial probability field with the high-resolution image tensor to obtain a joint feature tensor. The high-precision alpha prediction module is used to input the joint feature tensor into the second-level high-resolution refinement neural network and perform local high-frequency detail prediction based on the unknown dynamic width region to obtain a high-precision alpha channel prediction map. The image synthesis module is used to synthesize the high-precision Alpha channel prediction image with the original image to be processed, and output the image data processing result including the transparency channel; The closed-loop fine-tuning module is used to update the parameters of the first-level visual large model and the second-level high-resolution refined neural network based on labeled business samples.
[0144] Specifically, the closed-loop fine-tuning module is used to continuously collect image samples, manual correction results, or semi-manual annotation results in specific business scenarios, and to perform low-rank fine-tuning or full-parameter fine-tuning on the first-level visual large model and the second-level high-resolution refined neural network based on the samples, so that the system can adapt to long-tail materials, special lighting and shadows, and complex edge distribution in specific business scenarios.
[0145] As a preferred implementation method, the high-precision alpha prediction module of this solution preferably includes: The focus region determination unit is used to determine the dynamically unknown region based on the adaptive spatial probability field, and to generate a hard focus mask and a context focus mask; The computation scheduling unit is used to perform loss domain masking, feature gating, attention masking or region pruning scheduling based on the hard focusing mask or context focusing mask, so that the computation focus of the second-level high-resolution refinement neural network is concentrated on the region with unknown dynamic width. The Alpha synthesis unit is used to synthesize the refined transparency value output by the second-level high-resolution refinement neural network with the coarse transparency value corresponding to a defined region to obtain the full-image Alpha channel.
[0146] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0148] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A graph data processing method based on the collaboration of large visual models and neural networks, characterized in that, It includes: S1. Obtain the original image to be processed, and perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor suitable for neural network calculation; The semantic processing scale is determined based on the resolution of the original image to be processed, and a semantic processing tensor is generated for input to the first-level visual large model, while retaining the high-resolution image tensor corresponding to the resolution of the original image to be processed. S2. Input the semantic processing tensor into the first-level visual large model, extract global semantic features through the first-level visual large model, and output a foreground probability map representing the probability that each pixel belongs to the target foreground, and use the foreground probability map as a global semantic mask. S3. Based on the foreground probability map, the semantic processing scale, and the local gradient information of the foreground probability map, calculate the classification uncertainty and boundary uncertainty of each pixel, and generate an adaptive spatial probability field according to the classification uncertainty and boundary uncertainty. The adaptive spatial probability field is used to divide the image region into a foreground region, a background region, and a region with unknown dynamic width. S4. Spatially align the adaptive spatial probability field with the high-resolution image tensor and fuse them in the channel dimension to obtain a joint feature tensor; S5. Input the joint feature tensor into the second-level high-resolution refinement neural network, generate a focus mask based on the dynamic width unknown region, and perform multi-scale feature fusion, attention inference and high-frequency detail prediction within the spatial range corresponding to the focus mask to obtain a high-precision Alpha channel prediction map. S6. Combine the high-precision Alpha channel prediction image with the original image to be processed to output the image data processing result including the transparency channel.
2. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 1, characterized in that, In step S1, the semantic processing scale is determined as follows: Let the tensor of the original image to be processed be... Its size is The longest edge threshold is The pixel area threshold is Then determine the scaling factor: The working resolution is obtained based on the scale factor: in, and These represent the height and width of the original image, respectively. and These represent the working resolution height and width, respectively. This indicates a round-down operation. Indicates the scaling factor; When the original image to be processed meets the conditions for ultra-high resolution processing, the tensor of the original image is divided into blocks with a side length of... The overlap width is Multiple image patches, with a step size of for adjacent image patches. The model outputs of each image patch are normalized and fused according to the overlap weights to obtain a full-image prediction result that is aligned with the spatial position of the original image.
3. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 1, characterized in that, In step S2, the first-level visual large model includes a general pre-trained visual backbone network and a segmentation-decoding network; during the model training phase, the training samples include RGB images and their corresponding binary masks, and the binary masks are normalized to obtain the mask ground truth. ,in: This represents a marker mask with pixel values ranging from 0 to 255; The first-level visual large model outputs a logits tensor. The foreground probability map is obtained through the Sigmoid function. : The training loss of the first-level visual large model includes at least binary cross-entropy loss, soft Dice loss, and edge-weighted binary cross-entropy loss, wherein the soft Dice loss... Defined as: in, Indicates the first Foreground prediction probability of 1 pixel, Indicates the first The true value of the mask in pixels. To prevent smoothing factors with a denominator of zero.
4. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 1, characterized in that, In step S3, the classification uncertainty is calculated as follows: Based on foreground probability map Calculate the marginal confidence level: Calculate the binary Shannon entropy: And calculate the normalized classification uncertainty: in, Indicates the marginal confidence level. Represents the binary Shannon entropy. This indicates uncertainty in normalized classification; The boundary uncertainty is calculated as follows: Foreground probability map P Perform horizontal and vertical Sobel convolutions respectively to obtain the horizontal gradient. and vertical gradient And calculate the probability gradient magnitude: Further, we obtain boundary uncertainties: in, Indicates the magnitude of the probability gradient. This represents the gradient normalization scaling parameter. This indicates uncertainty at the boundary.
5. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 4, characterized in that, In step S3, based on the classification uncertainty... and boundary uncertainty Calculate the overall uncertainty : or in, and Let represent the weights of classification uncertainty and boundary uncertainty, respectively. ; Based on the comprehensive uncertainty And the semantic processing scale determines the pixel-level unknown bandwidth. and make the pixel-level unknown bandwidth It increases with the increase of the overall uncertainty; Pixels that meet the high confidence condition for the foreground are marked as defined foreground regions, and pixels that meet the high confidence condition for the background are marked as defined background regions. Pixels located in the transition range between the foreground and background and with corresponding pixel-level unknown bandwidth are also marked. The covered pixels are marked as regions of unknown dynamic width, thereby generating the adaptive spatial probability field.
6. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 1, characterized in that, In step S4, let the high-resolution image tensor be... Its spatial dimensions are The number of channels is The adaptive spatial probability field is Its spatial dimensions are The number of channels is Then, it is first obtained through spatial alignment mapping: in, This represents the bilinear resampling space alignment mapping function. This represents a probabilistic prior tensor with the same tensor space size as the high-resolution image. Then, by splicing the channels, we get: in, Let the joint feature tensor have the number of channels as . ; This is a channel splicing function; and through Convolution performs learnable channel projection on the joint feature tensor to obtain the projected joint features. : in, for Convolution operation function.
7. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 6, characterized in that, In step S5, the focusing mask includes a hard focusing mask and a context focusing mask; A hard focusing mask is generated based on the dynamically wide unknown region in the adaptive spatial probability field. ,in, Represents pixels It belongs to the region with an unknown dynamic width. Represents pixels It does not belong to the region with unknown dynamic width; For the hard focusing mask Perform a radius of The extended processing yields the context focus mask. Its definition is: in, Indicates the context expansion radius, Indicates an extension operation; The second-level high-resolution refinement neural network performs gating processing on intermediate features based on the context focusing mask: in, Indicates the features after gating. This represents element-wise multiplication. This indicates that the mask will be expanded to match the number of feature channels.
8. The graph data processing method based on the collaboration of large visual models and neural networks as described in claim 7, characterized in that, In step S5, the final alpha channel is synthesized as follows: in, Represents pixels The final alpha transparency value, Represents pixels Context focus mask, This represents the refinement transparency value output by the second-level high-resolution refinement neural network within the context-focused region. This represents a rough transparency value obtained from the output of determining the foreground region, determining the background region, or the first-level visual large model. Step S6 further includes updating the parameters of the first-level visual large model and the second-level high-resolution refined neural network based on the labeled business samples.
9. A graph data processing system based on the collaboration of a large visual model and a neural network, wherein the graph data processing method based on the collaboration of a large visual model and a neural network as described in any one of claims 1 to 8 is characterized in that, It includes: The image preprocessing module is used to acquire the original image to be processed, perform size transformation, color space conversion and normalization on the original image to be processed to obtain an image tensor, and generate a semantic processing tensor for input to the first-level visual large model and a high-resolution image tensor corresponding to the resolution of the original image to be processed. The feature extraction module is used to input the semantic processing tensor into the first-level visual large model to extract global semantic features and output a foreground probability map; The spatial probability field generation module is used to calculate classification uncertainty and boundary uncertainty based on the foreground probability map, semantic processing scale and local gradient information, and generate an adaptive spatial probability field containing a foreground region, a background region and a region with unknown dynamic width. The feature alignment and fusion module is used to spatially align and channel-fuse the adaptive spatial probability field with the high-resolution image tensor to obtain a joint feature tensor. The high-precision alpha prediction module is used to input the joint feature tensor into the second-level high-resolution refinement neural network and perform local high-frequency detail prediction based on the unknown dynamic width region to obtain a high-precision alpha channel prediction map. The image synthesis module is used to synthesize the high-precision Alpha channel prediction image with the original image to be processed, and output the image data processing result including the transparency channel; The closed-loop fine-tuning module is used to update the parameters of the first-level visual large model and the second-level high-resolution refined neural network based on labeled business samples.
10. The graph data processing system based on the collaboration of a large visual model and a neural network as described in claim 9, characterized in that, The high-precision Alpha prediction module includes: The focus region determination unit is used to determine the dynamically unknown region based on the adaptive spatial probability field, and to generate a hard focus mask and a context focus mask; The computation scheduling unit is used to perform loss domain masking, feature gating, attention masking or region pruning scheduling based on the hard focusing mask or context focusing mask, so that the computation focus of the second-level high-resolution refinement neural network is concentrated on the region with unknown dynamic width. The Alpha synthesis unit is used to synthesize the refined transparency value output by the second-level high-resolution refinement neural network with the coarse transparency value corresponding to a defined region to obtain the full-image Alpha channel.