Image sampling processing method and device, computer, storage medium and program product
Through the methods of guided filtering and similarity detection, the problem of insufficient restoration of image detail information in the existing technology is solved, and the accuracy and performance of high-resolution features are improved.
Patent Information
- Application Number
- CN202410317411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-19
AI Technical Summary
In the prior art, when high-resolution features are obtained by interpolating low-resolution features, image detail information cannot be effectively restored, resulting in poor HR feature effects.
By obtaining the guiding features and the features to be sampled, guided filtering is performed to determine the adjacent semantic similarity and detail similarity, and upsampling is performed in combination with the adjacent similarity to generate target features to predict the detection results of the image processing task.
It achieves controllable alignment of features of different resolutions in semantic content, restores image detail information, and improves the accuracy and performance of sampling processing.
Smart Images

Figure CN120672600A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image sampling and processing method, device, computer, storage medium, and program product. Background Art
[0002] As a fundamental element of deep network architectures, feature upsampling aims to restore the spatial resolution of low-resolution (LR) features to produce high-resolution (HR) features. This is widely used in image processing fields such as image segmentation. Currently, LR features are typically interpolated to produce HR features. However, this approach directly supplements LR features without adding detail. This results in a higher resolution HR feature, but lacks accurate image detail, making the HR feature less effective. Summary of the Invention
[0003] The embodiments of the present application provide an image sampling processing method, device, computer, storage medium, and program product, which can improve the accuracy and sampling performance of sampling processing.
[0004] On the one hand, an embodiment of the present application provides an image sampling processing method, the method comprising:
[0005] Obtaining the guiding features and the features to be sampled of the initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled;
[0006] Performing guided filtering on the guided features and the features to be sampled to obtain guided filtering features;
[0007] Perform mutual similarity detection on the guided filter feature and the feature to be sampled to determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels;
[0008] Denoising the guide feature to obtain the denoised feature, performing detail-aware self-alignment on the denoised feature and the guide feature to obtain the adjacent detail similarity corresponding to the feature to be sampled;
[0009] The adjacent semantic similarity and the adjacent detail similarity are combined into the adjacent similarity, and the adjacent similarity is used to upsample the unsampled features to obtain the target features. The image detection results of the image processing task are predicted based on the target features.
[0010] On the one hand, an embodiment of the present application provides an image sampling processing method, the method comprising:
[0011] Obtain image samples and image processing annotations of image samples in image processing tasks;
[0012] The image sample is input into the initial image processing model, the sample guiding feature of the image sample is obtained through the initial guiding network in the initial image processing model, and the sample to-be-sampled feature of the image sample is obtained through the initial task main network in the initial image processing model; the resolution of the sample guiding feature is greater than the resolution of the sample to-be-sampled feature;
[0013] Performing guided filtering on the sample guided features and the sample to-be-sampled features to obtain the sample guided filtering features;
[0014] Perform mutual similarity detection on the sample guided filter feature and the sample to be sampled feature to determine the sample adjacent semantic similarity corresponding to the sample to be sampled feature; the sample adjacent semantic similarity is used to represent the semantic similarity between each pixel in the sample to be sampled feature and the adjacent pixel of the pixel;
[0015] Denoising the sample guide feature to obtain the sample denoising feature, performing detail-aware self-alignment on the sample denoising feature and the sample guide feature to obtain the sample adjacency detail similarity corresponding to the sample to-be-sampled feature;
[0016] The sample adjacency semantic similarity and the sample adjacency detail similarity are combined to form the sample adjacency similarity, and the sample adjacency similarity is used to upsample the sample features to obtain the target sample features;
[0017] The target sample features are predicted through the initial task prediction network in the initial image processing model to obtain the sample detection results of the image processing task;
[0018] Through image processing annotation and sample detection results, the parameters of the initial image processing model are adjusted to obtain the image processing model.
[0019] In one aspect, an embodiment of the present application provides an image sampling and processing device, the device comprising:
[0020] A feature acquisition module is used to acquire the guiding features and the features to be sampled of the initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled;
[0021] A guided filtering module is used to perform guided filtering on the guided features and the features to be sampled to obtain guided filtering features;
[0022] The semantic parsing module is used to perform mutual similarity detection between the guided filter feature and the feature to be sampled, and determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels;
[0023] The detail parsing module is used to denoise the guide features to obtain denoised features, perform detail-aware self-alignment on the denoised features and the guide features, and obtain the adjacent detail similarity corresponding to the features to be sampled;
[0024] The feature sampling module is used to combine the adjacent semantic similarity and the adjacent detail similarity into adjacent similarity, and use the adjacent similarity to upsample the sampled features to obtain the target features;
[0025] The image processing module is used to predict the image detection results of the image processing task based on the target features.
[0026] The feature acquisition module is specifically used to:
[0027] Performing convolution processing on the initial image to obtain a first convolution feature, and performing normalization processing on the first convolution feature to obtain a normalized feature;
[0028] Perform convolution processing on the normalized features to obtain the guiding features of the initial image;
[0029] In the task main network corresponding to the image processing task, the features to be sampled of the initial image are obtained.
[0030] The guided filtering module is specifically used for:
[0031] Perform linear projection on the guide feature to obtain the first projection feature, and perform linear projection on the feature to be sampled to obtain the second projection feature;
[0032] Upsampling the second projection feature to obtain a first sampling feature, constructing a guided filtering model based on the first projection feature and the first sampling feature, and parsing the guided filtering model to obtain a guided filtering feature;
[0033] The semantic parsing module is specifically used to:
[0034] A mutual similarity test is performed on the guided filter feature and the first sampling feature to obtain the adjacent semantic similarity corresponding to the feature to be sampled.
[0035] When upsampling the second projection feature to obtain the first sampling feature, the semantic parsing module is used to:
[0036] Obtaining a first resolution of the first projection feature, and obtaining a second resolution of the second projection feature;
[0037] Determine a ratio of the first projection feature to the second resolution as a first upsampling factor;
[0038] The second projection feature is upsampled by bilinear interpolation based on the first upsampling factor to obtain a first sampling feature.
[0039] When a guided filtering model is constructed based on the first projection feature and the first sampling feature, and the guided filtering model is parsed to obtain the guided filtering feature, the guided filtering module is used to:
[0040] Constructing a semantic adjacency window, constructing an adjacency difference parameter in the semantic adjacency window, and constructing a guided filtering model according to the adjacency difference parameter, the first projection feature, and the first sampling feature;
[0041] Analyze the guided filtering model to obtain a first value of the adjacent difference parameter;
[0042] Performing averaging processing on the first value of the adjacent difference parameter based on the semantic adjacent window to obtain a filtering parameter;
[0043] The filtering parameters are fused with the first projection features to obtain the guided filtering features.
[0044] When the guided filtering feature and the first sampling feature are subjected to feature fusion processing to obtain the adjacent semantic similarity corresponding to the feature to be sampled, the semantic parsing module is used to:
[0045] Obtain the i-th adjacent sampling feature of the first adjacent pixel of the i-th pixel from the first sampling feature, and combine the i-th adjacent sampling features corresponding to the i-th pixel into an i-th adjacent sampling feature matrix; i is a positive integer; the interval between any two adjacent pixels of the i-th pixel and the first adjacent pixel of the i-th pixel is the second upsampling multiple;
[0046] The product of the i-th pixel filter feature of the i-th pixel in the guided filter feature and the transpose of the i-th adjacent sampling feature matrix is determined as the adjacent semantic similarity corresponding to the i-th pixel.
[0047] The detail parsing module is specifically used to:
[0048] Construct a denoising convolution kernel, convolve the denoising convolution kernel with the guide feature to obtain the denoising feature;
[0049] From the denoising features, obtain the adjacent denoising features of the first adjacent pixel points corresponding to M pixels respectively, and form the adjacent denoising features corresponding to each pixel point into an adjacent denoising feature matrix of the pixel point; M is a positive integer;
[0050] The pixel guidance feature of each pixel in the guidance feature is fused with the adjacent denoising feature matrix of the pixel to obtain the adjacent detail similarity corresponding to M pixels.
[0051] When upsampling the to-be-sampled features using adjacency similarity to obtain target features, the feature sampling module is used to:
[0052] Activate the adjacent similarity corresponding to the i-th pixel to obtain the i-th activation similarity; i is a positive integer;
[0053] Using the i-th activation similarity, perform weighted summation on the initial pixel features of the second adjacent pixel of the i-th pixel in the feature to be sampled to obtain the i-th sampled pixel feature of the i-th pixel;
[0054] When the post-sampling pixel features of all pixels are obtained, the post-sampling pixel features of all pixels are combined into target features.
[0055] The image processing module is specifically used for:
[0056] Input the target features into the task prediction network corresponding to the image processing task, predict the target features through the task prediction network, and obtain the image detection results of the image processing task;
[0057] The device also includes:
[0058] The result output module is used to output the image detection results and perform the business corresponding to the image processing task based on the image detection results.
[0059] In one aspect, an embodiment of the present application provides an image sampling and processing device, the device comprising:
[0060] A sample acquisition module is used to obtain image samples and image processing annotations of image samples in image processing tasks;
[0061] A sample processing module is used to input the image sample into the initial image processing model, obtain the sample guiding features of the image sample through the initial guiding network in the initial image processing model, and obtain the sample to-be-sampled features of the image sample through the initial task main network in the initial image processing model; the resolution of the sample guiding features is greater than the resolution of the sample to-be-sampled features;
[0062] The sample guidance module is used to perform guided filtering processing on the sample guidance features and the sample to be sampled features to obtain the sample guidance filtering features;
[0063] The first parsing module is used to perform mutual similarity detection on the sample guided filter feature and the sample to be sampled feature, and determine the sample adjacent semantic similarity corresponding to the sample to be sampled feature; the sample adjacent semantic similarity is used to represent the semantic similarity between each pixel in the sample to be sampled feature and the adjacent pixel of the pixel;
[0064] The second parsing module is used to perform denoising on the sample guide feature to obtain the sample denoising feature, perform detail-aware self-alignment on the sample denoising feature and the sample guide feature to obtain the sample adjacent detail similarity corresponding to the sample to-be-sampled feature;
[0065] The sample sampling module is used to combine the sample adjacency semantic similarity and the sample adjacency detail similarity to form the sample adjacency similarity, and use the sample adjacency similarity to perform upsampling processing on the sample to be sampled to obtain the target sample feature;
[0066] The sample prediction module is used to predict the target sample features through the initial task prediction network in the initial image processing model to obtain the sample detection results of the image processing task;
[0067] The model training module is used to adjust the parameters of the initial image processing model through image processing annotation and sample detection results to obtain the image processing model.
[0068] The sample acquisition module is specifically used to:
[0069] Parse the image processing task, obtain the task processing parameters, and obtain the image labels of the candidate images;
[0070] Determine the candidate images whose image labels meet the task processing parameters as the initial samples, perform quality detection on the initial samples, and obtain the sample quality corresponding to the initial samples;
[0071] An initial sample whose sample quality is greater than or equal to a standard quality threshold is determined as an image sample, the image sample and the image processing task are sent to a management device, and the image processing annotation sent by the management device is obtained.
[0072] The model training module is specifically used to:
[0073] Generate a difference loss function based on the difference data between the image processing annotation and the sample detection results;
[0074] Perform quality detection on the target sample features to obtain image quality parameters, and combine the image quality parameters with the difference loss function to form a loss function;
[0075] The loss function is used to adjust the parameters of the initial image processing model to obtain the image processing model.
[0076] On the one hand, an embodiment of the present application provides a computer device, including a processor, a memory, and an input and output interface;
[0077] The processor is connected to the memory and the input and output interface respectively, wherein the input and output interface is used to receive and output data, the memory is used to store the computer program, and the processor is used to call the computer program so that the computer device including the processor executes the image sampling and processing method in one aspect of the embodiment of the present application.
[0078] In one aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is suitable for being loaded and executed by a processor so that a computer device having the processor performs the image sampling and processing method in one aspect of an embodiment of the present application.
[0079] In one aspect, an embodiment of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional embodiments of the present application. In other words, when the computer instructions are executed by the processor, the methods provided in various optional embodiments of the present application are implemented.
[0080] Implementing the embodiments of this application will have the following beneficial effects:
[0081] In an embodiment of the present application, the guiding features and the features to be sampled of the initial image can be obtained; the resolution of the guiding features is greater than the resolution of the features to be sampled; the guiding features and the features to be sampled are subjected to guided filtering processing to obtain guided filtering features, and the adjacent semantic similarity corresponding to the features to be sampled is determined through the guided filtering features and the features to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel point in the features to be sampled and the adjacent pixels of the pixel point; the guiding features are denoised to obtain denoised features, and the denoised features and the guiding features are detail-aware self-aligned to obtain the adjacent detail similarity corresponding to the features to be sampled; the adjacent semantic similarity and the adjacent detail similarity are combined into adjacent similarity, the features to be sampled are upsampled using the adjacent similarity to obtain target features, and the image detection results of the image processing task are predicted based on the target features. Through the above process, a guided filter feature aligned with the feature to be sampled can be generated by guided filtering. Semantic perception mutual alignment can be performed based on the guided filter feature to determine the semantic similarity between the guided feature and the feature to be sampled. At the same time, the guided feature is self-aligned to obtain the similarity of detail fidelity perception. By combining these two similarities, controllable alignment of features at different resolutions in terms of semantic content can be achieved, and the image detail information in upsampling can be restored, thereby improving the accuracy of the sampling process and the sampling performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0083] Figure 1 This is a network interaction architecture diagram for image sampling and processing provided by an embodiment of the present application;
[0084] Figure 2 This is a schematic diagram of an image sampling and processing scenario provided by an embodiment of the present application;
[0085] Figure 3 This is a schematic diagram of an image segmentation scenario provided by an embodiment of the present application;
[0086] Figure 4 This is a flow chart of a method for image sampling and processing provided by an embodiment of the present application;
[0087] Figure 5 This is a schematic diagram of a feature alignment scenario provided by an embodiment of the present application;
[0088] Figure 6 This is a flow chart of another method for image sampling and processing provided by an embodiment of the present application;
[0089] Figure 7 This is a schematic diagram of an image processing scenario provided by an embodiment of the present application;
[0090] Figure 8 This is a flow chart of an image sampling training method provided in an embodiment of the present application;
[0091] Figure 9 This is a schematic diagram of an interaction scenario provided by an embodiment of the present application;
[0092] Figure 10 This is a schematic diagram of upsampling optimization provided by an embodiment of the present application;
[0093] Figure 11 Schematic diagram of an image sampling and processing device provided in an embodiment of the present application;
[0094] Figure 12 is a schematic diagram of another image sampling and processing device provided in an embodiment of the present application;
[0095] Figure 13 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0096] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0097] Among them, if it is necessary to collect object (such as user, etc.) data in this application, a prompt interface or pop-up window will be displayed before or during the collection. The prompt interface or pop-up window is used to prompt the user that certain data is currently being collected. Only after the user issues a confirmation operation on the prompt interface or pop-up window, the relevant steps for data acquisition will be executed, otherwise the process will end. Moreover, the acquired user data (such as initial images and image samples, etc.) will be used in reasonable and legal scenarios or purposes. Optionally, in some scenarios where user data needs to be used but the user has not authorized it, authorization can be requested from the user, and the user data can be used when the authorization is passed. In other words, this application will use the acquired user data reasonably and legally, that is, the use of user data complies with the relevant provisions of laws and regulations.
[0098] Among them, some of the data involved in this application are explained as follows:
[0099] Feature map: A feature map obtained by convolving an image with a filter. A feature map can be convolved with a filter to generate a new feature map, which can be referred to as feature or feature for short.
[0100] Feature upsampling: Feature upsampling is the process of increasing the resolution of low-resolution features through an interpolation-like operation.
[0101] Upsampling kernel: Upsampling kernel, usually an upsampled high-resolution pixel is obtained by weighted summation of the low-resolution pixels around it, and this weighting coefficient is the upsampling kernel.
[0102] Optionally, this application may be technically implemented using artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence, such as the image processing model trained in this application. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0103] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0104] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0105] In the examples of this application, see Figure 1 , Figure 1 This is a network interaction architecture diagram of an image sampling process provided by an embodiment of the present application, such as Figure 1As shown, the computer device 101 can obtain an initial image, perform feature extraction on the initial image, and obtain guiding features and features to be sampled, wherein the guiding features can be considered as high-resolution (HR) features, and the features to be sampled can be considered as low-resolution (LR) features. The computer device 101 can perform guided filtering (Guided Filter) on the guiding features and the features to be sampled to obtain guided filtering features, and the guided filtering features are aligned with the guiding features, so that semantic-aware similarity detection can be performed on this basis; at the same time, the guiding features are self-aligned to achieve detail-fidelity-aware similarity detection. By combining the above two aspects and obtaining the target features, controllable feature alignment can be achieved, thereby achieving controllable alignment of features at different resolutions in terms of semantic content, and recovery of image detail information in upsampling, thereby improving the accuracy of the sampling process and the sampling performance.
[0106] Furthermore, the computer device 101 can predict the image detection result of the image processing task based on the target features. Optionally, the computer device 101 can obtain the initial image from the local memory, and when the image detection result is obtained, the image detection result can be output; or the computer device 101 can also respond to the image processing request of any business device, obtain the initial image sent by the business device, and when the image detection result is obtained, send the image detection result to the business device, wherein the business device can be any business device, such as Figure 1 The service device 102a, the service device 102b or the service device 102c shown in FIG.
[0107] For details, see Figure 2 , Figure 2 This is a schematic diagram of an image sampling and processing scenario provided by an embodiment of the present application. Figure 2As shown, a computer device can obtain an initial image 201, obtain a guiding feature and a feature to be sampled of the initial image 201, wherein the guiding feature is an HR feature, the feature to be sampled is an LR feature, and the resolution of the guiding feature is greater than the resolution of the feature to be sampled. Further, the computer device can perform guided filtering on the guiding feature and the feature to be sampled to obtain a guided filtering feature, and determine the adjacent semantic similarity corresponding to the feature to be sampled through the guided filtering feature and the feature to be sampled. Through the guided filtering process, an HR feature aligned with the LR feature (i.e., the feature to be sampled), i.e., the guided filtering feature, can be generated. Based on the feature alignment, the adjacent semantic similarity between the feature to be sampled and the guided filtering feature is determined, thereby achieving controllable alignment of features at different resolutions in terms of semantic content. Further, the guiding feature is denoised to obtain a denoised feature, and the denoised feature and the guiding feature are detail-aware self-aligned to obtain the adjacent detail similarity corresponding to the feature to be sampled, so that the adjacent detail similarity can achieve detail fidelity perception and restore image detail information during upsampling. By combining adjacent semantic similarity with adjacent detail similarity to form adjacent similarity, the to-be-sampled features are upsampled using adjacent similarity to obtain target features. This improves the accuracy and performance of the sampling process by integrating controllable alignment of semantic content and restoration of image detail information during upsampling. Furthermore, the image detection results of image processing tasks can be predicted based on target features. The improvements in sampling accuracy and performance can improve the accuracy of image processing tasks. The image processing task can be any task requiring upsampling, such as image segmentation tasks, panoptic segmentation tasks, and target detection tasks, without limitation.
[0108] For example, see Figure 3 , Figure 3 This is a schematic diagram of an image segmentation scenario provided by an embodiment of the present application. Figure 3 As shown, the computer device can perform shallow feature extraction on the initial image 301 to obtain the guiding feature 302; and perform feature extraction on the initial image 301 through the task main network corresponding to the image processing task to obtain the feature to be sampled 303. Through the solution of the present application, the guiding feature 302 and the feature to be sampled 303 are semantically aligned with each other, and detail-aware self-aligned, to achieve controllable alignment of the guiding feature 302 and the feature to be sampled 303, and upsample the feature to be sampled 303 to obtain the target feature 304. The image detection result of the image processing task is predicted based on the target feature 304. Here, taking the image processing task as an image segmentation task as an example, the target feature 304 can be predicted to obtain the image detection result 305 of the image processing task. The image detection result 305 is used to represent the image segmentation result after the initial image 301 is segmented.
[0109] It is understandable that the computer devices mentioned in the embodiments of the present application include but are not limited to terminal devices or servers. In other words, the computer device can be a server or a terminal device, or it can be a system composed of a server and a terminal device. Among them, the terminal device mentioned above can be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, PDAs, vehicle-mounted equipment, augmented reality / virtual reality (AR / VR) equipment, helmet displays, smart TVs, wearable devices, smart speakers, digital cameras, cameras and other mobile Internet devices (mobile internet device, MID) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. As Figure 1 As shown in , the terminal device can be a laptop (as shown by the service device 102b), a mobile phone (as shown by the service device 102c) or a vehicle-mounted device (as shown by the service device 102a), etc. Figure 1 Only some of the devices are listed as examples. Optionally, the business device 102a refers to a device located in the vehicle 103. The business device 102a can be used to output the initial image and image detection results, and has an image output function, such as outputting image 1021. The server mentioned above can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, content delivery network (CDN), and big data and artificial intelligence platforms.
[0110] Optionally, the data involved in the embodiments of the present application can be stored in a computer device, or the data can be stored based on cloud storage technology or a blockchain network, which is not limited here.
[0111] Further, see Figure 4 , Figure 4 This is a flow chart of a method for image sampling and processing provided by an embodiment of the present application. Figure 4 As shown, the image sampling process includes the following steps:
[0112] Step S401: Obtain guiding features and features to be sampled of the initial image.
[0113] In an embodiment of the present application, the computer device can perform shallow feature extraction on the initial image to obtain guiding features; and perform deep feature extraction on the initial image to obtain features to be sampled of the initial image. The resolution of the guiding features is greater than the resolution of the features to be sampled. The shallow feature extraction focuses mainly on the local and detailed information of the image, and usually uses a smaller convolution kernel to extract detailed information such as color, texture, edge and corner points in the image to obtain local details of the image; the deep feature extraction focuses mainly on the global and abstract features of the image, and obtains more global information of the image by increasing the depth of the network. It usually uses a larger convolution kernel and pooling operation to gradually abstract and compress image information, thereby obtaining higher-level feature representations. These features have stronger semantics and can better describe the overall content and category of the image.
[0114] Specifically, the computer device can input the initial image into the image processing model, obtain the guiding features of the initial image through the guiding network in the image processing model; and obtain the image to be sampled of the initial image through the task main network in the image processing model. Optionally, the computer device can perform convolution processing on the initial image to obtain a first convolution feature, perform normalization processing on the first convolution feature to obtain a normalized feature; perform convolution processing on the normalized feature to obtain the guiding features of the initial image; and obtain the features to be sampled of the initial image in the task main network corresponding to the image processing task, that is, the guiding network includes a convolution layer, a normalization layer, and a convolution layer, etc. The number of convolution layers included in the guiding network and the type of normalization layer are not limited here. The normalization layer can be a group normalization layer (GroupNorm) or a batch normalization layer (Batch Normalization, BN), etc. The size of the convolution kernel used in the convolution layer in the guiding network is less than or equal to the first size threshold, which is used to perform shallow feature extraction on the image. For example, the guided network may also include "convolutional layer (Conv) - group normalization layer (GroupNorm) - activation layer (ReLU) - convolutional layer (Conv)", that is, the guided network may be any network that can perform shallow feature extraction, and the specific composition is not limited here, such as the "convolutional layer - normalization layer - convolutional layer" listed above, or "convolutional layer - normalization layer - activation layer - convolutional layer", etc., and there is no restriction here, and the number of convolutional layers used in any composition, as well as the type of normalization layer, are not restricted. For example, in one possible network composition, assuming that the guided network includes "Conv3×3-GroupNorm-ReLU-Conv3×3", the computer device can perform convolution processing on the initial image through a 3×3 convolution layer to obtain a first convolution feature, which is a feature with 64 channels; the first convolution feature is normalized through a group normalization layer and an activation layer to obtain a normalized feature, and then the normalized feature is convolved through a 3×3 convolution layer to obtain a guided feature of the initial image, which can be a feature with a channel size of 32.
[0115] Optionally, the computer device can perform deep feature extraction on the initial image to obtain the features to be sampled of the initial features. Specifically, the features to be sampled of the initial image can be obtained through the task main network in the image processing model. The guide feature can be denoted as y, and the feature to be sampled can be denoted as x, where x∈R hw×C , y∈R HW×c, where H, W, c, h, w, and C are all positive integers. HW represents the size of the guiding feature and can be considered as its resolution. hw represents the size of the feature to be sampled and can be considered as its resolution. c represents the number of channels of the guiding feature and C represents the number of channels of the feature to be sampled. For each pixel i∈{1,…,HW} in the guiding feature, for each pixel i∈{1,…,hw} in the feature to be sampled. The guiding feature usually contains more image detail information, while the feature to be sampled usually contains more semantic information.
[0116] Step S402: performing guided filtering processing on the guided features and the features to be sampled to obtain guided filtering features.
[0117] In an embodiment of the present application, the computer device may perform a linear projection on the guide feature to obtain a first projection feature, and perform a linear projection on the feature to be sampled to obtain a second projection feature. The linear projection may be a normalization plus convolution process, such as a group normalization plus convolution process, such as Group Normalization + 1x1 convolution. For example, the guide feature is group normalized and then convolved to obtain the first projection feature, which can be denoted as q∈R HW×D ; Perform group normalization on the sampled features and then perform convolution to obtain the second projection feature, which can be recorded as k∈R hw×D . Wherein, D is used to represent the number of channels of the first projection feature and the second projection feature, and D is a positive integer. Alternatively, the linear projection can be a convolution, normalization plus convolution process, etc. For example, the guiding feature can be convolved to obtain the second convolution feature, the second convolution feature can be normalized to obtain the first activation feature, the first activation feature can be convolved to obtain the first projection feature; the unsampled feature can be convolved to obtain the third convolution feature, the third convolution feature can be normalized to obtain the second activation feature, the second activation feature can be convolved to obtain the second projection feature, etc. In other words, the number of convolution processes included in the linear projection is not limited, that is, the linear projection includes at least one or more convolutions on the corresponding data (such as the guiding feature or the unsampled feature), and multiple times means at least twice. Based on the characteristics of the guiding feature and the unsampled feature, the first projection feature usually includes more image detail information, and the second projection feature usually includes more semantic information. The space where k is located can be called the semantic space, and the space where q is located can be called the detail space.
[0118] Furthermore, the second projection feature may be upsampled to obtain a first sampling feature, a guided filtering model may be constructed based on the first projection feature and the first sampling feature, and the guided filtering model may be analyzed to obtain a guided filtering feature.
[0119] Optionally, when upsampling the second projection feature to obtain the first sampling feature, the first resolution of the first projection feature and the second resolution of the second projection feature can be obtained; the ratio of the first projection feature to the second resolution can be determined as the first upsampling multiple. Further, the second projection feature can be upsampled based on the first upsampling multiple to obtain the first sampling feature. The upsampling can be nearest neighbor interpolation upsampling, bilinear interpolation upsampling or bicubic interpolation upsampling, etc., which are not limited here. For example, the second projection feature can be upsampled by bilinear interpolation based on the first upsampling multiple to obtain the first sampling feature. The first sampling feature can be recorded as By upsampling the second projection features, the semantic space and the detail space can be aligned in a controllable manner. Based on the alignment, subsequent processing can fully consider the content differences between the guiding features and the features to be sampled, thereby improving the accuracy of upsampling.
[0120] Furthermore, a guided filtering model can be constructed based on the first projection feature and the first sampling feature, and the guided filtering model can be parsed to obtain the guided filtering feature. Specifically, the computer device can construct a semantic adjacent window, and the semantic adjacent window can be denoted as I j , I j It is a square window centered on pixel j and having a window width of r, where j is a positive integer and r is a positive integer. r can be considered a hyperparameter and can be obtained from experience or manually set. Alternatively, a first upsampling multiple can be obtained, and the window width of the adjacent window is determined based on the first upsampling multiple. That is, when the window width is determined based on the first upsampling multiple, the window width can be between the first integer multiple and the second integer multiple of the first upsampling multiple, and there is no restriction here. For example, if the first upsampling multiple is 4, if the first integer multiple is 1 and the second integer multiple is 3, then the window width r∈[4,12], which can be 6 or 8, etc. Further, an adjacent difference parameter can be constructed in the semantic adjacent window, and a guided filtering model can be constructed based on the adjacent difference parameter, the first projection feature, and the first sampling feature. The adjacent difference parameter includes a coefficient adjacent difference parameter and an offset adjacent difference parameter. The guided filtering model is used to make the first sampling feature and the guided filtering feature as close as possible.
[0121] The guided filtering model can be shown as formula (1) and formula (2):
[0122]
[0123] As shown in formula (1), m jd and n jd is the affine transformation coefficient of the jth pixel and the dth channel, that is, m jd is the coefficient adjacent difference parameter, n jd is the offset adjacent difference parameter, j is a positive integer, and d is a positive integer. Wherein, the subscript j is used to represent the jth pixel, the subscript d is used to represent the dth channel, and the subscript i is used to represent the i-th pixel, as shown in Used to represent the features of the first sampling feature at the i-th pixel and the d-th channel, m jd Used to represent the value of the coefficient adjacent difference parameter at the jth pixel and the dth channel, n jd It is used to represent the value of the offset adjacent difference parameter at the jth pixel, the dth channel, etc. ∈ is a regularization coefficient with a small value. It is used to limit the upper limit of the value of the coefficient adjacent difference parameter. It is a hyperparameter that can be an empirical value used to generate the regularization term or a manual setting, such as ∈ = 0.001. Formula (2) is as follows:
[0124]
[0125] in," " is a universal quantification symbol, which is a mathematical symbol used to represent universal quantifiers. The meaning shown in formula (2) is that for the feature of the first projection feature q at any pixel in the semantic adjacency window centered on the jth pixel, after the adjacency difference parameter corresponding to the jth pixel is calculated as shown in formula (2), the guided filter feature obtained after the guided filter model is parsed at the pixel has the same value. Among them, q id It is used to represent the feature of the first projection feature at the i-th pixel and the d-th channel. As shown in formula (1) and formula (2), it is used to represent the guided filtering model.
[0126] Specifically, the computer device can project the first projection feature into the semantic space where the second projection feature is located to obtain a first guided sub-model, as shown in formula (2); use the first sampling feature to perform guided semantic constraints to obtain a first constraint condition, perform value constraints on the adjacent difference parameter to obtain a second constraint condition, and combine the first constraint condition and the second constraint condition to form a second guided sub-model, as shown in formula (1); and combine the first guided sub-model and the second guided sub-model to form a guided filtering model. The guided semantic constraint is used to represent the similarity between the guided filtering feature obtained by the constrained guided filtering model and the first sampling feature in semantic information, that is, the constrained guided filtering feature is as similar as possible to the first sampling feature in semantic information.
[0127] Furthermore, the guided filter model is analyzed to obtain the first value of the adjacent difference parameter, that is, the value of the adjacent difference parameter that minimizes the result of formula (1) based on formula (2). Specifically, the first guided sub-model and the second guided sub-model are integrated to obtain the function to be analyzed, which can be written as The analytical function to be processed is solved to obtain the first value of the adjacent difference parameter, including the first value of the coefficient adjacent difference parameter and the first value of the offset adjacent difference parameter. The computer device can solve the analytical function to obtain a parameter analytical expression, as shown in formula (3-1) and formula (3-2), and determine the result of the parameter analytical expression as the first value of the adjacent difference parameter. The parameter analytical expression can include the coefficient parameter analytical expression and the offset parameter analytical expression. The analytical process of the guided filtering model can be referred to formula (3-1) and formula (3-2). Formula (3-1) is as follows:
[0128]
[0129] As shown in formula (3-1), the analytical result used to represent the coefficient adjacent difference parameter is the coefficient parameter analytical formula of the coefficient adjacent difference parameter. Solving it to obtain the first value of the coefficient adjacent difference parameter. Formula (3-2) is as follows:
[0130]
[0131] As shown in formula (3-2), the analytical result for expressing the offset adjacent difference parameter is the offset parameter analytical formula of the offset adjacent difference parameter, which is solved to obtain the first value of the offset adjacent difference parameter. Used to represent q id In the semantic adjacency window I j The mean of , that is, the mean of the features of all pixels of the first projection feature in the semantic adjacency window centered on the j-th pixel; Used to represent q id In the semantic adjacency window I j , that is, the variance of the features of all pixels in the semantic adjacency window centered on the jth pixel. yes in I jThe mean of the first sampling feature, that is, the mean of the features of all pixels in the semantic adjacency window centered on the j-th pixel. Among them, |I| is used to represent the number of pixels included in each semantic adjacency window, which can be artificially r*r. By performing the analytical process shown in formulas (3-1) and (3-2) on the guided filtering model, the first value of the adjacency difference parameter is obtained, including the first value of the coefficient adjacency difference parameter and the first value of the offset adjacency difference parameter.
[0132] Furthermore, the computer device can perform mean processing on the first value of the adjacency difference parameter based on the semantic adjacency window to obtain a filtering parameter; and fuse the filtering parameter with the first projection feature to obtain a guided filtering feature. The adjacency difference parameters of each pixel point obtained by the analytical process shown in the above formula (3-1) and formula (3-2) can exist in different semantic adjacency windows at the same time. By performing mean processing on the adjacency difference parameters included in any semantic adjacency window, the filtering parameters of the central pixel point corresponding to the semantic adjacency window are obtained. The filtering parameters include coefficient filtering parameters and offset filtering parameters, wherein the coefficient filtering parameters of the i-th pixel point can be recorded as The offset filter parameter of the i-th pixel can be expressed as The central pixel point corresponding to the semantic adjacent window refers to the pixel point located at the center position of the semantic adjacent window, such as the semantic adjacent window I j The central pixel of is the jth pixel, etc. The filtering parameters are fused with the first projection feature to obtain the guided filtering feature. Specifically, the filtering parameters can be used to project the first projection feature to the semantic space where the second projection feature is located to obtain the guided filtering feature. This process can be shown in formula (4):
[0133]
[0134] Optionally, the process can also be recorded as follows: the computer device can determine the average of the feature similarities between the first projection feature and the first sampling feature, at each pixel in the j-th semantic adjacency window, as the j-th pixel similarity corresponding to the j-th pixel, as shown in formula (3-1): Get the first mean value of the features of each pixel point included in the jth semantic adjacent window of the first projection feature, as shown in formula (3-1) Get the second mean of the features of each pixel point included in the jth semantic adjacent window of the first sampling feature, as shown in formula (3-1) The product of the jth first mean and the jth second mean is determined as the jth mean similarity; the difference between the jth pixel similarity and the jth mean similarity is determined as the jth deviation similarity. Obtain the variance of the features of each pixel included in the jth semantic adjacency window of the first projection feature, as shown in formula (3-1) The jth variance is adjusted using the regularization coefficient to obtain the adjusted variance. The ratio of the jth deviation similarity to the adjusted variance is determined as the first value of the coefficient adjacency difference parameter corresponding to the jth pixel. Based on the jth second mean, the first value of the coefficient adjacency difference parameter corresponding to the jth pixel, and the jth first mean, the first value of the offset adjacency difference parameter corresponding to the jth pixel is determined. This process is described in formula (3-2). Similarly, the adjacency difference parameters for all pixels included in the first projection feature and the first sampling feature can be obtained.
[0135] Furthermore, the first value of the adjacency difference parameter can be averaged based on the semantic adjacency window to obtain the filter parameter; the filter parameter can be fused with the first projection feature to obtain the guided filter feature. Specifically, the first value of the adjacency difference parameter of the pixel point included in the i-th semantic adjacency window can be averaged to obtain the filter parameter of the i-th pixel point. For example, the first value of the coefficient adjacency difference parameter of the pixel point included in the i-th semantic adjacency window can be averaged to obtain the coefficient filter parameter of the i-th pixel point, which is recorded as The first value of the offset adjacency difference parameter of the pixel point included in the i-th semantic adjacency window is averaged to obtain the offset filtering parameter of the i-th pixel point, which is recorded as The coefficient filter parameters and offset filter parameters of the i-th pixel point are used to update the features of the first projection feature at the i-th pixel point to obtain the pixel filter feature corresponding to the i-th pixel point, as shown in formula (4): GF ) id Similarly, the pixel filter features corresponding to all the pixel points included in the first projection feature can be obtained, and the pixel filter features corresponding to all the pixel points included in the first projection feature are combined into a guided filter feature.
[0136] Step S403 : performing mutual similarity detection on the guided filtering feature and the feature to be sampled, and determining the adjacent semantic similarity corresponding to the feature to be sampled.
[0137] In an embodiment of the present application, the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and the adjacent pixel of the pixel. The computer device can perform feature fusion processing on the guided filter feature and the first sampling feature to obtain the adjacent semantic similarity corresponding to the feature to be sampled. Specifically, from the first sampling feature, the i-th adjacent sampling feature of the first adjacent pixel of the i-th pixel is obtained, and the i-th adjacent sampling feature corresponding to the i-th pixel is composed into the i-th adjacent sampling feature matrix; i is a positive integer. Among them, the i-th adjacent sampling feature is used to represent the feature of the first adjacent pixel of the i-th pixel in the first sampling feature. The interval between any two adjacent pixels in the i-th pixel and the first adjacent pixel of the i-th pixel is the second upsampling multiple. The second upsampling multiple refers to the multiple used when upsampling the feature to be sampled. For example, if the feature to be sampled needs to be upsampled 4 times, the second upsampling multiple is 4. Furthermore, the similarity test can be performed on the guided filter feature and the adjacent sampling feature matrix to obtain the adjacent semantic similarity corresponding to the feature to be sampled. Specifically, taking the i-th pixel as an example, the product of the i-th pixel filter feature of the i-th pixel in the guided filter feature and the transpose of the i-th adjacent sampling feature matrix can be determined as the adjacent semantic similarity corresponding to the i-th pixel. The process of obtaining the adjacent semantic similarity corresponding to the i-th pixel can be shown in formula (5):
[0138]
[0139] As shown in formula (5), the (s s ) i Used to represent the adjacent semantic similarity corresponding to the i-th pixel point, N(i) is used to represent the first adjacent pixel corresponding to the i-th pixel. The first adjacent pixel corresponding to the i-th pixel refers to the K*K pixels centered at the i-th pixel. K is a positive integer used to represent the number of unit adjacent points (i.e., the number of adjacent pixels required for each row and column), and the interval between each two adjacent pixels is the second upsampling multiple. At this time, the dimension of the adjacent semantic similarity corresponding to the i-th pixel can be considered to be K 2 , recorded as K can be a hyperparameter, such as 3, or manually set, and can also be updated as needed. Similarly, the adjacency semantic similarity of all pixels included in the guiding feature can be obtained. The adjacency semantic similarity of all pixels included in the guiding feature constitutes the adjacency semantic similarity used to upsampling the unsampled feature.
[0140] The guided filtering feature q obtained by guided filtering GFCan be used with Controllable alignment, see Figure 5 , Figure 5 This is a schematic diagram of a feature alignment scenario provided by an embodiment of the present application, such as Figure 5 As shown, for the guided filter feature 501, the dimension of the guided filter feature 501 is the same as the dimension of the guided feature. Therefore, there will be a difference in the second upsampling multiple in the correspondence with the feature to be sampled 502 in the pixel point, as shown in FIG. Figure 5 As shown in , assuming that the second upsampling multiple is 2, the pixel point 1a in the guided filtering feature 501 is equivalent to the pixel point 1b in the feature to be sampled 502, the pixel point 2a in the guided filtering feature 501 is equivalent to the pixel point 2b in the feature to be sampled 502, the pixel point 3a in the guided filtering feature 501 is equivalent to the pixel point 3b in the feature to be sampled 502, the pixel point 4a in the guided filtering feature 501 is equivalent to the pixel point 4b in the feature to be sampled 502, and the pixel point 5a in the guided filtering feature 501 is equivalent to the pixel point 5b in the feature to be sampled 502. b, pixel 6a in the guiding filter feature 501 is equivalent to pixel 6b in the feature to be sampled 502, etc. Therefore, the interval between adjacent pixels in the first adjacent pixel of any pixel is taken as the second upsampling multiple, such as the first adjacent pixel of pixel 5a includes pixel 1a, pixel 3a, pixel 4a, pixel 2a and pixel 6a, etc., so that the obtained adjacent semantic similarity of each pixel can correspond to the feature to be sampled, thereby reducing the content difference between the guiding feature and the feature to be sampled and improving the accuracy of the similarity determination. Where q GF In terms of semantic information It is very similar, which is achieved by formula (1) and retains the detail information in q. It is achieved by formula (2) to project the first projection feature into the semantic space where the second projection feature is located, and to achieve feature alignment in the semantic space. This allows the adjacent semantic similarity to more accurately supplement the detail information for the feature to be sampled, thereby improving the accuracy of upsampling.
[0141] Step S404: Denoising the guiding feature to obtain a denoised feature, performing detail-aware self-alignment on the denoised feature and the guiding feature to obtain the adjacent detail similarity corresponding to the feature to be sampled.
[0142] In the embodiment of the present application, the computer device can construct a denoising convolution kernel, perform convolution processing on the denoising convolution kernel and the guide feature, and obtain the denoising feature, which can be recorded as Among them, q blur It is used to represent the denoising feature, and f is used to represent the denoising convolution kernel. For example, when the dimension of the denoising convolution kernel is 3×3, the denoising convolution kernel can be recorded as f∈R 3×3 , Used to represent convolution processing. For example, the denoising convolution kernel can be a Gaussian convolution kernel, etc., and the Gaussian convolution kernel is used to perform Gaussian blur denoising. At this time, the computer device can perform convolution processing on the Gaussian convolution kernel and the guiding feature to obtain a denoised feature. Among them, when the denoising convolution kernel is a Gaussian convolution kernel, the variance of the Gaussian convolution kernel can be 1. By denoising the guiding feature, some local tiny noise in the guiding feature can be removed to improve the quality of the feature, thereby improving the accuracy of the subsequent determination of the adjacent similarity to a certain extent, and then improving the accuracy of upsampling. Further, the computer device can obtain the adjacent denoising feature matrix from the denoising feature, perform similarity detection on the adjacent denoising feature matrix and the guiding feature, and obtain the adjacent detail similarity corresponding to the feature to be sampled. Specifically, the adjacent denoising features of the first adjacent pixels corresponding to M pixels can be obtained from the denoising features, and the adjacent denoising features corresponding to each pixel can be combined into the adjacent denoising feature matrix of the pixel; M is a positive integer; the pixel guide feature of each pixel in the guide feature is fused with the adjacent denoising feature matrix of the pixel to obtain the adjacent detail similarity corresponding to the M pixels. The process of obtaining the adjacent detail similarity can be shown in formula (6):
[0143]
[0144] As shown in formula (6), the pixel guiding feature of the guiding feature at the i-th pixel point can be obtained, which is recorded as the i-th pixel guiding feature, and the adjacent denoising feature of the denoising feature at the first adjacent pixel point of the i-th pixel point can be obtained, which is recorded as the i-th adjacent denoising feature; the i-th adjacent denoising feature is composed of the adjacent denoising feature matrix of the i-th pixel point; the product of the transpose of the i-th pixel guiding feature and the adjacent denoising feature matrix of the i-th pixel point is determined as the adjacent detail similarity corresponding to the i-th pixel point. At this time, the dimension of the adjacent detail similarity corresponding to the i-th pixel point can be considered to be K 2 , recorded as Similarly, the adjacent detail similarity of all pixels included in the guiding feature can be obtained. The adjacent detail similarity of all pixels included in the guiding feature constitutes the adjacent detail similarity used for upsampling the feature to be sampled. Since two pixels that are close in detail space are usually similar in semantic space in a part of the image, this method can be used to detect the detail-perceived similarity of the feature to be sampled and obtain the detail-fidelity-perceived similarity. This process can also be seen in Figure 6 Steps S602 and S605 are shown.
[0145] Step S405 , combining the adjacent semantic similarity and the adjacent detail similarity into adjacent similarity, upsampling the to-be-sampled features using the adjacent similarity to obtain target features, and predicting the image detection result of the image processing task based on the target features.
[0146] In the embodiment of the present application, the computer device may combine the adjacency semantic similarity and the adjacency detail similarity into an adjacency similarity. Specifically, the sum of the adjacency semantic similarity and the adjacency detail similarity of the i-th pixel may be determined as the adjacency similarity corresponding to the i-th pixel. This process can be shown in formula (7):
[0147] s i =(s s ) i +(s d ) i (7)
[0148] As shown in formula (7), (s s ) i It is used to represent the semantic similarity of the adjacent pixels of the i-th pixel, (s d ) i Used to represent the similarity of adjacent details of the i-th pixel, s i Used to represent the adjacency similarity of the i-th pixel, That is, the dimension of the adjacent similarity of the i-th pixel is K 2 In this way, the adjacency similarity of each pixel in the post-sampling dimension can be obtained. The post-sampling dimension is used to represent the number of pixels corresponding to the feature that can be obtained after upsampling the feature to be sampled. For example, if the dimension of the feature to be sampled is hw×C and it needs to be upsampled to a feature of HW×c dimension, then the post-sampling dimension is HW, and in this case, i is a positive integer less than or equal to HW.
[0149] Furthermore, the computer device can use the adjacency similarity to upsample the feature to be sampled to obtain the target feature. The adjacency similarity can be activated to obtain the activation similarity, and the activation similarity can be used to upsample the feature to be sampled to obtain the target feature. Specifically, the adjacency similarity corresponding to the i-th pixel point can be activated to obtain the i-th activation similarity; i is a positive integer. The i-th activation similarity is used to perform weighted summation on the initial pixel features of the second adjacent pixel point of the i-th pixel point in the feature to be sampled to obtain the i-th post-sampling pixel feature of the i-th pixel point. The acquisition process of the i-th post-sampling pixel feature can be shown in formula (8):
[0150]
[0151] As shown in formula (8), It is used to represent the pixel feature after the i-th sampling, and Softmax() is used to represent the activation function, which is used to activate the adjacent similarity corresponding to the i-th pixel. N′(i) is used to represent the second adjacent pixel of the i-th pixel. Any two adjacent pixels in the second adjacent pixel of the i-th pixel are also adjacent in the feature to be sampled. The pixel interval between the first adjacent pixel and the second adjacent pixel realizes the controllable alignment between the guide feature and the feature to be sampled, such as Figure 5 As shown, the accuracy of upsampling and sampling performance are improved. When the post-sampling pixel features of all pixels are obtained, the post-sampling pixel features of all pixels are combined into target features.
[0152] Furthermore, the computer device can input the target features into the task prediction network corresponding to the image processing task, and predict the target features through the task prediction network to obtain the image detection result of the image processing task. Figure 3 As shown, the target feature 304 can be input into the task prediction network corresponding to the image processing task, and the target feature 304 can be predicted by the task prediction network to obtain the image detection result 305 of the image processing task. Figure 3 Taking the image segmentation task as an example, the task prediction network can be a segmentation head and can include one or more convolutional layers. For example, if the segmentation head is a 1x1 convolution, the 1x1 convolution can be used in the task prediction network to predict the target feature 304 to obtain the image segmentation result.
[0153] Optionally, the computer device can also output the image detection result, and execute the business corresponding to the image processing task based on the image detection result. For example, if the image processing task is an image segmentation task, and the business corresponding to the image processing task is an object display business, the computer device can divide the initial image into multiple segmented areas based on the image detection result, separate the object display images corresponding to the multiple segmented areas from the initial image, and output the object display images corresponding to the multiple segmented areas. For example, if the image processing task is a target detection task, and the business corresponding to the image processing task is an anomaly detection business, the computer device can obtain the object area to be detected corresponding to the target object based on the image detection result, perform anomaly detection on the object area to be detected, and obtain the object state of the object to be detected, which includes the normal state of the object and the abnormal state of the object, etc. The above are several possible application scenarios. The present application can be applied to any image processing task that requires upsampling processing, and the business corresponding to the subsequent image processing task can be any business that the image processing task can be applied to.
[0154] Further, see Figure 6 , Figure 6This is a flow chart of another method for image sampling and processing provided by an embodiment of the present application. Figure 6 As shown, the image sampling process includes the following steps:
[0155] Step S601: Obtain guiding features and features to be sampled of the initial image.
[0156] In the embodiments of this application, please refer to Figure 4 The relevant description of step S401 in will not be repeated here.
[0157] Step S602: linearly project the guide feature to obtain a first projection feature, and linearly project the feature to be sampled to obtain a second projection feature.
[0158] In the embodiments of this application, please refer to Figure 4 The relevant description of step S402 in will not be repeated here.
[0159] Step S603 : performing guided filtering processing on the first projection feature and the second projection feature to obtain a guided filtering feature.
[0160] In the embodiments of this application, please refer to Figure 4 The relevant description of step S402 in will not be repeated here.
[0161] Step S604 : performing mutual similarity detection on the guided filter feature and the second projection feature to determine the adjacent semantic similarity corresponding to the feature to be sampled.
[0162] In the embodiment of the present application, the first sampling feature corresponding to the guided filtering feature and the second projection feature is subjected to mutual similarity detection to determine the adjacent semantic similarity corresponding to the feature to be sampled. For details, see Figure 4 The relevant description of step S403 in will not be repeated here.
[0163] Step S605 , performing denoising processing on the first projection feature to obtain a denoised feature, performing detail-aware self-alignment on the denoised feature and the first projection feature to obtain the adjacent detail similarity corresponding to the feature to be sampled.
[0164] In the embodiments of this application, please refer to Figure 4 The relevant description of step S404 in the above step will not be repeated here. Specifically, the computer device can construct a denoising convolution kernel, perform convolution processing on the denoising convolution kernel and the first projection feature, and obtain a denoising feature, which can be recorded as Wherein, q is the first projection feature. Further, the adjacent denoising feature matrix can be obtained from the denoising feature, and the adjacent denoising feature matrix and the first projection feature are tested for similarity to obtain the adjacent detail similarity corresponding to the feature to be sampled. Specifically, from the denoising feature, the adjacent denoising features of the first adjacent pixel points corresponding to M pixels are obtained, and the adjacent denoising features corresponding to each pixel point are combined into the adjacent denoising feature matrix of the pixel point; M is a positive integer; the pixel projection feature of each pixel point in the first projection feature is fused with the adjacent denoising feature matrix of the pixel point to obtain the adjacent detail similarity corresponding to the M pixels. The process of obtaining the adjacent detail similarity can be shown in formula (9):
[0165]
[0166] Step S606 , combining the adjacent semantic similarity and the adjacent detail similarity into adjacent similarity, upsampling the to-be-sampled features using the adjacent similarity to obtain target features, and predicting the image detection result of the image processing task based on the target features.
[0167] In the embodiments of this application, please refer to Figure 4 The relevant description of step S405 in will not be repeated here.
[0168] For further information, see Figure 7 , Figure 7 This is a schematic diagram of an image processing scenario provided by an embodiment of the present application. Figure 7 As shown, the computer device can perform deep feature extraction on the initial image to obtain the feature to be sampled 7011 (LR DeepFeature), and perform shallow feature extraction on the initial image to obtain the guidance feature 7012 (HR Guidance Feature), see Figure 4 The sampling feature 7011 is linearly projected to obtain the second projection feature, and the second projection feature is up-sampled to obtain the first sampling feature 7021; the guide feature 7012 is linearly projected to obtain the first projection feature 7022. This process can be referred to Figure 4 The first projection feature 7022 is used to perform a guided filter process (Guided-Filter) on the first sampling feature 7021 to obtain a guided filter feature 7031, and a mutual similarity test (Mutual-Similarity) is performed on the guided filter feature 7031 and the first sampling feature 7021 to obtain the adjacent semantic similarity S corresponding to the feature to be sampled. s , the process can be seen in Figure 4The first projection feature 7022 is subjected to denoising, such as Gaussian blurring, to obtain a denoised feature 7032. The denoised feature 7032 is subjected to self-similarity detection (Self-Similarity) with the first projection feature 7022 to obtain the adjacent detail similarity S corresponding to the feature to be sampled. d , the process can be seen in Figure 4 The related description shown in step S403 is further described. s Similarity S with adjacent details d Composition adjacency similarity ( Figure 7 ⊕ process shown in ), the adjacent similarity is used to upsample the sampled feature 7041 to obtain the target feature 7042, as shown in Figure 7 shown process, which can be found in Figure 4 The relevant description shown in step S404.
[0169] Further optional, above Figures 4 to 6 The steps shown can be implemented by an image processing model. Specifically, in step S401, the computer device can input an initial image into the image processing model, obtain guiding features of the initial image through the guiding network in the image processing model, and obtain features to be sampled of the initial image through the task main network in the image processing model. In step S402, the computer device can use the image processing model to perform guided filtering on the guiding features and the features to be sampled to obtain guided filtered features. The guided filtered features and the features to be sampled are used to determine the adjacent semantic similarity corresponding to the features to be sampled. In step S403, the computer device can use the image processing model to perform denoising on the guiding features to obtain denoised features. The denoised features and the guiding features are then subjected to detail-aware self-alignment to obtain the sample adjacent detail similarity corresponding to the features to be sampled. In step S404, the computer device can use the image processing model to combine the adjacent semantic similarity and the adjacent detail similarity to form an adjacent similarity. The adjacent similarity is used to upsample the features to be sampled to obtain target features. The target features are predicted by the task prediction network in the image processing model to obtain the image detection result of the image processing task.
[0170] Furthermore, the training process of the image processing model can be found in Figure 8 , Figure 8 This is a flow chart of a sampling training method provided by an embodiment of the present application. Figure 8 As shown, the sampling training process includes the following steps:
[0171] Step S801: Obtain an image sample and an image processing annotation of the image sample in an image processing task.
[0172] In an embodiment of the present application, a computer device can parse an image processing task, obtain task processing parameters, and acquire image labels for candidate images. The candidate image whose image label meets the task processing parameters is determined as an initial sample, and the initial sample is quality-checked to obtain the sample quality corresponding to the initial sample; the initial sample whose sample quality is greater than or equal to the standard quality threshold is determined as an image sample, and the image sample and the image processing task are sent to a management device to obtain the image processing annotation sent by the management device. Alternatively, the computer device can directly obtain manually provided image samples, as well as image processing annotations of image samples in image processing tasks, etc., which are not limited here.
[0173] Step S802: input the image sample into the initial image processing model, obtain the sample guidance features of the image sample through the initial guidance network in the initial image processing model, and obtain the sample to-be-sampled features of the image sample through the initial task main network in the initial image processing model.
[0174] In the embodiment of the present application, the resolution of the sample guide feature is greater than the resolution of the sample to be sampled feature. Figure 4 The process of obtaining the guiding features and the features to be sampled in step S401 will not be described in detail here.
[0175] Step S803: performing guided filtering processing on the sample guided features and the sample to-be-sampled features to obtain sample guided filtered features.
[0176] In the embodiment of the present application, the generation process of the sample-guided filtering feature can be specifically referred to Figure 4 The process of obtaining the guided filtering features in step S402 will not be described in detail here.
[0177] Step S804 : performing mutual similarity detection on the sample guided filtering feature and the sample to-be-sampled feature to determine the sample adjacency semantic similarity corresponding to the sample to-be-sampled feature.
[0178] In the embodiment of the present application, the sample adjacency semantic similarity is used to represent the semantic similarity between each pixel in the sample to be sampled and its adjacent pixels. The determination process of the sample adjacency semantic similarity can be specifically referred to. Figure 4 The process of obtaining the adjacent semantic similarity in step S403 will not be described in detail here.
[0179] Step S805 , performing denoising processing on the sample guiding feature to obtain a sample denoised feature, performing detail-aware self-alignment on the sample denoised feature and the sample guiding feature to obtain the sample adjacent detail similarity corresponding to the sample to-be-sampled feature.
[0180] In the embodiments of this application, please refer to Figure 4 The process of obtaining the adjacent detail similarity in step S404 will not be described in detail here.
[0181] Step S806 , combining the sample adjacency semantic similarity and the sample adjacency detail similarity to form the sample adjacency similarity, and using the sample adjacency similarity to perform upsampling processing on the sample to be sampled features to obtain the target sample features.
[0182] In the embodiments of this application, please refer to Figure 4 The process of obtaining the target features in step S405 will not be described in detail here.
[0183] Step S807: predict the target sample features through the initial task prediction network in the initial image processing model to obtain the sample detection result of the image processing task.
[0184] In the embodiments of this application, please refer to Figure 4 The process of obtaining the image detection result in step S405 will not be described in detail here.
[0185] Step S808: Adjust the parameters of the initial image processing model based on the image processing annotation and sample detection results to obtain an image processing model.
[0186] In an embodiment of the present application, a difference loss function is generated based on the difference data between the image processing annotation and the sample detection result; the quality detection of the target sample features is performed to obtain image quality parameters, and the image quality parameters and the difference loss function are combined into a loss function; the loss function is used to adjust the parameters of the initial image processing model to obtain the image processing model. Alternatively, the difference loss function is directly determined as the loss function, and the loss function is used to adjust the parameters of the initial image processing model to obtain the image processing model. Among them, the parameters of the image processing model converge, and the image processing model includes a guidance network corresponding to the initial guidance network, a task main network corresponding to the initial task main network, and a task prediction network corresponding to the initial task prediction network.
[0187] Among them, in this embodiment, please refer to Figure 4 The steps in the process are equivalent to adding samples of the nouns involved on the basis of the reference steps to obtain the implementation process of the step, such as step S803. Figure 4 In step S402, "sample" is added to the nouns involved in the relevant description of step S402 (such as updating the first projection feature to the first sample projection feature, updating the second convolution feature to the second sample convolution feature, etc., which are not listed here one by one), and the implementation process of the sample-guided filtering feature in step S803 can be obtained.
[0188] Optionally, the training configuration adopted in this embodiment can be a default one or manually provided, which is not limited here. The training configuration can include but is not limited to a model optimization algorithm, a learning rate, a loss function, and the number of training iterations.
[0189] In an embodiment of the present application, the guiding features and the features to be sampled of the initial image can be obtained; the resolution of the guiding features is greater than the resolution of the features to be sampled; the guiding features and the features to be sampled are subjected to guided filtering processing to obtain guided filtering features, and the adjacent semantic similarity corresponding to the features to be sampled is determined through the guided filtering features and the features to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel point in the features to be sampled and the adjacent pixels of the pixel point; the guiding features are denoised to obtain denoised features, and the denoised features and the guiding features are detail-aware self-aligned to obtain the adjacent detail similarity corresponding to the features to be sampled; the adjacent semantic similarity and the adjacent detail similarity are combined into adjacent similarity, the features to be sampled are upsampled using the adjacent similarity to obtain target features, and the image detection results of the image processing task are predicted based on the target features. Through the above process, a guided filter feature aligned with the feature to be sampled can be generated by guided filtering. Semantic perception mutual alignment can be performed based on the guided filter feature to determine the semantic similarity between the guided feature and the feature to be sampled. At the same time, the guided feature is self-aligned to obtain the similarity of detail fidelity perception. By combining these two similarities, controllable alignment of features at different resolutions in terms of semantic content can be achieved, and the image detail information in upsampling can be restored, thereby improving the accuracy of the sampling process and the sampling performance.
[0190] For example, see Figure 9 , Figure 9 This is a schematic diagram of an interactive scenario provided by an embodiment of the present application. Figure 9 As shown, front-end A can provide an initial image to the back-end, which includes a deep network including the solution of the present application. The back-end obtains an image detection result corresponding to the initial image through the deep network, and feeds the image detection result back to front-end A, which can output the image detection result. Alternatively, the service device can send the initial image to the computer device, which predicts the image detection result of the initial image through the deep network including the solution of the present application, and feeds the image detection result back to the service device.
[0191] Since the solution of the present application realizes controllable feature alignment, shallower features can be directly selected as guiding features to help the upsampling process recover more image details, so that no matter what resolution the feature to be sampled is upsampled, the guiding features can be directly used to upsample the feature to be sampled. For example, see Figure 10 , Figure 10This is a schematic diagram of upsampling optimization provided by an embodiment of the present application. Figure 10 As shown in FIG, when image C4 needs to be upsampled, either upsample C4 directly, which results in poor image quality due to the lack of image detail information, or Figure 10 As shown in (a), upsampling is performed layer by layer on the basis of image C1, that is, "image C1 -> image C2 -> image C3 -> image C4", and the image after upsampling of image C4 is obtained. When the resolution of the image to be sampled is too low, the process needs to be upsampled multiple times, which makes the upsampling efficiency low and wastes resources. By adopting the solution of the present application, when upsampling the image C4, the image C4 can be directly upsampled. Figure 10 In the method shown in (b), image C1 is used to upsample image C4 to obtain image Since the solution of the present application realizes controllable alignment of features at different resolutions, direct upsampling processing can be performed directly under large resolution differences, and the supplementation of image detail information in the upsampling processing can be guaranteed, thereby improving the quality of the upsampling processing. At the same time, only one upsampling is required, thereby improving the efficiency of the upsampling processing.
[0192] The present application and its existing upsampling algorithm were verified on the same data set, and the performance of the present application and the existing upsampling algorithm were shown in Table 1 below:
[0193] Table 1
[0194]
[0195] As shown in Table 1 above, it can be seen that compared with other upsampling algorithms, all indicators of this application have been improved. The upsampling algorithm implemented in this application has made great progress compared with other upsampling algorithms.
[0196] Further, see Figure 11 , Figure 11 Schematic diagram of an image sampling and processing device provided in an embodiment of the present application. The image sampling and processing device can be a computer program (including program code, etc.) running on a computer device. For example, the image sampling and processing device can be an application software; the device can be used to execute the corresponding steps of the method provided in an embodiment of the present application. Figure 11 As shown, the image sampling processing device 1100 can be used to Figure 4 The computer device in the corresponding embodiment, specifically, the device may include: a feature acquisition module 11, a guided filtering module 12, a semantic parsing module 13, a detail parsing module 14, a feature sampling module 15 and an image processing module 16.
[0197] The feature acquisition module 11 is used to acquire the guiding features and the features to be sampled of the initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled;
[0198] A guided filtering module 12 is configured to perform guided filtering on the guided features and the features to be sampled to obtain guided filtering features;
[0199] The semantic parsing module 13 is used to perform mutual similarity detection between the guided filter feature and the feature to be sampled, and determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels;
[0200] The detail analysis module 14 is used to perform denoising on the guide feature to obtain a denoised feature, perform detail-aware self-alignment on the denoised feature and the guide feature to obtain the adjacent detail similarity corresponding to the feature to be sampled;
[0201] A feature sampling module 15 is configured to combine the adjacent semantic similarity and the adjacent detail similarity into an adjacent similarity, and perform upsampling processing on the to-be-sampled features using the adjacent similarity to obtain target features;
[0202] The image processing module 16 is used to predict the image detection result of the image processing task based on the target features.
[0203] The feature acquisition module 11 is specifically configured to:
[0204] Performing convolution processing on the initial image to obtain a first convolution feature, and performing normalization processing on the first convolution feature to obtain a normalized feature;
[0205] Perform convolution processing on the normalized features to obtain the guiding features of the initial image;
[0206] In the task main network corresponding to the image processing task, the features to be sampled of the initial image are obtained.
[0207] The guide filtering module 12 is specifically configured to:
[0208] Perform linear projection on the guide feature to obtain the first projection feature, and perform linear projection on the feature to be sampled to obtain the second projection feature;
[0209] Upsampling the second projection feature to obtain a first sampling feature, constructing a guided filtering model based on the first projection feature and the first sampling feature, and parsing the guided filtering model to obtain a guided filtering feature;
[0210] The semantic parsing module 13 is specifically used to:
[0211] The guided filtering feature and the first sampling feature are subjected to feature fusion processing to obtain the adjacent semantic similarity corresponding to the feature to be sampled.
[0212] When upsampling the second projection feature to obtain the first sampling feature, the semantic parsing module 12 is used to:
[0213] Obtaining a first resolution of the first projection feature, and obtaining a second resolution of the second projection feature;
[0214] Determine a ratio of the first projection feature to the second resolution as a first upsampling factor;
[0215] The second projection feature is upsampled by bilinear interpolation based on the first upsampling factor to obtain a first sampling feature.
[0216] When a guided filtering model is constructed based on the first projection feature and the first sampling feature, and the guided filtering model is parsed to obtain the guided filtering feature, the guided filtering module 12 is used to:
[0217] Constructing a semantic adjacency window, constructing an adjacency difference parameter in the semantic adjacency window, and constructing a guided filtering model according to the adjacency difference parameter, the first projection feature, and the first sampling feature;
[0218] Analyze the guided filtering model to obtain a first value of the adjacent difference parameter;
[0219] Performing averaging processing on the first value of the adjacent difference parameter based on the semantic adjacent window to obtain a filtering parameter;
[0220] The filtering parameters are fused with the first projection features to obtain the guided filtering features.
[0221] When the guided filtering feature and the first sampling feature are subjected to feature fusion processing to obtain the adjacent semantic similarity corresponding to the feature to be sampled, the semantic parsing module 13 is used to:
[0222] Obtain the i-th adjacent sampling feature of the first adjacent pixel of the i-th pixel from the first sampling feature, and combine the i-th adjacent sampling features corresponding to the i-th pixel into an i-th adjacent sampling feature matrix; i is a positive integer; the interval between any two adjacent pixels of the i-th pixel and the first adjacent pixel of the i-th pixel is the second upsampling multiple;
[0223] The product of the i-th pixel filter feature of the i-th pixel in the guided filter feature and the transpose of the i-th adjacent sampling feature matrix is determined as the adjacent semantic similarity corresponding to the i-th pixel.
[0224] The detail analysis module 14 is specifically used to:
[0225] Construct a denoising convolution kernel, convolve the denoising convolution kernel with the guide feature to obtain the denoising feature;
[0226] From the denoising features, obtain the adjacent denoising features of the first adjacent pixel points corresponding to M pixels respectively, and form the adjacent denoising features corresponding to each pixel point into an adjacent denoising feature matrix of the pixel point; M is a positive integer;
[0227] The pixel guidance feature of each pixel in the guidance feature is fused with the adjacent denoising feature matrix of the pixel to obtain the adjacent detail similarity corresponding to M pixels.
[0228] When upsampling the to-be-sampled features by using the adjacency similarity to obtain the target features, the feature sampling module 15 is used to:
[0229] Activate the adjacent similarity corresponding to the i-th pixel to obtain the i-th activation similarity; i is a positive integer;
[0230] Using the i-th activation similarity, perform weighted summation on the initial pixel features of the second adjacent pixel of the i-th pixel in the feature to be sampled to obtain the i-th sampled pixel feature of the i-th pixel;
[0231] When the post-sampling pixel features of all pixels are obtained, the post-sampling pixel features of all pixels are combined into target features.
[0232] The image processing module 16 is specifically configured to:
[0233] Input the target features into the task prediction network corresponding to the image processing task, predict the target features through the task prediction network, and obtain the image detection results of the image processing task;
[0234] The device also includes:
[0235] The result output module 17 is used to output the image detection result and perform the service corresponding to the image processing task based on the image detection result.
[0236] An embodiment of the present application provides an image sampling and processing device, which can obtain guiding features and features to be sampled of an initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled; guided filtering is performed on the guiding features and the features to be sampled to obtain guided filtering features, and the adjacent semantic similarity corresponding to the features to be sampled is determined through the guided filtering features and the features to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel point in the features to be sampled and the adjacent pixels of the pixel point; the guiding features are denoised to obtain denoised features, and the denoised features and the guiding features are detail-aware self-aligned to obtain the adjacent detail similarity corresponding to the features to be sampled; the adjacent semantic similarity and the adjacent detail similarity are combined into adjacent similarity, the features to be sampled are upsampled using the adjacent similarity to obtain target features, and image detection results of image processing tasks are predicted based on the target features. Through the above process, a guided filter feature aligned with the feature to be sampled can be generated by guided filtering. Semantic perception mutual alignment can be performed based on the guided filter feature to determine the semantic similarity between the guided feature and the feature to be sampled. At the same time, the guided feature is self-aligned to obtain the similarity of detail fidelity perception. By combining these two similarities, controllable alignment of features at different resolutions in terms of semantic content can be achieved, and the image detail information in upsampling can be restored, thereby improving the accuracy of the sampling process and the sampling performance.
[0237] Further, see Figure 12 , Figure 12 Schematic diagram of another image sampling and processing device provided in an embodiment of the present application. The image sampling and processing device can be a computer program (including program code, etc.) running on a computer device. For example, the image sampling and processing device can be an application software; the device can be used to execute the corresponding steps of the method provided in an embodiment of the present application. Figure 12 As shown, the image sampling processing device 1200 can be used to Figure 8 The computer device in the corresponding embodiment, specifically, the device may include: a sample acquisition module 21, a sample processing module 22, a sample guidance module 23, a first analysis module 24, a second analysis module 25, a sample sampling module 26, a sample prediction module 27 and a model training module 28.
[0238] A sample acquisition module 21 is used to acquire image samples and image processing annotations of image samples in image processing tasks;
[0239] The sample processing module 22 is configured to input the image sample into the initial image processing model, obtain the sample guiding features of the image sample through the initial guiding network in the initial image processing model, and obtain the sample to-be-sampled features of the image sample through the initial task main network in the initial image processing model; the resolution of the sample guiding features is greater than the resolution of the sample to-be-sampled features;
[0240] The sample guiding module 23 is used to perform guided filtering processing on the sample guiding feature and the sample to be sampled feature to obtain the sample guided filtering feature;
[0241] The first parsing module 24 is used to perform mutual similarity detection on the sample guided filter feature and the sample to be sampled feature, and determine the sample adjacent semantic similarity corresponding to the sample to be sampled feature; the sample adjacent semantic similarity is used to represent the semantic similarity between each pixel in the sample to be sampled feature and its adjacent pixels;
[0242] The second parsing module 25 is used to perform denoising on the sample guide feature to obtain a sample denoising feature, perform detail-aware self-alignment on the sample denoising feature and the sample guide feature to obtain the sample adjacent detail similarity corresponding to the sample to-be-sampled feature;
[0243] The sample sampling module 26 is configured to combine the sample adjacency semantic similarity and the sample adjacency detail similarity to form the sample adjacency similarity, and use the sample adjacency similarity to perform upsampling processing on the sample to be sampled to obtain the target sample feature;
[0244] The sample prediction module 27 is used to predict the target sample features through the initial task prediction network in the initial image processing model to obtain the sample detection results of the image processing task;
[0245] The model training module 28 is used to adjust the parameters of the initial image processing model through image processing annotation and sample detection results to obtain an image processing model.
[0246] The sample acquisition module 21 is specifically configured to:
[0247] Parse the image processing task, obtain the task processing parameters, and obtain the image labels of the candidate images;
[0248] Determine the candidate images whose image labels meet the task processing parameters as the initial samples, perform quality detection on the initial samples, and obtain the sample quality corresponding to the initial samples;
[0249] An initial sample whose sample quality is greater than or equal to a standard quality threshold is determined as an image sample, the image sample and the image processing task are sent to a management device, and the image processing annotation sent by the management device is obtained.
[0250] The model training module 28 is specifically used to:
[0251] Generate a difference loss function based on the difference data between the image processing annotation and the sample detection results;
[0252] Perform quality detection on the target sample features to obtain image quality parameters, and combine the image quality parameters with the difference loss function to form a loss function;
[0253] The loss function is used to adjust the parameters of the initial image processing model to obtain the image processing model.
[0254] Further, see Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 13 As shown, the computer device in the embodiment of the present application may include: one or more processors 1301, a memory 1302, and an input / output interface 1303. The processor 1301, the memory 1302, and the input / output interface 1303 are connected via a bus 1304. The memory 1302 is used to store computer programs, which include program instructions. The input / output interface 1303 is used to receive and output data, such as for data exchange between the computer device and business equipment. The processor 1301 is used to execute the program instructions stored in the memory 1302.
[0255] The processor 1301 is deployed in a prediction device for an image processing task and can perform the following operations:
[0256] Obtaining the guiding features and the features to be sampled of the initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled;
[0257] Performing guided filtering on the guided features and the features to be sampled to obtain guided filtering features;
[0258] Perform mutual similarity detection on the guided filter feature and the feature to be sampled to determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels;
[0259] Denoising the guide feature to obtain the denoised feature, performing detail-aware self-alignment on the denoised feature and the guide feature to obtain the adjacent detail similarity corresponding to the feature to be sampled;
[0260] The adjacent semantic similarity and the adjacent detail similarity are combined into the adjacent similarity, and the adjacent similarity is used to upsample the unsampled features to obtain the target features. The image detection results of the image processing task are predicted based on the target features.
[0261] When the processor 1301 is used to train the image processing model, it can perform the following operations:
[0262] Obtain image samples and image processing annotations of image samples in image processing tasks;
[0263] The image sample is input into the initial image processing model, the sample guiding feature of the image sample is obtained through the initial guiding network in the initial image processing model, and the sample to-be-sampled feature of the image sample is obtained through the initial task main network in the initial image processing model; the resolution of the sample guiding feature is greater than the resolution of the sample to-be-sampled feature;
[0264] Performing guided filtering on the sample guided features and the sample to-be-sampled features to obtain the sample guided filtering features;
[0265] Perform mutual similarity detection on the sample guided filter feature and the sample to be sampled feature to determine the sample adjacent semantic similarity corresponding to the sample to be sampled feature; the sample adjacent semantic similarity is used to represent the semantic similarity between each pixel in the sample to be sampled feature and the adjacent pixel of the pixel;
[0266] Denoising the sample guide feature to obtain the sample denoising feature, performing detail-aware self-alignment on the sample denoising feature and the sample guide feature to obtain the sample adjacency detail similarity corresponding to the sample to-be-sampled feature;
[0267] The sample adjacency semantic similarity and the sample adjacency detail similarity are combined to form the sample adjacency similarity, and the sample adjacency similarity is used to upsample the sample features to obtain the target sample features;
[0268] The target sample features are predicted through the initial task prediction network in the initial image processing model to obtain the sample detection results of the image processing task;
[0269] Through image processing annotation and sample detection results, the parameters of the initial image processing model are adjusted to obtain the image processing model.
[0270] In some feasible implementations, the processor 1301 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0271] The memory 1302 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1301 and the input / output interface 1303. A portion of the memory 1302 may also include a non-volatile random access memory. For example, the memory 1302 may also store device type information.
[0272] In a specific implementation, the computer device can execute the following operations through its built-in functional modules: Figure 4 or Figure 8 For details on the implementation methods provided in each step, please refer to the Figure 4 or Figure 8 The implementation methods provided in each step are not repeated here.
[0273] The embodiment of the present application provides a computer device, including: a processor, an input and output interface, and a memory, wherein the processor obtains a computer program in the memory and executes the computer program. Figure 4The various steps of the method shown in are used to perform image sampling and processing operations. The embodiment of the present application realizes obtaining the guiding features and the features to be sampled of the initial image; the resolution of the guiding features is greater than the resolution of the features to be sampled; the guiding features and the features to be sampled are subjected to guided filtering processing to obtain the guiding filtering features, and the adjacent semantic similarity corresponding to the features to be sampled is determined through the guiding filtering features and the features to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel point in the features to be sampled and the adjacent pixel points of the pixel point; the guiding features are subjected to denoising processing to obtain denoising features, and the denoising features and the guiding features are subjected to detail-aware self-alignment to obtain the adjacent detail similarity corresponding to the features to be sampled; the adjacent semantic similarity and the adjacent detail similarity are combined into adjacent similarity, and the features to be sampled are upsampled using the adjacent similarity to obtain the target features, and the image detection results of the image processing task are predicted based on the target features. Through the above process, a guided filter feature aligned with the feature to be sampled can be generated by guided filtering. Semantic perception mutual alignment can be performed based on the guided filter feature to determine the semantic similarity between the guided feature and the feature to be sampled. At the same time, the guided feature is self-aligned to obtain the similarity of detail fidelity perception. By combining these two similarities, controllable alignment of features at different resolutions in terms of semantic content can be achieved, and the image detail information in upsampling can be restored, thereby improving the accuracy of the sampling process and the sampling performance.
[0274] The present invention also provides a computer-readable storage medium storing a computer program suitable for being loaded and executed by the processor. Figure 4 The image sampling and processing methods provided in each step are detailed in this Figure 4 The implementation methods provided in each step will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the computer program can be deployed to be executed on one computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected by a communication network.
[0275] The computer-readable storage medium may be the image sampling and processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0276] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs Figure 4 The method provided in the various optional methods realizes the generation of a guided filter feature that is aligned with the feature to be sampled through guided filtering, and semantic perception mutual alignment can be performed based on the guided filter feature to determine the semantic similarity between the guided feature and the feature to be sampled. At the same time, the guided feature is self-aligned to obtain the similarity of detail fidelity perception. By combining these two similarities, controllable alignment of features at different resolutions in terms of semantic content is achieved, and image detail information is restored in upsampling, thereby improving the accuracy of sampling processing and sampling performance.
[0277] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0278] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0279] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in this description according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0280] The methods and related devices provided by the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided by the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable image sampling and processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable image sampling and processing device generate instructions for implementing the process in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable image sampling and processing equipment to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable image sampling processing device, so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.
[0281] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0282] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0283] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An image sampling and processing method, characterized in that: The method comprises: Acquire a guiding feature and a feature to be sampled of an initial image; wherein the resolution of the guiding feature is greater than the resolution of the feature to be sampled; Performing guided filtering on the guided feature and the feature to be sampled to obtain a guided filtering feature; Performing mutual similarity detection on the guided filter feature and the feature to be sampled to determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels; Denoising the guiding feature to obtain a denoised feature, performing detail-aware self-alignment on the denoised feature and the guiding feature to obtain adjacent detail similarity corresponding to the feature to be sampled; The adjacent semantic similarity and the adjacent detail similarity are combined into an adjacent similarity, the feature to be sampled is upsampled using the adjacent similarity to obtain a target feature, and an image detection result of an image processing task is predicted based on the target feature.
2. The method according to claim 1, wherein The step of obtaining the guiding features and the features to be sampled of the initial image includes: Performing convolution processing on the initial image to obtain a first convolution feature, and performing normalization processing on the first convolution feature to obtain a normalized feature; Performing convolution processing on the normalized features to obtain guiding features of the initial image; In a task main network corresponding to the image processing task, features to be sampled of the initial image are obtained.
3. The method according to claim 1, wherein The performing guided filtering on the guided feature and the feature to be sampled to obtain the guided filtering feature includes: Performing linear projection on the guide feature to obtain a first projection feature, and performing linear projection on the feature to be sampled to obtain a second projection feature; performing upsampling processing on the second projection feature to obtain a first sampling feature, constructing a guided filtering model based on the first projection feature and the first sampling feature, and parsing the guided filtering model to obtain a guided filtering feature; The performing mutual similarity detection on the guided filtering feature and the feature to be sampled to determine the adjacent semantic similarity corresponding to the feature to be sampled includes: A mutual similarity test is performed on the guided filtering feature and the first sampling feature to obtain an adjacent semantic similarity corresponding to the feature to be sampled.
4. The method according to claim 3, wherein The upsampling of the second projection feature to obtain the first sampling feature includes: Obtaining a first resolution of the first projection feature, and obtaining a second resolution of the second projection feature; Determine a ratio of the first projection feature to the second resolution as a first upsampling factor; The second projection feature is up-sampled by bilinear interpolation based on the first up-sampling multiple to obtain a first sampling feature.
5. The method according to claim 3, wherein The step of constructing a guided filtering model according to the first projection feature and the first sampling feature, and analyzing the guided filtering model to obtain a guided filtering feature includes: Constructing a semantic adjacency window, constructing an adjacency difference parameter in the semantic adjacency window, and constructing a guided filtering model according to the adjacency difference parameter, the first projection feature, and the first sampling feature; Analyzing the guided filtering model to obtain a first value of the adjacency difference parameter; Performing averaging processing on the first value of the adjacency difference parameter based on the semantic adjacency window to obtain a filtering parameter; The filtering parameters are fused with the first projection features to obtain guided filtering features.
6. The method according to claim 3, wherein The performing mutual similarity detection on the guided filtering feature and the first sampling feature to obtain the adjacent semantic similarity corresponding to the feature to be sampled includes: Obtaining, from the first sampling features, an i-th adjacent sampling feature of a first adjacent pixel of an i-th pixel, and combining the i-th adjacent sampling features corresponding to the i-th pixel into an i-th adjacent sampling feature matrix; i is a positive integer; and an interval between any two adjacent pixels of the i-th pixel and the first adjacent pixel of the i-th pixel is a second upsampling multiple; The product of the i-th pixel filtering feature of the i-th pixel point in the guided filtering feature and the transpose of the i-th adjacent sampling feature matrix is determined as the adjacent semantic similarity corresponding to the i-th pixel point.
7. The method according to claim 1, wherein The denoising process is performed on the guiding feature to obtain a denoised feature, and detail-aware self-alignment is performed on the denoised feature and the guiding feature to obtain the adjacent detail similarity corresponding to the feature to be sampled, including: Constructing a denoising convolution kernel, and performing convolution processing on the denoising convolution kernel and the guide feature to obtain a denoising feature; From the denoising features, obtain the adjacent denoising features of the first adjacent pixel points corresponding to M pixels respectively, and form the adjacent denoising features corresponding to each pixel point into an adjacent denoising feature matrix of the pixel point; M is a positive integer; The pixel guidance feature of each pixel in the guidance feature is fused with the adjacent denoising feature matrix of the pixel to obtain the adjacent detail similarity corresponding to M pixels.
8. The method according to claim 1, wherein The upsampling process of the feature to be sampled by using the adjacency similarity to obtain the target feature includes: Activate the adjacent similarity corresponding to the i-th pixel to obtain the i-th activation similarity; i is a positive integer; Using the i-th activation similarity, weighted summation is performed on the initial pixel features of the second adjacent pixel of the i-th pixel in the feature to be sampled to obtain the i-th sampled pixel feature of the i-th pixel; When the sampled pixel features of all the pixels are acquired, the sampled pixel features of all the pixels are combined into target features.
9. The method according to claim 1, wherein The image detection result of the image processing task predicted based on the target feature includes: Inputting the target feature into a task prediction network corresponding to the image processing task, predicting the target feature through the task prediction network, and obtaining an image detection result of the image processing task; The method further comprises: Output the image detection result, and execute the business corresponding to the image processing task based on the image detection result.
10. An image sampling and processing method, characterized in that: The method comprises: Obtaining an image sample and an image processing annotation of the image sample in an image processing task; Inputting the image sample into an initial image processing model, obtaining a sample guiding feature of the image sample through an initial guiding network in the initial image processing model, and obtaining a sample feature to be sampled of the image sample through an initial task main network in the initial image processing model; the resolution of the sample guiding feature is greater than the resolution of the sample feature to be sampled; Performing guided filtering on the sample guided features and the sample to-be-sampled features to obtain sample guided filtering features; Performing a mutual similarity test on the sample guided filtering feature and the sample feature to be sampled to determine the sample adjacent semantic similarity corresponding to the sample feature to be sampled; the sample adjacent semantic similarity is used to represent the semantic similarity between each pixel in the sample feature to be sampled and its adjacent pixel points; Denoising the sample guide feature to obtain a sample denoising feature, performing detail-aware self-alignment on the sample denoising feature and the sample guide feature to obtain a sample adjacency detail similarity corresponding to the sample feature to be sampled; The sample adjacency semantic similarity and the sample adjacency detail similarity are combined to form a sample adjacency similarity, and the sample adjacency similarity is used to perform upsampling processing on the sample to be sampled features to obtain target sample features; Predicting the target sample features through the initial task prediction network in the initial image processing model to obtain a sample detection result of the image processing task; The parameters of the initial image processing model are adjusted according to the image processing annotation and the sample detection result to obtain an image processing model.
11. The method according to claim 10, wherein The obtaining of image samples and image processing annotation of the image samples in the image processing task includes: Parse the image processing task, obtain the task processing parameters, and obtain the image labels of the candidate images; Determine the candidate image whose image label meets the task processing parameter as an initial sample, perform quality detection on the initial sample, and obtain the sample quality corresponding to the initial sample; The initial sample whose sample quality is greater than or equal to the standard quality threshold is determined as an image sample, the image sample and the image processing task are sent to a management device, and the image processing annotation sent by the management device is obtained.
12. The method according to claim 10, wherein The step of adjusting parameters of the initial image processing model based on the image processing annotation and the sample detection result to obtain the image processing model includes: generating a difference loss function based on difference data between the image processing annotation and the sample detection result; Performing quality detection on the target sample features to obtain image quality parameters, and combining the image quality parameters with the difference loss function to form a loss function; The loss function is used to adjust the parameters of the initial image processing model to obtain an image processing model.
13. An image sampling and processing device, characterized in that: The device comprises: A feature acquisition module, configured to acquire guiding features and features to be sampled from an initial image; the resolution of the guiding features being greater than the resolution of the features to be sampled; A guided filtering module, configured to perform guided filtering on the guided feature and the feature to be sampled to obtain a guided filtering feature; A semantic parsing module is used to perform mutual similarity detection on the guided filter feature and the feature to be sampled, and determine the adjacent semantic similarity corresponding to the feature to be sampled; the adjacent semantic similarity is used to represent the semantic similarity between each pixel in the feature to be sampled and its adjacent pixels; a detail parsing module, configured to perform denoising on the guiding feature to obtain a denoised feature, perform detail-aware self-alignment on the denoised feature and the guiding feature, and obtain a similarity of adjacent details corresponding to the feature to be sampled; a feature sampling module, configured to combine the adjacent semantic similarity and the adjacent detail similarity into an adjacent similarity, and perform upsampling processing on the feature to be sampled using the adjacent similarity to obtain a target feature; An image processing module is used to predict the image detection result of the image processing task based on the target features.
14. An image sampling and processing device, characterized in that: The device comprises: A sample acquisition module, used to acquire image samples and image processing annotations of the image samples in image processing tasks; a sample processing module, configured to input the image sample into an initial image processing model, obtain sample guiding features of the image sample through an initial guiding network in the initial image processing model, and obtain sample features to be sampled of the image sample through an initial task main network in the initial image processing model; the resolution of the sample guiding features is greater than the resolution of the sample features to be sampled; A sample guiding module, configured to perform guided filtering processing on the sample guiding feature and the sample to-be-sampled feature to obtain a sample guided filtering feature; A first parsing module is configured to perform mutual similarity detection on the sample guided filtering feature and the sample feature to be sampled, and determine the sample adjacency semantic similarity corresponding to the sample feature to be sampled; the sample adjacency semantic similarity is used to represent the semantic similarity between each pixel in the sample feature to be sampled and its adjacent pixels; A second parsing module is configured to perform denoising on the sample guiding feature to obtain a sample denoising feature, perform detail-aware self-alignment on the sample denoising feature and the sample guiding feature to obtain a sample adjacency detail similarity corresponding to the sample feature to be sampled; a sample sampling module, configured to combine the sample adjacency semantic similarity and the sample adjacency detail similarity to form a sample adjacency similarity, and use the sample adjacency similarity to perform upsampling processing on the sample to be sampled features to obtain target sample features; A sample prediction module is used to predict the target sample features through the initial task prediction network in the initial image processing model to obtain the sample detection result of the image processing task; The model training module is used to adjust the parameters of the initial image processing model through the image processing annotation and the sample detection results to obtain an image processing model.
15. A computer device, characterized in that: Includes processor, memory, input and output interfaces; The processor is connected to the memory and the input / output interface, respectively, wherein the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method described in any one of claims 1 to 9, or executes the method described in any one of claims 10 to 12.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 9, or executes the method according to any one of claims 10 to 12.
17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented, or the method according to any one of claims 10 to 12 is executed.