An image processing method, apparatus, device and medium

By combining UAV image acquisition and feature extraction models with surface and edge saliency features, the problem of insufficient grayscale range in UAV remote sensor images was solved, enabling high-precision segmentation and detection of features of power transmission line components.

CN116883876BActive Publication Date: 2026-02-10GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310884319.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2026-02-10
Estimated Expiration
2043-07-18

AI Technical Summary

Technical Problem

The insufficient grayscale range of images acquired by UAV remote sensors results in poor visual contrast in power transmission line images, making it difficult to accurately extract features of power transmission line components for segmentation and defect detection.

Method used

Images are acquired using an image acquisition device, and surface salient features and edge salient features of the target part are extracted using a feature extraction model. The edge enhancement salient features and surface fusion salient features are combined to ensure the integrity of the target part features.

Benefits of technology

It improves the accuracy of target part segmentation, ensuring the integrity and precision of power transmission line inspection, especially in complex terrain and environments with large-scale grayscale variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883876B_ABST
    Figure CN116883876B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, device, equipment and medium. The method obtains a collection image of a detection area through an image collector, extracts surface saliency features and edge saliency features of a target part in the collection image based on a feature extraction model, determines edge enhanced saliency features of the target part according to the surface saliency features and the edge saliency features, determines surface fusion saliency features of the target part according to the edge enhanced saliency features and the surface saliency features, and determines target features of the target part according to the edge enhanced saliency features and the surface fusion saliency features. The technical scheme ensures the integrity of the target part features and improves the accuracy of target part segmentation by utilizing the complementarity between the surface feature information and the edge feature information of the target part.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power transmission line detection, and in particular to an image processing method and device, equipment and a medium. BACKGROUND

[0002] Due to the vast territory and complex topography of China, overhead power transmission lines often need to pass through forests, mountains, hills and other complex terrains. Overhead power transmission lines are exposed to the wild for a long time and need to withstand the natural wear and tear of equipment, human damage, lightning, snow and rain and other external risks; at the same time, the power transmission line also long-term withstands the impact of high current. With the increase of the operation time of the power transmission line, the above problems are easy to cause damage or fault hazards to the main components in the power transmission line.

[0003] For a long time, in order to ensure the safe and stable operation of power, the power inspection work is mainly in the form of manual inspection. This method not only has low inspection efficiency, but also has personal safety hazards for the inspection personnel. In recent years, the development of unmanned aerial vehicles provides a new method and means for the stable operation and maintenance of power transmission lines. Using unmanned aerial vehicle remote sensing technology to automatically detect power transmission lines has become an important way to improve work efficiency, protect the safety of workers and maintain the stable operation of power.

[0004] However, the gray scale range of the image obtained by the remote sensor of the unmanned aerial vehicle cannot cover the entire gray scale range that it can reach, so that the visual contrast of the image is poor, and it is difficult to accurately extract the features of the power transmission line parts from the image to segment the parts for defect detection, and the power transmission line cannot be well monitored. SUMMARY

[0005] The present application provides an image processing method, device, equipment and medium, which utilizes the complementarity between the surface feature information and the edge feature information of the target part to ensure the integrity of the target part features and improve the accuracy of target part segmentation.

[0006] According to an aspect of the present application, an image processing method is provided, which comprises:

[0007] acquiring a collection image of a detection area by an image collector, and extracting surface saliency features and edge saliency features of a target part in the collection image based on a feature extraction model;

[0008] determining edge-enhanced saliency features of the target part according to the surface saliency features and the edge saliency features;

[0009] determining surface-fused saliency features of the target part according to the edge-enhanced saliency features and the surface saliency features;

[0010] According to the edge reinforced saliency feature and the surface fused saliency feature, a target feature of the target part is determined.

[0011] According to another aspect of the present application, an image processing apparatus is provided, the apparatus comprising:

[0012] A part feature extraction module is configured to acquire a captured image of a region to be detected by an image collector, and extract a surface saliency feature and an edge saliency feature of a target part in the captured image based on a feature extraction model, respectively;

[0013] An edge feature reinforcement module is configured to determine an edge reinforced saliency feature of the target part according to the surface saliency feature and the edge saliency feature;

[0014] A surface feature fusion module is configured to determine a surface fused saliency feature of the target part according to the edge reinforced saliency feature and the surface saliency feature;

[0015] A target feature determination module is configured to determine a target feature of the target part according to the edge reinforced saliency feature and the surface fused saliency feature.

[0016] According to another aspect of the present application, an electronic device is provided, the device comprising:

[0017] At least one processor; and

[0018] A memory connected in communication with the at least one processor; wherein,

[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the image processing method according to any one of the embodiments of the present application.

[0020] According to another aspect of the present application, a computer readable storage medium is provided, the computer readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image processing method according to any one of the embodiments of the present application when executed by the processor.

[0021] The technical scheme provided in the application obtains the acquisition image of the to-be-detected region through the image collector, and extracts the surface saliency feature and the edge saliency feature of the target part in the acquisition image based on a feature extraction model respectively; determines the edge reinforced saliency feature of the target part according to the surface saliency feature and the edge saliency feature; determines the surface fusion saliency feature of the target part according to the edge reinforced saliency feature and the surface saliency feature; and determines the target feature of the target part according to the edge reinforced saliency feature and the surface fusion saliency feature. The technical scheme uses the complementarity between the surface feature information and the edge feature information of the target part to ensure the integrity of the target part feature and improve the accuracy of target part segmentation.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0024] Figure 1 A flow chart of an image processing method provided for the first embodiment of the application;

[0025] Figure 2 A flow chart of an image processing method provided for the second embodiment of the application;

[0026] Figure 3 A structural schematic diagram of a feature extraction model provided for the second embodiment of the application;

[0027] Figure 4 A flow chart of an image processing method provided for the third embodiment of the application;

[0028] Figure 5 A structural schematic diagram of an image processing device provided for the fourth embodiment of the application;

[0029] Figure 6 A structural schematic diagram of an image processing method device for implementing the embodiments of the application. DETAILED DESCRIPTION

[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0031] It should be noted that the terms "first," "second," "third," "target," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] Example 1

[0033] Figure 1 This is a flowchart of an image processing method provided in Embodiment 1 of this application. This embodiment is applicable to the segmentation of targets in an image. The method can be executed by an image processing device, which can be implemented in hardware and / or software, and can be configured in a device with data processing capabilities. Figure 1 As shown, the method includes:

[0034] S110. Acquire images of the area to be detected using an image acquisition device, and extract the surface salient features and edge salient features of the target part in the acquired images based on a feature extraction model.

[0035] The image acquisition device can be a CMOS image sensor (Complementary Metal-Oxide-Semiconductor), a high-resolution CCD (Charge Coupled Device) digital camera, an optical camera, a multispectral imager, an infrared scanner, a laser scanner, a magnetometer, a synthetic aperture radar, etc. The acquired images can be optical photographs, image frames, remote sensing images, etc. The area to be inspected can be determined according to a pre-defined inspection task and route. Specifically, when the location information of the inspection equipment equipped with the image acquisition device is detected to be within the area to be inspected, the image acquisition device acquires images of that area.

[0036] For example, when drones inspect power transmission lines, remote sensing information of the area to be inspected can be obtained through remote sensing sensors mounted on the drone, and the remote sensing information can be processed by a computer and made into remote sensing images according to certain requirements.

[0037] The feature extraction model can be a neural network model. Sample images can be pre-input into the neural network to be trained to obtain prediction results. A loss function is formed based on the prediction results and the labels of the sample images to train the neural network.

[0038] The target part can be determined based on the actual inspection needs. For example, when inspecting a power transmission line, the target part can be the power transmission line itself; or, when inspecting a fishing area, the target part can be the fishing boat.

[0039] Among them, surface saliency features can be surface features of salient regions in the acquired image, and edge saliency features can be edge features of salient regions in the acquired image.

[0040] In this embodiment of the invention, the acquired image can be input into a pre-trained feature extraction model, and after processing through convolutional layers, max pooling layers, activation function layers, etc., the surface salient features and edge salient features of the target part in the acquired image can be obtained.

[0041] Optionally, surface saliency features and edge saliency features of the target part in the acquired image are extracted based on the feature extraction model, including: extracting edge path features of the target part in the acquired image based on the backbone network in the feature extraction model to obtain at least one first edge path feature and at least one second edge path feature; wherein, the level of the first edge path feature is lower than the level of the second edge path feature; extracting edge saliency features of the target part based on the at least one first edge path feature; and extracting surface saliency features of the target part based on the at least one second edge path feature.

[0042] For example, the backbone network includes convolutional layers, max pooling layers, and three IRB (Improved Residual Block) layers. The convolutional layers are used to obtain the first side path features, the max pooling layers are used to reduce the parameters in the feature extraction process, and the three IRB layers are used to obtain three second side path features.

[0043] Since lower-level edge path features retain more edge attributes than higher-level edge path features, in this embodiment of the invention, the edge saliency features of the target part can be extracted through the first edge path features, and the surface saliency features of the target part can be extracted through the second edge path features, thus obtaining an initial feature set f = {f1, f2, f3, f4}, where f1 is the edge saliency feature of the first level, f2 is the surface saliency feature of the second level, f3 is the surface saliency feature of the third level, and f4 is the surface saliency feature of the fourth level.

[0044] In this embodiment of the invention, a nonlinear activation function layer may be added after each convolutional layer of the backbone network 100 to ensure the nonlinearity of the feature extraction model.

[0045] S120. Determine the edge reinforcement salience feature of the target part based on the surface salience feature and the edge salience feature.

[0046] Edge saliency features are used to model salient edge information of target parts, but local feature information alone is insufficient to fully represent edge information. The larger the receptive field mapped by higher-level edge path features, the more accurate the localization. Therefore, in this embodiment of the invention, surface saliency features can be added to edge saliency features to suppress non-edge saliency information in the edge saliency features.

[0047] Optionally, determining the edge enhancement saliency features of the target part based on the surface saliency features and the edge saliency features includes: determining the edge enhancement saliency features of the target part based on the surface saliency features extracted by the edge saliency features and the highest-level second edge path features using a first saliency feature extraction network in the feature extraction model.

[0048] Since the receptive field mapped by the second side path feature at the highest level is the largest, the surface saliency features extracted by the second side path feature at the highest level can be added to the edge saliency features to determine the edge enhancement saliency features of the target part.

[0049] Specifically, the edge enhancement saliency features can be determined using the following formula:

[0050] F1=ψ(f1+Φ(Γ(Trans(f m,ε(f1))),μ(f1)));

[0051] Where F1 represents the edge enhancement saliency feature, f m Let f(f1) represent the highest level of surface saliency features, ε(f1) represent the number of feature channels of f1, Trans(f4, ε(f1)) represents changing the number of feature channels of f4 to ε(f1), Γ(Trans(f4, ε(f1))) represents a non-linear activation operation, μ(f1) represents the size of f1, Φ(Γ(Trans(f4, ε(f1))), μ(f1)) represents changing the size of Γ(Trans(f4, ε(f1))) to μ(f1), and ψ(f1+Φ(Γ(Trans(f4, ε(f1))), μ(f1))) represents a series of convolutions and non-linear operations.

[0052] S130. Determine the surface fusion salience feature of the target part based on the edge enhancement salience feature and the surface salience feature.

[0053] Guided by edge information, edge enhancement saliency features are fused into surface saliency features to obtain surface fusion saliency features, which enriches surface details and achieves more accurate surface prediction results.

[0054] Specifically, edge enhancement saliency features can be fused into surface saliency features at each level to obtain surface fusion saliency features at each level; alternatively, edge enhancement saliency features can be fused into the features obtained after fusion of surface saliency features at each level to obtain total surface fusion saliency features.

[0055] Optionally, determining the surface saliency features of the target part based on the edge enhancement saliency features and the surface saliency features includes: determining the surface enhancement saliency features of the current level based on the surface saliency features extracted by the second edge path features of the current level and the surface saliency features extracted by the second edge path features of the next level, using the second saliency feature extraction network in the feature extraction model; wherein the next level is higher than the current level; and determining the surface fusion saliency features of the target part based on the guiding network in the feature extraction model, using the edge enhancement saliency features and the surface enhancement saliency features of each level.

[0056] Since a larger receptive field in a higher-level feature map leads to more accurate localization, a high-to-low positional information propagation mechanism can be used to progressively transfer surface salient features from higher levels to lower levels. Specifically, the surface enhancement salient features at the current level can be determined using the following formula:

[0057] F n =ψ(fn +Φ(Γ(Trans(f n+1 , ε(f n ))), μ(f n )));

[0058] Among them, F n Let n be the surface enhancement saliency feature of the nth level, where n is the level other than the highest level;

[0059] F m =ψ(f m );

[0060] Among them, F m Let m be the surface enhancement saliency feature at level m, where m is the highest level.

[0061] In this embodiment of the invention, to achieve more accurate surface prediction by fusing the edge enhancement saliency features of the target part and the surface enhancement saliency features of each edge path under the guidance of the target part's edge information, a guiding network is established in the feature extraction model to accurately predict the surface position of the part and the edges of the lines under the guidance of edge information. More importantly, for the prediction of the surface position of the part for each edge path, by adding edge enhancement saliency features, the details of the segmented part surface can be richer, and the high-level prediction can be more accurate.

[0062] Specifically, the surface fusion saliency features of the target part can be determined using the following formula:

[0063] G n =ψ(Φ(Γ(Trans(F) n ,ε(F1))),μ(F1))+F1);

[0064] Among them, G n This is a significant feature of surface fusion.

[0065] S140. Determine the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features.

[0066] The target features include edge enhancement saliency features and surface fusion saliency features. In this embodiment of the invention, new images can be constructed by modeling using target features, which can preserve the complete line edges in the image, laying the foundation for surface defect detection of target parts.

[0067] This invention provides an image processing method. The method acquires an image of the region to be detected using an image acquisition device, and extracts surface saliency features and edge saliency features of the target part from the acquired image based on a feature extraction model. Based on the surface saliency features and edge saliency features, edge enhancement saliency features of the target part are determined. Based on the edge enhancement saliency features and surface saliency features, surface fusion saliency features of the target part are determined. Based on the edge enhancement saliency features and surface fusion saliency features, the target features of the target part are determined. This technical solution ensures the integrity of the target part's features and improves the accuracy of target part segmentation by complementing the surface feature information and edge feature information of the target part.

[0068] Example 2

[0069] Figure 2 This is a flowchart of an image processing method provided in Embodiment 2 of this application. This embodiment is an optimization based on the above embodiment. Figure 2 As shown, the method in this embodiment specifically includes the following steps:

[0070] S210. Acquire images of the area to be detected using an image acquisition device, and extract the surface salient features and edge salient features of the target part in the acquired images based on a feature extraction model.

[0071] S220. Based on the surface saliency features and the edge saliency features, determine the edge reinforcement saliency features of the target part.

[0072] S230. Based on the edge enhancement saliency feature and the surface saliency feature, determine the surface fusion saliency feature of the target part.

[0073] S240. Determine the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features.

[0074] S250. Based on the saliency supervision network in the feature extraction model, determine the binary cross-entropy loss value, structural similarity index loss value, and intersection-union ratio loss value of each output in the feature extraction model. The output of the feature extraction model includes at least one of the edge saliency feature, the surface saliency feature, the edge enhancement saliency feature, and the surface fusion saliency feature.

[0075] The binary cross-entropy loss value is the loss function in binary classification, and it can be determined by the following formula:

[0076] l BCE =-∑ (i,j)[GT(i,j)log(S(i,j))+(1-GT(i,j))log(1-S(i,j))];

[0077]

[0078] Among them, l BCE GT(i,j)∈{0,1} represents the true value of pixel (i,j), and S(i,j) is the predicted probability that pixel (i,j) becomes the surface or edge of the target part. It is a 3x3 transition convolutional layer with padding of 1 and 1 channel, used to transform multi-channel input feature maps into single-channel maps; F = {F1, F2, ..., F...} n} represents the set of edge-enhanced saliency features and surface-enhanced saliency features; G represents the set of surface-blended saliency features, Edge represents saliency edge ground truth, and Object represents saliency surface ground truth.

[0079] Among them, the structural similarity index loss value can capture the structural information in the acquired images. Introducing the structural similarity index into the training loss can yield very attractive results. Therefore, the structural similarity index loss value is introduced into the loss function of the feature extraction model to learn the structural information of salient edge ground truths and salient part surface ground truths. Specifically, it can be determined by the following formula:

[0080]

[0081] Among them, l SSMM Let x be the structural similarity index loss value, where x = {x} i,j :i,j=1,...,N}, y = {y i,j :i,j=1,...,N}, μ x Let σ be the mean of x. x Let x be the standard deviation, and μ be the standard deviation of x. y Let σ be the mean of y. y Let σ be the standard deviation of y. xy Let C1 and C2 be the covariance, and C1 and C2 be empirical constants. C1 can be 0.01. 2 C2 can be 0.03 2 .

[0082] The Cross-Union Ratio (CUI) loss value is used to represent the similarity measure between two sets, and is then used as a standard evaluation metric for target part segmentation and detection. Specifically, the CUI loss value can be determined by the following formula:

[0083]

[0084] Among them, l loU This is the crossover and union ratio loss value.

[0085] S260. Determine the total loss value of the feature extraction model based on the binary cross-entropy loss value, the structural similarity index loss value, and the cross-union ratio loss value.

[0086] The total loss value can be determined using the following formula:

[0087]

[0088]

[0089] Where Loss is the total loss value, l (n) Let be the loss value of the output of the nth edge path feature. The binary cross-entropy loss value is the output of the nth edge path feature. The structural similarity index loss value is the output of the nth edge path feature. This is the cross-union ratio loss value of the output of the nth edge path feature.

[0090] Since the feature extraction model structure proposed in this embodiment does not require loading any pre-trained model trained on the ImageNet dataset, and optimizes the part surface segmentation task by promoting mutual assistance between part surfaces and line edges, the accuracy of part surface prediction is significantly improved. Simultaneously, a hybrid loss function integrating binary cross-entropy, structural similarity index, and cross-union ratio is introduced to monitor the salient part surface and line edge predictions of the entire network at three levels: pixel level, region level, and mapping level, laying the foundation for part surface defect detection.

[0091] Figure 3 This is a schematic diagram of the structure of a feature extraction model provided in Embodiment 2 of the present invention. Figure 3 As shown, the feature extraction model includes a backbone network 100, a first salient feature extraction network 200, a second salient feature extraction network 300, a guiding network 400, and a salient supervision network.

[0092] The backbone network 100 includes a convolutional layer 101, a max pooling layer 102, a first improved residual block 103, a second improved residual block 104, and a third improved residual block 105. The convolutional layer 101 is used to generate edge saliency features f1, the max pooling layer 102 is used to reduce model parameters, the first improved residual block 103 is used to generate first surface saliency features f2, the second improved residual block 104 is used to generate second surface saliency features f3, and the third improved residual block 105 is used to generate third surface saliency features f4.

[0093] The first saliency feature extraction network 200 is used to first perform bilinear interpolation and convolution on the third surface saliency feature f4, then add it pixel by pixel to the edge saliency feature f1, and finally perform convolution to obtain the edge enhancement saliency feature F1.

[0094] The second saliency feature extraction network 300 is used to perform convolution processing on the third surface saliency feature f4 to obtain the third surface enhanced saliency feature F4; the third surface saliency feature f4 is first subjected to bilinear interpolation and convolution processing, then added to the second surface saliency feature f3 pixel by pixel, and finally subjected to convolution processing to obtain the second surface enhanced saliency feature F3; the second surface enhanced saliency feature F3 is first subjected to bilinear interpolation and convolution processing, then added to the first surface saliency feature f2 pixel by pixel, and finally subjected to convolution processing to obtain the first surface enhanced saliency feature F2.

[0095] The guiding network 400 is used to add the edge enhancement saliency feature F1 to the surface enhancement saliency features F2, F3, and F4 (which have undergone bilinear interpolation and convolution) pixel by pixel, and then perform convolution processing on each of them to obtain the first surface fusion saliency feature G2, the second surface fusion saliency feature G3, and the third surface fusion saliency feature G4. Finally, the first surface fusion saliency feature G2, the second surface fusion saliency feature G3, and the third surface fusion saliency feature G4 are added pixel by pixel to obtain the total surface fusion saliency feature G.

[0096] A saliency supervision network is used to calculate the loss value for the convolutionally processed edge enhancement saliency feature F1, first surface enhancement saliency feature F2, second surface enhancement saliency feature F3, third surface enhancement saliency feature F4, first surface fusion saliency feature G2, second surface fusion saliency feature G3, third surface fusion saliency feature G4 and total surface fusion saliency feature G.

[0097] This invention provides an image processing method. The method acquires an image of the region to be detected using an image acquisition device, and extracts surface saliency features and edge saliency features of the target part from the acquired image based on a feature extraction model. Based on the surface and edge saliency features, edge enhancement saliency features of the target part are determined. Based on the edge enhancement and surface saliency features, surface fusion saliency features of the target part are determined. Based on the edge enhancement and surface fusion saliency features, target features of the target part are determined. Based on the saliency supervision network in the feature extraction model, the binary cross-entropy loss value, structural similarity index loss value, and cross-union ratio (CUP) loss value of each output in the feature extraction model are determined. Based on the binary cross-entropy loss value, structural similarity index loss value, and CUP loss value, the total loss value of the feature extraction model is determined. This technical solution monitors the surface and edge predictions of the feature extraction model at three levels: pixel level, region level, and mapping level, to ensure the integrity of the target part features and improve the accuracy of target part segmentation.

[0098] Example 3

[0099] Figure 4 This is a flowchart of an image processing method provided in Embodiment 3 of this application. This embodiment is an optimization based on the above embodiment. Figure 4 As shown, the method in this embodiment specifically includes the following steps:

[0100] S310. Acquire images of the area to be detected using an image acquisition device.

[0101] S320. Divide the acquired image into regions to obtain at least one region; wherein the division criterion is the degree of pixel grayscale change.

[0102] Since the grayscale range of the image directly acquired by the image acquisition device cannot cover the entire grayscale range it can reach, the visual contrast of the acquired image is poor, making it difficult to segment the target parts in the image. Therefore, this embodiment of the invention divides the acquired image according to the degree of grayscale change to distinguish between areas with large grayscale changes and areas with small grayscale changes, thereby adjusting the grayscale of the acquired image.

[0103] The degree of pixel grayscale variation can be determined by pre-estimating the grayscale of sample image sets for different region labels to establish the region division criteria. For example, during the inspection of power transmission lines, the ground feature information in the images collected from different regions will vary. Some images are of plains areas dominated by houses and roads, while others are of mountainous areas dominated by trees. Therefore, the window size needs to be appropriately selected during image enhancement. A window that is too large will result in poor region allocation, while a window that is too small will lead to excessive data processing and significantly reduce computational efficiency. Therefore, this embodiment of the invention uses the following method to divide the regions of the collected images.

[0104] Optionally, the acquired image is divided into regions, including: dividing the acquired image according to a preset window size to obtain at least one region, and determining the average gray value and gray-level dispersion of each region; determining the region features of each region based on the average gray value, the gray-level dispersion, and a preset standard threshold; wherein the preset standard threshold is determined according to each region feature; if there are other regions in the acquired image whose region features have not been determined, the preset window size is adjusted, and the other regions are divided according to the adjusted preset window size, until all regions in the acquired image have their region features determined.

[0105] First, a variable window of size M×N can be selected based on the grayscale distribution characteristics of the acquired image. The window is then moved left and right with a step size M, and moved up and down with a step size N. The average grayscale value and grayscale dispersion of the region corresponding to each window are calculated.

[0106] Specifically, the average grayscale value can be determined using the following formula:

[0107]

[0108] Where f is the average gray value, and f(i,j) is the gray value at pixel (i,j) in the acquired image.

[0109] Specifically, the degree of grayscale dispersion can be determined using the following formula:

[0110]

[0111] Among them, T 2 This represents the degree of grayscale dispersion.

[0112] Secondly, the average gray value and gray dispersion of each region are compared with the preset standard threshold to determine the regional characteristics of each region.

[0113] For example, images acquired during power transmission line inspections are divided into farmland / road areas with significant grayscale changes and forest areas with relatively small grayscale changes. Labeled sample data are selected accordingly, and the average grayscale value and grayscale dispersion are calculated for each area using the formulas for calculating average grayscale value and grayscale dispersion. These values ​​are then used as preset standard thresholds. If the average grayscale value and grayscale dispersion of a certain area fall within the preset standard threshold for farmland / road areas, that area is identified as a farmland / road area; similarly, if the average grayscale value and grayscale dispersion of a certain area fall within the preset standard threshold for forest areas, that area is identified as a forest area.

[0114] Of course, dividing the acquired image into regions using only a window of size M×N may result in some regions having average gray values ​​and gray-level dispersion that do not fall within the preset standard threshold corresponding to any pre-determined regional feature. Therefore, in this embodiment of the invention, for the remaining regions in the acquired image whose regional features have not been determined, the preset window size can be reduced to m×n, and then the window can be translated horizontally with a step size m and vertically with a step size n. The average gray value and gray-level dispersion of the region corresponding to each window are calculated, and then the regions are divided according to the preset standard threshold. The above steps can be repeated until the regional features of all regions in the acquired image have been determined.

[0115] S330. Enhancement processing is performed on each region in the acquired image.

[0116] In this embodiment of the invention, different degrees of enhancement processing can be applied to different regions based on their characteristics. For example, the enhancement effect can be stronger for pixels in forest areas with small grayscale changes, and weaker for pixels in farmland and road areas with significant grayscale changes.

[0117] Specifically, a weighting index α and a smoothing parameter β can be set for the pixel that needs to be enhanced. α controls the influence of surrounding pixels of different gray levels on the gray value of the pixel to be enhanced; β is used to smooth the transition between other areas and this pixel.

[0118] Optionally, enhancement processing is performed on each region of the acquired image, including: determining the grayscale histogram of each region in the acquired image, and performing histogram equalization on each grayscale histogram to obtain the equalization transformation function of each grayscale histogram; traversing each pixel in the acquired image, and determining the adjustment weight of the at least one neighboring sampling point on the target pixel based on the distance between the target pixel and at least one neighboring sampling point during the traversal; and determining the enhanced grayscale value of the target pixel based on the equalization transformation function of the at least one neighboring sampling point and the adjustment weight.

[0119] For example, Figure 4This is a schematic diagram of an image enhancement method provided as an embodiment of the present invention. For example... Figure 4 As shown, m×n is a small region, and point (x,y) is the target pixel. The grayscale of the target pixel is enhanced by sampling points of the four neighboring image blocks.

[0120] First, the gray-level histograms of the four neighboring image blocks of the target pixel can be determined, and histogram equalization can be performed on each gray-level histogram to obtain the equalization transformation function for each gray-level histogram. Specifically, the gray levels of the acquired image are considered as random variables in the interval [0, 1], and this random variable can be expressed by its probability density function. Let the gray-level value of the target pixel be r (0≤r≤1), and the enhanced gray level after transformation be s, let P r (r) and P s Let (s) represent the probability density functions of random variables r and s, respectively, and let T(r) be the equilibrium transformation function. Then, the following equation holds:

[0121] Each pixel's grayscale value *r* in the original image will generate a grayscale value *s*0 after transformation. The transformation function *T(r)* should satisfy: 1) *T(r)* is a single value and monotonically increasing in the interval 0 ≤ *r* ≤ 1; 2) 0 ≤ *T(r)* ≤ 1. This ensures that the output grayscale level has the same range as the input grayscale level. Let:

[0122]

[0123] Then there is

[0124] Substituting the above equation into the probability density function, we get:

[0125]

[0126] When applied to digital image processing, if the digital image has L gray levels, it becomes as follows:

[0127]

[0128] In the formula: k represents the grayscale of the digital image; n represents the total number of pixels; n j P represents the number of pixels on the j-th grayscale layer; r (r j T(r) represents the probability density on the j-th gray level; k ) represents the equalization transformation function for the pixel on the k-th gray level; s k This is the final transformation result.

[0129] Secondly, the adjustment weights of the four neighboring sampling points on the target pixel can be determined based on the distance between the target pixel and its four neighboring sampling points. Specifically, the adjustment weights of the four neighboring sampling points on the target pixel can be determined using the following formula:

[0130]

[0131] Among them, W k (x, y) represents the adjustment weight of the k-th neighboring sampling point on the target pixel, d k (x, y) represents the distance from the k-th neighboring sampling point to the target pixel.

[0132] Finally, the enhanced grayscale value of the target pixel is determined based on the equalization transformation function of the four-neighborhood and the adjustment weights. Specifically, the enhanced grayscale value of the target pixel can be determined using the following formula:

[0133]

[0134] Where g(x, y) is the enhanced grayscale value of the target pixel, and T k (I(x, y)) is the equalization transformation function of the image block corresponding to the k-th neighboring sampling point, and I(x, y) is the original gray value of the k-th neighboring sampling point.

[0135] Furthermore, by performing the above steps on each region of the acquired image, the enhancement processing of the acquired image can be completed.

[0136] S340. Based on the feature extraction model, extract the surface salient features and edge salient features of the target part in the enhanced acquired image respectively.

[0137] S350. Determine the edge reinforcement salience feature of the target part based on the surface salience feature and the edge salience feature.

[0138] S360. Based on the edge enhancement saliency feature and the surface saliency feature, determine the surface fusion saliency feature of the target part.

[0139] S370. Determine the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features.

[0140] This invention provides an image processing method. The method involves acquiring an image of a region to be detected using an image acquisition device; dividing the acquired image into regions to obtain at least one region; performing enhancement processing on each region of the acquired image; extracting surface saliency features and edge saliency features of a target part from the acquired image based on a feature extraction model; determining edge enhancement saliency features of the target part based on the surface and edge saliency features; determining surface fusion saliency features of the target part based on the edge enhancement and surface saliency features; and determining the target features of the target part based on the edge enhancement and surface fusion saliency features. This technical solution, by performing region-based enhancement processing on the acquired image, not only makes subsequent image matching easier but also provides higher-quality image data for image matching and orthophoto production, especially showing a more significant effect in forested areas where UAV images have significant distortion and small grayscale variations.

[0141] Example 4

[0142] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in Embodiment 4 of this application. Figure 5 As shown, the device includes:

[0143] The part feature extraction module 410 is used to acquire images of the area to be detected through an image acquisition device, and extract the surface salient features and edge salient features of the target part in the acquired images based on the feature extraction model.

[0144] Edge feature enhancement module 420 is used to determine the edge enhancement salience features of the target part based on the surface salience features and the edge salience features;

[0145] The surface feature fusion module 430 is used to determine the surface fusion saliency features of the target part based on the edge enhancement saliency features and the surface saliency features;

[0146] The target feature determination module 440 is used to determine the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features.

[0147] This invention provides an image processing apparatus that acquires an image of a region to be detected using an image acquisition device, and extracts surface saliency features and edge saliency features of a target part from the acquired image based on a feature extraction model. Based on the surface saliency features and edge saliency features, the apparatus determines edge enhancement saliency features of the target part; based on the edge enhancement saliency features and surface saliency features, it determines surface fusion saliency features of the target part; and based on the edge enhancement saliency features and surface fusion saliency features, it determines the target features of the target part. This technical solution, by complementing the surface feature information and edge feature information of the target part, ensures the integrity of the target part features and improves the accuracy of target part segmentation.

[0148] Furthermore, the part feature extraction module 410 includes:

[0149] The side path feature extraction unit is used to extract features of the target part in the enhanced image based on the backbone network in the feature extraction model, to obtain at least one first side path feature and at least one second side path feature; wherein, the level of the first side path feature is lower than the level of the second side path feature.

[0150] An edge saliency feature extraction unit is used to extract the edge saliency features of the target part based on the at least one first edge path feature;

[0151] A surface saliency feature extraction unit is used to extract the surface saliency features of the target part based on the at least one second side path feature;

[0152] Furthermore, the edge feature enhancement module 420 includes:

[0153] An edge feature enhancement unit is used to determine the edge enhancement saliency features of the target part based on the first saliency feature extraction network in the feature extraction model, according to the edge saliency features and the surface saliency features extracted according to the second edge path features of the highest level.

[0154] Accordingly, the surface feature fusion module 430 includes:

[0155] A surface feature enhancement unit is used to determine the surface enhancement saliency features of the current level based on the surface saliency features extracted by the second saliency feature extraction network in the feature extraction model, according to the surface saliency features extracted by the second side path features of the current level and the surface saliency features extracted by the second side path features of the next level; wherein, the next level is higher than the current level;

[0156] The surface feature fusion unit is used to determine the surface fusion saliency features of the target part based on the guiding network in the feature extraction model, according to the edge enhancement saliency features and the surface enhancement saliency features at each level.

[0157] Furthermore, the device also includes:

[0158] The sub-loss value determination module is used to determine the binary cross-entropy loss value, structural similarity index loss value, and intersection-union ratio loss value of each output in the feature extraction model, based on the saliency supervision network in the feature extraction model after determining the target features of the target part according to the edge enhancement saliency features and the surface fusion saliency features; wherein, the output of the feature extraction model includes at least one of the edge saliency features, the surface saliency features, the edge enhancement saliency features, and the surface fusion saliency features;

[0159] The total loss value determination module is used to determine the total loss value of the feature extraction model based on the binary cross-entropy loss value, the structural similarity index loss value, and the intersection-union ratio loss value.

[0160] Furthermore, the device also includes:

[0161] An image region segmentation module is used to segment the acquired image of the region to be detected by an image acquisition device to obtain at least one region; wherein the segmentation criterion is the degree of pixel grayscale change.

[0162] The image region enhancement module is used to enhance each region in the acquired image separately.

[0163] Furthermore, the image region segmentation module includes:

[0164] The sub-region grayscale determination unit is used to divide the acquired image according to a preset window size to obtain at least one region, and to determine the average grayscale value and grayscale dispersion of each region.

[0165] The first image region segmentation unit is used to determine the regional features of each region based on the average gray value, the gray value dispersion, and a preset standard threshold; wherein the preset standard threshold is determined according to the features of each region.

[0166] The second image region segmentation unit is used to adjust the preset window size and divide the remaining regions according to the adjusted preset window size if there are other regions in the acquired image whose regional features have not been determined, until all regions in the acquired image have had their regional features determined.

[0167] Furthermore, the image region enhancement module includes:

[0168] The histogram equalization unit is used to determine the gray-level histogram of each region in the acquired image, and to perform histogram equalization on each gray-level histogram to obtain the equalization transformation function of each gray-level histogram.

[0169] The adjustment weight determination unit is used to traverse each pixel in the acquired image and, during the traversal, determine the adjustment weight of the at least one neighboring sampling point on the target pixel based on the distance between the target pixel and at least one neighboring sampling point.

[0170] An image region enhancement unit is used to determine the enhanced grayscale value of the target pixel based on the equalization transformation function of the at least one neighborhood and the adjustment weight.

[0171] The image processing apparatus provided in this application embodiment can execute an image processing method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method.

[0172] Example 5

[0173] Figure 6 A schematic diagram of the structure of a device 10 that can be used to implement embodiments of this application is shown. The device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0174] like Figure 6 As shown, device 10 includes at least one processor 11 and a memory, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc., communicatively connected to at least one processor 11. The memory stores computer programs executable by at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of device 10. The processor 11, ROM 12, and RAM 13 are interconnected via bus 14. Input / output (I / O) interface 15 is also connected to bus 14.

[0175] Multiple components in device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0176] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image processing methods.

[0177] In some embodiments, the image processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image processing method by any other suitable means (e.g., by means of firmware).

[0178] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0179] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0180] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including sound input, voice input, or haptic input).

[0182] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0183] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0184] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.

[0185] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. An image processing method, characterized in that, The method includes: The image of the area to be detected is acquired by an image acquisition device, and the surface salience features and edge salience features of the target part in the acquired image are extracted based on the feature extraction model. Based on the surface saliency features and the edge saliency features, determine the edge reinforcement saliency features of the target part; Based on the edge enhancement saliency features and the surface saliency features, the surface fusion saliency features of the target part are determined; Based on the edge enhancement saliency feature and the surface fusion saliency feature, the target features of the target part are determined; The feature extraction model includes: The network consists of a backbone network, a first salient feature extraction network, a second salient feature extraction network, a guiding network, and a salient supervision network. The backbone network includes: Convolutional layers are used to generate salient edge features; Max pooling layers are used to reduce model parameters; The first improved residual block is used to generate the first surface salient feature; The second improved residual block is used to generate the second surface salient feature; The third improved residual block is used to generate the third surface salient feature; The first saliency feature extraction network is used to first perform bilinear interpolation and convolution on the third surface saliency features, then add them pixel by pixel to the edge saliency features, and finally perform convolution to obtain the edge enhanced saliency features. The second saliency feature extraction network is used to perform convolution processing on the third surface saliency features to obtain the third surface enhanced saliency features; the third surface saliency features are first subjected to bilinear interpolation and convolution processing, then added to the second surface saliency features pixel by pixel, and finally subjected to convolution processing to obtain the second surface enhanced saliency features; the second surface enhanced saliency features are first subjected to bilinear interpolation and convolution processing, then added to the first surface saliency features pixel by pixel, and finally subjected to convolution processing to obtain the first surface enhanced saliency features; The guiding network is used to add the edge enhancement saliency features pixel-by-pixel to the first surface enhancement saliency features, the second surface enhancement saliency features, and the third surface enhancement saliency features that have undergone bilinear interpolation and convolution processing, and then perform convolution processing on each of them to obtain the first surface fusion saliency features, the second surface fusion saliency features, and the third surface fusion saliency features; finally, the first surface fusion saliency features, the second surface fusion saliency features, and the third surface fusion saliency features are added pixel-by-pixel to obtain the total surface fusion saliency features; The saliency supervision network is used to calculate the loss value for the convolution-processed edge-enhanced saliency features, first surface-enhanced saliency features, second surface-enhanced saliency features, third surface-enhanced saliency features, first surface fusion saliency features, second surface fusion saliency features, third surface fusion saliency features, and total surface fusion saliency features, respectively.

2. The method according to claim 1, characterized in that, Based on the feature extraction model, the surface salient features and edge salient features of the target part in the acquired image are extracted respectively, including: Based on the backbone network in the feature extraction model, the side path features of the target part in the acquired image are extracted to obtain at least one first side path feature and at least one second side path feature; wherein, the level of the first side path feature is lower than the level of the second side path feature. Based on the at least one first edge path feature, extract the edge saliency features of the target part; Based on the at least one second side path feature, extract the surface salient features of the target part.

3. The method according to claim 2, characterized in that, Based on the surface saliency features and the edge saliency features, the edge reinforcement saliency features of the target part are determined, including: Based on the first saliency feature extraction network in the feature extraction model, and according to the surface saliency features extracted by the edge saliency features and the second edge path features of the highest level, the edge enhancement saliency features of the target part are determined. Accordingly, based on the edge enhancement saliency feature and the surface saliency feature, the surface fusion saliency feature of the target part is determined, including: Based on the second saliency feature extraction network in the feature extraction model, the surface saliency features of the current level are determined according to the surface saliency features extracted by the second side path features of the current level and the surface saliency features extracted by the second side path features of the next level; wherein, the next level is higher than the current level; Based on the guided network in the feature extraction model, the surface fusion saliency features of the target part are determined according to the edge enhancement saliency features and the surface enhancement saliency features at each level.

4. The method according to claim 1, characterized in that, After determining the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features, the method further includes: Based on the saliency supervision network in the feature extraction model, determine the binary cross-entropy loss value, structural similarity index loss value, and cross-union ratio loss value of each output in the feature extraction model; wherein, the output of the feature extraction model includes at least one of the edge saliency feature, the surface saliency feature, the edge enhancement saliency feature, and the surface fusion saliency feature; The total loss value of the feature extraction model is determined based on the binary cross-entropy loss value, the structural similarity index loss value, and the cross-union ratio loss value.

5. The method according to claim 1, characterized in that, After acquiring an image of the area to be detected using an image acquisition device, the method further includes: The acquired image is divided into regions to obtain at least one region; wherein the division criterion is the degree of pixel grayscale change; Enhancement processing is performed on each region in the acquired image.

6. The method according to claim 5, characterized in that, The acquired image is divided into regions, including: The acquired image is divided into at least one region according to a preset window size, and the average gray value and gray dispersion of each region are determined. The regional characteristics of each region are determined based on the average gray value, the gray level dispersion, and a preset standard threshold; wherein the preset standard threshold is determined separately for each regional characteristic. If there are other regions in the acquired image whose regional features have not been determined, the preset window size is adjusted, and the other regions are divided according to the adjusted preset window size until all regions in the acquired image have their regional features determined.

7. The method according to claim 5, characterized in that, Enhancement processing is performed on each region of the acquired image, including: Determine the grayscale histogram of each region in the acquired image, and perform histogram equalization on each grayscale histogram to obtain the equalization transformation function of each grayscale histogram; Traverse each pixel in the acquired image, and during the traversal, determine the adjustment weight of the at least one neighboring sampling point on the target pixel based on the distance between the target pixel and at least one neighboring sampling point. The enhanced grayscale value of the target pixel is determined based on the equalization transformation function of the at least one neighborhood and the adjustment weight.

8. An image processing apparatus, characterized in that, The device includes: The part feature extraction module is used to acquire images of the area to be detected through an image acquisition device, and extract the surface salient features and edge salient features of the target part in the acquired images based on the feature extraction model. An edge feature enhancement module is used to determine the edge enhancement saliency features of the target part based on the surface saliency features and the edge saliency features; A surface feature fusion module is used to determine the surface fusion saliency features of the target part based on the edge enhancement saliency features and the surface saliency features; The target feature determination module is used to determine the target features of the target part based on the edge enhancement saliency features and the surface fusion saliency features; The feature extraction model includes: The network consists of a backbone network, a first salient feature extraction network, a second salient feature extraction network, a guiding network, and a salient supervision network. The backbone network includes: Convolutional layers are used to generate salient edge features; Max pooling layers are used to reduce model parameters; The first improved residual block is used to generate the first surface salient feature; The second improved residual block is used to generate the second surface salient feature; The third improved residual block is used to generate the third surface salient feature; The first saliency feature extraction network is used to first perform bilinear interpolation and convolution on the third surface saliency features, then add them pixel by pixel to the edge saliency features, and finally perform convolution to obtain the edge enhanced saliency features. The second saliency feature extraction network is used to perform convolution processing on the third surface saliency features to obtain the third surface enhanced saliency features; the third surface saliency features are first subjected to bilinear interpolation and convolution processing, then added to the second surface saliency features pixel by pixel, and finally subjected to convolution processing to obtain the second surface enhanced saliency features; the second surface enhanced saliency features are first subjected to bilinear interpolation and convolution processing, then added to the first surface saliency features pixel by pixel, and finally subjected to convolution processing to obtain the first surface enhanced saliency features; The guiding network is used to add the edge enhancement saliency features pixel-by-pixel to the first surface enhancement saliency features, the second surface enhancement saliency features, and the third surface enhancement saliency features that have undergone bilinear interpolation and convolution processing, and then perform convolution processing on each of them to obtain the first surface fusion saliency features, the second surface fusion saliency features, and the third surface fusion saliency features; finally, the first surface fusion saliency features, the second surface fusion saliency features, and the third surface fusion saliency features are added pixel-by-pixel to obtain the total surface fusion saliency features; The saliency supervision network is used to calculate the loss value for the convolution-processed edge-enhanced saliency features, first surface-enhanced saliency features, second surface-enhanced saliency features, third surface-enhanced saliency features, first surface fusion saliency features, second surface fusion saliency features, third surface fusion saliency features, and total surface fusion saliency features, respectively.

9. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method and device, model training method and device, equipment and storage medium

    CN115908982A

  • Salient target detection method based on edge feature guidance

    CN116229104A