Image detection method and device, nonvolatile storage medium and electronic equipment

By removing the hair structure in the skin image and utilizing the global information capture method of the image detection model, the problem of accuracy in identifying lesion areas in skin images is solved, and the accuracy and reliability of skin lesion detection are improved.

CN120689586APending Publication Date: 2025-09-23CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573579.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies lack the ability to effectively capture global information in skin images, resulting in the inability to accurately identify lesion areas.

Method used

By determining the curvature information of the initial image to remove the hair structure, the encoder branch, decoder branch and pyramid pooling module in the image detection model are used, combined with the attention mechanism, to capture global information and reduce interference, thereby improving the accuracy of lesion area recognition.

Benefits of technology

It achieves accurate identification of lesion areas in skin images, improves segmentation accuracy and robustness, and provides a more reliable diagnostic tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689586A_ABST
    Figure CN120689586A_ABST
Patent Text Reader

Abstract

The invention discloses an image detection method and device, a nonvolatile storage medium and electronic equipment. The method comprises the steps that curvature information of all areas in an initial image is determined, a hair structure in the initial image is removed according to the curvature information, and a target image is obtained, and the initial image is a skin image; an abnormal area in a target image is determined through an image detection model, the image detection model comprises an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, a plurality of up-sampling units in the encoder branch are connected with the pyramid pooling module, and a plurality of up-sampling units in the decoder branch are connected with the pyramid pooling module. The abnormal region is a region where the skin may have lesions; and determining skin lesion conditions according to the abnormal region. The technical problem that the lesion area in the skin image cannot be accurately recognized due to the lack of effective capture of global information in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to an image detection method, device, non-volatile storage medium, and electronic device. Background Art

[0002] When determining the lesion area in the skin image, the related technology focuses too much on local features and lacks effective capture of global information, resulting in the inability of the related technology to accurately identify the possible lesion area in the skin image.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide an image detection method, device, non-volatile storage medium and electronic device to at least solve the technical problem of being unable to accurately identify the diseased area in the skin image due to the lack of effective capture of global information in the related art.

[0005] According to one aspect of an embodiment of the present application, an image detection method is provided, including: determining curvature information of each area in an initial image, and removing hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; determining abnormal areas in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal areas are areas where skin lesions may exist; and determining the skin lesion condition based on the abnormal areas.

[0006] Optionally, determining the curvature information of each region in the initial image and removing the hair structure in the initial image based on the curvature information includes: determining the Hessian matrix of the initial image, wherein the Hessian matrix is ​​used to describe the curvature information of each region in the initial image; determining the eigenvalues ​​of the Hessian matrix, wherein the eigenvalues ​​are used to reflect the principal curvature of each region; and removing the hair structure in the initial image based on the eigenvalues.

[0007] Optionally, the encoder branch includes multiple downsampling units, wherein each downsampling unit includes a parallel multi-branch module, the parallel multi-branch module includes a parallel first branch and a second branch, the first branch includes multiple parallel convolution layers, and the second branch also includes multiple parallel convolution layers; the convolution kernel sizes corresponding to the multiple parallel convolution layers in the first branch are different from each other, and the convolution kernel sizes corresponding to the multiple parallel convolution layers in the second branch are different from each other.

[0008] Optionally, the resolutions of the feature maps output by each upsampling unit in the decoder branch are different from each other, wherein the pyramid pooling module is used to fuse the feature maps output by each upsampling unit to determine the abnormal area.

[0009] Optionally, the attention mechanism is a hybrid spatial channel attention mechanism that includes a channel attention mechanism and a spatial attention mechanism, wherein the channel attention mechanism includes a small-order statistics module and a large-order statistics module, the small-order statistics module is used to determine the local detail features of the channel, and the large-order statistics module is used to determine the global dependency between channels.

[0010] Optionally, the small-order statistics module includes a global average pooling layer and a first fully connected layer, wherein the global average pooling layer is used to extract small-order features of the image input into the small-order statistics module, and the fully connected layer is used to process the small-order features to obtain small-order attention weights corresponding to the small-order statistics module, and the small-order features include statistical features of the image; the large-order statistics module includes a bilinear pooling layer and a second fully connected layer, wherein the bilinear pooling layer is used to perform pairwise outer product processing on the feature vectors of each channel in the image input into the large-order statistics module to obtain a feature matrix, and the second fully connected layer is used to process the feature matrix to obtain the large-order attention weights corresponding to the large-order statistics module.

[0011] Optionally, the method also includes: determining the small-order attention weight determined by the small-order statistics module; determining the large-order attention weight determined by the large-order statistics module; and performing element-by-element addition processing on the small-order attention weight and the large-order attention weight to obtain the channel attention weight.

[0012] According to another aspect of an embodiment of the present application, an image detection device is also provided, including: a first processing module, used to determine the curvature information of each area in an initial image, and remove the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; a second processing module, used to determine the abnormal area in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; a third processing module, used to determine the skin lesion area based on the abnormal area.

[0013] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is executed, a device where the non-volatile storage medium is located is controlled to execute an image detection method.

[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the image detection method is executed when the program is run.

[0015] According to another aspect of an embodiment of the present application, a computer program product is further provided, including a computer program, which implements the image detection method when executed by a processor.

[0016] In an embodiment of the present application, curvature information of each area in an initial image is determined, and the hair structure in the initial image is removed based on the curvature information to obtain a target image, wherein the initial image is a skin image; abnormal areas in the target image are determined by an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected by an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal areas are areas where the skin may have lesions; the skin lesion condition is determined based on the abnormal area, and interference is eliminated by removing the hair structure based on the curvature information, and a pyramid pooling module to which one or more upsampling units are respectively connected is added to the model, thereby achieving the purpose of accurately capturing the global information of the image and reducing interference information, thereby achieving the technical effect of improving the recognition accuracy of the lesion area, and thus solving the technical problem of being unable to accurately identify the lesion area in the skin image due to the lack of effective capture of global information in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 1 is a schematic diagram of the structure of a computer terminal (mobile terminal) provided according to an embodiment of the present application;

[0019] Figure 2 1 is a flow chart of an image detection method provided according to an embodiment of the present application;

[0020] Figure 3 This is a structural diagram of a parallel multi-branch module provided according to an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of a feature map processing flow in an image detection model provided according to an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of a segmentation result provided according to an embodiment of the present application;

[0023] Figure 6 It is a structural schematic diagram of an image detection device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] In the field of medical image analysis, the automatic segmentation and detection of skin lesions is a crucial and challenging task. Early detection and diagnosis of skin lesions, especially malignant melanoma, is crucial for patient treatment outcomes. Skin lesions typically appear as irregular spots or protrusions on the skin, and their morphology, size, color, and other characteristics play an important role in their classification and diagnosis. Traditional skin lesion segmentation methods rely heavily on manual feature extraction and traditional image processing techniques, such as thresholding and edge detection. However, these methods are often limited by noise, illumination variations, and morphological irregularities when processing complex skin lesion images, resulting in low segmentation accuracy. In recent years, with the development of deep learning technology, particularly the widespread application of convolutional neural networks (CNNs) in image classification and segmentation tasks, research in medical image segmentation has made significant progress. The U-Net model, in particular, has become a widely used classic model due to its outstanding performance in medical image segmentation. Through its encoder-decoder architecture, the U-Net model effectively extracts multi-layered features from images and gradually restores their spatial resolution, demonstrating exceptionally high accuracy in medical image segmentation tasks. However, a limitation of the traditional U-Net model is that its segmentation capabilities rely solely on local features and lack effective capture of global information. Specifically, U-Net primarily extracts local features through convolutional layers and combines information from different layers through skip connections. However, this approach has limited ability to capture complex image backgrounds and details of lesions. Furthermore, U-Net does not effectively utilize spatial and channel correlation information, which can lead to performance bottlenecks when processing images with complex backgrounds.

[0027] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0028] According to an embodiment of the present application, a method embodiment of an image detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0029] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image detection method. Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0030] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0031] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image detection method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned image detection method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0032] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0033] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0034] In the above operating environment, the embodiment of the present application provides an image detection method, such as Figure 2 As shown, the method includes the following steps:

[0035] Step S202, determining curvature information of each region in the initial image, and removing the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image;

[0036] In the technical solution provided in step S202, the steps of determining curvature information of each region in the initial image and removing the hair structure in the initial image based on the curvature information include: determining the Hessian matrix of the initial image, wherein the Hessian matrix is ​​used to describe the curvature information of each region in the initial image; determining the eigenvalues ​​of the Hessian matrix, wherein the eigenvalues ​​are used to reflect the principal curvature of each region; and removing the hair structure in the initial image based on the eigenvalues.

[0037] In skin lesion image analysis, hair is a common interfering factor that can affect accurate segmentation and subsequent analysis of the lesion area. Therefore, hair removal is a key step in skin lesion image preprocessing. First, to analyze the local shape of the image, the image is smoothed using a Gaussian kernel function. The Hessian matrix describes the curvature information of the local image region. Next, by calculating the eigenvalues ​​of the Hessian matrix, the principal curvature of the local image region can be obtained. By analyzing this curvature information, hair structure in the image can be more accurately detected, not just the hair boundaries, thereby achieving the goal of hair removal.

[0038] It should be noted that in the analysis of skin lesion images, hair is an extremely difficult and common interference factor, posing a significant challenge to the accurate segmentation of the lesion area and subsequent in-depth analysis. The random distribution of hair, its complex texture, and its visual interweaving with skin lesions often lead to a significant reduction in the accuracy of the segmentation algorithm, thereby affecting the doctor's judgment of the nature of the lesion and the accurate assessment of the course of the disease. Therefore, hair removal is not only a key step in the preprocessing stage, but also a necessary link to improve the accuracy and robustness of skin lesion image analysis. In order to effectively identify and remove these interferences, a hair recognition method based on local curvature analysis is adopted. By deeply understanding the geometric characteristics of the image, the hair structure is accurately located, thereby ensuring the accuracy of skin lesion segmentation.

[0039] First, to analyze the local shape of the image and reduce the effects of noise, the original image is smoothed using a Gaussian kernel function, a fundamental step in multiscale image processing. The Gaussian kernel, weighted by its spatial position within the image, performs a weighted average of pixels in a local region of the image, effectively smoothing the image and reducing the effects of noise while preserving the key shape information of the lesion. This smoothing process, by adjusting the standard deviation parameter of the Gaussian kernel, allows for controlled smoothing intensity, ensuring a clear boundary between hair and other skin structures without excessively blurring the details of the lesion.

[0040] Next, by calculating the Hessian matrix of the smoothed image, the curvature information of the local area of ​​the image is further extracted. The Hessian matrix is ​​a second-order derivative matrix that can describe the local curvature of each pixel in the image, including the primary curvature and secondary curvature. In skin lesion images, the shape of hair usually appears as a linear structure, and its curvature characteristics are significantly different from the curvature characteristics of the surrounding skin tissue or lesion area. This characteristic provides a theoretical basis for hair recognition. By calculating the Hessian matrix of each pixel, two eigenvalues ​​of the point can be obtained, corresponding to the primary curvature and secondary curvature, which can reflect the shape characteristics of the local area of ​​the image. Specifically, high curvature values ​​correspond to sharp turning points of lines or edges in the image, while low curvature values ​​correspond to relatively flat areas. The hair structure, due to its slenderness and significant curvature changes, is usually reflected in the eigenvalues ​​of the Hessian matrix.

[0041] To identify hair structure from the Hessian matrix, we focus on the distribution of its eigenvalues. The distribution of hair eigenvalues ​​typically exhibits strong linearity and contrast. In a Gaussian-smoothed image, the eigenvalues ​​of hair regions exhibit high local extremes, corresponding to the principal curvatures of the hair. Based on this observation, hair regions can be distinguished from other areas by setting an appropriate threshold. This process involves in-depth analysis of the eigenvalue distribution to identify regions with high curvature variations, which often coincide with hair structure. By analyzing the eigenvalues ​​of the Hessian matrix for each pixel, we can construct an information map of the local curvature of the image. This map allows for more accurate detection of hair structure in the image, not just its boundaries. This provides precise positioning information for subsequent, accurate hair removal.

[0042] Hair removal is an iterative and optimized process. A hair structure recognition algorithm was designed based on the distribution characteristics of the Hessian matrix eigenvalues. This algorithm identifies potential hair regions by analyzing the distribution of eigenvalues. It then applies a series of image processing techniques, including morphological operations like dilation and erosion, as well as connected component analysis, to further refine and confirm whether these regions are truly hair, as well as their boundaries and shapes. After identifying hair regions, curvature-based edge recognition techniques are used, combined with hair's geometric characteristics, such as slenderness and continuity, to precisely locate the hair contours. This process fully utilizes the eigenvalue information of the Hessian matrix, ensuring accurate hair contour recognition and avoiding over-segmentation or mis-segmentation.

[0043] After confirming the specific location of the hair, background replacement techniques used in image processing are used to replace the hair area with the average value of the surrounding skin or skin texture predicted through deep learning, thereby visually eliminating the interference of the hair. This replacement process is based on the precise recognition of hair structure, ensuring the complete preservation of skin lesion information in the original image while removing the influence of hair on image analysis. To further improve the naturalness and reliability of the replacement, deep learning models, such as variants of the U-Net network, are also introduced to predict more accurate textures of skin lesions in hairy areas. This makes the images after hair removal more realistic and less likely to affect the doctor's judgment of the lesion area.

[0044] The key to using the eigenvalues ​​of the Hessian matrix for hair structure identification lies in eigenvalue analysis and threshold setting. The magnitude of the eigenvalue reflects the strength of the local curvature. High curvature values ​​correspond to turning points at boundaries or lines, which are particularly pronounced in hair structures. By statistically analyzing the distribution of the eigenvalues, a reasonable threshold can be set to identify areas with significant and continuous curvature changes as hair. This threshold setting must take into account the average width of the hair, the curvature range of the skin lesion, and the overall contrast of the image to ensure that hair removal occurs without accidentally damaging the lesion or important skin structures.

[0045] After the threshold is set, a series of mathematical morphological operations, such as dilation and erosion, are performed to further refine and enhance the connected components of the hair, ensuring clear and continuous boundaries. Dilation fills any small gaps within the hair, while erosion removes noise along the boundaries, resulting in a purer hair structure. This process is repeated until the boundaries of the hair structure are optimized, neither too wide nor too narrow, to ensure segmentation accuracy and robustness.

[0046] Furthermore, connected component analysis can be used to identify continuous high-curvature regions as a single hair structure, avoiding the problem of segmenting a single hair into multiple parts and ensuring the consistency and integrity of hair recognition. This curvature-based connected component analysis technique not only identifies the main trunk of the hair but also captures the details of the hair ends and bifurcations, thereby more comprehensively removing the impact of hair on skin lesion segmentation.

[0047] Finally, the aforementioned analysis based on the eigenvalues ​​of the Hessian matrix, combined with mathematical morphological operations and connected component analysis, accurately identifies and locates hair structures in images, effectively removing hair before skin lesion segmentation, significantly improving the accuracy and reliability of subsequent segmentation. This method is not only applicable to images of conventional skin lesions, but also to images with densely distributed and complex hair structures, demonstrating its wide applicability and high robustness in practical applications.

[0048] In skin lesion image analysis, a hair recognition method based on local curvature analysis was employed. Through Gaussian smoothing, eigenvalue calculation and analysis of the Hessian matrix, threshold setting, mathematical morphological operations, and connected component analysis, a systematic hair removal framework was constructed. This framework not only accurately identifies and locates hair structures but also effectively eliminates the interference of hair on skin lesion segmentation, ensuring the accuracy and interpretability of segmentation results. Hair removal more clearly reveals the shape, size, and boundaries of the lesion area, providing doctors with more reliable and intuitive image information for diagnosis, thus playing a vital role in diagnosing and treating skin diseases.

[0049] In summary, hair is a confounding factor in skin lesion image analysis that requires careful attention. Through analysis and processing based on the eigenvalues ​​of the Hessian matrix, hair can be effectively identified and removed, ensuring the accuracy and reliability of skin lesion segmentation. This method not only advances the state of the art in skin lesion image analysis but also provides doctors with a more precise diagnostic tool, contributing to the early diagnosis and treatment of skin diseases.

[0050] Alternatively, the skin lesion image can be represented as a function u(x), where x = (x1, x2)∈R 2 Represents the pixel position in the image. 2 Represents a two-dimensional Euclidean space, that is, the set of all possible pixel positions. In order to analyze the local shape of the image, the image is smoothed using a Gaussian kernel function to obtain a smoothed image ω. The smoothed image ω can be expressed as:

[0051] ω=G σ (y)*u

[0052] Among them, G σ (y) is the Gaussian kernel function, and * represents a two-dimensional convolution operation.

[0053] The Gaussian kernel function is defined as:

[0054]

[0055] Where y = (y1, y2) is the pixel position in the image, and σ is the standard deviation parameter that controls the smoothness of the Gaussian kernel.

[0056] Calculate the Hessian matrix H of the smoothed image w (x,σ) The Hessian matrix is ​​a second-order derivative matrix defined as:

[0057]

[0058] In the above matrix, is the second-order partial derivative of the image ω in the x1 direction, reflecting the change in the curvature of the image in the horizontal direction. is the second-order partial derivative of the image ω in the x2 direction, reflecting the change in the curvature of the image in the vertical direction. is the mixed second-order partial derivative of the image ω with respect to x1 and x2, describing the correlation between the curvatures in the two directions.

[0059] Step S204: determining abnormal areas in the target image using an image detection model, wherein the image detection model includes an encoder branch, a decoder branch, and a pyramid pooling module. The encoder branch and the decoder branch are connected via an attention mechanism. Multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module. The abnormal areas are areas where skin lesions may be present.

[0060] In the technical solution provided in step S204, the encoder branch includes multiple downsampling units, wherein each downsampling unit includes a parallel multi-branch module, the parallel multi-branch module includes a parallel first branch and a second branch, the first branch includes multiple parallel convolution layers, and the second branch also includes multiple parallel convolution layers; the convolution kernel sizes corresponding to the multiple parallel convolution layers in the first branch are different from each other, and the convolution kernel sizes corresponding to the multiple parallel convolution layers in the second branch are different from each other.

[0061] In some embodiments of the present application, by using convolution kernels of different sizes, the receptive field range of image feature extraction can be expanded, thereby significantly enhancing the ability of the convolution layer to extract image features. The structure of the parallel multi-branch module (PMBS, Parallel Multi-Branch Structure) is as follows: Figure 3 As shown, it includes a first branch PMBS1 and a second branch PMBS2, wherein, Figure 3 The first branch is on the left. Figure 3 The second branch is on the right. PMBS1 consists of three parallel convolutional layers: Conv3, Conv5, and Conv7, while PMBS2 consists of three parallel convolutional layers: Conv1, Conv3, and Conv5. This design enables the network to simultaneously capture local and global features at different scales, thereby improving the diversity of feature representation. As an optional implementation, PMBS1 is also connected to a dense block, and PMBS2 is also connected to an upsampling module: UPSAMPLE.

[0062] It should be noted that the number of parallel convolutional layers in the first branch and the second branch does not have to be three, which is only shown here for illustration.

[0063] As an optional implementation, the resolutions of the feature maps output by each upsampling unit in the decoder branch are different from each other, wherein the pyramid pooling module is used to fuse the feature maps output by each upsampling unit to determine the abnormal area.

[0064] like Figure 4As shown in the figure, the Pyramid Pooling Block (PPB) effectively combines shallow and deep features by fusing image features at four different scales. The PPB architecture consists of four different layers of outputs: the top 1×1 layer generates a global-scale feature map through global pooling, while the remaining three layers divide the feature map into multiple subregions and generate independent feature representations within each subregion. Since the output feature maps of different layers have different channel dimensions, a 1×1 convolutional layer is added after each layer to reduce the dimension to ensure balanced weighting of features at different scales in the global feature representation. Specifically, if the PPB contains N layers, the number of output channels of each layer is compressed to 1 / N of the original dimension. Subsequently, the resolution of the feature map is restored to the original size of the input image through bilinear interpolation. Finally, the initial input feature map is concatenated with the multi-scale fused feature map processed by the PPB to form the final global feature representation. This design not only effectively extracts multi-scale features but also enhances the network's ability to capture image details and global structure through feature fusion, thereby improving the model's performance in complex tasks.

[0065] In addition, from Figure 4 It can be seen that the image feature processing process in this application is compared with the image feature processing process of the U-NET model with jump connections in the related art, and an additional step of fusing the output features of each upsampling unit of the decoding branch is added, so as to better utilize the feature information in the feature maps of different resolutions and capture global image features. The image detection model provided in the embodiment of the present application is also based on the U-NET model with jump connections, and adds multi-size parallel branches in each downsampling unit, and adds a pyramid pooling module connected to one or more upsampling units respectively, and adds a hybrid spatial channel attention mechanism (HSCA, Hybrid Channel Attention Mechanism) between the upsampling unit and the downsampling unit, wherein the channel attention mechanism includes a small-order statistics module and a large-order statistics module. In addition, from Figure 4 It can be seen that when the decoder branch of the present application performs upsampling and feature fusion, the fused feature maps will be processed by HSCA.

[0066] In some embodiments of the present application, the attention mechanism is a hybrid spatial channel attention mechanism that includes a channel attention mechanism and a spatial attention mechanism, wherein the channel attention mechanism includes a small-order statistics module and a large-order statistics module, the small-order statistics module is used to determine the local detail features of the channel, and the large-order statistics module is used to determine the global dependency between channels.

[0067] As an optional implementation, the following method can be used to determine the channel attention weight: determine the small-order attention weight determined by the small-order statistics module; determine the large-order attention weight determined by the large-order statistics module; and perform element-by-element addition of the small-order attention weight and the large-order attention weight to obtain the channel attention weight.

[0068] Alternatively, in related segmentation methods, the receptive field is limited to a local area, which limits the ability to obtain extensive and rich contextual information. Therefore, the embodiment of the present application introduces a channel and spatial attention mechanism. The following figure shows the structure of the channel attention module. This module can directly obtain the actual feature A∈R C×H×W Calculate the channel attention map X∈R C×C The original features are reshaped into R C×N , and then perform a product operation between A and its transpose A′. Where C is the number of channels, H is the image height, and W is the image width. N is the number of features after flattening. Finally, a SoftMax layer is used to generate the channel attention map X∈R C×C , the formula is as follows:

[0069]

[0070] Among them, x ij represents the influence of the i-th channel on the j-th channel. Next, a multiplication operation is performed between the transpose X′ of X and A, and the result is reshaped into R C×H×W Finally, the result is multiplied by the scale parameter β and added element-wise to A to get the final output:

[0071]

[0072] Where β represents a learnable weight. It consists of a lightweight compression and excitation block and uses layer normalization before the activation function ReLU. This layer normalization block improves the performance of object detection and segmentation. The module can be expressed as:

[0073]

[0074] Among them, A is the feature map (also the input), Np represents the number of positions in the feature map, A and Z represent the input and output respectively, represents the weight of global attention, δ(·)=W v2 ReLU(LN(W v1 (·))) represents the bottleneck transformation. Finally, the final feature map is obtained by element-by-element addition, which is the final channel-space mixed attention weight.

[0075] As an optional implementation, the small-order statistics module includes a global average pooling layer and a first fully connected layer, wherein the global average pooling layer is used to extract the small-order features of the image input into the small-order statistics module, and the fully connected layer is used to process the small-order features to obtain the small-order attention weights corresponding to the small-order statistics module, and the small-order features include the statistical features of the image; the large-order statistics module includes a bilinear pooling layer and a second fully connected layer, wherein the bilinear pooling layer is used to perform pairwise outer product processing on the feature vectors of each channel in the image input into the large-order statistics module to obtain a feature matrix, and the second fully connected layer is used to process the feature matrix to obtain the large-order attention weights corresponding to the large-order statistics module.

[0076] In some embodiments of the present application, small-order statistics and large-order statistics can be used in the channel attention module to calculate the correlation between attributes to obtain a power representation of the attributes. It comprises two modules: a small-order statistics module (SOSCA, Small Order Statistics Channel Attention) and a large-order statistics module (LOSCA, Large Order Statistics Channel Attention). The SOSCA module uses a global average pooling method to extract small-order features, and then processes these features through a small fully connected layer. The results are further rescaled to obtain intermediate features. On the other hand, the LOSCA attention mechanism uses bilinear pooling to obtain intermediate features of different channels and uses these features as input to the fully connected network. Finally, the final output of the hybrid model, that is, the channel attention weight, is obtained by adding the results of LOSCA and SOSCA element by element. This helps to explore the correlation between features.

[0077] The small-order statistics module uses global average pooling to compress the spatial dimension of the gated attribute. c ,…,X C ]∈R H×W×C is an attribute set, and its corresponding gate feature is y=[Y1,…,Y c ,…,Y C ]∈R H×W×C The global average pooling operation can be expressed as:

[0078]

[0079] These features are then fed into the fully connected layer. Let θ1(W1,b1) be the first fully connected layer FC1, and θ2(W1,b1) be the second fully connected layer FC2. The resulting channel attention weight f can be expressed as:

[0080]

[0081] Here, δ represents the ReLU activation function and σ represents the Sigmoid activation function. The ReLU activation function enhances the nonlinearity of the network, while the Sigmoid function scales the input to the range (0, 1) to obtain the weights. This design helps to establish correlations between different channels, thereby learning intermediate features. Using the attention weights, the final feature can be expressed as:

[0082]

[0083] Large-order statistics module: The above attention module calculates the mean of the gated attributes, which are used to rescale the input attributes. However, it fails to establish the correlation between the attributes in the middle of the channel. Therefore, the large-order statistics calculation is introduced in the embodiment of the present application. In this module, the input is also X = [X1, ..., X c ,…,X C ]∈R H×W×C is an attribute set, and its corresponding gate feature is y=[Y1,…,Y c ,…,Y C ]∈R H×W×C Bilinear pooling is applied to the gated features, where B represents the pooling matrix, which can be expressed as:

[0084]

[0085] Where vec represents a vectorized operation. The resulting pooling matrix is ​​semisymmetric and positive definite, containing the inner product of two different feature channels. This inner product helps introduce high-order features. Furthermore, the matrix logarithmic normalization model is applied to normalize the SPD matrix data. This matrix can be expressed in eigenvalue decomposition form:

[0086] B=UεU T

[0087] Among them, U represents the eigenvector, ε=diag(σ1,…,σ c ,…,σ C ) represents a diagonal matrix. By normalizing the matrix Get the channel attribute d∈R C , which is calculated as follows:

[0088]

[0089] The obtained d is processed by the fully connected network θ3(W3,b3) and θ4(W4,b4) to generate the attention weight s∈R C , and its learning method is:

[0090]

[0091] Step S206: determining the skin lesion condition based on the abnormal area.

[0092] In some embodiments of the present application, binary cross entropy (BCE) can be used as a loss function to optimize the segmentation results so that they are closer to the actual lesion area boundary, thereby improving the accuracy and usability of skin lesion detection. Figure 5 shown.

[0093] The target image is obtained by determining the curvature information of each area in the initial image and removing the hair structure in the initial image based on the curvature information, wherein the initial image is a skin image; the abnormal area in the target image is determined by an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected by an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; the skin lesion condition is determined based on the abnormal area, and interference is eliminated by removing the hair structure based on the curvature information, and a pyramid pooling module connected to one or more upsampling units is added to the model, so as to achieve the purpose of accurately capturing the global information of the image and reducing interference information, thereby achieving the technical effect of improving the recognition accuracy of the lesion area, and thus solving the technical problem of not being able to accurately identify the lesion area in the skin image due to the lack of effective capture of global information in the related technology.

[0094] The image detection method provided in the embodiment of the present application has at least the following beneficial effects compared to the related art:

[0095] First: Use the Gaussian kernel function to smooth the image, calculate the eigenvalues ​​of the Hessian matrix, analyze the local curvature of the image, and accurately detect the hair structure to remove interfering hair and ensure the accuracy of the segmentation results.

[0096] Second: By adopting a network structure with multi-scale feature fusion and introducing a parallel multi-branch structure and pyramid pooling module, the convolutional neural network's ability to extract local and global features is effectively improved, thereby enhancing the network's image analysis performance.

[0097] Third, a hybrid spatial channel attention mechanism is introduced, combining small-order and large-order statistical attention modules. Through operations such as global average pooling and bilinear pooling, the weights of different channels and spatial regions are dynamically adjusted, further improving the model's sensitivity to image features and segmentation accuracy. Using binary cross-entropy as the loss function, the segmentation results are optimized to more closely approximate the boundaries of the true lesion area, thereby improving the accuracy and usability of skin lesion detection.

[0098] Furthermore, the image detection method provided in the present embodiment achieved a 77.6% probability of correct keypoints when verified using the mean intersection over union (MIOU) estimation, and the RMSE and MSE of the algorithm model reached 4.57 and 5.65, respectively. These performances are significantly higher than those of image detection methods in related technologies.

[0099] The embodiment of the present application provides an image detection device, Figure 6 It is a structural diagram of the device. Figure 6 As can be seen from the figure, the device includes: a first processing module 60, which is used to determine the curvature information of each area in the initial image, and remove the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; a second processing module 62, which is used to determine the abnormal area in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; a third processing module 64, which is used to determine the skin lesion area based on the abnormal area.

[0100] In some embodiments of the present application, the first processing module 60 determines the curvature information of each area in the initial image, and the step of removing the hair structure in the initial image based on the curvature information includes: determining the Hessian matrix of the initial image, wherein the Hessian matrix is ​​used to describe the curvature information of each area in the initial image; determining the eigenvalues ​​of the Hessian matrix, wherein the eigenvalues ​​are used to reflect the principal curvature of each area; and removing the hair structure in the initial image based on the eigenvalues.

[0101] In some embodiments of the present application, the encoder branch includes multiple downsampling units, wherein each downsampling unit includes a parallel multi-branch module, the parallel multi-branch module includes a parallel first branch and a second branch, the first branch includes multiple parallel convolution layers, and the second branch also includes multiple parallel convolution layers; the convolution kernel sizes corresponding to the multiple parallel convolution layers in the first branch are different from each other, and the convolution kernel sizes corresponding to the multiple parallel convolution layers in the second branch are different from each other.

[0102] In some embodiments of the present application, the resolutions of the feature maps output by each upsampling unit in the decoder branch are different from each other, wherein the pyramid pooling module is used to fuse the feature maps output by each upsampling unit to determine the abnormal area.

[0103] In some embodiments of the present application, the attention mechanism is a hybrid spatial channel attention mechanism that includes a channel attention mechanism and a spatial attention mechanism, wherein the channel attention mechanism includes a small-order statistics module and a large-order statistics module, the small-order statistics module is used to determine the local detail features of the channel, and the large-order statistics module is used to determine the global dependency between channels.

[0104] In some embodiments of the present application, the small-order statistics module includes a global average pooling layer and a first fully connected layer, wherein the global average pooling layer is used to extract the small-order features of the image input into the small-order statistics module, and the fully connected layer is used to process the small-order features to obtain the small-order attention weights corresponding to the small-order statistics module, and the small-order features include the statistical features of the image; the large-order statistics module includes a bilinear pooling layer and a second fully connected layer, wherein the bilinear pooling layer is used to perform pairwise outer product processing on the feature vectors of each channel in the image input into the large-order statistics module to obtain a feature matrix, and the second fully connected layer is used to process the feature matrix to obtain the large-order attention weights corresponding to the large-order statistics module.

[0105] In some embodiments of the present application, the second processing module 62 is also used to: determine the small-order attention weight determined by the small-order statistics module; determine the large-order attention weight determined by the large-order statistics module; and perform element-by-element addition processing on the small-order attention weight and the large-order attention weight to obtain the channel attention weight.

[0106] It should be noted that the various modules in the above-mentioned image detection device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0107] According to an embodiment of the present application, a non-volatile storage medium is also provided, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following image detection method: determining the curvature information of each area in the initial image, and removing the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; determining the abnormal area in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; determining the skin lesion condition based on the abnormal area.

[0108] According to an embodiment of the present application, an electronic device is also provided, including a memory and a processor, the processor being used to run a program stored in the memory, wherein the following image detection method is executed when the program is running: determining the curvature information of each area in an initial image, and removing the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; determining the abnormal area in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; determining the skin lesion condition based on the abnormal area.

[0109] According to an embodiment of the present application, a computer program product is also provided, including a computer program. When the computer program is executed by a processor, it implements the following image detection method: determining the curvature information of each area in an initial image, and removing the hair structure in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; determining the abnormal area in the target image through an image detection model, wherein the image detection model includes an encoder branch, a decoder branch and a pyramid pooling module, the encoder branch and the decoder branch are connected through an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, and the abnormal area is an area where the skin may have lesions; determining the skin lesion condition based on the abnormal area.

[0110] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0112] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0113] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0115] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image detection method, characterized in that: include: determining curvature information of each region in an initial image, and removing hair structures in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; Determining an abnormal area in the target image using an image detection model, wherein the image detection model includes an encoder branch, a decoder branch, and a pyramid pooling module, the encoder branch and the decoder branch are connected via an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, wherein the abnormal area is an area where skin lesions may exist; The skin lesion condition is determined based on the abnormal area.

2. The image detection method according to claim 1, wherein: Determining curvature information of each region in the initial image and removing the hair structure in the initial image according to the curvature information includes: Determining a Hessian matrix of the initial image, wherein the Hessian matrix is ​​used to describe curvature information of each region in the initial image; Determining eigenvalues ​​of the Hessian matrix, wherein the eigenvalues ​​are used to reflect the principal curvatures of the respective regions; The hair structure in the initial image is removed according to the feature value.

3. The image detection method according to claim 1, wherein: The encoder branch includes multiple downsampling units, wherein: Each of the downsampling units includes a parallel multi-branch module, the parallel multi-branch module includes a parallel first branch and a parallel second branch, the first branch includes a plurality of parallel convolutional layers, and the second branch also includes a plurality of parallel convolutional layers; The convolution kernel sizes corresponding to the multiple parallel convolution layers in the first branch are different from each other, and the convolution kernel sizes corresponding to the multiple parallel convolution layers in the second branch are different from each other.

4. The image detection method according to claim 1, wherein: The resolutions of the feature maps output by the upsampling units in the decoder branch are different from each other, where The pyramid pooling module is used to fuse the feature maps output by each of the upsampling units, so as to determine the abnormal area.

5. The image detection method according to claim 1, wherein: The attention mechanism is a hybrid spatial channel attention mechanism that includes a channel attention mechanism and a spatial attention mechanism, wherein, The channel attention mechanism includes a small-order statistics module and a large-order statistics module. The small-order statistics module is used to determine the local detail features of the channel, and the large-order statistics module is used to determine the global dependency between channels.

6. The image detection method according to claim 5, characterized in that: The small-order statistics module includes a global average pooling layer and a first fully connected layer, wherein the global average pooling layer is used to extract small-order features of the image input to the small-order statistics module, and the fully connected layer is used to process the small-order features to obtain small-order attention weights corresponding to the small-order statistics module, wherein the small-order features include statistical features of the image; The large-order statistics module includes a bilinear pooling layer and a second fully connected layer, wherein the bilinear pooling layer is used to perform pairwise outer product processing on the feature vectors of each channel in the image input into the large-order statistics module to obtain a feature matrix, and the second fully connected layer is used to process the feature matrix to obtain the large-order attention weight corresponding to the large-order statistics module.

7. The image detection method according to claim 5, characterized in that: The method further comprises: Determine the small-order attention weight determined by the small-order statistics module; Determining the large-order attention weight determined by the large-order statistics module; The small-order attention weight and the large-order attention weight are added element by element to obtain the channel attention weight.

8. An image detection device, characterized in that: include: a first processing module, configured to determine curvature information of each region in an initial image, and remove hair structures in the initial image based on the curvature information to obtain a target image, wherein the initial image is a skin image; a second processing module, configured to determine an abnormal area in the target image using an image detection model, wherein the image detection model includes an encoder branch, a decoder branch, and a pyramid pooling module, the encoder branch and the decoder branch are connected via an attention mechanism, and multiple upsampling units in the encoder branch are respectively connected to the pyramid pooling module, wherein the abnormal area is an area where skin lesions may be present; The third processing module is used to determine the skin lesion area based on the abnormal area.

9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the image detection method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the image detection method according to any one of claims 1 to 7 is executed when the program is run.

11. A computer program product, characterized in that The invention comprises a computer program, which implements the image detection method according to any one of claims 1 to 7 when being executed by a processor.