Lightweight hysteroscope image diseased region detection model construction method
By constructing a lightweight hysteroscopic image lesion site detection model, using multi-scale feature pyramid and lightweight backbone network design, the computational complexity of lesion detection on resource-constrained medical devices is solved, and efficient and accurate real-time diagnosis is achieved.
Patent Information
- Application Number
- CN202510429449.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
The existing hysteroscopic image lesion detection model has complex calculations and large calculations, making it difficult to operate on medical equipment with limited resources, and the detection efficiency and accuracy are insufficient, which cannot meet the needs of real-time diagnosis.
A lightweight hysteroscopic image lesion site detection model is constructed, and multi-scale feature pyramid, lightweight backbone network design, object detection head design, model compression and quantization technology is adopted, including multi-scale feature extraction, depth separable self-association mechanism, dynamic regional attention mechanism and adaptive anchor box generation to reduce computing complexity and storage requirements.
It improves the accuracy and efficiency of lesion site detection, adapts to the resource limitations of different medical equipment, and achieves fast and accurate real-time diagnosis.
Smart Images

Figure CN120339702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing and computer vision, and particularly relates to a method for constructing a lightweight hysteroscopy image lesion detection model, aiming to improve the efficiency and accuracy of lesion detection in hysteroscopy examinations, and is particularly applicable to medical devices with limited resources and real-time diagnosis scenarios. Background Art
[0002] Hysteroscopy examination is an important means for diagnosing intrauterine lesions. It can directly observe the internal situation of the uterus and provide key information for disease diagnosis and treatment. However, in actual clinical applications, the interpretation of hysteroscopy images depends on doctors' experience and professional knowledge, and there are problems such as strong subjectivity and low efficiency. Moreover, due to factors such as complex backgrounds of hysteroscopy images and unclear characteristics of lesion sites, it is challenging to accurately detect lesion sites.
[0003] Although traditional object detection models based on machine learning and deep learning can assist doctors in lesion detection to a certain extent, these models often have complex structures, large computational amounts, and high requirements for hardware devices. In medical scenarios, especially in some primary medical institutions, device resources are relatively limited and it is difficult to support the operation of complex models. In addition, real-time performance is also an important requirement in hysteroscopy examinations. The inference time of complex models is relatively long and cannot meet the requirements of rapid diagnosis. Therefore, developing a lightweight, efficient, and accurate hysteroscopy image lesion detection model has important clinical significance. Summary of the Invention
[0004] The purpose of the present invention is to construct a lightweight hysteroscopy image lesion detection model. Through innovative architecture design and technical means, while ensuring the detection accuracy of lesion sites, the computational complexity and storage requirements of the model are significantly reduced to adapt to the situation of limited medical device resources and achieve rapid and accurate real-time diagnosis.
[0005] The technical solution adopted by the present invention: A method for constructing a lightweight hysteroscopy image lesion detection model includes the following steps: Step 1: Construction of a multi-scale feature pyramid (1) Hierarchical feature extraction: Perform hierarchical feature extraction on the input hysteroscopy image I; divide the image into multiple non-overlapping image units, and each image unit is converted into a feature vector through feature mapping; use a series of network modules with different processing capabilities to process these feature vectors, and gradually extract feature information at different levels to obtain a basic feature map ; Subsequently, use network modules with different downsampling strategies to for further processing to generate feature maps of different scales , ,..., ; (2)Cross - layer feature fusion: Perform fusion operations on feature maps of adjacent scales; for feature maps and , first use the upsampling operation U to perform upsampling so that it has the same spatial size as ; then use the downsampling operation D to perform downsampling to enhance the diversity of features; add the processed feature map to to obtain the fused feature map , and the formula is ; through this cross - layer fusion method, the model can make full use of the information in feature maps of different scales and improve the detection ability for lesion sites of different sizes; Step 2: Design of lightweight backbone network (1)Depth - separable self - correlation mechanism: Improve the traditional feature correlation mechanism by introducing a depth - separable self - correlation mechanism, which decomposes the feature correlation calculation into two steps: depth self - correlation and point - wise self - correlation ; the depth self - correlation operation is performed independently within each channel to calculate the self - correlation relationship of the input feature X in the channel dimension, obtaining ; the point - wise self - correlation operation then performs 1x1 convolution on to adjust the channel dimension and obtain the final self - correlation output ; this decomposition method effectively reduces the complexity of feature correlation calculation and reduces the computational amount of the model; (2)Dynamic region attention mechanism: Introduce the dynamic region attention mechanism DRA, which dynamically adjusts the size and position of the attention region according to the possible distribution and scale of the lesion sites in the hysteroscopy image; at different positions and different levels of the feature map, adaptively determine the range of the attention region based on the common features and distribution rules of the lesion sites; Step 3: Design of object detection head (1)Multi - scale feature matching detection: Use the fused feature maps of different scales output by the backbone network to perform multi - scale detection of lesion sites; adopt the feature pyramid network (FPN) structure to perform further upsampling and downsampling operations on feature maps of different scales to make the feature maps have better consistency in semantics and space; for each scale of feature map, process it through a series of convolutional layers and fully - connected layers to predict the class probability of the lesion site and the bounding box regression parameters , ; the class probability is calculated using the softmax function, and the formula is , where is the score of category c, and K is the total number of lesion categories; the regression loss uses the smooth L1 loss function: where is the true offset; (2) Adaptive anchor box generation: To improve the accuracy and recall rate of lesion detection, an adaptive anchor box generation mechanism is introduced; according to the feature maps of different scales and the distribution of lesion sites, anchor boxes are generated adaptively; on high-resolution feature maps, smaller-scale anchor boxes are generated for detecting small lesion sites; on low-resolution feature maps, larger-scale anchor boxes are generated for detecting large lesion sites; at the same time, according to the common aspect ratio distribution of lesion sites, the aspect ratios of the anchor boxes are adjusted to better match the shape of the lesion sites; Step Four: Model Compression and Quantization (1) Pruning strategy: During the model training process, calculate the contribution of each parameter to the loss function as the importance score of the parameter , and the formula is where L is the loss function, is the model parameter; after training is completed, set a pruning threshold T , and remove the parameters with importance scores lower than T to reduce the number of model parameters; (2) Quantization method: Convert the model parameters from a high-precision data type to a low-precision data type.
[0006] Furthermore, in the first step, network modules with different downsampling strategies are used to perform further processing to generate feature maps of different scales , ,..., ; the downsampling rate of the first downsampling strategy is 2, the second is 4, and so on, so as to obtain multi-scale feature representations to capture the features of lesion sites of different sizes.
[0007] Furthermore, in the second step, at different positions and different levels of the feature map, according to the common features and distribution rules of lesion sites, the range of the region of interest is determined adaptively; among them, in places where suspected lesion regions are concentrated, the region of interest is narrowed to focus on local features; in the background region, the region of interest is enlarged to capture more extensive context information. Through this dynamic adjustment, the model can more effectively focus on the features of lesion sites and improve the detection accuracy.
[0008] Further, in step four, the model parameters are converted from a high-precision data type to a low-precision data type; specifically, 32-bit floating-point parameters are quantized into 8-bit integer parameters. By statistically analyzing the parameter distribution, the quantization factor s is determined, and the original parameters are quantized into , which reduces the storage requirement and computational complexity of the model while ensuring the model performance.
[0009] Advantages of the present invention: (1) Efficient feature utilization: Through the construction of a multi-scale feature pyramid and the design of a lightweight backbone network, it is possible to fully integrate feature information of different scales, effectively improve the detection accuracy of lesion sites of different sizes, and assist doctors in more accurately detecting lesions.
[0010] (2) Lightweight design: The application of the depthwise separable self-correlation mechanism, the dynamic region attention mechanism, and the model compression and quantization technology significantly reduces the computational complexity and storage requirement of the model, enabling the model to operate efficiently on resource-constrained medical devices and expanding the application scope of the technology.
[0011] (3) Strong adaptability: The adaptive anchor box generation and the dynamic region attention mechanism enable the model to adaptively adjust the detection strategy according to the characteristics of hysteroscopy images and the distribution of lesion sites, improving the adaptability and generalization ability of the model to complex hysteroscopy images and meeting the requirements of different clinical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the overall architecture diagram of the lightweight hysteroscopy image lesion site detection model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The present invention will be further described below with reference to the drawings and specific embodiments.
[0014] As Figure 1 shown, the overall structure of the lightweight hysteroscopy image lesion site detection model constructed by the present invention, from top to bottom, is an input layer, a multi-scale feature pyramid module, a lightweight backbone network module, a target detection head module, and an output layer. The input layer receives the hysteroscopy image, and through the multi-scale feature pyramid module, hierarchical feature extraction and cross-layer feature fusion are performed to obtain fused feature maps of different scales. These fused feature maps are input into the lightweight backbone network module, and are further processed through the depthwise separable self-correlation mechanism and the dynamic region attention mechanism to enhance the ability to capture lesion site features. The processed feature maps are then input into the target detection head module for multi-scale feature matching detection and adaptive anchor box generation, and the category and location of the lesion site are predicted. Finally, the output layer outputs the detection results of the lesion site, including the lesion category and the corresponding bounding box information.
[0015] A method for constructing a lightweight detection model for pathological regions in hysteroscopy images, comprising the following steps: Step 1: Construction of a multi-scale feature pyramid (1) Hierarchical feature extraction: The input hysteroscopy image I is segmented into 16x16 image units, and each image unit is mapped to a feature vector with a dimension of 768; 12 network modules with specific processing capabilities are used to process these feature vectors to obtain a basic feature map ; Then, network modules with downsampling rates of 2, 4, and 8 are used to process respectively to generate three feature maps with different scales; (2) Cross-layer feature fusion: For adjacent-scale feature maps, such as and , bilinear interpolation is used as the upsampling operation U to upsample so that it has the same spatial size as ; average pooling is used as the downsampling operation D to downsample ; the processed is added to to obtain a fused feature map ; similarly, and are obtained in turn; Step 2: Design of a lightweight backbone network (1) Depthwise separable self-correlation mechanism: In the depthwise self-correlation operation, the input features are evenly divided into 8 groups along the channel dimension, and each group independently performs self-correlation calculation; the pointwise self-correlation operation uses a convolutional layer with 768 1x1 convolutional kernels to adjust the channel dimension of the output of the depthwise self-correlation; (2) Dynamic region attention mechanism: According to the prior distribution information of pathological regions in a large number of hysteroscopy images, the image is divided into different regions; in the regions with a high incidence of suspected lesions, the size of the attention region is set to 8x8; in the background region, the size of the attention region is set to 16x16; at the same time, the center position of the attention region is dynamically adjusted according to the preliminarily detected suspected lesion positions in the image; Step 3: Design of an object detection head (1) Multi-scale feature matching detection: An FPN structure is adopted to process the , and feature maps of three scales output by the backbone network; bilinear interpolation is used for upsampling, and max pooling is used for downsampling; each scale of feature map is processed through 3 convolutional layers and 2 fully connected layers to predict the class probability and bounding box regression parameters of the pathological region (2) Adaptive anchor box generation: On (high-resolution feature map), generate anchor boxes with scales of 32x32 and aspect ratios of 1:1, 1:2, and 2:1 for detecting smaller lesion sites; on , generate anchor boxes with a scale of 64x64 and the same aspect ratios; on (low-resolution feature map), generate anchor boxes with a scale of 128x128 and the same aspect ratios for detecting larger lesion sites; Step 4: Model compression and quantization (1) Pruning strategy: During the model training process, calculate the importance scores of the parameters every 5 epochs ; after the training is completed, determine the pruning threshold T through experiments, and remove the parameters with importance scores lower than T ; (2) Quantization method: Quantize the trained model parameters. By statistically analyzing the parameter distribution, determine the quantization factor s, and quantize 32-bit floating-point parameters into 8-bit integer parameters to obtain the final lightweight hysteroscopy image lesion site detection model.
[0016] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A method for constructing a lightweight detection model for pathological regions in hysteroscopy images, characterized in that: It includes the following steps: Step 1: Construction of multi-scale feature pyramid (1)Hierarchical feature extraction: Perform hierarchical feature extraction on the input hysteroscopy image I; divide the image into multiple non-overlapping image units, and each image unit is converted into a feature vector through feature mapping; use a series of network modules with different processing capabilities to process these feature vectors, and gradually extract feature information at different levels to obtain the basic feature map ; Subsequently, network modules with different downsampling strategies are used to further process it to generate feature maps of different scales , , ..., ; (2)Cross-layer feature fusion: Perform a fusion operation on feature maps of adjacent scales; for the feature maps and , first use the upsampling operation U to perform upsampling so that it has the same spatial dimensions as ; then use the downsampling operation D to perform downsampling to enhance the diversity of features; add the processed feature map to to obtain the fused feature map , and the formula is ; Step 2: Design of lightweight backbone network (1) Depthwise separable self-correlation mechanism: Decompose the feature correlation calculation into depthwise self-correlation and pointwise self-correlation in two steps; The depthwise self-correlation operation is performed independently within each channel to calculate the self-correlation relationship of the input feature X in the channel dimension, resulting in ; The pointwise self-correlation operation then performs a 1x1 convolution on to adjust the channel dimension and obtain the final self-correlation output ; (2) Dynamic region attention mechanism: Introduce the dynamic region attention mechanism DRA, and dynamically adjust the size and position of the attention region according to the possible distribution and scale of the lesion sites in the hysteroscopy images; at different positions and different levels of the feature map, adaptively determine the range of the attention region based on the common features and distribution rules of the lesion sites; Step 3: Design of object detection head (1)Multi-scale feature matching detection: Use the fused feature maps of different scales output by the backbone network to perform multi-scale detection of lesion sites; Using a feature pyramid network structure, further upsampling and downsampling operations are performed on feature maps of different scales to make the feature maps have better consistency in semantics and space; for each scale of feature map, it is processed through a series of convolutional layers and fully connected layers to predict the class probabilities of the lesion sites and the bounding box regression parameters , ; The class probabilities are calculated using the softmax function, with the formula , where is the score for class c, K is the total number of lesion classes; the regression loss uses the smooth L1 loss function: wherein is the true offset; (2) Adaptive anchor box generation: To improve the accuracy and recall rate of lesion site detection, introduce the adaptive anchor box generation mechanism; adaptively generate anchor boxes according to the feature maps of different scales and the distribution of lesion sites; generate smaller-scale anchor boxes on the high-resolution feature map for detecting small lesion sites; generate larger-scale anchor boxes on the low-resolution feature map for detecting large lesion sites; at the same time, adjust the aspect ratio of the anchor boxes according to the common aspect ratio distribution of the lesion sites to better match the shape of the lesion sites; Step 4: Model compression and quantization (1)Pruning strategy: During the model training process, calculate the contribution of each parameter to the loss function as the importance score of the parameter , and the formula is , where L is the loss function and are the model parameters; after training, set a pruning threshold T , and remove the parameters with importance scores lower than T to reduce the number of model parameters; (2) Quantization method: Convert the model parameters from high-precision data types to low-precision data types.
2. A method for constructing a lightweight hysteroscopy image lesion site detection model according to claim 1, characterized in that: In the first step, a network module with different downsampling strategies is used to perform further processing to generate feature maps of different scales , ,..., ; among which, the downsampling rate of the first downsampling strategy is 2, the second is 4, and so on, so as to obtain multi-scale feature representations to capture the features of lesion sites of different sizes.
3. A method for constructing a lightweight hysteroscopy image lesion site detection model according to claim 2, characterized in that: In Step 2, at different positions and different levels of the feature map, adaptively determine the range of the attention region based on the common features and distribution rules of the lesion sites; among them, in the places where the suspected lesion regions are concentrated, narrow the attention region to focus on local features; in the background region, enlarge the attention region to capture more extensive context information.
4. A method for constructing a lightweight hysteroscopy image lesion site detection model according to claim 3, characterized in that: In the fourth step, the model parameters are converted from a high-precision data type to a low-precision data type; among them, 32-bit floating-point parameters are quantized into 8-bit integer parameters; by statistically analyzing the parameter distribution, the quantization factor s is determined, and the original parameters are quantized to .
Citation Information
Cited By
Abnormal lesion area detection method of endoscope image
CN120747032A