Multi-resolution-channel attention network construction method capable of detecting high-heterogeneity X-ray skull side image key points

By adopting a multi-resolution-channel attention network in the key point positioning of skull lateral image, the problems of low feature extraction efficiency and lack of effective attention mechanism in the prior art are solved, and high-precision and fast key point positioning are achieved, which is suitable for the detection of high heterogeneous images.

CN120047668AActive Publication Date: 2025-05-27AFFILIATED STOMATOLOGICAL HOSPITAL OF NANJING MEDICAL UNIV +1

Patent Information

Application Number
CN202510101168.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art has problems such as low feature extraction efficiency, lack of effective attention mechanism, single feature fusion method and difficulty in dealing with high heterogeneity images in the key points positioning of skull lateral image.

Method used

Using a multi-resolution-channel attention network, the channel attention mechanism module, lightweight fusion convolution design and multi-scale attention-guided feature fusion network are introduced to highlight important channel characteristics, realize multi-scale feature expression, and improve the generalization ability of the model through multi-level aggregation loss function.

Benefits of technology

It significantly improves the model's ability to capture details in the head X-ray image, improves the accuracy and speed of key point positioning, and can quickly complete real-time key point detection tasks, providing accurate and efficient auxiliary support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047668A_ABST
    Figure CN120047668A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-resolution-channel attention network construction method capable of detecting high-heterogeneity X-ray skull side image key points. Firstly, a feature extraction network composed of lightweight fusion convolution and a channel attention mechanism is constructed, the calculation complexity is reduced, and non-important channel features are inhibited. And then a multi-resolution progressive feature extraction sub-network is constructed to extract multi-resolution features, and the feature multi-scale expression capability is enhanced. And then a multi-scale attention guidance feature fusion network is constructed to be connected with multiple points of the sub-networks, so that the extracted multi-resolution features are fully fused. The constructed prediction network aligns and splices multi-scale features through up-sampling to fuse different levels of information, uses an independent detection head to process different levels of features, automatically outputs prediction results of key point positions in an image, and improves the model generalization ability through a multi-level aggregation loss function. According to the invention, automatic detection of the key points in the oral clinical high-heterogeneity X-ray skull side image can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision prediction, and particularly to a method for constructing a multi-resolution-channel attention network capable of detecting key points in high-heterogeneity X-ray lateral cephalometric images. Background Art

[0002] The measurement and analysis of lateral cephalometric images are key steps in the research of craniofacial growth and development, malformation diagnosis, and formulation of correction plans in the modern orthodontics field. The analysis results will directly affect the doctor's judgment of the patient's overall condition. When analyzing lateral cephalometric images, doctors need to mark the key points of anatomical structures such as dentition and craniofacial on the images, and conduct quantitative and qualitative analysis on the distance and angular relationships between these key points. These analysis results will provide an important basis for formulating diagnosis and treatment and surgical plans. Therefore, the accuracy of key point coordinate positioning is directly related to the effectiveness of diagnosis and treatment.

[0003] Currently, most doctors in clinical practice still use the method of manually positioning key points. The speed and accuracy of this method largely depend on the doctor's professional level and physical condition, and may lead to diagnostic errors or even surgical failures due to positioning deviations. Therefore, it is particularly urgent to develop an automatic detection method for key points in high-heterogeneity X-ray lateral cephalometric images to assist doctors in accurately positioning the key point positions. This not only helps to improve the positioning accuracy but also significantly improves the clinical diagnosis efficiency.

[0004] In recent years, deep learning has demonstrated excellent performance in fields such as image object detection and recognition. Some scholars have attempted to introduce deep learning models into the orthodontics field to promote the development of key point automatic detection technology.

[0005] Yang et al. proposed a U-shaped network CETransNet based on convolutional enhanced Transformer for key point positioning in lateral cephalometric images. While establishing global context connections, this network retains the ability of convolutional neural networks to obtain local information, introduces an exponential weighted loss function, and further improves the positioning accuracy by suppressing the loss of distant pixels. Experimental results show that this method has certain clinical application value.

[0006] Wang et al. proposed the high-resolution network HRNet, which shows excellent performance in key point detection and human pose estimation tasks by gradually introducing multi-scale feature fusion while maintaining high-resolution features. Its superior multi-scale processing ability makes it particularly suitable for clinical application scenarios that require high-precision key point positioning.

[0007] In summary, in the fields of medical image analysis and facial feature key point detection, high-precision key point localization is of great significance. Although the traditional HRNet performs well in maintaining high-resolution features, there are still the following problems:

[0008] 1. The feature extraction efficiency is not high, and the consumption of computing resources is large;

[0009] 2. Lack of an effective attention mechanism, making it difficult to highlight key features;

[0010] 3. A single feature fusion method, without making full use of multi-scale information;

[0011] 4. All the processed images have high consistency, and the detection of highly heterogeneous images has not been achieved.

[0012] To address these deficiencies, a method for constructing a multi-resolution-channel attention network for detecting key points in highly heterogeneous lateral cephalometric X-ray images is proposed. By introducing a channel attention mechanism module, it highlights the features of important channels and suppresses the interference of irrelevant channels. The design of lightweight fusion convolution reduces the number of model parameters and thus the computational complexity while maintaining the effectiveness of feature extraction. In the multi-scale attention-guided feature fusion network, the features extracted by the lightweight fusion feature extraction network and the multi-resolution progressive feature extraction sub-network are repeatedly fused to achieve more effective multi-scale feature representation. The introduction of a multi-level aggregation loss function in the prediction network improves the generalization ability of the model. The optimized network structure and efficient feature extraction method effectively improve the processing speed of the model, enabling it to complete the real-time key point detection task of lateral cephalometric images faster, thereby providing accurate and efficient auxiliary support for doctors in oral and maxillofacial surgery to accurately locate the key point positions of highly heterogeneous lateral cephalometric images. Summary of the Invention

[0013] To solve the above technical problems, the present invention provides a method for constructing a multi-resolution-channel attention network for detecting key points in highly heterogeneous lateral cephalometric X-ray images, including the following steps:

[0014] Step 1: Construct a lightweight fusion feature extraction network: The lightweight fusion feature extraction network is composed of several combinations of lightweight fusion convolution-batch normalization-ReLU activation function layers, several bottleneck units connected in sequence, and finally passes through a channel attention mechanism module;

[0015] Step 2: Construct a multi-resolution progressive feature extraction sub-network: The multi-resolution progressive feature extraction sub-network is based on a simplified feature pyramid network as the basic structure, performs multi-resolution feature extraction on the original input image to construct features of different resolutions, and fuses them in the main part of the network through skip connections to achieve more effective multi-scale feature representation;

[0016] Step 3: Construct a multi-scale attention-guided feature fusion network: The multi-scale attention-guided feature fusion network consists of several combinations of convolutional-batch normalization-ReLU activation function layers and several multi-scale attention-guided feature fusion sub-networks connected in sequence;

[0017] Step 4: Construct a prediction network: The prediction network aligns the multi-scale features obtained from the multi-scale attention-guided feature fusion network to the same scale through upsampling, and then splices them to fuse information at different levels; then the features before splicing and the features after splicing are processed by independent detection heads respectively to process features at different levels and generate prediction results, and the generalization ability of the model is improved by supervising through a multi-level aggregation loss function;

[0018] Step 5: Network training and prediction: Use the constructed multi-resolution-channel attention key point detection network to train the labeled lateral cephalometric images; input the newly acquired lateral cephalometric images to be detected into the trained network, automatically predict the positions of the key points in the images, and output the error between the prediction results and the manual annotations.

[0019] The further limited technical solution of the present invention is:

[0020] In the method for constructing the multi-resolution-channel attention key point detection network described above, in step 1, the lightweight fusion feature extraction network includes a lightweight fusion convolutional-batch normalization-ReLU activation function combination layer (LBR), a bottleneck unit, and a channel attention mechanism module:

[0021] First, the input image passes through a lightweight fusion convolutional-batch normalization-ReLU activation function combination layer (LBR) to complete preliminary feature extraction, and the extracted feature size is H×W; then it passes through a second lightweight fusion convolutional-batch normalization-ReLU activation function combination layer (LBR) to complete downsampling while performing feature extraction, and the feature size becomes H / 2×W / 2;

[0022] The obtained features pass through four cascaded bottleneck units and then through a channel attention mechanism module.

[0023] In the lightweight fusion feature extraction network described above, in the lightweight fusion convolution, the input feature passes through a lightweight convolution and then is spliced with the feature passing through a standard convolution, and the splicing result passes through the ReLU activation function to obtain the output of the lightweight fusion convolution.

[0024] In the aforementioned lightweight fusion feature extraction module, in the bottleneck unit, the input features sequentially pass through two convolutional-batch normalization-activation function combination layers (CBR), a convolution, and a batch normalization, and then perform an element-wise addition operation with the input features. The result of the addition is passed through the ReLU activation function to obtain the output of the bottleneck unit.

[0025] In the aforementioned lightweight fusion feature extraction module, the channel attention mechanism module captures global information at the channel level, multiplies the obtained weights with the original features, realizes adaptive feature recalibration at the channel level, highlights the features of important channels, suppresses the interference of irrelevant channels, and significantly improves the model's ability to capture detailed features in head X-ray images. The specific process is as follows:

[0026] First, for the input features with the size of B×C×H×W, the information of the spatial dimension H×W is compressed into a global channel feature vector through global average pooling, and the size after compression is B×C. Then, a two-layer fully connected network is used to generate the weight value for each channel: the first layer reduces the number of channels from C to C / 16, and the dimensionality reduction reduces the number of parameters and computational complexity and learns the compressed features between channels, and then the ReLU activation function is used to introduce non-linearity; the second layer restores to the original number of channels C, and at the same time generates a weight value for each channel, and then passes through the Sigmoid activation function to be normalized to [0,1]. The generated channel weights are multiplied with the input features channel by channel to complete the enhancement or suppression of the channel features.

[0027] In the aforementioned method for constructing a multi-resolution-channel attention key point detection network, in step 2, the multi-resolution progressive feature extraction sub-network includes a convolutional-batch normalization-ReLU activation function combination layer (CBR), a convolutional-bilinear interpolation upsampling combination layer, and a weighted fusion module:

[0028] First, the input image has the dimension of H×W. After passing through a CBR module, the feature X1 is obtained, and the feature size is H / 2×W / 2; after X1 passes through a CBR module, the feature X2 is obtained, and the feature size is H / 4×W / 4; after X2 passes through a CBR module, the feature X3 is obtained, and the feature size is H / 8×W / 8; after X3 passes through a CBR module, the feature X4 is obtained, and the feature size is H / 16×W / 16;

[0029] Feature X2 is weighted and fused with feature X1 after passing through the upsampling combination layer to obtain feature Y1; feature X3 is weighted and fused with feature X2 after passing through the upsampling combination layer to obtain feature Y2; feature X4 is weighted and fused with feature X3 after passing through the upsampling combination layer to obtain feature Y3; feature Y4 is directly passed from feature X4;

[0030] The finally output features are Y1, Y2, Y3, and Y4, where the feature size of Y1 is H / 2×W / 2, the feature size of Y2 is H / 4×W / 4, the feature size of Y3 is H / 8×W / 8, and the feature size of Y4 is H / 16×W / 16.

[0031] In the method for constructing the multi-resolution-channel attention key point detection network described above, in step 3, the multi-scale attention-guided feature fusion network includes a convolution-batch normalization-ReLU activation function combination layer (CBR), a multi-scale attention-guided feature fusion sub-network one, a multi-scale attention-guided feature fusion sub-network two, and a multi-scale attention-guided feature fusion sub-network three:

[0032] First, the input feature passes through a CBR module to change the current number of channels while maintaining the original resolution branch, obtaining feature A1; the input feature passes through a CBR module for downsampling to generate a new resolution branch, obtaining feature A2; A1 and A2 pass through the multi-scale attention-guided feature fusion sub-network one, and the output features are D1 and D2; among them, the feature sizes of the input feature, feature A1, and feature D1 are H / 2×W / 2; the feature sizes of feature A2 and feature D2 are H / 4×W / 4;

[0033] The output features D1 and D2 respectively pass through a CBR module to obtain features A3 and A4; D2 passes through a CBR module for downsampling to generate a new resolution branch, obtaining feature A5; A3, A4, and A5 pass through four cascaded multi-scale attention-guided feature fusion sub-network twos, and the output features are D3, D4, and D5; among them, the feature sizes of feature A3 and feature D3 are H / 2×W / 2; the feature sizes of feature A4 and feature D4 are H / 4×W / 4; the feature sizes of feature A5 and feature D5 are H / 8×W / 8;

[0034] The output features D3, D4, and D5 respectively pass through a CBR module to obtain features A6, A7, and A8; D5 passes through a CBR module for downsampling to generate a new resolution branch, obtaining feature A9; A6, A7, A8, and A9 pass through three cascaded multi-scale attention-guided feature fusion sub-network threes, and the output features are D6, D7, D8, and D9; among them, the feature sizes of feature A6 and feature D6 are H / 2×W / 2, the feature sizes of feature A7 and feature D7 are H / 4×W / 4, the feature sizes of feature A8 and feature D8 are H / 8×W / 8, and the feature sizes of feature A9 and feature D9 are H / 16×W / 16.

[0035] In the aforementioned multi-scale attention-guided feature fusion network, the multi-scale attention-guided feature fusion sub-network one includes a basic unit, a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0036] In the multi-scale attention-guided feature fusion sub-network one, the input features are A1 and A2 respectively; feature A1 passes through four cascaded basic units to obtain feature B1; feature A2 passes through four cascaded basic units to obtain feature B2; where the feature size of feature B1 is H / 2×W / 2, and the feature size of feature B2 is H / 4×W / 4;

[0037] Feature B1 is added element-wise to B2 that has passed through the upsampling combination layer to obtain feature C1; feature B2 is added element-wise to B1 that has passed through the downsampling combination layer 1 to obtain feature C2; the feature size of feature C1 is H / 2×W / 2, and the feature size of feature C2 is H / 4×W / 4;

[0038] After feature C1 passes through the ReLU activation function, it is weighted and fused with Y1 output in step 2, and then passes through the channel attention mechanism to obtain feature D1; after feature C2 passes through the ReLU activation function, it is weighted and fused with Y2 output in step 2, and then passes through the channel attention mechanism to obtain feature D2.

[0039] In the aforementioned multi-scale attention-guided feature fusion network, the multi-scale attention-guided feature fusion sub-network two includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0040] In the multi-scale attention-guided feature fusion sub-network two, the input features are A3, A4, and A5 respectively; feature A3 passes through eight cascaded basic units to obtain feature B3; feature A4 passes through eight cascaded basic units to obtain feature B4; feature A5 passes through eight cascaded basic units to obtain feature B5; the feature size of feature B3 is H / 2×W / 2, the feature size of feature B4 is H / 4×W / 4, and the feature size of feature B5 is H / 8×W / 8;

[0041] Feature B3 is sequentially added element-wise to B4 and B5 that have passed through the upsampling combination layer to obtain feature C3; feature B4 is sequentially added element-wise to B3 that has passed through the downsampling combination layer 1 and B5 that has passed through the upsampling combination layer to obtain feature C4; feature B5 is sequentially added element-wise to B3 that has passed through the downsampling combination layer 2 and B4 that has passed through the downsampling combination layer 1 to obtain feature C5; the feature size of feature C3 is H / 2×W / 2, the feature size of feature C4 is H / 4×W / 4, and the feature size of feature C5 is H / 8×W / 8;

[0042] After the feature C3 passes through the ReLU activation function, it is weighted and fused with Y1 output in Step 2, and then passes through the channel attention mechanism to obtain the feature D3; after the feature C4 passes through the ReLU activation function, it is weighted and fused with Y2 output in Step 2, and then passes through the channel attention mechanism to obtain the feature D4; after the feature C5 passes through the ReLU activation function, it is weighted and fused with Y3 output in Step 2, and then passes through the channel attention mechanism to obtain the feature D5.

[0043] In the multi-scale attention-guided feature fusion network described above, the multi-scale attention-guided feature fusion sub-network three includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, a CBR-CBR-convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0044] In the multi-scale attention-guided feature fusion sub-network three, the input features are A6, A7, A8, and A9 respectively; the feature A6 passes through eight cascaded basic units to obtain the feature B6; the feature A7 passes through eight cascaded basic units to obtain the feature B7; the feature A8 passes through eight cascaded basic units to obtain the feature B8; the feature A9 passes through eight cascaded basic units to obtain the feature B9; the feature size of the feature B6 is H / 2×W / 2, the feature size of the feature B7 is H / 4×W / 4, the feature size of the feature B8 is H / 8×W / 8, and the feature size of the feature B9 is H / 16×W / 16;

[0045] The feature B6 is element-wise added to B7, B8, and B9 that have passed through the upsampling combination layer in sequence to obtain the feature C6; the feature B7 is element-wise added to B6 that has passed through the downsampling combination layer 1, B8, and B9 that have passed through the upsampling combination layer in sequence to obtain the feature C7; the feature B8 is element-wise added to B6 that has passed through the downsampling combination layer 2, B7 that has passed through the downsampling combination layer 1, and B9 that have passed through the upsampling combination layer in sequence to obtain the feature C8; the feature B9 is element-wise added to B6 that has passed through the downsampling combination layer 3, B7 that has passed through the downsampling combination layer 2, and B8 that has passed through the downsampling combination layer 1 in sequence to obtain the feature C9; the feature size of the feature C6 is H / 2×W / 2, the feature size of the feature C7 is H / 4×W / 4, the feature size of the feature C8 is H / 8×W / 8, and the feature size of the feature C9 is H / 16×W / 16;

[0046] After the feature C6 passes through the ReLU activation function, it is weighted and fused with Y1 output in step 2, and then passes through the channel attention mechanism to obtain the feature D6; after the feature C7 passes through the ReLU activation function, it is weighted and fused with Y2 output in step 2, and then passes through the channel attention mechanism to obtain the feature D7; after the feature C8 passes through the ReLU activation function, it is weighted and fused with Y3 output in step 2, and then passes through the channel attention mechanism to obtain the feature D8; after the feature C9 passes through the ReLU activation function, it is weighted and fused with Y4 output in step 2, and then passes through the channel attention mechanism to obtain the feature D9.

[0047] In the basic unit described above, the input feature sequentially passes through a combination of a CBR, a convolution, and a batch normalization, and then performs an element-wise addition operation with the input feature. The result of the addition is passed through the ReLU activation function to obtain the output of the basic unit.

[0048] In the method for constructing the multi-resolution-channel attention key point detection network described above, in step 4, the prediction network includes a bilinear interpolation upsampling module, a splicing module, and a detection head module:

[0049] First, the multi-scale features obtained by the multi-scale attention-guided feature fusion network are aligned to the same scale H / 2×W / 2 through bilinear interpolation upsampling, and the obtained features are Z1, Z2, Z3, and Z4. Four prediction results are obtained by passing them through four independent detection heads; the features Z1, Z2, Z3, and Z4 are spliced along the channel dimension to obtain the feature Z0, and a prediction result is obtained by passing through a detection head;

[0050] The loss of the five obtained prediction results is calculated, and the five generated losses are fused through a multi-level aggregation loss function. The feature learning ability of the model is improved through multi-scale supervision, so as to provide a stronger supervision signal, optimize the feature learning process, and enhance the stability of the detection result.

[0051] In the prediction network described above, in the detection head module, the input feature sequentially passes through a combination of a CBR and a convolution to obtain the output of the detection head.

[0052] In the method for constructing the multi-resolution-channel attention key point detection network described above, in step 5, while the network prediction part outputs the key point coordinate positions in the output image, it compares the prediction result with the manually marked result and outputs the error between them, so as to obtain the qualification rate of the test data under different evaluation indexes, so as to better observe the accuracy of the prediction result.

[0053] Compared with the prior art, the technical effects of the present invention are as follows: First, the present invention inputs the newly acquired lateral cephalometric image into the trained multi-resolution-channel attention key point detection network. During the test process, only one scan of the entire lateral cephalometric image is required to quickly and automatically detect the positions of the key points in the image, and give the error between the prediction result and the manual annotation and the qualification rate evaluation index. Second, in the present invention, the channel attention mechanism module and lightweight fusion convolution are used to adaptively adjust the weights of each channel to highlight the features of important channels while reducing the network complexity and the number of parameters. Second, in the present invention, the multi-scale attention-guided feature fusion network is used to repeatedly fuse the features extracted by the lightweight fusion feature extraction network and the multi-resolution progressive feature extraction sub-network, realizing a more effective multi-scale feature expression. In addition, in the prediction network of the present invention, a multi-level aggregation loss function is introduced to supervise and improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a schematic flow chart of the specific implementation steps of the present invention;

[0055] Figure 2 is a schematic diagram of the overall structure of the network model of the present invention;

[0056] Figure 3 is a schematic diagram of some basic structures in the network model of the present invention;

[0057] Figure 4 is a schematic diagram of the multi-resolution progressive feature extraction sub-network in the network model of the present invention;

[0058] Figure 5 is a schematic diagram of the multi-scale attention-guided feature fusion sub-network in the network model of the present invention;

[0059] Figure 6 is a schematic diagram of the result of detecting the key points of the lateral cephalometric image in the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The present embodiment provides a method for constructing a multi-resolution-channel attention network capable of detecting key points of highly heterogeneous X-ray lateral cephalometric images, as Figure 1 shown, including the following steps:

[0061] Step 1: Data collection and processing. Collect lateral cephalometric images, and mark the position coordinate information of 12 key points in the images from the lateral cephalometric images. Perform data augmentation processing on the lateral cephalometric images to generate diverse training data to improve the generalization ability of the model. Integrate the processed image data to form a training set for model training.

[0062] Data collection and processing include the following two parts:

[0063] (1) Mark the position coordinate information of 12 key feature points from the lateral cephalometric image. The specific steps are as follows: First, mark within the physical coordinate system with the measurement unit set to millimeters. Subsequently, use the established conversion algorithm to convert the position coordinates marked in the physical coordinate system into the corresponding coordinates in the pixel coordinate system, thus completing the mapping from the physical space to the pixel space.

[0064] (2) Use data augmentation techniques for the training lateral cephalometric images. The specific steps are as follows: Through diverse processing such as random scaling and rotation of the lateral cephalometric images, enable the model to better adapt to different angles and background changes, while making up for the defect of less image data and improving the robustness and performance of the detection method.

[0065] Step 2: Construct a multi-resolution-channel attention key-point detection network model. As Figure 2 shown, this model includes a lightweight fusion feature extraction network, a multi-resolution progressive feature extraction sub-network, a multi-scale attention-guided feature fusion network, and a prediction network.

[0066] The lightweight fusion feature extraction network consists of several lightweight fusion convolution-batch normalization-activation function combination layers and several bottleneck units connected in sequence, and finally consists of a channel attention mechanism module; the multi-resolution progressive feature extraction sub-network extracts multi-resolution features and fuses them in the main part of the network through skip connections to achieve more effective multi-scale feature expression; the multi-scale attention-guided feature fusion network consists of several convolution-batch normalization-activation function combination layers and several multi-scale attention-guided feature fusion sub-networks connected in sequence. The convolution combination layer transfers the features of the original branches and generates new-resolution branches, and the multi-scale attention-guided feature fusion sub-network repeatedly fuses the information of all current branches; the prediction network aligns the multi-scale features to the same scale through upsampling and splices them to fuse information at different levels. The features before splicing and the features after splicing are processed by independent detection heads to separately process features at different levels and generate prediction results, and the generalization ability of the model is supervised and improved through a multi-level aggregation loss function.

[0067] As Figure 2 shown, the lightweight fusion feature extraction network includes lightweight fusion convolution-batch normalization-ReLU activation function combination layers (LBR), bottleneck units, and a channel attention mechanism module:

[0068] First, the input image passes through a lightweight fusion convolutional-batch normalization-ReLU activation function combination layer (LBR) to complete preliminary feature extraction, and the extracted feature size is H×W; then it passes through a second lightweight fusion convolutional-batch normalization-ReLU activation function combination layer (LBR), and while performing feature extraction, it completes the downsampling operation, and the feature size becomes H / 2×W / 2;

[0069] After passing the obtained features through four cascaded bottleneck units, they are then passed through a channel attention mechanism module.

[0070] As Figure 3 shown, in the lightweight fusion convolution of the lightweight fusion feature extraction network, the input feature passes through a lightweight convolution and then is concatenated with the feature passing through a standard convolution, and the concatenated result passes through the ReLU activation function to obtain the output of the lightweight fusion convolution.

[0071] As Figure 3 shown, in the bottleneck unit of the lightweight fusion feature extraction network, the input feature passes through two convolutional-batch normalization-activation function combination layers (CBR), a convolution, and a combination of batch normalization in sequence, and then performs an element-wise addition operation with the input feature, and the added result passes through the ReLU activation function to obtain the output of the bottleneck unit.

[0072] The channel attention mechanism module captures global information at the channel level, multiplies the obtained weights with the original features, realizes channel-level adaptive feature recalibration, highlights the features of important channels, suppresses the interference of irrelevant channels, and significantly improves the model's ability to capture detailed features in head X-ray images; the specific process is as follows:

[0073] First, for the input feature with the size of B×C×H×W, the information of the spatial dimension H×W is compressed into a global channel feature vector through global average pooling, and the size after compression is B×C; then a two-layer fully connected network is used to generate the weight value of each channel: the first layer reduces the number of channels from C to C / 16, and the dimensionality reduction reduces the number of parameters and computational amount and learns the compressed features between channels, and then uses the ReLU activation function to introduce non-linearity; the second layer restores to the original number of channels C, and at the same time generates a weight value for each channel, and then normalizes it to [0,1] through the Sigmoid activation function; multiply the generated channel weights with the input features channel by channel to complete the enhancement or suppression of the channel features.

[0074] As Figure 4 shown, the multi-resolution progressive feature extraction sub-network includes a convolutional-batch normalization-ReLU activation function combination layer (CBR), a convolutional-bilinear interpolation upsampling combination layer, and a weighted fusion module:

[0075] First, the input image has dimensions H×W. After passing through a CBR module, feature X1 is obtained, with a feature size of H / 2×W / 2; X1 passes through a CBR module to obtain feature X2, with a feature size of H / 4×W / 4; X2 passes through a CBR module to obtain feature X3, with a feature size of H / 8×W / 8; X3 passes through a CBR module to obtain feature X4, with a feature size of H / 16×W / 16;

[0076] Feature X2 passes through an upsampling combination layer and is weighted and fused with feature X1 to obtain feature Y1; feature X3 passes through an upsampling combination layer and is weighted and fused with feature X2 to obtain feature Y2; feature X4 passes through an upsampling combination layer and is weighted and fused with feature X3 to obtain feature Y3; feature Y4 is directly obtained by passing feature X4;

[0077] Finally, the output features are Y1, Y2, Y3, and Y4, where the feature size of Y1 is H / 2×W / 2, the feature size of Y2 is H / 4×W / 4, the feature size of Y3 is H / 8×W / 8, and the feature size of Y4 is H / 16×W / 16.

[0078] As Figure 2 shown, the multi-scale attention-guided feature fusion network includes a convolution-batch normalization-ReLU activation function combination layer (CBR), a multi-scale attention-guided feature fusion sub-network one, a multi-scale attention-guided feature fusion sub-network two, and a multi-scale attention-guided feature fusion sub-network three:

[0079] First, the input feature passes through a CBR module to change the current number of channels while maintaining the original resolution branch, obtaining feature A1; the input feature passes through a CBR module for downsampling to generate a new resolution branch, obtaining feature A2; A1 and A2 pass through the multi-scale attention-guided feature fusion sub-network one, and the output features are D1 and D2; among them, the feature sizes of the input feature, feature A1, and feature D1 are: H / 2×W / 2; the feature sizes of feature A2 and feature D2 are H / 4×W / 4;

[0080] The output features D1 and D2 respectively pass through a CBR module to obtain features A3 and A4; D2 passes through a CBR module for downsampling to generate a new resolution branch, obtaining feature A5; A3, A4, and A5 pass through four cascaded multi-scale attention-guided feature fusion sub-network twos, and the output features are D3, D4, and D5; among them, the feature sizes of feature A3 and feature D3 are H / 2×W / 2; the feature sizes of feature A4 and feature D4 are H / 4×W / 4; the feature sizes of feature A5 and feature D5 are H / 8×W / 8;

[0081] The output features D3, D4, and D5 respectively pass through a CBR module to obtain features A6, A7, and A8; D5 passes through a CBR module for downsampling to generate a new resolution branch and obtain feature A9; A6, A7, A8, and A9 pass through three cascaded multi-scale attention-guided feature fusion subnets three, and the output features are D6, D7, D8, and D9; among them, the feature sizes of feature A6 and feature D6 are H / 2×W / 2, the feature sizes of feature A7 and feature D7 are H / 4×W / 4, the feature sizes of feature A8 and feature D8 are H / 8×W / 8, and the feature sizes of feature A9 and feature D9 are H / 16×W / 16.

[0082] As Figure 5 shown, the multi-scale attention-guided feature fusion subnet one includes a basic unit, a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0083] In the multi-scale attention-guided feature fusion subnet one, the input features are A1 and A2 respectively; feature A1 passes through four cascaded basic units to obtain feature B1; feature A2 passes through four cascaded basic units to obtain feature B2; among them, the feature size of feature B1 is H / 2×W / 2, and the feature size of feature B2 is H / 4×W / 4;

[0084] Feature B1 and B2 that has passed through the upsampling combination layer are added element-wise to obtain feature C1; feature B2 and B1 that has passed through the downsampling combination layer 1 are added element-wise to obtain feature C2; the feature size of feature C1 is H / 2×W / 2, and the feature size of feature C2 is H / 4×W / 4;

[0085] After feature C1 passes through the ReLU activation function, it is weighted and fused with Y1 output in step 2, and then passes through the channel attention mechanism to obtain feature D1; after feature C2 passes through the ReLU activation function, it is weighted and fused with Y2 output in step 2, and then passes through the channel attention mechanism to obtain feature D2.

[0086] As Figure 5 shown, the multi-scale attention-guided feature fusion subnet two includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0087] In the multi-scale attention-guided feature fusion sub-network two, the input features are A3, A4, and A5 respectively; feature A3 passes through eight cascaded basic units to obtain feature B3; feature A4 passes through eight cascaded basic units to obtain feature B4; feature A5 passes through eight cascaded basic units to obtain feature B5; the feature size of feature B4 is H / 4×W / 4, and the feature size of feature B5 is H / 8×W / 8;

[0088] Feature B3 is sequentially added element-wise to B4 and B5 that have passed through the upsampling combination layer to obtain feature C3; feature B4 is sequentially added element-wise to B3 that has passed through the downsampling combination layer 1 and B5 that has passed through the upsampling combination layer to obtain feature C4; feature B5 is sequentially added element-wise to B3 that has passed through the downsampling combination layer 2 and B4 that has passed through the downsampling combination layer 1 to obtain feature C5; the feature size of feature C3 is H / 2×W / 2, the feature sizes of feature B4 and feature C4 are H / 4×W / 4, and the feature sizes of feature B5 and feature C5 are H / 8×W / 8;

[0089] After feature C3 passes through the ReLU activation function, it is weighted and fused with Y1 output in step 2, and then passes through the channel attention mechanism to obtain feature D3; after feature C4 passes through the ReLU activation function, it is weighted and fused with Y2 output in step 2, and then passes through the channel attention mechanism to obtain feature D4; after feature C5 passes through the ReLU activation function, it is weighted and fused with Y3 output in step 2, and then passes through the channel attention mechanism to obtain feature D5.

[0090] As Figure 5 shown, the multi-scale attention-guided feature fusion sub-network three includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, a CBR-CBR-convolution-batch normalization combination layer, an element-wise addition module, a weighted fusion module, and a channel attention mechanism module:

[0091] In the multi-scale attention-guided feature fusion sub-network three, the input features are A6, A7, A8, and A9 respectively; feature A6 passes through eight cascaded basic units to obtain feature B6; feature A7 passes through eight cascaded basic units to obtain feature B7; feature A8 passes through eight cascaded basic units to obtain feature B8; feature A9 passes through eight cascaded basic units to obtain feature B9; the feature size of feature B6 is H / 2×W / 2, the feature size of feature B7 is H / 4×W / 4, the feature size of feature B8 is H / 8×W / 8, and the feature size of feature B9 is H / 16×W / 16;

[0092] Feature B6 is successively added element-wise to B7, B8, and B9 that have passed through the upsampling combination layer to obtain feature C6; Feature B7 is successively added element-wise to B6 that has passed through the downsampling combination layer 1, B8, and B9 that have passed through the upsampling combination layer to obtain feature C7; Feature B8 is successively added element-wise to B6 that has passed through the downsampling combination layer 2, B7 that has passed through the downsampling combination layer 1, and B9 that have passed through the upsampling combination layer to obtain feature C8; Feature B9 is successively added element-wise to B6 that has passed through the downsampling combination layer 3, B7 that has passed through the downsampling combination layer 2, and B8 that has passed through the downsampling combination layer 1 to obtain feature C9; The feature size of feature C6 is H / 2×W / 2, the feature size of feature C7 is H / 4×W / 4, the feature size of feature C8 is H / 8×W / 8, and the feature size of feature C9 is H / 16×W / 16;

[0093] After feature C6 passes through the ReLU activation function, it is weighted and fused with Y1 output in step 2, and then passes through the channel attention mechanism to obtain feature D6; After feature C7 passes through the ReLU activation function, it is weighted and fused with Y2 output in step 2, and then passes through the channel attention mechanism to obtain feature D7; After feature C8 passes through the ReLU activation function, it is weighted and fused with Y3 output in step 2, and then passes through the channel attention mechanism to obtain feature D8; After feature C9 passes through the ReLU activation function, it is weighted and fused with Y4 output in step 2, and then passes through the channel attention mechanism to obtain feature D9.

[0094] As Figure 3 shown, in the basic unit, the input feature successively passes through a combination of a CBR, a convolution, and a batch normalization, and then performs an element-wise addition operation with the input feature. The result of the addition is passed through the ReLU activation function to obtain the output of the basic unit.

[0095] As Figure 2 shown, the prediction network includes a bilinear interpolation upsampling module, a splicing module, and a detection head module:

[0096] First, the multi-scale features obtained by the multi-scale attention-guided feature fusion network are aligned to the same scale H / 2×W / 2 through bilinear interpolation upsampling, and the obtained features are Z1, Z2, Z3, and Z4. Four prediction results are obtained by passing them through four independent detection heads; Features Z1, Z2, Z3, and Z4 are spliced along the channel dimension to obtain feature Z0, and one prediction result is obtained by passing through one detection head;

[0097] Loss calculations are performed on the obtained five prediction results, and the five generated losses are fused through a multi-level aggregation loss function. The feature learning ability of the model is improved through multi-scale supervision, thereby providing a stronger supervision signal, optimizing the feature learning process, and enhancing the stability of the detection results.

[0098] As Figure 3As shown in the figure, in the detection head module of the prediction network, after the input features pass through a combination of a CBR and a convolution in sequence, the output of the detection head can be obtained.

[0099] Step 3: Network training and prediction. Use the constructed multi-resolution-channel attention key point detection network to train the labeled lateral cephalometric images. Input the newly acquired lateral cephalometric images to be detected into the trained network, automatically predict the position coordinate information of 12 key points, and at the same time, output the error between the prediction result and the manual annotation to better observe the accuracy of the prediction data. Thus, obtain the qualified rate of the test data under different evaluation metrics to better observe the accuracy of the prediction data.

[0100] The loss function of each output head uses the face key point detection evaluation metric NME. The NME of each image is defined as: the Euclidean distance between all predicted points and the labeled points, divided by the number of key points. The above loss function expression is as follows:

[0101]

[0102] Among them, P and G are the predicted value and the true value of the key point coordinates of each image respectively, p i is the predicted value of the i-th key point, g i is the true value of the i-th key point, M is the number of key points, and d is the normalization factor used to eliminate the influence of scale, which is set to 1 here.

[0103] When the multi-level aggregation loss function performs weighted fusion on the losses of each detection head: the loss of the prediction result obtained after feature concatenation is recorded as loss 0, accounting for 0.6; the losses of the four prediction results obtained before feature concatenation are recorded as loss 1, loss 2, loss 3, and loss 4, each accounting for 0.1.

[0104] In the network prediction part, while outputting the coordinates of 12 key points, the prediction result can be compared with the manually annotated result, and the error between them can be output, so as to obtain the qualified rate of the test data under different evaluation metrics to better observe the accuracy of the prediction result.

[0105] For a certain X-ray lateral cephalometric image to be detected, in practical applications, it specifically includes the following steps:

[0106] (1) The training set includes high-heterogeneity X-ray lateral cephalometric images of 336 research cases.

[0107] (2) The coordinate positions of 12 key points were marked on 336 lateral cephalometric X-ray images. The specific steps were as follows: First, the marking was carried out in the physical coordinate system with the measurement unit set to millimeters. Subsequently, using the established conversion algorithm, the position coordinates marked in the physical coordinate system were converted into the corresponding coordinates in the pixel coordinate system, thus completing the mapping from the physical space to the pixel space.

[0108] (3) Data augmentation techniques, including random scaling, rotation and other diverse processes, were applied to the 336 cases of data for training, obtaining 1680 cases of data. They were divided into a training set and a validation set according to the ratio of 0.85:0.15, and finally 1344 cases of training data and 336 cases of validation data were obtained.

[0109] (4) The initial values of the hyperparameters of the key point detection network model were set. The number of training rounds was set to 50, and the initial learning rate was set to 0.0002. The learning rate was decayed at rounds 10, 20, and 30 respectively, and the Adam optimizer was adopted.

[0110] (5) Steps (3) to (4) were repeated as the training stage of the multi-resolution-channel attention key point detection network model to obtain the final model weights.

[0111] (6) The multi-resolution-channel attention key point detection network model trained in step (5) was used to test the newly acquired lateral cephalometric X-ray images, and finally the coordinate positions of the 12 key points predicted in the newly acquired lateral cephalometric X-ray images were obtained. At the same time, the error between the predicted result and the manual marking was output to better observe the accuracy of the predicted data.

[0112] As Figure 6 shown, the final coordinate positions of the 12 key points output by the present invention are shown, and the numbers 1-12 in the figure indicate the serial numbers of the points.

[0113] The lightweight fusion feature extraction network of the present invention performs feature extraction and enhancement operations. Through lightweight design, the number of model parameters is reduced, significantly reducing the computational complexity while maintaining the effectiveness of feature extraction. The channel attention mechanism module adaptively adjusts the weights of each channel, highlighting the features of important channels and suppressing the interference of irrelevant channels. The multi-resolution progressive feature extraction sub-network extracts multi-resolution features, which are fused in the main part of the network through skip connections to achieve more effective multi-scale feature representation. The multi-scale attention-guided feature fusion network generates branches with new resolutions and repeatedly fuses the information of all current branches. The prediction network aligns the multi-scale features obtained by the multi-scale attention-guided feature fusion network to the same scale through upsampling, and splices them to fuse information at different levels. The features before splicing and the features after splicing are processed by independent detection heads to separately process features at different levels and generate prediction results, and the generalization ability of the model is supervised and improved through a multi-level aggregation loss function. In the test stage, the newly acquired lateral cephalometric image is input into the trained multi-resolution-channel attention key point detection network, and the positions of 12 key feature points in the image can be quickly and automatically located by scanning the entire lateral cephalometric image only once.

[0114] In addition to the above embodiments, the present invention may have other embodiments. Any technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.

Claims

1. A method for constructing a multi-resolution-channel attention network that can detect key points in highly heterogeneous X-ray head lateral images, characterized in that The following steps are involved: Step 1: Construct a lightweight fusion feature extraction network: The lightweight fusion feature extraction network consists of several lightweight fusion convolution-batch normalization-ReLU activation function combination layers, several bottleneck units connected in sequence, and finally a channel attention mechanism module; Step 2: Construct a multi-resolution progressive feature extraction subnetwork: The multi-resolution progressive feature extraction subnetwork uses a simplified feature pyramid network as its basic structure to perform multi-resolution feature extraction on the original input image to construct features of different resolutions, which are then fused in the main part of the network through skip connections to achieve more effective multi-scale feature expression. Step 3: Construct a multi-scale attention-guided feature fusion network: The multi-scale attention-guided feature fusion network consists of several convolution-batch normalization-ReLU activation function combination layers and several multi-scale attention-guided feature fusion sub-networks connected in sequence; Step 4: Construct a prediction network: The prediction network aligns the multi-scale features obtained by the multi-scale attention-guided feature fusion network to the same scale through upsampling, and then splices them to fuse information at different levels; then the features before and after splicing are passed through independent detection heads to process features at different levels and generate prediction results, and the generalization ability of the model is improved through multi-level aggregation loss function supervision; Step 5: Network training and prediction: Use the constructed multi-resolution-channel attention key point detection network to train the annotated head lateral images; input the newly collected head lateral images to be detected into the trained network, automatically predict the positions of key points in the image, and output the error between the predicted results and the manual annotations.

2. The method according to claim 1, characterized in that: In step 1, the constructed lightweight fusion feature extraction network includes a lightweight fusion convolution-batch normalization-ReLU activation function combination layer (LBR), a bottleneck unit and a channel attention mechanism module: First, the input image passes through a lightweight fusion convolution-batch normalization-ReLU activation function combination layer (LBR) to complete preliminary feature extraction, and the extracted feature size is H×W; then it passes through a second lightweight fusion convolution-batch normalization-ReLU activation function combination layer (LBR) to complete the downsampling operation while performing feature extraction, and the feature size becomes H / 2×W / 2; The obtained features are passed through four cascaded bottleneck units and then through a channel attention mechanism module.

3. The method according to claim 2, characterized in that: In the lightweight fused convolution, the input feature undergoes a lightweight convolution and then is concatenated with the feature that has undergone a standard convolution. The concatenated result is passed through a ReLU activation function to obtain the output of the lightweight fused convolution.

4. The method according to claim 2, characterized in that: In the bottleneck unit, the input features are sequentially passed through two convolution-batch normalization-activation function combination layers (CBR), a convolution, and a batch normalization combination, and then element-by-element addition operation is performed with the input features. The addition result is passed through the ReLU activation function to obtain the output of the bottleneck unit.

5. The method according to claim 2, characterized in that: The channel attention mechanism module captures the global information at the channel level, multiplies the obtained weights with the original features, and realizes adaptive feature recalibration at the channel level, which highlights the features of important channels, suppresses the interference of irrelevant channels, and significantly improves the model's ability to capture detailed features in skull X-ray images. The specific process is as follows: First, for the input features of size B×C×H×W, the information of spatial dimension H×W is compressed into a global channel feature vector through global average pooling, and the compressed size is B×C; Then a two-layer fully connected network is used to generate the weight value of each channel: the first layer reduces the number of channels from C to C / 16, reduces the dimension to reduce the number of parameters and calculations and learns the compression features between channels, and then uses the ReLU activation function to introduce nonlinear capabilities; The second layer restores the original number of channels C and generates a weight value for each channel, which is then normalized to [0, 1] through the Sigmoid activation function. The generated channel weights are multiplied by the input features channel by channel to enhance or suppress the channel features.

6. The method according to claim 1, characterized in that: In step 2, the multi-resolution progressive feature extraction subnetwork includes a convolution-batch normalization-ReLU activation function combination layer (CBR), a convolution-bilinear interpolation upsampling combination layer and a weighted fusion module: First, the input image dimension is H×W. After passing through a CBR module, feature X1 is obtained, and the feature size is: H / 2×W / 2; X1 passes through a CBR module to obtain feature X2, and the feature size is: H / 4×W / 4; X2 passes through a CBR module to obtain feature X3, and the feature size is: H / 8×W / 8; X3 passes through a CBR module to obtain feature X4, and the feature size is: H / 16×W / 16; After the upsampling combination layer, feature X2 is weightedly fused with feature X1 to obtain feature Y1; after the upsampling combination layer, feature X3 is weightedly fused with feature X2 to obtain feature Y2; after the upsampling combination layer, feature X4 is weightedly fused with feature X3 to obtain feature Y3; feature Y4 is directly transferred from feature X4; The final output features are Y1, Y2, Y3 and Y4, where the feature size of Y1 is H / 2×W / 2, the feature size of Y2 is H / 4×W / 4, the feature size of Y3 is H / 8×W / 8, and the feature size of Y4 is H / 16×W / 16.

7. The method according to claim 1, characterized in that: In step 3, the multi-scale attention-guided feature fusion network includes a convolution-batch normalization-ReLU activation function combination layer (CBR), a multi-scale attention-guided feature fusion sub-network 1, a multi-scale attention-guided feature fusion sub-network 2, and a multi-scale attention-guided feature fusion sub-network 3: First, the input feature passes through a CBR module, maintaining the original resolution branch while changing the current number of channels to obtain feature A1; The input features are downsampled through a CBR module to generate a new resolution branch and obtain feature A2; A1 and A2 are guided by the multi-scale attention feature fusion sub-network 1, and the output features obtained are D1 and D2; in, The feature sizes of the input feature, feature A1, and feature D1 are H / 2×W / 2; the feature sizes of feature A2 and feature D2 are H / 4×W / 4; The output features D1 and D2 are passed through a CBR module respectively to obtain features A3 and A4; D2 is downsampled through a CBR module to generate a new resolution branch to obtain feature A5; A3, A4 and A5 are passed through four cascaded multi-scale attention-guided feature fusion sub-networks 2 to obtain output features D3, D4 and D5; among which, the feature size of features A3 and D3 is H / 2×W / 2; the feature size of features A4 and D4 is H / 4×W / 4; the feature size of features A5 and D5 is H / 8×W / 8; The output features D3, D4 and D5 are respectively passed through a CBR module to obtain features A6, A7 and A8; D5 is downsampled through a CBR module to generate a new resolution branch to obtain feature A9; A6, A7, A8 and A9 are passed through three cascaded multi-scale attention-guided feature fusion sub-networks, and the output features obtained are D6, D7, D8 and D9; among which the feature sizes of features A6 and D6 are H / 2×W / 2, the feature sizes of features A7 and D7 are H / 4×W / 4, the feature sizes of features A8 and D8 are H / 8×W / 8, and the feature sizes of features A9 and D9 are H / 16×W / 16.

8. The method according to claim 7, characterized in that: The multi-scale attention-guided feature fusion subnetwork includes a basic unit, a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, an element-by-element addition module, a weighted fusion module, and a channel attention mechanism module: In the multi-scale attention-guided feature fusion subnetwork 1, the input features are A1 and A2 respectively; feature A1 passes through four cascaded basic units to obtain feature B1; Feature A2 is passed through four cascaded basic units to obtain feature B2; the feature size of feature B1 is H / 2×W / 2, and the feature size of feature B2 is H / 4×W / 4; Feature B1 is added element by element to B2 after the upsampling combination layer to obtain feature C1; feature B2 is added element by element to B1 after the downsampling combination layer 1 to obtain feature C2; ​​the feature size of feature C1 is H / 2×W / 2, and the feature size of feature C2 is H / 4×W / 4; After the feature C1 passes through the ReLU activation function, it is weightedly fused with the output Y1 of step 2, and then passes through the channel attention mechanism to obtain the feature D1; after the feature C2 passes through the ReLU activation function, it is weightedly fused with the output Y2 of step 2, and then passes through the channel attention mechanism to obtain the feature D2; The multi-scale attention-guided feature fusion subnetwork 2 includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, an element-by-element addition module, a weighted fusion module, and a channel attention mechanism module: In the multi-scale attention-guided feature fusion subnetwork 2, the input features are A3, A4 and A5 respectively; feature A3 passes through eight cascaded basic units to obtain feature B3; feature A4 passes through eight cascaded basic units to obtain feature B4; feature A5 passes through eight cascaded basic units to obtain feature B5; the feature size of feature B3 is H / 2×W / 2, the feature size of feature B4 is H / 4×W / 4, and the feature size of feature B5 is H / 8×W / 8; Feature B3 is sequentially added element-by-element with B4 and B5 after the upsampling combination layer to obtain feature C3; feature B4 is sequentially added element-by-element with B3 after the downsampling combination layer 1 and B5 after the upsampling combination layer to obtain feature C4; feature B5 is sequentially added element-by-element with B3 after the downsampling combination layer 2 and B4 after the downsampling combination layer 1 to obtain feature C5; the feature size of feature C3 is H / 2×W / 2, the feature size of feature C4 is H / 4×W / 4, and the feature size of feature C5 is H / 8×W / 8; After the feature C3 is activated by the ReLU function, it is weightedly fused with the output Y1 of step 2, and then passes through the channel attention mechanism to obtain feature D3; after the feature C4 is activated by the ReLU function, it is weightedly fused with the output Y2 of step 2, and then passes through the channel attention mechanism to obtain feature D4; after the feature C5 is activated by the ReLU function, it is weightedly fused with the output Y3 of step 2, and then passes through the channel attention mechanism to obtain feature D5; The multi-scale attention-guided feature fusion subnetwork 3 includes a convolution-bilinear interpolation upsampling combination layer, a convolution-batch normalization combination layer, a CBR-convolution-batch normalization combination layer, a CBR-CBR-convolution-batch normalization combination layer, an element-by-element addition module, a weighted fusion module, and a channel attention mechanism module: In the multi-scale attention-guided feature fusion subnetwork 3, the input features are A6, A7, A8 and A9 respectively; feature A6 passes through eight cascaded basic units to obtain feature B6; feature A7 passes through eight cascaded basic units to obtain feature B7; feature A8 passes through eight cascaded basic units to obtain feature B8; feature A9 passes through eight cascaded basic units to obtain feature B9; the feature size of feature B6 is H / 2×W / 2, the feature size of feature B7 is H / 4×W / 4, the feature size of feature B8 is H / 8×W / 8, and the feature size of feature B9 is H / 16×W / 16; Feature B6 is sequentially added element-by-element with B7, B8 and B9 that have passed through the upsampling combination layer to obtain feature C6; feature B7 is sequentially added element-by-element with B6 that has passed through the downsampling combination layer 1, B8 and B9 that have passed through the upsampling combination layer to obtain feature C7; feature B8 is sequentially added element-by-element with B6 that has passed through the downsampling combination layer 2, B7 that has passed through the downsampling combination layer 1 and B9 that has passed through the upsampling combination layer to obtain feature C8; feature B9 is sequentially added element-by-element with B6 that has passed through the downsampling combination layer 3, B7 that has passed through the downsampling combination layer 2 and B8 that has passed through the downsampling combination layer 1 to obtain feature C9. The feature size of feature C6 is H / 2×W / 2, the feature size of feature C7 is H / 4×W / 4, the feature size of feature C8 is H / 8×W / 8, and the feature size of feature C9 is H / 16×W / 16; After the feature C6 passes through the ReLU activation function, it is weightedly fused with Y1 output from step 2, and then passes through the channel attention mechanism to obtain feature D6; after the feature C7 passes through the ReLU activation function, it is weightedly fused with Y2 output from step 2, and then passes through the channel attention mechanism to obtain feature D7; after the feature C8 passes through the ReLU activation function, it is weightedly fused with Y3 output from step 2, and then passes through the channel attention mechanism to obtain feature D8; after the feature C9 passes through the ReLU activation function, it is weightedly fused with Y4 output from step 2, and then passes through the channel attention mechanism to obtain feature D9; In the basic unit, the input features are sequentially subjected to a combination of a CBR, a convolution, and a batch normalization, and then element-by-element addition operation is performed with the input features. The addition result is passed through a ReLU activation function to obtain the output of the basic unit.

9. The method according to claim 1, characterized in that: In step 4, the prediction network includes a bilinear interpolation upsampling module, a splicing module and a detection head module: First, the multi-scale features obtained by the multi-scale attention-guided feature fusion network are aligned to the same scale H / 2×W / 2 through bilinear interpolation upsampling. The obtained features are Z1, Z2, Z3 and Z4, which are passed through four independent detection heads to obtain four prediction results; the features Z1, Z2, Z3 and Z4 are concatenated along the channel dimension to obtain feature Z0, which is passed through one detection head to obtain one prediction result; The loss is calculated for the five prediction results, and the five losses generated are fused through a multi-level aggregation loss function. The feature learning ability of the model is improved through multi-scale supervision, thereby providing a stronger supervision signal, optimizing the feature learning process, and enhancing the stability of the detection results. In the detection head module, the input features are sequentially combined by a CBR and a convolution to obtain the output of the detection head.

10. The method according to claim 1, characterized in that: In step 5, the network prediction part outputs the coordinate positions of the key points in the image, compares the prediction results with the manually labeled results, and outputs the error between them, thereby obtaining the pass rate of the test data under different evaluation indicators, so as to better observe the accuracy of the prediction results.

Citation Information

Patent Citations

  • Lightweight face detection method and system based on mixed attention feature pyramid structure

    CN113591795A

  • 2D attitude detection method based on depth separable convolution and channel attention

    CN117558061A

  • Indoor article detection method based on YOLOv7

    CN117789012A

  • Medical image lesion area automatic segmentation method based on improved lightweight convolutional network, medium and equipment

    CN118781347A

  • Method and device for detecting positional relationship between molar root and neural tube

    CN119228764A

Cited By

  • Chemical experiment device identification method, device and equipment based on multilayer feature fusion

    CN121904723A