A nasopharyngeal lesion detection method based on dynamic feature fusion and difference perception
Patent Information
- Application Number
- CN202611022416.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
然而,鼻咽内镜图像中病灶区域往往呈现形态多样、边界模糊、与周围正常组织对比度低等特点,且不同患者、不同病变阶段的病灶在大小、形状、纹理上存在显著差异,仅依靠医师肉眼观察容易出现漏诊和误诊
[0048]1.本发明设计了一种动态特征融合与差异感知的鼻咽部病灶检测方法的目标检测网络,为了适应鼻咽癌目标的边界模糊以及尺度差异显著的特点,在网络中加入了TFDGM模块,该模块通过多尺度卷积结构获取不同感受野下的特征信息,该方式能够提升模型对边缘模糊或尺度变化较大的病灶的感知能力;并在P3、P4检测头前加入DRFM模块通过构建语义特征与结构特征,对病灶区域的高层语义信息和局部纹理结构信息进行协同建模,并利用两类特征之间的差异响应生成引导权重,实现检测特征的动态重分配;
Smart Images

Figure CN122820643A_ABST
Abstract
Description
[0001] This invention relates to the fields of deep learning and computer vision (lesion detection), specifically to a method for detecting nasopharyngeal lesions based on dynamic feature fusion and differential perception, which is an improvement on the YOLO26 network. Background Technology
[0002] Nasopharyngeal carcinoma (NPC) is a malignant tumor originating from the nasopharyngeal mucosal epithelium, with a high incidence rate in southern my country and Southeast Asian countries. Electronic nasopharyngeal endoscopy, due to its non-invasive, intuitive, and repeatable advantages, has become the preferred method for early screening and clinical diagnosis of NPC. However, lesions in nasopharyngeal endoscopy images often exhibit diverse morphologies, blurred boundaries, and low contrast with surrounding normal tissues. Furthermore, lesions vary significantly in size, shape, and texture among different patients and at different stages of the disease, making it easy to miss or misdiagnose NPC relying solely on visual observation. Therefore, researching automatic detection methods for nasopharyngeal endoscopy images based on computer vision and deep learning is of great significance for improving the accuracy and efficiency of early diagnosis of NPC. In recent years, deep learning technology, especially convolutional neural networks (CNNs), has achieved widespread success in the field of medical image analysis. Many researchers have applied general object detection frameworks (SSD, YOLO series, etc.) to lesion detection tasks in nasopharyngeal endoscopy images, achieving some progress. However, lesions in nasopharyngeal endoscopy images often vary drastically in scale: early, tiny lesions may only occupy a few dozen pixels, while advanced lesions can cover a large area of mucosa.
[0003] To address the challenges of large scale variations, irregular morphology, and low contrast with normal mucosa in nasopharyngeal endoscopic images, a TFDGM (Tri-frequency Dynamic Guidance Module) feature fusion module is proposed. This module jointly models high-frequency, low-frequency, and edge difference information to enhance key lesion features and fuse multi-scale information. Simultaneously, a DFRM (Difference-guided Feature Reallocation Module) module is introduced before the P3 and P4 detection heads. Through a difference-guided feature reallocation mechanism, it adaptively redistributes detection features, enhancing the expression of lesion edges and fine-grained features. Summary of the Invention
[0004] The purpose of this invention is to design a network for detecting lesion regions in nasopharyngeal endoscopic images, which can detect lesions with large scale changes in nasopharyngeal endoscopic images and has high detection accuracy.
[0005] This invention provides a method for detecting nasopharyngeal lesions based on dynamic feature fusion and differential perception. It mainly includes adding a dual-branch attention-guided and multi-scale frequency feature enhancement module (TFDGM) to the neck network. This module selects and fuses features extracted from the original image by the backbone network in the neck network. Simultaneously, through multi-scale feature processing and attention adjustment, the module guides the network to focus more on potential lesion areas. This method enhances the expression of details and contextual information without changing the feature resolution. Furthermore, a DFRM module is introduced before the P3 and P4 detection heads to provide more stable and discriminative input features for subsequent detection heads. The detection heads are used to output the final detection results.
[0006] This invention mainly consists of the following steps:
[0007] 1. The detection dataset is constructed and labeled by professional doctors, and the dataset is divided into training set, validation set, and test set;
[0008] 2. Construct an object detection network based on the YOLO series of networks, currently using YOLO26 as an example network. The object detection network includes a backbone network, a neck network, and a detection head. The backbone network extracts spatial domain features of different scales through multi-layer convolution and downsampling operations.
[0009] 3. Using the YOLO26 network model as the base network, the network was optimized based on the complex and diverse characteristics of nasopharyngeal carcinoma target detection images. The constructed nasopharyngeal carcinoma target detection network mainly includes a YOLO26 backbone network based on C3K2, a dual-branch attention-guided and multi-scale frequency feature enhancement module (TFDGM), which includes a three-frequency dynamic modeling (TFDM) submodule and a dynamic weight selection (DSW) submodule. The data flow of the entire network is as follows: first, the output features of the backbone network are sent to the neck, and the high-level semantic features (P5) and mid-level detail features (P4) after upsampling in the neck are fused. Since the fused features contain both target semantic information and spatial structural information, the fused features are further enhanced to highlight the lesion edge and texture features. The enhanced features are then passed to the subsequent detection branches, and a differential guided feature redistribution module (DFRM) is used before Decet_P3 and Decet_P4 to improve the detection accuracy of lesion targets.
[0010] 4. Train the completed target detection model on the created vocal cord laryngoscope image dataset, and save the weight file of the best-performing trained model;
[0011] 5. During the inference phase, the optimal weight file is called into the object detection model to achieve automated detection of video streams or static images.
[0012] In the above steps, the constructed nasopharyngeal endoscopy image target detection network mainly includes a backbone network, a neck network structure with added bi-branch attention guidance and multi-scale frequency feature enhancement, a differential guidance feature redistribution module, and a detection head.
[0013] The backbone network is built based on the C3k2 module, which has strong feature extraction capabilities. The output feature maps of different levels of the backbone network are selected for subsequent networks to continue extracting lesion features.
[0014] After the first C3K2 module of the neck network, i.e., after the fusion output of the P5 and P4 layers of the backbone network, a dual-branch attention-guided and multi-scale frequency feature enhancement module is added to enhance the model's ability to extract multi-scale features. The construction of the dual-branch attention-guided and multi-scale frequency feature enhancement module includes the following steps:
[0015] Step 1: Since the upsampled high-level semantic features (P5) and mid-level detail features (P4) are fused, the fused features contain both target semantic information and spatial structure information. Therefore, the original feature map output by this layer is obtained first. It is divided into three branches, that is, the number of channels is divided into And respectively processed by convolution kernels as The expansion rates are respectively Deep convolutions are used to extract local texture information and contextual semantic information from different receptive fields. The outputs of the three branches are then concatenated to obtain multi-scale fused features.
[0016] Step 2: Input the original features Global max pooling and global average pooling are used to extract significant response information and global statistical information, respectively. The two pooling features are then concatenated, and attention weights are generated using a 1×1 convolution and a sigmoid activation function. Finally, these weights are multiplied element-wise with the multi-scale features from step 1 to obtain the final result. This enhances the characteristics of key lesion areas and suppresses background interference;
[0017] Step 3: Transfer the feature map obtained in Step 2 pass Convolution restores the channel dimension, and the residual is added to the original input features to obtain the output. This enhances feature representation capabilities while ensuring training stability.
[0018] Step 4: Apply the enhanced feature map obtained in Step 3 The data is fed into two parallel branches for processing, with one context branch using two serial expansion rates respectively. and of Deep convolutions acquire semantic information with a large receptive field, thus obtaining contextual features. Another detailed branch is through serial Depth convolution and GSConv extract local edge and texture information to obtain detailed features. ;
[0019] Step 5: Apply the enhanced feature map obtained in Step 3 The input is fed into the TFDM (Tri-frequency Dynamic Modeling) module and processed within the TFDM. After the transformation, two branches, one high-frequency and one low-frequency, are separated. Both branches pass through... Transformation. The high-frequency branch passes through a convolution kernel of size [missing value]. Depthwise convolution and a convolution kernel size of The depthwise convolution is then passed through CA to finally output the high-frequency lesion response. The low-frequency branch first passes through a... Depthwise convolution and a convolution kernel size of The expansion rate is The depthwise convolution is then processed. Adjusting the output low-frequency semantic response Since lesion boundary regions are often accompanied by significant local structural abrupt changes, the high-frequency branch effectively responds to texture and edge details, while the low-frequency branch focuses more on global smooth semantic information. Therefore, there are often significant structural differences between high- and low-frequency responses in the boundary region. Based on this edge difference, the high-frequency lesion response output described above is used in the branch. and low-frequency semantic response To achieve this, by using Edge feature modeling is obtained, and then two serial dilation rates are respectively... and of Deep convolution is used to enhance edge feature responses and ultimately generate edge difference responses. ;
[0020] Step 6: The product generated in Step 5 The responses from the three lesions are input into the DSW (Dynamic Selective Weighting) module, where they are first processed through a... The depthwise convolution is then compressed and the computational cost is reduced through BN normalization and SiLU activation, followed by... After convolution adjustment, dynamic competitive fusion is achieved through Softmax, ultimately outputting three dynamic competitive coefficients. ;
[0021] Step 7: Assign the three dynamic coefficients generated in Step 6. For the context feature branch, since this branch focuses more on low-frequency semantic information, use the low-frequency competing coefficients to generate the context guiding factor. Context branch features Enhanced guidance, including For the other detailed branch, since the lesion region relies more on high-frequency and edge information, the detailed guiding factor is generated using both high-frequency competition coefficients and edge competition coefficients. And detailed branch features Enhanced guidance, including Ultimately, this enhances edge and high-frequency information. For context branches, first use... and The features are then fused and enhanced using residual connections, ultimately generating enhanced contextual features. Its calculation formula is ,in It is a hyperparameter that can be learned by the network, initially set to... For detailed branches, first use and The data is then fused and further enhanced using residual connections to generate enhanced detail features. Its calculation formula is ,in It is also a hyperparameter that can be learned by the network, initially set to The setting of hyperparameters and their initial values is to avoid excessive interference from prior lesions on the original features, and an enhancement coefficient is introduced. and The intensity of lesion guidance is adaptively adjusted. For the enhancement coefficient, a larger coefficient can enhance the response of the lesion area, while a smaller coefficient helps to preserve the original contextual semantic information, thereby achieving a balance between lesion enhancement and background preservation.
[0022] Step 8: Utilize the enhanced contextual features obtained in Step 7 With detailed features Splicing and then... After convolutional fusion, a channel attention mechanism (SE) is introduced. Channel weights are generated through global average pooling and fully connected layers to enhance key channels, ultimately yielding the output features. .
[0023] Output The deep semantic features extracted from the backbone network are used as subsequent differential-guided feature redistribution modules to improve the detection accuracy of lesion targets. The steps for constructing the differential-guided feature redistribution module are as follows:
[0024] Step 1: Input feature map (The output of the fusion of deep semantic features and shallow detail features extracted from the backbone network by the TFDGM module) is fed into the semantic feature extraction branch and the structural feature extraction branch, respectively. The semantic branch uses... Convolution performs channel compression and linear mapping to obtain preliminary semantic features. These features are then processed non-linearly using the GELU activation function, and then... Convolution is used to restore channel dimensions in order to extract high-level semantic information of the lesion region; structural branches are employed. Depth convolution, Depth convolution and Deep convolutions are used to construct an orientation-aware feature encoder to enhance the representation of lesion edges and local texture structure features, thereby obtaining semantic features. and structural features ;
[0025] Step 2: Process the semantic features obtained in Step 1 and structural features To perform differential modeling, a differential response map is obtained by calculating the element-wise squared difference between the two, and its expression is as follows: Then, average aggregation is performed along the channel dimension, i.e. This yields an initial differential feature map that characterizes the degree of difference between the lesion area and the background area. And the initial difference feature map Through a Enhanced differential feature maps are obtained by performing local differential convolutions using depthwise convolution. ;
[0026] Step 3: The enhanced differential feature map obtained in Step 2... The final differential guided weight map is generated using the Sigmoid activation function. ,in The value is close to This indicates that the lesion area is of high importance and is close to... Indicates the background area;
[0027] Step 4: Apply the difference-guided weighting graph obtained in Step 3. Semantic features used to guide the acquisition in step 1 and structural features Dynamic reallocation is performed to enhance semantic features; the calculation formula is as follows: The structural features retain structural information, and the calculation formula is as follows: That is, the lesion area through Enhanced, background area through Preserve structural information to achieve adaptive feature redistribution;
[0028] Step 5: Reassign the semantic features obtained in Step 4 With structural features Concatenate along the channel dimension, then pass through Convolutional feature fusion, and finally combined with input features Perform residual connections to output the enhanced feature map. .
[0029] Output the enhanced feature map This module serves as input to object detection heads or classification networks to perform the task of detecting and identifying nasopharyngeal lesions. By integrating multi-scale contextual information, local detail information, and a lesion region guidance mechanism, the entire module can effectively improve the model's detection accuracy and recall for small target lesions, lesions with blurred edges, and complex background regions.
[0030] The object detection module employs a decoupled detection head structure, consisting of a classification branch, a bounding box regression branch, a one-to-many branch, and a one-to-one branch. YOLO26 uses a dual-branch detection head structure, including a one-to-many branch and a one-to-one branch. The one-to-many branch provides rich positive sample supervision information, while the one-to-one branch is used for end-to-end object detection inference. Unlike YOLOv8 and YOLO11, YOLO26 eliminates the Direct Bounding Box Regression (DFL) and uses direct bounding box regression to predict object locations, thereby reducing the number of detection head parameters and computational complexity. The loss function for the classification branch is the BCE loss function. :
[0031]
[0032] in Predicting class probabilities Indicates the true label, The number of positive samples;
[0033] In YOLOv26, the bounding box regression branch removed the DFL module, no longer using a discrete probability distribution to predict bounding boxes, but instead directly regressing the target bounding box coordinates. The bounding box loss... Depend on and Common components:
[0034]
[0035] in This represents the intersection-union ratio (IoU) between the predicted bounding box and the ground truth bounding box. For the true bounding box, To predict the bounding box, This is the balance coefficient;
[0036] The One-to-Many branch employs a TAL label assignment strategy, using a large number of Top-K candidate samples for supervised learning, and its loss function is... Represented as:
[0037]
[0038] Classification loss is used Regression loss usage and ;
[0039] The One-to-One branch employs a strict one-to-one label matching mechanism, where each target corresponds to only one predicted bounding box. Its loss function is... Represented as:
[0040]
[0041] Similar to the One-to-Many branch, classification loss is used. Regression loss usage and ;
[0042] To address the issue of unbalanced supervision intensity during two-branch training, YOLOv26 proposes a Progressive Loss dynamic weight strategy. In the early stages of training, the focus is on optimizing the One-to-Many branch, while in the later stages, the focus gradually shifts to the One-to-One branch, with its weights adjusted accordingly. The update formula is:
[0043]
[0044] in For the current training round, For the total number of training rounds, As the initial weights, For the final weight;
[0045] Total loss function It can be represented as:
[0046]
[0047] By adopting the above technical solution, the present invention has the following advantages:
[0048] 1. This invention designs a target detection network for a dynamic feature fusion and difference perception method for nasopharyngeal lesion detection. To adapt to the characteristics of blurred boundaries and significant scale differences in nasopharyngeal carcinoma targets, a TFDGM module is added to the network. This module obtains feature information under different receptive fields through multi-scale convolutional structures, which can improve the model's ability to perceive lesions with blurred edges or large scale changes. Furthermore, a DRFM module is added before the P3 and P4 detection heads to construct semantic features and structural features, and to collaboratively model the high-level semantic information and local texture structure information of the lesion region. The difference response between the two types of features is used to generate guiding weights to achieve dynamic redistribution of detection features.
[0049] 2. The TFDGM module introduces a three-frequency dynamic modeling mechanism. The module separates features into high-frequency and low-frequency branches using FFT. The high-frequency branch primarily focuses on texture, edges, and detail regions, while the low-frequency branch focuses on the overall semantic structure. Furthermore, it utilizes the difference between high and low-frequency responses to construct an edge difference branch, thereby enhancing the expressive power of lesion boundary regions.
[0050] 3. The TFDGM module uses a DSW dynamic selection weighting mechanism to achieve adaptive competition between responses from different lesions. The module dynamically generates weight coefficients for high-frequency, low-frequency, and marginal branches using Softmax, allowing the model to automatically adjust its focus based on the feature responses of different regions, thus exhibiting better flexibility and adaptability.
[0051] 4. The DRFM module accurately characterizes the feature differences between the lesion area and the background area, enabling the network to focus on the lesion area with strong differential response, thereby effectively enhancing the ability to express lesion edges, texture details and fine-grained semantic features. Attached Figure Description
[0052] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:
[0053] Figure 1 This is a schematic diagram of the nasopharyngeal lesion detection method based on dynamic feature fusion and differential perception of the present invention;
[0054] Figure 2 This is a schematic diagram of the TFDGM module of the present invention;
[0055] Figure 3 This is a schematic diagram of the TFDM module of the present invention;
[0056] Figure 4 This is a schematic diagram of the DSW module of the present invention;
[0057] Figure 5 This is a schematic diagram of the DFRM module of the present invention;
[0058] Figure 6 This is a schematic diagram of the target detection network for the nasopharyngeal lesion detection method based on dynamic feature fusion and differential perception of the present invention. Detailed Implementation Plan
[0059] The present invention will be further described in detail below with reference to specific implementations. The description is for explanation and not limitation of the present invention.
[0060] This invention proposes a method for detecting nasopharyngeal lesions based on dynamic feature fusion and differential perception, such as... Figure 1 As shown, the specific steps include:
[0061] 1. The detection dataset is constructed and labeled by professional doctors, and the dataset is divided into training set, validation set, and test set;
[0062] 2. Construct an object detection network based on the YOLO series of networks, currently using YOLO26 as an example network. The object detection network includes a backbone network, a neck network, and a detection head. The backbone network extracts spatial domain features of different scales through multi-layer convolution and downsampling operations.
[0063] 3. Using the YOLO26 network model as the base network, the network was optimized according to the complex and diverse characteristics of nasopharyngeal carcinoma target detection images. The constructed nasopharyngeal carcinoma target detection network mainly includes a YOLO26 backbone network based on C3K2, and a dual-branch attention-guided and multi-scale frequency feature enhancement module (TFDGM). The data flow of the entire network is as follows: first, the output features of the backbone network are sent to the neck, and the high-level semantic features (P5) and mid-level detail features (P4) after upsampling in the neck are fused. Since the fused features contain both target semantic information and spatial structural information, the fused features are further enhanced to highlight the lesion edge and texture features. The enhanced features are then passed to the subsequent detection branches, and a differential guided feature redistribution module (DFRM) is used before Decet_P3 and Decet_P4 to improve the detection accuracy of lesion targets.
[0064] 4. Train the completed target detection model on the created vocal cord laryngoscope image dataset, and save the weight file of the best-performing trained model;
[0065] 5. During the inference phase, the optimal weight file is called into the object detection model to achieve automated detection of video streams or static images.
[0066] The specific implementation method provides the steps for constructing a target detection network for a dynamic feature fusion and differential perception method for nasopharyngeal lesion detection. Specifically, it includes: a backbone network, a neck network structure with TFDGM dual-branch attention guidance and multi-scale frequency feature enhancement modules, a differential guidance feature redistribution module, and a target detection module.
[0067] Step 1: A professional doctor constructs and labels the detection dataset, and divides the dataset into a training set, a validation set, and a test set;
[0068] Step 2: Construct an object detection network based on the YOLO series of networks, currently using YOLO26 as an example network. The object detection network includes a backbone network, a neck network, and a detection head. The backbone network extracts spatial domain features at different scales through multi-layer convolution and downsampling operations.
[0069] Step 3: Introduce the TFDGM module into the YOLO26 neck network. Perform dual-branch attention guidance and multi-scale frequency feature enhancement on the original features output from the backbone network. Place the TFDGM module after the fusion of upsampled high-level semantic features (P5) and mid-level detail features (P4). Since the fused features simultaneously contain target semantic and spatial structural information, further enhance the fused features to highlight lesion edges and texture features. The enhanced features are then passed to subsequent detection branches to improve the detection accuracy of lesions. The steps for constructing the dual-branch attention guidance and multi-scale frequency feature enhancement module are as follows:
[0070] Step 3-1: Since the upsampled high-level semantic features (P5) and mid-level detail features (P4) have been fused, the fused features contain both target semantic information and spatial structure information. Therefore, the original feature map output by this layer is obtained first. It is divided into three branches, that is, the number of channels is divided into And respectively processed by convolution kernels as The expansion rates are respectively Deep convolutions are used to extract local texture information and contextual semantic information from different receptive fields. The outputs of the three branches are then concatenated to obtain multi-scale fused features.
[0071] Step 3-2: Input the original features Global max pooling and global average pooling are used to extract significant response information and global statistical information, respectively. The two pooling features are then concatenated, and attention weights are generated using a 1×1 convolution and a sigmoid activation function. Finally, these weights are multiplied element-wise with the multi-scale features from step 3-1 to obtain the final value. This enhances the characteristics of key lesion areas and suppresses background interference;
[0072] Step 3-3: Convert the feature map obtained in step 3-2 into a single image. pass Convolution restores the channel dimension, and the residual is added to the original input features to obtain the output. This enhances feature representation capabilities while ensuring training stability.
[0073] Step 3-4: The enhanced feature map obtained in Step 3-3... The data is fed into two parallel branches for processing, with one context branch using two serial expansion rates respectively. and of Deep convolutions acquire semantic information with a large receptive field, thus obtaining contextual features. Another detailed branch is through serial Depth convolution and GSConv extract local edge and texture information to obtain detailed features. ;
[0074] Step 3-5: The enhanced feature map obtained in Step 3-3... The input is fed into the TFDM (Tri-frequency Dynamic Modeling) module and processed within the TFDM. After the transformation, two branches, one high-frequency and one low-frequency, are separated. Both branches pass through... Transformation. The high-frequency branch passes through a convolution kernel of size [missing value]. Depthwise convolution and a convolution kernel size of The depthwise convolution is then passed through CA to finally output the high-frequency lesion response. The low-frequency branch first passes through a... Depthwise convolution and a convolution kernel size of The expansion rate is The depthwise convolution is then processed. Adjusting the output low-frequency semantic response Since lesion boundary regions are often accompanied by significant local structural abrupt changes, the high-frequency branch effectively responds to texture and edge details, while the low-frequency branch focuses more on global smooth semantic information. Therefore, there are often significant structural differences between high- and low-frequency responses in the boundary region. Based on this edge difference, the high-frequency lesion response output described above is used in the branch. and low-frequency semantic response To achieve this, by using Edge feature modeling is obtained, and then two serial dilation rates are respectively... and of Deep convolution is used to enhance edge feature responses and ultimately generate edge difference responses. ;
[0075] Step 3-6: The product generated in step 3-5 The responses from the three lesions are input into the DSW (Dynamic Selective Weighting) module, where they are first processed through a... The depthwise convolution is then compressed and the computational cost is reduced through BN normalization and SiLU activation, followed by... After adjusting the output dimension through convolution, dynamic competitive fusion is achieved using Softmax, ultimately outputting three dynamic competitive coefficients. ;
[0076] Steps 3-7: Assign the three dynamic coefficients generated in Step 3-6. For the context feature branch, since this branch focuses more on low-frequency semantic information, low-frequency competing coefficients are used to generate the context guiding factor. Context branch features Enhanced guidance, including For the other detailed branch, since the lesion region relies more on high-frequency and edge information, the detailed guiding factor is generated using both high-frequency competition coefficients and edge competition coefficients. And detailed branch features Enhanced guidance, including Ultimately, this enhances edge and high-frequency information. For context branches, first use... and The features are then fused and enhanced using residual connections, ultimately generating enhanced contextual features. Its calculation formula is ,in It is a hyperparameter that can be learned by the network, initially set to... For detailed branches, first use and The data is then fused and further enhanced using residual connections to generate enhanced detail features. Its calculation formula is ,in It is also a hyperparameter that can be learned by the network, initially set to The setting of hyperparameters and their initial values is to avoid excessive interference from prior lesions on the original features, and an enhancement coefficient is introduced. and The intensity of lesion guidance is adaptively adjusted. For the enhancement coefficient, a larger coefficient can enhance the response of the lesion area, while a smaller coefficient helps to preserve the original contextual semantic information, thereby achieving a balance between lesion enhancement and background preservation.
[0077] Step 3-8: Obtain the enhanced contextual features from Step 3-7 With detailed features Splicing and then... After convolutional fusion, a channel attention mechanism (SE) is introduced. Channel weights are generated through global average pooling and fully connected layers to enhance key channels, ultimately yielding the output features. ;
[0078] Step 4: A differentially guided feature redistribution module (DFRM) is used before Decet_P3 and Decet_P4 to improve the detection accuracy of lesion targets. The steps for constructing the differentially guided feature redistribution module are as follows:
[0079] Step 4-1: Input feature map (The output of the fusion of deep semantic features and shallow detail features extracted from the backbone network by the TFDGM module in step 3) is fed into the semantic feature extraction branch and the structural feature extraction branch, respectively. The semantic branch uses... Convolution performs channel compression and linear mapping to obtain preliminary semantic features. These features are then processed non-linearly using the GELU activation function, and then... Convolution is used to restore channel dimensions in order to extract high-level semantic information of the lesion region; structural branches are employed. Depth convolution, Depth convolution and Deep convolutions are used to construct an orientation-aware feature encoder to enhance the representation of lesion edges and local texture structure features, thereby obtaining semantic features. and structural features ;
[0080] Step 4-2: Apply the semantic features obtained in Step 4-1 and structural features To perform differential modeling, a differential response map is obtained by calculating the element-wise squared difference between the two, and its expression is as follows: Then, average aggregation is performed along the channel dimension, i.e. This yields an initial differential feature map that characterizes the degree of difference between the lesion area and the background area. And the initial difference feature map Through a Enhanced differential feature maps are obtained by performing local differential convolutions using depthwise convolution. ;
[0081] Step 4-3: The enhanced differential feature map obtained in Step 4-2... The final differential guided weight map is generated using the Sigmoid activation function. ,in The value is close to This indicates that the lesion area is of high importance and is close to... Indicates the background area;
[0082] Step 4-4: Apply the difference-guided weighting graph obtained in Step 4-3 Used to guide the semantic features obtained in step 4-1 and structural features Dynamic reallocation is performed to enhance semantic features; the calculation formula is as follows: The structural features retain structural information, and the calculation formula is as follows: That is, the lesion area through Enhanced, background area through Preserve structural information to achieve adaptive feature redistribution;
[0083] Step 4-5: Reassign the semantic features obtained in Step 4-4 With structural features Concatenate along the channel dimension, then pass through Convolutional feature fusion, and finally combined with input features Perform residual connections to output the enhanced feature map. ;
[0084] Step 5: Enhance the feature map output from Step 4. Input to the detection head;
[0085] Step 6: The object detection module adopts a decoupled detection head structure, consisting of a classification branch, a bounding box regression branch, a one-to-many branch, and a one-to-one branch. YOLO26 uses a dual-branch detection head structure, including a one-to-many branch and a one-to-one branch; the one-to-many branch is responsible for providing rich positive sample supervision information, while the one-to-one branch is used to implement end-to-end object detection inference; unlike YOLOv8 and YOLO11, YOLO26 eliminates DFL and uses direct bounding box regression to predict the target location, thereby reducing the number of detection head parameters and computational complexity; the loss function of the classification branch is the BCE loss function. :
[0086]
[0087] in Predicting class probabilities Indicates the true label, The number of positive samples;
[0088] In YOLOv26, the bounding box regression branch removed the DFL module, no longer using a discrete probability distribution to predict bounding boxes, but instead directly regressing the target bounding box coordinates. The bounding box loss... Depend on and Common components:
[0089]
[0090] in This represents the intersection-union ratio (IoU) between the predicted bounding box and the ground truth bounding box. For the true bounding box, To predict the bounding box, This is the balance coefficient;
[0091] The One-to-Many branch employs a TAL label assignment strategy, using a large number of Top-K candidate samples for supervised learning, and its loss function is... Represented as:
[0092]
[0093] Classification loss is used Regression loss usage and ;
[0094] The One-to-One branch employs a strict one-to-one label matching mechanism, where each target corresponds to only one predicted bounding box. Its loss function is... Represented as:
[0095]
[0096] Similar to the One-to-Many branch, classification loss is used. Regression loss usage and ;
[0097] To address the issue of unbalanced supervision intensity during two-branch training, YOLOv26 proposes a Progressive Loss dynamic weight strategy. In the early stages of training, the focus is on optimizing the One-to-Many branch, while in the later stages, the focus gradually shifts to the One-to-One branch, with its weights adjusted accordingly. The update formula is:
[0098]
[0099] in For the current training round, For the total number of training rounds, As the initial weights, For the final weight;
[0100] Total loss function It can be represented as:
[0101]
[0102] Step 7: Use the prepared training set data and corresponding object detection labels to train and construct an object detection network for a dynamic feature fusion and differential perception method for nasopharyngeal lesion detection;
[0103] Step 8: Test the trained network model using the prepared test set data and corresponding object detection labels.
[0104] To verify the effectiveness of the improved algorithm proposed in this invention for nasopharyngeal carcinoma target detection, the data used were annotated by physicians with extensive clinical experience. Images of nasopharyngeal carcinoma, including training, validation, and test sets, are arranged according to... The experimental platform used Ubuntu 16.04, PyCharm Professional Edition 2020.1 x64 for Python development, and PyTorch and Anaconda3 for deep learning applications. Model training employed parallel training with four NVIDIA GeForce GTX 1080 Ti graphics cards, each with 11 GB of VRAM. Multi-GPU acceleration was achieved through Distributed Data Parallel (DDP) to improve training efficiency and ensure the stability of experimental results. The server configuration included an Intel Xeon processor and 128 GB of RAM, sufficient to meet the training requirements of the proposed improved algorithm's object detection model. This invention uses mAP@0.5, a commonly used metric in object detection, and its method has shown promise on the YOLO26s benchmark model. Improvement in mAP.
Claims
1. A method for detecting nasopharyngeal lesions based on dynamic feature fusion and differential perception, including feature construction. The steps are as follows: Step 1: First, obtain the nasopharyngeal carcinoma target detection dataset annotated by professional physicians; Step 2: Using the YOLO26 network model as the base network, the network is optimized according to the complex and diverse characteristics of nasopharyngeal carcinoma target detection images. The constructed nasopharyngeal carcinoma target detection network mainly includes a YOLO26 backbone network based on C3K2, a dual-branch attention-guided and multi-scale frequency feature enhancement module (TFDGM), which includes a tri-frequency dynamic modeling (TFDM) sub-module and a dynamic selective weighting (DSW) sub-module. The data flow of the entire network is as follows: first, the output features of the backbone network are sent to the neck, and a TFDGM module is added to the neck and placed after the high-level semantic features P5 and the mid-level detail features P4 are fused after upsampling. The enhanced features are then passed to the subsequent detection branches, and a differential guided feature redistribution module (DFRM) is passed before Decet_P3 and Decet_P4 to improve the detection accuracy of lesion targets. Step 3: Train the completed detection model based on dual-branch attention guidance and multi-scale frequency feature enhancement on the corresponding dataset for several rounds, and adjust the parameters to obtain the optimal model; Step 4: Use the obtained optimal model for subsequent detection tasks.
2. The method for detecting nasopharyngeal lesions guided by three-frequency dynamic feature fusion and differential perception according to claim 1, characterized in that... The construction steps of the dual-branch attention-guided and multi-scale frequency feature enhancement module are as follows: Step 1: First, obtain the fused feature map of P5 layer and P4 layer. It is divided into three branches, that is, the number of channels is divided into And respectively processed by convolution kernels as The expansion rates are respectively The deep convolution is used to extract local texture information and contextual semantic information under different receptive fields; the outputs of the three branches are concatenated to obtain multi-scale fusion features; Step 2: Input the original features Global max pooling and global average pooling are used to extract significant response information and global statistical information, respectively. The two pooling features are then concatenated, and attention weights are generated using a 1×1 convolution and a sigmoid activation function. Finally, these weights are multiplied element-wise with the multi-scale features from step 1 to obtain the final value. This enhances the characteristics of key lesion areas and suppresses background interference; Step 3: Transfer the feature map obtained in Step 2 pass Convolution restores the channel dimension, and the residual is added to the original input features to obtain the output. This enhances feature representation capabilities while ensuring training stability. Step 4: Apply the enhanced feature map obtained in Step 3 The data is fed into two parallel branches for processing, with one context branch using two serial expansion rates respectively. and of Deep convolutions acquire semantic information with a large receptive field, thus obtaining contextual features. Another detailed branch is through serial Depth convolution and GSConv extract local edge and texture information to obtain detailed features. ; Step 5: Apply the enhanced feature map obtained in Step 3 Input is fed into the TFDM module and processed within the TFDM module. After the transformation, two branches, one high-frequency and one low-frequency, are separated. Both branches pass through... Transformation. The high-frequency branch passes through a convolution kernel of size [missing value]. Depthwise convolution and a convolution kernel size of The depthwise convolution is then passed through CA to finally output the high-frequency lesion response. The low-frequency branch first passes through a... Depthwise convolution and a convolution kernel size of The expansion rate is The depthwise convolution is then processed. Adjusting the output low-frequency semantic response The high-frequency branch effectively responds to texture and edge details, while the low-frequency branch focuses more on global smooth semantic information. There are often significant structural differences between the high-frequency and low-frequency responses. Based on this, the edge difference branch uses the high-frequency lesion response output mentioned above. and low-frequency semantic response To achieve this, by using Edge feature modeling is obtained, and then two serial dilation rates are respectively... and of Deep convolution is used to enhance edge feature responses and ultimately generate edge difference responses. ; Step 6: The product generated in Step 5 The responses from the three lesions are input into the DSW (Dynamic Selective Weighting) module, where they are first processed through a... The depthwise convolution is then compressed and the computational cost is reduced through BN normalization and SiLU activation, followed by... After convolution adjustment, dynamic competitive fusion is achieved through Softmax, ultimately outputting three dynamic competitive coefficients. ; Step 7: Assign the three dynamic coefficients generated in Step 6. For the context feature branch, since this branch focuses more on low-frequency semantic information, use the low-frequency competing coefficients to generate the context guiding factor. Context branch features Enhanced guidance, including For the other detailed branch, since the lesion area relies more on high-frequency and edge information, the high-frequency competition coefficient and the edge competition coefficient are used together to generate the detailed guiding factor. And detailed branch features Enhanced guidance, including Ultimately, this enhances edge and high-frequency information. For context branches, first use... and The features are then fused and enhanced using residual connections, ultimately generating enhanced contextual features. Its calculation formula is ,in It is a hyperparameter that can be learned by the network, initially set to... For detailed branches, first use and The data is then fused and further enhanced using residual connections to generate enhanced detail features. Its calculation formula is ,in It is also a hyperparameter that can be learned by the network, initially set to The setting of hyperparameters and their initial values is to avoid excessive interference from prior lesions on the original features, and an enhancement coefficient is introduced. and The intensity of lesion guidance is adaptively adjusted; for the enhancement coefficient, a larger coefficient can enhance the response of the lesion area, while a smaller coefficient helps to preserve the original contextual semantic information, thereby achieving a balance between lesion enhancement and background preservation. Step 8: Utilize the enhanced contextual features obtained in Step 7 With detailed features Splicing and then... After convolutional fusion, a channel attention mechanism (SE) is introduced. Channel weights are generated through global average pooling and fully connected layers to enhance key channels, ultimately yielding the output features. .
3. The method for detecting nasopharyngeal lesions by three-frequency dynamic fusion and differential perception according to claim 1, characterized in that... The construction steps for the semantic structure boundary difference guidance module are as follows: Step 1: The deep semantic features extracted by the backbone network and the shallow detail features are fused together by the TFDGM module to output the feature map. The data are fed into the semantic feature extraction branch and the structural feature extraction branch, respectively. The semantic branches are used Convolution performs channel compression and linear mapping to obtain preliminary semantic features; then, it undergoes non-linear processing using the GELU activation function, and finally... Convolution is used to restore channel dimensions in order to extract high-level semantic information of the lesion region; structural branches are employed. Depth convolution, Depth convolution and Deep convolutions are used to construct an orientation-aware feature encoder to enhance the representation of lesion edges and local texture structure features, thereby obtaining semantic features. and structural features ; Step 2: Process the semantic features obtained in Step 1 and structural features To perform differential modeling, a differential response map is obtained by calculating the element-wise squared difference between the two, and its expression is as follows: Then, average aggregation is performed along the channel dimension, i.e. This yields an initial differential feature map that characterizes the degree of difference between the lesion area and the background area. And the initial difference feature map Through a Enhanced differential feature maps are obtained by performing local differential convolutions using depthwise convolution. ; Step 3: The enhanced differential feature map obtained in Step 2... The final differential guided weight map is generated using the Sigmoid activation function. ,in The value is close to This indicates that the lesion area is of high importance and is close to... Indicates the background area; Step 4: Apply the difference-guided weighting graph obtained in Step 3. Semantic features used to guide the acquisition in step 1 and structural features Dynamic reallocation is performed to enhance semantic features; the calculation formula is as follows: The structural features retain structural information, and the calculation formula is as follows: That is, the lesion area through Enhanced, background area through Preserve structural information to achieve adaptive feature redistribution; Step 5: Reassign the semantic features obtained in Step 4 With structural features Concatenate along the channel dimension, then pass through Convolutional feature fusion, and finally combined with input features Perform residual connections to output the enhanced feature map. It is used for subsequent P3 and P4 detection heads to perform target detection and classification of lesions.