CBCT 3D Image Recognition Method for Adenoid Size Assessment
Through the CBCT three-dimensional image recognition method, the adenoid data is automatically processed using sinusoidal convolution and multi-scale feature fusion modules, solving the problems of high radiation risk, expensive and relying on manual operation of the existing adenoid evaluation method, and achieving high-precision, low-cost and standardized evaluation of adenoid size.
Patent Information
- Application Number
- CN202510431600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing adenoid evaluation methods have problems such as high radiation risk, expensive, complex operation, relying on manual operations and strong subjectivity of the results. It is difficult to accurately evaluate the size and hypertrophy of the adenoid, resulting in misdiagnosis or misdiagnosis.
The CBCT three-dimensional image recognition method is adopted to automatically process the three-dimensional data of the adenoid through sinusoidal bar convolution and multi-scale feature fusion module, calculate the volume of the adenoid and nasopharyngeal airway, and use the AI system to perform full-process automated operations to reduce manual operations and improve accuracy and consistency.
It realizes high-precision, low-radiation and low-cost evaluation of adenoid size, reduces operational complexity and manual operation deviation, improves the standardization of work efficiency and results, and is suitable for large-scale screening.
Smart Images

Figure CN119941738B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical practice technology, in particular to a CBCT three-dimensional image recognition method for adenoid size assessment. Background Art
[0002] As people pay more attention to their health, the need to accurately assess the size of adenoids has become more urgent. Adenoids are lymphatic tissues located above and behind the pharynx. They are an important part of the human immune system. Their main function is to protect the upper respiratory tract from invasion by external pathogens. Adenoids usually shrink gradually with age, but in some cases, such as infection or allergic reaction, adenoids may hypertrophy, causing a series of health problems.
[0003] Adenoid hypertrophy may cause a variety of clinical symptoms, including obstructive sleep apnea, otitis media, nasal congestion, mouth breathing, etc. These symptoms may affect the patient's normal breathing, sleep quality, hearing and other physiological functions, and even affect the patient's overall quality of life. Therefore, accurate assessment of the size and degree of hypertrophy of the adenoids is crucial for early diagnosis and treatment.
[0004] In the clinical diagnosis of adenoids, accurate evaluation methods are directly related to the treatment effect. Through scientific evaluation methods, doctors can more accurately judge whether the adenoids are enlarged, the degree of enlargement, and whether clinical symptoms will be caused. Correct evaluation can help doctors formulate more reasonable treatment plans, thereby improving efficacy and reducing side effects. On the contrary, inaccurate evaluation may lead to misdiagnosis or missed diagnosis, affecting the pertinence and effect of treatment.
[0005] As the disease progresses, adenoids hypertrophy not only affects the patient's short-term health, but may also have a profound negative impact on his or her growth and development, mental health, social ability and other aspects. Therefore, accurate assessment of the size of the adenoids is of great significance for early diagnosis and timely intervention. Modern medical imaging technology, such as cone beam computed tomography equipment and three-dimensional reconstruction technology, can provide accurate adenoids size assessment, thereby helping doctors develop personalized treatment plans and improve treatment outcomes.
[0006] The current adenoids assessment method is to judge whether the adenoids are enlarged based on the patient's clinical symptoms, such as nasal congestion, snoring, sleep apnea and hearing loss. Doctors make a preliminary judgment based on the patient's symptoms. However, the severity of symptoms in this method varies from individual to individual, and the assessment results are relatively subjective. The size and degree of enlargement of the adenoids cannot be accurately assessed based on symptoms alone, which can easily lead to missed diagnosis or misdiagnosis.
[0007] Currently, existing methods for adenoid assessment include X-ray imaging. By taking X-ray images of the adenoids, their size, shape, and whether they compress the airway are observed. However, the disadvantage of this method is the radiation risk, and the image is two-dimensional and lacks all-round three-dimensional information.
[0008] Currently, existing methods for adenoid assessment include CT scans. CT scans generate high-resolution three-dimensional images of the adenoids, providing clearer information about the size of the adenoids and their compression of surrounding structures, such as the upper airway and nasal cavity. However, the disadvantages of this method are relatively high radiation exposure, high cost, complex operation, and the need for professional equipment and technical support.
[0009] Currently, existing methods for adenoid assessment include magnetic resonance imaging. By using strong magnetic fields and radio waves, detailed images of the adenoids are generated, which is particularly suitable for imaging soft tissues and can clearly show the shape and size of the adenoids. However, the disadvantages of this method are long examination time, high cost, complex operation, and relatively high requirements for patient tolerance.
[0010] Currently, existing methods for adenoid assessment include endoscopy. By using an endoscope to directly observe the adenoids through the nasal cavity or oral cavity, it is commonly used to examine symptoms such as nasal congestion and snoring caused by adenoid hypertrophy. However, the disadvantage of this method is that it is limited to local observation, and some patients may not be suitable.
[0011] In view of the above limitations, there is an urgent need to introduce more scientific and accurate assessment methods to improve the accuracy, efficiency, and standardization level of diagnosis. Summary of the Invention
[0012] The purpose of the present invention is to provide a CBCT three-dimensional image recognition method for adenoid size assessment, which solves the problems raised in the above background technology.
[0013] To achieve the above purpose, the present invention provides the following technical solution: A CBCT three-dimensional image recognition method for adenoid size assessment, including the following steps:
[0014] Step S1: After setting the boundary based on the CBCT image, calculate the volume of the adenoids.
[0015] Step S11: Acquisition and preprocessing of CBCT image data.
[0016] Step S12: Correct the head position and determine the FH plane.
[0017] Step S13: Determine the boundaries of the adenoids and the nasopharyngeal airway.
[0018] Step S2: Enhancement and correction of the boundaries of the adenoids and the nasopharyngeal airway based on strip convolution.
[0019] Step S21: Construct strip convolution in the form of a sine function.
[0020] Step S22, adenoid and nasopharyngeal airway segmentation;
[0021] Step S23, extraction of the bounding boxes of the adenoid and nasopharyngeal airway;
[0022] Step S24, edge correction of the adenoid and nasopharyngeal airway;
[0023] Step S25, calculation of the volume of the adenoid and nasopharyngeal airway;
[0024] Step S3, output volume of the multi-scale feature fusion module;
[0025] Step S31, preprocessing of the input image, adaptation of the low-resolution image input;
[0026] Step S32, lightweight multi-scale feature extraction;
[0027] Step S33, branch processing module, specialization with low computational cost;
[0028] Step S34, dynamic convolution module, fast feature integration and quantization of dynamic convolution;
[0029] Step S35, attention fusion module;
[0030] Step S4, evaluated size of the adenoid.
[0031] Optionally, the specific process of obtaining and preprocessing the CBCT image data in step S11 is as follows:
[0032] Through a CBCT scanning device, three-dimensional image data of the patient is obtained, and then the three-dimensional image data is cleaned, denoised, and three-dimensionally reconstructed to generate a visualized three-dimensional model.
[0033] Optionally, the specific process of correcting the head position and determining the FH plane in step S12 is as follows:
[0034] Based on the visualized three-dimensional model output in step S11, the connecting line of the left and right orbital points on the coronal plane is parallel to the horizontal line, and the orbitomeatal plane on the sagittal plane is parallel to the horizontal plane;
[0035] The specific process of determining the boundaries of the adenoid and nasopharyngeal airway in step S13 is as follows:
[0036] The boundaries of the adenoid and nasopharyngeal airway are determined in the three-dimensional model, including the anterior boundary in the sagittal plane, the posterior boundary in the sagittal plane, the upper boundary in the sagittal plane, the lower boundary in the sagittal plane, the anterior boundary in the coronal plane, the lateral boundary in the coronal plane, the posterior boundary in the coronal plane, the anterior boundary in the axial plane, the posterior boundary in the axial plane, and the lateral boundary in the axial plane. By annotating the boundaries of the adenoid and nasopharyngeal airway in the coronal, sagittal, and axial planes, a cuboid model is constructed in three-dimensional space.
[0037] Optionally, the specific process of constructing the bar convolution using the form of the sine function in step S21 is as follows:
[0038] The sine bar convolution is an operation where a one-dimensional sine bar convolution kernel slides along one direction of the image to extract edges with linear and periodic features. The sine bar convolution combines the periodicity of the sine function and enhances edges, textures, and periodic structures in the image through the periodic characteristics of the sine wave;
[0039] The sine bar convolution kernel can be represented by the sine function, and the calculation formula is as follows:
[0040] Where:
[0041] W(x) denotes the sine bar convolution kernel;
[0042] A denotes the amplitude, which determines the intensity of the sine wave;
[0043] f denotes the frequency, which controls the periodicity of the sine wave and affects the scale of the image features to be detected;
[0044] denotes the phase, which controls the starting position of the sine wave;
[0045] x denotes the spatial variable, representing the position coordinates along a certain direction in the image;
[0046] When performing convolution on the input image, the sine bar convolution kernel W(x) will slide along one direction of the image, perform local weighted summation on the image, and thus extract the edge, texture, and periodic structure information in the image;
[0047] The calculation formula for the sine bar convolution is:
[0048] ;
[0049] Where:
[0050] S( x, y ) denotes the convolution result of the image at position ( x, y );
[0051] I ( x + i, y ) denotes the pixel value of the image at position ( x + i, y );
[0052] denotes the sine bar convolution kernel sliding along multiple directions;
[0053] k Refers to the length of the sine bar convolution kernel, which determines the convolution operation;
[0054] Sine bar convolution kernel Has periodicity, which can help highlight the periodically changing parts in the image and enhance the edge detection effect;
[0055] When performing sine bar convolution, the sine bar convolution kernel slides in the three-dimensional CBCT image to calculate the weighted sum of each position in the image. The specific steps are as follows:
[0056] For each three-dimensional pixel point I ( x, y, z), use the sine bar convolution kernel to perform weighted calculation with the surrounding pixels of this point;
[0057] The output result S ( x, y, z)is a three-dimensional pixel value in the image, representing the edge intensity at this position;
[0058] The calculation formula for three-dimensional sine bar convolution is:
[0059] ;
[0060] Where:
[0061] S ( x, y, z)refers to the three-dimensional image pixel value, representing the edge intensity at the position of ( x, y, z);
[0062] I ( x, y, z) refers to the pixel value of the image at the position of ( x, y, z);
[0063] W ( i, j, l ) refers to the three-dimensional sine bar convolution kernel.
[0064] Optionally, the adenoid and nasopharyngeal airway edge correction in step S24 includes: image input, edge expansion, sine bar convolution kernel of the expanded image, and edge correction;
[0065] Image input:
[0066] Designate the input original image pixel value as I ( x, y ), and its size is M × N , and edge processing requires expanding the edge pixels:
[0067] Edge expansion:
[0068] In the original image pixel valueI ( x, y ) Perform mirror extension around it, and the extension width is w , w being the filter radius to obtain the pixel values of the extended image I ext ( x, y );
[0069] The calculation formula for the pixel values of the extended image is:
[0070] Where:
[0071] I ext ( x, y ) refers to the pixel values of the extended image;
[0072] I ( x, y ) refers to the pixel values of the original image;
[0073] M refers to the width of the image;
[0074] N refers to the height of the image;
[0075] I ( x, y ), 0 ≤ x < M refers to when I ( x, y ) is within the valid range of the image, that is, when 0 ≤ x < M , directly use the original intensity value of I ( x, y );
[0076] I ( w - x, y ), x < 0 refers to processing the boundary pixels on the left side of the original image, and replacing them with the intensity values of the corresponding pixel points on the right side of the image ( w - x, y ) through mirror extension to maintain the continuity and smoothness of the image edge during edge correction;
[0077] I ( 2M - x - 1, y ), x ≥ M refers to processing the boundary pixels on the right side of the image, and replacing them with the intensity values of the corresponding pixel points on the left side of the image ( 2M - x - 1, y ) through mirror extension to maintain the continuity and smoothness of the image edge during edge correction;
[0078] The calculation formula for the sine bar convolution kernel of the extended image is as follows:
[0079] ; ;
[0080] Where:
[0081] K ( x, y ) refers to the pixel value of the extended image I ext ( x, y ) is the sine bar convolution kernel;
[0082] G ( x, y ) refers to the Gaussian weighting function;
[0083] σ refers to the standard deviation of the Gaussian kernel, which is used to control the smoothness degree.
[0084] Optionally, the specific process of calculating the adenoid and nasopharyngeal airway volume in step S25 is as follows:
[0085] Based on the three-dimensional reconstruction model, calculate the number of voxels in the adenoid and nasopharyngeal airway regions. A voxel V voxel is the basic unit in a three-dimensional image, and each voxel represents a spatial position in the image;
[0086] The calculation formula for the adenoid and nasopharyngeal airway volume is as follows:
[0087] ; ; Where:
[0088] V refers to the total volume of the adenoid and nasopharyngeal airway. By accumulating all the voxels in the adenoid and nasopharyngeal airway regions, the total volume of the adenoid is obtained;
[0089] V voxel refers to the volume of each voxel;
[0090] d x ×d y ×d z refers to the volume of each voxel. The voxel volume is usually determined by the resolution of the image, that is, the actual size of each voxel. If the resolution of the image is d x ×d y ×d z , then the volume of each voxel isV voxel 。
[0091] Optionally, the specific process in the output volume of the multi-scale feature fusion module in step S3 is as follows:
[0092] For the input image preprocessing in step S31, the specific process of adapting the low-resolution image input is as follows:
[0093] Reduce the resolution of the input image through downsampling operations to reduce the computational volume at the source, completed using strided convolution or bilinear interpolation. The calculation formula is as follows:
[0094] ; where: I 、 Refers to the low-resolution image after downsampling; I x,y Refers to the original image; s Refers to the reduction ratio; The specific process of lightweight multi-scale feature extraction in step S32 is as follows:
[0095] The core design of the multi-scale feature fusion module is to reduce the feature map volume while extracting much richer multi-scale information, specifically including multi-scale parallel paths and compression operations;
[0096] The multi-scale parallel paths are specifically:
[0097] When extracting features with different receptive fields, reduce redundancy by sharing the basic calculation of the feature map and attaching lightweight convolutional kernels;
[0098] The compression operation is specifically:
[0099] After multi-scale feature fusion, use 1×1 convolution to compress the channels. The calculation formula is as follows:
[0100] ;
[0101] Where: Refers to the compressed feature map;
[0102] F MSFM ( x, y )Refers to the original feature map input to the multi-scale feature fusion module;
[0103] Conv 1×1 Refers to 1×1 convolution, used to reduce the number of channels;
[0104] For the branch processing module in step S33, the specific process of low-computation specialization is as follows:
[0105] Optimize the computational cost further through a specialized lightweight processing method that performs detail enhancement, context capture, and global information through branch processing, including: a detail enhancement branch, a context capture branch, and a global information branch;
[0106] Detail enhancement branch:
[0107] Use a high-pass filter to simulate gradient operations through convolution, extract edge detail information, and compress the volume;
[0108] The calculation formula of the detail enhancement branch is as follows:
[0109] ;
[0110] Where: Refers to the enhanced detail feature map;
[0111] HighPass Refers to the high-pass filter for extracting edge details;
[0112] Context capture branch:
[0113] Use dilated convolution to expand the receptive field and avoid reducing the resolution of the feature map;
[0114] The calculation formula of the context capture branch is as follows:
[0115] ;
[0116] Where: Refers to the detail feature map after context capture;
[0117] DilatedConv Refers to dilated convolution;
[0118] rate Refers to the dilation rate for controlling the size of the receptive field;
[0119] rate = 2 means the convolution kernel interval is 2 pixels;
[0120] Global information branch:
[0121] Use global average pooling to directly extract global features, with low computational cost. After lightweight operations on the output of each branch, the computational volume is further reduced;
[0122] The calculation formula of the global information branch is as follows:
[0123] ;
[0124] Where: Refers to the detail feature map after the global information branch;
[0125] GlobalAvgPooling Refers to global average pooling, which is used to compress the feature map of each channel into a single value.
[0126] Optionally, the specific process in the output volume of the multi-scale feature fusion module in step S3 further includes:
[0127] The specific process of the dynamic convolution module in step S34, fast feature integration quantization dynamic convolution, is as follows:
[0128] Weight generation: Using the compressed feature input, the weight generation network is designed as a lightweight version, including a single-layer linear mapping. The calculation formula for the dynamic weight is as follows:
[0129] ;
[0130] Where:
[0131] W i Refers to the dynamic weight;
[0132] Linear Refers to the single-layer fully connected network;
[0133] Softmax Refers to the normalized weight; Refers to the i compressed feature of the i th branch, including the detail branch, the context branch, and the global branch;
[0134] Feature integration: The dynamic convolution weight performs weighted summation on the features. The lightweight design of the dynamic convolution module reduces the computational cost while retaining the adaptive adjustment ability for the branch features. The calculation formula for feature integration is as follows:
[0135] ;
[0136] Where:
[0137] F DCM ( x, y ) Refers to the feature map output by the dynamic convolution module, representing the adaptive fusion result of the multi-scale features; Refers to the i th branch's feature value at position ( x, y );
[0138] The specific process of the attention fusion module in step S35 is as follows:
[0139] Spatial attention is calculated by a lightweight convolution with a convolution kernel size of 1×1. The calculation formula is as follows;
[0140] ;
[0141] Among them:
[0142] F 空间注意力 ( x, y ) refers to the spatial attention weight map;
[0143] Sigmoid refers to the activation function, which normalizes the weights to the interval [0, 1] and represents the importance of spatial positions;
[0144] MaxPool refers to the max pooling operation, which extracts the maximum value at each position in the feature map;
[0145] AvgPool refers to the average pooling operation, which extracts the mean value at each position in the feature map;
[0146] The channel attention mechanism is completed through global average pooling and a single-layer perceptron, and the calculation formula is as follows:
[0147] ;
[0148] Among them: F 通道注意力 ( x, y ) refers to the channel attention vector;
[0149] Linear refers to the single-layer fully connected network;
[0150] GAP refers to the global average pooling;
[0151] F DCM refers to the multi-scale features output by the dynamic convolution module;
[0152] The calculation formula of the final fused features is as follows:
[0153] ;
[0154] Among them:
[0155] F AFM ( x, y ) refers to the optimized feature map;
[0156] By simplifying the calculation path, the attention mechanism achieves a balance between low computational volume and efficient feature enhancement.
[0157] Optionally, the calculation process of the evaluated size of the adenoid in step S4 is as follows:
[0158] The size of the adenoid is evaluated by calculating the ratio of the adenoid volume to the nasopharyngeal cavity volume. The calculation formula is as follows:
[0159] The size index of the adenoid = adenoid volume / (adenoid volume + nasopharyngeal airway volume).
[0160] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0161] The present invention has an automated operation throughout the process. From the input of CBCT data to the calculation of the adenoid volume, the whole process is automatically completed by the AI system, reducing cumbersome manual operations, greatly improving work efficiency, saving a large amount of time, and being able to process more data. Different from the traditional technology, the posture movements and calculation processes in the traditional technology usually require doctors to manually calibrate the adenoid boundary, adjust the reference layout, and perform cumbersome workflow such as boundary correction.
[0162] The present invention automatically processes the three-dimensional data of the adenoid through a sine bar convolution model, uses a precise algorithm to obtain the adenoid boundary recognition and correction and calculate the volume, greatly reducing the operations performed manually, providing higher accuracy and consistency. Through the automated process, the measurements of all patients will be highly consistent, avoiding the deviation caused by operation deviation. Different from the existing methods such as three-dimensional model analysis method, boundary drawing and boundary correction, which often rely on manual operation and empirical judgment, and the measurement accuracy is greatly affected by the operator's skills and additional factors.
[0163] The system of the present invention seamlessly integrates modules such as data input, three-dimensional model reconstruction, boundary correction, and volume calculation, providing a unified and comprehensive platform, reducing the complexity of data processing, greatly improving the convenience of operation, and quickly completing the entire measurement process. Different from the existing technology, there may be problems of data incompatibility or poor information transmission between multiple software tools in the existing technology, resulting in a cumbersome operation process and an inability to analyze. BRIEF DESCRIPTION OF THE DRAWINGS
[0164] Figure 1 It is a schematic structural diagram of the present invention connecting bilateral infraorbital points in the coronal plane;
[0165] Figure 2 It is a schematic structural diagram of the present invention in the sagittal plane FH plane;
[0166] Figure 3 It is a schematic structural diagram of the present invention with the adenoid and nasopharyngeal airway boundary markings in the sagittal plane;
[0167] Figure 4 It is a schematic structural diagram of the present invention with the adenoid and nasopharyngeal airway boundary markings in the coronal plane;
[0168] Figure 5Schematic structural diagram of the axial position adenoid and nasopharyngeal airway boundary annotation of the present invention;
[0169] Figure 6 Schematic structural diagram of the three-dimensional cuboid model of the present invention for jointly annotating the sagittal, coronal, and axial position boundaries;
[0170] Figure 7 Schematic structural diagram of the multi-scale feature fusion module of the present invention.
[0171] Figure 8 Flowchart of the steps of the CBCT three-dimensional image recognition method for evaluating the size of the present adenoid. Detailed implementation manners
[0172] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0173] Example 1, please refer to Figures 1 to 8 , this embodiment provides a technical solution: a CBCT three-dimensional image recognition method for evaluating the size of the adenoid, including the following steps:
[0174] Step S1, after setting the boundary based on the CBCT image, calculate the volume of the adenoid;
[0175] Step S11, acquisition and preprocessing of CBCT image data;
[0176] Step S12, correct the head position and determine the FH plane;
[0177] Step S13, determine the adenoid and nasopharyngeal airway boundaries;
[0178] Step S2, strengthen and correct the adenoid and nasopharyngeal airway boundaries based on strip convolution;
[0179] Step S21, construct strip convolution in the form of a sine function;
[0180] Step S22, segment the adenoid and nasopharyngeal airway;
[0181] Step S23, extract the bounding boxes of the adenoid and nasopharyngeal airway;
[0182] Step S24, correct the edges of the adenoid and nasopharyngeal airway;
[0183] Step S25, calculate the volume of the adenoid and nasopharyngeal airway;
[0184] Step S3. The multi-scale feature fusion module outputs a volume.
[0185] Step S31. Preprocess the input image and adapt the input of the low-resolution image.
[0186] Step S32. Lightweight multi-scale feature extraction.
[0187] Step S33. Branch processing module, specialized for low computational load.
[0188] Step S34. Dynamic convolution module, fast feature integration and quantization of dynamic convolution.
[0189] Step S35. Attention fusion module.
[0190] Step S4. Evaluate the size of the adenoid.
[0191] In this embodiment: In the imaging method of the present invention, different from the prior art that uses rays and CT, but this method is two-dimensional or low-resolution three-dimensional with poor soft tissue contrast, the present invention uses a CBCT device, which is three-dimensional high-resolution imaging and can accurately restore the spatial relationship between the adenoid and the airway; different from the prior art when dealing with boundaries, which may rely more on manual annotation and threshold segmentation and may lead to blurred edges, the present invention uses sinusoidal bar convolution to strengthen the anatomical boundary and weight correction to eliminate artifacts, with higher segmentation accuracy; different from the prior art in terms of computational efficiency, the prior art may use manual measurement, which is time-consuming and has high computational redundancy in traditional networks. The present invention uses MSFM through dynamic convolution and attention mechanism, with faster inference speed; different from the prior art in terms of clinical applicability, the prior art may use MRI, resulting in high cost and strong invasiveness of the endoscope. The present invention samples non-invasive CBCT and full-automatic analysis, thereby reducing the cost to a certain extent and being suitable for large-scale screening; different from the prior art in terms of result standardization, the prior art may be based on the analysis of physicians, but differences in experience lead to evaluation biases. The present invention samples a standardized formula for comprehensive analysis, enhancing the comparability of results among different medical institutions.
[0192] Please refer to Figures 1 to 8 , the specific process of obtaining and preprocessing the CBCT image data in step S11 is as follows:
[0193] Through a CBCT scanning device, obtain the three-dimensional image data of the patient, and then clean, denoise and perform three-dimensional reconstruction on the three-dimensional image data to generate a visual three-dimensional model.
[0194] In this embodiment: In the process of three-dimensional image analysis, the first step is the data acquisition stage. By using a cone beam computed tomography (CBCT) device, detailed three-dimensional images of the patient can be obtained. To ensure image clarity and low noise, the setting of scanning parameters is crucial. The resolution is usually set to 0.3 mm or higher to capture fine structures; the voltage is adjusted between 70 - 120 kVp to obtain appropriate X-ray penetration; the current is controlled at 5 - 10 mA to balance image quality and patient radiation dose. The precise setting of these parameters is essential for subsequent diagnosis and treatment planning.
[0195] Secondly is the image preprocessing stage. First, to remove the noise introduced during scanning, median filtering or Gaussian filtering techniques are applied. These methods can effectively reduce noise while retaining edge information.
[0196] Finally, it enters the three-dimensional reconstruction stage. The CBCT images are three-dimensionally reconstructed to generate a visualization model. This step involves complex image processing and computer vision techniques, including image segmentation, feature extraction, and three-dimensional modeling, etc. The final three-dimensional model can visually display the structure, providing strong visual support for clinical diagnosis and treatment. By realizing details through these technologies, precise three-dimensional information can be extracted from the original CBCT images.
[0197] Please refer to Figures 1 to 8 , the specific process of correcting the head position and determining the FH plane in step S12 is as follows:
[0198] Based on the visualized three-dimensional model output in step S11, make the connection line between the left and right orbital points on the coronal plane parallel to the horizontal line, and make the orbitomeatal plane on the sagittal plane parallel to the horizontal plane.
[0199] In this embodiment: First, by connecting the bilateral suborbital points on the coronal plane and making it parallel to the horizontal plane, specifically as Figure 1 shown, mark the bilateral suborbital points on the image, and then adjust the coronal plane to pass through these two points to ensure that the connection line between the bilateral suborbital points is parallel to the horizontal plane.
[0200] Secondly, on the sagittal plane, by connecting the suborbital point and the ear point and ensuring that this plane (the orbitomeatal plane) is parallel to the horizontal plane, as Figure 2 shown, connect the suborbital point and the ear point and adjust the head position to make this connection line parallel to the horizontal plane, ensuring that the head is in a standardized posture during the imaging examination. Correcting the head position before measurement can help ensure that during the CBCT imaging examination, the head maintains a consistent angle and position, thereby improving the accuracy and repeatability of the image data and reducing measurement errors caused by angular deviation, which is especially suitable for the precise evaluation of facial anatomical structures, adenoids, and other regions.
[0201] The specific process of determining the boundaries of the adenoid and the nasopharyngeal airway in step S13 is as follows:
[0202] Determine the boundaries of the adenoid and the nasopharyngeal airway in the three-dimensional model, including the anterior boundary in the sagittal plane, the posterior boundary in the sagittal plane, the upper boundary in the sagittal plane, the lower boundary in the sagittal plane, the anterior boundary in the coronal plane, the lateral boundary in the coronal plane, the posterior boundary in the coronal plane, the anterior boundary in the axial plane, the posterior boundary in the axial plane, and the lateral boundary in the axial plane. By marking the boundaries of the adenoid and the nasopharyngeal airway in the coronal, sagittal, and axial planes, a cuboid model is constructed in three-dimensional space.
[0203] In this embodiment: First, mark the boundaries of the adenoid and the nasopharyngeal airway in the sagittal plane. The marking method of the sagittal plane boundaries can be detailed as follows Figure 3 as shown:
[0204] The anterior boundary in the sagittal plane: The anterior boundary is defined as the edge of the adenoid. Figure 3 The marking line 1A in
[0205] marks this boundary, indicating the position of the anterior edge of the adenoid in the sagittal plane; Figure 3 The posterior boundary in the sagittal plane: The posterior boundary is a line passing through the spheno-occipital synchondrosis and perpendicular to the FH plane. In Figure 3
[0206]
[0207] The upper boundary in the sagittal plane: The upper boundary passes through the spheno-occipital synchondrosis and is parallel to the FH plane. The marking line 2A marks the upper boundary;
[0208] The lower boundary in the sagittal plane: The lower boundary passes through the tip of the uvula and is parallel to the FH plane. The marking line 4A marks the position of the lower boundary.
[0209] The upper and lower boundaries of the nasopharyngeal airway have the same definition as the adenoid boundary, and the perimeter is the edge of the upper airway; Figure 4 as shown:
[0210] The anterior boundary in the coronal plane: The anterior boundary is defined as the anterior edge of the adenoid. Figure 4 The anterior boundary is marked as the marking line 1B in
[0211] The lateral boundary in the coronal plane: The lateral boundary is composed of the medial plates of the two sphenoid bones. Figure 4 The marking line 3B in
[0212] Figure 4 The posterior boundary in the coronal plane: The posterior boundary is determined by a line passing through the sphenoid cartilage union and perpendicular to the FH plane. The position of the posterior boundary is marked by the marking line 2B in
[0213] Next is the axial adenoid and nasopharyngeal airway boundary annotation: The annotation method of the axial boundary can be detailedly divided into the following parts as Figure 5 shown:
[0214] The anterior boundary in the axial plane: The definition of the anterior boundary is the edge of the adenoid, and the marking line 81C in Figure 5 marks this boundary;
[0215] The posterior boundary in the axial plane: The posterior boundary is the line passing through the spheno-occipital synchondrosis and perpendicular to the FH plane. The marking line 93C in Figure 5 marks the position of the posterior boundary, indicating that the posterior boundary passes through the sphenoid cartilage junction and is perpendicular to the FH plane;
[0216] The lateral boundaries in the axial plane: The lateral boundaries are composed of the medial plates of the sphenoid bones on both sides. In Figure 5 the marking line 102C marks the position of the lateral boundaries;
[0217] By accurately annotating the adenoid and nasopharyngeal airway boundaries in the coronal, sagittal, and axial planes, as Figure 6 shown, a cuboid-shaped model can be constructed in three-dimensional space. Among them, the 1D area is used to clearly identify the position and scope of the adenoid, and the 2D area is used to clearly identify the position and scope of the upper airway. This multi-dimensional annotation method can not only accurately reflect the morphological characteristics of the adenoid in different directions but also help better understand its spatial position and volume size in the anatomical structure. The complex three-dimensional visualization technology provides an important reference basis for medical diagnosis and treatment planning, enabling doctors to more accurately evaluate the pathological state of the adenoid;
[0218] The traditional method is to use lateral radiographs for two-dimensional line distance measurement. However, this method innovatively uses CBCT to measure the adenoid and nasopharyngeal volumes and evaluates the size of the adenoid by calculating their volume ratio. Secondly, through automatic measurement by the AI system, the cumbersome manual operation is reduced, the work efficiency is greatly improved, a large amount of time is saved, and more data can be processed.
[0219] Please refer to Figures 1 to 8 for the specific process of constructing the strip convolution in the form of a sine function in step S21:
[0220] The sine strip convolution is an operation of sliding a one-dimensional sine strip convolution kernel along one direction of the image to extract edges with linear and periodic characteristics. The sine strip convolution combines the periodicity of the sine function and enhances the edges, textures, and periodic structures in the image through the periodic characteristics of the sine wave;
[0221] The sine strip convolution kernel can be represented by a sine function, and the calculation formula is as follows:
[0222] Wherein:
[0223] W(x) represents a sine bar convolution kernel;
[0224] A represents the amplitude, which determines the intensity of the sine wave;
[0225] f represents the frequency, which controls the periodicity of the sine wave and affects the scale of the image features to be detected;
[0226] represents the phase, which controls the starting position of the sine wave;
[0227] x represents the spatial variable, which represents the position coordinates in a certain direction in the image;
[0228] When performing convolution on the input image, the sine bar convolution kernel W(x) will slide along one direction of the image, perform local weighted summation on the image, and thus extract edge, texture, and periodic structure information in the image;
[0229] The calculation formula for sine bar convolution is:
[0230] ;
[0231] Wherein:
[0232] S( x, y ) represents the convolution result of the image at the position ( x, y );
[0233] I ( x + i, y ) represents the pixel value of the image at the position ( x + i, y );
[0234] represents the sine bar convolution kernel sliding along multiple directions;
[0235] k represents the length of the sine bar convolution kernel, which determines the convolution operation;
[0236] The sine bar convolution kernel has periodicity, which can help highlight the periodically changing parts in the image and enhance the edge detection effect;
[0237] When performing sine bar convolution, the sine bar convolution kernel will slide in the three-dimensional CBCT image and calculate the weighted sum of each position in the image. The specific steps are as follows:
[0238] For each three-dimensional pixel point I ( x, y,z), perform weighted calculation on the surrounding pixels of this point using a sine bar convolution kernel;
[0239] The output result S ( x, y, z) is a three-dimensional pixel value in the image, representing the edge intensity at this position;
[0240] The calculation formula for three-dimensional sine bar convolution is:
[0241] ;
[0242] Where:
[0243] S ( x, y, z) refers to the three-dimensional image pixel value, representing ( x, y, z) the edge intensity at the position;
[0244] I ( x, y, z) refers to the pixel value of the image at the position ( x, y, z);
[0245] W ( i, j, l ) refers to the three-dimensional sine bar convolution kernel.
[0246] Enhance the gray-scale contrast between the adenoid and the nasopharyngeal airway through three-dimensional sine bar convolution;
[0247] It should be noted that a further explanation is made for the segmentation of the adenoid and the nasopharyngeal airway in step S22;
[0248] Step S22 includes threshold segmentation: Use the threshold method to extract the regions of the adenoid and the nasopharyngeal airway, and set an appropriate gray-scale threshold based on the comparison of the gray-scale value with the surrounding tissues to distinguish the adenoid and the nasopharyngeal airway from other tissues;
[0249] Step S22 includes edge detection: Use methods such as sine bar convolution to enhance the edges in the image and accurately extract the boundaries.
[0250] Step S22 includes three-dimensional morphological operations: Apply morphological operations such as dilation, erosion, and opening to correct the shapes of the regions of the adenoid and the nasopharyngeal airway, fill in discontinuous edges or remove noise to ensure the coherence of the boundaries;
[0251] Step S22 includes deep learning segmentation: In complex situations, deep learning algorithms, that is, U-Net, can be used to automatically segment the regions of the adenoid and the nasopharyngeal airway, which can better handle complex situations in the image and obtain higher-precision segmentation results.
[0252] It should be noted that the extraction of the adenoid and nasopharyngeal airway bounding boxes in step S23 will be further described;
[0253] Using the segmented adenoid and nasopharyngeal airway regions and the corresponding two-dimensional slice data, a three-dimensional model is reconstructed, and at this time, its contour and volume in the three-dimensional space have been formed;
[0254] Through image processing software such as Mimics and 3D Slicer, three-dimensional reconstruction is performed to generate a three-dimensional volume model of the adenoid and nasopharyngeal airway.
[0255] The edge correction of the adenoid and nasopharyngeal airway in step S24 includes: image input, edge expansion, sine bar convolution kernel of the expanded image, and edge correction;
[0256] Image input:
[0257] The pixel value of the input original image is designated as I ( x, y ), and its size is M × N , and edge processing requires expanding the edge pixels:
[0258] Edge expansion:
[0259] Mirror expansion is performed around the pixel value of the original image I ( x, y ), and the expansion width is w , w is the filter radius to obtain the pixel value of the expanded image I ext ( x, y );
[0260] The calculation formula for the pixel value of the expanded image is:
[0261] Where:
[0262] I ext ( x, y ) refers to the pixel value of the expanded image;
[0263] I ( x, y ) refers to the pixel value of the original image;
[0264] M refers to the width of the image;
[0265] N refers to the height of the image;
[0266] I ( x, y ), 0 ≤ x < MIt refers to when I ( x, y ) is within the valid range of the image, that is, when 0 ≤ x < M , directly use the original intensity value of I ( x, y );
[0267] I ( w - x, y ), x < 0 refers to processing the boundary pixels on the left side of the original image, and replacing them with the intensity values of the corresponding pixel points on the right side of the image through mirror extension ( w - x, y ) to maintain the continuity and smoothness of the image edge during edge correction;
[0268] I ( 2M - x - 1, y ), x ≥ M refers to processing the boundary pixels on the right side of the image, and replacing them with the intensity values of the corresponding pixel points on the left side of the image through mirror extension ( 2M - x - 1, y ) to maintain the continuity and smoothness of the image edge during edge correction;
[0269] The calculation formula for the sine bar convolution kernel of the extended image is as follows:
[0270] ; ;
[0271] Where:
[0272] K ( x, y ) refers to the pixel value of the extended image I ext ( x, y ) of the sine bar convolution kernel;
[0273] G ( x, y ) refers to the Gaussian weighting function;
[0274] σ refers to the standard deviation of the Gaussian kernel, which is used to control the smoothness;
[0275] Edge correction:
[0276] In the edge part, since the values of the extended image are not real image values, the response results are inaccurate. Therefore, weight correction needs to be introduced;
[0277] The calculation formula for the weight in edge correction is as follows:
[0278] ; ; Where:
[0279] W ( x,y ) refers to the weight map, which is the weighted effective area value at the coordinate ( x,y ). This value is used to correct the convolution response later to reduce the interference of the extended area on edge detection;
[0280] M ( x,y ) refers to the binary mask matrix, which is used to define the range of the effective area. M ( x,y ) functions to only retain the original image area, with a size of M × N , and mask the invalid area introduced by image expansion;
[0281] Otherwise refers to other situations;
[0282] G ( u, v ) refers to the Gaussian kernel function, which is used to perform weighted averaging on the masked area. G ( u, v ) functions to suppress high-frequency noise and smooth the edge transition area;
[0283] ( u, v ) refers to the support range of the convolution kernel, that is, the offset of the convolution kernel in the horizontal and vertical directions;
[0284] Corrected response:
[0285] The calculation formula for the corrected convolution response is as follows:
[0286] ; where:
[0287] R 、 ( x, y ) refers to the corrected convolution response map;
[0288] By dividing by the weight value W(x, y) , the response error in the edge area caused by image expansion is eliminated. The M (x + u, y + v) value in the extended area is 0, and its original convolution response value R ( x, y ) is unreliable due to the lack of real pixel values. Such invalid responses can be suppressed through weight correction;
[0289] R ( x, y ) refers to the original convolution response map;
[0290] ε refers to a very small constant to prevent the denominator from being zero;
[0291] Remove the extended region:
[0292] Crop the corrected response map R 、 ( x, y ), and retain the original image region 0 ≤ x < M , 0 ≤ y < N;
[0293] Thus, artifact interference can be avoided. The response values in the extended region may introduce noise or artifacts due to the lack of real data. After cropping, only the data of the effective anatomical region is retained. In cooperation with head position correction, combined with the coordinate system of head position correction, the sagittal plane, coronal plane, and axial plane are used to ensure that the three-dimensional volume calculation is only based on real anatomical structures. Through this edge correction, the boundary blurring problem that may be caused by traditional methods such as threshold segmentation is solved, thereby improving the segmentation accuracy.
[0294] The specific process of calculating the volume of the adenoid and nasopharyngeal airway in step S25 is as follows:
[0295] Based on the three-dimensional reconstruction model, calculate the number of voxels in the adenoid and nasopharyngeal airway regions. A voxel V voxel is the basic unit in a three-dimensional image, and each voxel represents a spatial position in the image;
[0296] The calculation formula for the volume of the adenoid and nasopharyngeal airway is as follows:
[0297] ; ; where:
[0298] V refers to the total volume of the adenoid and nasopharyngeal airway. By accumulating all the voxels in the adenoid and nasopharyngeal airway regions, the total volume of the adenoid is obtained;
[0299] V voxel refers to the volume of each voxel;
[0300] d x ×d y ×d z refers to the volume of each voxel. The voxel volume is usually determined by the resolution of the image, that is, the actual size of each voxel. If the resolution of the image is d x ×d y ×d z , then the volume of each voxel is V voxel .
[0301] In this embodiment: For the enhancement and correction steps of the adenoid and nasopharyngeal airway boundaries based on sinusoidal strip convolution, during the training process of cone beam computed tomography (CBCT) images, in order to meet the real-time computing requirements of on-site deployment, an improved sinusoidal strip convolution (MSC) model is adopted. By introducing the sinusoidal strip convolution model, it facilitates the medical image segmentation technology. Through multi-directional strip convolution operations, it can effectively enhance and correct the boundaries of the adenoid and nasopharyngeal airway. Specifically, the MSC model can better locate the target area by introducing multi-scale features. At the same time, combined with multi-directional convolution operations, it further optimizes the accuracy of boundary recognition and correction. This technology performs excellently in the adenoid and nasopharyngeal airway segmentation task, significantly improving the segmentation accuracy and robustness, providing more reliable imaging support for medical diagnosis and treatment. Strip convolution is an operation where a one-dimensional convolution kernel slides along one direction of the image, aiming to extract edges with linear or periodic features. Sinusoidal strip convolution combines the periodicity of the sine function to enhance edges, textures, and periodic structures in the image through the periodic features of the sine wave. This method constructs a three-dimensional structure model of the adenoid based on cone beam computed tomography (CBCT) technology through an accurate adenoid and nasopharyngeal airway boundary recognition and volume calculation method, and realizes the efficient calculation of volume through the voxel accumulation method. This method combines resolution calibration and error correction to ensure the accuracy of the volume calculation results, providing standardized data support for the diagnosis, evaluation, and treatment of adenoid hyperplasia. And through the automated adenoid and nasopharyngeal airway boundary correction and enhancement technology, using a deep learning model combined with morphological algorithms, the boundaries in the adenoid images are adaptively adjusted and the details are enhanced to ensure the integrity and clarity of the boundaries. Through feature extraction and local contrast enhancement, the problems of blurred and discontinuous boundaries in traditional methods are eliminated, the accuracy of boundary recognition is improved, and the subjective errors of manual operations are reduced.
[0302] Embodiment 2, based on the above embodiment: Please refer to Figures 1 to 8 , for the preprocessing of the input image in step S31, the specific process of adapting the low-resolution image input is as follows:
[0303] Reduce the resolution of the input image through downsampling operations to reduce the computational volume from the source, which is completed using stride convolution or bilinear interpolation. The calculation formula is as follows:
[0304] ; where: I 、 refers to the low-resolution image after downsampling; I x,y refers to the original image; s refers to the reduction ratio; The specific process of lightweight multi-scale feature extraction in step S32 is as follows:
[0305] The core of the design of the multi-scale feature fusion module is to reduce the volume of the feature map while extracting much richer multi-scale information, specifically including multi-scale parallel paths and compression operations;
[0306] The multi-scale parallel paths are specifically as follows:
[0307] When extracting features with different receptive fields, redundancy is reduced by sharing the basic calculation of the feature map and attaching lightweight convolutional kernels;
[0308] The compression operation is specifically as follows:
[0309] After multi-scale feature fusion, a 1×1 convolution is used to compress the channels, and the calculation formula is as follows:
[0310] ;
[0311] Where: refers to the compressed feature map;
[0312] F MSFM ( x, y ) refers to the original feature map input to the multi-scale feature fusion module;
[0313] Conv 1×1 refers to the 1×1 convolution, which is used to reduce the number of channels;
[0314] For the branch processing module in step S33, the specific process of specialization with low computational complexity is as follows:
[0315] Through branch processing for detail enhancement, context capture, and global information, a specialized lightweight processing method is used to further optimize the computational complexity, including: detail enhancement branch, context capture branch, and global information branch;
[0316] Detail enhancement branch:
[0317] A high-pass filter is used to simulate gradient operations through convolution, extract edge detail information, and compress the volume;
[0318] The calculation formula of the detail enhancement branch is as follows:
[0319] ;
[0320] Where: refers to the enhanced detail feature map;
[0321] HighPass refers to the high-pass filter, which is used to extract edge details;
[0322] Context capture branch:
[0323] Use dilated convolution to expand the receptive field and avoid reducing the resolution of the feature map;
[0324] The calculation formula of the context capture branch is as follows:
[0325] ;
[0326] Where: Refers to the detailed feature map after context capture;
[0327] DilatedConv Refers to dilated convolution;
[0328] rate Refers to the dilation rate, which is used to control the size of the receptive field;
[0329] rate =2 means the convolution kernel interval is 2 pixels;
[0330] Global information branch:
[0331] Use global average pooling to directly extract global features, with low computational cost. After lightweight operations on the output of each branch, the computational volume is further reduced;
[0332] The calculation formula of the global information branch is as follows:
[0333] ;
[0334] Where: Refers to the detailed feature map after the global information branch;
[0335] GlobalAvgPooling Refers to global average pooling, which is used to compress the feature map of each channel into a single value.
[0336] The specific process of the dynamic convolution module in step S34, namely fast feature integration quantization dynamic convolution, is as follows:
[0337] Weight generation: Using the compressed feature input, the weight generation network is designed as a lightweight version, including a single-layer linear mapping. The calculation formula of the dynamic weight is as follows:
[0338] ;
[0339] Where:
[0340] W i Refers to the dynamic weight;
[0341] Linear Refers to the single-layer fully connected network;
[0342] Softmax Refers to the normalized weight; Referring to the i compressed features of the i th branch, including the detail branch, the context branch, and the global branch;
[0343] Feature integration: The dynamic convolution weights perform weighted summation on the features. The lightweight design of the dynamic convolution module reduces the computational cost while retaining the ability to adaptively adjust the branch features. The calculation formula for feature integration is as follows:
[0344] ;
[0345] Where:
[0346] F DCM ( x, y ) refers to the feature map output by the dynamic convolution module, representing the adaptive fusion result of multi-scale features; Refers to the i th branch's feature value at position ( x, y );
[0347] The specific process of the attention fusion module in step S35 is as follows:
[0348] Spatial attention is calculated through a lightweight convolution with a kernel size of 1×1. The calculation formula is as follows;
[0349] ;
[0350] Where:
[0351] F 空间注意力 ( x, y ) refers to the spatial attention weight map;
[0352] Sigmoid Refers to the activation function, which normalizes the weights to the [0,1] interval, representing the importance of spatial positions;
[0353] MaxPool Refers to the max pooling operation, which extracts the maximum value at each position in the feature map;
[0354] AvgPool Refers to the average pooling operation, which extracts the mean value at each position in the feature map;
[0355] The channel attention mechanism is completed through global average pooling and a single-layer perceptron. The calculation formula is as follows:
[0356] ;
[0357] Where: F 通道注意力 ( x, y ) refers to the channel attention vector;
[0358] Linear Refers to a single-layer fully connected network;
[0359] GAP Refers to global average pooling;
[0360] F DCM Refers to the multi-scale features output by the dynamic convolution module;
[0361] The calculation formula for the final fused features is as follows:
[0362] ;
[0363] Where:
[0364] F AFM ( x, y ) Refers to the optimized feature map;
[0365] By simplifying the calculation path, the attention mechanism achieves a balance between low computational volume and efficient feature enhancement.
[0366] In this embodiment: For the multi-scale feature fusion module, that is, the MSFM module, the output volume is a multi-channel feature map that fuses multi-scale features. By fusing features from different scales, the model can capture detailed information and global information, enhancing the understanding of the input data. The output volume can maintain the same spatial dimension as the input feature map, or change due to operations such as pooling and upsampling. Due to the fusion of multi-scale features, the number of channels in the output volume usually increases, thus containing more hierarchical feature information. The structure is as Figure 7 shown. These feature information helps to improve the performance of the model in object detection tasks, provides richer high-level features, and helps the model better perform tasks such as classification and localization. In short, the output volume of the multi-scale feature fusion module effectively enhances the model's ability in complex multi-scale tasks through multi-scale feature fusion;
[0367] By introducing a multi-scale feature fusion module, the output efficiency and accuracy of adenoid and nasopharyngeal airway volume calculation are significantly improved. The multi-scale feature fusion module can extract the global structural features and local detail information of the adenoid under different receptive fields, realizing the efficient segmentation and recognition of the adenoid and nasopharyngeal airway regions. Through strip convolution operations and feature splicing mechanisms, the module can make full use of the multi-level information of the image, improve the computational efficiency of the model, and reduce the resource occupation caused by redundant calculations. In addition, the module adaptively weights and integrates the features, highlighting the saliency of the adenoid and nasopharyngeal airway regions, ensuring the accuracy and consistency of the 3D reconstruction and volume calculation results. The key lies in using the ability of multi-scale feature fusion to accelerate the adenoid and nasopharyngeal airway volume calculation process, including multi-scale feature extraction, the design of an efficient feature fusion algorithm, and a fast volume output method based on this module. This method is applicable to a variety of medical imaging scenarios, providing efficient and reliable technical support for clinical diagnosis and treatment.
[0368] The calculation process of the evaluated size of the adenoid in step S4 is as follows:
[0369] The size of the adenoid is evaluated by calculating the ratio of the adenoid volume to the nasopharyngeal cavity volume. The calculation formula is as follows:
[0370] The size index of the adenoid = adenoid volume / (adenoid volume + nasopharyngeal airway volume).
[0371] In this embodiment: Through this calculation formula, it can be judged whether the adenoid completely obstructs the nasopharyngeal cavity or the adenoid volume is normal, so as to quantify the degree of airway obstruction caused by adenoid hypertrophy, providing an objective basis for clinical grading, that is, mild, moderate, and severe, replacing the subjective evaluation of traditional endoscopy or MRI, and realizing standardized and repeatable quantitative diagnosis.
[0372] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A CBCT three-dimensional image recognition method for evaluating the size of the adenoid, characterized in that, It includes the following steps: Step S1: After setting the boundary based on the CBCT image, calculate the volume of the adenoid. Step S1 includes: acquisition and preprocessing of CBCT image data, head position correction, determination of the FH plane, and determination of the boundaries of the adenoid and nasopharyngeal airway; Step S2: Enhancement and correction of the boundaries of the adenoid and nasopharyngeal airway based on strip convolution. Step S2 includes: constructing strip convolution in the form of a sine function, segmentation of the adenoid and nasopharyngeal airway, extraction of the bounding box of the adenoid and nasopharyngeal airway, edge correction of the adenoid and nasopharyngeal airway, and calculation of the volume of the adenoid and nasopharyngeal airway; Step S3: The multi-scale feature fusion module outputs the volume. Step S3 includes: preprocessing of the input image, adaptation of the low-resolution image input, lightweight multi-scale feature extraction, branch processing module, low computational complexity specialization, dynamic convolution module, fast feature integration and quantization of dynamic convolution, and attention fusion module; Step S4: The evaluated size of the adenoid; The specific process of the acquisition and preprocessing of CBCT image data in step S1 is as follows: Through the CBCT scanning device, obtain the three-dimensional image data of the patient, and then perform cleaning, denoising, and three-dimensional reconstruction on the three-dimensional image data to generate a visualized three-dimensional model; The specific process of head position correction and determination of the FH plane in step S1 is as follows: Based on the visualized three-dimensional model output from the acquisition and preprocessing of CBCT image data, make the connecting line of the left and right orbital points on the coronal plane parallel to the horizontal line, and make the orbitomeatal plane on the sagittal plane parallel to the horizontal plane; The specific process of determining the boundaries of the adenoid and nasopharyngeal airway in step S1 is as follows: Determine the boundaries of the adenoid and nasopharyngeal airway in the three-dimensional model, including the anterior boundary in the sagittal plane, the posterior boundary in the sagittal plane, the upper boundary in the sagittal plane, the lower boundary in the sagittal plane, the anterior boundary in the coronal plane, the lateral boundary in the coronal plane, the posterior boundary in the coronal plane, the anterior boundary in the axial plane, the posterior boundary in the axial plane, and the lateral boundary in the axial plane. By annotating the boundaries of the adenoid and nasopharyngeal airway in the coronal, sagittal, and axial planes, construct a cuboid model in three-dimensional space; The specific process of constructing strip convolution in the form of a sine function in step S2 is as follows: Strip convolution is an operation in which a one-dimensional sine strip convolution kernel slides along one direction of the image to extract edges with linear and periodic features. Sine strip convolution combines the periodicity of the sine function and enhances edges, textures, and periodic structures in the image through the periodic features of the sine wave; The sine strip convolution kernel can be represented by a sine function, and the calculation formula is as follows: W(x) = Asin(2πfx + φ); Where: W(x) refers to the sine strip convolution kernel; A refers to the amplitude, which determines the intensity of the sine wave; f refers to the frequency, which controls the periodicity of the sine wave and affects the scale of the image features to be detected; φ refers to the phase, which controls the starting position of the sine wave; x refers to the spatial variable, representing the position coordinate along a certain direction in the image; When performing convolution on the input image, the sine strip convolution kernel W(x) will slide along one direction of the image, perform local weighted summation on the image, and thus extract the edge, texture, and periodic structure information in the image; The calculation formula of sine strip convolution is: Where: S(x, y) refers to the convolution result of the image at the position (x, y); I(x + i, y) refers to the pixel value of the image at the position (x + i, y); Asin(2πfi + φ) refers to the sine bar convolution kernel sliding in multiple directions; k refers to the length of the sine bar convolution kernel, which determines the convolution operation; The sine bar convolution kernel W(x) = Asin(2πfx + φ) is periodic and can help highlight the periodically varying parts in the image, enhancing the edge detection effect; When performing sine bar convolution, the sine bar convolution kernel slides in the three-dimensional CBCT image to calculate the weighted sum of each position in the image. The specific steps are as follows: For each three-dimensional pixel point I(x, y, z), use the sine bar convolution kernel to perform weighted calculation with the surrounding pixels of this point; The output result S(x, y, z) is a three-dimensional pixel value in the image, representing the edge intensity at this position; The calculation formula for three-dimensional sine bar convolution is: Where: S(x, y, z) refers to the three-dimensional image pixel value, representing the edge intensity at the position (x, y, z); I(x, y, z) refers to the pixel value of the image at the position (x, y, z); W(i, j, l) refers to the three-dimensional sine bar convolution kernel; In step S2, the edge correction of the adenoid and nasopharyngeal airway includes image input, edge expansion, sine bar convolution kernel of the expanded image, and edge correction; Image input: The input original image pixel value is designated as I(x, y), and its size is M × N. Edge processing requires expanding the edge pixels: Edge expansion: Mirror-expand around the original image pixel value I(x, y) with an expansion width of w, where w is the filter radius, to obtain the expanded image pixel value I ext (x, y); The calculation formula for the pixel value of the expanded image is: Where: I ext (x, y) refers to the pixel value of the expanded image; I(x, y) refers to the original image pixel value; M refers to the width of the image; N refers to the height of the image; I(x, y), 0 ≤ x < M means that when I(x, y) is within the valid range of the image, that is, when 0 ≤ x < M, directly use the original intensity value of I(x, y); I(w - x, y), x < 0 refers to processing the boundary pixels on the left side of the original image. By mirror expansion, use the intensity value of the corresponding pixel point (w - x, y) on the right side of the image to replace it, maintaining the continuity and smoothness of the image edge during edge correction; I(2M - x - 1, y), x ≥ M refers to processing the boundary pixels on the right side of the image. By mirror expansion, use the intensity value of the corresponding pixel point (2M - x - 1, y) on the left side of the image to replace it, maintaining the continuity and smoothness of the image edge during edge correction; The calculation formula for the sine bar convolution kernel of the expanded image is as follows: K(x,y) = sin(2f x )·G(x,y); Where: K(x, y) refers to the extended image pixel value I ext The sine bar convolution kernel of (x, y); G(x, y) refers to the Gaussian weighting function; σ refers to the Gaussian kernel standard deviation, which is used to control the smoothness; The specific process of calculating the volume of the adenoid and nasopharyngeal airway in step S2 is as follows: Based on the three-dimensional reconstruction model, calculate the number of voxels in the adenoid and nasopharyngeal airway regions, where voxel V voxel is the basic unit in the three-dimensional image, and each voxel represents a spatial position in the image; The calculation formula for the volume of the adenoid and nasopharyngeal airway is as follows: V voxel = d x × d y × d z ; Where: V refers to the total volume of the adenoid and nasopharyngeal airway. By accumulating all the voxels in the adenoid and nasopharyngeal airway region, the total volume of the adenoid is obtained; V voxel Denotes the volume of each voxel; d x × d y × d z Refers to the volume of each voxel. The voxel volume is usually determined by the resolution of the image, that is, the actual size of each voxel. If the resolution of the image is d x × d y × d z , then the volume of each voxel is V voxel ; The specific process of the output volume in the multi-scale feature fusion module in step S3 is as follows: In the input image preprocessing in step S3, the specific process of adapting the input of the low-resolution image is as follows: Reduce the resolution of the input image through downsampling operations, reducing the computational volume at the source, accomplished using strided convolution or bilinear interpolation. The calculation formula is as follows: I' = DownSample(I x,y , scale = s); Where: I refers to the low-resolution image after downsampling; I x,y Refers to the original image; s refers to the reduction ratio; The specific process of lightweight multi-scale feature extraction in step S3 is as follows: The core design of the multi-scale feature fusion module is to reduce the volume of the feature map while extracting much richer multi-scale information, specifically including multi-scale parallel paths and compression operations; The multi-scale parallel paths are specifically: When extracting features with different receptive fields, reduce redundancy by sharing the basic calculation of the feature map and attaching lightweight convolutional kernels; The compression operation is specifically: After multi-scale feature fusion, use 1×1 convolution to compress the channels; Where: Refers to the compressed feature map; F MSFM (x, y) refers to the original feature map input into the multi-scale feature fusion module; Conv 1×1 Refers to 1×1 convolution, used to reduce the number of channels; For the branch processing module in step S3, the specific process of low-computation specialization is as follows: Through branch processing for detail enhancement, context capture, and global information, a specialized lightweight processing method is used to further optimize the computational volume, including: detail enhancement branch, context capture branch, and global information branch; Detail enhancement branch: Use a high-pass filter, simulate gradient operations through convolution, extract edge detail information, and compress the volume; The calculation formula for the detail enhancement branch is as follows: Where: Refers to the enhanced detailed feature map; HighPass refers to the high-pass filter used to extract edge details; Context capture branch: Use dilated convolution to expand the receptive field and avoid reducing the resolution of the feature map; The calculation formula for the context capture branch is as follows: Where: The detailed feature map after context capture by reference; DilatedConv refers to dilated convolution; rate refers to the dilation rate used to control the size of the receptive field; rate = 2 means the convolution kernel spacing is 2 pixels; Global information branch: Use global average pooling to directly extract global features, with low computational volume. After lightweight operations on the output of each branch, the computational volume is further reduced; The calculation formula for the global information branch is as follows: Where: Detail feature map after referring to the global information branch; GlobalAvgPooling refers to global average pooling, used to compress the feature map of each channel into a single value.
2. The CBCT three-dimensional image recognition method for adenoid size assessment according to claim 1, characterized in that: The specific process in the output volume of the multi-scale feature fusion module in step S3 also includes: For the dynamic convolution module in step S3, the specific process of fast feature integration and quantization of dynamic convolution is as follows: Weight generation: Using the compressed feature input, the weight generation network is designed as a lightweight version, including a single-layer linear mapping; Feature integration: The dynamic convolution weights perform weighted summation on the features. The lightweight design of the dynamic convolution module reduces the computational cost while retaining the ability to adaptively adjust the branch features; The specific process of the attention fusion module in step S3 is as follows: Spatial attention is calculated through lightweight convolution with a convolution kernel size of 1×1; The channel attention mechanism is completed through global average pooling and a single-layer perceptron.
Citation Information
Patent Citations
Adenoid hypertrophy degree automatic detection method based on oral cavity CBCT image
CN118710669A
Hierarchical Transform attention feature fusion-based iPSCs set pluripotency evaluation method
CN119672711A