Intelligent Identification Method and System for Jawbone Lesions Based on Oral CBCT Images
By registering multi-temporal CBCT image data and using a fully connected neural network with an adaptive weighting mechanism, the problems of lack of dynamic information and neglect of anatomical heterogeneity in jawbone lesion identification in existing technologies are solved, achieving more accurate and interpretable lesion identification.
Patent Information
- Application Number
- CN202511359232.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing intelligent identification methods for jawbone lesions based on oral CBCT images rely on static analysis of image data at a single time point, which cannot capture the dynamic evolution information of lesions, ignores the physiological differences of different anatomical regions of the jawbone, and lacks effective feature importance assessment and interpretability analysis.
By acquiring multi-temporal CBCT image data, registration processing and jaw region segmentation are performed to extract shape, texture and statistical features. A fully connected neural network with adaptive weighting and attention mechanisms is used for classification and recognition, highlighting key feature combinations.
It enables the effective capture and utilization of dynamic evolution information of jawbone lesions, improves the accuracy and robustness of identification, and provides clinicians with interpretable diagnostic evidence.
Smart Images

Figure CN120852891B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to a method and system for intelligent recognition of jaw lesions based on oral CBCT images. Background Technology
[0002] Existing intelligent identification methods for jawbone lesions based on oral CBCT images mainly rely on static analysis of image data from a single time point. They extract image features and perform classification and identification using deep learning networks or traditional machine learning algorithms. These methods typically employ convolutional neural networks for automatic segmentation and feature extraction of CBCT images, and then use classifiers such as support vector machines, random forests, or fully connected neural networks to identify lesions based on the extracted features, thus achieving automated detection of jawbone lesions to a certain extent.
[0003] However, single-phase image analysis cannot capture the dynamic evolution of lesions over time and lacks an effective description of lesion development trends. Secondly, the use of equal weighting in feature extraction ignores the physiological differences in lesion sensitivity among different anatomical regions of the jawbone (cortical bone and cancellous bone). Thirdly, existing methods lack an effective mechanism for assessing feature importance and cannot identify the key feature combinations most relevant to lesion identification. Finally, the identification results lack interpretability analysis and are difficult to provide decision-making basis for clinicians.
[0004] At the data utilization level, single-phase analysis limits the acquisition of dynamic information about lesions, while the effective fusion of multi-phase difference features has become a key technical challenge. At the feature processing level, equal-weighting ignores anatomical heterogeneity, and an adaptive weighting mechanism based on physiological characteristics needs to be established. At the model optimization level, there is a lack of methods for quantifying feature importance, and attention mechanisms need to be introduced to achieve automatic identification and weighted processing of key features. At the result interpretation level, black-box identification lacks clinical acceptability, and a comprehensive diagnostic analysis framework with interpretability needs to be constructed. Summary of the Invention
[0005] This application provides a method and system for intelligent identification of jaw lesions based on oral CBCT images, which solves the technical problems in the prior art where single-phase CBCT image analysis cannot effectively utilize the dynamic evolution information of lesions and the equal-weight feature processing ignores the anatomical heterogeneity of the jaw. It improves the pertinence of multi-phase difference feature fusion and the accuracy of intelligent identification of jaw lesions.
[0006] In a first aspect, this application provides a method for intelligent identification of jawbone lesions based on oral CBCT images, the method comprising:
[0007] Obtain oral CBCT image data of the patient's initial examination and follow-up examination, perform registration processing on the image data of the initial examination and follow-up examination to obtain spatially aligned multi-temporal CBCT images, and perform jaw region segmentation on the multi-temporal CBCT images to obtain a jaw region of interest mask.
[0008] Based on the jawbone region of interest mask, shape features, texture features, and statistical features are extracted from the images of the initial examination and follow-up examination to obtain the feature vectors of the initial examination and the follow-up examination.
[0009] The region of interest in the jawbone is divided into a cortical bone region and a cancellous bone region based on a bone mineral density threshold. Adaptive weighting coefficients are set for the cortical bone region and the cancellous bone region respectively based on their physiological characteristics. The weighted difference between the initial examination feature vector and the follow-up examination feature vector is calculated using the adaptive weighting coefficients to obtain a multi-temporal difference feature vector. A variance screening algorithm is used to select features from the multi-temporal difference feature vector to obtain an effective difference feature vector. The initial examination feature vector and the effective difference feature vector are concatenated and fused to generate a fused feature vector.
[0010] The fused feature vector is input into a fully connected neural network with attention mechanism weighting for classification and recognition. The key feature combination related to jawbone lesions is highlighted by the attention weight calculation to obtain the jawbone lesion recognition result.
[0011] Secondly, this application provides an intelligent identification system for jawbone lesions based on oral CBCT images, the intelligent identification system for jawbone lesions based on oral CBCT images comprising:
[0012] The registration module is used to acquire oral CBCT image data of the patient's initial examination and follow-up examination, perform registration processing on the image data of the initial examination and follow-up examination to obtain spatially aligned multi-temporal CBCT images, and perform jaw region segmentation on the multi-temporal CBCT images to obtain a jaw region of interest mask.
[0013] The extraction module is used to extract shape features, texture features, and statistical features from the images of the initial examination and follow-up examination based on the mask of the region of interest of the jawbone, so as to obtain the feature vector of the initial examination and the feature vector of the follow-up examination.
[0014] The fusion module is used to divide the region of interest of the jawbone into a cortical bone region and a cancellous bone region based on a bone density threshold. Adaptive weighting coefficients are set for the cortical bone region and the cancellous bone region respectively based on their physiological characteristics. The weighted difference between the initial examination feature vector and the follow-up examination feature vector is calculated using the adaptive weighting coefficients to obtain a multi-temporal difference feature vector. A variance screening algorithm is used to select features from the multi-temporal difference feature vector to obtain an effective difference feature vector. The initial examination feature vector and the effective difference feature vector are concatenated and fused to generate a fused feature vector.
[0015] The classification module is used to input the fused feature vector into a fully connected neural network with attention mechanism weighting for classification and recognition. The key feature combination related to jawbone lesions is highlighted by the attention weight calculation to obtain the jawbone lesion recognition result.
[0016] Thirdly, a smart identification device for jaw lesions based on oral CBCT images is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the smart identification device for jaw lesions based on oral CBCT images to execute the aforementioned smart identification method for jaw lesions based on oral CBCT images.
[0017] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the above-described intelligent identification method for jaw lesions based on oral CBCT images.
[0018] The technical solution provided in this application achieves effective capture and utilization of dynamic evolution information of jawbone lesions through multi-temporal CBCT image registration processing and adaptive weighted difference feature fusion, overcoming the limitations of existing technologies that rely solely on single-temporal static analysis. The rigid body registration algorithm based on mutual information ensures spatial consistency between initial and follow-up images, laying the foundation for accurate difference feature calculation. Multi-dimensional feature extraction covers shape, texture, and statistical features, comprehensively quantifying the morphological and textural characteristics of jawbone tissue. The adaptive weighting mechanism sets weight coefficients according to the physiological differences between cortical and cancellous bone, highlighting the importance of the cortical bone region in lesion identification and addressing the technical deficiency of equal-weight processing ignoring anatomical heterogeneity. The variance screening algorithm effectively removes noise interference by retaining significantly changing difference features, improving the discriminative power of feature vectors. The fused feature vectors organically combine static morphological information with dynamic evolution information, providing richer and more accurate feature descriptions for subsequent intelligent identification, significantly enhancing the accuracy and robustness of jawbone lesion identification.
[0019] By automatically learning feature importance weights, intelligent identification and highlighting of key features highly correlated with jawbone lesions were achieved. Attention weight calculation quantifies the contribution of each feature dimension to lesion identification, enabling the network to adaptively focus on the most discriminative feature combinations and effectively suppress interference from irrelevant features. This feature importance assessment mechanism not only improves recognition performance but also provides clinicians with interpretable diagnostic evidence. Through weight analysis and anatomical region contribution assessment, the characteristic patterns and evolutionary laws of different types of jawbone lesions were revealed. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of an embodiment of the intelligent identification method for jawbone lesions based on oral CBCT images in this application.
[0022] Figure 2 This is a schematic diagram illustrating the comparison of multi-temporal feature extraction in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of the fully connected neural network hierarchy in the embodiments of this application;
[0024] Figure 4 This is a schematic diagram of an embodiment of the intelligent recognition system for jawbone lesions based on oral CBCT images in this application.
[0025] Figure 5 This is a schematic block diagram of the structure of the intelligent identification device for jawbone lesions based on oral CBCT images in an embodiment of the present invention. Detailed Implementation
[0026] This application provides a method and system for intelligent identification of jawbone lesions based on oral CBCT images. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the intelligent identification method for jawbone lesions based on oral CBCT images in this application includes:
[0028] Step S1: Obtain oral CBCT image data of the patient's initial examination and follow-up examination, perform registration processing on the image data of the initial examination and follow-up examination to obtain spatially aligned multi-temporal CBCT images, and perform jaw region segmentation on the multi-temporal CBCT images to obtain the jaw region of interest mask.
[0029] Step S2: Extract shape features, texture features, and statistical features from the images of the initial examination and follow-up examination based on the region of interest mask of the jawbone to obtain the feature vector of the initial examination and the feature vector of the follow-up examination.
[0030] Step S3: Divide the region of interest of the jawbone into cortical bone region and cancellous bone region according to the bone mineral density threshold. Set adaptive weight coefficients for cortical bone region and cancellous bone region respectively based on the physiological characteristics of cortical bone region and cancellous bone region. Calculate the weighted difference between the feature vector of the initial examination and the feature vector of the follow-up examination by the adaptive weight coefficient to obtain the multi-temporal difference feature vector. Use the variance screening algorithm to select features from the multi-temporal difference feature vector to obtain the effective difference feature vector. Concatenate and fuse the feature vector of the initial examination and the effective difference feature vector to generate the fused feature vector.
[0031] Step S4: Input the fused feature vector into a fully connected neural network with attention mechanism weighting for classification and recognition. The key feature combination related to jawbone lesions is highlighted by the attention weight calculation to obtain the jawbone lesion recognition result.
[0032] It is understood that the executing entity of this application can be an intelligent recognition system for jawbone lesions based on oral CBCT images, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiments use a server as an example for illustration.
[0033] Specifically, the registration process employs a rigid body registration algorithm based on mutual information. Mutual information measures the statistical dependence between two images, quantifying the correlation between images by calculating the joint probability distribution and marginal probability distribution. The registration process seeks the optimal spatial transformation parameters, including three translation parameters and three rotation parameters, by maximizing the mutual information values between the initial and follow-up images. The Powell optimization algorithm is used for iterative solution until the convergence threshold is reached with a preset accuracy. The registered multi-temporal CBCT images have the same voxel size and spatial coordinate system, solving the feature calculation error problem caused by the spatial inconsistency of temporal images in existing technologies. For jaw region segmentation, a deep convolutional neural network is used. The encoder extracts multi-scale features stepwise through convolutional and pooling layers, while the decoder upsamples and reconstructs the spatial resolution through transposed convolution. Skip connections fuse feature maps from different layers of the encoder with corresponding layers of the decoder. The network outputs a pixel-level segmentation probability map. Thresholding is used to mark pixels with a probability value greater than 0.5 as jaw regions. Morphological post-processing removes small noise regions, ultimately yielding an accurate jaw region of interest mask.
[0034] Shape features are calculated using a 3D shape analysis algorithm. Volume is obtained by counting the number of non-zero voxels within the mask and multiplying it by the physical size of a single voxel. Surface area is calculated by extracting a 3D surface mesh using the Marching Cubes algorithm. Sphericity is defined as the ratio of the surface area of an ideal sphere equal to the mask volume to the actual surface area. Compactness reflects the complexity of the shape. Texture features are calculated based on the gray-level co-occurrence matrix (GLCM), which records the co-occurrence frequency of gray-level values for pixel pairs at specific distances and directions. Contrast measures the degree of variation in gray-level differences between adjacent pixels. Correlation quantifies the linear correlation between pixel values and gray-level values. Energy represents the uniformity of the texture, and uniformity reflects the consistency of the gray-level distribution. Statistical features are obtained through gray-level histogram statistical analysis. The mean is the arithmetic mean of the gray-level values of all voxels within the region of interest in the jawbone. Standard deviation measures the dispersion of gray-level values relative to the mean. Skewness reflects the degree of skewness of the gray-level distribution relative to a symmetrical distribution, and kurtosis describes the sharpness of the gray-level distribution. All feature parameters are Z-score standardized to eliminate the influence of differences in the units of different features. The standardized features extracted from the initial examination images and follow-up examination images are concatenated and combined according to the categories to form the corresponding feature vectors.
[0035] Based on the physiological differences in jawbone anatomy, a dual-threshold segmentation algorithm labels high-density areas with a HU value greater than 1000 as cortical bone regions, medium-density areas with a HU value between 300 and 1000 as cancellous bone regions, and low-density areas with a HU value less than 300 are excluded. Cortical bone, characterized by high density and low porosity, is highly sensitive to lesions, and an adaptive weighting coefficient of 1.5 is set. Cancellous bone, with relatively lower density and higher porosity, exhibits relatively slow changes, and a weighting coefficient of 1.0 is set. The weighted difference calculation iterates through each feature parameter in the initial and follow-up feature vectors, selecting the appropriate adaptive weighting coefficient based on the cortical or cancellous bone region corresponding to the voxel. The weighted difference feature is obtained by multiplying the weighting coefficient by the absolute difference between the initial and follow-up feature values. A variance screening algorithm calculates the variance of each weighted difference feature across all samples. Features with large variance values indicate significant changes across different samples and strong discriminative ability; features with variance values exceeding 0.1 are retained as effective difference feature vectors, addressing the lack of specificity in difference feature selection in existing technologies. The fused feature vector is formed by concatenating the initial check feature vector and the effective difference feature vector in dimensional order, and contains both static morphological information and dynamic evolutionary information.
[0036] The attention-weighted fully connected neural network's data processing comprises a four-layer structure. The input layer receives all dimensions of the fused feature vector. The first hidden layer contains 256 neurons that process the input features through linear transformation and the ReLU activation function. The second hidden layer contains 128 neurons, and the third hidden layer contains 64 neurons. The output layer contains two neurons, corresponding to the lesion category and the normal category, respectively. The attention mechanism calculates feature importance between the first and second hidden layers. A learnable weight matrix maps the output of the first hidden layer to an importance score, which is then processed by a softmax normalization function to obtain attention weights. These attention weights are then multiplied element-wise with the output of the first hidden layer to achieve feature weighting, highlighting important feature dimensions highly correlated with jawbone lesions. The weighted feature representation is input to subsequent hidden layers for further nonlinear transformation. Finally, the output layer calculates the probability distribution of the lesion and normal categories using a softmax activation function. The probability value corresponding to the lesion category is used as the jawbone lesion identification probability. When the identification probability exceeds a preset threshold of 0.5, the jawbone lesion is identified as positive.
[0037] In one specific embodiment, shape features, texture features, and statistical features are extracted from images of the initial and follow-up examinations based on a region of interest mask of the jawbone, including:
[0038] Within the region of interest of the jawbone, a three-dimensional shape analysis algorithm is used to calculate the volume, surface area, sphericity, and compactness to obtain shape feature parameters;
[0039] The gray-level co-occurrence matrix is calculated in multiple directions in three-dimensional space using the sliding window technique. Contrast, correlation, energy and uniformity are extracted to obtain texture feature parameters. The sliding window size is 8×8×8 voxels.
[0040] The statistical characteristic parameters are obtained by calculating the mean, standard deviation, skewness, and kurtosis of the voxel gray values within the region of interest in the jawbone through gray-level histogram statistical analysis.
[0041] The feature parameters extracted from the initial examination images and follow-up examination images were subjected to Z-score standardization. The standardized shape feature parameters, texture feature parameters, and statistical feature parameters of the initial examination images were sequentially concatenated to generate the initial examination feature vector. The standardized shape feature parameters, texture feature parameters, and statistical feature parameters of the follow-up examination images were sequentially concatenated to generate the follow-up examination feature vector.
[0042] Specifically, multi-dimensional feature extraction and adaptive weighting mechanisms address the technical problems of insufficient comprehensiveness of single feature descriptions and the neglect of anatomical heterogeneity in jawbone processing in existing technologies. In the data processing of the 3D shape analysis algorithm, volume calculation is achieved by traversing all voxels within the mask of the jawbone region of interest, counting the number of voxels with non-zero grayscale values, and multiplying this number by the physical size of each voxel. Surface area calculation employs the Marching Cubes algorithm to reconstruct triangular meshes at the mask boundaries, obtaining the surface area value by summing the areas of all triangular faces. Sphericity calculation requires first determining the surface area of an ideal sphere with the same volume, which is calculated as 4π multiplied by the square of the sphere's radius. The sphere's radius is calculated by taking the cube root of 4 / 3π multiplied by the volume. Sphericity equals the ratio of the ideal sphere's surface area to the actual surface area; a ratio closer to 1 indicates a more regular shape. Compactness reflects the complexity of the shape and is calculated by dividing the square of the surface area by the 2 / 3 power of the volume; a smaller value indicates a more compact and regular shape. The sliding window technique plays a crucial role in texture feature extraction. An 8×8×8 voxel 3D window moves within the region of interest (ROI) of the jawbone at a fixed step size. At each window position, a set of gray-level co-occurrence matrix (GLCM) features is calculated. The voxel gray values within the window are used to construct the GLCM according to specific spatial distances and directional relationships. Matrix elements record the frequency of pixel pairs with specific gray-level value combinations. Contrast is obtained by summing the squares of the differences between row and column indices and the product of matrix elements, reflecting the drastic degree of gray-level variation between adjacent pixels. Correlation is obtained by weighted averaging the products of the row and column indices and their mean deviations, measuring the linear correlation between pixels and gray values. Energy is equal to the sum of the squares of all matrix elements, and uniformity is equal to the sum of the absolute values of the row and column index differences divided by 1.
[0043] The data processing steps for grayscale histogram statistical analysis include traversing all voxels within the region of interest in the jawbone, counting the frequency of each grayscale value to construct a grayscale distribution histogram, calculating the mean by dividing the sum of all voxel grayscale values by the total number of voxels, calculating the standard deviation by summing the squares of the differences between each voxel grayscale value and the mean, dividing by the total number of voxels, and taking the square root, calculating the skewness to measure the degree of skewness of the grayscale distribution relative to a symmetrical distribution, calculated by dividing the third central distance by the cube of the standard deviation, and calculating the kurtosis to describe the sharpness of the grayscale distribution, calculated by dividing the fourth central distance by the fourth power of the standard deviation and then subtracting 3. The Z-score standardization process is performed separately for each feature parameter. First, the mean and standard deviation of the feature are calculated across all samples. Then, the original feature value of each sample is subtracted from the mean and divided by the standard deviation to obtain the standardized feature value. The standardized feature value has a mean of 0 and a standard deviation of 1, eliminating differences in units and numerical ranges between different features. The generation process of the initial and follow-up feature vectors involves concatenating the standardized feature parameters in order of feature type, with shape feature parameters at the beginning, texture feature parameters in the middle, and statistical feature parameters at the end. The dimension of the concatenated feature vector is equal to the total number of all feature parameters.
[0044] Figure 2 This is a schematic diagram illustrating the comparison of multi-temporal feature extraction in this application embodiment. The diagram shows the numerical changes of seven main feature parameters extracted from patients with jawbone lesions during initial and follow-up examinations. The horizontal axis represents different feature types, including volume features, surface area features, sphericity features, contrast features, correlation features, mean features, and standard deviation features; the vertical axis represents the corresponding feature values. As can be seen from the diagram, volume features, surface area features, and standard deviation features significantly increased during follow-up examinations compared to the initial examination, while sphericity features and mean features showed a decreasing trend, and contrast features and correlation features slightly increased. These feature change patterns reflect the typical feature evolution patterns of jawbone lesions over time.
[0045] In one specific embodiment, the region of interest in the jawbone is divided into a cortical bone region and a cancellous bone region based on a bone mineral density threshold, including:
[0046] A dual-threshold segmentation algorithm was adopted, setting HU value of 1000 as the upper threshold of cortical bone and HU value of 300 as the lower threshold of cancellous bone. Voxel regions with HU value greater than 1000 were marked as cortical bone regions, and voxel regions with HU value between 300 and 1000 were marked as cancellous bone regions.
[0047] The adaptive weighting coefficient for the cortical bone region was set to 1.5, and the adaptive weighting coefficient for the cancellous bone region was set to 1.0.
[0048] Specifically, this paper addresses the technical problem of neglecting the anatomical heterogeneity of the jawbone in existing technologies through a dual-threshold segmentation algorithm and adaptive weight coefficient settings. The data processing of the dual-threshold segmentation algorithm is based on the distribution characteristics of HU values of different tissues in CBCT images. HU values reflect the relative relationship between tissue density and water; a HU value of 0 indicates the same density as water, a HU value of 1000 indicates the same density as cortical bone, and a HU value of -1000 indicates the same density as air. The algorithm first traverses each voxel within the mask of the region of interest in the jawbone, reads the corresponding HU value, and then classifies and labels it according to a preset threshold. Voxels with HU values greater than 1000 are labeled as cortical bone regions, voxels with HU values between 300 and 1000 are labeled as cancellous bone regions, and voxels with HU values less than 300 are labeled as soft tissue or cavity regions and excluded from subsequent processing. The threshold of 1000 for upper cortical bone is based on the high density characteristics of cortical bone, which exhibits a high HU value due to its high mineral content and low porosity. The threshold of 300 for lower cancellous bone is based on the medium density characteristics of cancellous bone, which exhibits a medium HU value due to its trabecular structure and high porosity. This segmentation method based on differences in tissue density can accurately distinguish different anatomical structures within the jawbone.
[0049] The data processing for adaptive weighting coefficient settings is based on the different sensitivity characteristics of cortical bone and cancellous bone in disease development. Cortical bone, due to its dense structure and low metabolic activity, exhibits significant density and morphological changes in the early stages of disease; therefore, a higher weighting coefficient of 1.5 is set to enhance its contribution to the difference feature calculation. Cancellous bone, due to its porous trabecular structure and high metabolic activity, changes relatively slowly and less significantly during disease progression; therefore, a lower weighting coefficient of 1.0 is set as the baseline weight. The determination of the weighting coefficient values is based on statistical analysis of a large amount of clinical CBCT imaging data. By comparing the characteristic change amplitudes of different types of jawbone lesions in the cortical and cancellous bone regions, it was found that the characteristic change amplitude in the cortical bone region was on average about 1.5 times higher than that in the cancellous bone region. Therefore, the weighting coefficient for cortical bone was set to 1.5, and the weighting coefficient for cancellous bone was set to 1.0. The weighting coefficient plays a crucial role in the subsequent weighted difference calculation. For feature parameters located in the cortical bone region, the difference is multiplied by a weighting coefficient of 1.5 to amplify the value and highlight the importance of cortical bone changes. For feature parameters located in the cancellous bone region, the difference is multiplied by a weighting coefficient of 1.0 to maintain the original value and form a benchmark comparison.
[0050] For example, in the CBCT image analysis of a patient with a maxillary periapical cyst, the dual-threshold segmentation algorithm, when processing the region of interest mask of the jawbone, found that the cyst was mainly located in the cancellous bone region, with the cyst edge involving part of the cortical bone region. The algorithm iterated through the HU value of each voxel within the mask. The HU value of the central region of the cyst was close to the air density, approximately -800 to -400, and was excluded from the analysis range. The HU value of the cancellous bone region surrounding the cyst was between 400 and 800, and was marked as a cancellous bone region. The HU value of the cortical bone region contacted by the cyst edge was between 1200 and 1500, exceeding the cortical bone threshold, and was marked as a cortical bone region. During feature extraction, volumetric features, texture features, and statistical features located in the cancellous bone region received a weight coefficient of 1.0, while the corresponding features located in the cortical bone region received a weight coefficient of 1.5. During the calculation of the difference feature, the feature changes in the cortical bone region are manifested as decreased density and morphological changes due to bone resorption caused by cyst compression. These changes are amplified by a weighting coefficient of 1.5, making the lesion signal in the cortical bone region occupy a more important position in the difference feature vector. Meanwhile, the feature changes in the cancellous bone region are relatively small and maintain their original intensity by a weighting coefficient of 1.0. The entire weighting process ensures that different anatomical regions receive corresponding importance evaluations according to their lesion sensitivity.
[0051] In one specific embodiment, a multi-temporal difference feature vector is obtained by weighting the feature vector from the initial examination and the feature vector from the follow-up examination using adaptive weighting coefficients, including:
[0052] Iterate through each feature parameter in the initial examination feature vector and the follow-up examination feature vector, and select the corresponding adaptive weight coefficient for weighting based on the cortical bone region or cancellous bone region to which the jawbone voxel belongs.
[0053] The weighted difference feature is obtained by calculating the absolute difference between the weighted initial examination feature value and the follow-up examination feature value.
[0054] All weighted difference features are arranged in order to generate a multi-temporal difference feature vector;
[0055] The variance of each weighted difference feature in the whole sample was calculated using the analysis of variance method, and the weighted difference features with a variance value greater than 0.1 were selected as effective difference feature vectors.
[0056] Specifically, adaptive weighted difference calculation and variance screening address the technical problems of lack of specificity in difference feature calculation and low efficiency in feature selection in existing technologies. The data processing of traversing feature parameters requires establishing a correspondence between feature parameters and anatomical regions. Each feature parameter carries the anatomical region identifier of its source voxel during calculation. Shape features such as volume and surface area determine their primary region by statistically analyzing the voxel contribution of different anatomical regions. Texture features determine their region identifier by analyzing the anatomical region attribution of the dominant voxel within the sliding window. Statistical features determine their region attribution by calculating the weighted contribution of voxels from different anatomical regions. The data processing logic for weight selection is based on the anatomical region identifier of the feature parameters. The algorithm reads the region identifier information of the feature parameters; if the identifier is a cortical bone region, a weight coefficient of 1.5 is selected; if the identifier is a cancellous bone region, a weight coefficient of 1.0 is selected. For feature parameters in mixed regions, the final weight coefficient is calculated by weighted averaging based on the proportion of cortical and cancellous bone voxels.
[0057] The data processing for weighted difference feature calculation involves processing each feature parameter individually. The algorithm extracts the i-th feature parameter value from the initial examination feature vector, denoted as Fi_initial, and extracts the corresponding i-th feature parameter value from the follow-up examination feature vector, denoted as Fi_followup. The absolute difference between the two is calculated to obtain the original difference feature. Then, the original difference feature is multiplied by the corresponding adaptive weight coefficient to obtain the weighted difference feature. During the weighting process, the feature difference in the cortical bone region is amplified by a factor of 1.5 to highlight its importance, while the feature difference in the cancellous bone region remains unchanged as the baseline. The generation of the multi-temporal difference feature vector involves arranging all weighted difference features in the same order as the original feature vector. The weighted difference of shape features is placed at the beginning of the vector, the weighted difference of texture features is placed in the middle, and the weighted difference of statistical features is placed at the end, ultimately forming a multi-temporal difference feature vector with the same dimension as the original feature vector.
[0058] The data processing steps of the analysis of variance (ANOVA) method are used to evaluate the discriminative power of weighted difference features. Features with large variance values indicate significant variation across different samples and strong discriminative power, while features with small variance values indicate weak variation across different samples and poor discriminative power. The algorithm calculates the variance of each weighted difference feature across all training samples. First, it calculates the mean of the feature across all samples. Then, it calculates the squared difference between the feature value and the mean for each sample. The sum of all squared differences is then divided by the sample size minus 1 to obtain the variance value. The variance threshold of 0.1 is set based on statistical rules of thumb: a variance value less than 0.1 indicates that the feature varies very little across samples and contributes limitedly to the classification task; a variance value greater than 0.1 indicates that the feature has sufficient variability to support effective classification. The generation of effective difference feature vectors involves filtering and retaining weighted difference features whose variance exceeds a threshold. The algorithm iterates through each feature in the multi-phase difference feature vector, checking whether its variance is greater than 0.1. Features that meet the condition are retained and reorganized into effective difference feature vectors, while features that do not meet the condition are removed to reduce noise interference.
[0059] In one specific embodiment, the fused feature vector is input into a fully connected neural network weighted by an attention mechanism for classification and recognition, including:
[0060] A four-layer fully connected neural network architecture is constructed. The number of neurons in the input layer is equal to the dimension of the fused feature vector. The first hidden layer contains 256 neurons, the second hidden layer contains 128 neurons, the third hidden layer contains 64 neurons, and the output layer contains 2 neurons corresponding to the lesion and normal categories.
[0061] Attention weights are introduced between the first and second hidden layers. The importance score of each feature is calculated using the learnable parameter matrix. The importance score is then normalized and used as the attention weight to weight the output of the first hidden layer.
[0062] The weighted feature representation is input into the second and third hidden layers for further feature transformation, and the jawbone lesion identification result is obtained through the Softmax activation function of the output layer.
[0063] Specifically, the attention-weighted fully connected neural network addresses the technical problems of insufficient feature importance assessment and recognition accuracy in existing technologies. The data processing of the four-layer fully connected neural network architecture begins at the input layer. The number of neurons in the input layer is dynamically set to the dimension of the fused feature vector. If the fused feature vector contains 140 feature dimensions, the input layer contains 140 neurons, each receiving the feature value at the corresponding position in the fused feature vector as the input signal. The first hidden layer contains 256 neurons, each fully connected to all neurons in the input layer through a weight matrix (256×140) and a bias vector (256×1). The output of the first hidden layer is calculated using a linear transformation and the ReLU activation function. The linear transformation multiplies the input features by the weight matrix and adds the bias vector. The ReLU activation function sets all negative values to 0, leaving positive values unchanged. The activated output has a size of 256×1. The second hidden layer contains 128 neurons, the third hidden layer contains 64 neurons, and the output layer contains 2 neurons corresponding to the lesion category and the normal category, respectively. The connection method and calculation process between each layer are similar to those of the first hidden layer. The size of the weight matrix and the bias vector are determined according to the number of neurons in the adjacent layers.
[0064] The data processing for attention weight calculation is introduced between the first and second hidden layers. The attention mechanism evaluates the importance score of each feature through a learnable parameter matrix. The learnable parameter matrix has a size of 256×1, matching the output size of the first hidden layer. The importance score is calculated by matrix multiplication of the output of the first hidden layer with the learnable parameter matrix, resulting in 256 importance score values corresponding to the outputs of the 256 neurons in the first hidden layer. Normalization is performed using the softmax function, which converts all scores into a probability distribution. Each score is divided by the sum of the exponents of all scores, resulting in a normalized attention weight vector where the sum of all elements equals 1. A larger weight value indicates a more important feature. Weighting combines the attention weights with the output of the first hidden layer through element-wise multiplication. Each element of the first hidden layer output is multiplied by its corresponding attention weight value, resulting in a weighted feature representation. This weighted feature representation highlights the feature dimensions most relevant to jawbone lesion identification and suppresses the influence of unimportant features.
[0065] The feature transformation data processing involves inputting the weighted feature representation into the second and third hidden layers for further processing. The second hidden layer receives the weighted 256-dimensional feature representation, performs a linear transformation using a 128×256 weight matrix and a 128×1 bias vector, and then applies the ReLU activation function to obtain a 128-dimensional output. The third hidden layer receives the 128-dimensional output from the second hidden layer, performs a linear transformation using a 64×128 weight matrix and a 64×1 bias vector, and then applies ReLU activation to obtain a 64-dimensional high-level feature representation. The output layer receives the 64-dimensional feature representation from the third hidden layer, performs a linear transformation using a 2×64 weight matrix and a 2×1 bias vector, and obtains a 2-dimensional output vector representing the predicted scores for the lesion and normal categories, respectively. The Softmax activation function transforms the 2-dimensional score vector of the output layer into a probability distribution. The probability value for each category is calculated by dividing the exponent of the category score by the sum of the exponents of all category scores. The sum of the probabilities of the lesion and normal categories equals 1, and the category with the larger probability value is taken as the jawbone lesion identification result.
[0066] Figure 3 This is a schematic diagram of the hierarchical structure of the fully connected neural network in an embodiment of this application. The diagram illustrates the complete architecture of the attention-weighted fully connected neural network in this application. The horizontal axis represents the five hierarchical structures of the network, and the vertical axis represents the number of neurons in each layer. As shown in the diagram, the input layer contains 140 neurons for receiving fused feature vectors, the first hidden layer contains 256 neurons for feature dimension expansion, the second hidden layer contains 128 neurons for feature compression, the third hidden layer contains 64 neurons for further refining high-level features, and the output layer contains 2 neurons corresponding to the lesion and normal categories, respectively. The entire network architecture adopts a progressively decreasing layer design pattern. An attention mechanism is introduced between the first and second hidden layers, and hierarchical feature transformation achieves gradual abstraction from low-level features to high-level semantic features, completing the intelligent recognition task of jawbone lesions.
[0067] In one specific embodiment, attention weighting is used to calculate a combination of key features relevant to jawbone lesions, including:
[0068] Calculate the attention weight value corresponding to each feature dimension in the fused feature vector, and sort the attention weight values in descending order to obtain the feature importance ranking;
[0069] Set the weight threshold to 0.05, and mark features with attention weight values greater than the weight threshold as key features;
[0070] Based on the distribution of key features in the cortical and cancellous bone regions, the contribution of different anatomical regions to lesion identification was analyzed.
[0071] The top 20 key features with the highest weight values are grouped according to their anatomical location and feature type to generate key feature combinations that are highly correlated with jawbone lesions.
[0072] Specifically, attention weight analysis and key feature combination generation address the technical problems of insufficient feature importance quantification and interpretability in existing technologies. The data processing for calculating attention weight values is based on the output of the attention mechanism in the neural network. Each feature dimension corresponds to an attention weight value, ranging from 0 to 1; a larger value indicates a more significant contribution of the feature to jawbone lesion identification. Feature importance ranking is achieved through a descending order algorithm. The algorithm traverses all feature dimensions in the fused feature vector, reads the attention weight value corresponding to each feature, and uses quicksort or heapsort algorithms to sort the weight values from largest to smallest. Simultaneously, it records the original feature index corresponding to each weight value. The sorted result forms a feature importance ranking list, with the first element corresponding to the feature with the highest weight value and the last element corresponding to the feature with the lowest weight value. The weight threshold of 0.05 is set based on the statistical significance test principle. A weight value less than 0.05 indicates that the feature has a weak impact on the classification decision, while a weight value greater than 0.05 indicates that the feature has significant discriminative power. The key feature marking process involves traversing the feature importance ranking list and checking whether the weight value of each feature is greater than the threshold of 0.05. Features that meet the condition are marked as key features and saved in the key feature list.
[0073] The data processing for anatomical region distribution analysis requires combining the anatomical attribution information and weight values of features. Each key feature carries an anatomical region identifier, including cortical bone region identifiers, cancellous bone region identifiers, or mixed region identifiers. The analysis process statistically analyzes the sum of the weight values of key features in the cortical bone region, the sum of the weight values of key features in the cancellous bone region, and the distribution of the number of key features in each region. The contribution of each region to lesion identification is assessed by comparing the cumulative weight values and feature quantity distributions across different anatomical regions. The contribution of the cortical bone region is equal to the sum of the weight values of all key features in that region divided by the total weight values of all key features. The contribution of the cancellous bone region is calculated in the same way, and the sum of the contributions of the two regions equals 1. A higher contribution indicates that the anatomical region plays a more important role in identifying jawbone lesions. The selection of the first 20 key features is based on the weight value ranking results. Starting from the beginning of the feature importance ranking list, the 20 features with the highest weight values are selected sequentially. If the total number of key features is less than 20, all key features are selected.
[0074] The data processing for key feature combinations involves dual grouping based on anatomical location and feature type. Anatomical location grouping categorizes the top 20 key features into cortical bone and cancellous bone feature groups according to their anatomical region identifiers. Feature type grouping further divides features within each anatomical region into shape, texture, and statistical feature subgroups. This grouping results in a hierarchical key feature combination structure: the first layer is based on anatomical location, and the second layer is subdivided by feature type. Each subgroup contains key features with the same anatomical affiliation and feature type. The key feature combinations are weighted to calculate the overall importance score for each subgroup. The subgroup score equals the average of all feature weights within that group multiplied by the square root of the number of features in that group. This square root factor balances the impact of feature quantity on combination importance, preventing combinations with a large number of features from receiving excessively high scores.
[0075] In one specific embodiment, the jawbone lesion identification result is obtained by calculating attention weights to highlight key feature combinations related to jawbone lesions, including:
[0076] The fused feature vector, weighted by attention weights, is input into two neurons in the output layer. The probability distributions of lesion and normal categories are calculated through linear transformation and the Softmax activation function.
[0077] Extract the probability value corresponding to the lesion category as the jawbone lesion identification probability; when the jawbone lesion identification probability is greater than the preset threshold of 0.5, the sample is determined to be a positive jawbone lesion, otherwise it is determined to be negative;
[0078] Simultaneously, a confidence score is calculated for the identification results. The confidence score is equal to the maximum of the lesion probability and the normal probability, and a comprehensive diagnostic report is generated that includes the jawbone lesion identification results, the confidence score, and attention weight analysis.
[0079] Specifically, this method addresses the technical problem of lacking quantitative recognition results and reliability evaluation in existing technologies by employing probability calculation and confidence assessment. The data processing for output layer probability calculation takes a fused feature vector weighted by attention as input. This feature vector has already undergone feature transformation and weighting by the attention mechanism in the first three hidden layers, highlighting the features most relevant to jawbone lesion recognition. The output layer contains two neurons corresponding to the lesion category and the normal category, respectively. A linear transformation maps the 64-dimensional input features to a 2-dimensional output vector using a 2×64 weight matrix and a 2×1 bias vector. Each element in the weight matrix represents the connection strength between the input feature and the output category, and the bias vector is used to adjust the activation threshold of the neurons. The linear transformation calculation process performs matrix multiplication between the input feature vector and the weight matrix, and adds the bias vector to obtain an unnormalized category score. The lesion category score reflects the original prediction strength that the input sample belongs to the lesion category, and the normal category score reflects the original prediction strength that it belongs to the normal category. The Softmax activation function transforms unnormalized class scores into a probability distribution. The calculation process first performs an exponential operation on each class score, then divides the exponential value of the lesion class by the sum of the exponential values of all classes to obtain the probability of the lesion class, and divides the exponential value of the normal class by the sum to obtain the probability of the normal class. The sum of the two probability values equals 1.
[0080] The data processing for jawbone lesion identification probability extraction directly reads the probability value corresponding to the lesion category from the Softmax output. This probability value ranges from 0 to 1; the closer the value is to 1, the higher the confidence level that the sample belongs to the lesion category, and the closer the value is to 0, the lower the confidence level. The judgment rule is based on a preset threshold of 0.5 for binary classification. The threshold of 0.5 is chosen based on the principle of equal probability boundary in statistics. When the lesion probability is greater than 0.5, it indicates that the model considers the sample more likely to be a lesion and classifies it as a positive jawbone lesion; when the lesion probability is less than or equal to 0.5, it indicates that the model considers the sample more likely to be a normal sample and classifies it as a negative. The judgment process is implemented through simple numerical comparison. The algorithm reads the lesion identification probability value and compares it with the threshold of 0.5, outputting a positive or negative classification result. The confidence score is calculated using the maximum probability value as the evaluation index. The confidence score is equal to the larger of the lesion probability and the normal probability. When the lesion probability is 0.8 and the normal probability is 0.2, the confidence score is 0.8. When the lesion probability is 0.3 and the normal probability is 0.7, the confidence score is 0.7. The closer the confidence score is to 1, the higher the reliability of the identification result. The closer it is to 0.5, the greater the uncertainty of the identification result.
[0081] The data processing for generating the comprehensive diagnostic report integrates three components: jawbone lesion identification results, confidence scores, and attention weight analysis. The identification results section includes positive or negative classifications and corresponding probability values. The confidence score provides a quantitative indicator of the reliability of the identification results. The attention weight analysis includes the importance ranking of key features and anatomical region contribution analysis. The report format uses structured text and includes four parts: basic patient information, imaging examination time, identification result summary, and detailed analysis. The identification result summary directly provides the positive or negative conclusion and confidence score. The detailed analysis lists the top 10 key features with the highest weight values and their corresponding anatomical region affiliations. The anatomical region contribution analysis demonstrates the relative importance of the cortical bone region and the cancellous bone region in the identification process.
[0082] The above describes the intelligent identification method for jawbone lesions based on oral CBCT images in the embodiments of this application. The following describes the intelligent identification system for jawbone lesions based on oral CBCT images in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 4 One embodiment of the intelligent recognition system for jawbone lesions based on oral CBCT images in this application includes:
[0083] The registration module is used to acquire oral CBCT image data of the patient's initial examination and follow-up examination, perform registration processing on the image data of the initial examination and follow-up examination to obtain spatially aligned multi-temporal CBCT images, and perform jaw region segmentation on the multi-temporal CBCT images to obtain a jaw region of interest mask.
[0084] The extraction module is used to extract shape features, texture features, and statistical features from the images of the initial examination and follow-up examination based on the mask of the region of interest of the jawbone, so as to obtain the feature vector of the initial examination and the feature vector of the follow-up examination.
[0085] The fusion module is used to divide the region of interest of the jawbone into a cortical bone region and a cancellous bone region based on a bone density threshold. Adaptive weighting coefficients are set for the cortical bone region and the cancellous bone region respectively based on their physiological characteristics. The weighted difference between the initial examination feature vector and the follow-up examination feature vector is calculated using the adaptive weighting coefficients to obtain a multi-temporal difference feature vector. A variance screening algorithm is used to select features from the multi-temporal difference feature vector to obtain an effective difference feature vector. The initial examination feature vector and the effective difference feature vector are concatenated and fused to generate a fused feature vector.
[0086] The classification module is used to input the fused feature vector into a fully connected neural network with attention mechanism weighting for classification and recognition. The key feature combination related to jawbone lesions is highlighted by the attention weight calculation to obtain the jawbone lesion recognition result.
[0087] above Figure 4 The intelligent recognition system for jaw lesions based on oral CBCT images in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The intelligent recognition device for jaw lesions based on oral CBCT images in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0088] Reference Figure 5 This invention also provides an intelligent identification device for jawbone lesions based on oral CBCT images. This intelligent identification device for jawbone lesions based on oral CBCT images can be a server, and its internal structure can be as follows: Figure 5 As shown, the intelligent recognition device for jawbone lesions based on oral CBCT images includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computational and control capabilities. The memory of the intelligent recognition device for jawbone lesions based on oral CBCT images includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the intelligent recognition device for jawbone lesions based on oral CBCT images stores the data corresponding to this embodiment. The network interface of the intelligent recognition device for jawbone lesions based on oral CBCT images is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0089] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the intelligent recognition device for jaw lesions based on oral CBCT images to which the present invention is applied.
[0090] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the intelligent identification method for jaw lesions based on oral CBCT images.
[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a smart jawbone lesion recognition device based on oral CBCT images (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0093] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent identification of jaw bone lesions based on oral CBCT images, characterized in that, The method comprises: Obtaining CBCT image data of the patient's initial examination and follow-up examination, registering the image data of the initial examination and follow-up examination to obtain a spatially aligned multi-phase CBCT image, and segmenting the jawbone region of the multi-phase CBCT image to obtain a jawbone region of interest mask; Extracting shape features, texture features and statistical features in the image of the initial examination and follow-up examination according to the jawbone region of interest mask to obtain an initial examination feature vector and a follow-up examination feature vector; According to the bone density threshold, the jawbone region of interest is divided into cortical bone region and cancellous bone region, and adaptive weight coefficients are set based on the physiological characteristics of the cortical bone region and the cancellous bone region, and the initial examination feature vector and the follow-up examination feature vector are weighted and difference calculated by the adaptive weight coefficients to obtain a multi-phase difference feature vector, including: traversing each feature parameter in the initial examination feature vector and the follow-up examination feature vector, selecting the corresponding adaptive weight coefficient for weighting processing according to the cortical bone region or the cancellous bone region to which the feature parameter corresponds; the absolute difference between the weighted initial examination feature value and the follow-up examination feature value is calculated to obtain a weighted difference feature; all weighted difference features are arranged in order to generate the multi-phase difference feature vector; The variance screening algorithm is used for feature selection of the multi-phase difference feature vector, the variance value of each weighted difference feature in the whole sample is calculated, and the weighted difference features with a variance value greater than 0.1 are selected as effective difference feature vectors, and the initial examination feature vector and the effective difference feature vector are spliced and fused to generate a fusion feature vector; The fusion feature vector is input into an attention mechanism weighted fully connected neural network for classification and recognition, and the key feature combination related to the jawbone lesion is highlighted through attention weight calculation to obtain a jawbone lesion recognition result.
2. The method of claim 1, wherein, The shape features, texture features and statistical features extracted in the image of the initial examination and follow-up examination according to the jawbone region of interest mask comprise: The three-dimensional shape analysis algorithm is used to calculate the volume, surface area, sphericity and compactness in the jawbone region of interest to obtain shape feature parameters; The sliding window technology is used to calculate the gray level co-occurrence matrix in multiple directions in three-dimensional space, and the contrast, correlation, energy and uniformity are extracted to obtain texture feature parameters, and the sliding window size is 8*8*8 voxels; The mean, standard deviation, skewness and kurtosis of the voxel gray value in the jawbone region of interest are calculated by gray histogram statistical analysis to obtain statistical feature parameters; The feature parameters extracted from the initial examination image and the follow-up examination image are subjected to Z-score standardization processing, the shape feature parameters, texture feature parameters and statistical feature parameters of the initial examination image after standardization are sequentially spliced and combined to generate the initial examination feature vector, and the shape feature parameters, texture feature parameters and statistical feature parameters of the follow-up examination image after standardization are sequentially spliced and combined to generate the follow-up examination feature vector.
3. The method of claim 1, wherein, The jawbone region of interest is divided into cortical bone region and cancellous bone region according to the bone density threshold, comprising: The double-threshold segmentation algorithm is adopted, HU value 1000 is set as the upper threshold of cortical bone, HU value 300 is set as the lower threshold of cancellous bone, the voxel region with HU value greater than 1000 is marked as the cortical bone region, and the voxel region with HU value between 300 and 1000 is marked as the cancellous bone region; The adaptive weight coefficient of the cortical bone region is set as 1.5, and the adaptive weight coefficient of the cancellous bone region is set as 1.
0.
4. The method of claim 1, wherein, The fusion feature vector is input into the attention mechanism weighted full connection neural network for classification and recognition, which comprises: A four-layer full connection neural network architecture is constructed, the number of input layer neurons is equal to the dimension of the fusion feature vector, the first hidden layer contains 256 neurons, the second hidden layer contains 128 neurons, the third hidden layer contains 64 neurons, and the output layer contains 2 neurons corresponding to the lesion and normal categories; Attention weight calculation is introduced between the first hidden layer and the second hidden layer, the importance score of each feature is calculated through a learnable parameter matrix, and the importance score is normalized and used as the attention weight to weight the output of the first hidden layer; The weighted feature representation is input into the second hidden layer and the third hidden layer for further feature transformation, and the jaw bone lesion recognition result is obtained through the Softmax activation function of the output layer.
5. The method of claim 1, wherein, The key feature combination related to the jaw bone lesion is highlighted through attention weight calculation, which comprises: The attention weight value corresponding to each feature dimension in the fusion feature vector is calculated, the attention weight values are arranged in descending order to obtain the feature importance ranking; The feature with an attention weight value greater than the weight threshold of 0.05 is marked as a key feature; According to the distribution of the key features in the cortical bone region and the cancellous bone region, the contribution degree of different anatomical regions to the lesion recognition is analyzed; The top 20 key features with the highest weight values are grouped according to their anatomical positions and feature types to generate a key feature combination highly related to the jaw bone lesion.
6. The method of claim 1, wherein, The jaw bone lesion recognition result is obtained by highlighting the key feature combination related to the jaw bone lesion through attention weight calculation, which comprises: The fusion feature vector weighted by the attention weight is input into the 2 neurons of the output layer, and the probability distribution of the lesion category and the normal category is calculated through linear transformation and Softmax activation function; The probability value corresponding to the lesion category is extracted as the jaw bone lesion recognition probability; when the jaw bone lesion recognition probability is greater than the preset threshold 0.5, the sample is determined as jaw bone lesion positive, otherwise it is determined as negative; The confidence score of the recognition result is calculated at the same time, the confidence score is equal to the maximum value of the lesion probability and the normal probability, and a comprehensive diagnosis report containing the jaw bone lesion recognition result, the confidence score and the attention weight analysis is generated.
7. A jaw bone lesion intelligent identification system based on oral CBCT images, characterized in that, The jaw bone lesion intelligent recognition system based on oral CBCT images is used to implement the jaw bone lesion intelligent recognition method based on oral CBCT images as claimed in any one of claims 1 to 6, which comprises: The registration module is configured to acquire oral CBCT image data of a first examination and a follow-up examination of a patient, perform registration processing on the image data of the first examination and the follow-up examination to obtain multi-phase CBCT images that are spatially aligned, and perform jawbone region segmentation on the multi-phase CBCT images to obtain a jawbone region of interest mask. The extraction module is configured to extract shape features, texture features, and statistical features from the image of the first examination and the image of the follow-up examination according to the jawbone region of interest mask to obtain a first examination feature vector and a follow-up examination feature vector. The fusion module is configured to divide the jawbone region of interest into a cortical bone region and a cancellous bone region according to a bone density threshold, set adaptive weight coefficients based on physiological characteristics of the cortical bone region and the cancellous bone region, and perform weighted difference calculation on the first examination feature vector and the follow-up examination feature vector by using the adaptive weight coefficients to obtain a multi-phase difference feature vector. The weighted difference calculation includes: traversing each feature parameter in the first examination feature vector and the follow-up examination feature vector, selecting a corresponding adaptive weight coefficient for weighted processing according to whether the jawbone voxel corresponding to the feature parameter belongs to the cortical bone region or the cancellous bone region, calculating an absolute difference between the weighted first examination feature value and the weighted follow-up examination feature value to obtain a weighted difference feature, and arranging all the weighted difference features in sequence to generate the multi-phase difference feature vector. The variance screening algorithm is used to perform feature selection on the multi-phase difference feature vector, the variance value of each weighted difference feature in all samples is calculated, and the weighted difference features with a variance value greater than 0.1 are selected as effective difference feature vectors. The first examination feature vector and the effective difference feature vector are spliced and fused to generate a fusion feature vector. The classification module is configured to input the fusion feature vector into an attention mechanism weighted fully connected neural network for classification and recognition, calculate key feature combinations related to jawbone lesions by using attention weights, and obtain a jawbone lesion recognition result.
8. A jaw bone lesion intelligent identification device based on oral CBCT images, characterized in that, The computer program is run on the processor to implement the jawbone lesion intelligent recognition method based on oral CBCT images according to any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is run on the processor to implement the jawbone lesion intelligent recognition method based on oral CBCT images according to any one of claims 1 to 6.
Citation Information
Patent Citations
Oral cavity CT mandibular neural tube segmentation method based on neural network
CN110738661A
Three-dimensional stomatognathic model reconstruction system based on multi-modal data fusion
CN120411404A