A cervical vertebrae segmentation and key point detection method based on diffusion model

Through the diffusion model-based cervical vertebra segmentation and key point detection methods, combined with biomechanics and infrared thermal imaging technology, the problem of insufficient reliance on subjective experience and image processing accuracy in traditional methods is solved, and more accurate cervical vertebrae age assessment and the formulation of personalized treatment plans are achieved.

CN119180954BActive Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411219880.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-05-06
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

The traditional cervical vertebrae age staging method relies on the subjective experience of doctors, is inefficient and difficult to quantify, and existing medical image processing techniques are difficult to effectively process noise-free cervical vertebrae images, resulting in limited accuracy of segmentation and detection.

Method used

The cervical vertebra segmentation and key point detection method based on the diffusion model are adopted. By obtaining the skull lateral film image dataset and pre-processing, combining biomechanical parameters and infrared thermal imaging technology, multimodal information is integrated, and the Transformer and CNN models are used for segmentation and detection, and the segmentation results are iteratively optimized through the diffusion model.

Benefits of technology

It improves the quality of cervical vertebrae images, enhances the accuracy of segmentation and key point detection, provides more accurate data support for cervical vertebrae age staging, helps doctors objectively evaluate patients' growth and development level and formulate personalized orthodontic treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180954B_ABST
    Figure CN119180954B_ABST
Patent Text Reader

Abstract

The present invention provides a cervical vertebra segmentation and key point detection method based on a diffusion model, which relates to the technical field of medical image processing. The present invention performs denoising on cervical vertebra images through a diffusion model, thereby improving image quality and laying a foundation for subsequent segmentation and key point detection; the model combined with deep learning technology can accurately extract cervical vertebrae areas and key point positions, providing reliable data support for bone age staging; bone age staging based on these results can objectively and quantitatively evaluate the patient's growth and development level, and assist doctors in formulating personalized orthodontic treatment plans. In addition, the adaptive segmentation model of the stress distribution map generated by the biomechanical model and the pre-processed image can adjust the segmentation strategy according to individual differences, thereby improving accuracy and robustness; at the same time, parameters such as blood flow velocity and temperature distribution are incorporated into the machine learning model, providing new indicators for cervical bone age staging and enhancing the personalization and accuracy of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a cervical vertebra segmentation and key point detection method based on a diffusion model. Background Art

[0002] In the field of medical diagnosis and orthodontic treatment, the patient's growth and development level plays a vital role. It is an indispensable key to formulating personalized treatment plans. Accurate assessment of growth and development not only determines the best time for orthodontic treatment, but is also directly related to the choice of treatment strategy and the durability of treatment effects. However, traditional assessment methods rely too much on the patient's actual age, which obviously has limitations, because the growth and development rate of each individual is different, and patients in the same age group may show completely different growth and development levels.

[0003] In contrast, cervical vertebrae bone age, as a direct and objective indicator of growth and development, can more accurately reflect the patient's growth status; therefore, accurate judgment of cervical vertebrae bone age staging is of irreplaceable importance for evaluating the patient's growth and development and formulating treatment plans; but unfortunately, the traditional cervical vertebrae bone age staging method relies too much on the subjective experience of doctors, is inefficient and difficult to quantify, which greatly reduces its effectiveness in practical applications; although the vigorous development of deep learning technology has brought revolutionary progress in the field of medical image analysis; however, when processing cervical vertebrae images, due to the particularity of lateral cranial radiographs, images often have problems such as low contrast, blurred boundaries and artifacts; in addition, the anatomical structure of the cervical spine region is complex, and the vertebrae overlap and contact frequently, making it particularly difficult to accurately segment the subtle structures and blurred boundaries in the image; at present, the existing cervical vertebrae bone image segmentation and key point detection methods often perform unsatisfactorily when dealing with noisy images, and are difficult to directly apply to cervical vertebrae bone age staging.

[0004] At the same time, existing technologies mainly rely on traditional image processing methods and deep learning models, but usually lack effective combination of context information and local features; traditional image segmentation methods such as threshold method and edge detection method are effective in simple cases, but often have poor effects when processing complex or blurred images. At the same time, the segmentation model based on deep learning is highly dependent on the data set during training, and is prone to overfitting, resulting in reduced accuracy of the segmentation results;

[0005] In addition, these methods usually fail to effectively combine the dynamic biomechanical information and microcirculatory blood flow characteristics of the cervical spine region, resulting in limited accuracy of segmentation and detection. In particular, the existing models are difficult to fully reflect the physiological state of the cervical spine region in the absence of biomechanical parameters and microcirculatory blood flow information, thus affecting bone age staging and other important clinical evaluations. In addition, traditional methods usually ignore the fact that dynamic information such as vertebral stress distribution and motion trajectory often have an important impact on the segmentation results. The lack of integration of biomechanical parameters leads to the model's inability to fully reflect the complexity of the actual biological structure, thus affecting subsequent key point detection and bone age staging analysis, and ignoring the potential of infrared thermal imaging technology in revealing the microcirculatory blood flow characteristics of the cervical spine region; infrared thermal imaging can non-invasively provide key parameters such as microcirculatory blood flow velocity and temperature distribution, which are crucial in evaluating the functional status and pathological changes of the cervical spine. However, these key temperature gradient and blood flow velocity characteristics are often ignored or underutilized in existing segmentation and detection models;

[0006] The above information disclosed in the above Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to one of ordinary skill in the art. Summary of the invention

[0007] The purpose of the present invention is to provide a cervical vertebra segmentation and key point detection method based on a diffusion model to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A method for cervical vertebrae segmentation and key point detection based on a diffusion model, the specific steps comprising:

[0010] Acquire a lateral head radiograph image dataset, preprocess the image dataset, and divide the preprocessed image dataset into a training set and a test set;

[0011] Collect biomechanical parameters of the cervical spine area, including stress distribution, rotation angle and displacement data between vertebrae, and integrate the collected biomechanical data with the preprocessed image data set to form an enhanced feature set;

[0012] Obtain an enhanced feature set and input it into a cervical vertebra segmentation model based on Transformer and CNN. This model uses the global context information of Transformer and the local feature extraction capability of CNN, and introduces a diffusion model to gradually refine the segmentation results through iterative optimization, thereby segmenting the cervical vertebra.

[0013] At the same time, the stress distribution and motion trajectory of the vertebrae are combined with the segmentation model, and the model parameters are updated in real time using an adaptive weight adjustment mechanism;

[0014] Perform multi-scale feature extraction on the training set and test set of preprocessed images, upsample and fuse features of different scales to generate key point heat maps;

[0015] Infrared thermal imaging technology is used to obtain the microcirculatory blood flow characteristics of the cervical spine area, and key parameters in the microcirculatory blood flow characteristics are extracted. The key parameters include blood flow velocity and temperature distribution. The key parameters are integrated into the key point thermal map to obtain an enhanced thermal map.

[0016] A multivariate regression model based on temperature gradient and blood flow velocity was used in combination with image segmentation results to perform multi-scale analysis of cervical vertebrae bone age staging.

[0017] Uncertainty distribution is performed based on the enhanced heat map. The Gaussian distribution parameters of each key point are fitted from the enhanced heat map using the maximum likelihood estimation method. The Gaussian distribution parameters of multiple key points are used to construct a Gaussian mixture model as the initial uncertainty distribution.

[0018] Construct a learnable Gaussian mixture model to simulate the diffusion process, capture the uncertainty evolution characteristics of key point positions, use the parameters of the Gaussian mixture model as input, and iteratively simulate the dynamic evolution of uncertainty through the learnable diffusion model;

[0019] The training set and test set included in the preprocessed image are input into the background encoder to extract the global semantic feature vector and local spatial features, and the attention mechanism is used to fuse the features to obtain the final fused features;

[0020] The fused features are used as input, and the global and local context information of the image is used to constrain the generation process of key points through back diffusion, thereby realizing key point detection;

[0021] The current lateral head radiograph image dataset to be tested is input into the diffusion model constructed above, and the model is used to perform cervical vertebrae segmentation and key point detection on the image dataset to generate segmentation results and key point positions; at the same time, corresponding biomechanical analysis and health assessment are performed based on the detection results to provide accurate auxiliary information for clinical diagnosis.

[0022] Furthermore, a lateral head radiograph image dataset is obtained and preprocessed, specifically including:

[0023] The preprocessing includes converting the lateral head radiograph image into a unified RGB format; then the lateral head radiograph image is normalized;

[0024] The cervical vertebrae of each lateral head radiograph image sample in the image data set are labeled, including the labeling of the key points and segmented areas of the cervical vertebrae, to obtain a plurality of data including labeled cervical vertebrae area images.

[0025] Furthermore, the collected biomechanical data are integrated with the preprocessed image dataset to form an enhanced feature set, which specifically includes:

[0026] For the collection of biomechanical parameters:

[0027] Finite element analysis software was used to simulate the biomechanical behavior of the cervical spine region;

[0028] Using the geometric information extracted from the image data set, a three-dimensional cervical spine model is constructed; and the mechanical properties of each material are set, and the mechanical properties are Young's modulus and Poisson's ratio;

[0029] Set boundary conditions and loading conditions to simulate the motion and load of the cervical spine under actual physiological conditions;

[0030] Finite element analysis software was used to perform finite element analysis on the three-dimensional cervical spine model to calculate the stress distribution, rotation angle, and displacement of each vertebra under different load and motion conditions;

[0031] Stress distribution calculation formula:

[0032]

[0033] Among them, σ′ is stress, Fz is force, and As is the action area;

[0034] The displacement and rotation angle are calculated using the following formulas:

[0035]

[0036] Among them, Δx and Δy are displacement differences, θ is the rotation angle, and x new is the horizontal coordinate of the current position, x original is the horizontal coordinate of the original position, y new is the ordinate of the current position, y original is the ordinate of the original position; arctan is the inverse tangent function;

[0037] For the integration of biomechanical parameter data and image data:

[0038] Projecting the three-dimensional biomechanical data obtained by finite element analysis onto the pre-processed two-dimensional lateral head radiograph image;

[0039] For the generation of enhanced feature sets:

[0040] Each preprocessed image pixel is fused with the corresponding biomechanical parameters to generate an enhanced feature vector v containing multimodal information, which is defined as:

[0041] v=[I(x,y),σ′(x,y),θ(x,y),Δx(x,y),Δy(x,y)]

[0042] Where I(x,y) is the image pixel value, σ′(x,y) is the stress value, θ(x,y) is the rotation angle, Δx(x,y) and Δy(x,y) are the displacement differences;

[0043] The enhanced feature set is normalized so that all features have the same scale.

[0044] Furthermore, the key parameters of microcirculatory blood flow characteristics are extracted and integrated into the key point heat map; a multivariate regression model based on temperature gradient and blood flow velocity is used, combined with image segmentation results, to perform multi-scale analysis of cervical vertebrae bone age staging; specifically, the following are included:

[0045] For infrared thermal imaging data acquisition and preprocessing:

[0046] Use high-resolution infrared thermal imaging equipment to image the patient's cervical spine area and obtain thermal imaging images of microcirculatory blood flow characteristics;

[0047] Register the thermal imaging image with the cervical vertebrae segmentation result image;

[0048] For extracting microcirculatory blood flow characteristics:

[0049] Extract the temperature distribution data of the cervical spine area from the thermal imaging image; for each pixel point, define the temperature matrix T(x′,y′) to represent the temperature value at the position (x′,y′);

[0050] The blood flow velocity is estimated based on the gradient of temperature distribution; first calculate the temperature gradient

[0051]

[0052] A blood flow velocity estimation model based on temperature gradient is established, and the blood flow velocity v(x′, y′) is set to be proportional to the temperature gradient, expressed as:

[0053]

[0054] Among them, k′ is the proportionality coefficient, which represents the relationship between temperature gradient and blood flow velocity;

[0055] For integrating key parameters into key point heatmaps:

[0056] A multivariate regression model based on temperature gradient and blood flow velocity was constructed. The formula of the multivariate regression model was:

[0057] R(x′,y′)=α·v(x′,y′)+β·T(x′,y′)+γ

[0058] Among them, R(x′, y′) is the thermal map response value, v(x′, y′) is the blood flow velocity, T(x′, y′) is the temperature value, and α, β, and γ are regression coefficients;

[0059] The output value of the regression model is fused with the previously generated key point heat map to obtain the final enhanced heat map; the enhanced heat map is expressed as:

[0060] H enhanced (x′,y′)=Sigmoid(H original (x′,y′)+R(x′,y′))

[0061] Among them, H enhanced (x′, y′) is the enhanced heat map, H original (x′, y′) is the initial key point heat map, R(x′, y′) is the output value of the multivariate regression model; the Sigmoid function is used to normalize the output value of the final heat map to between 0 and 1;

[0062] Multi-scale analysis of cervical vertebrae bone age staging:

[0063] The enhanced thermal map is combined with the segmented cervical vertebra image features to form a joint feature vector; the vector includes the geometric shape features, microcirculation blood flow features and temperature distribution features of the cervical vertebra; it is combined by the following formula:

[0064] F combined =[F bone ,H enhanced ,T(x′,y′),v(x′,y′)]

[0065] Among them, F bone is the shape feature of the cervical vertebra segmentation image, F combined is the joint eigenvector;

[0066] The formed joint feature vector was used to predict cervical vertebrae bone age stage using a regression model based on multi-scale analysis;

[0067] The formula of the multiscale regression model is:

[0068] Age predicted =w1·F bone +w2·H enhanced +w3·T(x′,y′)+w4·v(x′,y′)+b″

[0069] Among them, Age predicted is the predicted bone age, w1, w2, w3, w4 are the learned positive weight coefficients, b″ is the bias term,

[0070] Furthermore, uncertainty distribution is performed based on the enhanced heat map. The Gaussian distribution parameters of each key point are fitted from the enhanced heat map using the maximum likelihood estimation method. The Gaussian distribution parameters of multiple key points are combined into a Gaussian mixture model as the initial uncertainty distribution, which specifically includes:

[0071] Extract the candidate position of the key point from the enhanced heat map, that is, the position with the largest probability value in the enhanced heat map. The coordinates of the candidate position (x′, y′) are the preliminary estimated positions of the key points;

[0072] Collect the 2D coordinates of all key points into a data set, where n′ is the total number of key points;

[0073] The obtained Gaussian distribution parameters of each key point are combined into a Gaussian mixture model, hereinafter referred to as GMM, to obtain the initial estimate H of the uncertainty distribution of the key point position extracted from the enhanced heat map K .

[0074] Furthermore, a learnable Gaussian mixture model is constructed to simulate the diffusion process and capture the uncertainty evolution characteristics of key point positions. The initial Gaussian mixture model is used as input and the learnable diffusion model is used to iteratively simulate the dynamic evolution of uncertainty, including:

[0075] Taking the initial Gaussian mixture model as input, the learnable diffusion model is used to iteratively simulate the dynamic evolution of uncertainty.

[0076] Using the initial uncertainty distribution H K Parameters, run the expectation maximization algorithm of Gaussian mixture noise, and then fit the covariance matrix of each key point position;

[0077] The number of Gaussian components of the Gaussian mixture model is defined as the number of mixed components used to represent the uncertainty distribution of key points. Secondly, the EM algorithm is used to optimize the Gaussian mixture model parameter φ GMM ;

[0078] Get the manually annotated key point distribution H0 and set the position of each key point They are all determined cervical key point coordinate values. The position of each key point is calculated using the fitted GMM model φGMM The corresponding Gaussian distribution parameters;

[0079] Perform a forward diffusion process on H0, and after K steps, define the final key point uncertainty distribution Equal to the fitted GMM distribution φ GMM .

[0080] The preprocessed image is input into the background encoder to extract the global semantic feature vector and local spatial features, and the attention mechanism is used to fuse the features to obtain the final fused features, which include:

[0081] The background encoder uses a Transformer-based network structure to convert the preprocessed image into 2D patches through the embedding layer and flatten each patch to form an input sequence X = [X1, X2, ..., X N2 ], N2 is the total number of patches, each sequence is first mapped through a linear layer to convert it into a feature vector of fixed dimension d2, and the sine and cosine functions are used to add position encoding to obtain the final embedded input sequence E = [e1, e2, ..., e N2 ];

[0082] The above embedding sequence E is used as input and fed into a standard Transformer encoder, which contains multiple layers of Transformer encoding blocks.

[0083] Use multiple attention heads, each head calculates different attention weights, concatenates the outputs of multiple heads, and passes through a linear layer to get the final attention output; Through the multi-head attention mechanism, different attention weights are obtained, the outputs of multiple heads are concatenated, and the final attention output Y is obtained through a linear layer;

[0084] The attention output Y is fed into a feedforward neural network, and the model’s ability to model local features is enhanced through two fully connected layers and ReLU activation functions.

[0085] The Transformer encoding blocks are repeatedly stacked in multiple layers. Each encoding block includes a multi-head attention mechanism and a feedforward neural network, and uses layer normalization and residual connections.

[0086] From the last layer output of the Transformer encoder, take the hidden state of the last time step as the global semantic feature vector g, and from the middle layer output of the Transformer encoder, take the hidden state of each time step as the local spatial feature sequence S;

[0087] The global semantic feature vector g extracted from the input data is input into a fully connected layer, and through the linear transformation of this layer, a vector a with a dimension of N3 is obtained:

[0088] Multiply the local feature sequence S by the attention weight a element by element to obtain the weighted local feature:

[0089] Sum the weighted local features to get the final fusion feature f ST .

[0090] Furthermore, the fused features are used as input, and the global and local context information of the image is used to constrain the generation process of key points through back diffusion, including:

[0091] Using the reverse diffusion method, the uncertainty distribution from the key point The determined key point distribution H0 is restored in the process, and the uncertainty distribution is gradually reduced, and finally an accurate key point prediction result is obtained; as shown below:

[0092]

[0093] in, is the sampled noise sample, f ST is the fusion feature extracted by the encoder, which represents the key point context feature. is the current step embedding, the kth diffusion step is represented by a sine function, K represents the diffusion step length, g θ is a trained diffusion model;

[0094] In the last step of back diffusion, a 1x1 convolutional layer is used to output the heat map of the key points;

[0095] Post-process the generated key point heat map.

[0096] Compared with the prior art, the present invention has the following beneficial effects: by using the diffusion model to perform denoising on the cervical vertebrae image, the image quality is improved, and a better basis is provided for subsequent segmentation and key point detection; the cervical vertebrae segmentation and key point detection model combined with deep learning technology can accurately extract the cervical vertebrae area and key point positions, and provide accurate data support for subsequent bone age staging; bone age staging based on the results of cervical vertebrae segmentation and key point detection can objectively and quantitatively evaluate the patient's growth and development, provide auxiliary decision-making information for doctors, and help formulate personalized orthodontic treatment plans;

[0097] The biomechanical model combines the stress distribution map generated by finite element analysis (FEA) with the preprocessed image to introduce physiological characteristics into the segmentation process. This adaptive segmentation model can adjust the segmentation strategy in real time according to the individual differences of patients, thereby improving the accuracy and robustness of segmentation.

[0098] Parameters such as blood flow velocity and temperature distribution are extracted and incorporated into the machine learning model to provide new indicators for cervical vertebrae bone age staging. This combination can effectively improve the limitations of traditional methods for growth and development assessment and provide more personalized and accurate treatment plans.

[0099] In summary, this method has the advantages of good denoising effect, high segmentation accuracy, accurate key point detection, and the ability to assist doctors in judging the degree of growth and development and formulating orthodontic treatment plans. It has important application value in the field of medical image analysis and orthodontic treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 It is a schematic diagram of the process of the present invention;

[0101] Figure 2 It is a framework diagram of the cervical vertebra segmentation and key point detection method based on the diffusion model of the present invention;

[0102] Figure 3 The framework diagram of the heat map extraction module of the present invention is

[0103] Figure 4 The framework diagram of the basic residual block, upsampling and downsampling modules of the present invention is shown in FIG.

[0104] Figure 5 It is a framework diagram of the local feature and global feature extraction and fusion module in the present invention. DETAILED DESCRIPTION

[0105] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.

[0106] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0107] Embodiment 1:

[0108] See also Figures 1 to 5, the present invention provides a technical solution:

[0109] Attached Figure 2 In the text, NBF-Filter (Neighborhood-Based Filter) refers to a neighborhood-based filtering method; this filtering method is mainly used in image processing, image segmentation, computer vision and other fields, and is an existing technology, so it will not be described in detail;

[0110] A method for cervical vertebrae segmentation and key point detection based on a diffusion model, the specific steps comprising:

[0111] Step S1: a data preprocessing step, obtaining a lateral head radiograph image dataset, preprocessing the image dataset, and dividing the preprocessed image dataset into a training set and a test set;

[0112] Step S2: Based on data preprocessing, biomechanical parameters of the cervical spine area are collected through biomechanical simulation software or experimental measurement equipment. The biomechanical parameters include stress distribution, rotation angle and displacement data between vertebrae. The collected biomechanical data are integrated with the image data set preprocessed in step S1 to form an enhanced feature set to facilitate subsequent model segmentation;

[0113] Step S3: cervical vertebra segmentation step, obtaining an enhanced feature set, and inputting the enhanced feature set into a cervical vertebra segmentation model based on Transformer and CNN. The model uses the global context information of Transformer and the local feature extraction capability of CNN, and introduces a diffusion model to gradually refine the segmentation results through iterative optimization to segment the cervical vertebra;

[0114] At the same time, the stress distribution and motion trajectory of the vertebrae are combined with the segmentation model, and the model parameters are updated in real time using an adaptive weight adjustment mechanism to improve the accuracy and adaptability of the segmentation.

[0115] Step S4: multi-scale feature extraction step, performing multi-scale feature extraction on the training set and the test set of the preprocessed image, and upsampling and fusing the features of different scales to generate a key point heat map;

[0116] Step S5: After the multi-scale feature extraction, the microcirculation blood flow characteristics of the cervical spine area are obtained by combining infrared thermal imaging technology, and key parameters in the microcirculation blood flow characteristics are extracted. The key parameters include blood flow velocity and temperature distribution. The key parameters are integrated into the key point thermal map to obtain an enhanced thermal map;

[0117] Using a multivariate regression model based on temperature gradient and blood flow velocity, combined with image segmentation results, a multi-scale analysis of cervical vertebrae bone age staging was performed to enhance the model's predictive ability.

[0118] Step S6: uncertainty distribution fitting step, performing uncertainty distribution based on the enhanced heat map, fitting the Gaussian distribution parameters of each key point from the enhanced heat map using the maximum likelihood estimation method, and constructing a Gaussian mixture model using the Gaussian distribution parameters of multiple key points as the initial uncertainty distribution;

[0119] Step S7: uncertainty evolution simulation step, constructing a learnable Gaussian mixture model to simulate the diffusion process, capturing the uncertainty evolution characteristics of the key point positions, taking the parameters of the initial Gaussian mixture model as input, and iteratively simulating the dynamic evolution of uncertainty through the learnable diffusion model;

[0120] Step S8: feature fusion step, inputting the training set and test set included in the preprocessed image into the background encoder to extract the global semantic feature vector and local spatial features, and performing feature fusion using the attention mechanism to obtain the final fused features;

[0121] Step S9: a key point generation constraint step, taking the fused features as input, using the global and local context information of the image, and constraining the key point generation process by back diffusion;

[0122] Step S10: Input the current lateral head radiograph image data set to be tested into the diffusion model constructed above, and use the model to perform cervical vertebrae segmentation and key point detection on the image data set to generate segmentation results and key point positions; at the same time, perform corresponding biomechanical analysis and health assessment based on the detection results to provide accurate auxiliary information for clinical diagnosis.

[0123] Further explanation: obtaining a lateral head radiograph image dataset and preprocessing the image dataset specifically includes:

[0124] The preprocessing includes unifying the format of the lateral head radiograph images and converting the images into RGB format; normalizing the lateral head radiograph images using the Z-score standardization method; converting the image pixel values ​​into a standard normal distribution centered at 0 and with a standard deviation of 1 to reduce the impact of differences between image domains. The preprocessing formula is as follows;

[0125]

[0126] Among them, p5 represents the gray value of the pixel in the original image, m3 represents the average gray value of the entire sample image, s is the standard deviation, which reflects the degree of dispersion of the pixel value relative to the average value, and n p5 Represents the new pixel value after normalization, sum(all p5 ) is the sum of all pixel values, num p5 Indicates the total number of pixels in the image sample;

[0127] The normalized image is subjected to noise reduction processing by using bilateral filtering; the image preprocessing also includes one or more preprocessing operations such as resizing, cropping, and histogram equalization, which are not limited here;

[0128] Annotate the cervical vertebrae of each lateral head radiograph image sample in the image data set, including the annotation of the key points and segmented areas of the cervical vertebrae, and obtain multiple data including annotated cervical vertebrae area images;

[0129] It should be noted that the cervical spine image data can be obtained from the lateral skull radiograph database, specifically from the database of relevant hospitals or public data on the Internet. The ratio of the training set and the test set are 70%, 30% or 80%, 20% respectively, and there is no restriction here.

[0130] Further, the collected biomechanical data is integrated with the image data set preprocessed in step S1 to form an enhanced feature set, which specifically includes:

[0131] For the collection of biomechanical parameters:

[0132] Finite Element Analysis (FEA) software was used to simulate the biomechanical behavior of the cervical spine area;

[0133] Establish a 3D cervical spine model: Use the geometric information extracted from the image data set to construct a 3D cervical spine model; the model needs to include relevant structures such as vertebrae, intervertebral discs, ligaments, and set the mechanical properties of each material, which are Young's modulus and Poisson's ratio;

[0134] Define boundary conditions and loading conditions: Set boundary conditions and loading conditions. The boundary conditions are head fixation and lower constraints, and the loading conditions are applying torque and pressure to simulate the movement and load of the cervical spine under actual physiological conditions. Specific conditions that need to be considered include: applying a certain head gravity and simulating common head movements, including flexion, extension and rotation, to obtain complete stress and displacement data;

[0135] Finite element analysis software was used to perform finite element analysis on the three-dimensional cervical spine model to calculate the stress distribution, rotation angle, and displacement of each vertebra under different load and motion conditions;

[0136] Stress distribution calculation formula:

[0137]

[0138] Where σ′ is stress, Fz is force, and As is the area of ​​action; stress distribution can help identify vulnerable areas in the cervical spine;

[0139] The displacement and rotation angle are calculated using the following formulas:

[0140]

[0141] Among them, Δx and Δy are displacement differences, θ is the rotation angle, and x new is the horizontal coordinate of the current position, x original is the horizontal coordinate of the original position, y new is the ordinate of the current position, y original is the ordinate of the original position; the displacement data reflects the deformation of the cervical spine under different loads, and arctan is the inverse tangent function;

[0142] Among them, when the patient stands naturally, the normal physiological curvature of the cervical spine corresponds to the original position. In this position, the relative position and posture of each vertebra of the cervical spine conform to the normal anatomical structure and are not affected by external forces;

[0143] For the integration of biomechanical parameter data and image data:

[0144] Projecting the three-dimensional biomechanical data obtained by finite element analysis onto the pre-processed two-dimensional lateral head radiograph image;

[0145] This process requires the use of an image registration-based method to align the output of the biomechanical model with the medical image data; the purpose of image registration is to find the actual position information of the cervical spine in three-dimensional space so that each pixel corresponds to a specific biomechanical parameter;

[0146] Image registration method selection: Select the image registration method based on mutual information (MI), which is defined as:

[0147] MI(X,Y)=H(X)+H(Y)-H(X,Y)

[0148] Among them, H(X) and H(Y) are the entropies of images X and Y respectively, and H(X,Y) is the joint entropy. This method can effectively handle the alignment problem between multimodal images, such as CT / MRI and lateral radiographs, and ensure the accurate integration of data.

[0149] For the generation of enhanced feature sets:

[0150] Each preprocessed image pixel is fused with the corresponding biomechanical parameters to generate an enhanced feature vector v containing multimodal information, which is defined as:

[0151] v=[I(x,y),σ′(x,y),θ(x,y),Δx(x,y),Δy(x,y)]

[0152] Where I(x,y) is the image pixel value at (x,y), σ′(x,y) is the stress value at (x,y), θ(x,y) is the rotation angle at (x,y), Δx(x,y) and Δy(x,y) are the displacements at (x,y);

[0153] The enhanced feature set is standardized so that all features have the same scale. The standardization method uses z-score standardization to convert the features into a standard normal distribution to ensure that the contributions of different features are balanced during model training.

[0154] The generated enhanced feature set is divided into training set and test set in a ratio of 70% / 30% or 80% / 20% to ensure that the model can generalize effectively on unseen data.

[0155] Further explanation: the enhanced feature set is obtained and input into the cervical vertebra segmentation model based on Transformer and CNN. The model uses the global context information of Transformer and the local feature extraction capability of CNN, and introduces the diffusion model to gradually refine the segmentation results through iterative optimization to segment the cervical vertebra. Specifically, it includes:

[0156] The noisy annotated mask is input into the encoder of the diffusion model. At the same time, the segmentation features of the preprocessed lateral head radiograph are extracted through the encoder-decoder and integrated into the encoding features of the diffusion model. The multi-scale features extracted by the diffusion model and the encoder-decoder are fused through the Transformer module based on the spectral space. The Transformer module uses the cross-attention mechanism to learn the interaction between noise features and semantic features.

[0157] The divided training set images are input into the encoder-decoder to extract the decoded segmentation features, and the extracted decoded segmentation features are integrated into the encoding features of the diffusion model; the details are as follows:

[0158]

[0159] M=(F(c 0 )Wq)(F(d 0 )W k ) T

[0160] in, The last feature decoded by the U-network, is the first feature of the diffusion model, * represents the sliding window kernel operation, represents the conventional element-level operation, and k GaussA learnable Gaussian kernel is used for smooth activation, and the maximum value between the smoothed map and the original feature map is selected through the Max operation to retain the most relevant information, thereby obtaining a smooth anchor feature map. represents a 1×1 convolution kernel, which reduces the number of anchor feature channels to 1 and activates it using the sigmoid activation function, adding it to In each channel of F(c 0 ) and F(d 0 ) represents the feature c extracted by the encoder 0 and the characteristic d in the diffusion model 0 The features obtained by transferring to Fourier space (FFT), W q and W k are the learnable query weights and key weights in Fourier space;

[0161] The semantic features of the image are extracted through a U-shaped network encoder and injected into the embedding vector of the diffusion model using a fusion module to assist in fusing the bottom-level global semantic features in the segmentation model with the inverse recovery features of the diffusion model. The fusion module is a Transformer based on the spectrum space.

[0162] Specifically, firstly, the feature c extracted by the encoder 0 and the characteristic d in the diffusion model 0 Transfer to Fourier space (FFT), and get F(c 0 ) and F(d 0 ); Then, the attention mechanism is used to calculate the affinity weight matrix in Fourier space as follows:

[0163] M=(F(c 0 )Wq)(F(d 0 )W k ) T

[0164] Where W q and W k are the learnable query weights and key weights in Fourier space;

[0165] Use a bandpass filter (NBP-Filter) to align the frequency representation, filter out unnecessary frequency components, and enhance relevant features;

[0166] The filtered affinity map was transformed back to Euclidean space using inverse fast Fourier transform;

[0167] The fused final features are further refined through a multi-layer perceptron (MLP) to enhance the decoding ability of the model and generate the final cervical vertebra segmentation results; the final feature c′0 is further obtained;

[0168] The final fused features are input into the decoder, and the final cervical vertebra segmentation results are obtained through implicit iterative optimization.

[0169] It should be noted that the image encoder and decoder adopt existing technologies and will not be described in detail.

[0170] To further illustrate, multi-scale feature extraction is performed on the preprocessed image, and features of different scales are upsampled and fused to generate a key point heat map, specifically including:

[0171] By inputting the preprocessed image into multiple parallel high-resolution networks and processing the image at 1 / 4, 1 / 8, and 1 / 16 resolutions respectively, feature extraction is performed at different scales, and feature information is exchanged between multiple sub-networks through cross-scale connections;

[0172] The high-resolution sub-network will pass its feature map to the low-resolution sub-network, and the low-resolution sub-network will pass its feature map to the high-resolution sub-network, so that sub-networks of different resolutions can share each other's feature information, and obtain multiple feature maps with different resolutions;

[0173] High-resolution subnetwork setup: The preprocessed image is downsampled 4 times through two basic convolution modules, each of which includes a convolution layer, a batch normalization layer, and a ReLU activation function;

[0174] The specific implementation method of multi-scale feature extraction is as follows Figure 4 As shown:

[0175] High-resolution sub-network settings: downsample by 4 times through two 3×3 basic convolution modules, each of which includes a convolution layer, a batch normalization layer, and a ReLU activation function; then adjust the number of channels through a linear layer; then, based on the linear layer, downsample by 4 times and 8 times through two parallel 3×3 convolution layers respectively to obtain two branches of different scales; finally, downsample the branch downsampled by 8 times again through a 3×3 convolution layer to 16 times, thereby obtaining multiple feature maps with different resolutions;

[0176] Upsample and concatenate the feature maps of different resolutions to obtain the final fused multi-scale features;

[0177] The specific implementation method of multi-scale feature fusion is as follows: Figure 4As shown in , for each scale branch in the above, four basic residual blocks are passed through, and then the output of the 8-fold downsampling branch is upsampled twice, and the output of the 16-fold downsampling branch is upsampled 4 times, and then added with the output of the 4-fold downsampling. Finally, a ReLu activation function is used to obtain the fusion output of the 4-fold downsampling branch, and the fusion of other branches is similar; where each residual block is as follows Figure 4 As shown in (a), it includes two 3×3 convolution modules, a residual connection block with a convolution kernel size of 1×1, and a ReLu activation function; the downsampling module is as follows Figure 4 As shown in (c), it includes a 3×3 convolution layer, a batch normalization layer and a ReLu activation function. A 3×3 convolution layer needs to be added for each downsampling by 2 times. The upsampling module is as follows: Figure 4 As shown in (b), a convolution layer with a convolution kernel size of 1×1 is used, followed by batch normalization, and finally the nearest neighbor interpolation is used to directly amplify the required multiple to obtain the up-sampled result;

[0178] The fused multi-scale feature map is input into a 1×1 convolutional layer to map the fused feature map to a single-channel heat map representation; then the activation function is used to normalize the output value to between 0 and 1 to obtain a key point heat map with key point response intensity information.

[0179] Further explanation: extract key parameters from microcirculatory blood flow characteristics and integrate them into key point heat maps; use a multivariate regression model based on temperature gradient and blood flow velocity, combined with image segmentation results, to perform multi-scale analysis of cervical vertebrae bone age staging; specifically include:

[0180] For infrared thermal imaging data acquisition and preprocessing:

[0181] Use high-resolution infrared thermal imaging equipment to image the patient's cervical spine area and obtain thermal imaging images of microcirculatory blood flow characteristics; the pixel value of the thermal imaging image represents the temperature information at each point;

[0182] The thermal imaging image is registered with the cervical vertebrae segmentation result image; specifically, an image registration method based on mutual information (MI) is used to achieve accurate alignment of multimodal images;

[0183] For extracting microcirculatory blood flow characteristics:

[0184] Extract the temperature distribution data of the cervical spine area from the thermal imaging image; for each pixel point, define the temperature matrix T(x′,y′) to represent the temperature value at the position (x′,y′);

[0185] The blood flow velocity is estimated based on the gradient of temperature distribution; first calculate the temperature gradient

[0186]

[0187] A blood flow velocity estimation model based on temperature gradient is established, and the blood flow velocity v(x′, y′) is set to be proportional to the temperature gradient, expressed as:

[0188]

[0189] Wherein, k′ is a proportionality coefficient, which indicates the relationship between temperature gradient and blood flow velocity; k′ is determined by the expert group through experimental data;

[0190] For integrating key parameters into key point heatmaps:

[0191] A multivariate regression model based on temperature gradient and blood flow velocity was constructed. The formula of the multivariate regression model was:

[0192] R(x′,y′)=α·v(x′,y′)+β·T(x′,y′)+γ

[0193] Among them, R(x′, y′) is the response value of the thermal map, v(x′, y′) is the blood flow velocity, T(x′, y′) is the temperature value, α, β, γ are regression coefficients; α, β, γ are determined by model training;

[0194] The output value of the regression model is fused with the previously generated key point heat map to obtain the final enhanced heat map; the enhanced heat map is expressed as:

[0195] H enhanced (x′,y′)=Sigmoid(H original (x′,y′)+R(x′,y′))

[0196] Among them, H enhanced (x′, y′) is the value of the enhanced heat map at position (x′, y′), H original (x′, y′) is the value of the initial key point heat map at the position (x′, y′), R(x′, y′) is the output value of the multivariate regression model; the Sigmoid function is used to normalize the output value of the enhanced heat map to between 0 and 1;

[0197] Multi-scale analysis of cervical vertebrae bone age staging:

[0198] The enhanced thermal map is combined with the segmented cervical vertebra image features to form a joint feature vector; the vector includes the geometric shape features, microcirculation blood flow features and temperature distribution features of the cervical vertebra; it is combined by the following formula:

[0199] F combined =[F bone ,H enhanced,T(x′,y′),v(x′,y′)]

[0200] Among them, F bone is the geometric shape feature of the cervical vertebra segmentation image, F combined is the joint eigenvector;

[0201] Using the above joint feature vector, a regression model based on multi-scale analysis is used to predict cervical vertebrae bone age stages; the formula of the multi-scale regression model is:

[0202] Age predicted =w1·F bone +w2·H enhanced +w3·T(x′,y′)+w4·v(x′,y′)+b″

[0203] Among them, Age predicted is the predicted bone age, w1, w2, w3, w4 are the learned positive weight coefficients, and b″ is the bias term; w1, w2, w3, w4 and b″ are determined by the expert group through experimental data and will not be described in detail;

[0204] The regression model is optimized using the training data set, and the model parameters are updated by minimizing the mean square error; the formula for the mean square error is:

[0205]

[0206] Where N is the number of samples, Age true,i is the true bone age of the i-th sample;

[0207] An independent validation set was used to evaluate the accuracy and robustness of the model; the model performance was further verified by the cross-validation method to ensure its generalization ability to different datasets.

[0208] Further explanation: uncertainty distribution is performed based on the enhanced heat map. The Gaussian distribution parameters of each key point are fitted from the enhanced heat map using the maximum likelihood estimation method. The Gaussian distribution parameters of multiple key points are combined into a Gaussian mixture model as the initial uncertainty distribution, which specifically includes:

[0209] Extract the candidate position of the key point from the enhanced heat map, that is, the position with the largest probability value in the enhanced heat map. The coordinates of the candidate position are the preliminary estimated position of the key point;

[0210] Collect the 2D coordinates of all key points into a data set, where n′ is the total number of key points;

[0211] For each initially selected key point position, a 2D Gaussian distribution is fitted to describe the uncertainty of the position. The expression of the 2D Gaussian distribution function is:

[0212] gaussian2d(x′,y′,m′ ux′ ,m′ uy′ ,σ x′x′ ,σ x′y′ ,σ y′y′ )

[0213] Among them, x′, y′ represent the position of the key point in the enhanced heat map, and the two parameters m′ux′, m′uy′ represent the mean of the 2D Gaussian distribution, that is, the center position of the distribution, corresponding to the predicted coordinates of the key point; σ x′x′ ,σ x′y′ ,σ y′y′ The three parameters define the covariance matrix of the 2D Gaussian distribution, representing the variance along the x-axis, the y-axis, and the covariance between the x-axis and the y-axis;

[0214] Optimize the above Gaussian distribution parameters, and obtain the Gaussian distribution parameters of each key point position by minimizing the negative log-likelihood loss of the heat map at a series of positions;

[0215] The obtained Gaussian distribution parameters of each key point are combined into a Gaussian mixture model (GMM), which is referred to as GMM. The Gaussian distribution parameters include mean m′, covariance σ and weight, and the initial estimate H of the uncertainty distribution of the key point position extracted from the enhanced heat map is obtained. K .

[0216] To further illustrate, a learnable Gaussian mixture model is constructed to simulate the diffusion process and capture the uncertainty evolution characteristics of the key point positions. The initial Gaussian mixture model is used as input and the learnable diffusion model is used to iteratively simulate the dynamic evolution of uncertainty. Specifically, it includes:

[0217] Taking the initial Gaussian mixture model as input, the learnable diffusion model is used to iteratively simulate the dynamic evolution of uncertainty.

[0218] Using the initial uncertainty distribution H K Parameters, run the expectation maximization algorithm of Gaussian mixture noise, and then fit the covariance matrix of each key point position;

[0219] The number of Gaussian components of the Gaussian mixture model is defined as M′, which is used to represent the number of mixed components of the uncertainty distribution of key points. Secondly, the EM algorithm is used to optimize the Gaussian mixture model parameter φ GMM ; including the mean μ of each Gaussian component m″ , covariance Σ m″ and the mixing coefficient π m″ ; These parameters fit the keypoint uncertainty distribution of the target by maximizing the log-likelihood function as follows:

[0220]

[0221] in, Represents the uncertainty distribution H K N of the samples GMM Key points, μ m″ and Σ m″ is the mean and covariance matrix of the m″th Gaussian component, π m″ ∈[0,1] means from the m″th mixed component Any sample drawn from The probability of represents the probability density function of the m″th Gaussian component; that is:

[0222]

[0223] Where, d|Σ m″ | is the dimension of the key point, T represents the transpose of the matrix;

[0224] The EM algorithm consists of the following two steps:

[0225] Step E: Calculate each sample as follows The posterior probability of belonging to the m″th Gaussian component

[0226]

[0227] Among them, π m″ represents the mixing coefficient of the m″th Gaussian component, represents the probability density function of the m″th Gaussian component, the denominator Representation sample The sum of the probabilities belonging to all Gaussian components, also called marginal probability;

[0228] M step: Use the posterior probability calculated in the E step to update the GMM parameter φ GMM :

[0229]

[0230] Among them, the above three formulas represent the mean μ of the m″th Gaussian component respectively. m″ , covariance Σ m″ and the mixing coefficient π m″ All samples According to their posterior probability of belonging to the m″th component Obtained by weighted average, T represents the transpose of the matrix;

[0231] By iterating the E step and the M step until the GMM parameters converge, a fitted GMM model is obtained, which contains the covariance matrix ∑ of each key point position m″ ;

[0232] Get the manually annotated key point distribution H0 and set the position of each key point are all determined cervical key point coordinate values, using the fitted GMM model φ GMM , calculate the position of each key point The corresponding Gaussian distribution parameters;

[0233] The forward diffusion process is performed on H0. After K steps, the generated noise distribution is Equal to the fitted GMM distribution φ GMM , thus obtaining the final key point uncertainty distribution The specific formula is:

[0234]

[0235] in, is the distribution generated from the fit The samples generated in , h0 represents the original sample of the heat map, is the mean μ of the selected Gaussian component m″ ,∈ G is the noise variable, 1 m″ is a binary indicator indicating whether the m″th Gaussian component is selected. If it is selected, it is 1, otherwise it is 0. k is the coefficient for the decay over time in the diffusion model.

[0236] To further illustrate, the preprocessed image is input into the background encoder to extract the global semantic feature vector and local spatial features, and the attention mechanism is used to fuse the features to obtain the final fused features, which specifically include:

[0237] The specific implementation steps are as follows: Figure 5 As shown;

[0238] The encoder uses a Transformer-based network structure to convert the preprocessed image into 2D patches through the embedding layer and flatten each patch to form an input sequence X = [X1, X2, ..., X N2 ], N2 is the total number of patches, each sequence is first mapped through a linear layer to convert it into a feature vector of fixed dimension d2, and the sine and cosine functions are used to add position encoding to obtain the final embedded input sequence E = [e1, e2, ..., e N2 ];

[0239] The above embedding sequence E is used as input and fed into a standard Transformer encoder, which contains multiple layers of Transformer encoding blocks.

[0240] Specifically, the input sequence E passes through three independent linear layers to obtain Query (Q2), Key (K2), Value (V2), and calculate the weight of the attention map:

[0241]

[0242] Where W q ,W k″ ,W v is the learnable weight matrix, d2 k is the dimension of Key, used to scale the attention weight;

[0243] Use multiple attention heads, each head calculates different attention weights, concatenates the outputs of multiple heads, and passes through a linear layer to get the final attention output; Through the multi-head attention mechanism, different attention weights are obtained, the outputs of multiple heads are concatenated, and the final attention output Y is obtained through a linear layer;

[0244] The attention output Y is fed into a feedforward neural network, and the model's ability to model local features is enhanced through two fully connected layers and ReLU activation functions; the details are as follows:

[0245] FFN (X) =Max(0,YW1+b1)W2+b2

[0246] Among them, W1, W2, b1, and b2 are learnable parameters, which are determined by an expert group based on experimental data;

[0247] The Transformer encoding blocks are repeatedly stacked in multiple layers. Each encoding block includes a multi-head attention mechanism and a feedforward neural network, and uses layer normalization and residual connections.

[0248] From the last layer output of the Transformer encoder, take the hidden state of the last time step as the global semantic feature vector g, and from the intermediate layer output of the Transformer encoder, take the hidden state of each time step as the local spatial feature sequence S = [s1, s2, ..., s N ];

[0249] The global semantic feature vector g extracted from the input data is input into a fully connected layer, and through the linear transformation of this layer, a vector a with a dimension of N3 is obtained:

[0250] a=softmax(gWa +b a )

[0251] Multiply the local feature sequence S by the attention weight a element by element to obtain the weighted local feature:

[0252] weighted S =[a1s1,a2s2,…,a N3 s N3 ]

[0253] Sum the weighted local features to get the final fusion feature f ST :

[0254] f ST =a1s1+a2s2+…+a N3 s N3

[0255] Among them, W a , b a It is a learnable parameter determined by an expert group based on experimental data.

[0256] To further illustrate, the fused features are used as input, and the global and local context information of the image is used to constrain the generation process of key points through back diffusion, which specifically includes:

[0257] Using the reverse diffusion method, the uncertainty distribution from the key point Recover the determined key point distribution H0, gradually reduce the uncertainty distribution, and finally get an accurate key point prediction result As shown below:

[0258]

[0259] in, is the sampled noise sample, represents the next step of iterative denoising, f ST The key point context features extracted by the encoder, is the current step embedding, the kth diffusion step is represented by a sine function, K represents the diffusion step length, g θ is a trained diffusion model;

[0260] In the last step of back diffusion, a 1x1 convolutional layer is used to output the heat map of the key points;

[0261] Post-process the generated key point heat map; the specific implementation steps are as follows:

[0262] First, the heat map is scanned pixel by pixel to find the maximum value in a fixed window around each pixel. If the response value of the pixel is the maximum value in the window, it is marked as a candidate key point. Then, non-maximum suppression is performed on the obtained candidate key points. After non-maximum suppression, the remaining key points are the final detection results, and the coordinates of these key points are output.

[0263] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0264] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. Those skilled in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software methods depends on the specific application and design constraints of the technical solution.

[0265] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0266] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for cervical vertebrae segmentation and key point detection based on a diffusion model, characterized in that: The specific steps include: Obtain a lateral head radiograph image dataset, preprocess the image dataset, and divide the preprocessed image dataset into a training set and a test set; specifically including: The preprocessing includes converting the lateral head radiograph image into a unified RGB format; then the lateral head radiograph image is normalized; Annotate the cervical vertebrae of each lateral head radiograph image sample in the image data set, including the annotation of the key points and segmented areas of the cervical vertebrae, and obtain multiple data including annotated cervical vertebrae area images; Collect biomechanical parameters of the cervical spine area, including stress distribution, rotation angle and displacement data between vertebrae, and integrate the collected biomechanical data with the preprocessed image data set to form an enhanced feature set; specifically, it includes: For the collection of biomechanical parameters: Finite element analysis software was used to simulate the biomechanical behavior of the cervical spine region; Using the geometric information extracted from the image data set, a three-dimensional cervical spine model is constructed; and the mechanical properties of each material are set, and the mechanical properties are Young's modulus and Poisson's ratio; Set boundary conditions and loading conditions to simulate the motion and load of the cervical spine under actual physiological conditions; Finite element analysis software was used to perform finite element analysis on the three-dimensional cervical spine model to calculate the stress distribution, rotation angle, and displacement of each vertebra under different load and motion conditions; Stress distribution calculation formula: Among them, σ′ is stress, Fz is force, and As is the action area; The displacement and rotation angle are calculated using the following formulas: Where Δx and Δy are displacement differences, θ is the rotation angle, xnew is the horizontal coordinate of the current position, xoriginal is the horizontal coordinate of the original position, ynew is the vertical coordinate of the current position, yoriginal is the vertical coordinate of the original position; arctan is the inverse tangent function; Among them, when the patient stands naturally, the normal physiological curvature of the cervical spine corresponds to the original position; For the integration of biomechanical parameter data and image data: Projecting the three-dimensional biomechanical data obtained by finite element analysis onto the pre-processed two-dimensional lateral head radiograph image; For the generation of enhanced feature sets: Each preprocessed image pixel is fused with the corresponding biomechanical parameters to generate an enhanced feature vector v containing multimodal information, which is defined as: v=[I(x,y),σ′(x,y),θ(x,y),Δx(x,y),Δy(x,y)] Where I(x,y) is the image pixel value at (x,y), σ′(x,y) is the stress value, θ(x,y) is the rotation angle at (x,y), Δx(x,y) and Δy(x,y) are the displacement differences at (x,y); The enhanced feature set is normalized so that all features have the same scale; Obtain an enhanced feature set, input the enhanced feature set into the cervical vertebra segmentation model based on Transformer and CNN, introduce a diffusion model, gradually refine the segmentation results through iterative optimization, and then segment the cervical vertebra; Perform multi-scale feature extraction on the training set and test set of preprocessed images, upsample and fuse features of different scales to generate key point heat maps; Infrared thermal imaging technology is used to obtain the microcirculatory blood flow characteristics of the cervical spine area, and key parameters in the microcirculatory blood flow characteristics are extracted. The key parameters include blood flow velocity and temperature distribution. The key parameters are integrated into the key point thermal map to obtain an enhanced thermal map. A multivariate regression model based on temperature gradient and blood flow velocity was used in combination with image segmentation results to perform multi-scale analysis of cervical vertebrae bone age staging. Specifically, it includes: For infrared thermal imaging data acquisition and preprocessing: Use high-resolution infrared thermal imaging equipment to image the patient's cervical spine area and obtain thermal imaging images of microcirculatory blood flow characteristics; Register the thermal imaging image with the cervical vertebrae segmentation result image; For extracting microcirculatory blood flow characteristics: Extract the temperature distribution data of the cervical spine area from the thermal imaging image; for each pixel point, define the temperature matrix T(x′,y′) to represent the temperature value at the position (x′,y′); The blood flow velocity is estimated based on the gradient of temperature distribution; first calculate the temperature gradient A blood flow velocity estimation model based on temperature gradient is established, and the blood flow velocity v(x′, y′) is set to be proportional to the temperature gradient, expressed as: Among them, k′ is the proportionality coefficient, which represents the relationship between temperature gradient and blood flow velocity; For integrating key parameters into key point heatmaps: A multivariate regression model based on temperature gradient and blood flow velocity was constructed. The formula of the multivariate regression model was: R(x′,y′)=α·v(x′,y′)+β·T(x′,y′)+γ Among them, R(x′, y′) is the thermal map response value, v(x′, y′) is the blood flow velocity, T(x′, y′) is the temperature value, and α, β, and γ are regression coefficients; The output value of the regression model is fused with the previously generated key point heat map to obtain the final enhanced heat map; the enhanced heat map is expressed as: Henhanced(x′,y′)=Sigmoid(Horiginal(x′,y′)+R(x′,y′)) Among them, Henhanced(x′,y′) is the value of the enhanced heat map at the position (x′,y′), Horiginal(x′,y′) is the value of the initial key point heat map at the position (x′,y′), and R(x′,y′) is the output value of the multivariate regression model; the Sigmoid function is used to normalize the output value of the enhanced heat map to between 0 and 1; Multi-scale analysis of cervical vertebrae bone age staging: The enhanced thermal map is combined with the segmented cervical vertebra image features to form a joint feature vector; the vector includes the geometric shape features, microcirculation blood flow features and temperature distribution features of the cervical vertebra; it is combined by the following formula: Fcombined=[Fbone,Henhanced,T(x′,y′),v(x′,y′)] Among them, Fbone is the geometric shape feature of the cervical vertebra segmentation image, and Fcombined is the joint feature vector; The formed joint feature vector was used to predict cervical vertebrae bone age stage using a regression model based on multi-scale analysis; The formula of the multiscale regression model is: Agepredicted=w1·Fbone+w2·Henhanced+w3·T(x′,y′)+w4·v(x′,y′)+b″ Among them, Agepredicted is the predicted bone age, w1, w2, w3, w4 are the learned positive weight coefficients, b″ is the bias term, Uncertainty distribution is performed based on the enhanced heat map. The Gaussian distribution parameters of each key point are fitted from the enhanced heat map using the maximum likelihood estimation method. The Gaussian distribution parameters of multiple key points are used to construct a Gaussian mixture model as the initial uncertainty distribution. Construct a learnable Gaussian mixture model to simulate the diffusion process, capture the uncertainty evolution characteristics of key point positions, use the parameters of the Gaussian mixture model as input, and iteratively simulate the dynamic evolution of uncertainty through the learnable diffusion model; Specifically include: Taking the initial Gaussian mixture model as input, the learnable diffusion model is used to iteratively simulate the dynamic evolution of uncertainty. Using the initial uncertainty distribution H K Parameters, run the expectation maximization algorithm of Gaussian mixture noise, and then fit the covariance matrix of each key point position; The number of Gaussian components of the Gaussian mixture model is defined as M′, which is used to represent the number of mixed components of the uncertainty distribution of key points. Secondly, the EM algorithm is used to optimize the Gaussian mixture model parameter φ GMM ; Get the manually annotated key point distribution H0 and set the position of each key point are all determined cervical key point coordinate values, using the fitted GMM model φ GMM , calculate the position of each key point The corresponding Gaussian distribution parameters; Perform a forward diffusion process on H0, and after K steps, define the final key point uncertainty distribution Equal to the fitted GMM distribution φ GMM ; The training set and test set included in the preprocessed image are input into the background encoder to extract the global semantic feature vector and local spatial features, and the attention mechanism is used to fuse the features to obtain the final fused features; The fused features are used as input, and the global and local context information of the image is used to constrain the generation process of key points through back diffusion, thereby realizing key point detection; The current lateral head radiograph image data set to be tested is input into the diffusion model constructed above, and the model is used to perform cervical vertebra segmentation and key point detection on the image data set to generate segmentation results and key point positions.

2. The method for cervical vertebrae segmentation and key point detection based on diffusion model according to claim 1, characterized in that: Obtain an enhanced feature set and input it into a cervical vertebra segmentation model based on Transformer and CNN. This model uses the global context information of Transformer and the local feature extraction capability of CNN, and introduces a diffusion model to gradually refine the segmentation results through iterative optimization to segment the cervical vertebra. Specifically, it includes: The noisy annotated mask is input into the encoder of the diffusion model; at the same time, the segmentation features of the preprocessed lateral head radiograph image are extracted through the encoder-decoder and integrated into the encoding features of the diffusion model; The divided training set images are input into the encoder-decoder to extract the decoded segmentation features, and the extracted decoded segmentation features are integrated into the encoding features of the diffusion model; The semantic features of the image are extracted through a U-shaped network encoder and injected into the embedding vector of the diffusion model using a fusion module to assist in fusing the bottom-level global semantic features in the segmentation model with the inverse recovery features of the diffusion model. The fusion module is a Transformer based on the spectrum space. Align the frequency representations using a bandpass filter and filter the frequency components to obtain an affinity map; The filtered affinity map was transformed back to Euclidean space using inverse fast Fourier transform; The final fused features are further refined through a multi-layer perceptron to enhance the decoding capability of the model and generate the final cervical vertebrae segmentation results; The final fused features are input into the decoder, and the final cervical vertebra segmentation results are obtained through implicit iterative optimization.

3. The method for cervical vertebra segmentation and key point detection based on diffusion model according to claim 2, characterized in that: Multi-scale feature extraction is performed on the preprocessed image, and features of different scales are upsampled and fused to generate a key point heat map, including: By inputting the preprocessed image into multiple parallel high-resolution networks and processing the image at 1 / 4, 1 / 8, and 1 / 16 resolutions respectively, feature extraction is performed at different scales, and feature information is exchanged between multiple sub-networks through cross-scale connections; The high-resolution sub-network passes its feature map to the low-resolution sub-network, and the low-resolution sub-network passes its feature map to the high-resolution sub-network, so that sub-networks with different resolutions can share each other's feature information; thus, multiple feature maps with different resolutions are obtained; Upsample and concatenate the feature maps of different resolutions to obtain the final fused multi-scale features; The fused multi-scale feature map is input into a 1×1 convolutional layer to map the fused feature map to a single-channel heat map representation; then the activation function is used to normalize the output value to between 0 and 1 to obtain a key point heat map with key point response intensity information.

4. The method for cervical vertebra segmentation and key point detection based on diffusion model according to claim 3, characterized in that: Based on the enhanced heat map, uncertainty distribution is performed. The Gaussian distribution parameters of each key point are fitted from the enhanced heat map using the maximum likelihood estimation method. The Gaussian distribution parameters of multiple key points are combined into a Gaussian mixture model as the initial uncertainty distribution, which includes: Extract the candidate position of the key point from the enhanced heat map, that is, the position with the largest probability value in the enhanced heat map. The coordinates of the candidate position are the preliminary estimated positions of the key point; Collect the 2D coordinates of all key points into a dataset; The obtained Gaussian distribution parameters of each key point are combined into a Gaussian mixture model, hereinafter referred to as GMM, to obtain the initial estimate H of the uncertainty distribution of the key point position extracted from the enhanced heat map K .

5. The method for cervical vertebra segmentation and key point detection based on diffusion model according to claim 4, characterized in that: The preprocessed image is input into the background encoder to extract the global semantic feature vector and local spatial features, and the attention mechanism is used to fuse the features to obtain the final fused features, which include: The background encoder uses a Transformer-based network structure to convert the preprocessed image into 2D patches through the embedding layer and flatten each patch to form an input sequence X = [X1, X2, ..., X N2 ], N2 is the total number of patches, each sequence is first mapped through a linear layer to convert it into a feature vector of fixed dimension d2, and the sine and cosine functions are used to add position encoding to obtain the final embedded input sequence E = [e1, e2, ..., e N2 ]; The above embedding sequence E is used as input and fed into a standard Transformer encoder, which contains multiple layers of Transformer encoding blocks. Use multiple attention heads, each head calculates different attention weights, concatenates the outputs of multiple heads, and passes through a linear layer to get the final attention output; Through the multi-head attention mechanism, different attention weights are obtained, the outputs of multiple heads are concatenated, and the final attention output Y is obtained through a linear layer; The attention output Y is fed into a feedforward neural network, and the model’s ability to model local features is enhanced through two fully connected layers and ReLU activation functions. The Transformer encoding blocks are repeatedly stacked in multiple layers. Each encoding block includes a multi-head attention mechanism and a feedforward neural network, and uses layer normalization and residual connections. From the last layer output of the Transformer encoder, take the hidden state of the last time step as the global semantic feature vector g, and from the middle layer output of the Transformer encoder, take the hidden state of each time step as the local spatial feature sequence S; The global semantic feature vector g extracted from the input data is input into a fully connected layer, and through the linear transformation of this layer, a vector a with a dimension of N3 is obtained: Multiply the local feature sequence S by the attention weight a element by element to obtain the weighted local feature: Sum the weighted local features to get the final fusion feature f ST .

6. The method for cervical vertebra segmentation and key point detection based on diffusion model according to claim 5, characterized in that: The fused features are used as input, and the global and local context information of the image is used to constrain the generation process of key points through back diffusion, including: Using the reverse diffusion method, the uncertainty distribution from the key point Recover the determined key point distribution H0, gradually reduce the uncertainty distribution, and finally get an accurate key point prediction result The iteration formula is as follows: in, is the sampled noise sample, represents the next step of iterative denoising, f ST is the fusion feature extracted by the encoder, which represents the key point context feature. is the current step embedding, represented by the sine function as the current k-th diffusion step, K represents the total number of diffusion steps, g θ is a trained diffusion model; In the last step of back diffusion, a 1x1 convolutional layer is used to output the heat map of the key points; Post-process the generated key point heat map.

Citation Information

Patent Citations

  • Cervical vertebra MRI image self-supervised segmentation method using diffusion model to generate data

    CN117036386A

  • Key point detection method for skull side position film

    CN117372425A