Intelligent assessment method for cervical lateral deviation based on multi-modal fusion and deep learning

The cervical spine lateral displacement assessment method based on multimodal fusion and deep learning solves the problems of single data source and weak feature generalization ability in traditional methods, and achieves higher accuracy and stability in cervical spine lateral displacement assessment.

CN121746355APending Publication Date: 2026-03-27INNER MONGOLIA NEUSOFT INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional methods for assessing cervical spine lateral displacement rely on a single data source, ignore multi-dimensional information, have weak feature generalization ability, and cannot capture the true mechanical state under dynamic conditions, resulting in large errors in the assessment results.

Method used

A multimodal fusion and deep learning approach is adopted to acquire multimodal data for preprocessing. Heatmaps are generated through feature weighted fusion. Key points are tracked by combining optical flow and Kalman filtering. A cervical spine-trunk coupled mechanical model is constructed, the gravity line reference axis is optimized, the pelvic tilt posture deviation is corrected, and the optimal offset is solved.

Benefits of technology

It improves the accuracy and reliability of cervical spine lateral deviation assessment, reduces the error of assessment results, and enhances the detection accuracy and stability under complex body postures and dynamic conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746355A_ABST
    Figure CN121746355A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent assessment method for cervical lateral deviation based on multi-modal fusion and deep learning, and solves the technical problems of single data source, weak feature generalization ability and large assessment result error caused by only analysis of a static single frame in existing cervical lateral deviation assessment. The method comprises the steps of obtaining multi-modal data, performing preprocessing, attitude normalization and feature weighted fusion to obtain fused multi-scale features, performing mapping to generate a heat map, and performing peak retrieval on the heat map to output key points; tracking time sequence tracks of the key points, correcting abnormal key points, and collecting continuous multi-frame key points for key point detection to obtain target key points; and constructing a cervical vertebra trunk coupling mechanical model, dynamically optimizing a gravity line reference axis, correcting pelvic inclination attitude deviation, and optimally solving the optimal offset through a Lagrange multiplication method and a least square method. The method can be widely applied to the technical field of medical image analysis and computer-aided diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of medical image analysis and computer-aided diagnosis technology, and more specifically, it relates to an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning. Background Technology

[0002] Cervical lateral deviation is a key biological parameter for assessing spinal health. The magnitude, direction, and whether it is accompanied by vertebral sequence disorder are core characteristics for measuring the biomechanical balance of the cervical spine and the overall stability of the spine. It is also an important basis for early identification of cervicogenic diseases.

[0003] However, traditional methods for assessing cervical spine lateral displacement rely on a single data source, using RGB images or point cloud data from the back, neglecting multi-dimensional information such as surface texture and thermal distribution, and failing to locate key anatomical points. Furthermore, traditional geometric feature detection methods rely on manually designed features, resulting in weak generalization ability and a high likelihood of assessment failure when dealing with examinees with significant differences in body posture. Moreover, traditional methods can only analyze single-frame static data such as RGB images or point cloud data from the back, ignoring the micro-dynamic features of human posture and multi-frame temporal correlation information, failing to capture the true biomechanical state and morphological changes of the cervical spine under dynamic conditions, leading to large errors in the assessment results. Summary of the Invention

[0004] The purpose of this application is to provide an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning, so as to solve the technical problems in the prior art that the cervical spine lateral displacement assessment has a single data source, weak feature generalization ability, and can only analyze static single frames, resulting in large errors in the assessment results.

[0005] To achieve the above objectives, this application provides an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning, comprising the following steps: Acquire multimodal data, perform preprocessing to obtain preprocessed multimodal data, perform pose normalization to obtain feature maps of each modality, perform feature weighted fusion to obtain fused multiscale features, perform mapping to generate heatmaps, and perform peak retrieval on the heatmaps based on anatomical constraints to output key points; The temporal trajectory of key points is traced by optical flow and Kalman filtering. Temporal consistency constraints are constructed to correct abnormal key points. Key points are collected in multiple consecutive frames for key point detection to obtain the target key points. A cervical spine-trunk coupled mechanical model is constructed based on the target key points. The gravity line reference axis is dynamically optimized, the pelvic tilt posture deviation is corrected, and the optimal offset is solved by Lagrange multiplication and least squares method.

[0006] Preferably, the process of performing feature weighted fusion includes: constructing a feature pyramid, capturing the detailed texture features of each modality feature map in the shallow layer, capturing the semantic information of each modality feature map in the deep layer, and using the U-Net structure to perform feature weighted fusion of the detailed texture features and semantic information to obtain the fused multi-scale features.

[0007] Preferably, before constructing the feature pyramid, it is necessary to obtain the weight coefficients of each modality. The process of obtaining the weight coefficients of each modality includes: The channel attention module dynamically adjusts the weights of each modality feature map through the channel attention module and the spatial attention module. The channel attention module performs global average pooling and max pooling on each modality feature map to extract global information and generates the weight coefficients of each modality through two fully connected network layers. The spatial attention module generates spatial weight maps through convolutional layers to enhance the feature response of key regions.

[0008] Preferably, the formula for solving the optimal offset is: ; In the formula, This is the lateral offset distance vector of the cervical spine. The three-dimensional coordinates of the key point at C7 of the cervical spine. The three-dimensional coordinates of the posterior center reference point of the pelvis. The direction vector of the gravity line. , The regularization coefficient is . For biomechanically constrained energy terms, This is a timing consistency constraint.

[0009] Preferably, the process of obtaining target key points through key point detection includes: acquiring key points in multiple consecutive frames, performing frame-by-frame detection to obtain multi-frame results, performing weighted fusion to obtain multi-frame detection results, taking the standard deviation of the multi-frame detection results as an indicator, and judging whether the quality of the current multiple consecutive frames of key points is qualified. If so, the target key point is obtained; otherwise, multiple consecutive frames of key points are reacquired.

[0010] Preferably, the heatmap generation process includes: mapping the fused multi-scale features to high resolution using an upsampling network, and generating a heatmap of the corresponding key points using a Gaussian kernel algorithm with the real location of the key points as the center.

[0011] Preferably, the upsampling network performs peak retrieval on the heatmap based on anatomical constraints to obtain the initial 2D coordinates of key points, and performs Taylor expansion to output the 2D coordinates and confidence scores of the key points; Anatomical constraints include intervertebral spacing constraints, symmetry constraints, and connectivity constraints, which are used to constrain key points and generate key points that conform to human physiological laws.

[0012] Preferably, the process of correcting pelvic tilt posture deviation includes: dynamically adjusting the reference axis direction based on the pelvic tilt angle, optimizing the objective function with mechanical balance, solving the position of the gravity line in the equilibrium state, correcting the tilt angle of the pelvis in the coronal plane, and matching the reference coordinate system with the physiological posture height of the human body when standing naturally. The formula for the objective function is: ; In the formula, It is the force of gravity. For muscle tension, For ligament restraint, This is the normalization coefficient.

[0013] Preferably, the formula for correcting the pelvic tilt angle in the coronal plane is: ; In the formula, The corrected measurement results, For the measurement results, The degree of inclination of the pelvis in the coronal plane.

[0014] Preferably, the multimodal data includes: depth images, RGB images, and thermal imaging data; Preprocessing includes: using bilateral filtering to remove noise from depth images while preserving edge features; performing color normalization on RGB images to unify colors under different lighting conditions; and calibrating the thermal imaging data to improve the accuracy of temperature data.

[0015] The beneficial effects of this application are as follows: This application provides an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning. First, multimodal data is acquired and preprocessed sequentially, and pose normalization is performed to obtain feature maps for each modality. Then, feature maps of each modality are fused by weighted feature aggregation, which enriches the feature dimensions and improves the feature generalization ability, breaking the information limitations of single-modal data. Second, optical flow and Kalman filtering are combined to track the temporal trajectory of key points, and temporal consistency constraints are constructed to constrain key points, correct abnormal key points, and simultaneously collect key points from multiple consecutive frames to detect and locate target key points, ensuring the accuracy of target key point selection. Finally, a cervical spine-torso coupling mechanical model and posture compensation mechanism are constructed to solve for the optimal displacement, making the acquired cervical spine lateral displacement more in line with physiological laws, and improving the accuracy and reliability of cervical spine lateral displacement assessment results. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the overall process of an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning, provided as an embodiment of this application. Detailed Implementation

[0018] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.

[0019] Please see Figure 1 This application provides an embodiment of an intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning, comprising: S1: Acquire multimodal data and perform preprocessing to obtain preprocessed multimodal data.

[0020] First, an extrinsic parameter calibration method based on a calibration board was used to unify the spatial coordinates of the infrared depth camera (1280×720 resolution, 30fps), the RGB camera (1920×1080 resolution), and the thermal imaging camera (640×480 resolution). Then, the back of the subject was photographed using the three calibrated cameras to obtain multimodal data.

[0021] Specifically, in one optional embodiment, a checkerboard calibration board is placed within the acquisition scene. Three cameras are controlled to capture images of the calibration board from multiple angles, detecting the checkerboard corner points in the images captured by each camera. The rotation matrix R and translation vector t of each camera are solved by minimizing the reprojection error, establishing a spatial mapping relationship for the three-modal data. The registration accuracy is verified by backprojection, controlling the error within 0.3 pixels to ensure that the data of each modality are aligned in spatial position. Then, using the calibrated infrared depth camera, RGB camera, and thermal imaging camera, the back of the subject is captured within the acquisition scene to obtain multimodal data, including depth images, RGB images, and thermal imaging data.

[0022] Furthermore, the multimodal data (depth image, RGB image, and thermal imaging data) are preprocessed to obtain preprocessed multimodal data. This application employs bilateral filtering to remove noise from the depth image while preserving edge features. Color normalization is performed on the RGB image to unify the colors of images under different lighting conditions. Temperature calibration is performed on the thermal imaging data to improve the accuracy of the temperature data.

[0023] S2: The preprocessed multimodal data is normalized to obtain feature maps of each modality. Feature weighting and fusion are performed to obtain fused multi-scale features. Mapping is then performed to generate a heatmap. Based on anatomical constraints, peak retrieval is performed on the heatmap to output key points.

[0024] The encoder employs a three-branch ResNet-50 structure, receiving preprocessed multimodal data and sharing some weights to reduce parameters and computational complexity. A Spatial Transformation Network (STN) module is inserted at the front end of the encoder. Based on the preprocessed multimodal data, the STN module automatically learns affine transformation parameters (rotation, translation, scaling), adaptively adjusting parameters through end-to-end training. This eliminates interference from ±15° body tilt and ±50mm positional displacement of the subject, automatically corrects deviations in human posture, eliminates the influence of standing posture, and normalizes human postures from different standing positions to a standard state, obtaining feature maps for each modality.

[0025] Furthermore, due to the prominent bones in the C7 cervical vertebrae and pelvic region, the surface temperature is relatively low (approximately 31-32°C, while the surrounding soft tissue area is approximately 33-34°C). This temperature difference can be used to accurately locate the skeletal region. The spatial derivative of temperature is calculated using temperature gradient detection (Sobel operator) to generate a temperature gradient distribution map, thereby enhancing the boundary features between key areas and other areas and improving the signal-to-noise ratio of key anatomical regions.

[0026] Then, in the fusion layer, the weights of each modality feature map are dynamically adjusted through the channel attention module and the spatial attention module to improve the feature response of key regions.

[0027] Specifically, channel attention extracts global information by performing global average pooling and max pooling on the feature maps of each modality, and generates weight coefficients for each modality through two fully connected layers (dimensionality reduction ratio of 16). Spatial attention generates spatial weight maps through convolutional layers (7×7 kernels) to enhance the feature responses of key regions (such as the cervical spine and pelvis).

[0028] Next, the decoder uses an upsampling network with a U-Net structure to generate a key point heatmap.

[0029] Specifically, a five-layer feature pyramid is constructed (resolutions ranging from 256×256 to 16×16), with each layer extracting features from different receptive fields. Shallow layers capture detailed texture features from RGB images, while deeper layers capture semantic information from depth images and thermal imaging data. Utilizing the skip connection characteristics of the U-Net structure, the detailed texture features from the shallow-layer RGB images are fused with the semantic information from the deep-layer infrared depth data and thermal imaging data through feature weighting, resulting in fused multi-scale features. This achieves multi-scale, cross-modal feature fusion, preserving both the details of the body surface texture and the integrity of the anatomical semantic information. The number of layers in the feature pyramid is not limited and can be set according to actual needs.

[0030] Then, the fused multi-scale features are mapped to a 256×256 resolution using an upsampling network, centered on the true location of the keypoints, and the Gaussian kernel algorithm (standard deviation) is applied. (pixel) to generate a heat map (probability distribution map) of 15 channels corresponding to 15 key points.

[0031] Furthermore, the output layer performs peak retrieval on the heatmaps of the 15 channels corresponding to the 15 keypoints generated by the decoder based on anatomical constraints, obtaining the peak position (the point with the highest probability) of each channel's heatmap and the initial 2D coordinates of the corresponding keypoint. Taylor expansion is performed on the initial 2D coordinates, and by fitting the pixel values ​​near the heatmap peak, sub-pixel-level precise positions are calculated, improving the 2D coordinate positioning accuracy to 0.2 pixels (approximately 0.6 mm), while simultaneously outputting the confidence score of the keypoint. The output layer embeds three anatomical constraint loss functions to constrain the keypoints, including intervertebral distance constraints, symmetry constraints, and connectivity constraints, ensuring that the keypoints conform to human physiological patterns.

[0032] Specifically, the intervertebral spacing is constrained to ensure that it conforms to physiological and anatomical principles. The formula is as follows: ; In the formula, The loss function is the vertebral spacing constraint. Intervertebral distance, For reference vertebral interbody spacing, This is the normalization coefficient.

[0033] The lateral coordinate difference between the left and right symmetrical key points is limited to <3mm by symmetry constraints, as shown in the formula: ; In the formula, For symmetry-constrained loss functions, Here are the horizontal coordinates of the key points that are symmetrical on the left. Here are the horizontal coordinates of the key points that are symmetrical on the right. This is the normalization coefficient.

[0034] By using connectivity constraints to keep the distance between adjacent key points within a reasonable range and preventing key points from jumping around, the formula is as follows: ; In the formula, The connectivity constraint loss function, Let i be the coordinates of the i-th key point. Let j be the coordinates of the j-th key point adjacent to the i-th key point. This is the normalization coefficient.

[0035] S3: Track the temporal trajectory of key points using optical flow and Kalman filtering, construct temporal consistency constraints to correct abnormal key points, collect continuous multimodal data for key point detection, and obtain the target key points.

[0036] Farneback dense optical flow is used to calculate the pixel displacement field between consecutive frames, tracking the motion trajectory of keypoints in the time dimension. A Kalman filter is constructed, incorporating the position and velocity of the keypoints into the state vector, and the process noise covariance matrix is ​​set. Observation noise covariance matrix The system predicts the location of key points in the next frame using a Kalman filter and fuses it with the real-time detection results. By balancing the weights of the predicted and observed values, the system ensures that the motion trajectory of the key points closely resembles the actual movement patterns of the human body, thus avoiding the problem of jumps in single-frame detection.

[0037] The periodic chest cavity movements caused by human respiration (frequency 0.2-0.3 Hz, amplitude 5-10 mm) were analyzed. The dominant frequency component (micro-dynamic feature) of the key point motion trajectory was extracted by Fourier transform to identify the respiratory cycle. Based on the micro-dynamic feature, key point measurements were performed in the stable phase of the respiratory cycle (end of respiration) to avoid dynamic interference and reduce the measurement fluctuation from ±1.5 mm to ±0.4 mm.

[0038] Based on the trajectory output by the continuous frame key point trajectory tracking algorithm, a temporal consistency constraint rule is constructed to limit the positional amplitude of key points in adjacent frames to not exceed the reasonable range of physiological motion. Then, through the prediction update mechanism of Kalman filtering, key points that deviate from the motion trajectory are corrected to ensure the smoothness and continuity of key point detection results between consecutive frames.

[0039] Keypoints are collected in multiple consecutive frames spanning 1-2 seconds and 3-5 frames, and keypoint detection is performed separately to obtain multi-frame results. A weighted average method is used to fuse the multi-frame detection results (the weight is proportional to the peak height of the heatmap corresponding to each keypoint in each frame; the higher the peak, the more reliable the detection result of that frame, and the greater the weight). At the same time, the standard deviation of the multi-frame results is calculated as an uncertainty indicator. If the standard deviation is >1.0mm, the quality of the currently collected consecutive frames of keypoints is considered insufficient, and re-collection is prompted to reduce the risk of false detection in a single frame and improve the stability of keypoint localization; if the standard deviation is <1.0mm, the target keypoints with stable localization are obtained.

[0040] In an optional embodiment, the Dropout layer is kept active during the inference phase (with a dropout rate of p=0.3), and 30 independent forward propagations are performed on the same subject's samples to obtain 30 sets of keypoints. The mean of the 30 sets of results is calculated as the final keypoint, and the standard deviation of the final keypoint represents cognitive uncertainty (the larger the standard deviation, the higher the model's prediction discrepancy for that sample and the less confidence it has). If the standard deviation is >1.0 mm, it is judged as high uncertainty, indicating insufficient sample information.

[0041] Keypoint detection was performed using single-modal networks (infrared, RGB, and thermal imaging only), resulting in three sets of independent predictions. A consistency index was calculated among the three predictions to measure the matching degree of the multimodal results. The formula is: ; In the formula, For multimodal consistency, The coordinates of key points obtained from infrared modal network detection. These are the keypoint coordinates obtained from the RGB modal network detection. These are the coordinates of key points obtained from thermal imaging modal network detection.

[0042] when When the thickness is <2mm, it is considered to have high consistency, with a confidence level >90%. when If the value is greater than 5mm, it is considered a modal conflict, indicating that the detection results of different modal data have too large a deviation and manual confirmation is required.

[0043] Based on 30 Monte Carlo Dropout samplings, assuming the predicted distribution follows a normal distribution, the 95% confidence interval is calculated using the following formula: ; In the formula, The 95% confidence interval is... The mean, , where is the standard deviation. Statistical analysis on the 1000-case validation set shows that the actual coverage of the 95% confidence interval reached 97.2%, higher than the theoretical value of 95%, indicating the reliability of the confidence interval prediction results.

[0044] The test-retest decision rule is designed by integrating three core quality indicators, namely Monte Carlo Dropout standard deviation. Multimodal consistency and peak height of heat map The decision rule is: if >1.0mm or >4mm or If the value is less than 0.6, the system will automatically flag insufficient data quality and suggest a retest. Clinical trials have shown that this rule can accurately identify key target points, reducing the invalid measurement rate from 25% to 5%, significantly improving the efficiency and stability of the testing process.

[0045] S4: Based on the target key points, construct a cervical spine-trunk coupled mechanical model, dynamically optimize the gravity line reference axis, correct the pelvic tilt posture deviation, and optimize the solution of the optimal offset through Lagrange multiplication and least squares method.

[0046] Based on the target key points, a simplified rigid-spring system model is established. The cervical spine is modeled as a rigid chain of 7 vertebrae, and the intervertebral connections (intervertebral discs, ligaments) are modeled as torsional springs (stiffness coefficient k = 5 N·m / rad). The trunk is modeled as a flexible support, considering the mechanical balance relationship of gravity, muscle tension, and ligament constraints. Model parameters (such as head mass of approximately 5 kg and center of gravity located at the C1-C2 level) are based on biomechanical calibration to accurately reproduce the mechanical coupling characteristics of the cervical spine and trunk.

[0047] Traditional methods use a simple vertical line as a reference axis, but the line of gravity adjusts according to posture when the human body is standing. This application is based on the pelvic tilt angle. (Calculate the pelvic tilt angle using the line connecting the anterior superior iliac spine and the posterior superior iliac spine) The reference axis direction is dynamically adjusted. The objective function is optimized using mechanical equilibrium to determine the position of the gravity line in equilibrium, ensuring the reference coordinate system closely matches the physiological posture of a person standing naturally. The formula for the objective function is as follows: ; In the formula, It is the force of gravity. For muscle tension, For ligament restraint, This is the normalization coefficient.

[0048] Detecting the degree of pelvic tilt in the coronal plane (Normal range is ±5°), when | When |> 3°, the measurement result is corrected using the following formula: ; In the formula, The corrected measurement results, The measurement results are as follows.

[0049] This correction process can eliminate spurious offsets caused by pelvic tilt (up to 2 mm), improving the consistency of measurement results by 60% across different standing postures.

[0050] The calculation of cervical spine offset distance is transformed into an optimization problem, and the objective function is designed as follows: ; In the formula, For geometric measurement constraints, For biomechanical constraints, when A value of 0.3 represents the regularization coefficient. For geometric measurement constraint matrix, For cervical spine offset parameters, The target vector is a geometrically constrained vector. Let be the target vector of the biomechanical constraints. This is the normalization coefficient.

[0051] By solving the optimization problem using Lagrange multiplication, the optimal offset that satisfies multiple constraints is obtained. Compared with simple geometric projection, the stability is improved by 8 times (the standard deviation is reduced from 1.8 mm to 0.2 mm when the attitude changes by ±10°).

[0052] The biomechanically optimized formula for calculating cervical spine offset distance is as follows: ; In the formula, This is the lateral offset distance vector of the cervical spine (2D: x, y directions). The three-dimensional coordinates of the key point at C7 of the cervical spine. The three-dimensional coordinates of the posterior center reference point of the pelvis. The direction vector of the gravity line. , The regularization coefficients are set to 0.3 and 0.2 in this embodiment, respectively. There is no limitation on the setting of the regularization coefficients, and they can be selected according to the actual situation. For biomechanically constrained energy terms, solutions that violate force equilibrium are penalized. The time-series consistency constraint penalizes solutions that deviate significantly from the previous frame. Iterative optimization using the least squares method converges after 5-8 iterations, yielding the optimal offset.

[0053] Example 1: Comparative experiment between the single-modal method and this application.

[0054] The test set consisted of 300 participants with different body postures (standard / hunchback / scoliosis) and data quality (excellent / good / average), covering all age groups from teenagers to young adults and the elderly.

[0055] Specifically, the control group was set up with a single-modal detection method (infrared depth only, RGB only, thermal imaging only), and the experimental group was set up with a three-modal adaptive fusion method (infrared depth + RGB + thermal imaging). The same key point detection network architecture was used for both groups, and the measurement accuracy and detection success rate under complex body shapes were compared.

[0056] The average measurement error of the single-modal method is 1.5-2.0 mm, while the measurement error of the three-modal fusion method is only 0.6±0.3 mm, representing a 65-70% improvement in accuracy compared to the single-modal approach. In complex body shape samples such as kyphosis and scoliosis, the single-modal method achieves a detection success rate of only 60-75%, while the three-modal fusion method achieves a success rate of 97%, demonstrating a 30% improvement in robustness. Therefore, this application can combine the texture details of RGB, the spatial structure of infrared depth, and the anatomical region features of thermal imaging to improve the identification of key anatomical points in complex scenes, thereby enhancing detection accuracy.

[0057] Example 2: Comparative experiment between traditional geometric feature detection methods and the present application.

[0058] The control group consisted of traditional geometric feature methods based on curvature analysis or symmetry detection, while the experimental group consisted of the end-to-end deep learning method based on three-branch ResNet-50 and U-Net proposed in this application. The measurement accuracy, abnormal body shape generalization ability, and processing efficiency of the two groups were compared.

[0059] Experimental results show that the measurement error of traditional geometric feature detection methods is 1.8±1.2 mm, while the error of this application is 0.6±0.3 mm, representing a 3-fold improvement in accuracy. Traditional geometric feature detection methods have a failure rate as high as 38% on abnormal body shape samples, while the failure rate of the end-to-end deep learning method in this application is only 3%, demonstrating a 12-fold improvement in generalization ability. Traditional methods, due to their complex point cloud processing flow, take 8-12 seconds per sample, while the lightweight deep learning model inference only requires 80 ms, representing an efficiency improvement of over 100 times.

[0060] Therefore, the end-to-end deep learning method of this application can automatically extract deep correlation features of multimodal data, and has significant advantages in accuracy, generalization and processing efficiency compared with traditional geometric feature detection methods.

[0061] Example 3: Comparison experiment between the single-frame static measurement method and the present application.

[0062] The control group used a single-frame static measurement method, while the experimental group used the temporal multi-frame fusion method of 3-frame or 5-frame fusion as described in this application. 10 seconds of continuous data were collected from the subjects while they were standing naturally and breathing normally. The measurement fluctuations and temporal consistency of the three groups were compared.

[0063] Experimental results show that the standard deviation of a single-frame measurement is 1.2 mm, which decreases to 0.5 mm after 3-frame fusion and further to 0.3 mm after 5-frame fusion, representing a 4-fold improvement in stability. Furthermore, the 5-frame fusion method reduces the coefficient of variation of continuous measurements from 18% to 4%, significantly improving temporal smoothness and effectively eliminating the interference of physiological micro-dynamics such as respiration on the measurement results.

[0064] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0065] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A smart assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning, characterized in that, Includes the following steps: Multimodal data is acquired, preprocessed to obtain preprocessed multimodal data, pose normalization is performed to obtain feature maps of each modality, feature weighted fusion is performed to obtain fused multiscale features, mapping is performed to generate a heatmap, and peak retrieval is performed on the heatmap based on anatomical constraints to output key points; The temporal trajectory of the key points is traced by optical flow and Kalman filtering. Temporal consistency constraints are constructed to correct abnormal key points. Key points are collected in multiple consecutive frames for key point detection to obtain the target key points. Based on the target key points, a cervical spine-trunk coupled mechanical model is constructed, the gravity line reference axis is dynamically optimized, the pelvic tilt posture deviation is corrected, and the optimal offset is solved by Lagrange multiplication and least squares method.

2. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The process of performing the feature weighted fusion includes: constructing a feature pyramid, capturing the detailed texture features of each modality feature map in a shallow layer, capturing the semantic information of each modality feature map in a deep layer, and using a U-Net structure to perform feature weighted fusion of the detailed texture features and the semantic information to obtain the fused multi-scale features.

3. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 2, characterized in that, Before constructing the feature pyramid, it is necessary to obtain the weight coefficients of each modality. The process of obtaining the weight coefficients of each modality includes: The weights of each modality feature map are dynamically adjusted by a channel attention module and a spatial attention module. The channel attention module performs global average pooling and max pooling on each modality feature map to extract global information and generates the weight coefficients of each modality through two fully connected network layers. The spatial attention module generates a spatial weight map through convolutional layers to enhance the feature response of key regions.

4. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The formula for solving the optimal offset is: ; In the formula, This is the lateral offset distance vector of the cervical spine. The three-dimensional coordinates of the key point at C7 of the cervical spine. The three-dimensional coordinates of the posterior center reference point of the pelvis. The direction vector of the gravity line. , The regularization coefficient is . For biomechanically constrained energy terms, This is a timing consistency constraint.

5. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The process of obtaining the target key point by performing the key point detection includes: collecting key points in multiple consecutive frames, performing frame-by-frame detection to obtain multi-frame results, performing weighted fusion to obtain multi-frame detection results, taking the standard deviation of the multi-frame detection results as an indicator, and judging whether the quality of the key points in the current multiple consecutive frames is qualified. If so, the target key point is obtained; otherwise, the key points in multiple consecutive frames are collected again.

6. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The process of generating the heatmap includes: mapping the fused multi-scale features to a high resolution using an upsampling network, and generating a heatmap corresponding to the key point using a Gaussian kernel algorithm with the real location of the key point as the center.

7. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 6, characterized in that, The upsampling network performs peak retrieval on the heatmap based on the anatomical constraints to obtain the initial 2D coordinates of the key points, and performs Taylor expansion to output the 2D coordinates and confidence of the key points. The anatomical constraints include intervertebral spacing constraints, symmetry constraints, and connectivity constraints, which are used to constrain the key points and generate key points that conform to human physiological laws.

8. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The process of correcting the pelvic tilt posture deviation includes: dynamically adjusting the reference axis direction based on the pelvic tilt angle, optimizing the objective function with mechanical balance, solving the position of the gravity line under equilibrium state, correcting the tilt angle of the pelvis in the coronal plane, and matching the reference coordinate system with the physiological posture height of the human body when standing naturally. The formula for the objective function is: ; In the formula, It is the force of gravity. For muscle tension, For ligament restraint, This is the normalization coefficient.

9. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 8, characterized in that, The formula for correcting the pelvic tilt angle in the coronal plane is as follows: ; In the formula, The corrected measurement results, For the measurement results, The degree of inclination of the pelvis in the coronal plane.

10. The intelligent assessment method for cervical spine lateral displacement based on multimodal fusion and deep learning as described in claim 1, characterized in that, The multimodal data includes: depth images, RGB images, and thermal imaging data; The preprocessing includes: using bilateral filtering to remove noise from the depth image while preserving edge features; performing color normalization on the RGB image to unify the colors of the image under different lighting conditions; and calibrating the thermal imaging data to improve the accuracy of the temperature data.

Citation Information

Cited By

  • Automatic cervical spine MRI image segmentation method and system based on attention mechanism

    CN122199581A