Hand rehabilitation trend analysis method and system based on image processing and key points
By constructing a hand simulation model and a Riemannian manifold model, the problems of subjectivity and crude quantification in the assessment of hand motor dysfunction were solved, and the accurate assessment of hand rehabilitation trends and abnormality detection were achieved.
Patent Information
- Application Number
- CN202511431422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-30
AI Technical Summary
Existing technologies for assessing hand motor dysfunction are highly subjective, have crude quantification, are difficult to capture subtle dynamic changes, and lack a framework for unified representation and fusion of time-series data, thus failing to meet clinical measurement requirements.
By acquiring hand video fusion input data and medical image data, a hand simulation model is constructed, key point sets and temporal cellular feature data are obtained, a Riemannian manifold model of the hand is established, Riemannian quantization feature data is obtained, and a trend analysis report is obtained based on the trend analysis model.
It provides more objective data support for the assessment of hand rehabilitation trends, with precise anatomical basis, and can assess the current rehabilitation status, future status and abnormal status, assisting physicians in formulating treatment plans.
Smart Images

Figure CN121237335A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to a method and system for analyzing hand rehabilitation trends based on image processing and key points. Background Technology
[0002] In the field of clinical rehabilitation, hand motor dysfunction is a common complication for patients with stroke, upper limb trauma, neurodegenerative diseases, and other conditions. After onset, varying degrees of impairment in hand grasping and extension functions occur, and the recovery of hand function directly affects the patient's ability to perform daily living activities and quality of life. Currently, clinical assessments rely heavily on traditional methods such as manual muscle strength testing by doctors, goniometer measurements of joint range of motion, and patient subjective function scales. While these methods are widely used, they inherently have limitations such as strong subjectivity, coarse quantification, and difficulty in capturing subtle dynamic changes. With the development of image processing technology, the use of visual sensing technology has become a major research direction in rehabilitation medicine, aiming to drive data-driven decision-making through objective data and achieve more accurate rehabilitation program analysis. However, existing technical solutions still have several bottlenecks.
[0003] On the one hand, two-dimensional pose estimation methods based on ordinary videos are easily affected by occlusion and viewpoint, and the accuracy of the extracted joints is limited. They also lack physical scale information, making it difficult to meet the measurement requirements in clinical practice. Furthermore, most analysis methods only focus on external kinematic parameters and fail to deeply integrate medical images that reveal the internal physiological state. On the other hand, there is a lack of a unified framework for representing and integrating time-series data, making it difficult to quantify the dynamic evolution patterns during the rehabilitation process. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method and system for analyzing hand rehabilitation trends based on image processing and key points. The purpose and effectiveness of this method and system for analyzing hand rehabilitation trends based on image processing and key points are achieved through the following specific technical means: Hand rehabilitation trend analysis methods based on image processing and key points include: S1: Acquire hand video fusion input data and medical imaging data, and pair the hand video fusion input data and medical imaging data based on the patient's rehabilitation nodes; S2: Generate an optimized segmentation mask based on hand video fusion input data and medical image data, and construct a hand simulation model based on the optimized segmentation mask, hand video fusion input data and medical image data; S3: Obtain the key point set and the key point set of the hand surface based on the hand simulation model, and at the same time, establish a temporal 3D honeycomb mesh model according to the honeycomb connection rules, and obtain temporal honeycomb feature data; S4: Construct a Riemannian manifold model of the hand based on temporal cellular feature data, and obtain Riemannian quantization feature data through the Riemannian manifold model of the hand; S5: Obtain trend analysis reports based on trend analysis models.
[0005] Furthermore, hand video fusion input data and medical imaging data are acquired, and the hand video fusion input data and medical imaging data are paired based on the patient's rehabilitation nodes, including: The system acquires hand imaging schemes, obtains multi-view hand video data based on these schemes, and simultaneously acquires hand electromyography signal data and hand inertial signal data. It also acquires patient rehabilitation milestones and, based on these milestones, acquires hand CT image data and hand MRI image data at different rehabilitation stages. Based on signal processing algorithms, time offset correction is performed on multi-view hand video data, hand electromyography signal data and hand inertial signal data. At the same time, distortion correction and epipolar correction are performed on multi-view hand video data. The corrected multi-view hand video data, hand electromyography signal data and hand inertial signal data are encapsulated into hand video fusion input data. Data parsing is performed on hand CT and hand MRI images to obtain CT image arrays and MRI image arrays. The CT and MRI image arrays are then resampled. The resampled MRI image arrays are corrected based on the N4ITK algorithm. At the same time, the resampled and corrected CT and MRI image arrays are denoised. The CT and MRI image arrays are then encapsulated into medical image data. Based on the patient's rehabilitation nodes, the input data of hand videos from different rehabilitation stages are fused and paired with the corresponding medical imaging data.
[0006] Furthermore, an optimized segmentation mask is generated based on the fusion input data of hand video and medical image data, including: Extract multi-view hand video data from the hand video fusion input data, and obtain single-view hand video data based on the multi-view hand video data. Calculate the hand optical flow field data for every two consecutive video frames in the single-view hand video data based on the optical flow algorithm. The hand optical flow amplitude is obtained based on the hand optical flow field data. At the same time, an optical flow amplitude map is constructed based on the hand optical flow amplitude. The ROI region in the optical flow amplitude map is located. The hand motion energy curve is constructed based on the ROI region and the hand optical flow amplitude. The hand motion energy curve is filtered, and the extreme points of the curve are located. The extreme points of the curve include the curve peak and the curve valley. Based on the extreme points of the curve, multi-view consistency verification is performed to obtain multi-view keyframe images. The multi-view keyframe images are keyframe images that have consistent curve extreme points in all views. The SAM model is used to segment multi-view keyframe images, generating an initial segmentation mask for each multi-view keyframe image. The initial segmentation mask is then optimized for connectivity to obtain an optimized segmentation mask.
[0007] Furthermore, a hand simulation model is constructed simultaneously based on the optimized segmentation mask, hand video fusion input data, and medical image data, including: Acquire multi-view keyframe images and corresponding optimized segmentation masks at each time point, obtain hand disparity maps at each time point through stereo matching algorithms, convert hand disparity maps into hand depth maps, obtain hand region point clouds at each time point based on hand depth maps, reconstruct hand region point clouds, and obtain 3D hand surface mesh. Adaptive segmentation is performed on CT image array data and MRI image array data in medical imaging data to generate CT image segmentation data and MRI image segmentation data. The MRI image segmentation data and CT image segmentation data are registered to generate MRI image registration data. The CT image segmentation data is registered with a 3D hand surface mesh to generate video surface data. Using the CT coordinate system as the model coordinate system, a hand simulation model was constructed based on MRI image registration data, video surface data, and CT surface data.
[0008] Furthermore, adaptive segmentation is performed on the CT image array data and MRI image array data in the medical imaging data to generate CT image segmentation data and MRI image segmentation data. The CT image segmentation data and MRI image segmentation data are then registered to generate MRI image registration data. Finally, the CT image segmentation data is registered with a 3D hand surface mesh to generate video surface data, including: Joint segmentation is performed on CT image array data based on threshold segmentation algorithm and connected component analysis, and soft tissue segmentation is performed on MRI image array data based on nnUnet network. The segmented regions are then marked to generate CT image segmentation data and MRI image segmentation data. CT surface data is extracted from CT image segmentation data. Simultaneously, edge detection is performed based on MRI image segmentation data to obtain MRI pseudo-surface data. Rigid registration of MRI image segmentation data and CT image segmentation data is performed based on CT surface data and MRI pseudo-surface data to obtain MRI image registration data. The CT surface data and the 3D hand surface mesh were initially registered using the principal component analysis algorithm, and then the video surface data were obtained by non-rigid registration of the initially registered CT surface data and the 3D hand surface mesh using the non-rigid ICP algorithm.
[0009] Furthermore, key point sets and hand surface key point sets are obtained based on the hand simulation model. Simultaneously, a temporal 3D honeycomb mesh model is established according to honeycomb connection rules, and temporal honeycomb feature data is obtained, including: Based on CT surface data in the hand simulation model, the center of the interphalangeal joint and the center of the wrist joint are located, and the fingertip point, phalangeal point, and wrist point are located to obtain a set of key points. Based on the non-rigid registration relationship between CT surface data and 3D hand surface mesh in the hand simulation model, the key point set is mapped to the 3D hand surface mesh in the hand simulation model, and the key point set of the hand surface is obtained based on the MRI image segmentation data in the hand simulation model. Connect the key points on the hand surface of the same finger, connect the key points on the hand surface of adjacent fingers, connect the key points on the hand surface of the wrist area, generate a triangular honeycomb mesh between the connected key points on the hand surface, and generate a temporal 3D honeycomb mesh model. Geometric and physical feature data of the temporal 3D cellular mesh model are extracted separately to generate temporal cellular feature data.
[0010] Furthermore, a Riemannian manifold model of the hand is constructed based on temporal cellular feature data, and Riemannian quantization feature data is obtained through the hand Riemannian manifold model, including: The temporal cellular feature data is standardized, and then linear dimensionality reduction is performed on the temporal cellular feature data based on principal component analysis algorithm. Then, nonlinear dimensionality reduction is performed on the temporal cellular feature data based on nonlinear dimensionality reduction algorithm to obtain a low-dimensional cellular embedding matrix. Riemannian manifold modeling is performed based on low-dimensional cellular embedding matrix to construct a hand Riemannian manifold model. At the same time, multiple initial motion trajectories are obtained. The normalized path of every two initial motion trajectories is calculated based on the dynamic time warping algorithm. Based on the normalized path, multiple initial motion trajectories are resampled to obtain aligned motion trajectories. Feature extraction was performed based on the Riemannian manifold model of the hand and the aligned motion trajectory to obtain Riemannian quantization feature data.
[0011] Furthermore, trend analysis reports are generated based on trend analysis models, including: Based on step S4, CT image segmentation data and MRI image segmentation data are obtained. Skeletal features are extracted from the CT image segmentation data to generate a skeleton feature sequence. Soft tissue features are extracted from the MRI image segmentation data to generate a soft tissue feature sequence. Based on the shooting time, the skeletal feature sequence and soft tissue feature sequence are reconstructed. Missing interpolation is performed on the reconstructed skeletal feature sequence and soft tissue feature sequence, and the soft tissue feature sequence and skeletal feature sequence are sorted. The hand electromyography signal data and hand inertial signal data, soft tissue feature sequence, bone feature sequence and Riemann quantization feature data from the hand video fusion input data are imported into the trend analysis model to obtain a trend analysis report; The trend analysis report includes a current rehabilitation status assessment report, a future rehabilitation trend prediction report, and an anomaly report.
[0012] Further trend analysis models include: The trend analysis model includes an input unit, a fusion unit, and an output unit; The input unit includes a functional feature branch, a skeletal feature branch, and a soft tissue feature branch. The functional feature branch is used to input Riemann quantization feature data, hand electromyography signal data, and hand inertial signal data. The skeletal feature branch is used to input skeletal feature sequences, and the soft tissue feature branch is used to input soft tissue feature sequences. The fusion unit fuses and splices Riemann quantized feature data, input skeletal feature sequences, and soft tissue feature sequences based on a cross-modal attention mechanism. The output unit includes a trend prediction unit, a status assessment unit, and an anomaly detection module. The trend prediction unit is used to generate a future rehabilitation trend prediction report. The status assessment unit is used to classify and generate a current rehabilitation status assessment report. The anomaly detection module is built based on a generative adversarial network and is used to judge the data of the input unit and generate anomaly reports.
[0013] A hand rehabilitation trend analysis system based on image processing and key points includes: An initial module is used to acquire hand video fusion input data and medical image data, and to pair the hand video fusion input data and medical image data based on the patient rehabilitation node; A simulation module is used to generate an optimized segmentation mask and construct a hand simulation model; A cellular module, which is used to establish a temporal 3D cellular mesh model and acquire temporal cellular feature data; A manifold module is used to construct a Riemannian manifold model of the hand and obtain Riemannian quantization feature data; The reporting module generates a trend analysis report based on a trend analysis model.
[0014] Based on the above, this application embodiment first acquires hand video fusion input data and medical image data, and pairs the hand video fusion input data and medical image data according to the patient's rehabilitation nodes. Then, an optimized segmentation mask is generated based on the hand video fusion input data and medical image data. Simultaneously, a hand simulation model is constructed based on the optimized segmentation mask, the hand video fusion input data, and the medical image data. Subsequently, a key point set and a key point set on the hand surface are obtained based on the hand simulation model. Simultaneously, a temporal 3D honeycomb mesh model is established according to honeycomb connection rules, and temporal honeycomb feature data is obtained. Finally, a Riemannian manifold model of the hand is constructed based on the temporal honeycomb feature data. Riemannian quantization feature data is obtained through the hand Riemannian manifold model, and a trend analysis report is obtained based on a trend analysis model. This invention transforms complex information such as hand posture and movement trajectory into geometric and kinematic feature parameters, thereby providing more objective data for assessing hand rehabilitation trends. Furthermore, it fuses CT / MRI image data with video streams from hand video fusion input data representing external motor function to establish a hand simulation model, providing more accurate anatomical evidence. Additionally, it extracts temporal cellular feature data, Riemann quantization feature data, and hand video fusion input data from the hand simulation model based on a 3D cellular mesh model, and establishes a unified representation and fusion framework based on a trend analysis model. This allows for the assessment of current rehabilitation status, future status, and abnormal status, shifting from passive to active assessment and better assisting physicians in evaluating patients' conditions and determining subsequent treatment and rehabilitation plans. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the hand rehabilitation trend analysis method based on image processing and key points provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the hand rehabilitation trend analysis system based on image processing and key points provided in an embodiment of the present invention. Detailed Implementation
[0016] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate the technical solutions of the present invention, but should not be used to limit the scope of protection of the present invention.
[0017] Example:
[0018] As attached Figure 1 , Figure 2 As shown: This invention provides a method for analyzing hand rehabilitation trends based on image processing and key points, applicable to hand rehabilitation trend analysis, and includes the following steps: Step S1: Obtain hand video fusion input data and medical imaging data, and pair the hand video fusion input data and medical imaging data based on the patient's rehabilitation node.
[0019] In this embodiment, step S1 includes: Step S11: Obtain hand imaging scheme, acquire multi-view hand video data based on hand imaging scheme, simultaneously acquire hand electromyographic signal data and hand inertial signal data, acquire patient rehabilitation nodes, and acquire hand CT image data and hand MRI image data at different rehabilitation stages based on patient rehabilitation nodes.
[0020] Specifically, before photographing the patient's hand, a hand photography plan is formulated, clearly defining the initial position of the patient's hand, the speed and amplitude of the movement, and the shooting angle, to ensure the consistency of the data collected at different rehabilitation stages. Two or more cameras can be used to photograph the hand from different angles, such as from the front view, the side view at 45 degrees, to ensure overlapping fields of view. At the same time, the binocular camera is calibrated, and the sensor group needs to be arranged to collect hand electromyographic signal data and hand inertial signal data. Then, the patient's arm can be placed on a support to maintain the consistency of photography at different stages.
[0021] Furthermore, patients can perform repetitive movements from fully extending to fully clenching their fists according to instructions. Simultaneously, hand electromyography signal data, hand inertial signal data, and multi-view hand video data are collected. At the same time, hand CT and hand MRI images are acquired at different rehabilitation nodes, i.e., different rehabilitation periods. For example, rehabilitation nodes can be set at 1 week, 2 weeks, 1 month, or 2 months. At each node, hand CT and hand MRI images are acquired with the patient's arm placed on a support. CT images focus on bone structure, while MRI images focus on soft tissue.
[0022] Step S12: Based on signal processing algorithms, time offset correction is performed on multi-view hand video data, hand electromyography signal data and hand inertial signal data. At the same time, distortion correction and epipolar correction are performed on multi-view hand video data. The corrected multi-view hand video data, hand electromyography signal data and hand inertial signal data are encapsulated into hand video fusion input data.
[0023] Understandably, when acquiring multi-view hand video data, hand electromyography signal data, and hand inertial signal data, due to the difference in clocks of the acquisition devices, it is necessary to use signal processing algorithms to align the time axes of each data point and eliminate time offset.
[0024] Specifically, the timestamp of each frame is extracted from the metadata of the multi-view hand video data, and the timestamp of each data sample is extracted from the original data files of the hand electromyography signal data and the hand inertial signal data. A start time Ts and an end time Te are defined, for example, starting from the first received start pulse signal and ending from the last end pulse signal. All data streams are cut to the same time interval from start time to end time. At this time, for any time T in the time interval from start time to end time, the frame with the timestamp closest to T can be found in the video. The frames with the timestamps closest to T in the hand electromyography signal data and the hand inertial signal data can be time aligned.
[0025] Specifically, for distortion correction and epipolar correction, distortion correction is based on the intrinsic parameter matrix obtained from camera calibration, and eliminates the optical distortion of the lens through pixel coordinate mapping. Epipolar correction can utilize the camera's extrinsic parameters, such as baseline distance and rotation matrix, to project video images from different viewpoints onto the same epipolar plane. That is, two projection transformation matrices are calculated to rotate the video image planes of the two cameras onto the same plane and align their epipolar lines. Both distortion correction and epipolar correction can be performed using relevant functions in OpenCV.
[0026] Furthermore, after completing time offset correction, distortion correction, and epipolar correction, the data can be encapsulated according to the classification and storage format of timestamp and data type to generate hand video fusion input data, which can be directly called in subsequent steps.
[0027] Step S13: Perform data parsing on the hand CT image data and hand MRI image data to obtain CT image array data and MRI image array data. Resample the CT image array data and MRI image array data. Correct the resampled MRI image array data based on the N4ITK algorithm. At the same time, perform noise reduction operation on the resampled CT image array data and the corrected MRI image array data. Encapsulate the CT image array data and MRI image array data into medical image data.
[0028] Specifically, the original DICOM files of hand CT and hand MRI image data are read, and the metadata in the original DICOM files is parsed to extract key metadata, such as pixel spacing data, slice spacing data, slice thickness data, MRI sequence information, CT HU value, etc., thereby outputting CT image array data and MRI image array data.
[0029] Furthermore, when acquiring hand MRI and hand CT image data, anisotropy is often present, with higher intra-slice resolution and lower inter-slice resolution. In this case, it is necessary to calculate the voxel size of the CT and MRI image array data, set the target resolution to the lowest resolution component, and use interpolation algorithms to process the CT image array data and trilinear interpolation or B-spline interpolation to process the MRI image array data. This resamples the CT and MRI image array data onto an isotropic voxel grid, avoiding size errors between CT and MRI.
[0030] Furthermore, bias field correction is performed on the resampled isotropic MRI image array data. It is understandable that MRI images will have slowly changing intensity shadows, i.e. bias fields, due to the inhomogeneity of the radio frequency field. Therefore, the N4ITK algorithm is used to correct them.
[0031] Understandably, the N4ITK algorithm is a field correction algorithm for MRI images. This algorithm uses a combination of fourth-order Gaussian filtering and local enhancement normalization to gradually eliminate the field effect. The N4ITK algorithm can effectively estimate and remove the bias field, thereby obtaining an image with uniform intensity.
[0032] Furthermore, denoising is performed on the resampled CT image array data and the corrected MRI image array data. Before denoising the MRI image array data, it needs to be normalized to a uniform standard range, such as 0-1 or 0-255. Then, denoising can be performed according to the noise characteristics of the MRI image array data, such as nonlocal mean filtering and dictionary learning denoising. For the CT image array data, the original pixel values are converted into standard HU values, and denoising is performed using methods such as nonlocal mean filtering and anisotropic diffusion filtering. The denoised CT image array data and MRI image array data are then packaged into medical image data.
[0033] Step S14: Based on the patient's rehabilitation nodes, the hand video fusion input data at different rehabilitation stages is paired with the corresponding medical image data.
[0034] Step S2: Generate an optimized segmentation mask based on the hand video fusion input data and medical image data, and construct a hand simulation model based on the optimized segmentation mask, the hand video fusion input data and medical image data.
[0035] In this embodiment, step S2: Step S21: Extract multi-view hand video data from the hand video fusion input data, and obtain single-view hand video data based on the multi-view hand video data. Calculate the hand optical flow field data for every two consecutive video frames in the single-view hand video data based on the optical flow algorithm.
[0036] Specifically, time-aligned multi-view hand video data is read frame by frame, and each single-view hand video data is extracted from the multi-view hand video data. In each single-view hand video data, for two consecutive frames Ft and Ft+1, the optical flow field of Ft and Ft+1 is calculated using the TV-L1 optical flow algorithm. The optical flow field is a dense optical flow field.
[0037] Furthermore, for two consecutive frames Ft and Ft+1, Ft and Ft+1 are converted into grayscale images, and the optical flow calculation function in the TV-L1 optical flow algorithm is called. The optical flow calculation function outputs a two-channel array with the same size as the original image. One channel stores the displacement component of each pixel in the X direction, i.e., the horizontal direction, and the other channel stores the displacement component of each pixel in the Y direction. Finally, an optical flow field sequence is output, and each optical flow field describes the motion from Ft to Ft+1.
[0038] It is understandable that the hand is not a fixed rigid body and its motion is relatively complex. Dense optical flow can capture relatively subtle global motion information such as finger bending, which can provide a comprehensive data calculation basis for the motion energy calculation in the subsequent step S22.
[0039] Step S22: Obtain the optical flow amplitude of the hand based on the optical flow field data of the hand, and construct an optical flow amplitude map through the optical flow amplitude of the hand. Locate the ROI region in the optical flow amplitude map, and construct the hand motion energy curve based on the ROI region and the optical flow amplitude of the hand.
[0040] Specifically, for each pixel in the optical flow field of the hand, the optical flow amplitude, i.e., the magnitude, is calculated, which represents the movement speed of the pixel from Ft to Ft+1. The motion amplitude of all pixels is calculated, and an optical flow amplitude map of the same size as the original frame is obtained. The optical flow amplitude map can intuitively represent the degree of motion in each region. If the value is larger, it means that the motion is more intense, while if the value is smaller, the motion is more static.
[0041] Furthermore, the average value in the optical flow amplitude map is calculated and set as a threshold. Pixel regions with amplitude values higher than the threshold are statistically analyzed to locate the Region of Interest (ROI) and obtain the optical flow amplitude map after ROI processing. In addition, the ROI region can also be located by a moving target detection algorithm based on background subtraction of consecutive frames. For each optical flow amplitude map after ROI processing, the median of the amplitude of all pixels in the ROI region is calculated and set as a scalar value. Then, this is repeated for each pair of consecutive frames to obtain multiple consecutive scalar values. The multiple consecutive scalar values are connected in chronological order to obtain the hand motion energy curve. The above operation is performed on each single-view hand video data to obtain the hand motion energy curve for each view.
[0042] Step S23: Filter the hand motion energy curve and locate the curve extreme points. The curve extreme points include curve peaks and curve valleys. Perform multi-view consistency verification based on the curve extreme points and obtain multi-view keyframe images. The multi-view keyframe images are keyframe images that have consistent curve extreme points in all views.
[0043] Understandably, the acquired hand motion energy curve may contain high-frequency jitter or noise. Gaussian filtering is used to filter it, thereby eliminating jitter noise and maintaining the main trend.
[0044] Specifically, after the filtering operation, local maxima and local minima in the hand motion energy curve are searched. Local maxima correspond to curve peaks, which represent the moment when the hand movement speed is fastest, usually during the transition period of the action, such as the moment when the speed is fastest from extending the hand to clenching the fist. Local minima correspond to curve troughs, which represent the moment when the hand movement speed is slowest, usually during the holding period of the action, such as the holding posture of fully extending the hand or fully clenching the fist. When searching for curve peaks and troughs, a minimum threshold needs to be set to filter out fluctuations that are too small. Then, multi-view consistency verification is performed.
[0045] For example, suppose there is a valley point Tv1 in the first view. Search for a time window of Tv1±a on the hand motion energy curve in another view, where a represents the tolerance, for example, 3. If a valley point is also detected in the time window set in the other view, then the valley points are consistent. If no valley point is detected in the time window set in the other view, then Tv1 may be a spurious signal and should be deleted. Finally, obtain the multi-view keyframe image based on the extreme points of the curve after the multi-view consistency verification is completed.
[0046] Step S24: Segment the multi-view keyframe images based on the SAM model, generate an initial segmentation mask corresponding to each multi-view keyframe image, perform connectivity optimization on the initial segmentation mask, and obtain an optimized segmentation mask.
[0047] Specifically, the SAM model is a model for image segmentation tasks. Multi-view keyframe images and cue points are input into the SAM model. The image encoder in the SAM model generates a one-time image embedding, and the cue encoder generates a cue embedding. The image embedding and cue embedding are input into the mask decoder to generate multiple candidate segmentation masks. Each candidate segmentation mask is accompanied by a prediction confidence score, which is used to represent the degree of confidence that the SAM model has in each candidate segmentation mask. Then, the candidate segmentation mask with the highest prediction confidence score is selected, which is the initial segmentation mask.
[0048] In some possible implementations, cue points can be generated based on simple and lightweight object detection algorithms, such as the YOLO object detection algorithm.
[0049] Furthermore, a closing operation, i.e., a dilation and erosion operation in image processing, is performed on the initial segmentation mask to fill the tiny gaps and holes between fingers or at joints, thereby ensuring that the hand region is a complete connected region. The edges are then smoothed. In some cases, if there is interference from multiple similar objects in the image background, multiple discrete mask regions may be generated. In this case, connected component analysis is required to retain only the hand region. Finally, the optimized segmentation mask is output.
[0050] Step S25: Obtain multi-view keyframe images and corresponding optimized segmentation masks at each time point; obtain hand disparity maps at each time point through stereo matching algorithm; convert hand disparity maps into hand depth maps; obtain hand region point clouds at each time point based on hand depth maps; reconstruct hand region point clouds to obtain 3D hand surface mesh.
[0051] Specifically, the hand disparity map at each time point can be obtained through a semi-global matching algorithm (SGM algorithm) or a derived matching algorithm. Compared with the local block matching algorithm, the SGM algorithm can perform cost aggregation on multiple paths, thereby better processing weaker texture areas or occlusions and outputting a more complete disparity map.
[0052] It should be noted that the time point is represented as the frame time point. Each time the stereo matching algorithm obtains the hand disparity map, it is based on the multi-view keyframe images at the same time point and the corresponding optimized segmentation mask. If it is based on the multi-view keyframe images at different time points and the corresponding optimized segmentation mask, the stereo matching algorithm will obtain incorrect disparity and depth.
[0053] Furthermore, based on the triangulation principle of stereo vision, depth is inversely proportional to parallax. The hand depth map can be obtained based on the hand parallax map. Then, the hand depth map is projected inversely into 3D space to generate a 3D point for each pixel. An optimized segmentation mask is used to filter background points, thereby obtaining the point cloud of the hand region.
[0054] Furthermore, the above steps are repeated for keyframe images of different angles and time points within a rehabilitation node and the corresponding optimized segmentation mask to obtain multiple hand region point clouds. The multiple hand region point clouds are registered to the same coordinate system, and Poisson surface reconstruction is used to construct a 3D hand surface mesh model, i.e., a 3D hand surface mesh, which is represented as the surface model of the hand.
[0055] Step S26: Adaptively segment the CT image array data and MRI image array data in the medical image data to generate CT image segmentation data and MRI image segmentation data. Register the MRI image segmentation data and CT image segmentation data to generate MRI image registration data. Register the CT image segmentation data with the 3D hand surface mesh to generate video surface data.
[0056] In this embodiment, step S26 includes: Step S261: Joint segmentation is performed on the CT image array data based on the threshold segmentation algorithm and connected component analysis, and soft tissue segmentation is performed on the MRI image array data based on the nnUnet network. The segmented regions are then marked to generate CT image segmentation data and MRI image segmentation data.
[0057] Specifically, the threshold of the threshold segmentation algorithm can be defined according to the physical definition of HU value. For example, the HU value of cortical bone is usually higher than 400, while the HU value of soft tissue is generally in the range of -100 to 100. A threshold between 250 and 300 HU can be defined, and all voxels above this threshold are marked as bones. Then, connected component analysis is applied to mark the interconnected voxel clusters as independent sets, and labels are assigned to them to output CT image segmentation data.
[0058] Furthermore, for MRI, nnUnet is used to segment soft tissues in MRI images. The uuUnet network is used for forward propagation to predict the probability of each voxel belonging to an anatomical type. Based on statistical operations, the class with the highest probability is used as the label of the voxel, thereby outputting MRI image segmentation data.
[0059] Step S262: Extract CT surface data from CT image segmentation data, and simultaneously perform edge detection based on MRI image segmentation data to obtain MRI pseudo-surface data. Based on the CT surface data and MRI pseudo-surface data, perform rigid registration of MRI image segmentation data and CT image segmentation data to obtain MRI image registration data.
[0060] Specifically, voxels labeled as any bone in the CT image segmentation data are set to 1, while the remaining voxels are set to 0, thereby extracting CT surface data, i.e., a binary mask of a bone, from the CT image segmentation data. In MRI images, bones usually appear as low signals, and the boundaries of bones are often blurred, making it impossible to directly extract a clear surface. Therefore, edge detection algorithms are needed to process the MRI to find a substitute image that can approximately reflect the bone boundary information, i.e., MRI pseudo-surface data. In MRI pseudo-surface data, the boundaries between tissues are high values, and the interior of tissues is low value. Then, a registration algorithm based on mutual information can be used to find a transformation matrix based on the MRI pseudo-surface data and CT surface data. After applying this transformation matrix to the MRI image segmentation data, the implicit bone boundaries can be aligned with the bone surface of the CT.
[0061] Step S263: The CT surface data and the 3D hand surface mesh are initially registered using the principal component analysis algorithm, and the initially registered CT surface data and the 3D hand surface mesh are non-rigidly registered using the non-rigid ICP algorithm to obtain video surface data.
[0062] Specifically, the CT surface data and the 3D hand surface mesh are processed using principal component analysis (PCA) to roughly estimate the main orientation of the hand and perform initial rotation and translation. However, since the skin will slide and deform relative to the bones, a non-rigid ICP algorithm is required to perform non-rigid registration between the initially registered CT surface data and the 3D hand surface mesh. The non-rigid ICP algorithm allows for smooth surface deformation while finding the nearest point pair, and calculates a displacement vector for each vertex of the 3D hand surface mesh so that it eventually matches the surface of the CT surface data. This establishes a dynamic correspondence between the skin and the internal skeletal structure, thereby obtaining the registered 3D hand surface mesh, i.e., the video surface data.
[0063] Step S27: Using the CT coordinate system as the model coordinate system, construct a hand simulation model based on MRI image registration data, video surface data, and CT surface data.
[0064] Specifically, when constructing the hand simulation model, the CT coordinate system is used as the model coordinate system to establish a unified spatiotemporal coordinate system, ensuring spatial consistency of data across all modalities. When building the hand simulation model using digital twins, the hierarchical structure of multimodal data is first established and cross-modal associations are performed. For example, the multimodal data hierarchy can include anatomical and functional layers. When constructing the anatomical layer, CT surface data is imported as the underlying structure and assigned attributes such as vertex or facet bone names. Simultaneously, MRI image registration data is imported, and tissue type attributes are added to voxels. When constructing the functional layer, video surface data is imported, and each... The motion trajectory of the vertices constitutes dynamic changes, which can then be used for cross-modal association. By using the correspondence between the non-rigid registration of video surface data and CT surface data, the association between skin vertices and MRI voxels is determined, thereby realizing the mapping between external skin motion and internal bones and soft tissues. Then, the data is encapsulated in chronological order and three-level models with different levels of detail are generated for different application scenarios. The first-level model contains complete data of all vertices, faces, voxels and attributes, which can be used for high-precision analysis. The second-level model simplifies the skin and bone mesh and reduces the number of faces. The third-level model is represented only by feature parameters.
[0065] Step S3: Obtain the key point set and the key point set of the hand surface based on the hand simulation model. At the same time, establish a temporal 3D cellular mesh model according to the cellular connection rules and obtain temporal cellular feature data.
[0066] In this embodiment, step S3 includes: Step S31: Based on the CT surface data in the hand simulation model, locate the center of the interphalangeal joint and the center of the wrist joint, and simultaneously locate the fingertip point, phalangeal point, and wrist point to obtain a set of key points.
[0067] Specifically, the CT surface data is processed based on the traveling cube algorithm to separate the merged bones and generate triangular mesh surfaces to represent the bones. Then, the triangular mesh surfaces are identified to identify the two bones that form a joint, such as the proximal phalanx and the middle phalanx connecting at the proximal interphalangeal joint. For each bone, at the end of the joint formed with the adjacent bone, the point cloud of the joint surface is automatically extracted by setting a geometric clipping box. For the two joint surface point clouds extracted from the two bones, a sphere is fitted using the least squares method. The spheres have similar radii and their centers are close. The two spheres are fitted together to form an optimal sphere, and its center is taken as the center of the joint, thereby locating the center of the interphalangeal joint and the center of the wrist joint.
[0068] Furthermore, for the triangular mesh surface of each bone, principal component analysis is used to calculate its principal direction. The principal direction usually represents the longitudinal axis of the bone. The two extreme points of the triangular mesh surface of the bone are found along the principal direction, namely the distal end and the proximal end. The midpoint of these two points can be defined as the phalanx point. For the phalangeal bones, the distal vertex of its principal direction is the phalanx point. For the carpal bones, the wrist point can be defined by shape features on specific carpal bones, such as the scaphoid tubercle and pisiform bones, according to standard anatomical atlases. At the same time, a local anatomical coordinate system is established for the located key bone points and organized into a set of key points.
[0069] Step S32: Based on the non-rigid registration relationship between CT surface data and 3D hand surface mesh in the hand simulation model, the key point set is mapped to the 3D hand surface mesh in the hand simulation model, and the key point set of the hand surface is obtained based on the MRI image segmentation data in the hand simulation model.
[0070] Understandably, through non-rigid registration relationships, namely the dynamic correspondence between the skin and the internal skeletal structure established in step S263, the key point set is mapped to the 3D hand surface mesh in the hand simulation model. Then, the mapped key points are processed along their normal direction to the bone surface, and their relationship with soft tissue is processed through MRI image segmentation data, thereby ensuring that the key points are always on the skin surface, thus generating a key point set for the hand surface.
[0071] Step S33: Connect the key points of the hand surface of the same finger, connect the key points of the hand surface of adjacent fingers, connect the key points of the hand surface of the wrist area, generate a triangular honeycomb mesh between the connected key points of the hand surface, and generate a temporal 3D honeycomb mesh model.
[0072] Specifically, for each finger, including the thumb, index finger, middle finger, ring finger, and little finger, the key points on it are connected in physiological order from proximal to distal to simulate the skeletal chain of the finger and reflect the mechanical transmission path between the phalanges. Then, the metacarpophalangeal joints between adjacent fingers are connected. Finally, multiple key points in the wrist area are connected to form a closed and stable polygon, and an edge set is output, which contains all the edges created based on the above connection rules.
[0073] Furthermore, the coordinates of the edge set and key points on the hand surface are input into a constrained triangulation algorithm. The triangulation algorithm forces the inclusion of the edge set, which forms the basic framework. Triangular honeycomb patches are then filled within the basic framework to ensure the generation of a continuous manifold mesh. Then, based on the position of the mesh edges and faces, a functional label is assigned to each triangular patch, i.e., each honeycomb cell. The above process is repeated for each frame in the time series of the hand simulation model to construct a 3D honeycomb network model.
[0074] Step S34: Extract the geometric feature data and physical feature data of the temporal 3D cellular mesh model respectively to generate temporal cellular feature data.
[0075] Specifically, for each cell in the 3D cellular mesh model, the edge length of the cell is calculated, reflecting the Euclidean distance between adjacent key points and embodying joint angles and extension. The three interior angles of the cell are calculated, and changes in these angles reflect local curvature. The surface area of the cell is calculated, and changes in this area reflect local skin stretching and relaxation. The normal vector of the cell is calculated, and its use can infer local rotation or orientation. The local surface curvature at the location of the cell is estimated, which can be used to identify highly curved areas such as finger joints. The edge length, interior angles, surface area, normal vector, and local surface curvature are organized into geometric feature data. Then, physical features are extracted for each cell, including the MRI signal intensity value of the cell, which reflects the physiological state of the tissue beneath it. For example, on T2-weighted images, a higher signal intensity may indicate inflammation or edema. The cell label is extracted to reflect the dominant tissue type of the cell, such as muscle. The distance from the cell to the CT bone surface directly below it is calculated, and this distance can indirectly determine whether soft tissue is swollen. The geometric and physical features are integrated into temporal cellular feature data based on time sequence.
[0076] Step S4: Construct a Riemannian manifold model of the hand based on time-series cellular feature data, and obtain Riemannian quantization feature data through the Riemannian manifold model of the hand.
[0077] In this embodiment, step S4 includes: Step S41: Standardize the time-series cellular feature data, perform linear dimensionality reduction on the time-series cellular feature data based on principal component analysis algorithm, and perform nonlinear dimensionality reduction on the time-series cellular feature data based on nonlinear dimensionality reduction algorithm to obtain the low-dimensional cellular embedding matrix.
[0078] Specifically, Z-score standardization is performed on each column, i.e. each feature dimension, in the time-series cellular feature data to ensure that the feature data in the time-series cellular feature data are of the same magnitude, and to avoid certain large numerical feature data dominating the dimensionality reduction process.
[0079] Furthermore, the temporal cellular feature data is initially reduced in dimensionality using principal component analysis to remove linear correlations and retain most of the variance. Then, a second dimensionality reduction can be performed on the temporal cellular feature data after the initial dimensionality reduction using nonlinear manifold learning algorithms such as t-SNE or UMAP.
[0080] Understandably, nonlinear manifold learning algorithms such as t-SNE or UMAP can better maintain the local proximity relationship between high-dimensional data points. A low-dimensional embedding coordinate matrix, i.e. a low-dimensional cellular embedding matrix, such as two-dimensional or three-dimensional, can be found so that hand poses in high-D space are also close to each other in low-D embedding space.
[0081] Understandably, high-dimensional feature data is often difficult to understand intuitively and distances can be calculated directly. Dimensionality reduction can reveal the inherent low-dimensional manifold hidden in high-dimensional feature data. The proximity of points in the low-dimensional space represents the similarity of hand shapes, which can provide an intuitive and computable space for subsequent Riemannian manifold analysis.
[0082] Step S42: Riemann manifold modeling is performed based on the low-dimensional cellular embedding matrix to construct a hand Riemann manifold model. At the same time, multiple initial motion trajectories are obtained. The normalized path of each pair of initial motion trajectories is calculated according to the dynamic time warping algorithm. Based on the normalized path, multiple initial motion trajectories are resampled to obtain aligned motion trajectories.
[0083] Specifically, by examining the point cloud distribution of the low-dimensional cellular embedding matrix, if the point cloud is a continuous, possibly curved, band-like or clustered structure rather than uniformly dispersed, it can be confirmed that the hand pose space does indeed constitute a manifold. Based on this, the distance between two points on the manifold is defined as the shortest path connecting them. In actual calculations, a K-nearest neighbor graph can be constructed, where nodes are embedding points, edge weights are their Euclidean distances in the embedding space, and the zero distance between two points is the shortest path distance between the two points on the graph, thus defining the Riemann metric. Then, a tangent space is defined for each point on the manifold, i.e., each hand pose. The tangent space is a linear approximation of the manifold near that point. Based on the Riemann metric, the embedding points, and the tangent space of each point, a Riemannian manifold model of the hand is constructed.
[0084] Understandably, an initial motion trajectory represents the coordinates of all hand posture points arranged in chronological order within a single rehabilitation period on the manifold, while multiple initial motion trajectories represent the coordinates of all hand posture points arranged in chronological order within different rehabilitation periods on the manifold.
[0085] Furthermore, in order to compare motion trajectories at different times, the Dynamic Time Warping (DTW) algorithm is used to find the optimal non-linear time alignment between two trajectories. By stretching or compressing the time axis, the two trajectories achieve the best morphological match, and the warped path is calculated. Then, based on the warped path calculated by DTW, the motion trajectories at all times are resampled so that the motion trajectories at all times have the same logical time length, and the corresponding points represent similar hand posture stages, thereby obtaining a sequence of motion trajectories of equal length after DTW alignment.
[0086] Step S43: Based on the Riemannian manifold model of the hand and the aligned motion trajectory, feature extraction is performed to obtain Riemannian quantization feature data.
[0087] Specifically, the shortest paths between consecutive points on the aligned trajectory are accumulated segment by segment to obtain the total trajectory length for the entire movement process. Instantaneous velocity and acceleration are calculated on the manifold. Instantaneous velocity is expressed as the shortest distance between adjacent points divided by the time interval, and acceleration is expressed as the rate of change of velocity. The distribution and smoothness of the entire movement process can be analyzed through instantaneous velocity and acceleration. The better the rehabilitation, the smoother the velocity curve and the fewer sudden acceleration jitters. The local curvature of the trajectory is calculated. High curvature may represent a sudden change in the direction of movement, while lower, smoother curvature usually represents a more fluid and coordinated movement. The distribution range of the starting and ending points of each movement trajectory in the manifold is analyzed. If the starting and ending postures of each movement are more consistent during the patient's rehabilitation process, it indicates that the control precision of the movement has improved. In addition, other indicators can be calculated based on specific circumstances during the implementation process to obtain a more comprehensive analysis and obtain Riemann quantization feature data. Riemann quantization feature data can be expressed as [trajectory length, average length, velocity variance, maximum acceleration, average curvature, ...].
[0088] Step S5: Obtain a trend analysis report based on the trend analysis model.
[0089] In this embodiment, step S5 includes: Step S51: Based on the CT image segmentation data and MRI image segmentation data obtained in step S4, extract bone features from the CT image segmentation data to generate a bone feature sequence, and extract soft tissue features from the MRI image segmentation data to generate a soft tissue feature sequence.
[0090] Specifically, feature extraction is performed on CT image segmentation data, including but not limited to calculating the total number of voxels for each bone in the CT image segmentation data and multiplying it by the physical volume of a single voxel to obtain the volume in cubic millimeters. Simultaneously, the total bone area is calculated, and the surface area to volume ratio is obtained based on the total area and physical volume. The average HU value of all voxels in the CT image segmentation data is calculated, and the standard deviation of the CT value is obtained. These data are integrated into a bone feature sequence, which provides a relatively objective basis for assessing the bone's healing status. For example, an increase in the average HU value may indicate good bone healing progress. Next, feature extraction is performed on MRI image segmentation data. Muscle volume is extracted, which can be used to assess changes in the overall muscle volume. Cross-sectional area is extracted, which can be used to assess muscle atrophy or hypertrophy. Simultaneously, the percentage of fat infiltration is estimated on T1-weighted images, and the average and standard deviation of the T2 signal intensity of the muscle or tendon region are calculated on T2-weighted images. A significantly higher signal intensity than normal may indicate inflammation or damage. T2 relaxation time is obtained, which can be used to assess muscle strain or lesions. All of these data are integrated into a soft tissue feature sequence.
[0091] Step S52: Based on the shooting time, the skeletal feature sequence and soft tissue feature sequence are reconstructed, the reconstructed skeletal feature sequence and soft tissue feature sequence are subjected to missing interpolation, and the soft tissue feature sequence and skeletal feature sequence are sorted.
[0092] Step S53: Import the hand electromyography signal data and hand inertial signal data, soft tissue feature sequence, bone feature sequence and Riemann quantization feature data from the hand video fusion input data into the trend analysis model to obtain the trend analysis report.
[0093] Specifically, the trend analysis model includes an input unit, a fusion unit, and an output unit.
[0094] Furthermore, the input unit includes a functional feature branch, a skeletal feature branch, and a soft tissue feature branch. The functional feature branch is used to input Riemann quantization feature data, hand electromyography signal data, and hand inertial signal data. The skeletal feature branch is used to input skeletal feature sequences, and the soft tissue feature branch is used to input soft tissue feature sequences.
[0095] It should be noted that before inputting the three different time-series features, including function, skeleton, and soft tissue, into the trend analysis model, the three time-series features need to be aligned on the time axis and all features need to be normalized. In addition, to avoid missing values, interpolation correlation methods need to be used to fill in the missing values.
[0096] Specifically, for functional feature branches, temporal convolutional networks can be used as encoders, as they can capture long-term temporal dependencies. For skeletal feature branches, fully connected layers can be used for encoding, and for soft tissue feature branches, 1D CNNs can be used for encoding.
[0097] Furthermore, the fusion unit fuses and splices Riemann quantized feature data, input skeletal feature sequences, and soft tissue feature sequences based on a cross-modal attention mechanism.
[0098] Specifically, by introducing a cross-modal attention mechanism, the data output from the three branches are concatenated. For example, the functional branch is used as the query, and the skeletal and soft tissue branches are used as the key and value. The magnitude of the attention weights can directly reflect the importance of different anatomical features to the current functional surface.
[0099] Furthermore, the output unit includes a trend prediction unit, a status assessment unit, and an anomaly detection module. The trend prediction unit is used to generate a future rehabilitation trend prediction report. The status assessment unit is used to classify and generate a current rehabilitation status assessment report. The anomaly detection module is built based on a generative adversarial network and is used to judge the data of the input unit and generate an anomaly report.
[0100] The trend analysis report includes a current rehabilitation status assessment report, a future rehabilitation trend prediction report, and an anomaly report. The trend analysis report can assist physicians in assessing the patient's condition and can better specify subsequent treatment and rehabilitation training plans.
[0101] Specifically, the output unit is connected after the fusion unit. The trend prediction unit usually consists of one or more fully connected layers to predict feature values or clinical scores for future rehabilitation stages, such as predicting the status after 4 weeks. The status assessment unit usually uses a fully connected layer + Softmax function to assess the current rehabilitation status, which can be divided into poor, average, good, and excellent. The anomaly detection module uses an autoencoder or generative adversarial network structure. For new inputs, it calculates the reconstruction error or anomaly score. If the error is too high, it is judged that an anomaly has occurred.
[0102] Furthermore, when training the trend analysis model, its loss function is first constructed and a training strategy is formulated. The training strategy includes using curriculum learning, time series cross-validation, and early stopping. Curriculum learning is used to pre-train the encoders of each branch, and then the entire model network is fine-tuned end-to-end. Time series cross-validation is used to divide the training set and the validation set, and early stopping is used to prevent overfitting.
[0103] This invention provides a hand rehabilitation trend analysis system based on image processing and key points, applicable to hand rehabilitation trend analysis, including: An initial module is used to acquire hand video fusion input data and medical image data, and to pair the hand video fusion input data and medical image data based on the patient rehabilitation node; A simulation module is used to generate an optimized segmentation mask and construct a hand simulation model; A cellular module, which is used to establish a temporal 3D cellular mesh model and acquire temporal cellular feature data; A manifold module is used to construct a Riemannian manifold model of the hand and obtain Riemannian quantization feature data; The reporting module generates a trend analysis report based on a trend analysis model.
[0104] The specific usage and function of this embodiment are as follows: First, hand video fusion input data and medical image data are acquired, and then paired according to the patient's rehabilitation nodes. Next, an optimized segmentation mask is generated based on the hand video fusion input data and medical image data. Simultaneously, a hand simulation model is constructed based on the optimized segmentation mask, the hand video fusion input data, and the medical image data. Then, a keypoint set and a keypoint set on the hand surface are obtained based on the hand simulation model. Simultaneously, a temporal 3D honeycomb mesh model is established according to honeycomb connection rules, and temporal honeycomb feature data is obtained. Finally, a Riemannian manifold model of the hand is constructed based on the temporal honeycomb feature data. Riemannian quantization feature data is obtained through the hand Riemannian manifold model, and a trend analysis report is obtained based on a trend analysis model. This invention [details about the hand]. Complex information such as posture and movement trajectory is transformed into geometric and kinematic feature parameters, thus providing more objective data for the assessment of hand rehabilitation trends. Furthermore, by fusing CT / MRI image data with video streams from hand video fusion input data that characterize external motor function, a hand simulation model is established, providing more accurate anatomical basis. In addition, based on a 3D cellular mesh model, temporal cellular feature data, Riemann quantization feature data, and hand video fusion input data extracted from the hand simulation model can be used to establish a unified representation and fusion framework based on a trend analysis model. This allows for the assessment of current rehabilitation status, future status, and abnormal status, transforming passive assessment into active assessment, and better assisting physicians in assessing the patient's status and determining subsequent treatment and rehabilitation plans.
[0105] Furthermore, embodiments of the present invention also provide an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method in Embodiment 1 described above.
[0106] The following is a detailed introduction to the various components of the electronic device: In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement Embodiment 1 of this invention, such as one or more digital signal processors (DSPs) or one or more field-programmable gate arrays (FPGAs).
[0107] The processor can perform various functions of an electronic device by running or executing software programs stored in memory and by calling data stored in memory.
[0108] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.
[0109] The memory can be a real-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through an interface circuit of an electronic device; this embodiment of the invention does not specifically limit this.
[0110] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via limited means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0111] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0112] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0113] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A hand rehabilitation trend analysis method based on image processing and key points, characterized in that, The method comprises: S1: Obtain hand video fusion input data and medical image data, and pair the hand video fusion input data and the medical image data based on a patient rehabilitation node; S2: Generate an optimized segmentation mask based on the hand video fusion input data and the medical image data, and simultaneously construct a hand simulation model according to the optimized segmentation mask, the hand video fusion input data and the medical image data; S3: Obtain a key point set and a hand surface key point set according to the hand simulation model, simultaneously establish a time sequence 3D honeycomb grid model according to a honeycomb connection rule, and obtain time sequence honeycomb feature data; S4: Construct a hand Riemann manifold model based on the time sequence honeycomb feature data, and obtain Riemann quantization feature data through the hand Riemann manifold model; S5: Obtain a trend analysis report based on a trend analysis model.
2. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, Obtain hand video fusion input data and medical image data, and pair the hand video fusion input data and the medical image data based on a patient rehabilitation node, comprising: Obtain a hand shooting scheme, obtain multi-view hand video data based on the hand shooting scheme, simultaneously obtain hand electromyographic signal data and hand inertial signal data, obtain a patient rehabilitation node, and obtain hand CT image data and hand MRI image data of different rehabilitation periods based on the patient rehabilitation node; Perform time offset correction on the multi-view hand video data, the hand electromyographic signal data and the hand inertial signal data based on a signal processing algorithm, simultaneously perform distortion correction and polar line correction on the multi-view hand video data, and encapsulate the corrected multi-view hand video data, the hand electromyographic signal data and the hand inertial signal data as hand video fusion input data; Perform data analysis on the hand CT image data and the hand MRI image data to obtain CT image array data and MRI image array data, resample the CT image array data and the MRI image array data, correct the resampled MRI image array data based on an N4ITK algorithm, simultaneously denoise the resampled CT image array data and the corrected MRI image array data, and encapsulate the CT image array data and the MRI image array data as medical image data; Pair the hand video fusion input data of different rehabilitation periods and the corresponding medical image data based on the patient rehabilitation node.
3. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, Generate an optimized segmentation mask based on hand video fusion input data and medical image data, comprising: Extract multi-view hand video data from the hand video fusion input data, obtain each single-view hand video data based on the multi-view hand video data, and calculate hand optical flow field data of each two consecutive video frames in the single-view hand video data based on an optical flow algorithm; Obtain hand optical flow amplitude based on the hand optical flow field data, simultaneously construct an optical flow amplitude graph through the hand optical flow amplitude, position an ROI region in the optical flow amplitude graph, and construct a hand motion energy curve according to the ROI region and the hand optical flow amplitude; Filter the hand motion energy curve and locate the curve extreme points, including the curve peak and the curve valley, perform multi-view consistency verification based on the curve extreme points, and obtain multi-view key frame images, which are represented as key frame images with consistent curve extreme points in all views; Segment the multi-view key frame images based on the SAM model to generate an initial segmentation mask corresponding to each multi-view key frame image, and perform connectivity optimization on the initial segmentation mask to obtain an optimized segmentation mask.
4. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, Meanwhile, a hand simulation model is constructed based on the optimized segmentation mask, hand video fusion input data, and medical image data, including: Obtain the multi-view key frame image and the corresponding optimized segmentation mask at each time point, obtain the hand disparity map at each time point through a stereo matching algorithm, convert the hand disparity map to a hand depth map, and obtain the hand region point cloud at each time point based on the hand depth map, and reconstruct the hand region point cloud to obtain a 3D hand surface mesh; Adaptively cut the CT image array data and MRI image array data in the medical image data to generate CT image segmentation data and MRI image segmentation data, register the MRI image segmentation data and the CT image segmentation data to generate MRI image registration data, and register the CT image segmentation data and the 3D hand surface mesh to generate video surface data; Take the CT coordinate system as the model coordinate system, and construct the hand simulation model based on the MRI image registration data, the video surface data, and the CT surface data.
5. The hand rehabilitation trend analysis method based on image processing and key points according to claim 4, characterized in that, Adaptively cut the CT image array data and MRI image array data in the medical image data to generate CT image segmentation data and MRI image segmentation data, register the CT image segmentation data and the MRI image segmentation data to generate MRI image registration data, and register the CT image segmentation data and the 3D hand surface mesh to generate video surface data, including: Jointly cut the CT image array data based on a threshold segmentation algorithm and connectivity domain analysis, and cut the soft tissue of the MRI image array data based on the nnUnet network, and label the regions after cutting to generate CT image segmentation data and MRI image segmentation data; Extract the CT surface data from the CT image segmentation data, and perform edge detection based on the MRI image segmentation data to obtain MRI pseudo-surface data, and perform rigid registration on the MRI image segmentation data and the CT image segmentation data based on the CT surface data and the MRI pseudo-surface data to obtain MRI image registration data; Preliminarily register the CT surface data and the 3D hand surface mesh through principal component analysis, and non-rigidly register the preliminarily registered CT surface data and the 3D hand surface mesh through a non-rigid ICP algorithm to obtain video surface data.
6. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, According to the hand simulation model, obtain a key point set and a hand surface key point set, and establish a time sequence 3D honeycomb mesh model according to the honeycomb connection rule, and obtain time sequence honeycomb feature data, including: Position the centers of the interphalangeal joints and the centers of the wrist joints based on the CT surface data in the hand simulation model, and position the fingertip points, the phalangeal points, and the wrist points to obtain a key point set; Map the key point set to the 3D hand surface grid in the hand simulation model based on the non-rigid registration relationship between the CT surface data and the 3D hand surface grid in the hand simulation model, and obtain a hand surface key point set based on the MRI image segmentation data in the hand simulation model; Connect the hand surface key points of the same finger, connect the hand surface key points of adjacent fingers, and connect the hand surface key points of the wrist region, and generate a triangular honeycomb grid between the connected hand surface key points to generate a time-series 3D honeycomb grid model; Extract the geometric feature data and the physical feature data of the time-series 3D honeycomb grid model respectively to generate time-series honeycomb feature data.
7. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, Based on the time-series honeycomb feature data, a hand Riemann manifold model is constructed, and Riemann quantization feature data is obtained through the hand Riemann manifold model, including: Standardize the time-series honeycomb feature data, perform linear dimension reduction on the time-series honeycomb feature data based on a principal component analysis algorithm, and perform non-linear dimension reduction on the time-series honeycomb feature data based on a non-linear dimension reduction algorithm to obtain a low-dimensional honeycomb embedding matrix; Based on the low-dimensional honeycomb embedding matrix, a Riemann manifold model is constructed, and a plurality of initial motion trajectories are obtained, a regular path of each two initial motion trajectories in the plurality of initial motion trajectories is calculated based on a dynamic time warping algorithm, the plurality of initial motion trajectories are resampled based on the regular path to obtain aligned motion trajectories; Based on the hand Riemann manifold model and the aligned motion trajectories, feature extraction is performed to obtain Riemann quantization feature data.
8. The hand rehabilitation trend analysis method based on image processing and key points according to claim 1, characterized in that, Based on the trend analysis model, a trend analysis report is obtained, including: Based on the CT image segmentation data and the MRI image segmentation data obtained in step S4, bone feature extraction is performed on the CT image segmentation data to generate a bone feature sequence, and soft tissue feature extraction is performed on the MRI image segmentation data to generate a soft tissue feature sequence; Based on the shooting time, the bone feature sequence and the soft tissue feature sequence are reorganized, the reorganized bone feature sequence and the soft tissue feature sequence are subjected to a missing interpolation operation, and the soft tissue feature sequence and the bone feature sequence are sorted; The hand electromyographic signal data and the hand inertial signal data in the hand video fusion input data, the soft tissue feature sequence, the bone feature sequence, and the Riemann quantization feature data are imported into the trend analysis model to obtain a trend analysis report; The trend analysis report includes a current rehabilitation state evaluation report, a future rehabilitation trend prediction report, and an abnormal report.
9. The hand rehabilitation trend analysis method based on image processing and key points according to claim 8, characterized in that, The trend analysis model includes: The trend analysis model includes an input unit, a fusion unit, and an output unit; The input unit includes a functional feature branch, a bone feature branch, and a soft tissue feature branch, the functional feature branch is used to input the Riemann quantization feature data, the hand electromyographic signal data, and the hand inertial signal data, the bone feature branch is used to input the bone feature sequence, and the soft tissue feature branch is used to input the soft tissue feature sequence; The fusion unit fuses and splices the riemannian quantization feature data, the input skeleton feature sequence and the soft tissue feature sequence based on a cross-modal attention mechanism; The output unit includes a trend prediction unit, a state evaluation unit and an anomaly detection module, the trend prediction unit is used to generate a future rehabilitation trend prediction report, the state evaluation unit is used to classify and generate a current rehabilitation state evaluation report, and the anomaly detection module is constructed based on a generative adversarial network and is used to judge the data of the input unit to generate an anomaly report.
10. A hand rehabilitation trend analysis system based on image processing and key points, for implementing the method of any one of claims 1 to 9, characterized in that, Comprise: An initial module, the initial module is used for obtaining hand video fusion input data and medical image data, and pairing the hand video fusion input data and the medical image data based on a patient rehabilitation node; A simulation module, the simulation module is used for generating an optimized segmentation mask and constructing a hand simulation model; A honeycomb module, the honeycomb module is used for establishing a time sequence 3D honeycomb grid model and obtaining time sequence honeycomb feature data; A manifold module, the manifold module is used for constructing a hand riemannian manifold model and obtaining riemannian quantization feature data; A report module, the report module obtains a trend analysis report based on a trend analysis model.