Articular cartilage image segmentation method and system based on mixed attention mechanism

By mixing attention mechanism and biomechanical constraints, a method for articular cartilage image segmentation based on dynamic modeling of multimodal data is proposed, which solves the problem that multimodal data fusion is not adaptable to changes in load conditions and achieves high-precision cartilage injury segmentation and grading.

CN120635441AActive Publication Date: 2025-09-12TIANJIN PEOPLE HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510684052.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-12
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing technologies lack dynamic modeling of the spatiotemporal correlation of multimodal data in articular cartilage injury assessment, resulting in feature fusion being unable to adapt to changes in load conditions and affecting the accuracy of injury grading.

Method used

A method based on a hybrid attention mechanism is used to identify the spatiotemporal correlation features between piezoelectric mechanical data and polarized light stress data. A multimodal fusion feature map is generated through feature enhancement and dynamic feature weighting processing. Combined with cross-scale feature aggregation and biomechanical constraints, high-precision segmentation of cartilage deformation areas is achieved.

Benefits of technology

The temporal and spatial consistency error of cartilage deformation area segmentation has been improved by 45%, which has increased the detection rate of minor injuries. The correlation coefficient between the segmentation results and the clinical mechanical evaluation results exceeds 0.95, supporting the accurate grading of cartilage injuries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635441A_ABST
    Figure CN120635441A_ABST
Patent Text Reader

Abstract

The invention provides an articular cartilage image segmentation method and system based on a mixed attention mechanism. The method comprises the following steps: acquiring piezoelectric mechanical data and polarized light stress data of an articular cartilage under a dynamic load condition; identifying space-time correlation characteristics between the piezoelectric mechanical data and the polarized light stress data; performing feature enhancement on the space-time correlation features to obtain enhanced space-time correlation features; the enhanced space-time correlation features are input into a pre-trained segmentation network, the pre-trained segmentation network uses dynamic feature weights to perform weighting processing on features of different modals, a segmentation result is output, the segmentation result comprises a deformation area of the articular cartilage, and the articular cartilage injury level is determined according to the deformation area of the articular cartilage. The dynamic feature weights are generated by a mixed attention mechanism. According to the technical scheme provided by the invention, accurate segmentation and grading evaluation of the articular cartilage injury are realized by fusing the multi-modal mechanical data and adopting dynamic weighting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and system for articular cartilage image segmentation based on a hybrid attention mechanism. Background Art

[0002] In the field of articular cartilage injury assessment, real-time monitoring of the mechanical response of cartilage under high dynamic load conditions is necessary to accurately identify early-stage damage and quantify the extent of damage. During this monitoring process, it is necessary to efficiently integrate multimodal mechanical data (such as piezoelectric mechanical and optical stress data) and establish a mapping relationship between mechanical characteristics and damaged areas to support accurate clinical diagnosis and treatment decisions.

[0003] Existing solutions, such as multimodal data fusion methods based on convolutional neural networks, are widely used. This method extracts the spatial features of piezoelectric mechanical data and polarized light stress data separately, fuses them through simple feature concatenation or weighted averaging, and finally inputs them into a segmentation network to predict the cartilage damage area.

[0004] However, existing solutions lack dynamic modeling of the spatiotemporal correlation of multimodal data in the feature fusion stage, resulting in fixed contributions of different modal data and an inability to adapt to changes in load conditions. In particular, feature suppression or noise amplification problems are prone to occur in complex mechanical environments, affecting the accuracy of damage grading. Summary of the Invention

[0005] The present application provides a method and system for articular cartilage image segmentation based on a hybrid attention mechanism to solve the problem of low accuracy in articular cartilage image segmentation in the prior art.

[0006] In a first aspect, the present application provides an articular cartilage image segmentation method based on a hybrid attention mechanism, comprising:

[0007] Acquire piezoelectric and electromechanical data and polarized light stress data of articular cartilage under dynamic loading conditions;

[0008] identifying spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data;

[0009] Performing feature enhancement on the spatiotemporal correlation feature to obtain enhanced spatiotemporal correlation feature;

[0010] The enhanced spatiotemporal correlation features are input into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs a segmentation result. The segmentation result includes the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

[0011] Optionally, the step of inputting the enhanced spatiotemporal correlation features into a pre-trained segmentation network, wherein the pre-trained segmentation network performs weighted processing on features of different modalities using dynamic feature weights and outputs a segmentation result, including:

[0012] Inputting the enhanced spatiotemporal correlation features into a pre-trained segmentation network, the segmentation network performs channel division on the enhanced spatiotemporal correlation features according to time series and spatial distribution to generate a time channel feature set and a spatial distribution channel feature set;

[0013] Based on the temporal channel feature set and the spatial distribution channel feature set, a hybrid attention mechanism is used to calculate the spatiotemporal correlation strength of each channel to generate dynamic feature weights;

[0014] According to the dynamic feature weights, the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set are weightedly fused channel by channel to generate a multimodal fusion feature map;

[0015] Performing cross-scale feature aggregation on the multimodal fusion feature map to obtain an aggregation result;

[0016] Based on the continuity constraint condition of the preset articular cartilage deformation boundary and the aggregation result, the deformation area is predicted to obtain a segmentation result of the deformation area containing the articular cartilage.

[0017] Optionally, the step of calculating the spatiotemporal correlation strength of each channel based on the temporal channel feature set and the spatial distribution channel feature set using a hybrid attention mechanism to generate a dynamic feature weight includes:

[0018] Performing a multi-dimensional correlation analysis between the temporal channel feature set and the spatial distribution channel feature set to generate a cross-modal correlation matrix;

[0019] Based on a preset articular cartilage principal stress direction consistency constraint and a preset articular cartilage deformation boundary continuity constraint, a directional correction is performed on the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix;

[0020] Normalizing the corrected spatiotemporal correlation strength of each channel to obtain the normalized spatiotemporal correlation strength of each channel;

[0021] According to the normalized spatiotemporal correlation strength of each channel and the preset viscoelastic property constraints of articular cartilage, the channel-level dynamic feature weight is generated.

[0022] Optionally, the directional correction of the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix based on a preset articular cartilage principal stress direction consistency constraint and a preset articular cartilage deformation boundary continuity constraint includes:

[0023] Extracting temporal channel features and spatial distribution channel features of each channel from the cross-modal correlation matrix;

[0024] The angle between the principal stress direction of the time channel characteristic and the deformation boundary direction of the spatial distribution channel characteristic of each channel is calculated to generate the directional deviation;

[0025] generating a directional deviation correction weight based on the directional deviation and the preset viscoelastic property constraint of the articular cartilage;

[0026] According to the directional deviation correction weight, the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix is ​​iteratively corrected until the spatial distribution relationship between the principal stress direction and the deformation boundary direction meets the preset viscoelastic property constraint of the articular cartilage, thereby obtaining the corrected spatiotemporal correlation strength of each channel.

[0027] Optionally, performing channel-by-channel weighted fusion of the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set according to the dynamic feature weights to generate a multimodal fusion feature map includes:

[0028] Based on the mechanical properties of the stress concentration area of ​​articular cartilage, the dynamic feature weight is modified;

[0029] According to the corrected dynamic feature weights, the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set are nonlinearly superimposed channel by channel to generate a multimodal fusion feature map.

[0030] Optionally, identifying the spatiotemporal correlation characteristics between the piezoelectric mechanical data and the polarized light stress data includes:

[0031] performing spatiotemporal registration of the piezoelectric mechanical data and the polarized light stress data along the grid cells on the surface of the articular cartilage to generate spatiotemporally synchronized piezoelectric mechanical data and polarized light stress data;

[0032] For each grid cell on the surface of the articular cartilage, extracting the spatiotemporal dynamic fluctuation characteristics of the piezoelectric force and the directional distribution characteristics of the polarized light stress from the combined piezoelectric force and polarized light stress data;

[0033] A dynamic mode decomposition algorithm is used to extract the coupled modes of the spatiotemporal dynamic fluctuation characteristics and the polarized light stress direction distribution characteristics in the frequency domain, and the spatiotemporal correlation characteristics are generated according to the energy proportion of the coupled modes.

[0034] Optionally, the performing feature enhancement on the spatiotemporal correlation feature to obtain the enhanced spatiotemporal correlation feature includes:

[0035] Performing wavelet packet transformation on the spatiotemporal correlation characteristics to obtain a high-frequency stress fluctuation component and a low-frequency viscoelastic relaxation component;

[0036] Performing feature enhancement on the high-frequency stress fluctuation component to obtain a feature-enhanced high-frequency stress fluctuation component;

[0037] The enhanced high-frequency stress fluctuation component and the low-frequency viscoelastic relaxation component are fused through the jump connection in the U-shaped convolutional network architecture to obtain the enhanced spatiotemporal correlation feature.

[0038] In a second aspect, the present application provides an articular cartilage image segmentation system based on a hybrid attention mechanism, comprising:

[0039] an acquisition module for acquiring piezoelectric mechanical data and polarized light stress data of articular cartilage under dynamic loading conditions;

[0040] an identification module, configured to identify spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data;

[0041] An enhancement module, configured to enhance the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features;

[0042] An output module is used to input the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs segmentation results. The segmentation results include the deformation area of ​​the articular cartilage, so as to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

[0043] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an articular cartilage image segmentation method based on a hybrid attention mechanism as described in any one of the first aspects.

[0044] In a fourth aspect, the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements an articular cartilage image segmentation method based on a hybrid attention mechanism as described in any one of the first aspects.

[0045] In the present application, a method for articular cartilage image segmentation based on a hybrid attention mechanism is provided, which includes: obtaining piezoelectric mechanical data and polarized light stress data of articular cartilage under dynamic load conditions; identifying the spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data; performing feature enhancement on the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features; inputting the enhanced spatiotemporal correlation features into a pre-trained segmentation network, the pre-trained segmentation network uses dynamic feature weights to weight features of different modalities, and outputs a segmentation result, which includes the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage, and the dynamic feature weights are generated by the hybrid attention mechanism.

[0046] This application constructs a multi-physics coupled sensing system by simultaneously collecting piezoelectric and polarized stress data of articular cartilage under dynamic loads. This system overcomes the limitations of single-modal data representation and enables the collaborative analysis of mechanical and optical properties. Feature enhancement algorithms, such as spatiotemporal convolution or graph neural networks, extract spatiotemporal correlations between the piezoelectric and polarized stress signals, such as stress-charge coupling phase difference and strain rate correlation, to improve the signal-to-noise ratio of weak damage signals. A hybrid attention mechanism dynamically weights the piezoelectric and polarized stress features, prioritizing mechanical data under high load and optical features under minimal damage. This enables the segmentation network to achieve submillimeter accuracy in localizing cartilage deformation regions, a 35% improvement over fixed-weight strategies. By combining segmentation results (deformation area and edge sharpness) with biomechanical prior knowledge (such as the Young's modulus attenuation gradient), a four-level classification of cartilage damage (normal / mild / moderate / severe) is achieved, with a clinical agreement rate exceeding 92%. Furthermore, after the enhanced spatiotemporal correlation features are input into the segmentation network, they are first segmented into channels based on their temporal sequence and spatial distribution, generating a temporal channel feature set and a spatial distribution channel feature set. A cross-modal correlation matrix is ​​calculated for these two feature sets using a hybrid attention mechanism. Directional correction and normalization are performed, combining constraints based on the consistency of the principal stress directions and the continuity of the deformation boundaries of the articular cartilage, generating channel-level dynamic feature weights. Subsequently, the temporal and spatial features are weightedly fused channel by channel to form a multimodal fusion feature map. After cross-scale feature aggregation, the cartilage deformation region is predicted under viscoelastic constraints, resulting in a high-precision segmentation result. This method achieves adaptive dynamic fusion of piezoelectric and polarization data through the synergistic effect of the hybrid attention mechanism and multiple constraints (such as directionality, continuity, and viscoelasticity), reducing the spatiotemporal consistency error in cartilage deformation region segmentation by 45%. Cross-scale feature aggregation, combined with biomechanical prior constraints, improves the detection rate of minor lesions. The correlation coefficient between the segmentation results and clinical mechanical assessment exceeds 0.95, providing a reliable basis for the accurate grading of articular cartilage lesions.

[0047] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0049] Figure 1 A flowchart of a method for segmenting articular cartilage images based on a hybrid attention mechanism provided in an embodiment of the present application;

[0050] Figure 2 A schematic diagram of the structure of an articular cartilage image segmentation system based on a hybrid attention mechanism provided in an embodiment of the present application;

[0051] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0053] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 11, 12, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0055] Figure 1A flow chart of a method for segmenting articular cartilage images based on a hybrid attention mechanism is provided in an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0056] S11. Obtain piezoelectric mechanical data and polarized light stress data of articular cartilage under dynamic loading conditions.

[0057] Dynamic loading conditions refer to the time-varying mechanical loads to which articular cartilage is subjected during movement, including periodic, impact, or random stresses. Unlike static loading (constant pressure), dynamic loading triggers viscoelastic responses (such as stress relaxation and creep) and electromechanical coupling effects (piezoelectric properties) in the cartilage, which are key factors leading to the accumulation of cartilage damage. Piezoelectric electromechanical data is the charge distribution / voltage signal generated by articular cartilage under dynamic mechanical loads, collected by piezoelectric sensors, reflecting the electromechanical coupling properties of cartilage tissue. Polarized light stress data is an image of the birefringence stress distribution within the cartilage, acquired using a polarization optical imaging system. It characterizes the mechanical response of the tissue microstructure and captures the birefringence characteristics of the stress distribution within the cartilage.

[0058] For example, in the mechanics of knee cartilage, ex vivo human knee cartilage (femoral condyle) was prepared and preserved in physiological saline. Cyclic pressure (0.5 Hz frequency, 0-5 MPa stress) was applied using a biomechanical testing machine to simulate the joint load during walking. An array of micro-piezoelectric sensors (with 1 mm spacing between sensors within the array) was implanted on the cartilage surface, and voltage signals were recorded in real time at a sampling rate of 1 kHz under load. At a pressure of 3 MPa, sensor A output voltage was 12 mV, and sensor B output voltage was 8 mV (reflecting uneven stress distribution). A polarized optical coherence tomography (OCT) system was used to scan the loaded cartilage, acquiring birefringence images with a resolution of 10 μm. Output data included piezoelectric mechanical data and polarized stress data.

[0059] S12. Identify the spatiotemporal correlation characteristics between piezoelectric mechanical data and polarized light stress data.

[0060] Among them, the spatiotemporal correlation characteristics refer to the cross-modal coupling relationship between the time-varying law of the piezoelectric signal (such as peak frequency offset) and the spatial distribution of polarization stress (such as the principal stress vector field).

[0061] S13. Perform feature enhancement on the spatiotemporal correlation feature to obtain enhanced spatiotemporal correlation feature.

[0062] Feature enhancement is achieved by enhancing the signal-to-noise ratio of damage-related features through nonlinear transformation, suppressing interference such as motion artifacts. The enhanced spatiotemporal correlation features are fused features obtained by enhancing the original correlation features of the piezoelectric mechanical data and polarized light stress data.

[0063] S14. Input the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs segmentation results. The segmentation results include the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

[0064] Among them, the dynamic feature weight is the channel weight coefficient generated by the hybrid attention mechanism, which reflects the contribution of different modal features to damage identification. In the segmentation network, segmentation can be performed in the following ways: extract features of different scales through convolutional layers or dilated convolutions to capture transient response and spatial continuity information respectively. In the decoder part of the segmentation network, a boundary detection module, such as edge detection or conditional random field, is introduced to optimize the segmentation boundary using the spatial continuity features of polarized light stress data. In the encoder part of the segmentation network, a feature processing module of the time dimension is introduced to enhance the dynamic characteristics using the transient response features of piezoelectric mechanical data. The features of different modalities are independent feature representations of multi-source data (such as piezoelectric and polarized light) within the segmentation network.

[0065] Here's a specific example:

[0066] Dynamic loads (0-500 N, 1 Hz frequency, simulating human walking load) were applied to knee cartilage samples measuring 10 cm × 8 cm × 3 mm. An implantable piezoelectric sensor array (resolution 0.1 Pa, sampling rate 1 kHz) was used to capture microstress fluctuations in the surface and deep layers of the cartilage under dynamic load, generating time series data (dimensions [T × 128], T = 3000 time slices). A polarization imaging system (wavelength 532 nm, spatial resolution 50 μm) was used to capture the principal stress directions of collagen fibers within the cartilage, generating a stress tensor matrix (dimensions [100 × 80 × 3], with a 100 × 80 grid representing the three principal stress directions). The piezoelectric data were divided into 300 time windows (10 ms per window) according to the load cycle (1 Hz). The polarization data were spatially registered within the same time slice to generate a spatiotemporally synchronized joint dataset. The high-frequency fluctuation amplitude (100-500 Hz energy fraction) and principal stress direction within each time window were extracted. The principal stress direction and stress gradient amplitude (derivative along the collagen fiber direction) were extracted for each mesh. The spatial angle between the piezoelectric principal stress direction and the polarized light principal stress direction was calculated (regions with an average angle ≤15° were marked as highly correlated). The Pearson correlation coefficient between the piezoelectric high-frequency fluctuation amplitude and the polarized light stress gradient was analyzed, and a spatiotemporal correlation feature set (dimensions [100 × 80 × 5], including angle, correlation coefficient, piezoelectric amplitude, polarization gradient amplitude, and viscoelastic relaxation rate) was output. Directional correction was performed by increasing the piezoelectric amplitude weight by the dynamic coupling coefficient (α = 1.2) for regions with angles < 15°, and compressing the polarization gradient amplitude by the viscoelastic relaxation coefficient (β = 0.6) for regions with angles > 30°. For deep regions (compression-dominated), if the piezoelectric amplitude exceeds the viscoelastic bearing threshold of 200 Pa, the weight is reduced by β = 0.8. For surface regions (shear-dominated), if the polarization gradient and the piezoelectric fluctuation trend are consistent, the cross-modal correlation weight is increased by α = 1.5. The enhanced spatiotemporal correlation feature set is then fed into an enhanced feature set ([100 × 80 × 5]), where multi-scale features are extracted through three convolutional layers. A hybrid attention module assigns a weight wc = 0.7 to deep compression features and wc = 1.3 to surface shear features. A probability map of deformed regions is output (dimension [100 × 80], with a threshold above 0.5 marking damaged regions). The area percentage of deformed regions is calculated (15% for mild damage, 30% for moderate damage, and >50% for severe damage). Combined with the stress amplitude of the deformed region (>300 Pa is upgraded to the first level of damage), the damage grade is concordant with histological findings at 90% (accuracy of 92% for mild, 88% for moderate, and 85% for severe damage, respectively).

[0067] By executing S11-S14, this embodiment of the present application achieves highly sensitive, automated quantitative assessment of articular cartilage damage through multimodal data fusion and dynamic feature optimization. This method is particularly suitable for early diagnosis and dynamic mechanical research, providing a new clinical tool. Future research can explore combining this method with 3D reconstruction technology to further optimize spatial resolution and thereby improve the accuracy of articular cartilage image segmentation.

[0068] In one possible embodiment, S14 inputs the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to perform weighted processing on features of different modalities and outputs a segmentation result, including:

[0069] Step 141: Input the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The segmentation network divides the enhanced spatiotemporal correlation features into channels according to time series and spatial distribution to generate a time channel feature set and a spatial distribution channel feature set.

[0070] Among them, the enhanced spatiotemporal correlation features refer to the deep feature representation that integrates time series information and spatial distribution information. The time channel feature set contains the feature expressions of different time points, and the spatial distribution channel feature set contains the feature expressions of different spatial positions.

[0071] Step 142: Based on the temporal channel feature set and the spatial distribution channel feature set, a hybrid attention mechanism is used to calculate the spatiotemporal correlation strength of each channel to generate a dynamic feature weight.

[0072] Dynamic feature weights refer to the importance coefficients of each channel calculated through the attention mechanism, which are used to characterize the contribution of different spatiotemporal features. The hybrid attention mechanism is a computational method that combines channel attention and spatial attention to dynamically assess the importance of different feature channels and spatial locations. Channel attention calculates the weights of different channels (such as temporal and spatial channels) to measure their contribution to the final task. Spatial attention calculates the weights of different spatial locations in the feature map to enhance the response of key areas. A hybrid approach combines the two attention methods in series, parallel, or weighted fusion, allowing the model to simultaneously focus on important channels and key spatial regions.

[0073] For example, in Magnetic Resonance Imaging (MRI) analysis of articular cartilage, a temporal channel feature set (e.g., 16 frames × 256 channels) and a spatial channel feature set (256 × 256 × 256) are input. Through mixed attention calculation, the channel attention is found to have higher weights for frames 8-12 (mid-deformation phase) and the cartilage contact region, facilitating the subsequent identification of the contact region between the femoral condyle and the tibial plateau as a key spatial location.

[0074] Step 143: Based on the dynamic feature weights, perform channel-by-channel weighted fusion on the temporal channel features in the temporal channel feature set and the spatial distribution channel features in the spatial distribution channel feature set to generate a multimodal fusion feature map.

[0075] The temporal channel feature set extracts temporal features from MRI sequences, representing the dynamic changes of articular cartilage at consecutive time points. For example, 16 temporal channel features are extracted from 16 MRI frames (one channel per frame). The spatial distribution channel feature set extracts spatial features from a single MRI frame, representing the anatomical distribution of cartilage. For example, a 256×256 resolution feature map contains spatial information such as cartilage thickness and curvature. A multimodal fusion feature map is a unified feature representation obtained through weighted fusion, while preserving the important information of temporal and spatial features.

[0076] For example, knee MRI cartilage analysis involves inputting data: a temporal feature set consisting of 16 temporal channels (corresponding to 16 MRI frames), with each frame's feature map sized at 256 × 256 × 1. A spatial feature set consisting of 16 spatial channels (e.g., different anatomical layers), with each channel sized at 256 × 256 × 1, is used. Temporal weighting is set at 0.9 for frames 8–12 (knee flexion phase) and 0.5 for the remaining frames. Spatial weighting is set at 0.8 for the cartilage contact region (femoral condyle and tibial plateau) and 0.3 for the remaining regions. Feature maps for frames 8–12 are multiplied by a weight of 0.9, and for the remaining frames by 0.5, to enhance feature responses during critical periods of motion. Pixels in the cartilage contact region are multiplied by a weight of 0.8, while pixels in the remaining regions are multiplied by a weight of 0.3 to highlight deformation-sensitive areas. The weighted temporal and spatial features (16 channels) are then concatenated according to the channel dimension to generate a 32-channel fused feature map (256 × 256 × 32). The output multimodal fusion feature map contains enhanced flexion period temporal features and high-weight contact area spatial features. When used for subsequent deformation area segmentation, the network will pay more attention to the cartilage contact deformation area during movement.

[0077] Step 144: Perform cross-scale feature aggregation on the multimodal fusion feature map to obtain an aggregation result.

[0078] Among them, cross-scale feature aggregation is to integrate feature maps of different resolutions (i.e., scales) to fuse local details and global contextual information, and solve the perception limitations of single-scale features for small targets (such as cartilage micro-damage) or large-scale deformations (such as full joint movement).

[0079] For example, the input multimodal fusion feature map is 256×256×32 (spatial resolution 256×256, 32 fusion channels). Channels 1–16 are weighted temporal dynamic features (e.g., cartilage deformation during flexion), and channels 17–32 are weighted spatial distribution features (e.g., cartilage thickness and contact area). Multiscale features are extracted from the original resolution of 256×256×32 (preserving high-frequency details). After downsampling by 2, the size is 128×128×32 (capturing mid-range context). After downsampling by 4, the size is 64×64×32 (extracting global semantic information). The downsampled features are upsampled to 128×128 and concatenated with the downsampled features (64 channels). The fused mid-scale features are compressed to 32 channels using a 3×3 convolution to reduce computational overhead. The fused mid-scale features are then upsampled to 256×256 and concatenated with the original resolution features (64 channels). The image is compressed to 32 channels using a 3×3 convolution to preserve details. The resulting aggregated output is 256×256×32 (the same resolution as the input) and contains the following information: low-scale global information (such as full joint alignment), mid-scale local context (such as the cartilage-bone interface), and pixel-level details at the original scale (such as microcracks on the cartilage surface).

[0080] Step 145 : Based on the continuity constraint condition of the preset articular cartilage deformation boundary and the aggregation result, the deformation region is predicted to obtain a segmentation result of the deformation region including the articular cartilage.

[0081] Continuity constraints are mathematical constraints based on the physiological properties of articular cartilage (such as smoothness and anatomical connectivity). They ensure that the segmentation results conform to actual biomechanical behavior and prevent unreasonable fractures or mutations in the predicted deformation region (for example, cartilage cracks should not suddenly cross discontinuous regions). Deformation region prediction maps aggregated features into a binary mask (0 / 1) through the segmentation network, marking cartilage deformation regions (such as thinning and fibrosis). The final segmentation result is a binary image of the cartilage deformation region, which is used to quantify the extent of damage or guide clinical intervention.

[0082] For example, knee MRI cartilage segmentation input data is based on the multi-scale feature map of the aggregation result of 256×256×32. The anatomical constraint is that cartilage deformation is only allowed to occur at the joint contact surface (such as the femoral condyle-tibial plateau). The aggregation result is input into a 3-layer convolutional network, and a 256×256×1 probability map is output (each pixel value is 0~1, indicating the possibility of deformation). Conditional random field optimization is used to penalize isolated noise points and non-smooth boundaries. Tiny holes are filled through morphological post-processing (such as closing operation) to ensure the connectivity of the cartilage area. The output segmentation result binary mask is a 256×256 matrix, with the white area (1) being the deformed cartilage and the black area (0) being the normal tissue.

[0083] The following is a specific example: To segment cartilage deformation in knee MRI, a dynamic knee MRI sequence (32 frames, 256×256 resolution) is input, containing cartilage deformation information throughout flexion-extension motion. The input is enhanced spatiotemporal correlation features. Frame-by-frame segmentation generates 32 temporal channels (each corresponding to a 256×256 feature map) to capture the dynamic changes of cartilage during motion. Frame-by-frame segmentation generates 32 spatial channels (such as cartilage thickness and curvature) to describe the static structural distribution. The output is a temporal channel feature set (32×256×256) and a spatial channel feature set (32×256×256). Using mixed attention, these temporal and spatial channel feature sets are weighted more heavily (0.9) to mid-flexion (frames 12–20) and less heavily (0.3) to the static phase (frames 1–5). The cartilage contact region (femoral condyle–tibial plateau) is weighted 0.8, while the non-contact region is weighted 0.2. The output dynamic feature weight matrix (32×256×256) identifies key spatiotemporal features. The input multimodal fusion feature map is 256×256×64, generated by weighted fusion of temporal and spatial features. Multi-scale fusion preserves microcrack details on the cartilage surface at the original scale (256×256), captures mid-range context at the cartilage-bone interface by downsampling to 1 / 2 (128×128), and integrates global semantics of the entire joint motion by downsampling to 1 / 4 (64×64). The output aggregated result (256×256×32) fuses features from multiple resolutions. The aggregated result and the initial segmentation network with continuity constraints are fed into the output probability map, indicating the deformation candidate regions. Morphological closing optimization is performed to eliminate isolated noise points and ensure continuous deformation boundaries. The final segmentation result (256×256 binary mask) is output, with the deformation region coefficient increased to 0.91.

[0084] By executing steps 141 to 145, the method of the embodiment of the present application achieves high-precision and high-robustness segmentation of the articular cartilage deformation area. The enhanced spatiotemporal correlation features are decoupled into a time channel feature set and a spatial distribution channel feature set through spatiotemporal feature channel division, respectively capturing the dynamic deformation process and static anatomical structure of the cartilage; a hybrid attention mechanism is used to calculate the spatiotemporal correlation strength and generate dynamic feature weights, so that the model can adaptively focus on key time frames and important spatial areas; multi-resolution information is integrated through cross-scale feature aggregation, which not only retains the local features of the microstructure of the cartilage surface, but also integrates the global context of the overall movement of the joint; the aggregation results are optimized and predicted under the guidance of continuity constraints to ensure that the segmentation results are consistent with the physiological characteristics of the cartilage, and output deformation area segmentation results with anatomical rationality. The overall process improves the sensitivity and segmentation accuracy of cartilage micro-deformation detection, and provides a reliable image analysis method for early diagnosis and precise treatment planning of osteoarthritis.

[0085] In one possible embodiment, step 142, based on the temporal channel feature set and the spatial distribution channel feature set, a hybrid attention mechanism is used to calculate the spatiotemporal correlation strength of each channel to generate a dynamic feature weight, including:

[0086] Step a1: Perform inter-channel multi-dimensional correlation analysis on the temporal channel feature set and the spatial distribution channel feature set to generate a cross-modal correlation matrix.

[0087] The spatial distribution channel feature set is a set of spatial features (e.g., 256×256×32) extracted from a single MRI frame, describing static properties such as cartilage thickness and curvature. The cross-modal correlation matrix is ​​a matrix (32×32) generated through correlation analysis, quantifying the coupling strength between the temporal and spatial channels.

[0088] For example, dynamic MRI analysis of the knee joint uses 32 temporal channels (corresponding to 32 frames of dynamic MRI), each containing a 256×256 feature map. The feature map for frame 15 (mid-flexion) highlights areas of cartilage compression. The 32 spatial channels include: channels 1–16: cartilage thickness distribution (e.g., value = 1.0 at the thickest point of the femoral condyle, 0.2 at the edge); channels 17–32: curvature features (e.g., curvature = 0.8 in the contact zone, 0.1 in the flat zone). For each pair of temporal channels (T_i) and spatial channels (S_j), the difference between their joint distribution and their independent distributions is calculated. MI(T_i, S_j)=∑ p(T_i, S_j) log [p(T_i, S_j) / (p(T_i)p(S_j))], MI(T_15, S_5)=0.75 (time frame 15 is strongly correlated with thickness channel 5, indicating localized thinning due to compression); MI(T_1, S_20)=0.10 (still frame 1 is unrelated to curvature channel 20). Generate a correlation matrix to construct a 32×32 matrix, with rows and columns corresponding to the temporal and spatial channels, respectively. High values ​​in the output cross-modal correlation matrix indicate strong coupling between cartilage deformation and specific anatomical attributes.

[0089] Step a2: Based on the preset articular cartilage principal stress direction consistency constraint and the preset articular cartilage deformation boundary continuity constraint, the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix is ​​directionally corrected.

[0090] The principal stress direction consistency constraint ensures that the dominant stress direction of cartilage collagen fibers under load is aligned with the anatomical structure (for example, the fibers of the femoral condyle are radially arranged). The deformation boundary continuity constraint ensures that the cartilage deformation area has a smooth transition to avoid sudden changes.

[0091] For example, in a dynamic analysis of knee cartilage, the input data consists of a raw cross-modal correlation matrix (32×32), where rows represent temporal channels (e.g., T_15 for mid-flexion); columns represent spatial channels (e.g., S_5 for femoral condyle thickness); and values ​​represent correlation strengths (ranging from 0 to 1). The pre-set constraint is that the principal stress directions are primarily radial (angles 0° to 30°) for collagen fibers in the femoral condyle. A continuity threshold is set, requiring the difference in correlation strengths between adjacent spatial channels to be less than 0.2. Correction for principal stress direction consistency revealed that the raw correlation value between T_15 (compression) and S_10 (transverse shear channel) was 0.7, but this direction is perpendicular to the fiber orientation (radial), indicating nonphysiological coupling. Multiplying this value by a damping factor of 0.3 yields a corrected correlation value of 0.7 × 0.3 = 0.21. The correlation value between S_5 (thickness channel) and T_15 is 0.8, but the correlation value for its adjacent channel, S_6, drops sharply to 0.3 (a difference of 0.5 > the threshold of 0.2). The correlation values ​​of S_5 and S_6 were Gaussian filtered, and after correction, S_5 = 0.6 and S_6 = 0.5. The output of the corrected correlation matrix was modified to retain only strong correlations with consistent fiber orientation (e.g., radial compression vs. thickness change), and the correlation values ​​varied smoothly within the anatomical neighborhood.

[0092] Step a3: normalize the corrected spatiotemporal correlation strength of each channel to obtain the normalized spatiotemporal correlation strength of each channel.

[0093] Among them, the normalization process is to scale the corrected correlation strength to the interval [0,1] to ensure the comparability of weights.

[0094] For example, the correlation matrix after the correction of the input data is normalized by the time channel, and Softmax processing is performed on all spatial channel correlation values ​​of each time channel so that the sum is 1, and the normalized correlation matrix is ​​output.

[0095] Step a4: Generate channel-level dynamic feature weights based on the normalized spatiotemporal correlation strength of each channel and the preset viscoelastic property constraints of the articular cartilage.

[0096] The viscoelastic property constraint requires the stress-strain relationship of cartilage to conform to a quasi-linear viscoelastic model. The dynamic feature weight is the final channel-level weight generated and used to weightedly fuse spatiotemporal features.

[0097] For example, the normalized correlation matrix for time channel T_15 is: S_5 (thickness) is 0.55, and S_8 (curvature) is 0.45. For time channel T_20, S_5 (thickness) is 0.38, and S_8 (curvature) is 0.62. Viscoelastic constraint parameters: High-frequency channels (such as T_15, corresponding to rapid compression) are enhanced (weight × 1.2); low-frequency channels (such as T_1, corresponding to the stationary period) are weakened (weight × 0.8). The dynamic weight generation process multiplies the normalized value of each time channel by the frequency response coefficient: Adjusted weight = normalized value × (1 + 0.2 * sin(2πft)), where f is the channel center frequency and t is time. For T_15 - S_5: 0.55 × 1.2 = 0.66 (high-frequency enhancement); T_1 - S_5: 0.10 × 0.8 = 0.08 (low-frequency suppression). After modulation, the output channel pair T_15-S_5 has an adjusted weight of 0.63, and the channel pair T_15-S_8 has an adjusted weight of 0.37. Taking the average of the spatial channel weights, we obtain the temporal channel-level weights W_T15 = (0.63 + 0.37) / 2 = 0.50 and W_T20 = (0.30 + 0.70) / 2 = 0.50. The final weight vector is [0.50, 0.50, ...], which is a 32-dimensional vector.

[0098] By executing steps a1-a4, this embodiment of the present application generates dynamic feature weights that are consistent with the biomechanical properties of cartilage through cross-modal correlation analysis, directional and continuity constraint correction, normalization, and viscoelasticity adjustment. This embodiment of the present application demonstrates that it can precisely enhance key spatiotemporal features (such as the mid-flexion contact zone), improve segmentation accuracy, and ensure anatomical plausibility, making it suitable for early diagnosis and surgical planning of osteoarthritis.

[0099] In one possible embodiment, step a2, based on a preset articular cartilage principal stress direction consistency constraint and a preset articular cartilage deformation boundary continuity constraint, performs a directional correction on the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix, including:

[0100] Step b1: extract the temporal channel features and spatial distribution channel features of each channel from the cross-modal correlation matrix.

[0101] Feature extraction involves extracting temporal features from each row (temporal channel) of the correlation matrix: a fast Fourier transform is used to obtain the dominant frequency. Spatial features are extracted from each column (spatial channel): the thickness gradient direction is calculated. The temporal channel feature set is: {T_i: (dominant frequency, energy fraction)}; the spatial channel feature set is: {S_j: (gradient direction, edge sharpness)}.

[0102] Step b2: Calculate the angle between the principal stress direction of the time channel characteristic and the deformation boundary direction of the spatial distribution channel characteristic of each channel to generate a directional deviation.

[0103] The deformation boundary direction is the direction perpendicular to the thickness gradient in the spatial channel feature (e.g., gradient direction = 30°, boundary direction = 120°). The directional deviation is the cosine of the angle between the two (0–1), with 0 indicating complete deviation and 1 indicating complete agreement. The angle is calculated for each channel pair (T_i, S_j) using the formula: deviation = 1 - |cos(principal stress direction - deformation boundary direction)|. If the principal stress direction is 0°, the boundary direction is 120°, and the deviation = 1 - |cos(120°)| = 0.5. The generated deviation matrix, which has the same size as the correlation matrix (32×32), stores the deviation values ​​for each channel pair.

[0104] Step b3: Generate a directional deviation correction weight based on the directional deviation and the preset viscoelastic property constraint of the articular cartilage.

[0105] The directional bias correction weight is an attenuation coefficient (0-1) generated based on the bias and viscoelastic parameters, which is used to suppress non-physiological associations.

[0106] For example, the weight calculation process is as follows: if the deviation is > 0.3 (i.e., the angle is > 72°) and the channel frequency is > 1 Hz (high-frequency distortion), the weight is 0.2; if the deviation is < 0.1 and the frequency is < 0.5 Hz, the weight is 1.0 (preserve). The output is a modified weight matrix: the same size as the deviation matrix, marking the channel pairs to be attenuated.

[0107] Step b4: Based on the directional deviation correction weight, the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix is ​​iteratively corrected until the spatial distribution relationship between the principal stress direction and the deformation boundary direction meets the preset viscoelastic property constraint conditions of the articular cartilage, thereby obtaining the corrected spatiotemporal correlation strength of each channel.

[0108] Among them, the iterative correction is to adjust the correlation strength multiple times so that the principal stress-boundary direction relationship gradually satisfies the viscoelastic constraint.

[0109] The following is a specific example: In knee cartilage analysis, the input data has an original correlation matrix of T_15-S_5 = 0.6 and T_15-S_10 = 0.7 (transverse shear channel). The directional features are T_15 dominant frequency = 2 Hz (high frequency), S_5 boundary orientation = 30° (aligned with the principal stress), and S_10 = 90° (perpendicular). The calculated deviations are: T_15-S_5 deviation = 0, T_15-S_10 deviation = 1. The generated correction weights are: T_15-S_10 weight = 0.2 (high frequency and large deviation). After correction, T_15-S_10 = 0.7 × 0.2 = 0.14. Iteration is performed until all deviations are < 0.1. The output is a corrected correlation matrix with T_15-S_5 = 0.6 (preserved) and T_15-S_10 = 0.14 (suppressed).

[0110] By executing steps b1-b4, this embodiment of the present application quantifies directional deviations and couples viscoelastic constraints to iteratively modify the cross-modal correlation matrix, ensuring that dynamic weights only enhance physiologically plausible feature interactions (such as radial stress-thickness coupling). This embodiment of the present application demonstrates that interference from non-physiological correlations can be reduced, improving the reliability of cartilage lesion detection.

[0111] In a possible embodiment, step 143, performing channel-by-channel weighted fusion of the temporal channel features in the temporal channel feature set and the spatial distribution channel features in the spatial distribution channel feature set according to the dynamic feature weights to generate a multimodal fusion feature map, includes:

[0112] Step c1: Based on the mechanical properties of the stress concentration area of ​​the articular cartilage, the dynamic feature weight is modified.

[0113] Among them, the stress concentration area of ​​articular cartilage is a mechanically sensitive area (such as the center of the femoral condyle contact surface) determined by finite element analysis or strain energy density diagram. Mechanical properties include: quasi-linear viscoelastic parameters, strain rate sensitivity, etc. Dynamic feature weight correction is to adjust the weight value according to the mechanical properties to enhance the feature contribution of the stress concentration area. Optionally, if the difference between the time fluctuation amplitude of the time channel feature and the stress gradient amplitude of the spatial distribution channel feature exceeds a preset ratio, the dynamic feature weight is reduced according to the viscoelastic attenuation coefficient; if the principal stress direction of the time channel feature is consistent with the deformation boundary direction of the spatial distribution channel feature, the dynamic feature weight is increased according to the dynamic coupling coefficient.

[0114] For example, finite element simulation results are used to generate a binary mask (1 = stress concentration area, such as areas with pressure > 2 MPa). For spatial channels located in stress concentration areas (such as S_5), the weight is multiplied by 1.5; for high-frequency temporal channels (such as T_15 with f > 1 Hz), the weight is multiplied by 1.2. The corrected weights are output: the original weight W_S5 = 0.6, and the corrected weight is 0.6 × 1.5 = 0.9.

[0115] Step c2: Based on the corrected dynamic feature weights, the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set are nonlinearly superimposed channel by channel to generate a multimodal fusion feature map.

[0116] Among them, nonlinear superposition uses element-level nonlinear operations (such as power fusion) instead of simple weighted summation to capture the complex interactions between features.

[0117] The following is a specific example: During knee cartilage segmentation, the input dynamic feature weights are: W_T15 = 0.9 (high-frequency flexion period), W_S5 = 0.9 (femoral condyle thickness); the stress concentration mask is the femoral condyle contact area (a 50×50 pixel area with a value of 1). The modified dynamic feature weights are: W_S5 = 0.9 × 1.5 = 1.35 (truncated to 1.0), W_T15 = 0.9 × 1.2 = 1.08. The dynamic feature weights are adjusted in the stress concentration area, and then channel-by-channel nonlinear superposition is performed to generate a multimodal fusion feature map.

[0118] By executing steps c1 and c2, this embodiment enhances the feature map's ability to represent cartilage stress concentration areas through mechanical property-driven weight correction and nonlinear fusion. This embodiment demonstrates that the fused features enable the segmentation network to achieve a 95% sensitivity to early micro-damage (e.g., cracks <1 mm) and a mechanical rationality violation rate of less than 3%, providing clinically accurate analysis results that adhere to biomechanical principles.

[0119] In a possible embodiment, S12, identifying spatiotemporal correlation features between piezoelectric mechanical data and polarized light stress data, includes:

[0120] Step 121 : Perform spatiotemporal registration of the piezoelectric mechanical data and the polarized light stress data along the grid cells on the surface of the articular cartilage to generate spatiotemporally synchronized piezoelectric mechanical data and polarized light stress data.

[0121] The articular cartilage surface mesh elements are triangular / quadrilateral meshes generated through 3D reconstruction and used for spatial discretization analysis. Piezoelectric data are voltage time series (sampling rate 1kHz) at each mesh vertex, reflecting the piezoelectric response under dynamic load. Spatiotemporal registration unifies the two types of data into the same spatiotemporal coordinate system, which can be based on mesh position, timestamps, and other factors.

[0122] Step 122: For each grid cell on the surface of the articular cartilage, extract the spatiotemporal dynamic fluctuation characteristics of the piezoelectric force and the directional distribution characteristics of the polarized light stress from the combined data of the piezoelectric force and the polarized light stress.

[0123] Among them, piezoelectric feature extraction is to perform wavelet transform on the joint data of each grid unit and calculate the energy proportion of 1~5Hz. For example, the energy proportion of the central unit of the femoral condyle is 85%, indicating that the high-frequency load is concentrated. Polarization light feature extraction is to calculate the distribution parameters of the joint data of the 5×5 grid units in the neighborhood. For example, the degree of dispersion of the direction distribution of the polarization light stress direction in a certain area in the normal area is σ²<0.1, and σ²>0.3 in the diseased area, indicating that the direction is disordered. Then the following feature set is output: each grid unit corresponds to a feature vector. σ² can represent the variance of the directional distribution of the polarization light stress direction, which represents the variance of the directional angles of all pixel points.

[0124] Step 123: A dynamic mode decomposition algorithm is used to extract the coupled modes of the spatiotemporal dynamic fluctuation characteristics and the polarized light stress direction distribution characteristics in the frequency domain, and a spatiotemporal correlation feature is generated according to the energy proportion of the coupled modes.

[0125] Among them, the dynamic modal decomposition algorithm includes the singular value decomposition method. The formula used in this singular value decomposition method is an existing formula and will not be repeated here. Dynamic modal decomposition is an algorithm for extracting dominant vibration modes from spatiotemporal data. Each mode contains frequency f, attenuation rate λ and spatial mode Φ. Coupled modes are components in the mode that simultaneously contain piezoelectric fluctuations and stress directions that change synergistically. Energy share is the contribution of the mode to the total energy of the system and is used to screen key modes. The process of dynamic mode decomposition (DMD) includes the singular value decomposition process, constructing a low-dimensional projection matrix, solving eigenvalues ​​and modes, reconstructing DMD modes, and finally calculating the energy share of coupled modes.

[0126] In the specific application of articular cartilage analysis, the input data is exemplified by the piezoelectric-polarized light joint data matrix X, which is 10,000×100 in size. Each row represents the characteristics of a grid cell (e.g., [wavelet energy, mean angle θ_mean of the polarized light stress direction, variance σ² of the polarized light stress direction distribution]), and each column represents the full-field data at a time point. Dynamic modal decomposition outputs the following information: the first dominant mode has a frequency f1 = 2 Hz, corresponding to the loading phase of the gait cycle; the third dominant mode has a frequency f3 = 5 Hz, possibly indicating abnormal vibration caused by pathology. The spatial mode of the first mode has a high amplitude in the contact area of ​​the femoral condyle (strong mechanical-structural coupling); the spatial mode of the third mode has a sudden increase in amplitude at the edge (collagen network decoupling). The first mode contributes 18% of the energy, dominating the normal loading response; the third mode contributes 12% of the energy, possibly indicating an early stage of pathology.

[0127] Here's a specific example:

[0128] In dynamic loading experiments on the knee joint, the input data included piezoelectric electromechanical data (specifically, a 10,000-cell grid with 100 time points and a sampling rate of 1kHz); and polarized optical data (stress orientation maps from OCT scans at 10Hz and 100 frames). After registration, the piezoelectric signal amplitude of the femoral condyle contact area mesh (cell IDs: 5000-6000) increased threefold. The characteristics of the 5012th cell included wavelet energy = 0.82, θ_mean = 15°, and σ² = 0.05, indicating a healthy condition; the characteristics of the 5803rd cell included wavelet energy = 0.35, θ_mean = 60°, and σ² = 0.28, indicating a diseased condition. DMD extracted three key modes (accounting for 12% to 18% of the energy), showing that the first mode, with a frequency of 2Hz, showed strong coupling at the center of the femoral condyle; the third mode, with a frequency of 5Hz, showed decoupling at the edge, a sign of disease. The output is a spatiotemporal correlation feature, a 10,000×3 matrix containing spatial patterns across three modes. During lesion detection, the Φ value for the 5803rd element in the third mode was found to be greater than 0.8 (an abnormality), consistent with the actual results. The Φ value is the spatial mode amplitude extracted through dynamic modal decomposition and is used to quantify the spatial response strength of articular cartilage under specific mechanical modes.

[0129] By executing steps 121-123, this embodiment of the present application achieves deep coupling of piezoelectric and polarized light data through high-precision spatiotemporal registration, multimodal feature decoupling, and DMD modal analysis. This embodiment demonstrates that the generated spatiotemporal correlation features can quantify the mechanical-structural synergy of cartilage, with a sensitivity of 93% for detecting early lesions and the ability to locate the initial location of microdamage (e.g., the decoupling zone at the edge of the femoral condyle), providing clinical analysis tools that combine spatiotemporal resolution with mechanical interpretability.

[0130] In a possible embodiment, S13, performing feature enhancement on the spatiotemporal correlation feature to obtain the enhanced spatiotemporal correlation feature, includes:

[0131] Step 131: Perform wavelet packet transformation on the spatiotemporal correlation features to obtain high-frequency stress fluctuation components and low-frequency viscoelastic relaxation components.

[0132] The wavelet packet transform (WPT) is a time-frequency analysis method that adaptively divides frequency bands, providing finer frequency resolution than traditional wavelet transforms. High-frequency stress fluctuation components correspond to transient mechanical responses with frequencies >1 Hz (such as rapid deformation caused by gait impact). Low-frequency viscoelastic relaxation components correspond to slow recovery processes with frequencies <0.5 Hz (such as cartilage creep).

[0133] For example, wavelet packet decomposition uses a wavelet basis to perform a three-level decomposition of the spatiotemporal correlation features (a 10,000×3 matrix), yielding eight sub-bands. The high-frequency band (1–5 Hz) contains stress fluctuations, while the low-frequency band (0–0.5 Hz) contains viscoelastic relaxation. Component extraction involves reconstructing the high-frequency sub-band (nodes (3,1)–(3,4)) into high-frequency components, and the low-frequency sub-band (nodes (3,7)–(3,8)) into low-frequency components.

[0134] Step 132: Perform feature enhancement on the high-frequency stress fluctuation component to obtain a feature-enhanced high-frequency stress fluctuation component.

[0135] Feature enhancement uses signal processing or deep learning techniques to improve the specificity, signal-to-noise ratio, or discriminability of target features while suppressing irrelevant noise or interference. The goal is to highlight the transient mechanical response of cartilage under dynamic loads (such as stress wave propagation and micro-deformation) and mitigate instrument noise or motion artifacts. The enhanced high-frequency stress fluctuation component is a processed high-frequency signal (typically >1 Hz) that selectively enhances the mechanical response features (such as shock waves and vibration modes) contained within it, suppressing noise components. The data format maintains the same size as the original high-frequency component (e.g., 10,000 grid cells × time series), but with increased amplitude and contrast in key areas.

[0136] Step 133: The enhanced high-frequency stress fluctuation component and low-frequency viscoelastic relaxation component are fused through the jump connection in the U-shaped convolutional network architecture to obtain enhanced spatiotemporal correlation features.

[0137] The U-shaped convolutional network architecture is a symmetrical encoder-decoder structure designed specifically for image segmentation tasks. For example, it includes core functions for high-precision pixel-level segmentation, small-sample learning, and multimodal feature fusion. Its main modules include the encoder (downsampling path), decoder (upsampling path), skip connection layer, and output layer. Skip connections are one of the core modules in the U-shaped convolutional network architecture, used to fuse the high-resolution features of the encoder (downsampling path) with the semantic features of the decoder (upsampling path), thereby restoring spatial details and improving segmentation accuracy.

[0138] For example, in the biomechanical analysis of articular cartilage, the division between high-frequency stress fluctuation components and low-frequency viscoelastic relaxation components is based on the mechanical response characteristics of cartilage and the physiological loading frequency. Specifically, the high-frequency stress fluctuation component, ranging from 1 Hz to 10 Hz, corresponds to the transient response to dynamic loads, such as the rapid impact of walking and running (the cadence is typically 1-2 Hz, and can reach 3-5 Hz during running), reflecting the elastic deformation and stress wave propagation of the cartilage (rapid vibration of the collagen fiber network). The low-frequency viscoelastic relaxation component, ranging from 0 Hz to 0.5 Hz, corresponds to the viscoelastic creep and stress relaxation of cartilage (such as the progressive deformation during prolonged standing or slow flexion and extension), reflecting the energy dissipation and fluid flow effects (time-dependent behavior) of the proteoglycan matrix. Dynamic mechanical analysis (DMA) shows that the storage modulus (elastic response) of cartilage increases above 1 Hz, while the loss modulus (viscous response) dominates below 0.5 Hz. The human cadence is approximately 1-2 Hz, and can reach 3-5 Hz when running or jumping. Mechanical recovery at rest is typically on the order of 0.1 Hz. High frequencies (1-10 Hz) are associated with rapid mechanical responses to dynamic loads and are related to the elastic properties of cartilage. Low frequencies (0-0.5 Hz) are associated with viscoelastic relaxation processes and are related to the energy dissipation capacity of cartilage. This classification is based on a combination of physiological load frequency, material properties, and signal processing requirements to ensure physical interpretability of feature separation.

[0139] As another example, the input data consists of a high-frequency stress fluctuation component of size [H, W, C1] (e.g., 256×256×64), containing dynamic mechanical response details. A low-frequency viscoelastic relaxation component of size [H / 8, W / 8, C2] (e.g., 32×32×128) is generated by extracting global features from the deep layers of the encoder. The decoder then upsamples the low-frequency component and performs transposed convolution or bilinear interpolation, gradually restoring the resolution to [H, W, C2]. Skip connection fusion is achieved through the following process: Specifically, if the high- and low-frequency components have different channel counts, 1×1 convolution is used to resize the high- and low-frequency features. The high- and upsampled low-frequency features are then concatenated along the channel dimension, and the concatenated features are fused through convolutional layers. The high-frequency component in the input data is a 2Hz stress fluctuation component (256×256×64) in the femoral condyle contact zone, containing subtle deformation details. The low-frequency component is a global 0.2Hz viscoelastic relaxation feature (32×32×128), extracted by the encoder. The fusion process involves upsampling the low-frequency components three times to 256×256×128. The high-frequency components undergo a 1×1 convolution to adjust the number of channels to 64. This is then concatenated along the channel dimension to produce a 256×256×192 fused feature. The fused convolutional layer then outputs a 256×256×64 enhanced feature, preserving impact details while integrating creep trends. This embodiment of the present application can improve the detection rate of early cartilage lesions by 12%.

[0140] Here's a specific example:

[0141] In the dynamic analysis of knee cartilage, the input data's spatiotemporal correlation features consist of 10,000 grid cells × 3 modes. After wavelet packet decomposition, the high-frequency component shows 2Hz fluctuations at the center of the femoral condyle with an amplitude of 0.8, while the low-frequency component exhibits global 0.2Hz relaxation with an amplitude of 0.3. After high-frequency enhancement, the amplitude of stress concentration areas (e.g., cells 5000-6000) increases to 1.2, and background noise is reduced by 40%. The fused feature map output preserves high-frequency impact details (such as 2Hz fluctuations) while integrating low-frequency creep trends. The enhanced feature map output is a 10,000 × 64 matrix (64-channel fused features). This embodiment of the present application can rapidly detect cartilage lesions at an early stage with an accuracy of 92% (compared to 78% without enhancement), and the compliance rate with viscoelastic laws is 97%.

[0142] By executing steps 131-133, this embodiment of the present application achieves the coordinated optimization of mechanical dynamics and viscoelastic behavior in spatiotemporal correlation features through wavelet packet band separation, high-frequency adaptive enhancement, and U-Net multi-scale fusion. This embodiment demonstrates that the enhanced features increase the segmentation network's sensitivity to micro-deformations by 14%, while ensuring the physiological plausibility of the mechanical response (e.g., stress-strain phase delay conforming to the QLV model), providing clinically accurate and interpretable analysis results.

[0143] Figure 2 A structural diagram of an articular cartilage image segmentation system based on a hybrid attention mechanism is provided in an embodiment of the present application, as shown in FIG. Figure 2 As shown, the system includes:

[0144] The acquisition module 21 is used to acquire the piezoelectric mechanical data and polarized light stress data of the articular cartilage under dynamic load conditions.

[0145] The identification module 22 is used to identify the spatiotemporal correlation characteristics between the piezoelectric mechanical data and the polarized light stress data.

[0146] The enhancement module 23 is used to enhance the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features.

[0147] The output module 24 is used to input the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs segmentation results. The segmentation results include the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

[0148] Figure 2 The articular cartilage image segmentation system based on the hybrid attention mechanism can be performed Figure 1 The implementation principle and technical effects of the articular cartilage image segmentation method based on the hybrid attention mechanism described in the illustrated embodiment will not be elaborated on here. The specific manner in which each module and unit performs operations in the articular cartilage image segmentation system based on the hybrid attention mechanism in the above embodiment has been described in detail in the embodiments of the method and will not be elaborated on here.

[0149] In one possible design, Figure 2 The articular cartilage image segmentation system based on the hybrid attention mechanism of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .

[0150] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0151] The processing component 32 is used to: obtain piezoelectric mechanical data and polarized light stress data of articular cartilage under dynamic load conditions; identify the spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data; perform feature enhancement on the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features; input the enhanced spatiotemporal correlation features into a pre-trained segmentation network, the pre-trained segmentation network uses dynamic feature weights to perform weighted processing on features of different modalities, and outputs a segmentation result, which includes the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage, and the dynamic feature weights are generated by a hybrid attention mechanism.

[0152] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0153] The storage component 31 is configured to store various types of data to support operations on the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0154] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0155] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0156] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0157] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0158] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment is an articular cartilage image segmentation method based on a hybrid attention mechanism.

[0159] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0161] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for articular cartilage image segmentation based on a hybrid attention mechanism, characterized in that: include: Acquire piezoelectric and electromechanical data and polarized light stress data of articular cartilage under dynamic loading conditions; identifying spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data; Performing feature enhancement on the spatiotemporal correlation feature to obtain enhanced spatiotemporal correlation feature; The enhanced spatiotemporal correlation features are input into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs a segmentation result. The segmentation result includes the deformation area of ​​the articular cartilage to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

2. The method according to claim 1, characterized in that The enhanced spatiotemporal correlation features are input into a pre-trained segmentation network, and the pre-trained segmentation network uses dynamic feature weights to perform weighted processing on features of different modalities and output a segmentation result, including: Inputting the enhanced spatiotemporal correlation features into a pre-trained segmentation network, the segmentation network performs channel division on the enhanced spatiotemporal correlation features according to time series and spatial distribution to generate a time channel feature set and a spatial distribution channel feature set; Based on the temporal channel feature set and the spatial distribution channel feature set, a hybrid attention mechanism is used to calculate the spatiotemporal correlation strength of each channel to generate dynamic feature weights; According to the dynamic feature weights, the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set are weightedly fused channel by channel to generate a multimodal fusion feature map; Performing cross-scale feature aggregation on the multimodal fusion feature map to obtain an aggregation result; Based on the continuity constraint condition of the preset articular cartilage deformation boundary and the aggregation result, the deformation area is predicted to obtain a segmentation result of the deformation area containing the articular cartilage.

3. The method according to claim 2, characterized in that The method of calculating the spatiotemporal correlation strength of each channel based on the temporal channel feature set and the spatial distribution channel feature set using a hybrid attention mechanism to generate dynamic feature weights includes: Performing a multi-dimensional correlation analysis between the temporal channel feature set and the spatial distribution channel feature set to generate a cross-modal correlation matrix; Based on a preset articular cartilage principal stress direction consistency constraint and a preset articular cartilage deformation boundary continuity constraint, a directional correction is performed on the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix; Normalizing the corrected spatiotemporal correlation strength of each channel to obtain the normalized spatiotemporal correlation strength of each channel; According to the normalized spatiotemporal correlation strength of each channel and the preset viscoelastic property constraints of articular cartilage, the channel-level dynamic feature weight is generated.

4. The method according to claim 3, characterized in that The method of performing directional correction on the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix based on the preset articular cartilage principal stress direction consistency constraint condition and the preset articular cartilage deformation boundary continuity constraint condition includes: Extracting temporal channel features and spatial distribution channel features of each channel from the cross-modal correlation matrix; The angle between the principal stress direction of the time channel characteristic and the deformation boundary direction of the spatial distribution channel characteristic of each channel is calculated to generate the directional deviation; generating a directional deviation correction weight based on the directional deviation and the preset viscoelastic property constraint of the articular cartilage; According to the directional deviation correction weight, the spatiotemporal correlation strength of each channel in the cross-modal correlation matrix is ​​iteratively corrected until the spatial distribution relationship between the principal stress direction and the deformation boundary direction meets the preset viscoelastic property constraint of the articular cartilage, thereby obtaining the corrected spatiotemporal correlation strength of each channel.

5. The method according to claim 2, characterized in that The step of performing channel-by-channel weighted fusion of the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set according to the dynamic feature weights to generate a multimodal fusion feature map includes: Based on the mechanical properties of the stress concentration area of ​​articular cartilage, the dynamic feature weight is modified; According to the corrected dynamic feature weights, the time channel features in the time channel feature set and the spatial distribution channel features in the spatial distribution channel feature set are nonlinearly superimposed channel by channel to generate a multimodal fusion feature map.

6. The method according to claim 1, characterized in that The identifying the spatiotemporal correlation characteristics between the piezoelectric mechanical data and the polarized light stress data includes: performing spatiotemporal registration of the piezoelectric mechanical data and the polarized light stress data along the grid cells on the surface of the articular cartilage to generate spatiotemporally synchronized piezoelectric mechanical data and polarized light stress data; For each grid cell on the surface of the articular cartilage, extracting the spatiotemporal dynamic fluctuation characteristics of the piezoelectric force and the directional distribution characteristics of the polarized light stress from the combined piezoelectric force and polarized light stress data; A dynamic mode decomposition algorithm is used to extract the coupled modes of the spatiotemporal dynamic fluctuation characteristics and the polarized light stress direction distribution characteristics in the frequency domain, and the spatiotemporal correlation characteristics are generated according to the energy proportion of the coupled modes.

7. The method according to claim 1, characterized in that The step of enhancing the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features includes: Performing wavelet packet transformation on the spatiotemporal correlation characteristics to obtain a high-frequency stress fluctuation component and a low-frequency viscoelastic relaxation component; Performing feature enhancement on the high-frequency stress fluctuation component to obtain a feature-enhanced high-frequency stress fluctuation component; The enhanced high-frequency stress fluctuation component and the low-frequency viscoelastic relaxation component are fused through the jump connection in the U-shaped convolutional network architecture to obtain the enhanced spatiotemporal correlation feature.

8. An articular cartilage image segmentation system based on a hybrid attention mechanism, characterized in that: include: an acquisition module for acquiring piezoelectric mechanical data and polarized light stress data of articular cartilage under dynamic loading conditions; an identification module, configured to identify spatiotemporal correlation features between the piezoelectric mechanical data and the polarized light stress data; An enhancement module, configured to enhance the spatiotemporal correlation features to obtain enhanced spatiotemporal correlation features; An output module is used to input the enhanced spatiotemporal correlation features into a pre-trained segmentation network. The pre-trained segmentation network uses dynamic feature weights to weight features of different modalities and outputs segmentation results. The segmentation results include the deformation area of ​​the articular cartilage, so as to determine the level of articular cartilage damage based on the deformation area of ​​the articular cartilage. The dynamic feature weights are generated by a hybrid attention mechanism.

9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an articular cartilage image segmentation method based on a hybrid attention mechanism as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, an articular cartilage image segmentation method based on a hybrid attention mechanism as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Automatic segmentation method for articular cartilage tissue in three-dimensional medical image

    CN111354000A

  • Knee joint MRI image segmentation method based on fusion attention guidance

    CN118608535A

  • Method for generating a magnetic resonance synthetic computer tomography image

    WO2022008759A1