Deep learning based method and system for identifying bone erosion in rheumatoid arthritis x-ray
By constructing a differentiable entropy-weighted topological deep learning network, the problems of texture vanity and noise sensitivity in early bone erosion identification in X-ray images of rheumatoid arthritis were solved, achieving high-precision lesion identification and robustness of the diagnostic system, and improving the efficiency of early diagnosis of rheumatoid arthritis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
Existing deep learning-based technologies for identifying bone erosion in rheumatoid arthritis X-ray images suffer from texture loss and noise sensitivity when dealing with early, minute bone erosion. This leads to the loss of texture details and noise confounding during feature extraction, resulting in false positives.
A differentiable entropy weighted topological deep learning network is adopted to construct a simple complex set through frequency domain texture chaos distribution data and geometric grayscale manifold data. The sorting logic of simplex entering the filtering sequence is reconstructed using differentiable entropy weights to generate a continuous homology landscape feature tensor. A lesion probability mask is generated through a multi-scale feature fusion decoding network. Combined with the degradation robust control logic of statistical dispersion index, noise is suppressed and the lesion identification accuracy is improved.
It effectively overcomes the limitation of traditional CNNs in losing high-frequency texture details during downsampling, improves the sensitivity and specificity of early micro-bone erosion lesions, generates high-precision lesion probability masks, and ensures the safety and reliability of the auxiliary diagnostic system under complex working conditions.
Smart Images

Figure CN122367977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis technology, specifically to a method and system for identifying bone erosion in X-ray images of rheumatoid arthritis based on deep learning. Background Technology
[0002] With the deep integration of computer vision and medical imaging technologies, deep learning-based assisted diagnostic systems have been widely applied in the clinical diagnosis and treatment of rheumatoid arthritis (RA). In particular, for X-ray images, which are considered the "gold standard" for early RA diagnosis, the automated identification and quantitative assessment of bone erosion lesions using image analysis technology is of great significance for improving diagnostic efficiency and early intervention. However, existing image assessment technologies for osteoarthritis or bone erosion still face significant challenges. Taking the existing technology with publication number CN121121201A as an example, although it achieves multi-scale semantic modeling through nnU-Net and ResNet-18 networks, effectively improving the localization accuracy of intervertebral joint lesions, this type of standard convolutional neural network (CNN)-based approach has limitations in handling early, small bone erosion in RA: namely, the "feature binary paradox." On the one hand, early bone erosion is mainly manifested as high-frequency fracture of bone trabecular texture. However, the deep continuous downsampling operation of the CNN architecture tends to preserve global semantic features, resulting in the loss of high-frequency texture details and the formation of "texture blind spots". On the other hand, when attempting to introduce topological data analysis (TDA) technology that can capture hole features, due to its high sensitivity to geometric structures, the imaging noise commonly present in low-dose X-rays is often misjudged as tiny pathological holes, leading to "noise confusion" in the feature extraction process and thus causing false positives.
[0003] To address the technical contradiction between "texture transience" and "noise sensitivity," this invention proposes a deep learning-based "differentiable entropy-weighted topological fusion" recognition approach. This approach abandons the single-dimensional feature extraction model and instead constructs a collaborative decision-making mechanism: utilizing the degree of disorder (entropy) of texture energy in the frequency domain as prior knowledge to dynamically weight and intervene in the simplex construction process in geometric space. By transforming non-differentiable topological cohomology computation into differentiable neural network layers, it achieves soft suppression of noise signals and hard enhancement of pathological textures at the feature extraction source. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for identifying bone erosion in X-ray images of rheumatoid arthritis based on deep learning, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A deep learning-based method for identifying bone erosion in X-ray images of rheumatoid arthritis, specifically including:
[0007] S1: Obtain the frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the trabecular texture signal in the local area of the X-ray image to be analyzed.
[0008] S2: Obtain the geometric grayscale manifold data of the X-ray image to be analyzed, and construct the corresponding simple complex set based on the geometric grayscale manifold data;
[0009] S3: The frequency domain texture chaos distribution data and the simplex set are input into a differentiable entropy-weighted topological deep neural network layer; the differentiable entropy-weighted topological deep neural network layer is configured to perform weighted filtering homology calculations containing learnable parameters; the calculation includes: using the frequency domain texture chaos distribution data as dynamic weight factors, reconstructing the sorting logic of each simplex in the simplex set entering the filtering sequence, so as to generate a persistent homology landscape feature tensor as a persistent homology landscape indicator signal; the persistent homology landscape indicator signal characterizes the pathological bone structure loss features under the dual constraints of frequency domain energy disorder and geometric spatial connectivity;
[0010] S4: Input the persistent homogeneous landscape feature tensor into a multi-scale feature fusion decoding network; the decoding network is configured to perform cross-domain stitching and upsampling convolution of the persistent homogeneous landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area;
[0011] S5: Monitor the statistical dispersion index of the continuously coherent landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, trigger the degradation robust control logic and automatically generate an environmental adaptability degradation status code that characterizes the degraded operation of the system.
[0012] A deep learning-based system for identifying bone erosion in X-ray images of rheumatoid arthritis, the system being used to execute the aforementioned method for identifying bone erosion in X-ray images of rheumatoid arthritis, comprising:
[0013] The frequency domain feature extraction module is configured to acquire frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the bone trabecular texture signal in a local region of the X-ray image to be analyzed.
[0014] The topology manifold construction module is configured to acquire the geometric grayscale manifold data of the X-ray image to be analyzed, and construct a corresponding simple complex set based on the geometric grayscale manifold data;
[0015] The core module of entropy weight topology fusion includes a differentiable entropy weight topology deep neural network layer, which is configured to perform weighted filtering homology calculation with learnable parameters; the calculation includes: using the frequency domain texture chaos distribution data as a dynamic weight factor, reconstructing the sorting logic of each simplex in the simplex set entering the filtering sequence, so as to generate a continuous homology landscape feature tensor as a continuous homology landscape indicator signal.
[0016] The multi-scale feature fusion decoding module is configured to perform cross-domain stitching and upsampling convolution of the continuous homology landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area.
[0017] The robust adaptive control module is configured to monitor the statistical dispersion index of the continuously homogeneous landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, the degradation robust control logic is triggered to automatically generate an environmental adaptive degradation status code that characterizes the degraded operation of the system.
[0018] Compared with the prior art, the beneficial effects of the present invention are:
[0019] By constructing a differentiable entropy-weighted topological deep neural network layer (steps S1-S3), and using frequency domain texture chaos distribution data as dynamic weight factors, the sorting logic of the simplex entering the filtering sequence is reconstructed. This technique effectively overcomes the limitation of traditional CNNs in losing high-frequency texture details during downsampling. Simultaneously, it utilizes texture entropy properties to perform nonlinear filtering of background noise, suppressing false topological features (false positive holes) generated by traditional TDA technology due to its sensitivity to noise. This improves the system's sensitivity and specificity in identifying early, minute bone erosion lesions.
[0020] Through a multi-scale feature fusion decoding network (step S4), cross-domain stitching and upsampling convolution of continuous homology landscape feature tensors and spatial features are performed. This mechanism solves the problem that abstract topological features are difficult to regress to physical spatial coordinates, maps high-dimensional topological landscape signals back to two-dimensional image space, generates high-precision lesion probability masks, achieves pixel-level accurate locking of lesion regions, and improves the interpretability of imaging assessment.
[0021] By introducing degradation-robust control logic based on statistical dispersion (step S5), a degraded operation mode can be automatically triggered when the statistical dispersion falls below the confidence threshold due to extremely poor input image quality or unclear features (including severe artifacts and low contrast). This mechanism avoids forced inference by the deep learning model under invalid input, prevents the output of erroneous diagnostic information, and ensures the operational safety and reliability of the auxiliary diagnostic system under complex clinical conditions. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the overall architecture of the deep learning-based method for identifying bone erosion in X-ray images of rheumatoid arthritis according to the present invention.
[0023] Figure 2 This is a schematic diagram of the continuous homology landscape feature tensor generation technology route of the present invention;
[0024] Figure 3 This is a schematic diagram illustrating the numerical verification of the spatial separation degree of lesion features based on entropy-weighted topological fusion.
[0025] Figure 4 Numerical verification diagrams of the system logic response under different environmental disturbances;
[0026] Figure 5 This is a schematic diagram of the calculation route from step S4 to step S5. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0028] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but unless otherwise stated, these elements are not limited by these terms. These terms are used only to distinguish one element from another.
[0029] Example 1:
[0030] Please see Figures 1 to 5 The present invention provides a technical solution:
[0031] A deep learning-based method for identifying bone erosion in X-ray images of rheumatoid arthritis, the method being executed by an electronic computing device, includes the following steps:
[0032] S1: Obtain the frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the trabecular texture signal in the local area of the X-ray image to be analyzed.
[0033] S2: Obtain the geometric grayscale manifold data of the X-ray image to be analyzed, and construct the corresponding simple complex set based on the geometric grayscale manifold data;
[0034] S3: The frequency domain texture chaos distribution data and the simplex set are input into a differentiable entropy-weighted topological deep neural network layer; the differentiable entropy-weighted topological deep neural network layer is configured to perform weighted filtering homology calculations containing learnable parameters; the calculation includes: using the frequency domain texture chaos distribution data as dynamic weight factors, reconstructing the sorting logic of each simplex in the simplex set entering the filtering sequence, so as to generate a persistent homology landscape feature tensor as a persistent homology landscape indicator signal; the persistent homology landscape indicator signal characterizes the pathological bone structure loss features under the dual constraints of frequency domain energy disorder and geometric spatial connectivity;
[0035] S4: Input the persistent homogeneous landscape feature tensor into a multi-scale feature fusion decoding network; the decoding network is configured to perform cross-domain stitching and upsampling convolution of the persistent homogeneous landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area;
[0036] S5: Monitor the statistical dispersion index of the continuously coherent landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, trigger the degradation robust control logic and automatically generate an environmental adaptability degradation status code that characterizes the degraded operation of the system.
[0037] Further specifying, in step S1, the frequency domain texture chaos distribution data of the X-ray image to be analyzed is obtained, specifically including:
[0038] Read the scale parameter set and orientation parameter set defined in the external configuration file;
[0039] A multi-scale Gabor filter bank is constructed based on the scale parameter set and the orientation parameter set; the X-ray image to be analyzed is edge-expanded using the reflection filling algorithm, and the multi-scale Gabor filter bank is used to perform a convolution operation on the expanded image to generate a multi-channel frequency domain response map aligned with the original image size;
[0040] For each pixel position in the multi-channel frequency domain response map, a local sliding window is defined, and a logarithmic probability weighted summation operation is performed based on the energy distribution probability within the local sliding window to generate a frequency domain structure entropy exponent matrix as the frequency domain texture chaos distribution data.
[0041] Further specifying, in step S2, constructing the corresponding simple complex set based on the geometric grayscale manifold data specifically includes:
[0042] A spatial-grayscale anisotropic scaling factor is introduced; the pixel coordinates of the X-ray image to be analyzed are mapped to planar positions, and the pixel grayscale values are multiplied by the spatial-grayscale anisotropic scaling factor and mapped to elevation, thereby constructing a three-dimensional point cloud space;
[0043] Define an increasing distance threshold sequence and set a maximum topological dimension constraint parameter; for each threshold step in the distance threshold sequence, construct a simplex (based on Vietoris-Rips construction logic) from the set of points in the three-dimensional point cloud space where the Euclidean distance between each pair of points is less than the current threshold step; under the constraint of the maximum topological dimension constraint parameter, generate a set of simplex complexes, where any simplex with a dimension higher than the maximum topological dimension constraint parameter is discarded.
[0044] Further specifying, in step S3, the differentiable entropy-weighted topological deep neural network layer is configured to perform the following operations:
[0045] Perform cross-domain attribute mapping: calculate the geometric center of each simplex in the set of simplex complexes, and project the geometric center onto the frequency domain texture chaos distribution data to index the corresponding frequency domain texture entropy attribute value;
[0046] Perform weighted filtering calculation: Calculate the filtering time value by using a joint filtering function containing learnable parameters, combining the gray values of each simplex with the frequency domain texture entropy attribute value;
[0047] Perform differentiable soft sorting: Introduce a soft sorting temperature coefficient, calculate the pairwise difference probability based on the filtering time value and accumulate it to generate a continuous floating-point rank, thereby reconstructing the sorting logic of the simplex entering the filtering sequence;
[0048] Generating a landscape tensor: Based on the continuous floating-point rank, a persistent homology barcode is calculated, which consists of a series of birth-death intervals; using a preset multi-resolution variance Gaussian kernel function set, the birth-death intervals are mapped to a multi-channel weighted persistent landscape map, thus forming the persistent homology landscape feature tensor.
[0049] It should be further explained that: a continuous floating-point rank is used as a differentiable index to generate the corresponding soft permutation matrix, and this matrix is then used to perform weighted interpolation on the original filtered values. This step aims to construct a differentiable approximation of the continuously homologous barcode generation process, thereby allowing gradients to propagate back through the sorting operation and overcoming the problem of non-differentiability in traditional discrete sorting.
[0050] Further specifying the cross-domain attribute mapping, the process includes: determining the geometric center coordinates of each simplex in the set of simplexes in the 3D point cloud space; projecting the geometric center coordinates back to the pixel coordinate system of the X-ray image to be analyzed, and indexing the corresponding value in the frequency domain texture chaos distribution data; assigning the indexed value to the simplex as its corresponding frequency domain texture entropy attribute value, which is then used for subsequent weighted filtering calculations.
[0051] The following is a detailed implementation description of the above content: In the early diagnosis of rheumatoid arthritis, tiny bone erosion lesions are submerged in a complex background of trabecular bone texture. Existing technologies directly process grayscale information in the spatial domain or use texture analysis with fixed parameters, making it difficult to simultaneously capture the dual features of the lesion's "grayscale geometric concavity" and "bone texture fracture." Furthermore, traditional topological data analysis (TDA) methods are non-differentiable and cannot be embedded into deep learning networks for end-to-end optimization. To address these issues, this embodiment introduces a differentiable entropy-weighted topological processing mechanism, achieving a deep adaptive fusion of frequency domain texture features and geometric topological features. The key control parameters involved in this embodiment are defined as follows:
[0052] Frequency domain scale configuration set (denoted as ): This is a set of multiple positive integers used to define the scale range of the Gabor filter bank when extracting texture features. In this embodiment, this parameter is read from an external configuration file and set to [2, 4, 8]. The technical consideration for this value is to cover different texture frequency bands, ranging from minor trabecular fractures to larger osteoporotic regions.
[0053] Multi-scale Gabor kernels generate configuration sets, denoted as When acquiring the frequency domain texture chaos of the X-ray image to be analyzed, directly using a single-frequency filter cannot simultaneously capture the multi-scale fracture features exhibited by bone trabeculae at different erosion stages. This embodiment constructs a filter bank covering specific frequency bands and directions. By executing the construction process of a multi-scale Gabor kernel generation configuration set, this process is configured to generate a series of discretized kernel matrices for convolution. Specifically, a preset wavelength parameter list is obtained, which corresponds to the scale parameters [2,4,8] of the frequency domain scale configuration set; and a direction parameter list (corresponding to 8 directions). For each pair of wavelength and direction combinations, a two-dimensional complex sinusoidal carrier plane is generated. The frequency of this plane is determined by the reciprocal of the wavelength, and the wavefront normal direction is determined by the direction parameter. A two-dimensional Gaussian envelope function is generated, and the standard deviation parameter of this function is set to be linearly proportional to the current wavelength (in this embodiment, the scaling factor is set to 0.5) to control the spatial localization degree of the filter. A point-to-point multiplication operation is performed on the two-dimensional complex sinusoidal carrier plane and the two-dimensional Gaussian envelope function to generate a complex Gabor kernel. The real part of the complex Gabor kernel is extracted and then subjected to L2 norm normalization to ensure that all generated kernel matrices have a uniform energy response benchmark. This set of normalized matrices constitutes the multi-scale Gabor kernel generation configuration set.
[0054] Entropy calculation window (denoted as ): This is a two-dimensional size parameter defining the local region's extent, characterizing the observation field when calculating the frequency domain texture chaos degree. In this embodiment, this parameter is set to The unit is pixels. This value represents a balance between "statistical sample adequacy" and "lesion localization precision".
[0055] Number of topology scan steps (denoted as) : is a scalar parameter defining the accuracy of simple complex construction, representing the segmentation density of the distance threshold sequence. In this embodiment, it is set to 50.
[0056] Define the texture entropy weight coefficient (denoted as ). ) and gray intensity weighting coefficient (denoted as : These are a pair of floating-point scalar parameters that can be automatically updated via the backpropagation algorithm. Physically, they represent the network model's weights for "texture disorder" and "geometric concavity," respectively. Initially, they are randomly initialized to non-zero small values.
[0057] Landscape resolution parameters (denoted as) ): This parameter defines the size of the output feature tensor space, used to discretize the abstract topological barcode into a tensor that can be processed by a neural network. In this embodiment, it is set to... .
[0058] 1) For step S1, the generation of frequency domain texture chaos distribution data is performed: specifically, a data loading operation is performed, reading the X-ray image data to be analyzed and the above configuration parameters from the external storage medium. The multi-scale frequency domain feature extraction sub-process is initiated. This process constructs a set of Gabor filters based on the frequency domain scale configuration set. For the input image, convolution operations are performed using this set of filters. The logic of the convolution operation is as follows: before performing the convolution operation, a boundary padding preprocessing step is performed. The size of the input X-ray image and the size parameters of the current Gabor kernel are obtained. The required padding width for the current Gabor kernel is calculated, which is equal to half the kernel size (rounded down). Using the reflection-padding algorithm, the edges of the input image are mirrored and expanded based on the calculated padding width to generate an expanded image matrix. This eliminates convolution edge effects and maintains the spatial continuity of frequency domain features. The expanded image matrix is then convolved with each Gabor kernel. During the computation, a "Same" mode is enforced (representing that only the central region corresponding to the original input size is retained), thereby generating a multi-channel frequency domain response map aligned in height and width with the original X-ray image to be analyzed. The multi-channel frequency domain response map physically separates texture components of different scales and orientations. Next, the structural entropy index calculation sub-process is executed. For each pixel coordinate position in the multi-channel frequency domain response map, data within its local neighborhood bounded by the entropy calculation window is extracted. This local neighborhood data is normalized to an energy distribution probability. Then, the Shannon entropy logic operation is performed: each energy distribution probability is multiplied by its base-2 logarithmic value, all product results are summed and their negatives are taken. This result is filled into a new matrix of the same size as the original image, generating the frequency domain structural entropy index (FSEI) matrix. This frequency domain structural entropy index matrix constitutes "frequency domain texture chaos distribution data," and its value directly maps the degree of fracture or disorder in the trabecular microstructure.
[0059] Specifically, in the digital signal processing unit of an electronic device, each kernel function of a multi-scale Gabor filter bank All are preferably constructed based on the following physical models:
[0060] Where (x, y): physically defined as the discrete pixel coordinates of the filter kernel in two-dimensional space. In this embodiment, the coordinate range is determined by the kernel size. i9: physically defined as the imaginary unit, satisfying... It is used to construct sinusoidal carrier signals in the complex domain. The physical definition is the spatial coordinates after rotational transformation, and the specific calculation formula is as follows: as well as . Physically defined as the wavelength parameter of a sinusoidal carrier wave, corresponding to a frequency domain scale configuration set. In this embodiment, The preferred set of values is The pixel units correspond to high-frequency, mid-frequency, and low-frequency texture features, respectively. Physically defined as the orientation angle parameter of the Gabor core, corresponding to the set of orientation parameters described in the original text. In this embodiment, it is preferable to use... to The interval is divided into 8 equal directions. In this embodiment, the step size is... . Physically defined as the standard deviation of the Gaussian envelope, it is used to control the degree of spatial localization of the filter. In this embodiment, it is preferred to set... This ensures that the ratio of envelope to wavelength remains constant across different scales. The physical definition is the spatial aspect ratio, and the preferred value in this embodiment is 0.5, in order to fit the elongated texture features of the trabecular bone. Physically defined as phase offset, this embodiment preferably uses 0. Further, when generating the multi-scale Gabor kernel generation configuration set, for each generated complex kernel matrix, the main control loop only extracts its real part and performs L2 norm normalization. The physical significance of this normalization is to unify the energy response benchmark of filters at different scales, preventing large-scale filters from generating nonlinear energy gains due to their large kernel area. It should be noted that although this embodiment preferably uses the specific Gabor formula form described above, those skilled in the art can also use Log-Gabor filters or DoG (Difference of Gaussians) operators as alternatives to achieve multi-scale texture extraction; these transformations are all within the scope of this application. This method can effectively filter out unstructured background noise caused by soft tissue while preserving the directional texture of bone trabeculae in X-ray films.
[0061] The frequency domain structure entropy index is denoted as FSEI; it is determined as follows: the input data is a multi-channel frequency domain response map after Gabor convolution and nonlinear activation. A preset size is cropped centered on the current pixel position to be calculated. The local data block is defined. All values within this local data block are summed to obtain the total energy value of the block. Each value within the local data block is divided by the total energy value of the block to obtain the energy distribution probability value at that location, ensuring that the sum of all probability values equals 1. For each energy distribution probability value greater than zero, a base-2 logarithmic operation is performed to obtain a logarithmic probability value. Each energy distribution probability value is multiplied by its corresponding logarithmic probability value to obtain an intermediate entropy component. All intermediate entropy components are summed, and the sum is multiplied by -1 to obtain the final frequency domain structure entropy index. This calculation logic ensures that the frequency domain structure entropy index reflects the "disorderliness" of the texture energy distribution, and is independent of the absolute brightness of the image. In this embodiment, the computational environment executing step S1 is configured with a digital signal processing unit or a deep learning computational framework. The convolution operator in the computational environment is configured to support the following computational logic:
[0062] Complex number operation support: Supports direct complex domain convolution, or performs parallel two-channel convolution by treating the real and imaginary parts as independent feature channels to simulate the amplitude and phase transformations in the complex domain;
[0063] Spatial dimension alignment: The Same fill mode is used to ensure that the output frequency domain response map has the same spatial resolution as the original input image, thereby ensuring that the subsequently generated frequency domain entropy attribute value can accurately correspond to the pixel coordinates of the original image.
[0064] Specifically, for each pixel (u,v), the processor performs the following discretization mathematical operations to generate frequency domain texture chaos distribution data: defining a local sliding window. Its coverage area is centered at (u,v) Pixel region, corresponding to the entropy calculation window area Preferred .
[0065] Calculate the total energy value of the block within this window. With the normalized probability distribution p(i,j): When calculating the normalized probability distribution, the processor introduces a pre-set Laplace smoothing term. (Preferred) To avoid numerical calculation divergence in low-texture background areas: ; ;
[0066] Calculate the frequency domain structure entropy exponent FSEI at this location based on the Shannon entropy definition: Where R(i,j) is physically defined as the response amplitude of the multi-channel frequency domain response map at coordinate (i,j). If multiple channels exist, the maximum or average value of the multi-channel amplitude is preferred. : Indicates absolute value or modulo operation. Physically defined as a numerical stability constant, used to prevent the logarithm from becoming negative infinity. In this embodiment, the preferred value is... FSEI(u,v): Physically defined as the scalar value of the frequency domain structural entropy exponent matrix at coordinates (u,v), this value is normalized to the [0,1] interval and stored in memory. Through the entropy calculation based on the local probability distribution, the image is mapped from a simple grayscale domain to a "disorder" domain, thus allowing early bone erosion lesions that are difficult to distinguish in grayscale but exhibit slight disorder in texture arrangement to be highlighted.
[0067] 2) Step S2, Construction of the Geometric Gray-Level Manifold and Simplex Set: Simultaneously or after completing frequency domain processing, a geometric manifold mapping operation is performed. The X-ray image to be analyzed is treated as a function surface, and a three-dimensional point cloud space is established. Specifically, each pixel of the image is traversed, its planar horizontal and vertical coordinates are extracted as the first two dimensions of the point cloud, and its pixel gray-level value is extracted as the third dimension. A simplex is constructed based on the Vietoris-Rips algorithm logic. The externally configured topological scan step size is read, generating a distance threshold sequence increasing from zero to the maximum gray-level difference. For each threshold in the sequence, all point pairs in the three-dimensional point cloud space are traversed. If the Euclidean distance between two points is less than the current threshold, a connecting edge (1-simplex) is established; if three points are connected pairwise, a triangle (2-simplex) is established. This series of generated point, edge, and face sets constitutes the simplex sequence. This process captures the porosity features of bone erosion areas corresponding to different gray-level cutting layers of the skeletal structure.
[0068] Furthermore, in the process of constructing the simplex set, a "maximum topological dimension constraint parameter" is introduced to constrain computational complexity. In this embodiment, the maximum topological dimension constraint parameter is set to an integer value of 2. The specific construction logic is as follows: 0-simplexes (points) and 1-simplexes (edges) are formed by connecting point pairs based on a distance threshold; a 2-simplex (triangle) is constructed only when there are three interconnected vertices; any candidate high-dimensional simplexes with dimensions higher than the parameter setting are forcibly discarded (in this embodiment, tetrahedrons are included). The technical consideration for setting this value (dimension = 2) is that, for the bone erosion identification task, the core pathological features are mainly reflected in "connectivity disruption" (characterized by 0-dimensional homology features) and "closure of trabecular bone cavities" (characterized by 1-dimensional homology features), while higher-dimensional cavity features lack clear clinical interpretability in the projection data of two-dimensional X-ray films, and introducing high-dimensional computation would cause unnecessary resource redundancy.
[0069] Spatial-grayscale anisotropy scaling factor, denoted as In step S2, when constructing the simple complex, the planar coordinates (pixel units, range 0-512) and grayscale values (intensity units, range 0-255 or 0-65535) of the X-ray image belong to completely different physical dimensions. If a point cloud is directly constructed, the changes in the grayscale dimension may be masked by the huge span of the spatial dimension, or vice versa, leading to the failure of topological feature extraction. This embodiment performs a spatial-grayscale anisotropic scaling operation to obtain the width W, height H, and maximum possible number of grayscale levels of the X-ray image to be analyzed. Perform grayscale normalization by dividing the grayscale value of all pixels by 0. This maps the data to the [0,1] interval. A configurable hyperparameter, "Z-axis enhancement coefficient characterized by spatial-grayscale anisotropy scaling factor," is introduced and set to 2.0 in this embodiment. The normalized grayscale value is multiplied by this Z-axis enhancement coefficient. The technical consideration for setting the Z-axis enhancement coefficient to 2.0 is that, in topological data analysis, artificially stretching the "elevation" (grayscale) axis can amplify the geometric saliency of bone erosion areas (grayscale depressions) in the point cloud space, allowing the Vietoris-Rips algorithm to capture clinically significant pore features at a relatively small distance threshold.
[0070] In a preferred embodiment, the processor transforms the two-dimensional image pixels using the following coordinate transformation formula. Mapped to a set of point clouds in a three-dimensional metric space The processor performs the following isotropic normalization operation to preserve the geometric topological proportions of the original anatomical structure during normalization, preventing simple complex distortion caused by differences in image aspect ratio: determining the maximum spatial dimension of the image. The three-dimensional point cloud coordinates are generated using this uniform scaling factor. : This embodiment adopts... As a unified denominator, regardless of the size ratio of the original X-ray, the unit distance in the feature space always corresponds to a pixel distance of equal length in the physical space. This ensures that the hole features extracted by the Vietoris-Rips algorithm based on Euclidean distance truly reflect the circular or irregular shape of the bone, rather than geometric artifacts caused by image cropping.
[0071] in : Physically defined as the three-dimensional coordinate vector of the k-th point in the three-dimensional point cloud space. W, H: Defined as the physical width and height (in pixels) of the X-ray image to be analyzed. Physical definition: The first in the original image The discrete coordinates of each pixel in a two-dimensional image coordinate system. Defined as the original image in The gray intensity value at that location. Defined as the maximum gray level determined by the image bit depth (65535 for 16-bit DICOM images). : Defined as the spatial-grayscale anisotropic scaling factor (in this embodiment, it is the Z-axis enhancement coefficient). In this embodiment, it is preferably set to 2.0, but it can work effectively in the range of 1.5 to 3.0.
[0072] Furthermore, based on the generated point cloud set The processor constructs the simplex based on Vietoris-Rips logic. For a given distance threshold th1 in the distance threshold sequence, the simplex... The necessary and sufficient condition for being included in the set is:
[0073] ;in, This represents the Euclidean distance. : Represents a simplex constructed at the current threshold. : Represents a simplex Topological dimensions ( As an edge, (It is a triangle). : Represents a simplex Any two vertices in the vertex set. Simultaneously, the processor executes filtering logic based on the maximum topological dimension limit parameter: retaining only the dimension. The simplex is obtained by actively discarding tetrahedral and higher-dimensional structures. This embodiment introduces an anisotropic scaling factor. The distribution of data on the elevation axis is artificially stretched, which magnifies the tiny gray-scale depressions caused by bone erosion into significant “deep pit” structures in the topological space, so that they can be captured by the topology algorithm at a relatively small distance threshold.
[0074] It should be further explained that the range of the distance threshold sequence in this embodiment is determined by the feature scale of the lesion to be detected and the normalization ratio of the point cloud space. Specifically, the maximum value of the distance threshold sequence is set to a scale that can cover the connectivity of local trabecular microstructures, and is no greater than 1 / 10 of the diagonal length of the image normalization space, so as to avoid connecting unrelated anatomical structures that are too far apart, thereby generating topological noise.
[0075] The step size of the distance threshold sequence is adjusted to ensure that the spatial-grayscale anisotropy scaling factor can be captured. The topological changes of the enhanced grayscale pit at different depth sections.
[0076] This embodiment is based on Normalization and The recommended distance threshold th1 range is [0, 0.5]. In the normalized space, a Euclidean distance of 0.5 is equivalent to half the image width. This is sufficient as an upper limit for local microstructure analysis. A preferred upper limit is 0.2 or 0.3; this corresponds to a normalized space span of 20%–30%, which is sufficient to capture medium-sized bone defects and strong grayscale texture variations. The sequence step count is 50 to 100 steps.
[0077] For step S3, perform the processing of the differentiable entropy-weighted topological deep neural network layer:
[0078] Construction and weighted calculation of the joint filtering function: specifically obtaining the texture entropy weight coefficients. With gray intensity weighting coefficient For each simplex S in the set of simplex complexes, obtain its corresponding original gray value. The frequency domain texture entropy attribute value obtained by mapping (denoted as...) ). Defined as the frequency domain texture entropy attribute value obtained by indexing the frequency domain texture chaos distribution data from the simplex S through a cross-domain attribute mapping step. and Performing multiplication yields geometric terms, and then... and The entropy term is obtained by performing a multiplication operation. The geometric term is subtracted from the entropy term, and a learnable bias parameter is added. Finally, the ReLU activation logic is applied. The result is the "filtering time value" of the simplex. This logic ensures that the simplex will only obtain a small filtering time value when the region simultaneously satisfies "low grayscale" (bone loss) and "high entropy" (texture breakage), thus appearing "earlier" in the topological sequence and highlighting lesion features.
[0079] In traditional topological computation, sorting operations are non-differentiable. To address this issue, this embodiment introduces a smoothing maximum approximation operator. When determining the order in which simplexes enter the filtering sequence, a weighted summation logic based on the Softmax function is used to approximate the position index of each simplex in the sequence. This makes the entire computation process more consistent with the weight parameters. and Maintaining continuous differentiability allows for lossless backpropagation of error gradients. Based on a defined filtering sequence, a persistently coherent barcode is computed. This barcode consists of a series of intervals, each recording the "birth time" and "death time" of a topological feature. A landscape mapping operation is performed to convert the unstructured barcode into a tensor that can be processed by a neural network. Specifically, a set of Gaussian kernel functions with different variance parameters is defined. For each interval in the barcode, it is mapped onto a two-dimensional grid defined by the landscape resolution parameter. The specific computation logic is as follows: multiple Gaussian distribution functions are superimposed with the midpoint of the interval as the center and the length of the interval as the amplitude. The mapping maps generated by Gaussian kernels with different variances are stacked along the channel dimension to generate a persistently coherent landscape feature tensor. This multi-Gaussian kernel mapping mechanism not only preserves the persistent information of topological features but also captures lesion features of varying saliency through multi-scale variance, constituting a high signal-to-noise ratio "persistently coherent landscape indicator signal."
[0080] Specifically, for the "differentiable entropy weighted topological deep neural network layer", the feature interval represented by each birth-death interval in the continuously homogeneous barcode is... The processor utilizes a set of multi-resolution variance Gaussian kernel functions to generate persistently homogeneous landscape feature tensors. Among them, the first Pixel values of each channel Calculation based on discretized tensor construction using Gaussian mixture models:
[0081] ;
[0082] The parameters are defined as follows: Defined as a persistently homogeneous landscape feature tensor. Defined as the number of channels of a tensor, its value is equal to the size of the multi-resolution kernel variance set (in this embodiment) ). Defined as the spatial resolution parameter of a landscape image. Preferably, the tensor size is 512, meaning the generated tensor size is... . Physically defined as discrete grid coordinates of a persistent landscape plane, with values ranging from [value range missing]. . Physical definition: the first in a continuously coherent barcode The birth time and death time of each topological feature. Defined as the first The projection center of each topological feature onto the landscape plane. Preferably, , (This is represented by the birth-to-death ratio on the horizontal axis and lifespan length on the vertical axis.) : Defined as the first in the multi-resolution kernel variance set The set contains the values of each element. In this embodiment, the set is preferably... (Unit: grid pixels), corresponding to 4 feature channels, used to capture topological features with different salience. Defined as the first The weighting coefficients for each feature are preferably selected in this embodiment. In other words, the longer the lifecycle of a feature, the stronger the response it produces in the feature map.
[0083] This embodiment uses a Gaussian kernel stacking with multi-scale variance to convert sparse, unstructured topological barcodes into dense, multi-channel image tensors, enabling the feature data to be directly processed by subsequent standard convolutional neural network (CNN) layers.
[0084] Set the soft sorting temperature coefficient, denoted as . Traditional simple complex filtering sorting is discrete, and this "hard sorting" prevents gradients from backpropagating, making it impossible for the neural network to learn the weights. and This invention introduces a differentiable sorting approximation mechanism. In step S3, the sorting logic for reconstructing the filtered sequence is implemented by executing a "soft sorting" algorithm; specifically, for any two simplexes in the set of simplex complexes... and The corresponding filtering time value is calculated based on the joint filtering function. and Calculate the absolute value of the difference between the two and divide it by the soft-sorting temperature coefficient. (In this embodiment, the value is 0.1). The sigmoid activation function is then applied to the quotient from the previous step, resulting in a "pairwise order probability value" between 0 and 1. This value characterizes... Ranked Previous probability strength. For the simplex. The summation of this summation with the "pair order relation probability values" of all other simplexes in the set is taken as the result. The "soft rank" in the sequence. This embodiment... The closer the sorting is to 0, the closer it is to a hard sort. The larger the value, the smoother the gradient. A value of 0.1 strikes the optimal balance between effective gradient propagation and ranking accuracy.
[0085] In this embodiment, the differentiable entropy-weighted deep neural network layer performs the following specific soft-rank generation steps: Calculate the initial filter value. Based on the joint filter function, obtain the original filter value vector for each simplex in the current simplex set. To prevent gradient conflicts caused by identical filter values, generate a small random noise vector (its magnitude is set to one ten-thousandth of the original filter value) with the same length as the number of simplexes, and superimpose this noise vector onto the original filter value vector to obtain the perturbed filter value vector. For the i2th and j2nd simplexes in the set, calculate the difference between their perturbed filter values and divide this difference by the soft-ranking temperature coefficient. Perform the Sigmoid activation mapping. Perform the Sigmoid function operation on each value in the pairwise difference matrix to obtain a probability matrix representing "the i2th simplex is ranked after the j2nd simplex". For the i-th simplex, sum all the probability values in the i2th row of the probability matrix; this summation result is the continuous floating-point rank of that simplex in the filtering sequence. This soft rank was subsequently used to determine the birth and death times of persistently homologous barcodes.
[0086] The multi-resolution kernel variance set, denoted as The method for determining this is as follows: Multi-scale mapping is employed when generating the persistently homogeneous landscape feature tensor. Specifically, a set of baseline variance values [0.5, 1.0, 2.0, 4.0] are defined, corresponding to the pixel scale on the feature map. For each feature interval in the persistently homogeneous barcode (defined by birth time b and death time d), its lifetime length L = db is calculated. The baseline variance value that best matches this lifetime length L is selected, or multiple feature maps are generated in parallel using all baseline variance values. A Gaussian function is superimposed on the two-dimensional landscape plane, centered at the midpoint of the feature interval and with a width of the selected variance value. This multi-resolution design ensures that this layer can simultaneously preserve both "short-lived but dense micro-texture features" and "long-lived and significant macroscopic skeletal features."
[0087] To demonstrate the effectiveness of the joint filtering function in step S3, this embodiment constructs verification logic: In the early stages of training, a... When assigned a larger weight, the model behaves similarly to traditional grayscale-based bone erosion detection. As training progresses, if the training set contains a large number of early lesions with indistinct grayscale changes but disordered texture, the backpropagation algorithm will automatically increase the weight. The numerical value. This dynamic balancing mechanism is the fundamental feature that distinguishes this invention from the traditional static TDA method, and it relies on the differentiable channel provided by the aforementioned "soft sorting". Specifically, in order to achieve backpropagation of gradients, the neural network layer performs the following continuous mapping operation when calculating the rank of the simplex:
[0088] Calculate the filtering time value vector of all simplexes in the simplex set based on the joint filtering function. , where M is the total number of simplexes. Calculate the pairwise difference matrix. , of which elements The calculation formula is: The continuous floating-point rank of each simplex is generated by row summation. : ; The physical definition is a vector of filtered time values, the values of which are determined by the joint filtering function. Defined as the Sigmoid activation function, i.e. . Defined as the soft sorting temperature coefficient. In this embodiment, this parameter is preferably set to 0.1. This parameter controls the "hardness" of the sorting approximation: when hour, A permutation matrix that approaches binarization; when As the gradient increases, the gradient distribution becomes smoother. The physical definition is the probability estimate of "the i2th simplex is placed after the j2nd simplex". The physical definition is the soft index position of the i2th simplex in the filtering sequence, which is a real number rather than an integer, and supports adjustments to the weight parameters. Differentiate it.
[0089] It should be noted that the soft sorting logic in this embodiment can also use the "Sinkhorn iteration" or "NeuralSort" operator to replace the above-mentioned Sigmoid pairwise difference accumulation method, as long as it can output a differentiable sorting index.
[0090] By constructing the aforementioned soft sorting operator based on the Sigmoid function, the originally discrete and non-differentiable sorting operation is transformed into a continuous matrix operation. This allows the neural network to automatically adjust the texture entropy weight coefficients through backpropagation based on the final pathological classification loss. and grayscale intensity weighting coefficient This enables end-to-end adaptive learning.
[0091] To ensure stability throughout the entire computation process: the FSEI matrix generated in step S1 and the grayscale data in step S2 are mapped to the closed interval [0,1] using minimum-maximum normalization logic before proceeding to step S3. This ensures... and The weight adjustment is based on numerical changes of the same magnitude, avoiding gradient bias caused by differences in units.
[0092] Further elaboration on the above content: Figure 3This is a schematic diagram illustrating the results of a numerical verification experiment on the separation performance of a "differentiable entropy-weighted topological deep neural network layer" in one embodiment of the present invention. This numerical verification aims to test the theoretical response characteristics of the present invention under preset standardized boundary conditions (Gabor filtering and Vietoris-Rips construction logic under a set of scale parameters). Figure 3 In the graph, the horizontal axis represents the calculated Normalized Frequency Domain Structure Entropy Index (FSEI), characterizing the degree of disorder in texture energy; the vertical axis represents the Weighted Continuous Homogeneous Landscape Norm (W-PHL) calculated based on weighted filtering logic, characterizing the topological saliency of the geometric space. Blue data points represent the calculation results for the baseline group (normal trabecular bone texture), red data points represent the calculation results for the validation group (pathological bone erosion areas), and the curve in the middle represents the convergence decision boundary formed after system optimization.
[0093] The weighted persistent homology landscape norm (W-PHL) determination logic is as follows: a persistent homology barcode is constructed based on the filtering time value output by the joint filtering function; the birth-death interval representing the life cycle of topological features is extracted; the interval is mapped and projected onto a discretized two-dimensional grid using a preset multi-resolution Gaussian kernel function to generate a multi-channel persistent homology landscape feature tensor; and then the global Euclidean norm is performed on this tensor. The -Norm operation is a square root operation performed on the sum of the squares of the response amplitudes of all channels and pixel positions in the tensor space. This quantifies the total energy of the topological features under the dual constraints of frequency domain texture entropy and geometric spatial connectivity as a single scalar index characterizing the significance of lesions.
[0094] like Figure 3 As shown, the numerical calculation results clearly reveal the differences in the distribution of different pathological states in the feature space after introducing the "entropy weight dynamic factor". This mathematically proves that the method of "reconstructing simplex sorting logic using frequency domain texture chaos" proposed in this invention can effectively decouple the feature overlap in traditional methods. Specifically, in the low-entropy region (normal bone structure), the calculated data series exhibits tight clustering characteristics, and the W-PHL value remains low; in contrast, in the bone erosion region, due to the increase in frequency domain entropy caused by trabecular fracture, and the significant topological features caused by geometric holes, the data series drifts significantly to the upper right corner of the graph. This significant quantitative difference (the increase in Euclidean distance) directly confirms that after introducing the "differentiable entropy weight", this scheme can construct a highly robust classification hyperplane.
[0095] Further specifying, in step S4, the multi-scale feature fusion decoding network performs the following processing:
[0096] The process involves obtaining the persistently homogeneous landscape feature tensor and the shallow spatial feature map generated during the depth feature extraction process of the X-ray image to be analyzed; upsampling the persistently homogeneous landscape feature tensor using the nearest neighbor interpolation algorithm; generating a global attention gating signal using a convolution kernel initialized with positive bias; and weighting the shallow spatial feature map using the global attention gating signal.
[0097] The upsampled continuous homogeneous landscape feature tensor is concatenated with the weighted shallow spatial feature map along the channel dimension; the concatenated feature data is upsampled using a transposed convolution operation to restore spatial resolution, and the lesion probability mask data is generated through a pixel-level classification layer.
[0098] Further specifying, before weighting the shallow spatial feature map using the global attention gating signal, the method also performs cross-domain dimensional alignment and semantic reweighting operations; the operations specifically include:
[0099] The spatial resolution of the persistently coherent landscape feature tensor is upsampled to match that of the shallow spatial features using the nearest neighbor interpolation algorithm. Figure 1 To; Constructing a channel compression system The convolutional layer maps the upsampled tensor to a single-channel saliency weight map and performs a sigmoid activation operation on the weight map to generate a normalized gating factor.
[0100] The normalized gating factor is used as the global attention gating signal and is multiplied element-wise with the shallow spatial feature map to generate a pre-weighted spatial feature map; the upsampled continuous homogeneous landscape feature tensor is concatenated with the pre-weighted spatial feature map along the channel dimension.
[0101] Further specifying, in step S5, the degradation robust control logic is configured as follows:
[0102] The statistical dispersion index is specifically the global variance statistic; calculate the global variance statistic of the persistent homogeneous landscape feature tensor;
[0103] When the global variance statistic is lower than the confidence threshold, the partial derivative update operation for the texture entropy weight coefficient and the multi-scale Gabor filter bank parameters is stopped while the texture entropy weight coefficient is automatically set to zero, so that its gradient value remains invalid in backpropagation. This forces the joint filtering function to degenerate into a single-modal processing mode that only depends on the geometric grayscale manifold data. The environmental adaptive degradation status code is generated to indicate that the current recognition result is based only on geometric structural features.
[0104] Further specifying, while automatically resetting the texture entropy weight coefficients to zero, the method also performs gradient propagation blocking and parameter freezing operations; the specific operations include:
[0105] In the backpropagation computation graph, the gradient calculation path for the texture entropy weight coefficient is cut off, and its gradient value is forcibly set to a non-updateable state (None or Zero); the frequency and direction parameters of the multi-scale Gabor filter bank are locked, and it is prohibited from making any fine-tuning updates based on the loss function in the current degradation operation cycle, until the global variance statistics value recovers to above the confidence threshold and maintains a preset stable period.
[0106] The following are specific implementation instructions for the above content:
[0107] In the detailed embodiment of step S4, multi-scale feature fusion and lesion probability mask generation are performed: This step aims to solve the technical problem of "how to accurately map abstract topological features back to physical pixel space". While continuous coherence features are sensitive to structural fractures, they lack spatial location information; while shallow spatial features have clear edges, they contain a large amount of non-pathological background noise (normal bone texture). This embodiment organically fuses the two through a "global attention gating" mechanism.
[0108] Topological-spatial feature tensor set (denoted as) It is a feature container containing two components. Component one is the persistently coherent landscape feature tensor output from step S3. Component 1 represents the global pathological topological semantics; Component 2 is a shallow spatial feature map from step S1 or the front end of the feature extraction network (denoted as ). This characterizes the edges and texture details of the original image. The determination method is as follows: directly read the output of step S3 and the intermediate convolutional feature map generated in step S1 from the memory cache.
[0109] Gated activation threshold (denoted as) ): This refers to the parameters of the non-linear activation function used to map the unbounded convolution output to a probability interval of [0,1]. In this embodiment, it specifically refers to the operational logic of the Sigmoid function.
[0110] Pixel-level classification decision probability (denoted as ): Represents the value of each pixel in the output mask, indicating the confidence level that the point belongs to the "bone erosion lesion". Value range: [0.0, 1.0], where 1.0 represents a confirmed lesion and 0.0 represents the background.
[0111] To achieve "global attention gating" and "channel-dimensional stitching," this embodiment designs a pipelined processing logic of "alignment-gating-weighting-stitching." This mechanism utilizes topological features as a "filter" to first remove noise from spatial features, and then fuses the advantages of both. The specific execution steps are as follows:
[0112] The nearest neighbor interpolation algorithm is used to upsample the low-resolution persistently homogeneous landscape feature tensor, making its spatial resolution completely consistent with the shallow spatial feature map, thus establishing a pixel-level correspondence. A tensor initialized with positive bias is constructed. The convolutional layer compresses and maps the aligned topological tensor into a single-channel saliency weight map, and generates a normalized global attention gating factor via a sigmoid activation function to characterize the gating signal. The global attention gating factor is then multiplied element-wise with the shallow spatial feature map to be analyzed. This step utilizes topological saliency to suppress background noise in the spatial features, generating a pre-weighted spatial feature map. The upsampled, continuously homogeneous landscape feature tensor is physically concatenated with the pre-weighted spatial feature map along the channel dimension to form a hybrid feature tensor, which is then input into the subsequent transposed convolutional network.
[0113] Global attention gating factor (denoted as This parameter is a dimensionless floating-point matrix whose value range is restricted to the interval [0,1], and its spatial size and shallow spatial characteristics are related to the characteristics of the shallow space. Figure 1 The physical representation of the "permission" of the "persistent homogeneous landscape feature tensor" to the corresponding pixel location in the "shallow spatial feature map" is as follows: a high value indicates the presence of significant topological anomalies (including bone erosion holes) in the region, allowing spatial features to pass through; a low value indicates a background region, suppressing noise in the spatial features. This embodiment further elaborates on the action of "generating a normalized gating factor" in step S4.
[0114] The specific input data is an intermediate image of single-channel landscape features after upsampling and channel compression (denoted as...). This data comes from the execution of the "Continuous Coherence Landscape Feature Tensor". The output after convolution. Reading intermediate images of single-channel landscape features. The original value of each pixel is multiplied by 4. The inverse of the value x4 is then multiplied by... Using the natural constant e (approximately 2.71828) as the base, the exponent of the opposite number is calculated to obtain the intermediate variable y4. The intermediate variable y4 is then added to the constant 1 to obtain the denominator z4. The quotient of the constant 1 divided by the denominator z4 is calculated. This quotient is the global attention gating factor corresponding to that pixel position. .
[0115] Step S4 executes a cross-domain dimensional alignment and semantic reweighting mechanism: aiming to solve the mismatch between topological features (low resolution, high semantics) and spatial features (high resolution, low semantics) in terms of physical dimension and information hierarchy.
[0116] The landscape tensor output in step S3 has a low resolution ( ), cannot be directly connected with The spatial feature map corresponds to this. This embodiment performs "nearest neighbor interpolation". Specifically, for each source pixel in the landscape tensor, it is directly copied and filled into a target region block with a size magnification factor of P1. If the magnification factor is 16, then one pixel in the original image becomes one pixel in the new image. The color blocks. Nearest neighbor interpolation preserves the "step" characteristic of the topological signal, avoiding lesion boundary blurring caused by smooth interpolation. The specific logic is as follows:
[0117] The upsampled tensor is convolved using a convolution kernel with C input channels and 1 output channel. This step compresses the multi-scale topological features into a single heatmap. A global attention gating factor is then calculated by applying the "global attention gating factor" logic to the heatmap. Global attention gating factor It is expanded to the same number of channels as the shallow spatial feature map and element-wise multiplication is performed. The output data is the "pre-weighted spatial feature map", in which the texture feature values of non-lesion areas are forcibly decayed to near zero.
[0118] Furthermore, conventional direct stitching makes it difficult for the network to distinguish between "lesion edges and normal trabecular bone edges" at the decoding end. This invention employs an "attention gating" mechanism, utilizing persistently homogeneous landscape feature tensors. High-response regions (corresponding to high-persistence topological holes) generate spatial weight maps, forcibly suppressing shallow spatial feature maps. The background texture response is independent of topological features, thereby achieving "semantic focus".
[0119] Specifically, spatial resolution alignment is performed: shallow spatial feature maps are obtained. height With width Parameters. For persistently homogeneous landscape feature tensors Perform bilinear interpolation or nearest neighbor interpolation to change its spatial size from the original low resolution. Forced stretching to the same level Consistency. This step ensures a one-to-one correspondence between the two in the pixel coordinate system.
[0120] For aligned persistent homogeneous landscape feature tensors This performs depthwise compressed convolution operations. Specifically, the kernel size is [size missing]. The convolution kernel will continuously cohomologically coordinate the landscape feature tensor. Multi-channel data is compressed into a single-channel feature map. A Sigmoid non-linear activation operation is performed on each pixel value of this single-channel feature map. The logic of the operation formula is: take the negative power of the natural constant e, add 1, and take the reciprocal. The feature map generated by this operation is the "global attention gating signal," and the high-value regions on it accurately indicate regions with strong topological saliency.
[0121] The "global attention gating signal" is combined with the original shallow spatial feature map. Perform element-wise multiplication based on a broadcast mechanism. During the operation, low-value pixels in the gating signal will automatically attenuate the shallow spatial feature map. The feature responses at corresponding locations are used to filter out background noise and generate a "pre-weighted spatial feature map". The upsampled, continuously homogeneous landscape feature tensor is then used to... The "pre-weighted spatial feature map" is physically concatenated with the "pre-weighted spatial feature map" in the channel dimension to form a hybrid feature tensor containing "topological semantics + purified spatial details".
[0122] For the aforementioned mixed feature tensor, a transposed convolution operation is performed. Prior to this, to prevent the loss of boundary information during the alignment process, the continuously homogeneous landscape feature tensor is upsampled using a nearest neighbor interpolation algorithm; and when generating the gating signal, a tensor initialized with a positive bias (preferably 1.0) is used. Convolutional kernels perform convolution to prevent gradient vanishing in the early stages of training. A transposed convolution operation with a kernel size of 4, stride of 2, and padding of 1 is performed on the concatenated feature data to ensure that the feature map size is precisely doubled with each upsampling, thus restoring spatial resolution. This is used to gradually restore the spatial resolution of the image until it matches the size of the original input X-ray film. Through a... The convolutional layer compresses the number of channels to 1 and performs sigmoid activation again to generate the final pixel-level classification decision probability. The matrix is the probability mask data of lesions.
[0123] For the detailed implementation of step S5, degradation robust control and environmental adaptive degradation are performed: This implementation actively cuts off unreliable signal sources by monitoring the statistical dispersion index of the continuously coherent landscape feature tensor.
[0124] In this embodiment, the multi-scale Gabor filter bank in step S1 is a differentiable Gabor convolutional layer constructed as the front end of a deep neural network. Specifically, the frequency and direction parameters are defined as learnable tensors of this convolutional layer, and gradient calculation is enabled by default in the deep learning framework, allowing the network to fine-tune the frequency domain response characteristics of the filter through backpropagation during normal training. The control logic in step S5 is to dynamically intervene for this specific attribute.
[0125] Global variance statistic (denoted as ): Represents the statistical dispersion index; characterizing the degree of dispersion of pixel value distribution in the persistent homogeneous landscape feature tensor, used to quantify the "information richness" of topological features. Further details on "calculating the global variance statistic": The input data is the persistent homogeneous landscape feature tensor output from step S3. This tensor contains C channels, each with a size of [size missing]. .
[0126] Will continuously resonate with landscape feature tensors The pixel values of all channels and all rows and columns in the dataset are considered as a one-dimensional data set D, which contains... Each element. For the set The summation operation is performed on all elements, and the result is divided by the total number of elements N3 to obtain the global mean. Iterate through each element in set D and calculate its value relative to the global mean. The difference is calculated, and the square of this difference is performed to obtain the squared deviation term. All squared deviation terms are summed to obtain the total sum of squared deviations. The total sum of squared deviations is divided by the total number of elements N³, and the quotient is the global variance statistic. .
[0127] Confidence threshold (denoted as) ): Represents a preset, dimensionless floating-point baseline value. Specifically, it is read from an external read-only configuration file; in this embodiment, it is set to 0.05. It is a critical scalar used to determine whether the currently input frequency domain entropy information possesses "statistical significance." It defines the energy variance boundary between "effective texture signals" and "pure system background noise." The determination method is as follows: collect a set of no fewer than 100 "invalid X-ray film samples." The samples include: air-exposure images without a subject (characterizing detector dark current noise) and completely occluded images (characterizing environmental background noise). These images do not contain any skeletal structure information. For each sample in this dataset, generate a corresponding "persistently coherent landscape feature tensor" according to the logic of steps S1 to S3. For each generated tensor, calculate its global variance statistic. Summarize these 100 variance values to construct a "noise baseline variance distribution histogram." Calculate the 99th percentile of this distribution. Find a value such that the variance of 99% of the noise samples is below this value. Multiply the 99th percentile value by a safety factor of 1.2, and then solidify the result as a confidence threshold. This information is written to the system configuration file. This method ensures that the threshold has objective statistical significance; if the variance calculated in real time is higher than this threshold, there is more than 99% confidence that the signal does not originate from pure device noise, but contains actual physical texture information.
[0128] Environmental Adaptability Degradation Status Code (denoted as) ): This is a hexadecimal system status identifier used to indicate the current operating mode to downstream decision-making modules. Example values: 0x01 (normal bimodal mode), 0x02 (geometric unimodal degraded mode).
[0129] This method employs a control logic of "hysteresis comparison and gradient blocking". Conventional methods only score the confidence level of the results at the output end, and cannot actively avoid the influence of erroneous features during the calculation process. If the image is completely black, the random high entropy generated by S1 may cause the network to forcibly search for "texture breaks" in the noise, resulting in false positives.
[0130] Step S5, which executes the gradient propagation blocking and parameter freezing mechanism, specifically involves:
[0131] when At this time, it indicates that the input frequency domain entropy data is noise. If the gradient is calculated further, the error signal generated by the noise will incorrectly update the weights through backpropagation. In the automatic differential calculation graph, locate the "texture entropy weight coefficient". The variable node is set to "". The state value of the "gradient calculation permission flag" of this node is forcibly changed from True to False. During the subsequent backpropagation, when the differential engine traverses to this node, it will stop the chain rule differentiation operation, and the gradient values of this node and its upstream (Gabor filter bank parameters) will remain None or Zero.
[0132] This logic also simultaneously locks the frequency and direction parameters of the "multi-scale Gabor filter bank" in step S1. Specifically, it obtains a list of parameter handles for the filter generation function. It iterates through this list, setting the update state of each parameter object to "untrainable." This is to prevent the filter from distorting under extreme conditions to fit noisy data, preserving its optimal texture extraction capability learned under normal conditions, so that it can seamlessly switch back to normal operating mode immediately after image quality restoration.
[0133] Specifically, when the degradation robust control logic detects a low signal-to-noise ratio state (meeting the degradation condition), it performs gradient propagation blocking by modifying the node attributes of the deep neural network's backpropagation computation graph. This step specifically includes the following state transition logic: the processor responds to the global variance statistics. Below the confidence threshold The determination result generates a "noise suppression interrupt signal". In response to the "noise suppression interrupt signal", the processor addresses the stored texture entropy weight coefficient in memory. The parameter object is initialized and its value is forcibly rewritten to a floating-point zero value (0.0). The processor traverses all learnable parameter nodes (including frequency and direction parameters) corresponding to the multi-scale Gabor filter bank in the computation graph, switching their "gradient update permission flag" from "True" to "False". In subsequent backpropagation iterations, when the automatic differentiation engine traverses to the marked parameter node, it skips the partial derivative calculation operation for that node, keeping the corresponding gradient register value zero or null, thereby preventing the error signal from propagating to that branch.
[0134] For noise suppression interrupt signals: the physical definition is a logic level or software interrupt instruction generated internally by the system, used to trigger topology reconstruction or attribute changes in the computation graph.
[0135] Gradient update permission flag: Physically defined as a boolean attribute bit in a deep learning framework used to control whether parameters participate in optimizer updates.
[0136] Backpropagation iteration cycle: Physically defined as one complete time step from calculating the gradient of the loss function to updating the network parameters during the training process of a neural network.
[0137] It should be noted that the embodiments of this application are not limited to directly modifying the flag bits. In a static graph-based computational framework, gradient paths leading to texture branches can also be physically severed at the data flow level using conditional branch operators.
[0138] This embodiment constructs a "learning firewall" for noisy data by dynamically locking the update attributes of specific parameter nodes in the backpropagation computation graph. This prevents invalid gradients from contaminating the filter and the historical learning state of the fusion weights under extreme imaging conditions, ensuring that the model can immediately and effectively identify the data by memorizing the parameters after restoring normal input.
[0139] Furthermore, the execution status monitoring and mode switching logic is as follows: obtain the preset confidence threshold. Perform numerical comparison operations:
[0140] Determine the path (normal mode): If Maintain the current network state and generate status code 0x01.
[0141] Path determination (degradation mode): If This triggers the degradation control sequence.
[0142] When the path is triggered (downgrade mode), perform the following actions:
[0143] Zeroing out weights: Forces the texture entropy weight coefficients in the differentiable entropy weighted topology layer to be zero. The current memory value is overwritten as the floating-point number 0.0. In the automatic differentiation engine of a deep learning framework, this is... The parameter object's attribute is marked as False. This operation physically severs the propagation path of the loss function's derivative with respect to that parameter, preventing noisy data from contaminating the model. A hexadecimal code 0x02 is generated and appended to the metadata of the output lesion probability mask data.
[0144] This section aims to reveal the underlying mathematical mechanisms of "global attention gating" in step S4 and "degeneracy robust control" in step S5, and to clarify how they work together to solve the two core technical problems of "difficulty in spatial localization of abstract topological features" and "noise misleading under extreme conditions".
[0145] Global attention gating factor Range analysis: This parameter is generated by the Sigmoid function, and its theoretical range is (0,1). When approaching a minimum: This indicates that the corresponding pixel location is in a "homologically trivial" state in the topological space, meaning there are no significant, persistently homological features (including bone erosion holes or fractures). Physically, it corresponds to normal dense bone or a background region. A "noise suppression" operation is performed. At this point, the shallow spatial feature map multiplied with it ( Texture details in the image are treated as non-pathological background interference (normal bone trabecular texture), and their response values are forcibly attenuated. This solves the false positive problem caused by direct stitching of traditional U-Net.
[0146] when When approaching a maximum: This indicates that the corresponding pixel location has a high energy response in a persistently homogeneous landscape, i.e., there is a significant long-period "birth-death" feature. Physically, it corresponds to a clear core of bone destruction or erosion lesions. Perform a "feature permeability" operation. Allow shallow spatial features (edges, shapes) at this location to pass through the gate without loss and be strongly fused with topological semantics to achieve accurate delineation of the lesion boundary.
[0147] Global variance statistics ( The theoretical range of ) is .when (lower than) When the input image is completely black (underexposed), completely white (overexposed), or purely noisy, the tensor representing the continuous homogeneous landscape features lacks information and exhibits a uniform distribution or a state of zero. This corresponds to the input image being completely black (underexposed), completely white (overexposed), or purely noisy, making it impossible to extract effective topological structure. This state is characterized by "confidence collapse." At this point, the degradation logic of S5 is triggered, cutting off the entropy weight path to prevent the network from learning features from incorrect noise gradients.
[0148] Frequency domain texture entropy weighting coefficients ( The relationship between the monitoring indicator and the gradient update permission flag is qualitatively characterized as a conditional binary dependency. This dependency only applies if the monitoring indicator... hour, Participate in gradient update; otherwise, The gradients are forced to zero and frozen. In deep learning training, if the input data is pure noise, the network will attempt to fit the noise by adjusting the weights (overfitting). This design directly supports the beneficial effect of "environmental adaptability degradation," ensuring that the model does not suffer catastrophic forgetting or parameter drift under harsh clinical imaging conditions, thus improving the robustness of the system.
[0149] Furthermore, this embodiment is configured in the following digital twin verification scenario: constructing an X-ray sequence stream containing standard lesions, micro-lesions, normal bone, and extreme imaging artifacts. It runs in an inference engine with automatic differentiation capabilities, monitoring changes in each intermediate variable in real time.
[0150] Input image state: quantized into normalized pixel intensity distribution and signal-to-noise ratio (SNR). Topological feature intensity ( ): The maximum activation value of the simulated S3 output tensor. Confidence threshold ( The initial bias is set to 1.0. See Table 1 below for details.
[0151] Table 1: Examples of system response calculations under different environmental disturbances
[0152] Scene Number Input image state description Topological feature strength Global variance statistics Judgment Result Texture entropy weight state Backpropagation gradient response Mean of gating factor Background noise suppression ratio status codes S-01 Standard exposure, significant bone erosion Strong (0.95) 0.18 Pass Normal update Active 0.72 28% 0x01 S-02 Standard exposure, early micro-lesions Medium (0.45) 0.08 Pass Normal update Active 0.35 65% 0x01 S-03 Standard exposure, normal healthy bones Weak(0.02) 0.06 Pass Normal update Active 0.05 95% 0x01 S-04 Severe underexposure (complete black / extremely low SNR) Extremely weak (0.00) 0.004 Fail Forced reset Null (Frozen) N / A 100% (Hard) 0x02 S-05 Sensor artifacts (high-frequency white noise) Random (pseudo-strong) 0.012 Fail Forced reset Null (Frozen) N / A 100% (Hard) 0x02 S-06 Restore normal exposure (late part of the sequence) Strong (0.90) 0.16 Pass Resume updates Active 0.68 32% 0x01
[0153] Background noise suppression ratio, denoted as This parameter represents the proportion of spatial feature energy blocked by the global attention gating mechanism. Its calculation logic is as follows: subtract the mean of the global attention gating factor at the current moment from the constant 1, and convert the difference into a percentage. This parameter quantitatively characterizes the filtering ability of this scheme for non-lesion areas (normal bone texture, background noise). A higher value indicates a lower evaluation of the "topological salience" of the current region by this scheme, thereby forcibly suppressing spatial feature input to that region and preventing false positives.
[0154] The S-03 dataset validated the accuracy of "semantic reweighting" (S4 core effect). Comparing S-01 (significant lesions) and S-03 (normal bone), both had "standard exposure" in terms of input image quality and passed the test. The initial threshold screening (0.18 and 0.06 are both greater than 0.05). S-03 The success rate is 95%, while S-01 only reaches 28%. This significant difference in data demonstrates that even in the presence of geometric structures (normal bone), the differentiable entropy-weighted topological network of this invention can accurately identify the lack of "continuous homology features" (no pathological cavities) and block the spatial features of that region from entering the decoding network through a low gating factor (0.05). This directly solves the technical problem in the prior art where the U-Net network cannot distinguish between "normal bone trabeculae" and "pathological bone erosion," achieving a qualitative leap from "morphological segmentation" to "semantic segmentation."
[0155] The S-04 / S-05 dataset validated the effectiveness of the "environmental adaptability degradation" mechanism (the core effect of S5). Under the extreme conditions of S-04 and S-05, the global variance statistics (0.004, 0.012) fell below the confidence threshold. At this point, the table showed a switch to the Null (Frozen) state, and the status code changed to 0x02. This response mechanism demonstrates that this scheme can proactively cut off the learning path when it detects insufficient input entropy (pure noise). Compared to existing technologies that still force gradient calculation under noisy input, leading to model parameter divergence, this invention constructs a model protection barrier at the physical level through a "gradient freezing" strategy, ensuring the survivability and parameter stability of this scheme in harsh imaging environments.
[0156] The S-06 data set verifies the "elastic recovery capability" of this solution. The data shows that when the input environment switches from S-05 (artifact) back to S-06 (normal), the main control loop of this solution does not need to be restarted. The response value rose back to 0.16, and the backpropagation gradient response immediately returned to Active. This demonstrates that the control logic of this invention possesses millisecond-level dynamic response capability, enabling it to adapt to intermittent equipment interference that may occur in clinical practice.
[0157] Furthermore, the following operating state ranges were determined:
[0158] Interval 1: Geometric single-mode degraded operation interval. Its boundary is determined by the statistical significance limit of information entropy. Within this interval, the global variance of the persistently homogeneous landscape feature tensor is lower than... Mathematically, this indicates that the numerical fluctuations within this tensor are not statistically significant, and are likely thermal noise or background noise rather than a valid topological signal. The corresponding operation is to adjust the texture entropy weighting coefficients in the backpropagation computation graph. And set the Gabor filter parameter properties to False. Force the Gabor filter to... The forward propagation value is set to 0.0. An environment adaptation degradation status code 0x02 is generated and output, indicating that the current segmentation result of the downstream module is based solely on geometric grayscale features.
[0159] Interval 2: Topology-geometric dual-modal fusion operation interval. Its boundaries are determined by the existence verification of valid texture features. Within this interval, if the variance statistic exceeds the threshold, it indicates that the frequency domain texture data contains a discriminative "birth-death" interval pattern, i.e., the existence of physically plausible trabecular microstructures or erosion pore features.
[0160] The corresponding operation is to allow. The filter parameters are then updated normally using the gradient based on the loss function. The entire process of "alignment-gating-weighting-segmentation" in step S4 is executed, utilizing... Perform semantic enhancement on spatial features. Generate and output status code 0x01, indicating that the current result is a high-confidence bimodal fusion result.
[0161] Figure 4 This is a schematic diagram illustrating the results of a numerical verification experiment conducted on the logic of "multi-scale feature fusion decoding (S4)" and "degradation robust control (S5)" in one embodiment of the present invention. The numerical verification aims to test the theoretical response characteristics of the present invention under preset standardized boundary conditions (i.e., test vectors S-01 to S-06 as described in Table 1). Figure 4 In the middle, the horizontal axis represents the standardized test scenario covering standard lesions, healthy bones, and extreme noise, and the upper vertical axis represents the global variance statistics calculated based on the S5 step. The lower vertical axis represents the background noise suppression ratio calculated based on step S4. The red dashed line represents the preset confidence threshold. The gray-filled bars represent the system protection state triggered by "gradient freeze" due to excessively low variance.
[0162] like Figure 4 As shown, the numerical calculation results clearly reveal the system's adaptive control logic for input information entropy. This mathematically proves that the "environmentally adaptive degradation" and "global attention gating" methods proposed in this invention can balance high sensitivity and high robustness. Specifically, in test scenarios S-04 (severe underexposure) and S-05 (artifact interference), the calculated global variance (0.004 and 0.012, respectively) is significantly lower than the confidence threshold (0.05), correctly triggering the degradation state (upper gray column) and forcibly locking the background suppression ratio to 100% (lower full column), thereby blocking the pollution of the network by invalid gradients. In stark contrast, in scenario S-03 (normal healthy skeleton), although the variance passed the threshold detection, the calculated background noise suppression ratio was 95%, which directly confirms that after introducing "global attention gating," this scheme can automatically suppress false positive responses of non-pathological bone texture by utilizing the lack of topological semantics.
[0163] Figure 1 This paper demonstrates the complete technical logic path of the present invention for processing X-ray images of rheumatoid arthritis. The left side shows the skeletal image object to be analyzed, and the right side shows the core processing flow. Figure 1 The process first executes the "frequency domain texture and manifold construction" step, while simultaneously acquiring the frequency domain texture chaos distribution data of the trabecular bone in the local area (corresponding to S1), and constructing a simple complex set based on geometric grayscale (corresponding to S2), thus completing the initialization of multimodal data.
[0164] The data stream then enters the "Entropy Weight Topological Homology Calculation" module (corresponding to S3). Through a differentiable entropy weight topological deep neural network layer, the frequency domain texture chaos degree is used as a dynamic weight to reconstruct the filtering and sorting logic of the simplex, calculate and generate a continuous homology landscape feature tensor. This tensor accurately represents the bone structure missing features under the dual constraints of texture entropy and geometric connectivity.
[0165] The generated feature tensor is then input into the "multi-scale feature fusion decoding" module (corresponding to S4). The decoding network performs cross-domain splicing and upsampling convolution to fuse the abstract topological landscape features with the shallow spatial features of the X-ray film, generating lesion probability mask data for accurately locating the bone erosion area.
[0166] The entire process is controlled by the "Robust Monitoring and Degradation Control" module (corresponding to S5), which calculates the statistical dispersion index of the continuously homogeneous landscape feature tensor in real time. If the index is lower than the preset confidence threshold, the degradation control logic is automatically triggered, generating an environmental adaptive degradation status code, indicating that the main control loop enters a degradation operation mode that relies solely on geometric features, thereby ensuring the robustness of this solution under low signal-to-noise ratio images.
[0167] Example 2:
[0168] A deep learning-based system for identifying bone erosion in X-ray images of rheumatoid arthritis, the system being used to execute the aforementioned method for identifying bone erosion in X-ray images of rheumatoid arthritis, comprising:
[0169] The frequency domain feature extraction module is configured to acquire frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the bone trabecular texture signal in a local region of the X-ray image to be analyzed.
[0170] The topology manifold construction module is configured to acquire the geometric grayscale manifold data of the X-ray image to be analyzed, and construct a corresponding simple complex set based on the geometric grayscale manifold data;
[0171] The core module of entropy weight topology fusion includes a differentiable entropy weight topology deep neural network layer, which is configured to perform weighted filtering homology calculation with learnable parameters; the calculation includes: using the frequency domain texture chaos distribution data as a dynamic weight factor, reconstructing the sorting logic of each simplex in the simplex set entering the filtering sequence, so as to generate a continuous homology landscape feature tensor as a continuous homology landscape indicator signal.
[0172] The multi-scale feature fusion decoding module is configured to perform cross-domain stitching and upsampling convolution of the continuous homology landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area.
[0173] The robust adaptive control module is configured to monitor the statistical dispersion index of the continuously homogeneous landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, the degradation robust control logic is triggered to automatically generate an environmental adaptive degradation status code that characterizes the degraded operation of the system.
[0174] The computational logic involved in this application can be constructed using algorithms such as regression analysis in machine learning, establishing a mathematical model by analyzing the inherent trends and interrelationships of the collected parameters. This process can be implemented using specialized computational tools (such as Python's Scikit-learn library or the R language environment). Throughout all calculations, to eliminate the influence of different physical dimensions and ensure that data is compared and analyzed on the same scale, the input parameters in each formula are dimensionless. The dimensionless techniques used include, but are not limited to, max-min normalization or Z-score standardization.
[0175] The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python) configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.
[0176] It should be emphasized that the foregoing embodiments are merely illustrative of preferred implementations of the present invention and are not intended to limit the scope of protection of the present invention. This application also provides a computer-readable storage medium having computer program instructions stored thereon.
Claims
1. A deep learning-based method for identifying bone erosion in X-ray images of rheumatoid arthritis, characterized in that, Specifically, it includes: S1: Obtain the frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the trabecular texture signal in the local area of the X-ray image to be analyzed. S2: Obtain the geometric grayscale manifold data of the X-ray image to be analyzed, and construct the corresponding simple complex set based on the geometric grayscale manifold data; S3: Input the frequency domain texture chaos distribution data and the simple complex set into a differentiable entropy weighted topological deep neural network layer; The differentiable entropy-weighted topological deep neural network layer is configured to perform weighted filtering homology computation with learnable parameters; the computation includes: using the frequency domain texture chaos distribution data as dynamic weight factors, reconstructing the sorting logic of each simplex entering the filtering sequence in the simplex set, so as to generate a persistent homology landscape feature tensor as a persistent homology landscape indicator signal; the persistent homology landscape indicator signal characterizes the pathological bone structure loss features under the dual constraints of frequency domain energy disorder and geometric spatial connectivity; S4: Input the persistent homogeneous landscape feature tensor into a multi-scale feature fusion decoding network; the decoding network is configured to perform cross-domain stitching and upsampling convolution of the persistent homogeneous landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area; S5: Monitor the statistical dispersion index of the continuously coherent landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, trigger the degradation robust control logic and automatically generate an environmental adaptability degradation status code that characterizes the degraded operation of the system.
2. The method for identifying bone erosion in X-ray films of rheumatoid arthritis based on deep learning according to claim 1, characterized in that: In step S1, obtaining the frequency domain texture chaos distribution data of the X-ray image to be analyzed specifically includes: reading the scale parameter set and orientation parameter set defined in the external configuration file; A multi-scale Gabor filter bank is constructed based on the scale parameter set and the orientation parameter set; the X-ray image to be analyzed is edge-expanded using the reflection filling algorithm, and the multi-scale Gabor filter bank is used to perform a convolution operation on the expanded image to generate a multi-channel frequency domain response map aligned with the original image size; For each pixel position in the multi-channel frequency domain response map, a local sliding window is defined, and a logarithmic probability weighted summation operation is performed based on the energy distribution probability within the local sliding window to generate a frequency domain structure entropy exponent matrix as the frequency domain texture chaos distribution data.
3. The method for identifying bone erosion in X-ray films of rheumatoid arthritis based on deep learning according to claim 2, characterized in that: In step S2, a set of corresponding simple complexes is constructed based on the geometric grayscale manifold data, specifically including: A spatial-grayscale anisotropic scaling factor is introduced; the pixel coordinates of the X-ray image to be analyzed are mapped to planar positions, and the pixel grayscale values are multiplied by the spatial-grayscale anisotropic scaling factor and mapped to elevation, thereby constructing a three-dimensional point cloud space; Define an increasing distance threshold sequence and set a maximum topological dimension constraint parameter; for each threshold step in the distance threshold sequence, construct a simplex from the set of points in the three-dimensional point cloud space where the Euclidean distance between each pair of points is less than the current threshold step; under the constraint of the maximum topological dimension constraint parameter, generate a set of simplex complexes, where any simplex with a dimension higher than the maximum topological dimension constraint parameter is discarded.
4. The method for identifying bone erosion in X-ray films of rheumatoid arthritis based on deep learning according to claim 3, characterized in that: In step S3, the differentiable entropy-weighted topological deep neural network layer is configured to perform the following operations: Perform cross-domain attribute mapping: calculate the geometric center of each simplex in the set of simplex complexes, and project the geometric center onto the frequency domain texture chaos distribution data to index the corresponding frequency domain texture entropy attribute value; Perform weighted filtering calculation: Calculate the filtering time value by using a joint filtering function containing learnable parameters, combining the gray values of each simplex with the frequency domain texture entropy attribute value; Perform differentiable soft sorting: Introduce a soft sorting temperature coefficient, calculate the pairwise difference probability based on the filtering time value and accumulate it to generate a continuous floating-point rank, thereby reconstructing the sorting logic of the simplex entering the filtering sequence; Generating a landscape tensor: Based on the continuous floating-point rank, a persistent homology barcode is calculated, which consists of a series of birth-death intervals; using a preset multi-resolution variance Gaussian kernel function set, the birth-death intervals are mapped to a multi-channel weighted persistent landscape map, thus forming the persistent homology landscape feature tensor.
5. The deep learning-based method for identifying bone erosion in X-ray films of rheumatoid arthritis according to claim 4, characterized in that: Performing cross-domain attribute mapping specifically includes: for each simplex in the set of simplexes, determining its geometric center coordinates in the three-dimensional point cloud space; projecting the geometric center coordinates back to the pixel coordinate system of the X-ray image to be analyzed, and indexing the corresponding position value in the frequency domain texture chaos distribution data; assigning the indexed value to the simplex as its corresponding frequency domain texture entropy attribute value.
6. The deep learning-based method for identifying bone erosion in X-ray films of rheumatoid arthritis according to claim 5, characterized in that: In step S4, the multi-scale feature fusion decoding network performs the following processing: The process involves obtaining the persistently homogeneous landscape feature tensor and the shallow spatial feature map generated during the depth feature extraction process of the X-ray image to be analyzed; upsampling the persistently homogeneous landscape feature tensor using the nearest neighbor interpolation algorithm; generating a global attention gating signal using a convolution kernel initialized with positive bias; and weighting the shallow spatial feature map using the global attention gating signal. The upsampled continuous homogeneous landscape feature tensor is concatenated with the weighted shallow spatial feature map along the channel dimension; the concatenated feature data is upsampled using a transposed convolution operation to restore spatial resolution, and the lesion probability mask data is generated through a pixel-level classification layer.
7. The deep learning-based method for identifying bone erosion in X-ray films of rheumatoid arthritis according to claim 6, characterized in that: Before applying the global attention gating signal to the shallow spatial feature map for weighting, the method also performs cross-domain dimensional alignment and semantic reweighting operations; the operations specifically include: The spatial resolution of the persistently homogeneous landscape feature tensor is upsampled to match that of the shallow spatial feature map using the nearest neighbor interpolation algorithm; a channel compression mechanism is then constructed. The convolutional layer maps the upsampled tensor to a single-channel saliency weight map and performs a sigmoid activation operation on the weight map to generate a normalized gating factor. The normalized gating factor is used as the global attention gating signal and is multiplied element-wise with the shallow spatial feature map to generate a pre-weighted spatial feature map; the upsampled continuous homogeneous landscape feature tensor is concatenated with the pre-weighted spatial feature map along the channel dimension.
8. The method for identifying bone erosion in X-ray films of rheumatoid arthritis based on deep learning according to claim 7, characterized in that: In step S5, the degradation robust control logic is configured as follows: The statistical dispersion index is specifically the global variance statistic; calculate the global variance statistic of the persistent homogeneous landscape feature tensor; When the global variance statistic is lower than the confidence threshold, the partial derivative update operation for the texture entropy weight coefficient and the multi-scale Gabor filter bank parameters is stopped while the texture entropy weight coefficient is automatically set to zero, so that its gradient value remains invalid in backpropagation. This forces the joint filtering function to degenerate into a single-modal processing mode that only depends on the geometric grayscale manifold data. The environmental adaptive degradation status code is generated to indicate that the current recognition result is based only on geometric structural features.
9. The method for identifying bone erosion in X-ray films of rheumatoid arthritis based on deep learning according to claim 8, characterized in that: While automatically resetting the texture entropy weight coefficients to zero, the method also performs gradient propagation blocking and parameter freezing operations; the specific operations include: In the backpropagation computation graph, the gradient calculation path for the texture entropy weight coefficient is cut off, and its gradient value is forcibly set to an unupdateable state; the frequency and direction parameters of the multi-scale Gabor filter bank are locked, and it is prohibited from making any fine-tuning updates based on the loss function in the current degradation operation cycle, until the global variance statistics value recovers to above the confidence threshold and maintains a preset stable period.
10. A system for identifying bone erosion on X-ray films of rheumatoid arthritis, characterized in that: The system is used to perform the method for identifying bone erosion on X-ray films of rheumatoid arthritis as described in any one of claims 1-8, including: The frequency domain feature extraction module is configured to acquire frequency domain texture chaos distribution data of the X-ray image to be analyzed. The frequency domain texture chaos distribution data characterizes the degree of frequency domain energy disorder of the bone trabecular texture signal in a local region of the X-ray image to be analyzed. The topology manifold construction module is configured to acquire the geometric grayscale manifold data of the X-ray image to be analyzed, and construct a corresponding simple complex set based on the geometric grayscale manifold data; The core module of entropy weight topology fusion includes a differentiable entropy weight topology deep neural network layer, which is configured to perform weighted filtering homology calculation with learnable parameters; the calculation includes: using the frequency domain texture chaos distribution data as a dynamic weight factor, reconstructing the sorting logic of each simplex in the simplex set entering the filtering sequence, so as to generate a continuous homology landscape feature tensor as a continuous homology landscape indicator signal. The multi-scale feature fusion decoding module is configured to perform cross-domain stitching and upsampling convolution of the continuous homology landscape feature tensor and the spatial features of the X-ray image to be analyzed, so as to generate lesion probability mask data for locking the bone erosion area. The robust adaptive control module is configured to monitor the statistical dispersion index of the continuously homogeneous landscape feature tensor in real time. When the statistical dispersion index is lower than the preset confidence threshold, the degradation robust control logic is triggered to automatically generate an environmental adaptive degradation status code that characterizes the degraded operation of the system.
Citation Information
Patent Citations
Intervertebral joint osteoarthritis image feature evaluation method based on deep learning
CN121121201A